Text-Free Prosody-Aware Generative Spoken Language Modeling

Eugene Kharitonov, Ann Lee, Adam Polyak, Yossi Adi, Jade Copet, Kushal Lakhotia, Tu-Anh Nguyen, Morgane Rivière, Abdelrahman Mohamed, Emmanuel Dupoux, Wei-Ning Hsu

Introduction

Natural language processing (NLP) has made tremendous progress recently. One of the most significant findings is that language models (LMs) are natural unsupervised multitask learners Radford et al. (2018, 2019); Brown et al. (2020) — by simply training a big neural network on next word prediction with a large amount of unlabeled text, it learns to comprehend, answer questions, summarize, and even translate Radford et al. (2019). Fine-tuning such pre-trained models further leads to the state-of-the-art performance on numerous benchmark tasks Brown et al. (2020), beating tailor-made models trained from scratch only on labeled data.

Given the impressive performance of pre-trained text language models, it is tempting to approach spoken language processing tasks by first transcribing speech into text with an automatic speech recognition (ASR) system and then utilizing text-based models for comprehension and generation. However, there are a number of caveats for such a framework. First, the majority of the world’s languages are primarily spoken and do not have associated texts in large quantities Lewis et al. (2016). In practice, this limits the reach of NLP techniques to a fraction of the world’s languages that have a large presence on the web and for which there exists a widely available high quality ASR system. Second, despite sharing the same vocabulary and syntactic rules, the spoken form and the written form of the same language still vary significantly in terms of sentence lengths, word distributions, presence of disfluencies and back-channelings, and so on Biber (1991). This makes language models pre-trained on web text not suitable for processing spoken languages. Third, text does not reflect the rich set of features conveyed by oral languages. Speech carries not only phonetic information, but also non-verbal vocalizations (laughter, voice clicks, filler vocalization, etc), rhythm and intonation (prosody), and emotional markers. All of these features could help, not only with generating more expressive speech Ren et al. (2020); Łańcucki (2021), but also with the semantic analysis of the content of the message Cutler et al. (1997); Tran et al. (2017).

To combat these deficiencies, more recently there is increasing interest in exploring speech pre-training using large quantities of unlabeled speech data Chung et al. (2019); Schneider et al. (2019); Kharitonov et al. (2021); Baevski et al. (2020); Hsu et al. (2021c); Liu et al. (2020); Ling and Liu (2020); Tjandra et al. (2020); Hsu et al. (2021b, a). However, most of the studies evaluate their models on discriminative tasks, such as ASR and those in the SUPERB benchmark Yang et al. (2021). To the best of our knowledge, generative spoken language modelling (GSLM) Lakhotia et al. (2021) is the only prior work that evaluates prompted speech completion, a generative tasks that is similar to the text completion task in GPT-2 Radford et al. (2019). To remove the reliance on text, GSLM exploits discovered units from self-supervised models to build a unit language model (uLM) and a unit-to-spectrogram (u2S) model. Speech completion can be achieved by first sampling a unit sequence from the uLM with a unit prompt inferred from a speech prompt, and then synthesizing the sampled sequence into speech with the u2S model. Unfortunately, because those discovered units encode mostly phonetic information Polyak et al. (2021), it suffers from the same prosodic information loss issue as text-based LMs. Therefore, when using that uLM for speech completion, it fails to continue with a coherent tone to the prompt.

In this paper, we introduce a prosody-aware generative spoken language model (pGSLM) that jointly models phonetic content and prosody, in order to leverage prosody for comprehension, and to generate speech coherent with the prompt, which is a precursor for building speech-based dialogue systems. In keeping with our aim of liberating NLP from its over-reliance on text, we follow GSLM and represent the phonetic content with self-supervised units discovered from raw audio. As for prosody, it is represented by the pattern of quantized fundamental frequency (F0) and duration. pGSLM is comprised of two separately trained components: an auto-regressive Multi-Stream Transformer Language Model (MS-TLM) that predicts the next phonetic and prosodic representation given the past ones, and a unit High-Fidelity Generative Adversarial Network (HiFi-GAN) adapted from Polyak et al. (2021) that converts the MS-TLM output into a waveform like a vocoder. To evaluate the proposed model, we adopt metrics from Lakhotia et al. (2021) for content evaluation, and devise a series of metrics for prosody evaluation. Experimental results demonstrate that 1) joint modeling of prosody improves phonetic content modeling, 2) pGSLM can generate speech continuation coherent with the prompt in term of the content and the prosody, and 3) proper choices of model and prosodic representation is crucial to synthesizing natural, coherent, and expressive speech.

Related Work

Our work is related to utilizing prosody for comprehension and predicting prosody for speech synthesis, which we discuss in the following sections.

Prosody, which is often characterized by the rhythm, intonation, and intensity of speech, carries useful information for comprehending speech in addition to the textual content Cutler et al. (1997). Prior studies have shown that including prosody information can improve the performance from text-only models on speech segmentation Shriberg et al. (2000), dialogue act classification Shriberg et al. (1998); Ward and Tsukahara (2000), syntactic parsing Tran et al. (2017), speech–language pathology Cohen et al. (2019), ASR Ostendorf et al. (2003); Shriberg and Stolcke (2004), and language modeling Huang and Renals (2007); Su and Jelinek (2008); Ward et al. (2012). These studies provide strong empirical evidences for the benefit of considering prosody in processing spoken languages, especially in the conversational scenarios.

This work shares the same motivation, but differs from the prior work in two crucial aspects. First, this work utilizes discrete units discovered from a self-supervised model and hence does not require any textual supervision, making it applicable to both written and unwritten languages, while in the prior work prosody information is used alongside text. Second, our model can be regarded as the speech version of GPT, which does not require any task-specific labels and can be pre-trained on large quantities of unlabeled speech data. The ability to leverage more data is shown to be the key to achieve good performance in text pre-training.

2 Prosody Prediction for Speech Synthesis

The proposed pGSLM model can be re-purposed as a text-to-speech (TTS) model when the phonetic content (represented as a unit sequence) is given and the prosody is generated by the MS-TLM model. This is similar to FastSpeech Ren et al. (2020) and FastPitch Łańcucki (2021) TTS models, where prosodic features are predicted from text and speech are generated conditioning on both the text and the predicted prosodic features. As FastSpeech and FastPitch are designed to improve the inference-time efficiency from auto-regressive models like Tacotron Wang et al. (2017), they predict prosodic features and spectrograms without introducing dependency between time steps. In other words, these models assume that the prosody features within an utterance are not correlated across time steps given the text, whereas our proposed MS-TLM does not make such an assumption. We will demonstrate empirically the conditional independence is not a realistic assumption and our model achieves better performance on prosody metrics with auto-regressive modeling.

As for analysis on prosody modeling, we present more extensive metrics by considering both teacher-forcing decoding and sampling, while prior work does not consider the multi-modal nature of prosody and only generate prosody deterministically (Ren et al., 2020). Moreover, we also evaluate prosody in a more disentangled manner by measuring the error of the prosody prediction module alone instead of measuring the error of the prosody extracted from the synthesized waveform: the latter conflates the impact from both the prosody prediction module and the vocoder.

Method

In this section, we first describe the phonetic and prosodic representations used in pGSLM, and then introduce the two components it is comprised of: a multi-stream transformer language model and an adapted unit HiFi-GAN.

We choose units with a vocabulary size of 100 derived from HuBERT Hsu et al. (2021a), a self-supervised speech model, as the phonetic representation. Specifically, these units are obtained through clustering the 6th transformer layer output of the base HuBERT model provided in Hsu et al. (2021a) using a k-means algorithm, following the recipe of HuBERT closely. A speech waveform can therefore be encoded into a sequence of discrete units at a frame rate of 50 units per second, or alternatively, into a sequence of (unit, duration) tuples using run-length encoding. HuBERT units were found to perform favorably compared to other self-supervised units such as wav2vec 2.0 Baevski et al. (2020) and VQ-VAE van den Oord et al. (2017) in terms of lexical content modeling Lakhotia et al. (2021) and disentangling prosodic information Polyak et al. (2021).

One may ask why F0 is only normalized by the speaker mean but not the variance. We argue that the variance encodes the “level of expressiveness” and it is desired to preserve it. This is demonstrated empirically in Figure A.2 in the appendix, where speakers from expressive datasets, EmoV and Blizzard 2013 SynSIG , exhibits larger speaker log F0 standard deviation than those in less expressive datasets, LJSpeech Ito and Johnson (2017) and VCTK Veaux et al. (2016). On the other hand, we also found that variance is more correlated mean in the linear space than in the log space, as shown in Figure A.3. Therefore, we argue that mean-normalized log F0 is a more suitable representation for prosody as it encodes less speaker information while preserving the level of expressiveness.

2 Multi-Stream Transformer LM

We adapt the Transformer LM from Lakhotia et al. (2021) to take multiple streams of input and predict multiple streams of output, and refer to it as the Multi-Stream Transformer Language Model (MS-TLM). An MS-TLM predicts a sequence of segment representations, which reduces the sequence length significantly and is found beneficial compared to predicting frame sequences Lakhotia et al. (2021). Each segment is represented with the unit uu, duration (in frames) dd, and normalized pitch lflf. The first two are obtained by run-length encoding the fixed frame rate unit sequence, while a segment lflf is computed by averaging those from voiced frames within a segment or set to 0 if the entire segment is unvoiced. An example is provide in Appendix C.

We see that the factorial assumption here may be too strong, because the duration and the pitch of a segment are highly correlated with the phonetic content of the same segment. To alleviate that without introducing intra-step dependency or interleaving streams (which increases the sequence length and requires determining an order for the three streams a priori), we introduce a delay factor Δ\Delta (Δ≥0\Delta\geq 0) for prosodic streams, which shift prosodic input and output streams backward by Δ\Delta steps, taking (ut−1,dt−Δ−1,lft−Δ−1)(u_{t-1},d_{t-\Delta-1},lf_{t-\Delta-1}) as input and outputting (ut,dt−Δ,lft−Δ)(u_{t},d_{t-\Delta},lf_{t-\Delta}). When Δ=1\Delta=1, each step of the LM predicts the unit of the current segment and the prosodic representations of the previous segment, of which the lexical unit has been observed already, as shown in Figure 1.

2.2 Quantizing prosodic representations

A straightforward solution to encode prosody streams dd and lflf is to represent them as continuous values and minimize an L1 or L2 loss for training, similar to FastSpeech2 Ren et al. (2020) and FastPitch Łańcucki (2021). Doing so assumes that the duration and the pitch of a segment follow a unimodal distribution (Laplace for L1 and Gaussian for L2) given the context. If the underlying distribution is multimodal with wide spread, the learned distribution would be significantly underfitting with a mean far from the modes. Empirically, we found that such modeling indeed leads to predicting lflf values very close to 0 for all segments, and the generated prosody sounds dull and boring.

2.3 Training objective

The training loss is a weighted sum of three per-stream losses. Omitting dependency on the context for brevity, MS-TLM defines a distribution p(ut,dt,lft)p(u_{t},d_{t},lf_{t}) of the potential values for a timestep tt. Then, denoting ground-truth per-channel values as ut∗,dt∗,lft∗u_{t}^{*},d_{t}^{*},lf_{t}^{*}, we get:

In all experiments, we use cross-entropy as the loss on the predictions of the unit channel (LuL_{u}). Whenever we operate on quantized prosody values (both duration and F0), we also use cross-entropy as losses LdL_{d} and LlfL_{lf}. In the case of continuous-valued prosody streams, we treat predicted values p(dt)p(d_{t}) and p(lft)p(lf_{t}) as the mode of Laplacian distributions and maximize the log likelihood of the model, which is equivalent to minimizing an L1 loss. In preliminary experiments, we found that the results are relatively robust to variations of the relative weights α\alpha and β\beta, hence we fix them α=β=0.5\alpha=\beta=0.5 in all our experiments.

2.4 Sampling from a model

To generate new utterances, potentially conditioned on a prompt, we run autoregressive generation where at each step we sample units, duration, and normalized log F0 values, append them to the context and feed them back. In the case of discrete channels (units, also duration/pitch in the case of discrete-valued models), we sample from the corresponding multinomial distribution. As commonly done in language modelling Lakhotia et al. (2021), we perform sampling with temperature by scaling the logits by the temperature parameter. We fine-tune the temperature on the validation data.

For MS-TLM that models normalized log F0 as continuous variables, we draw samples from a Laplacian distribution with its location parameter set to the predicted value, because the model assumes the output distribution is Laplacian (see §-3.2.3). For duration, to avoid sampling invalid values, we sample from a Laplacian distribution truncated at zero and round it to the nearest positive integer.

3 Waveform Generation with Unit Hifi-GAN

Given (u1:T,d1:T,lf1:T)(u_{1:T},d_{1:T},lf_{1:T}) generated from the MS-TLM, we adapt the discrete unit-based HiFi-GAN vocoder from (Polyak et al., 2021) to generate waveform. The original vocoder proposed in (Polyak et al., 2021) takes in frame-level discrete unit, pitch and speaker embedding as input and applies VQ-VAE quantization on the pitch. As MS-TLM predicts quantized speaker-mean normalized log F0 on the segment level, we modify the training of the vocoder so that it takes frame-level segment-average pitch as input, where the pitch values for frames within a segment are set to the same value. We apply the same quantization described in § 3.2.2 instead of VQ-VAE on the pitch. The unit Hifi-GAN and the MS-TLM are trained separately.

Experimental Setup

In our experiments, we train MS-TLM models on two English datasets: LibriSpeech Panayotov et al. (2015) and a 6K-hour subset Rivière and Dupoux (2020) of Libri-Light Kahn et al. (2020) which we refer to as LL-6K. Both datasets represent audio books and we use LibriSpeech dev-clean and test-clean as validation and test sets. As described in Section 3.1, we use HuBERT-based unit representations. However, to investigate whether our proposed models can work with other types of units, we also experiment with CPC Rivière and Dupoux (2020); Oord et al. (2018) and ground-truth phone representations. We experiment with a vocabulary of 100 units when working with Hubert and CPC, following the same protocol and using the same pre-trained models as Lakhotia et al. (2021). On the other hand, frame-level phone transcripts are obtained through forced-alignment using the tri6b model from Kaldi’s LibriSpeech recipe Povey et al. (2011). The position- and context-independent phones without lexical stress markers are used, which include 41 units (39 phones, one silence SIL, and one spoken noise SPN). The frame rate of CPC and phone units is 100Hz, and is 50Hz for HuBERT units.

We experiment with MS-TLM of two sizes: base and large. The base one has 6 layers, 8 attention heads per layer, embedding size of 512. Its FFN layer has 2048 units. The large variant has 12 layers, each with 16 heads, embedding size of 1024 and the FFN layer is of dimensionality 4096. We set attention dropout and dropout probabilities to 0.1 for both alternatives. On top of that, we apply sequence-level and span-level Baevski et al. (2020) input dropout to the two prosody streams. Specifically, each stream is zero-ed out with a probability of 0.2, and 2% of the steps are selected as starts, from which 5 steps of that stream is zero-ed out. Optimization is done using Adam Kingma and Ba (2014) with a peak learning rate of 5e-4. Learning rate ramps up linearly for the first 4K updates, and then decays to 0 with an inverse square-root schedule. We train the base model for 70 epochs, and large model for 100 epochs. Each GPU’s batch contains up to 3072 (u,d,lf)(u,d,lf) segments and we used 8 (16) GPUs to train base (large) MS-TLM. For each update, we aggregated gradients from 8 batches.

2 Prosody and Content Evaluation

Our overall goal is to find models that can freely generate meaningful content and consistent as well as diverse prosody. In this Section, we define a set of metrics that measure models’ performance over each stream individually and combined, in both the teacher-forcing mode and the inference mode.

A simple way to evaluate models is to measure its loss on hold-out data in a setup where for each step the full ground truth context is provided. For the unit stream, we measures Negative Log-Likelihood (NLL), equivalent to cross-entropy. For the duration and pitch streams we use Mean Absolute Error (MAE), equivalent to L1 loss. When the pitch values are quantized, we de-quantize predictions to the means of the respective buckets.

2.2 Per-stream prosody continuation

We next evaluate the model’s ability to complete a stream in isolation. Specifically, we provide a 3s prompt for all streams, and then sample auto-regressively the target stream while feeding the ground truth value for the other streams, as depicted in Figure 3. The prompts are inferred from the utterances in the validation set. When prosodic features are quantized, we sample with a temperature τ∈{0.0,0.25,0.5,0.7,1.0,1.3}\tau\in\{0.0,0.25,0.5,0.7,1.0,1.3\}, and when they are continuous, we sample with a scale b∈{0.0,0.05,0.125,0.25,0.5,0.7,1.0,1.3}b\in\{0.0,0.05,0.125,0.25,0.5,0.7,1.0,1.3\} for duration and b∈0.01×{2−6,2−5,⋯ ,20}b\in 0.01\times\{2^{-6},2^{-5},\cdots,2^{0}\} for pitch. The temperature/scale is chosen to minimize the Min-MAE for the corresponding stream, which we describe next. We chose different sweeping ranges for continuous pitch and duration because they have different inherent standard deviations.

A prompt might have multiple meaningful continuations in the content space Lakhotia et al. (2021). Similarly, a single sentence can have multiple correct prosodic profiles. To account for that, for each prompt we generate n=20n=20 samples so that a model has a chance to cover most modes of the underlying distribution, and report the minimal MAE (min-MAE) against the reference among the nn samples.

To quantify the models’ capability to generate consistent prosody, we measure Pearson correlation between the mean values of a stream in the prompt and in the generated continuation. Clearly, if the prompt has a distinct tempo or a pitch, a good continuation should reflect this. The same setup as the min-MAE metric is used (n=20n=20) with one exception: we only consider sequences that are at least 6s long.

To measure how expressive the generated prosody is, we calculate the standard deviation of the generated values and expect a good model to exhibit a similar level of that as the ground truth. The same setup as in “Min-MAE” is used.

2.3 Speech continuation

Lastly, we evaluate the model’s ability to carry out prompted speech completion, where all three streams are sampled given a 3s prompt using the temperature/scale parameter determined from per-stream continuation (§ 4.2.2) as illustrated in Figure 3. We sample the MS-TLM auto-regressively until it emits the EOS unit or reaches the length of the reference. The MS-TLM output is synthesized into a waveform using the adapted HiFi-GAN.

We re-use the maximum word-level continuation BLEU2 proposed by Lakhotia et al. (2021) to quantify how well a model can complete a prompt in terms of the textual content. We transcribe the waveform with an off-the-shelf wav2vec 2.0-based ASR Baevski et al. (2020) (same as Lakhotia et al. (2021)) and compute the BLEU2 score for each of the n=20n=20 continuations against the reference completion. The highest one is used as score for a prompt.

We ask humans to evaluate three aspects of speech continuation: sound quality, meaningfulness (how natural the text content is considering both grammar and meaning), and prosody (how consistent and natural the intonation and the rhythm is). We follow the human evaluation protocol used by Lakhotia et al. (2021) closely, where raters evaluate subjective quality of the recordings using headphones on a scale between 1 to 5 with an increment of 1, the higher the better. Only Native English speakers were recruited as raters for all three studies. The same 100 prompts as Lakhotia et al. (2021) from LibriSpeech test-other are used, and each system generates one continuation per prompt. Each continuation is evaluated by at least 5 raters for each aspect. The CrowdMOS package Ribeiro et al. (2011) was used for all experiments using the recommended recipes for outlier removal. All participants were recruited using the Amazon Mechanical Turk platform. The metrics on the three aspects are denoted as MOS, M-MOS, and P-MOS.

Results

In Table 1 we report teacher-forcing metric calculated on LibriSpeech dev-clean dataset for a diverse set of models. In rows 1-8, we report metric values for base MS-TLM models that are trained on LibriSpeech 960h transcribed into HuBERT-100 units. In rows 9-12 we consider large MS-TLM models trained on HuBERT transcripts of LL6k. Rows 13 & 14 and 15 & 16 contain metric values for models that are trained on LibriSpeech 960h transcribed using CPC and ground-truth phonetic units.Note: the metric values in this section are only comparable within the same unit type. To compare across unit types, one can synthesize the MS-TLM output into waveform and transcribe the speech with an ASR systems to compute metrics in the word or character space. The row 1 corresponds to the prosody-ignorant baseline model of Lakhotia et al. (2021).

On comparing two models that only predict units (rows 1 and 5) we see that by simply adding prosodic channels to the input of the model, we obtain considerably lower level of negative log-likelihood of the units (uu NLL: 1.522 vs. 1.336). The same trend persist for the models that predict prosodic channels, too. For instance, this holds in the case of the continuous-F0 models (rows 9 & 11: 1.513 vs. 1.421) and, equally for the quantized F0 HuBERT-based models (rows 10 and 12: 1.522 vs. 1.406). Moreover, this holds for the CPC-based models (row 13 & row 14) and even for the models trained on phone transcripts (rows 15 & 16). Hence we conclude that prosodic input universally improves speech “content” modelling.

Our results in Table 1 also allow us to investigate whether shifting prosody streams w.r.t. the unit stream (Δ>0\Delta>0) is useful. On comparing rows 6 & 7 we see that this is indeed the case: at an expense of some increase in uu NLL (e.g., 1.3371.337 vs. 1.4411.441) we obtain considerable relative improvement in dd MAE (0.722→0.5510.722\rightarrow 0.551). The trend follows when further increasing Δ\Delta. We also observe that having prosody in the context is beneficial when modeling prosody itself. Indeed, this is the case across all pairs of models (rows 9 & 11, 10 & 12) according to dd MAE and lflf MAE metrics. Moreover, this holds for the types of units that differ from HuBERT (CPC: rows 13 & 14, phonetic units: rows 15 & 16).

2 Prosodic Inputs Are Useful for Speech Generation

In our next experiment we study how the number of sampled prompt continuation affects prosody accuracy metrics (MAE). We report results for the four large models (rows 9-12) in Figure 4. From these results we observe that models that operate on quantized prosodic streams greatly benefit from sampling multiple candidates. In contrast, the two continuous-valued models seem to benefit little if at all (in the case of the F0 stream). We hypothesise that this striking difference is due to the ability of the multinomial-sampled trajectories to cover multiple mode of the underlying distribution, while the continuous-valued models produce samples that are “averaged” to the median of the underlying distribution due to the L1 loss.

In Table 2 we report the continuation metrics for four large MS-TLM models, trained on HuBERT transcripts of LL-6k (they correspond to rows 9-12 in Table 1).Audios samples of speech continuation are included in the supplementary material. These models differ in whether they have prosodic input or not (rows 11 & 12 vs. 9 & 10) and if the prosodic channels are discretized or not (10 & 12 vs. 9 & 11).

Firstly, on comparing models with and without prosodic input, we observe that having prosody in input improves the accuracy of the prosody continuation (in terms of MAE). This holds for predicting duration (e.g., 0.542 and 0.536 for rows 10 and 12). We see a higher relative difference for lflf (e.g., 0.096 vs. 0.077, same models). Our proposed models are also able to leverage provided prosody input to maintain high consistency of the prosody continuation, as measured by the correlation metrics. For example, for the continuous-prosody models the correlation values grows from 0.176 to 0.344 for the duration prediction and from 0.093 to 0.494 for the F0 channel. Having prosody input also turns out to be important for the word-level BLEU metric: models 11 and 12 outperform their counterparts without prosody inputs, 9 and 10.

Next, when contrasting discrete- and continuous-prosody models the following picture emerges. For both duration and F0 channels, discrete models achieve lower min-MAE errors. Further, both discrete models generate considerably more diverse F0 values than either of the continuous models (up to 2x higher std). Among the models with prosody inputs, the one with discrete prosody get higher variability in the dd channel. In contrast, the correlation metrics favor the prosody-aware continuous model. From the point of view of the word-level BLEU scores, both models are very close with the quantized model (row 12) being slightly ahead. We attribute this difference between the models to the ability of discrete-valued MS-TLM to better describe multi-modal distributions, as we saw above in the experiment reported in Figure 4.

Table 3 presents the human evaluation results. The model with prosody input and quantized prosody performs significantly better than the rest on MOS and M-MOS, and is on par with the variant with prosody input and continuous prosody on P-MOS. Note that when not having the prosody input, the model with quantized prosody performs significantly worse on all metrics, demonstrating the importance of auto-regressive generation for discrete representation.

To summarize, we conclude that (i) including prosody input allows better modelling of speech, and (ii) architectures that operate with quantized prosody values, generally, perform better on our introduced metrics.

Conclusion and Future Work

In this work, we propose a text-free prosody-aware generative spoken language model, pGSLM, which models textual content and prosodic information explicitly and does not use any text supervision by leveraging self-supervised units. Through extensive evaluation on a diverse set of metrics, we demonstrated that prosody not only improves content modeling, but also enables better prompted speech generation that is aware of both the content and the prosody from the prompt for the first time in the literature. We conducted a number of ablation studies to validate the effectiveness of model design choices.

As for broader impacts, this work serves as the foundation for building better conditional speech generation applications where prosody is essential, such as in the conversational scenarios. In addition, the proposed model could also serve as a pre-trained model for other classification tasks, such as emotion recognition or syntactic parsing from speech, or as a pre-trained model for generative tasks such as text-to-speech synthesis with more expressive and coherent prosody. Finally, the proposed prosody metrics (teacher-forcing duration and pitch MAE, continuation correctness/consistency/expressiveness) may also be used for evaluation of text-to-speech synthesis systems that can produce diverse prosody for a given text input.

References

Appendix A Analysis of Log F0 Distribution

Appendix B HiFi-GAN Adaptation Analysis

Table 4 presents an analysis of HiFi-GAN performance when using different quantized pitch representations. Similarly to Polyak et al. (2020) we report voice decision error (VDE) Nakatani et al. (2008), which measures the portion of frames with voicing decision error and F0 Frame Error (FFE) Chu and Alwan (2009), which measures the portion of frames that contain a deviation of more than 20% in pitch value or have a voicing decision error. Results show that the chosen quantizer achieve favorable performance in terms of VDE and comparable results in terms of FFE without having to pre-train a F0 VQ-VAE quantizer.

Appendix C Example of Converting Frame-Level to Segment-Level Representations

Assume we have an utterance of six frames: [(13, 1.5), (13, 2.5), (13, 0.0), (21, 0.0), (27, 1.3), (27, 3.5)] where the first number in each tuple denotes the unit of the frame and the second number denotes the speaker normalized log F0 of the frame. In particular, the third and the fourth frame are unvoiced and their lflf values are set to 0.0.

The segment level representation of the utterance is [(13, 3, 2.0), (21, 1, 0.0), (27, 2, 2.4)]. The first segment (13, 3, 2.0) is labeled with unit u=13u=13, duration d=3d=3 frames, and an average normalized log F0 lf=(1.5+2.5)/2=2.0lf=(1.5+2.5)/2=2.0 for the two voiced frames. The second segment contains only one unvoiced frame, and hence lflf is set to 0. Finally, the last segment contains two voiced frames, and therefore d=2d=2 and lf=(1.3+3.5)/2=2.4lf=(1.3+3.5)/2=2.4.

Appendix D Effects of F0 Representation on MS-TLM

Table 5 compares content modeling performance when using different pitch representations. Results show that using mean normalized pitch information is better than using raw pitch, and using log pitch is better than using linear pitch.

Appendix E More Details of Human Evaluation

The instruction page displayed to the raters are shown in Figure E.1. We modify the Introduction, Task Instruction, Example in the instruction page for MOS, MMOS, and PMOS correspondingly. The text used for each metric are detailed in Table 6