Prosodic Representation Learning and Contextual Sampling for Neural Text-to-Speech

Sri Karlapati, Ammar Abbas, Zack Hodari, Alexis Moinet, Arnaud Joly, Penny Karanasou, Thomas Drugman

Introduction

Neural text-to-speech (NTTS) techniques have significantly improved the naturalness of speech produced by TTS systems. We refer to NTTS systems as a subset of TTS systems that use neural networks to predict mel-spectrograms from phonemes, followed by the use of neural vocoder to generate audio from mel-spectrograms.

In order to improve the prosodyWe use the subtractive definition of prosody from . of speech obtained from NTTS systems, there has been considerable work in learning prosodic latent representations from ground truth speech. These methods use the target mel-spectrograms as input to an encoder which learns latent prosodic representations. These representations are used by the decoder in addition to the input phonemes, to generate mel-spectrograms. The latent representations obtained by encoding a target mel-spectrogram at the sentence level will have information that is not directly available from the phonemes, and by the subtractive definition of prosody, we may claim that these representations capture prosodic information. Several variational and non-variational methods have been proposed for learning prosodic latent representations. While these methods improve the prosody of synthesised speech, they need an input mel-spectrogram which is not available while running inference on unseen text. This gives rise to the problem of sampling from the learnt prosodic space. Sampling at random from the prior may result in the synthesised speech not having contextually appropriate prosody, as it has no relationship with the text being synthesised. In order to improve the contextual appropriateness of prosody in synthesised speech, there has been work on using textual features like contextual word embeddings and other grammatical information to directly condition NTTS systems. These methods require the NTTS model to learn an implicit correlation between the given textual features and the prosody of the sentence. One work also poses this sampling problem as a selection problem and uses both syntactic distance and BERT embeddings to select a latent prosodic representation from the ones seen at training time.

Bringing both the aforementioned ideas of using ground truth speech to learn prosodic latent representations and using textual information, we build Kathaka, a model trained using a two-stage training process to generate speech with contextually-appropriate prosody. In Stage \Romannum1, we learn the distribution of sentence-level prosodic representations from ground truth speech using a VAE. In Stage \Romannum2, we learn to sample from the learnt distribution using text. In this work, we introduce the BERT+Graph sampler, a novel sampling mechanism which uses both contextual word-piece embeddings from BERT and the syntactic structure of constituency parse trees through graph attention networks. We then compare Kathaka against a strong baseline and show that it obtains a relative improvement of 13.2%13.2\% in naturalness.

NTTS Baseline with Duration Modelling

We use a modified version of DurIAN as our baseline. The baseline learns to predict TT mel-spectrogram frames X=[x0,x1,…,xT−1]X=[\bm{{x}}_{0},\bm{{x}}_{1},\dots,\bm{{x}}_{T-1}], given PP phonemes p=[p0,p1,…,pP−1]\bm{{p}}=[p_{0},p_{1},\dots,p_{P-1}]. The baseline implicitly models the prosody of synthesised speech without any additional context or reference prosody. We have a phoneme encoder, which takes phonemes as input and returns phoneme encodings Y=[y0,y1,…,yP−1]Y=[\bm{{y}}_{0},\bm{{y}}_{1},\dots,\bm{{y}}_{P-1}]. At training, we use forced alignment to extract durations, d=[d0,d1,…,dP−1]\bm{{d}}=[d_{0},d_{1},\dots,d_{P-1}], where ∑di∈ddi=T\sum_{d_{i}\in\bm{{d}}}d_{i}=T. The phoneme encodings are upsampled by replication as per the durations d\bm{{d}} to obtain upsampled phoneme encodings Y↑=[y0↑,y1↑,…,yT−1↑]Y^{\uparrow}=[\bm{{y}}^{\uparrow}_{0},\bm{{y}}^{\uparrow}_{1},\dots,{\bm{{y}}^{\uparrow}_{T-1}}]. The decoder auto-regressively learns to predict the mel-spectrogram XX from Y↑Y^{\uparrow}. Unlike DurIAN, we do not use a post-net as this resulted in instabilities when training with a reduction factor of 11. Our baseline is shown within the bounding box in Fig. 1(a), and the set {θ}\{\theta\} represents its parameters.

We train a duration model d^=fν(p)\bm{{\hat{d}}}=f_{\nu}(\bm{{p}}), to predict the duration d^p\hat{d}_{p} in frames for each phoneme pp, as shown by the blocks within the bounding box in Fig. 1(b). We note that the distribution of phoneme durations contains multiple modes MM. We normalize durations for a group of tokens GmG_{m}, for each mode m∈Mm\in M separately using mean μm\mu_{m} and standard deviation σm\sigma_{m}. We use L2 loss to train the model as shown below, where ⟦⋅⟧\llbracket\cdot\rrbracket represent Iverson brackets.

Kathaka

Here, we introduce the two-stage approach to train Kathaka. In Section 3.1, we describe Stage \Romannum1 where the model learns a distribution over the prosodic latent space. In Section 3.2 we introduce our sampling mechanisms and finally discuss prosody dependent duration modelling in Section 3.3.

We add a variational reference encoder qϕq_{\phi}, with parameters {ϕ}\{\phi\}, to the NTTS architecture {θ}\{\theta\} as shown in Fig. 1(a), to learn prosodic representations from speech. This encoder qϕ(zo∣X)q_{\phi}(\bm{{z}}_{o}\mid X) takes a mel-spectrogram XX as input and predicts the parameters of a Gaussian distribution N(μo,σo2)\mathcal{N}(\bm{{\mu}}_{o},\bm{{\sigma}}^{2}_{o}), from which we sample a prosodic latent representation zo\bm{{z}}_{o}. We assume a prior distribution pθ(zo)=N(0,I)p_{\theta}(\bm{{z}}_{o})=\mathcal{N}(\bm{{0}},I) and train the model to maximize the evidence lower bound (ELBO) defined in Eq. 2, where α\alpha is used as the annealing factor to avoid posterior collapse.

2 Sampling Using Text

During inference, we need to generate the mel-spectrogram XX, and therefore will not have access to it. We note that prosody is driven by the contextual information available in a sentence . Therefore, we propose the usage of text or features derived from it to learn to predict a sentence-level prosodic latent representation zpred\bm{{z}}_{pred}. This is used in place of a mel-spectrogram based sentence-level prosodic latent representation zo\bm{{z}}_{o}. We define samplers sχ(zpred∣W)s_{\chi}(\bm{{z}}_{pred}\mid W), which take text or textual features WW as input, and learn to predict a sentence-level prosodic latent representation zpred\bm{{z}}_{pred}. Since the reference encoder qϕq_{\phi} predicts the parameters of a Gaussian distribution with a diagonal covariance matrix N(μo,σo2)\mathcal{N}(\bm{{\mu}}_{o},\bm{{\sigma}}^{2}_{o}), we treat this as a distribution matching problem, and train samplers to predict parameters of a Gaussian distribution N(μpred,σpred2)\mathcal{N}(\bm{{\mu}}_{pred},\bm{{\sigma}}^{2}_{pred}). We match these two distributions by minimising the KL Divergence:

where DD is the number of dimensions of the latent representation. During inference, we replace zo\bm{{z}}_{o} from the reference encoder qϕq_{\phi}, by zpred\bm{{z}}_{pred} obtained from the sampler sχs_{\chi}. Now we dive into various sampler architectures.

BERT is a masked language model known to capture contextual information in a given sentence . This contextual information is captured in the word-piece embeddings provided by BERT for a given text. Since there has been a lot of work showing correlations between the contextual information captured by BERT and the semantics of the sentence, we use BERT as a semantic sampler. As shown in Fig. 2(a), we use a BERT model pre-trained on long-form text to get contextual word-piece embeddings from the text. These word-piece embeddings are passed through a bidirectional LSTM. We concatenate the first and last hidden states to get a single sentence-level representation which is then projected to obtain μpred\bm{{\mu}}_{pred} and σpred2\bm{{\sigma}}^{2}_{pred}. The sampler’s parameters are denoted as {γ}⊆{χ}\{\gamma\}\subseteq\{\chi\}, and are trained using the loss in Eq. 3, while also fine-tuning BERT at a low learning rate.

2.2 Graph Sampler

Constituency parse trees have long been used to represent the grammatical structure of a piece of text, and their correlation to prosody is well known. As shown in Fig. 1 from , these parse trees TrTr have the words in a sentence as their leaves, and each leaf has only one parent node which represents the part-of-speech of that word. Upon removing the word nodes, the tree Tr′Tr^{{}^{\prime}} with it’s internal nodes can be considered as a representation of the syntax of the sentence. Since all trees Tr′Tr^{{}^{\prime}} can be represented as undirected acyclic graphs GG, we propose to use the tree as input to a Graph Neural Network. Graph Neural Networks have been used to learn representations at the node level, while exploiting the structure of the input . We train a Message-Passing based Graph Attention Network (MPGAT) with one attention head, which generates a node level representation based on the structure of the tree. In one pass of messages between connected nodes, a representation is learnt at every node depending on the connected nodes and itself. We pass NN such messages, where NN is the 75th percentile of the distribution of graph diameters, so that every node also has information pertaining to nodes that are not its immediate neighbours. As shown in Fig. 2(b), we extract GG from the text, and pass it through MPGAT. Upon obtaining a representation at each node, we extract the representations at the leaves in a depth-first order in order to obtain the representations at the parts-of-speech nodes in relation to the structure of the tree. We pass them through a bidirectional LSTM, concatenate the first and last hidden states to get a sentence-level representation, and project it to obtain μpred\bm{{\mu}}_{pred} and σpred2\bm{{\sigma}}^{2}_{pred}. The sampler’s parameters are denoted as {ω}⊆{χ}\{\omega\}\subseteq\{\chi\}, and are trained using the loss in Eq. 3.

2.3 BERT+Graph Sampler

In Fig. 2(c), we combine the BERT and Graph samplers by concatenating the sentence-level representations obtained from each of the methods, and projecting them to obtain μpred\bm{{\mu}}_{pred} and σpred2\bm{{\sigma}}^{2}_{pred}. We denote the set of all parameters in this sampler as {γ}∪{ω}∪{β}⊆{χ}{\{\gamma\}\cup\{\omega\}\cup\{\beta\}}\subseteq\{\chi\}, where {β}\{\beta\} is the set of projection parameters, and we train using Eq. 3.

3 Prosody-dependent duration modelling

We condition our duration model on the phoneme embeddings and the sampled prosodic latent representation, zpred ∼N(μpred,σpred2)\bm{{z}}_{pred}\ \sim\mathcal{N}(\bm{{\mu}}_{pred},\bm{{\sigma}}^{2}_{pred}). Thus, as shown in Fig. 1(b), we are modelling d=fν(p,zpred)\bm{{d}}=f_{\nu}(\bm{{p}},\bm{{z}}_{pred}).

4 Training & Inference

Our model is trained in 3 steps: 1) we train the model with a variational reference encoder and durations obtained from forced alignment as in Section 3.1; 2) we train a sampler on text as in Section 3.2; 3) we train a prosody-dependent duration model as in Section 3.3 using the prosodic latent space N(μpred,σpred2)\mathcal{N}(\bm{{\mu}}_{pred},\bm{{\sigma}}^{2}_{pred}) learnt in Step 2. During inference, we perform 3 steps: a) we use the sampler to predict the latent vector zpred\bm{{z}}_{pred}; b) we concatenate zpred\bm{{z}}_{pred} obtained in Step a, with the phoneme embeddings YY, to predict durations; c) we use upsampled phoneme encodings Y↑Y^{\uparrow}, zpred\bm{{z}}_{pred}, and the predicted durations, to get mel-spectrograms using the model trained in Step 1 of training. Speech is synthesised from mel-spectrograms by using a WaveNet vocoder .

Experiments

We used 38.5 hours of an internal long-form US English dataset recorded by a female narrator. 33 hours of this dataset was used as the training set and the remainder as the test set.

2 Evaluation

We conducted 2 separate MUSHRA tests for evaluating the naturalness of our system and the impact of different sampling techniques. Both MUSHRA tests were taken by 25 native US English listeners. Each test consisted of 50 samples which were 15 seconds in length on average. The listeners rated each system on a scale from 0 to 100 in terms of naturalness. We evaluated statistical significance using pairwise two-sided Wilcoxon signed-rank tests.

We compared each of the samplers mentioned in Section 3.2, with the baseline in a 4-system MUSHRA. Each of the samplers showed a statistically significant (all systems had a p-value<10−3\text{p-value}<10^{-3}, when compared to the baseline) improvement over the baseline, as seen from Fig. 3(a). There is no statistical significance in the differences in the scores between the samplers themselves. This shows that both the BERT and Graph samplers may be equally capable at the task of sampling from the learnt latent prosodic space, therefore, even upon combining them, there is no statistically significant difference in the MUSHRA scores of the samplers. We selected the BERT+Graph sampler as the sampler for the Kathaka model, and used it to measure the reduction in gap as it had the lowest minimum score among the three samplers.

2.2 Gap Reduction Study

In this MUSHRA test, we evaluated 4 systems, namely: 1) the baseline, 2) Kathaka, 3) Kathaka-Oracle, the VAE based NTTS model with oracle prosodic embeddings obtained from recordings, and 4) the original recordings from the narrator. Kathaka obtained a relative improvement over the baseline by a statistically significant 13.2%13.2\% (p-value<10−3\text{p-value}<10^{-3}), as shown in Fig. 3(b). The model using oracle prosodic representations, Kathaka-Oracle, reduces the gap to recordings by 29.19%29.19\% (p-value<10−4\text{p-value}<10^{-4}). We hypothesise that the gap between Kathaka and Kathaka-Oracle is due to the sampler being trained at the sentence level. A sentence may not contain all the context required to determine the appropriate prosody with which a piece of text should be rendered, especially in long-form speech synthesis. While using a sampler trained at the sentence level shows a significant relative improvement, we speculate that exploiting the context available beyond the sentence will help to further reduce this gap and is left as future work.

Conclusion

We presented Kathaka, an NTTS model trained using a novel two-stage training approach for generating speech with contextually appropriate prosody. In the first stage of training, we learnt a distribution of sentence-level prosodic representations. We then introduced a novel sampling mechanism of using trained samplers to sample from the learnt sentence-level prosodic distribution. We introduced two samplers, 1) the BERT sampler which uses contextual word-piece embeddings from BERT and 2) the Graph sampler where we interpret constituency parse trees as graphs and use a Message Passing based Graph Attention Network on them. We then combine both these samplers as the BERT+Graph sampler, which is used in Kathaka. We also modify the baseline duration model to incorporate the latent prosodic information. We conducted an ablation study of the samplers and showed a statistically significant improvement over the baseline in each case. Finally, we compared Kathaka against a baseline, and showed a statistically significant relative improvement of 13.2%13.2\%.

References