Prosodic Representation Learning and Contextual Sampling for Neural Text-to-Speech
Sri Karlapati, Ammar Abbas, Zack Hodari, Alexis Moinet, Arnaud Joly, Penny Karanasou, Thomas Drugman
Introduction
Neural text-to-speech (NTTS) techniques have significantly improved the naturalness of speech produced by TTS systems. We refer to NTTS systems as a subset of TTS systems that use neural networks to predict mel-spectrograms from phonemes, followed by the use of neural vocoder to generate audio from mel-spectrograms.
In order to improve the prosodyWe use the subtractive definition of prosody from . of speech obtained from NTTS systems, there has been considerable work in learning prosodic latent representations from ground truth speech. These methods use the target mel-spectrograms as input to an encoder which learns latent prosodic representations. These representations are used by the decoder in addition to the input phonemes, to generate mel-spectrograms. The latent representations obtained by encoding a target mel-spectrogram at the sentence level will have information that is not directly available from the phonemes, and by the subtractive definition of prosody, we may claim that these representations capture prosodic information. Several variational and non-variational methods have been proposed for learning prosodic latent representations. While these methods improve the prosody of synthesised speech, they need an input mel-spectrogram which is not available while running inference on unseen text. This gives rise to the problem of sampling from the learnt prosodic space. Sampling at random from the prior may result in the synthesised speech not having contextually appropriate prosody, as it has no relationship with the text being synthesised. In order to improve the contextual appropriateness of prosody in synthesised speech, there has been work on using textual features like contextual word embeddings and other grammatical information to directly condition NTTS systems. These methods require the NTTS model to learn an implicit correlation between the given textual features and the prosody of the sentence. One work also poses this sampling problem as a selection problem and uses both syntactic distance and BERT embeddings to select a latent prosodic representation from the ones seen at training time.
Bringing both the aforementioned ideas of using ground truth speech to learn prosodic latent representations and using textual information, we build Kathaka, a model trained using a two-stage training process to generate speech with contextually-appropriate prosody. In Stage \Romannum1, we learn the distribution of sentence-level prosodic representations from ground truth speech using a VAE. In Stage \Romannum2, we learn to sample from the learnt distribution using text. In this work, we introduce the BERT+Graph sampler, a novel sampling mechanism which uses both contextual word-piece embeddings from BERT and the syntactic structure of constituency parse trees through graph attention networks. We then compare Kathaka against a strong baseline and show that it obtains a relative improvement of in naturalness.
NTTS Baseline with Duration Modelling
We use a modified version of DurIAN as our baseline. The baseline learns to predict mel-spectrogram frames , given phonemes . The baseline implicitly models the prosody of synthesised speech without any additional context or reference prosody. We have a phoneme encoder, which takes phonemes as input and returns phoneme encodings . At training, we use forced alignment to extract durations, , where . The phoneme encodings are upsampled by replication as per the durations to obtain upsampled phoneme encodings . The decoder auto-regressively learns to predict the mel-spectrogram from . Unlike DurIAN, we do not use a post-net as this resulted in instabilities when training with a reduction factor of . Our baseline is shown within the bounding box in Fig. 1(a), and the set represents its parameters.
We train a duration model , to predict the duration in frames for each phoneme , as shown by the blocks within the bounding box in Fig. 1(b). We note that the distribution of phoneme durations contains multiple modes . We normalize durations for a group of tokens , for each mode separately using mean and standard deviation . We use L2 loss to train the model as shown below, where represent Iverson brackets.
Kathaka
Here, we introduce the two-stage approach to train Kathaka. In Section 3.1, we describe Stage \Romannum1 where the model learns a distribution over the prosodic latent space. In Section 3.2 we introduce our sampling mechanisms and finally discuss prosody dependent duration modelling in Section 3.3.
We add a variational reference encoder , with parameters , to the NTTS architecture as shown in Fig. 1(a), to learn prosodic representations from speech. This encoder takes a mel-spectrogram as input and predicts the parameters of a Gaussian distribution , from which we sample a prosodic latent representation . We assume a prior distribution and train the model to maximize the evidence lower bound (ELBO) defined in Eq. 2, where is used as the annealing factor to avoid posterior collapse.
2 Sampling Using Text
During inference, we need to generate the mel-spectrogram , and therefore will not have access to it. We note that prosody is driven by the contextual information available in a sentence . Therefore, we propose the usage of text or features derived from it to learn to predict a sentence-level prosodic latent representation . This is used in place of a mel-spectrogram based sentence-level prosodic latent representation . We define samplers , which take text or textual features as input, and learn to predict a sentence-level prosodic latent representation . Since the reference encoder predicts the parameters of a Gaussian distribution with a diagonal covariance matrix , we treat this as a distribution matching problem, and train samplers to predict parameters of a Gaussian distribution . We match these two distributions by minimising the KL Divergence:
where is the number of dimensions of the latent representation. During inference, we replace from the reference encoder , by obtained from the sampler . Now we dive into various sampler architectures.
BERT is a masked language model known to capture contextual information in a given sentence . This contextual information is captured in the word-piece embeddings provided by BERT for a given text. Since there has been a lot of work showing correlations between the contextual information captured by BERT and the semantics of the sentence, we use BERT as a semantic sampler. As shown in Fig. 2(a), we use a BERT model pre-trained on long-form text to get contextual word-piece embeddings from the text. These word-piece embeddings are passed through a bidirectional LSTM. We concatenate the first and last hidden states to get a single sentence-level representation which is then projected to obtain and . The sampler’s parameters are denoted as , and are trained using the loss in Eq. 3, while also fine-tuning BERT at a low learning rate.
2.2 Graph Sampler
Constituency parse trees have long been used to represent the grammatical structure of a piece of text, and their correlation to prosody is well known. As shown in Fig. 1 from , these parse trees have the words in a sentence as their leaves, and each leaf has only one parent node which represents the part-of-speech of that word. Upon removing the word nodes, the tree with it’s internal nodes can be considered as a representation of the syntax of the sentence. Since all trees can be represented as undirected acyclic graphs , we propose to use the tree as input to a Graph Neural Network. Graph Neural Networks have been used to learn representations at the node level, while exploiting the structure of the input . We train a Message-Passing based Graph Attention Network (MPGAT) with one attention head, which generates a node level representation based on the structure of the tree. In one pass of messages between connected nodes, a representation is learnt at every node depending on the connected nodes and itself. We pass such messages, where is the 75th percentile of the distribution of graph diameters, so that every node also has information pertaining to nodes that are not its immediate neighbours. As shown in Fig. 2(b), we extract from the text, and pass it through MPGAT. Upon obtaining a representation at each node, we extract the representations at the leaves in a depth-first order in order to obtain the representations at the parts-of-speech nodes in relation to the structure of the tree. We pass them through a bidirectional LSTM, concatenate the first and last hidden states to get a sentence-level representation, and project it to obtain and . The sampler’s parameters are denoted as , and are trained using the loss in Eq. 3.
2.3 BERT+Graph Sampler
In Fig. 2(c), we combine the BERT and Graph samplers by concatenating the sentence-level representations obtained from each of the methods, and projecting them to obtain and . We denote the set of all parameters in this sampler as , where is the set of projection parameters, and we train using Eq. 3.
3 Prosody-dependent duration modelling
We condition our duration model on the phoneme embeddings and the sampled prosodic latent representation, . Thus, as shown in Fig. 1(b), we are modelling .
4 Training & Inference
Our model is trained in 3 steps: 1) we train the model with a variational reference encoder and durations obtained from forced alignment as in Section 3.1; 2) we train a sampler on text as in Section 3.2; 3) we train a prosody-dependent duration model as in Section 3.3 using the prosodic latent space learnt in Step 2. During inference, we perform 3 steps: a) we use the sampler to predict the latent vector ; b) we concatenate obtained in Step a, with the phoneme embeddings , to predict durations; c) we use upsampled phoneme encodings , , and the predicted durations, to get mel-spectrograms using the model trained in Step 1 of training. Speech is synthesised from mel-spectrograms by using a WaveNet vocoder .
Experiments
We used 38.5 hours of an internal long-form US English dataset recorded by a female narrator. 33 hours of this dataset was used as the training set and the remainder as the test set.
2 Evaluation
We conducted 2 separate MUSHRA tests for evaluating the naturalness of our system and the impact of different sampling techniques. Both MUSHRA tests were taken by 25 native US English listeners. Each test consisted of 50 samples which were 15 seconds in length on average. The listeners rated each system on a scale from 0 to 100 in terms of naturalness. We evaluated statistical significance using pairwise two-sided Wilcoxon signed-rank tests.
We compared each of the samplers mentioned in Section 3.2, with the baseline in a 4-system MUSHRA. Each of the samplers showed a statistically significant (all systems had a , when compared to the baseline) improvement over the baseline, as seen from Fig. 3(a). There is no statistical significance in the differences in the scores between the samplers themselves. This shows that both the BERT and Graph samplers may be equally capable at the task of sampling from the learnt latent prosodic space, therefore, even upon combining them, there is no statistically significant difference in the MUSHRA scores of the samplers. We selected the BERT+Graph sampler as the sampler for the Kathaka model, and used it to measure the reduction in gap as it had the lowest minimum score among the three samplers.
2.2 Gap Reduction Study
In this MUSHRA test, we evaluated 4 systems, namely: 1) the baseline, 2) Kathaka, 3) Kathaka-Oracle, the VAE based NTTS model with oracle prosodic embeddings obtained from recordings, and 4) the original recordings from the narrator. Kathaka obtained a relative improvement over the baseline by a statistically significant (), as shown in Fig. 3(b). The model using oracle prosodic representations, Kathaka-Oracle, reduces the gap to recordings by (). We hypothesise that the gap between Kathaka and Kathaka-Oracle is due to the sampler being trained at the sentence level. A sentence may not contain all the context required to determine the appropriate prosody with which a piece of text should be rendered, especially in long-form speech synthesis. While using a sampler trained at the sentence level shows a significant relative improvement, we speculate that exploiting the context available beyond the sentence will help to further reduce this gap and is left as future work.
Conclusion
We presented Kathaka, an NTTS model trained using a novel two-stage training approach for generating speech with contextually appropriate prosody. In the first stage of training, we learnt a distribution of sentence-level prosodic representations. We then introduced a novel sampling mechanism of using trained samplers to sample from the learnt sentence-level prosodic distribution. We introduced two samplers, 1) the BERT sampler which uses contextual word-piece embeddings from BERT and 2) the Graph sampler where we interpret constituency parse trees as graphs and use a Message Passing based Graph Attention Network on them. We then combine both these samplers as the BERT+Graph sampler, which is used in Kathaka. We also modify the baseline duration model to incorporate the latent prosodic information. We conducted an ablation study of the samplers and showed a statistically significant improvement over the baseline in each case. Finally, we compared Kathaka against a baseline, and showed a statistically significant relative improvement of .