FastPitch: Parallel Text-to-speech with Pitch Prediction
Adrian Łańcucki
INTRODUCTION
Recent advances in neural text-to-speech (TTS) enabled real-time synthesis of naturally sounding, human-like speech. Parallel models are able to synthesize mel-spectrograms orders of magnitude faster than autoregressive ones, either by relying on external alignments , or aligning themselves . TTS models can be conditioned on qualities of speech such as linguistic features and fundamental frequency . The latter has been repeatedly shown to improve the quality of neural, but also concatenative models . Conditioning on is a common approach to adding singing capabilities . or adapting to other speakers .
In this paper we propose FastPitch, a feed-forward model based on FastSpeech that improves the quality of synthesized speech. By conditioning on fundamental frequency estimated for every input symbol, which we refer to simply as a pitch contour, it matches the state-of-the-art autoregressive TTS models. We show that explicit modeling of such pitch contours addresses the quality shortcomings of the plain feed-forward Transformer architecture. These most likely arise from collapsing different pronunciations of the same phonetic units in the absence of enough linguistic information in the textual input alone. Conditioning on fundamental frequency also improves convergence, and eliminates the need for knowledge distillation of mel-spectrogram targets used in FastSpeech. We would like to note that a concurrently developed FastSpeech 2 describes a similar approach.
Combined with WaveGlow , FastPitch is able to synthesize mel-spectrograms over 60 faster than real-time, without resorting to kernel-level optimizations . Because the model learns to predict and use pitch in a low resolution of one value for every input symbol, it makes it easy to adjust pitch interactively, enabling practical applications in pitch editing. Constant offsetting of with FastPitch produces naturally sounding low- and high-pitched variations of voice that preserve the perceived speaker identity. We conclude that the model learns to mimic the action of vocal chords, which happens during the voluntary modulation of voice.
MODEL DESCRIPTION
The architecture of FastPitch is shown in Figure 1. It is based on FastSpeech and composed mainly of two feed-forward Transformer (FFTr) stacks . The first one operates in the resolution of input tokens, the second one in the resolution of the output frames. Let be the sequence of input lexical units, and be the sequence of target mel-scale spectrogram frames. The first FFTr stack produces the hidden representation . The hidden representation is used to make predictions about the duration and average pitch of every character with a 1-D CNN
Ground truth and are used during training, and predicted and are used during inference. The model optimizes mean-squared error (MSE) between the predicted and ground-truth modalities
FastPitch is robust to the quality of alignments. We observe that durations extracted with distinct Tacotron 2 models tend to differ (Figure 2), where the longest durations have approximately the same locations, but may be assigned to different characters. Surprisingly, those different alignment models produce FastPitch models, which synthesize speech of similar quality.
2 Pitch of Input Symbols
We obtain ground truth pitch values through acoustic periodicity detection using the accurate autocorrelation method . The windowed signal is calculated using Hann windows. The algorithm finds an array of maxima of the normalized autocorrelation function, which become the candidate frequencies. The lowest-cost path through the array of candidates is calculated with the Viterbi algorithm. The path minimizes the transitions between the candidate frequencies. We set windows size to match the resolution of training mel-spectrograms, to get one value for every frame.
values are averaged over every input symbol using the extracted durations (Figure 3). Unvoiced values are excluded from the calculation. For training, the values are standardized to mean of 0 and standard deviation of 1. If there are no voiced estimates for a particular symbol, its pitch is being set to 0. We have not seen any improvements from modeling in log-domain.
Following , we tried averaging to three pitch values per every symbol, in the hope to capture the beginning, middle and ending pitch for every symbol. However, the model was judged inferior (Section 4.1).
RELATED WORK
Developed concurrently to our model, FastSpeech 2 has a different approach to conditioning on and uses phoneme inputs. The predicted contour has a resolution of one value for every mel-spectrogram frame, discretized to 256 frequency values. Additionally, the model is conditioned on energy. In FastPitch, the predicted contour has one value for every input symbol. In the experiments this lower resolution made it easier for the model to predict the contour, and for the user to later modify the pitch interactively. We found that this resolution is sufficient for the model to discern between different ways of pronouncing a grapheme during training. In addition, conditioning on higher formants might increase slightly the quality, capturable by pairwise comparisons (Section 4.1.1).
The predominant paradigm in text-to-speech is two-stage synthesis: first producing mel-scale spectrograms from text, and then the actual sound waves with a vocoder model . In attempts to speed up the synthesis, parallel models have been explored. In addition to Transformer models investigated in this work , convolutional GAN-TTS is able to synthesize raw audio waveforms with state-of-the-art quality. It is conditioned on linguistic and pitch features.
The efforts in parallelizing existing models include duration prediction similar to FastSpeech, applied to Tacotron , WaveRNN , or a flow-based model . Explicit modeling of duration typically use dynamic programming algorithms associated with inference and training of HMMs. Glow-TTS aligns with Viterbi paths, and FastSpeech has been improved with a variant of the forward-backward algorithm .
Explicit neural modeling of pitch was introduced alongside a neural TTS voice conversion model , which shares similarities with other models from IBM Research . An LSTM-based Variational Autoencoder generation network modeled prosody, and pitch was calculated with a separate tool prior to the training. Prosody information was encoded in vectors of four values: log-duration, start log-pitch, end log-pitch, and log-energy.
EXPERIMENTS
The source code with pre-trained checkpointshttps://github.com/NVIDIA/DeepLearningExamples/tree/master/PyTorch/SpeechSynthesis/FastPitch, and synthesized sampleshttps://fastpitch.github.io/ are available on-line. We synthesize waveforms for evaluation with pre-trained WaveGlow .
Parameters of the model mostly follow FastSpeech . Each FFTr layer is composed of a 1-D conv with kernel size 3 and 384/1536 input/output channels, ReLU activation, a 1-D conv with kernel size 3 and 1536/384 input/output filters, followed by Dropout and Layer Norm. Duration Predictor and Pitch Predictor have the same architecture: a 1-D conv with kernel size 3 and 384/256 channels, and a 1-D conv with 256/256 channels, each followed by ReLU, Layer Norm and Dropout layers. The last layer projects every 256-channel vector to a scalar. Dropout rate is 0.1, also on attention heads.
All described models were trained on graphemes. Training on phonemes leads to a similar quality of a model, with either Tacotron 2 or Montreal Forced Aligner durations. However, the mixed approach of training on phonemes and graphemes introduced unpleasant artifacts.
FastPitch has been trained on NVIDIA V100 GPUs with 32 examples per GPU and automatic mixed precision . The training converges after 2 hours, and full training takes 5.5 hours. We use the LAMB optimizer with learning rate , , , and 110-9$1000110-6$.
We have compared our FastPitch model with Tacotron 2 (Table 1). The samples have been scored on Amazon Turk with the Crowdsourced Audio Quality Evaluation Toolkit . We have generated speech from the first 30 samples from our development subset of the LJSpeech-1.1. At least 250 scores have been gathered per every model, with the total of 60 unique Turkers participating in the study. In order to qualify, the Turkers were asked to pass a hearing test.
Generative models pose difficulties for hyperparameter tuning. The quality of generated samples is subjective, and running large-scale studies time-consuming and costly. In order to efficiently rank multiple models, and avoid score drift when the developer scores samples over a long period of time, we have investigated the approach of blindly comparing pairs of samples. Pairwise comparisons allow to build a global ranking, assuming that skill ratings are transitive .
In an internal study over 50 participants scored randomly selected pairs of samples. Glicko-2 rating system , known from rating human players in chess, sports and on-line games, but also in the context of automatic scoring of generative models , was then used to build a ranking based on those scores (Figure 4). FastPitch variations with 1, 2 and 4 attention heads, 6 and 10 transformer layers, and pitch predicted in the resolution of one and three values per input token were compared. In addition, this rating method has proven useful during development in tracking multiple hyperparameter settings, even with a handful of evaluators.
1.2 Multiple Speakers
2 Pitch Conditioning and Inference Performance
A predicted pitch contour can be modified during inference to control certain perceived qualities of the generated speech. It can be used to increase or decrease , raise expressiveness and variance of pitch. The audio samples accompanying this paper demonstrate the effects of increasing, decreasing or inverting the frequency around the mean value for a single utterance, and interpolating between speakers for the multi-speaker model. We encourage the reader to listen to them.
Inference performance measurements were taken on NVIDIA A100 GPU in FP16 precision and TorchScript-serialized models. With batch size 1, the average real-time factor (RTF) for the first 2048 utterances from LJSpeech-1.1 training set is (Table 3). With WaveGlow, RTF for complete audio synthesis drops down to . RTF measured on Intel Xeon Gold 6240 CPU is . FastPitch is suitable to real-time editing of synthesized samples, and a single pitch value per input symbol is easy to interpret by a human, which we demonstrate in a video clip on the aforementioned website with samples.
CONCLUSIONS
We have presented FastPitch, a parallel text-to-speech model based on FastSpeech, able to rapidly synthesize high-fidelity mel-scale spectrograms with a high degree of control over the prosody. The model demonstrates how conditioning on prosodic information can significantly improve the convergence and quality of synthesized speech in a feed-forward model, enabling more coherent pronunciation across its independent outputs, and lead to state-of-the-art results. Our pitch conditioning method is simpler than many of the approaches known from the literature. It does not introduce an overhead, and opens up possibilities for practical applications in adjusting the prosody interactively, as the model is fast, highly expressive, and presents potential for multi-speaker scenarios.
ACKNOWLEDGEMENTS
The author would like to thank Dabi Ahn, Alvaro Garcia, and Grzegorz Karch for their help with the experiments and evaluation of the model, and Jan Chorowski, João Felipe Santos, Przemek Strzelczyk, and Rafael Valle for helpful discussions and support in preparation of this paper.