ParaTTS: Learning Linguistic and Prosodic Cross-sentence Information in Paragraph-based TTS
Liumeng Xue, Frank K. Soong, Shaofei Zhang, Lei Xie
I Introduction
The development of sequence-to-sequence (seq2seq) based neural acoustic models and neural vocoders , brought a significant improvement to the text-to-speech (TTS) synthesis quality, which allows to automatically synthesize human-like natural speech with high fidelity. TTS is widely applied to many scenarios, such as voice assistant, navigation, smart customer service, audiobook, just to name a few, due to its high efficiency and low cost compared to manual recordings. The audiobook is an important TTS application in rendering the text of stories into expressive voices.
The text of an audiobook is composed of successive sentences, paragraphs, sections and chapters in a coherent and hierarchical form to describe a fictional story or thoughts of an author . Moreover, such a coherent and hierarchical relationship in the text has an impact on how it is being uttered in human voice formation. A paragraph, consisting of one or more sentences is a self-contained unit of a discourse in writing for conveying a particular point or idea. Prosodic patterns in paragraph audios have been observed . For instance, pitch resets — higher pitch and increased pitch range are usually observed at the beginning of the new paragraph. Similar reset pattern has also been found in energy or RMS amplitude . Declination — the tendency of pitch and energy to decline over the paragraph, ending with low pitch and energy. Lengthening — speech rate peaks in the middle of a paragraph, lengthening in the initial and final position.
To synthesize speech on a paragraph basis, a straightforward way is to synthesize each sentence in a paragraph and then combined them together. A similar approach has been used in news and navigation broadcast applications . However, there are three disadvantages: (i) an additional step of post-processing is necessary to integrate individual sentences; (ii) during combining sentences, break duration length between two consecutive sentences should be considered to ensure the integrated paragraph speech sounds natural. Pause plays a significant role in the storytelling and pause prediction is studied for early statistical parametric speech synthesis (SPSS) ; (iii) the prosody in the combined paragraph speech may become inconsistent and not smooth perceptually, varying greatly from one sentence to the next and resulting in unnatural transitions. This issue can be alleviated by minimizing the acoustic variation and linguistic distance between a sentence and the previous one . All these issues make the process of paragraph speech synthesis complex and challenging. Furthermore, it is not a suitable method for paragraph-level speech synthesis because prosodic and acoustic difference exists between sentences spoken in isolation or in a paragraph . An alternative solution is to synthesize speech at a paragraph level directly. It is feasible in the current mainstream seq2seq TTS frameworks for long-form speech synthesis , but it still tends to bring listening fatigue to the listeners because the model does not provide appropriate prosody information in a paragraph. Evidence has shown that paragraph-level or discourse-level prosody improves the naturalness and expressiveness of SPSS , but annotations of discourse structure or prosodic properties are required. Preparing a large, well-annotated labeling speech dataset is usually too time-consuming and expensive to be practical.
In this paper, we aim at building a natural and expressive paragraph-based end-to-end TTS in a data-driven way. Constrained by the memory and computing resources, we train the model sentences-wise, incorporating paragraph-based contextual information into the training process. Without discourse-level structure annotation, we try to learn paragraph linguistic knowledge from the text itself. Specifically, we adopt a paragraph text encoder to extract high-level paragraph linguistic representation. Regarding prosody, we use typical prosody features including pitch, energy and duration extracted from the corresponding paragraph audios and capture paragraph prosodic information via a paragraph prosody encoder. Meanwhile, we apply multi-head attention mechanisms to learn inner linguistic and prosodic relations between sentences with their paragraph, a sentence-position network is used to further enhance the relevant context of sentences in the paragraph.
The contributions of this paper are summarized as follows:
This paper presents, to our knowledge, the first attempt to build a paragraph-based, end-to-end TTS with the linguistic and prosodic knowledge learned in a data-driven way. The sentence-position network adopted in our model provides simple yet effective features of individual sentences in multi-sentence paragraph generation and improves the naturalness of the synthesized paragraph.
Experimental results show that our proposed model can achieve good performance for generating natural, good-quality paragraph-based speech. We demonstrate that the proposed model can also learn more accurate pause duration between consecutive sentences and has a good generalization capability of producing natural speech for a given long paragraph which can be longer, or much longer than the paragraph used in the training data.
We also find that subjectively it is rather difficult to evaluate a long paragraph due to the relative short memory in a sequential audio listening test, which is indicated early in .
II Related works
Paragraph-related works A paragraph is a self-contained unit of a discourse in writing dealing with a particular point or idea, and the paragraph in spoken discourse carries a variety of information. Discourse relations (DR) expressing how different segments (i.e. elementary discourse units) of a text are logically connected have been studied and used to improve the naturalness of statistical parametric speech synthesis (SPSS) at sentence level . In addition to discourse structure itself, some works have studied the correlations between discourse and prosody , demonstrating that discourse structure contributes to overall discourse prosody. Furthermore, the discourse structure and prosody were applied to Hidden Markov Model (HMM) based TTS to improve the prosody of passage synthesized speech . Moreover, in addition to the intra-paragraph prosody patterns, inter-paragraph prosody patterns have also been investigated and then implemented into SPSS to improve the naturalness of the synthesized articles’ speech . Our work differs from these in the following two aspects: (i) they used SPSS while we use the current mainstream end-to-end TTS approach; (ii) they need discourse relation or prosody annotations for TTS modeling while we do not need any annotations for model training.
Recently, a chapter-wise understanding system for TTS in Chinese novels is related to our work, in which the chapter-wise understanding system realizes two text understanding tasks in Chinese novels - speaker determination and emotion classification for various voices and emotional expressions speech synthesis. The differences between this and ours are that: (i) they mainly focus on speakers and emotions from the text while we focus on linguistic and prosodic knowledge from both text and the corresponding audios; (ii) they capture information at the chapter level for audiobook speech synthesis while we learn information at the paragraph level for paragraph speech synthesis.
Context-related works There has been a wide range of research focused on learning or extracting contextual information to improve the performance of TTS. Multiple studies used textual context information extracted from text to improve the sentence prosody or cross-sentence prosody for sentence-based speech synthesis, or capture conversation information for conversational speech synthesis . Specifically, the textual context information can be semantics-related features extracted by pre-trained models , i.e., BERT or syntax-related features represented by parse trees or statistics . Apart from the textual context, it has been reported that acoustic features from the previous sentence can also lead to improvement of sentence-based TTS . Following this result, a comparison investigation on multiple context representation types of the previous sentence was studied , including textual and acoustic features, utterance-level and word-level features, and representations extracted with a large pre-trained model and learned jointly with the TTS training. Our work differs from these in two folds: (i) the contextual information used in these models was derived from either isolation sentences or consecutive sentences with a predefined length, while in our work, the contextual information is extracted from a variable length paragraph which is a self-contained unit of discourse and composed of several connective sentences; (ii) cross-sentence linguistic context were used to improve sentence-level or conversation-level speech synthesis in their models. By contrast, both linguistic and prosodic cross-sentence context information is used in our work to improve paragraph-level speech synthesis.
III The proposed model
The architecture of the proposed paragraph TTS model, or ParaTTS in short, is presented in Fig. 1. It consists of a modified Tacotron2 as the TTS backbone to generate mel spectrograms from a given phoneme sequence, a linguistics-aware network and a prosody-aware network to learn the linguistic and the prosodic knowledge of the whole paragraph and the relationship between sentences and the paragraph, and a sentence-position network to enhance the correlation context of sentences and its paragraph.
In this work, phoneme sequences are used as text inputs For Chinese, one Chinese character corresponds to one syllable, and one syllable is composed of at least one phonemes and can correspond to one or more Chinese characters.. Given a text, i.e., a Chinese character sequence, it is converted to the corresponding phoneme sequence by a front-end module which includes text normalization, part-of-speech tagging and grapheme-to-phoneme conversion. If not stated otherwise, we directly use the phoneme sequences as the inputs of the model for simplicity.
The architecture of the modified Tacotron2 is an attention-based encoder-decoder TTS model. Different from the vanilla Tacotron2 , we use a CBHG encoder instead of an LSTM encoder because the former is a powerful module for extracting representations from sequences. The encoder is composed of a phoneme embedding layer, a pre-net of 2 fully connected layers and a CBHG. The CBHG is composed of a bank of 1-D convolutional filters, followed by a highway network and a bidirectional gated recurrent unit (GRU) recurrent neural net (RNN). Moreover, we adopt the GMMv2b attention mechanism . The GMMv2b is robust for long sequences because it alleviates occasional catastrophic attention failures, such as repeating or skipping.
The encoder input is the phoneme sequence of a sentence in training or the phoneme sequence of a paragraph in inference. Although the different input, the way to be processed is the same. Given a phoneme sequence x = (, , , ), n is the length of the phoneme sequence, the CBHG encoder encodes it into a hidden state h = (, , ) :
where h represents the extracted high-level phoneme representation. Then the decoder generates the current output conditioned on the previous prediction at each step:
where is the context vector calculated by the attention mechanism which encourage the decoder to attend into the important encoder hidden states when generating the output:
Thus the speech sequence y = (, , ) is generated from the input phoneme sequence x based on conditional probability p (, , , , ), and the conditional probability can be formulated as:
A linear projection function f is used to predict acoustic features directly based on decoder outputs s.
III-B The Linguistics-aware Network
The linguistics-aware network consists of a paragraph text encoder with a multi-head attention mechanism, designed to learn the linguistic information of the entire paragraph and the relationship between the component sentence in a paragraph. The paragraph text encoder is also a CBHG as described in III-A, which is used to extract high-level paragraph linguistic representation, = (, , …, ), from the given paragraph phoneme sequence, = (, , , ), m is the length of the paragraph phoneme sequence.
In the attention mechanism, attention scores reflect the importance of the key vector with respect to the query vector, allowing the query vector to concentrate on the parts of the key vector . Inspired by this, a attention is employed in this work to capture the important parts of the paragraph phoneme sequence for each sentence phoneme sequence, providing an intrinsic characterisation between a sentence and its paragraph, and meanwhile compensating the lack of paragraph-related information when training model at each sentence unit. Furthermore, the multi-head attention mechanism is utilized to explore the dependencies in different representation subspaces of the vectors. During training, the sentence hidden representation, = (, , …, ), which is encoded from the sentence phoneme sequence = (, , , ) (n is the length of the sentence phoneme sequence) via the text encoder in the modified Tacotron2, is used as the query vector Q, and the paragraph hidden representation from the encoder of the linguistic-aware network is used as the key vector K and value vector V. Thus, a linguistic context vector = (, , , ), representing the linguistic correlation between the sentence and the paragraph, is calculated as follows:
where h is the head number of the multi-head attention mechanism, and i is the i-th head. The , , and are different projection matrices for output, query, key and value, and the is the dimension of the vector K.
The final state of the last bidirectional GRU layer in the paragraph text encoder accumulates the information forward and backward through the whole paragraph phoneme sequence. That is , the last element of the paragraph linguistic representation = (, , …, ). We view it as a condensed paragraph linguistic representation and add it with the context vector as the output of the linguistic-aware network. The output contains the linguistic information of the paragraph and the relationship between the sentence and its paragraph. Finally, we add the output and the encoder output in the TTS backbone model to make the model aware of the linguistic knowledge related to paragraphs.
III-C The Prosody-aware Network
The prosody-aware network consists of 4 submodules: paragraph prosody extractor, paragraph prosody predictor and paragraph prosody encoder with a multi-head attention mechanism.
The input of the paragraph prosody encoder is a 3-dimensional, phoneme-level, paragraph prosody feature vector. In training, these features are extracted from the corresponding paragraph speech by the paragraph prosody extractor. In inference, prosody features are predicted by the paragraph prosody predictor which is a 64-unit GRU layer and a 3-unit dense layer. The 3-dimensional prosody features include mean-variance normalized logarithmic fundamental frequency (LF0), intensity, and duration. The paragraph prosody encoder is similar to but different from the reference encoder depicted in . It consists of 6 convolution layers with batch normalization, followed by a 128-unit GRU layer. Here, we use the output of each time step in the GRU as a variable-length paragraph prosody representation, denoted as = (, , …, ), which has the same length (m) as the paragraph phoneme representation because the prosody feature input is at the phoneme level.
Similar to the linguistics-aware network, a multi-head attention mechanism is also used to associate the component sentences with the corresponding paragraph. Specifically, the query vector Q is the addition result of the encoder output of the TTS backbone and the output of the linguistics-aware network. The key vector K and the value vector V are both the paragraph prosody representation . Accordingly, the prosodic context vector = (, , , ), representing the prosodic relationship between a sentence and its parent paragraph, can be calculated by the Eqs. 6-8.
Additionally, the final state of GRU in the paragraph prosody encoder, i.e., the last element of the paragraph prosody representation = (, , …, ), can be viewed as a compressed paragraph prosody representation. We add it with the prosodic context vector to form the output of the prosody-aware network. The output represents the prosodic knowledge related to the paragraph and relations among component sentences in the paragraph. To make the TTS model aware of the prosodic information, we add it with the encoder output in the TTS backbone model.
In inference, the trained paragraph prosody predictor is used to predict the 3-dimensional, phone-level prosody features of the paragraph. The predictor input is the addition result of the linguistic-aware network and the encoder output of the TTS model. We conjecture that the input can predict the paragraph-level prosody features since it contains rich paragraph-relevant information. We train the paragraph prosody predictor using L1 loss between predicted and extracted prosody features and stop gradient flow to ensure the prosody prediction error does not affect the linguistics-aware network and the encoder of the TTS backbone.
III-D The Sentence-position Network
Distinctive differences in prosody can be found at the beginning, middle, and end sentences of a paragraph . We analyze the corresponding prosody patterns in paragraphs in section IV-A. To incorporate this prosodic information into the TTS model, we utilize a sentence-position network module which is composed of an up-sampling layer followed by a linear layer. The input to the sentence-position network is a 3-dimensional, one-hot code encoded as the first, middle or last sentence in the paragraph. The up-sampling layer up-samples the position code from sentence to phoneme level by replicating it. Finally, the linear layer is used to project the 3-dimensional, phone-level position code to the preset dimension of 256, facilitating the addition operation with the encoder output of the TTS backbone model.
An example of the sentence position code is illustrated in Fig. 2. Given a sentence, the sentence position in the paragraph can be determined and the corresponding phoneme sequence can also be obtained through the front-end module. The sentence position codes of the first, the last and the middle sentences are encoded as 0, 2 and 1, respectively. Then the sentence position is up-sampled from sentence level to phoneme level according to the corresponding phoneme sequence length.
The TTS backbone model, conditioned on (1) the sentence position code, and (2) the linguistic and prosodic information of the paragraph, generates the mel spectrograms of sentences from the phoneme sequences of sentences in training or the mel spectrograms of paragraphs from the phoneme sequences of the paragraphs in inference. The total loss L of the proposed model to be optimized is:
where is the mel spectrogram reconstruction loss of mean square error (MSE), is stop token loss of cross entropy, and is prosody prediction loss of MSE. , , are the weights of the corresponding losses, respectively.
IV Experiments
We first introduce the basic information of the corpus used in our experiments and then analyze the statistics of the corpus, intra-paragraph and inter-paragraph patterns. For experimental tests, we calculate objective metrics and also conduct subjective evaluations to measure the performance of the proposed model in generating paragraph speech.
Basic information In this work, we use a fairy-tale audio-book corpus to train and evaluate the proposed model. The information of the corpus is listed in Table I. The corpus contains 44 stories recorded by a Chinese female mimicking children’s voices, about 4.27 hours in total. We randomly select 40 stories as the training set and split the stories into 801 paragraphs and 2,525 utterances, about 4.08 hours. The rest of the 4 stories are used as the test set, which is split into 32 paragraphs.
Statistical analysis Then text length (in terms of sentences and Chinese characters) and speech duration (in seconds) statistics are presented in Fig. 3, in which (a), (b), and (c) are the distributions of the number of sentences in a paragraph, the number of Chinese characters in a sentence and a paragraph, respectively. On the average, each paragraph has 3 sentences, 55 Chinese characters, and each sentence has 17 Chinese characters. Additionally, (d) and (e) present the distributions of the duration length of sentences and paragraphs. The average duration lengths of a sentence and a paragraph are 6.5 and 11.4 seconds, respectivelyThe average number of sentences in a paragraph is 3, but the mode, referring to the value that occurs with the highest frequency is 2 due to the asymmetrical nature of the distribution. Consequently, the average duration length of a paragraph is 11.4 seconds, which is about twice as long as that of a sentence (6.5 seconds)..
Intra-paragraph prosody patterns analysis To understand the variation of prosody features within a paragraph, we perform a statistical analysis of the individual prosodic features in sentences in the first, middle and last positions in a paragraph. Specifically, we calculate the mean values of sentence-level prosody features in different positions and plot the variation curves as exemplified in Fig. 4, where two prosody patterns are observed. Declination: pitch and intensity decline along with the paragraph. Lengthening: speech rate is faster in the middle than in the initial and final position of a paragraph, lengthening in the initial and final position. We observe that the range of prosodic features among three different sentence positions is not big, which may be attributed to two reasons: the scale of the corpus is relatively small, and the speaking style is not very distinctively different.
Inter-paragraph prosody patterns analysis We also analyze the prosody feature variation across paragraph boundaries (break) and within a paragraph (no break), as shown in Table II. The values in the break are the mean discrepancy of prosody features between the last sentence in the current paragraph and the first sentence in the next paragraph. And the values in no break are the mean discrepancy of prosody features between the current sentence and the succeeding sentence within a paragraph. The positive values in no break suggest that there is a declination of prosody attributes within a paragraph. While, the negative values in break indicate that the phenomenon of prosody reset appears in paragraph boundaries, which means the LF0, intensity and speech rate increase at the beginning of a new paragraph. In this work, we focus on individual paragraph speech synthesis so that the inter-paragraph prosody patterns are not considered.
IV-B Experimental Setup
The relatively small size of the training data used in this study is a challenge to training a highly stable, end-to-end TTS model. Thus, we firstly pre-train the backbone of the modified Tacotron2 using a standard TTS corpus, containing 17.83-hour reading-style Chinese female speech data. Then we fine-tune the model for ParaTTS using the audio-book corpus. In the training stage, sentence phoneme sequences are fed into the model as input. Mel spectrograms are extracted from recordings, which are down-sampled from 44.1 kHz to 16kHz, and used as the target output. To obtain phone-level prosody features in paragraph prosody extractor, we first extract frame-level LF0 and intensity values using the Python library of ParselmouthParselmouth provides a complete and Pythonic interface to the internal Praat code that can be used for speech analysis of pitch and intensity. Parselmouth can be found athttps://github.com/YannickJadoul/Parselmouth and Praat can be found at https://www.fon.hum.uva.nl/praat/ . Meanwhile, we conduct force alignment using Hidden Markov Model Toolkit (HTK) to get the phone duration, and then calculate the phone-level LF0 and energy by averaging frame-level values based on the frame lengths of each phone. We train models on a single GPU with batch size of 16 up to 400k steps for the pre-trained model and 200k steps for the fine-tuned models, using Adam optimizer with = 0.9 and = 0.999. The hyper-parameters of the model used in our experiments are described in Table III. At the inference stage, paragraph phoneme sequence is fed into the model as input. The output mel spectrogram is transformed into a waveform using the multi-band WaveRNN vocoder , which is pre-trained to 500k steps using the standard corpus and adapted to 200k steps using the audio-book corpus.
In our evaluations, we compare the following five models for paragraph-level speech synthesis.
Baseline: the modified Tacotron2 as described in Section III-A.
LingTTS: Baseline with the linguistics-aware network.
ProsTTS: Baseline with the prosody-aware network.
ComTTS: Baseline with the combination of the linguistics-aware and prosody-aware networks.
ParaTTS: ComTTS with the sentence-position network.
IV-C Objective Evaluation
Naturalness We calculate mel-cepstrum distortion (MCD) to measure the naturalness objectively. Before computing MCD, we use dynamic time warping to align predicted and target mel spectrogram sequences because the lengths of the two sequences can be different. The MCD result is shown in the second column in Table IV. The model with linguistics-aware network (LingTTS) or prosody-aware network (ProsTTS) or both (ComTTS) gets a similar MCD and outperforms the Baseline. With the sentence-position network, the model (ParaTTS) decreases MCD further and achieves the lowest MCD.
Prosody To measure the synthesized prosody in a paragraph, we calculate the Pearson correlation coefficient on the prosody features at the syllable level, including LF0, intensity and duration. The LF0 and duration correlation results and the corresponding p-values are presented in Table IV. We do not list the intensity correlation because all models have a comparable high correlation value, around 0.90. Regarding the LF0 and duration correlations, all models achieve good correlations, and ParaTTS is the best. From the results, it is observed that even though ProsTTS is more related to the prosody, it does not achieve better results than the LingTTS. This may be because the prosody prediction heavily dependents on the linguistic information. Consequently, ComTTS which combines the linguistic-aware network and prosody-aware network does not achieve additional benefits compared with LingTTS. After adding sentence-position network, ParaTTS obtains improved performance, indicating the benefits of the sentence position information, which explicitly provides the simple yet effective features of individual sentences in multi-sentence paragraph generation.
We also generate the results by feeding the ground-truth prosody features to the ParaTTS, referred as ParaTTS (GT_prosody), to show the upper bound of the proposed model. The objective evaluations of ParaTTS (GT_prosody) are presented in Table IV as well. Obviously, ParaTTS (GT_prosody) achieves the best results. Additionally, we also observe that ParaTTS (GT_prosody) has a similar LF0 correlation with ParaTTS but obtains better MCD. The results indicate that the ground-truth prosody features benefit the naturalness and the prosody prediction in the proposed model is appropriate. To visualize the prosody prediction performance, we plot the pitch contours of a paragraph from the recording and the synthesized results of the Baseline, LingTTS, ProsTTS ComTTS and ParaTTS, as shown in Fig. 5. We can observe that the pitch contour of ParaTTS is the closest to that of the recordings compared to the other models.
Break It has been shown that pause duration is highly correlated to the discourse structure . The pause duration, particularly between successive sentences, is used to introduce suspense and climax in storytelling, which can enhance the audience’s attraction to the story and build some anticipation . We calculate the root mean square error (RMSE) of pause duration between consecutive sentences in a paragraph to explore if the models can learn the break across sentences. To be specific, the pause duration between two sentences is the difference value between the end time of the last word in the current sentence and the start time of the first word in the next sentence.
The pause duration RMSE result is shown in the last column in Table IV. The Baseline achieves the worst RMSE result. The other four models all achieve better results than the Baseline, in which ComTTS is slightly better than others. We observe that ParaTTS can learn a more accurate pause duration than the Baseline. Fig. 6 shows mel spectrograms of the recording and the synthesized paragraph speech by the Baseline, LingTTS, ProsTTS, ComTTS and ParaTTS, in which white boxes are the pause duration between two consecutive sentences. We can find that the pause duration in Baseline is too long to appropriate, which may cause unnatural perceived performance. The corresponding samples (numbered 1.3) can be found in our sound sample pagehttps://lmxue.github.io/paratts/. We conjecture that the multi-head mechanisms between the sentence and multi-sentence paragraph in linguistics-aware and prosody-aware networks can learn the cross-sentence context in training, including the pause duration.
IV-D Subjective Evaluation
A group of 20 listening subjects who are native Chinese speakers with normal hearing participates in the subjective tests and evaluates each paragraph as a whole rather than in its isolated sentences, a similar testing was conducted in .
Preference test The preference test between two models is to choose which model is preferable based upon the overall perceived impression. We first perform preference tests among Baseline, ComTTS and ParaTTS to compare the effectiveness of the linguistics-aware network, prosody-aware network and sentence-position network. The preference test results are presented in Fig. 7. ComTTS gets 30% more preference than the Baseline, indicating that the linguistic and prosodic information can indeed improve paragraph-based speech synthesis. With the additional sentence position information, ParaTTS gets an extra preference (4%) over ComTTS. This small preference gain is also consistent with the slightly better MCD and prosody correlations. As described in the intra-paragraph prosody patterns analysis, the prosody variation range among three different sentence positions is not large, hence leading to a relatively smaller improvement. In the following subjective tests, we compare Baseline, ParaTTS and recordings with mean opinion score (MOS) scores in in-domain and out-of-domain tests.
Detailed MOS test To further evaluate the perceived quality of the synthesized paragraph speech, we conduct 5-point mean opinion score (MOS) tests in four different dimensions: naturalness, pleasantness, pause and listening comfortThe explanations for four criteria are: naturalness: ”very natural” to ”unnatural”; pleasantness: ”very unpleasant” to ”very pleasant”; pause: “speech pauses confusing/unpleasant” to “speech pauses appropriate/pleasant”; listening comfort: “very exhausting” to “very easy”. .
The detailed MOS test scores are shown in Fig. 8. We observe that the MOS scores difference is not big, which was similarly observed in . We conjecture that may be due to the fact that long-form paragraph samples are too long for subjects to remember all the differences for a clearly distinguishable score. The results can still shed some light on the power of ParaTTS, which obtains scores consistently higher than the Baseline in all four testing fronts, where better naturalness and pleasantness are also reflected in the lower MCDs and higher LF0 correlations. The pause (break) MOS score of the ParaTTS is almost close to that of the recording, which is also confirmed with lower RMSE in pause duration. Listening to the multi-sentence paragraph audios synthesized by the ParaTTS does not increase the listening fatigue of the listeners and gets an on-par listening comfort score with the recordings, possibly helped by its naturalness and pleasantness close to the recordings.
IV-E Out-of-domain Test
To examine the proposed model’s generalization ability, we extend the test to out-of-domain, long paragraphs and extra-long paragraphs. The information on the in-domain and out-of-domain test sets is listed in Table V. In addition to the in-domain 38 short paragraphs, there are 12 long paragraphs and 6 extra-long paragraphs. On the average, each paragraph contains 5, 23 and 51 utterances, respectively. The overall impression MOS results are shown in Fig. 9. It is observed that the MOS scores of the recordings decrease with the length of the paragraph increases. In other words, even for the original recorded speech, long paragraphs tend to be getting lower MOS scores or causing more listening fatigue than shorter paragraphs. In the extra-long paragraphs testing set, there is still occasional skipping issue even though we adopt the robust GMMv2b attention mechanism, indicating that it is still challenging for long-sequence modeling in the attention mechanism.
ParaTTS yields higher scores than the Baseline, not only for the in-domain short paragraphs but for the out-of-domain, long and extra-long paragraphs. In inference, the multi-head attention in the linguistics-aware network plays a role of a self-attention mechanism. The encoders in both the TTS backbone and linguistic-aware network take the paragraph phoneme sequence as input to exploit the contextual embedding vector for rendering more natural speech. In this way, the advantage of long-range dependency utilized by the self-attention mechanism generalizes the model for synthesizing longer paragraphs.
V Conclusion
In this research, we propose to use a new, paragraph-based, end-to-end TTS model to model linguistic and prosodic information embedded in paragraph text with the corresponding acoustic data. We design both linguistics-aware and prosody-aware networks to learn the information via a paragraph encoder and its multi-head attention mechanism. Additionally, a sentence-position network is used to exploit the inter-sentence information in the paragraph. Trained on a storytelling, audio-book corpus (4.08 hours), recorded by a female Mandarin speaker, experimental results show that the proposed new paragraph-based model can produce TTS speech better than the conventional sentence-based TTS baseline system, both objectively and subjectively. The new model can learn the cross-sentence information well, e.g., the break durations between adjacent sentences, and generalize the learned information to longer or much longer paragraphs than those used in the training corpus.