A comparison of recent waveform generation and acoustic modeling methods for neural-network-based speech synthesis

Xin Wang, Jaime Lorenzo-Trueba, Shinji Takaki, Lauri Juvela, Junichi Yamagishi

Introduction

Text-to-speech (TTS) synthesis aims at converting a text string into a speech waveform. A conventional TTS pipeline normally includes a frond-end text-analyzer and a back-end speech synthesizer, each of which may include multiple modules for specific tasks. For example, the back-end based on statistical parametric speech synthesis (SPSS) uses statistical acoustic models to map linguistic features from the text-analyzer into compact acoustic features. It then uses a vocoder to generate a waveform from the acoustic features.

Recent TTS systems have well adopted to the impact of deep learning. First, frameworks different from the established TTS pipeline are defined. One framework merges the text-analyzer and acoustic models into a single neural network (NN), where the input text is directly mapped into acoustic features, e.g., by Tactron and Char2wav . Another framework is the unified back-end represented by Wavenet that directly converts the linguistic features into a waveform. There are also novel NN back-ends without the use of normal acoustic features, vocoders , or duration model . In parallel to these new frameworks, modules for the conventional TTS pipeline have also been developed. For instance, acoustic models based on the generative adversarial network (GAN) and autoregressive (AR) NN have been reported to alleviate the over-smoothing effect in acoustic modeling. There is also the Wavenet-based vocoder that uses deep learning rather than signal processing methods to attain waveform generation.

Given the recent progress with TTS, we think it is important to understand the pros and cons of the new methods on the basis of common conditions. As an initial step, we compared the new methods that could easily be plugged into the SPSS-based back-end of the classical TTS pipeline. These methods were divided into two groups, and thus the comparison included two parts. The first part compared acoustic models based on four types of NN, including the normal recurrent neural network (RNN), shallow AR-based RNN (SAR), and two GAN-based postfilter networks that were attached to the RNN and SAR. The second part compared waveforms generation techniques, including the high-quality source-filter vocoder WORLD , another synthesizer based on the log-domain pulse model (PML) , WORLD plus the Griffin-Lim algorithm for phase recovery , and the Wavenet-based vocoder.

Section 2 explains the details on the selected methods used for comparison. Section 3 explains the experiments and Section 4 draws conclusions.

Speech synthesis methods

This comparison study was based on perceptual evaluation of synthetic speech. The selected acoustic models and waveform generation methods were combined into the seven speech synthesis methods given in Fig. 1, where SAR-Wa, SAR-Pm, SAR-Pr, and SAR-Wo use the same acoustic models but different waveform generation modules, while SAR-Wo, GAS-Wo, GAR-Wo, and RNN-Wo use the same WORLD vocoder but differ in the acoustic model to generate the Mel-generalized coefficients (MGC) and band aperiodicity (BAP). The F0 for all the methods is generated by a deep AR (DAR) model . The oracle duration is used for all the methods. Neither formant enhancement nor MLPG was used.

Conventional vocoders assume that sn\bm{s}_{n} has a minimum phase in order to approximately recover sn\bm{s}_{n} from cn\bm{c}_{n}. However, the speech waveform is not merely a minimum-phase signal. One solution is to model the phase as additional acoustic features . Another way is to directly model complex-valued sn\bm{s}_{n} by using complex-valued NN or restricted Boltzman machine . Alternatively, a statistical model such as the Wavenet-based vocoder can be used rather than using a deterministic vocoder. Such a vocoder may learn the phase information not encoded by the acoustic features.

This work used a conventional vocoder called WORLD as the baseline. It then included a phase-recovery technique , a waveform synthesizer based on a log-domain pulse model , and a Wavenet-based vocoder for comparison. Complex-valued approaches may be included in future work.

1.2 Deterministic vocoders

The WORLD vocoder is similar to the legacy STRAIGHT vocoder and uses a source-filter-based speech production model with mixed excitation. It also assumes a minimum phase for the spectrum. By using the procedure in Sec. 2.1.1 reversely, WORLD converts the cepstral features, cn\bm{c}_{n}, given by the acoustic model into a linear amplitude spectrum. Then, WORLD creates the excitation signal by mixing a pulse and a noise signal in the frequency domain, where each frequency band is weighted by a BAP value. Finally, it generates a speech waveform based on the source-filter model.

Since WORLD assumes a minimum phase, this work included a phase-recovery method that may enhance the phase of the generated waveform by WORLD. This method re-computes the amplitude spectrum given the generated waveform and then applies the Griffin-Lim phase-recovery algorithm. Note that this method is different from that in our previous work where phase-recovery was directly conducted on the generated amplitude spectrum.

This work also included the pulse model in the log-domain (PML) as another waveform generator. One advantage of PML is that it avoids voicing decision per frame and may alleviate the perceptual degradation caused by voicing decision errors. In addition, PML may better represent noise in voiced segments. The PML and WORLD in this work used the same generated spectrum as input. The noise mask used by PML is generated by another RNN given input linguistic features, which it is not indicated in Fig. 1.

1.3 Data-driven vocoder

Besides the deterministic vocoders, this work included a statistical vocoder based on Wavenet . Suppose ot\bm{o}_{t} is a one-hot vector representing the quantized waveform sampling point at time tt, the Wavenet vocoder infers the distribution P(ot∣ot−R:t−1,at)P(\bm{o}_{t}|\bm{o}_{t-R:t-1},\bm{a}_{t}) given the acoustic feature at\bm{a}_{t} and previous observations ot−R:t−1={ot−R,⋯ ,ot−1}\bm{o}_{t-R:t-1}=\{\bm{o}_{t-R},\cdots,\bm{o}_{t-1}\}, where RR is the receptive field of the network. In this work, at{\bm{a}}_{t} contained the F0 and MGC. Note that, while natural at\bm{a}_{t} is used to train the vocoder, a^t\widehat{\bm{a}}_{t} generated by the acoustic model is used for waveform generation.

During generation, o^t\widehat{\bm{o}}_{t} is usually obtained by random sampling, i.e., o^t∼P(ot∣o^t−R:t−1,a^t)\widehat{\bm{o}}_{t}\sim P(\bm{o}_{t}|\widehat{\bm{o}}_{t-R:t-1},\widehat{\bm{a}}_{t}). However, we found that the synthetic waveform sounded better when samples in the voiced segments were generated as o^t=arg⁡max⁡otp(ot∣o^t−R:t−1,a^t)\widehat{\bm{o}}_{t}=\arg\max_{\bm{o}_{t}}p(\bm{o}_{t}|\widehat{\bm{o}}_{t-R:t-1},\widehat{\bm{a}}_{t}). This method is referred to as the greedy generation method. Fig. 4 plots the spectrogram and the instantaneous frequency (IF) of generated waveforms from the two methods, i.e., SAR-Wa (greedy) and SAR-Wa (random). Compared with random sampling, the greedy method generated more regular IF patterns in voiced segments. It perceptually made the waveforms sound less trembling. One reason for this may be that the ‘best’ generated o^t\widehat{\bm{o}}_{t} was temporally more consistent with o^t−R:t−1\widehat{\bm{o}}_{t-R:t-1}. Note that samples in the unvoiced parts were still randomly drawn, and the voiced/unvoiced state for each time could be inferred from the F0 in a^t\widehat{\bm{a}}_{t}.

2 Neural network acoustic models

The NN-based acoustic models included in this work aims at inferring the distribution of the acoustic feature sequence a1:N={a1,⋯ ,aN}\bm{a}_{1:N}=\{\bm{a}_{1},\cdots,\bm{a}_{N}\} conditioned on the linguistic feature l1:N={l1,⋯ ,lN}\bm{l}_{1:N}=\{\bm{l}_{1},\cdots,\bm{l}_{N}\} in NN frames. The baseline RNN, whether it has a recurrent output layer or not, defines the distribution as

where N(⋅)\mathcal{N}(\cdot) is the Gaussian distribution, I\bm{I} is the identity matrix, hn=HΘ(l1:N,n)\bm{h}_{n}=\mathcal{H}_{\bm{\Theta}}(\bm{l}_{1:N},n) is the outcome of the output layer at frame nn, and Θ\bm{\Theta} is the network’s weight. For generation, a^1:N\widehat{\bm{a}}_{1:N} can be acquired by using a mean-based method that sets a^1:N=h1:N\widehat{\bm{a}}_{1:N}=\bm{h}_{1:N}.

A shortcoming of RNN is that the across-frame dependency in a1:N\bm{a}_{1:N} is ignored. Hence, this work included a model that takes into account the dependency in a causal direction . This model, which is called shallow AR (SAR), defines the distribution as:

This model uses FΦ(an−K:n−1)=∑k=1Kβk⊙an−k+γ\mathcal{F}_{\bm{\Phi}}(\bm{a}_{n-K:n-1})=\sum_{k=1}^{K}\bm{\beta}_{k}\odot\bm{a}_{n-k}+\bm{\gamma} to merge the acoustic features in the previous KK frames and then changes the distribution for frame nn, which builds the causal across-frame dependency. Another deep AR (DAR) model similar to SAR can be defined . DAR in this work was only used to model the quantized F0 for all the experimental methods, and thus is not explained here.

As RNN and SAR are trained by using the maximum-likelihood criterion and then generate by using the mean-based method, they may not produce acoustic features with natural textures. Therefore, this work also included the GAN-based postfilter to enhance RNN and SAR. The GAN discriminators in this work did not crop the input acoustic features into small patches, which is different from original work. Instead, they made a true-false judgment every frame.

Experiments

This work used a Japanese speech corpus of neutral reading speech uttered by a female speaker. The duration is 50 hours, and the number of utterance is 30,016, out of which 500 were randomly selected as a validation set and another 500 were used as a test set. Linguistic features were extracted by using OpenJTalk . The dimension of these features vectors was 389. Acoustic features were extracted at a frame rate of 200 Hz (5 ms) from the 48 kHz waveforms. MGC and BAP were extracted by using WORLD. The dimensions for MGC were 60 and those for BAP were 25. The F0 was extracted by an assembly of multiple pitch trackers and then quantized into 255 levels for quantized F0 modeling .

2 Neural network configuration

The network structures of the acoustic models are plotted in Fig. 2. In DAR, RNN and SAR, the size of feedforward (FF), bi- and uni-directional LSTM layer was 512, 256, and 128. The size of the linear layer depended on the output features’ dimensions. For SAR, the output layer included a Gaussian distribution for MGC with the AR parameter K=1K=1 and another Gaussian for BAP with K=0K=0. In GAN, all the CNN layers had 256 output channels and conducted 1D convolution. Each FF layer in the GAN generator changed the dimension of a CNN’s output before skip-add operation. To reduce instability in low-dimensional MGCs, the input MGC to the discriminator was scaled element-wisely. The scaling weight was 0.0010.001 for the first five MGC dimensions, 0.010.01 for the next five dimensions, and 11 for the rest. The weight was for BAP.

All the acoustic models and Wavenet are implemented on a modified CURRENNT toolkit . This toolkit and training recipes can be found online (http://tonywangx.github.io).

3 Evaluation environment

The seven speech synthesis methods in Fig.1 were evaluated in terms of speech quality and speaker similarity. Natural and vocoded speech from WORLD and PML were also included in the evaluation. All the systems except for the Wavenet-based SAR-Wa were rated at sampling rates of 48 and 16 kHz. The speech samples of these systems were natural or generated at 48 kHz, and the 16 kHz samples were down-sampled from these 48 kHz samples. SAR-WA only generated samples of 16 kHz. All samples were normalized to -26 dBov, and a total of 19 groups of speech samples were evaluated.

The evaluation was crowdsourced online. Each evaluation set had 19 screens, i.e., one for each system. The order of systems in each set was randomly shuffled. The evaluators must answer two questions on each screen. First, they listened to a sample of the system under evaluation and rated the naturalness on a 1-to-5 MOS scale. They then rated the similarity of that test sample to the natural 48 kHz sample on a 1-to-5 MOS scale. The participants were allowed to replay the samples. The samples in one set were synthesized for the same text randomly selected from the test set.

4 Results and discussion

We collected a total of 1500 evaluation sets, i.e., 1500 scores for each system. A total of 235 native Japanese listeners participated, with an average of 6.4 sets per person. The statistical analysis was based on unpaired t-tests with a 95% confidence margin and Holm-Bonferroni compensation. The results are plotted in Fig. 3.

On acoustic modeling, the comparisons of the speech quality score among RNN-Wo, RGA-Wo, SAR-Wo, and SGA-Wo indicated that SAR (3.03) >> SGA (2.82) ∼\sim RGA (2.81) >> RNN (2.53) for both 48 kHz and 16 kHz. The same ordering can be seen in terms of speaker similarity. RNN’s unsatisfactory performance may be due to the over-smoothing effect on the generated MGC, which muffled the synthetic speech. This effect is indicated by the global variance (GV) plotted in Fig. 5, where the GV of the generated MGC from RNN was smaller than that of the other models. The performance of SAR is consistent with that in our previous work , which indicated that the AR model alleviated the over-smoothing effect. The GAN-based postfilter in both SGA and RGA also reduced the impact of over-smoothing. Further, GAN not only enhanced GV but also compensated for the modulation spectrum (MS) of the generated MGC throughout the whole frequency band, which can be seen from Fig. 6. However, neither SGA nor RGA outperformed SAR even though SAR did not boost the MS in the high frequency bands. After listening to the samples, we found that, while the samples of SGA and RGA had better spectrum details, they contained more artifacts. This indicated that generators in SGA and RGA may not optimally learn the distribution of natural MGC. Future work will look into details on SGA and RGA.

On waveform generation methods, the comparison among SAR-Wa, SAR-Pr, SAR-Pm, SAR-Wo, and vocoded speech showed that SAR-Wa significantly outperformed other waveform generation methods in terms of speech quality. More interestingly, SAR-Wa’s quality was even judged to be better than other synthetic methods by using a 48 kHz sampling rate. The quality of SAR-Wa was also close to the 16 kHz vocoded samples from Abs-Pm and Abs-Wo. A closes inspection of the samples showed that the Wavenet vocoder in SAR-Wa had fewer artifacts, such as amplitude modulation in the waveform and other effects possibly due to conventional vocoders. However, the gap between SAR-Wa and Abs-Wo/Abs-Pm was still large in terms of speaker similarity. One reason may be that, while the Wavenet vocoder trained by using natural acoustic features may well learn the mapping from acoustic features to waveforms, during waveform generation it cannot compensate for the degradation of generated acoustic features caused by the imperfect acoustic model SAR. Note that this problem may not be resolved by training the Wavenet vocoder using generated acoustic features if the acoustic model is not good enough.

Among other waveform generation methods, the PML vocoder, while being similar to WORLD for analysis-by-synthesis, it lagged behind when using the generated acoustic features. This result is somewhat different from that in , and future work will investigate the reasons for this. Comparison between SAR-Pr and SAR-Wo indicated that the phase recovery method did not improve the quality or similarity score. We suspect that this is because the iterative process in the Griffin-Lim algorithm may introduce some artifacts that degraded speech quality.

Conclusion

This work was our initial step in building a framework in which recent acoustic modeling and waveform generation methods could be compared on a common ground by using a large-scale perceptual evaluation. This work only considered a few methods that could easily be plugged into the common TTS pipeline. On acoustic models, the results showed that the autoregressive model SAR could achieve a better performance than a normal RNN. This SAR was also easier to train than a GAN-based model. On waveform generation methods, this work demonstrated the potential of the Wavenet vocoder and the advantage of statistical waveform modeling compared with conventional deterministic approaches.

We intend to investigate the reasons for the experiments’ results in more detail in future work. Meanwhile, we recently validated that SAR is a simple case of using normalizing flow to transform the density of target data. We will try to improve SAR and the overall performance of speech synthesis system. Complex-valued models that support the modeling of waveform phases may also be included.

References