Efficiently Trainable Text-to-Speech System Based on Deep Convolutional Networks with Guided Attention

Hideyuki Tachibana, Katsuya Uenoyama, Shunsuke Aihara

Introduction

Text-to-speech (TTS) is getting more and more common recently, and is getting to be a basic user interface for many systems. To further promote the use of TTS in various systems, it is significant to develop a manageable, maintainable, and extensible TTS component that is accessible to speech non-specialists, enterprising individuals and small teams.

Traditional TTS systems, however, are not necessarily friendly for them, since they are typically composed of many domain-specific modules. For example, a typical parametric TTS system is an elaborate integration of many modules e.g. a text analyzer, an F0F_{0} generator, a spectrum generator, a pause estimator, and a vocoder that synthesize a waveform from these data, etc.

Deep learning may integrate these internal building blocks into a single model, and connects the input and the output directly. This type of technique is sometimes called ‘end-to-end’ learning. Although such a technique is sometimes criticized as ‘a black box,’ an end-to-end TTS system named Tacotron , which directly estimates a spectrogram from an input text, has achieved promising performance recently, without using hand-engineered parametric models based on domain-specific knowledge. Tacotron, however, has a drawback that it uses many recurrent units which are quite costly to train. It is almost infeasible for ordinary labs who do not have luxurious machines to study and extend it further. In fact, some people tried to implement open clones of Tacotron , but they are struggling to reproduce the speech of satisfactory quality as clear as the original results.

The purpose of this paper is to show Deep Convolutional TTS (DCTTS), a novel fully convolutional neural TTS. The architecture is largely similar to Tacotron , but is based on a fully convolutional sequence-to-sequence learning model similar to the literature . We show that this fully convolutional TTS actually works in a reasonable setting. The contribution of this article is twofold: (1) Propose a fully CNN-based TTS system which can be trained much faster than an RNN-based state-of-the-art neural TTS system, while the sound quality is still acceptable. (2) An idea to rapidly train the attention module, which we call ‘guided attention,’ is also shown.

Recently, deep learning-based TTS systems have been intensively studied, and surprisingly high quality results are obtained in some of recent studies. The TTS systems based on deep neural networks include Zen’s work , the studies based on RNN e.g. , and recently proposed techniques e.g. WaveNet [13, sec. 3.2], Char2Wav , DeepVoice1&2 , and Tacotron .

Some of them tried to reduce the dependency on hand-engineered internal modules. The most extreme technique in this trend would be Tacotron , which depends only on mel and linear spectrograms, and not on any other speech features e.g. F0F_{0}. Our method is close to Tacotron in a sense that it depends only on these spectral representations of audio signals.

Most of the existing methods above use RNN, a natural technique of time series prediction. An exception is WaveNet, which is fully convolutional. Our method is also based only on CNN but our usage of CNN would be different from WaveNet, as WaveNet is a kind of a vocoder, or a back-end, which synthesizes a waveform from some conditioning information that is given by front-end components. On the other hand, ours is rather a front-end (and most of back-end processing). We use CNN to synthesize a spectrogram, from which a simple vocoder can synthesize a waveform.

1.2 Sequence to Sequence (((seq2seq))) Learning

Recently, recurrent neural networks (RNN) have become a standard technique for mapping a sequence to another sequence, especially in the field of natural language processing, e.g. machine translation , dialogue system , etc. See also [1, sec. 10.4].

RNN-based seq2seq, however, has some disadvantages. Firstly, a vanilla encode-decoder model cannot encode too long sequence into a fixed-length vector effectively. This problem has been resolved by a mechanism called ‘attention’ , and the attention mechanism now has become a standard idea of seq2seq learning techniques; see also [1, sec. 12.4.5.1].

Another problem is that RNN typically requires much time to train, since it is less suited for parallel computation using GPUs. To overcome this problem, several researchers proposed the use of CNN instead of RNN, e.g. . Some studies have shown that CNN-based alternative networks can be trained much faster, and sometimes can even outperform the RNN-based techniques.

Gehring et al. recently united these two improvements of seq2seq learning. They proposed an idea on how to use attention mechanism in a CNN-based seq2seq learning model, and showed that the method is quite effective for machine translation. The method we proposed is based on the similar idea to the literature .

Preliminary

2 Notation: Convolution and Highway Activation

In this paper, we denote the 1D convolution layer by a space saving notation Ck⋆δo←i(X)\mathsf{C}{}^{{o\leftarrow i}}_{{k}\star{\delta}}(X), where ii is the size of input channel, oo is the size of output channel, kk is the size of kernel, δ\delta is the dilation factor, and an argument XX is a tensor having three dimensions (batch, channel, temporal). The stride of convolution is always 11. Convolution layers are preceded by appropriately-sized zero padding, whose size is suitably determined by a simple arithmetic so that the length of the sequence is kept constant. Let us similarly denote the 1D deconvolution layer as Dk⋆δo←i(X)\mathsf{D}{{}^{{o\leftarrow i}}_{{k}\star{\delta}}}(X). The stride of deconvolution is always 22 in this paper. Let us write a layer composition operator as ⋅◃⋅\cdot\triangleleft\cdot, and let us write networks like F◃ReLU◃G(X):=F(ReLU(G(X))),\mathsf{F}\triangleleft\mathsf{ReLU}\triangleleft\mathsf{G}(X):=\mathsf{F}(\mathsf{ReLU}(\mathsf{G}(X))), and (F◃G)2(X):=F◃G◃F◃G(X)(\mathsf{F}\triangleleft\mathsf{G})^{2}(X):=\mathsf{F}\triangleleft\mathsf{G}\triangleleft\mathsf{F}\triangleleft\mathsf{G}(X), etc. ReLU\mathsf{ReLU} is an element-wise activation function defined by ReLU(x)=max⁡(x,0)\mathsf{ReLU}(x)=\max(x,0).

Convolution layers are sometimes followed by a Highway Network -like gated activation, which is advantageous in very deep networks: Highway(X;L)=σ(H1)⊙ReLU(H2)+(1−σ(H1))⊙X,\mathsf{Highway}(X;\mathsf{L})=\sigma(H_{1})\odot\mathsf{ReLU}(H_{2})+(1-\sigma(H_{1}))\odot X, where H1,H2H_{1},H_{2} are tensors of the same shape as XX, and are output as [H1,H2]=L(X)[H_{1},H_{2}]=\mathsf{L}(X) by a layer L\mathsf{L}. The operator ⊙\odot is the element-wise product, and σ\sigma is the element-wise sigmoid function. Hereafter let us denote HCk⋆δd←d(X):=Highway(X;Ck⋆δ2d←d)\mathsf{H}\mathsf{C}{}^{{d\leftarrow d}}_{{k}\star{\delta}}(X):=\mathsf{Highway}(X;\mathsf{C}{}^{{2d\leftarrow d}}_{{k}\star{\delta}}).

Proposed Network

Since some literature suggest that the staged synthesis from low- to high-resolution has advantages over the direct synthesis of high-resolution data, we synthesize the spectrograms using the following two networks. (1) Text2Mel, which synthesizes a mel spectrogram from an input text, and (2) Spectrogram Super-resolution Network (SSRN), which synthesizes a full STFT spectrogram from a coarse mel spectrogram. Fig. 2 shows the overall architecture.

The resultant Y1:F,2:T+1Y_{1:F,2:T+1} needs to approximate the temporally-shifted ground truth S1:F,2:T+1S_{1:F,2:T+1}. The error is evaluated by a loss function Lspec(Y1:F,2:T+1∣S1:F,2:T+1)\mathcal{L}_{\text{spec}}(Y_{1:F,2:T+1}|S_{1:F,2:T+1}), and is back-propagated to the network parameters. The loss function was the sum of L1 loss and a function Dbin\mathcal{D}_{\text{bin}} which we call binary divergence,

The networks are fully convolutional, and are not dependent on any recurrent units. In order to take into account the long contextual information, we used dilated convolution instead of RNN.

The top equation of Fig. 2 is the content of TextEnc\mathsf{TextEnc}. It consists of a character embedding and several 1D non-causal convolution layers. In the literature an RNN-based component named ‘CBHG’ was used, but we found this convolutional network also works well. AudioEnc\mathsf{AudioEnc} and AudioDec\mathsf{AudioDec}, shown in Fig. 2, are composed of 1D causal convolution layers. These convolution should be causal because the output of AudioDec\mathsf{AudioDec} is feedbacked to the input of AudioEnc\mathsf{AudioEnc} at the synthesis stage.

2 Spectrogram Super-resolution Network (SSRN)

The bottom equation of Fig. 2 shows SSRN\mathsf{SSRN}. In this paper, all convolutions of SSRN are non-causal, since we do not consider online processing. The loss function is the same as Text2Mel: the sum of L1 distance and Dbin\mathcal{D}_{\text{bin}}, defined by (5), between the synthesized spectrogram SSRN(S)\mathsf{SSRN}(S) and the ground truth ∣Z∣|Z|.

Guided Attention

In general, an attention module is quite costly to train. Therefore, if there is some prior knowledge, incorporating them into the model may be a help to alleviate the heavy training. We show that the following simple method helps to train the attention module.

Although this measure is based on quite a rough assumption, it improved the training efficiency. In our experiment, if we added the guided attention loss to the objective, the term began decreasing only after ∼\sim100 iterations. After ∼\sim5K iterations, the attention became roughly correct, not only for training data, but also for new input texts. On the other hand, without the guided attention loss, it required much more iterations. It began learning after ∼\sim10K iterations, and it required ∼\sim50K iterations to look at roughly correct positions, but the attention matrix was still vague. Fig. 3 compares the attention matrix, trained with and without guided attention loss.

2 Forcibly Incremental Attention at the Synthesis Stage

At the synthesis stage, the attention matrix AA sometimes fails to focus on the correct characters. Typical errors we observed were (1) skipping several letters, and (2) repeating a same word twice or more. In order to make the system more robust, we heuristically modified the matrix AA to be ‘nearly diagonal,’ by a simple rule as follows. We observed this method sometimes alleviated such failures.

Experiment

We used LJ Speech Dataset to train the networks. This is a public domain speech dataset consisting of ∼\sim13K pairs of text and speech, without phoneme-level alignment, ∼\sim24 hours in total. These speech data have a little reverberation. We preprocessed the texts by spelling out some of abbreviations and numeric expressions, decapitalizing the capitals, and removing less frequent characters not shown in Table 2, where NULL is a dummy character for zero-padding.

We implemented our neural networks using Chainer 2.0 . We trained the models using a household gaming PC equipped with two GPUs. The main memory of the machine was 62GB, which is much larger than the audio dataset. Both GPUs were NVIDIA GeForce GTX 980 Ti, with 6 GB memories.

For simplicity, we trained Text2Mel and SSRN independently and asynchronously using different GPUs. All network parameters were initialized using He’s Gaussian initializer . Both networks were trained by the ADAM optimizer . When training SSRN, we randomly extracted short sequences of length T=64T=64 for each iteration to save memory usage. To reduce the disk access, we reduced the frequency of creating the snapshot of parameters to only once per 5K iterations. Other parameters are shown in Table 2.

As it is not easy for us to reproduce the original results of Tacotron, we instead used a ready-to-use model for comparison, which seemed to produce the most reasonable sounds in the open implementations. It is reported that this model was trained using LJ Dataset for 12 days (877K iterations) on a GTX 1080 Ti, newer GPU than ours. Note, this iteration is still much less than the original Tacotron, which was trained for more than 2M iterations.

We evaluated mean opinion scores (MOS) of both methods by crowdsourcing on Amazon Mechanical Turk using crowdMOS toolkit . We used 20 sentences from Harvard Sentences List 1&2. The audio data were synthesized using five methods shown in Table 2. The crowdworkers rated these 100 clips from 1 (Bad) to 5 (Excellent). Each worker is supposed to rate at least 10 clips. To obtain more responses of higher quality, we set a few incentives shown in the literature. The results were statistically processed using the method shown in the literature .

2 Result and Discussion

In our setting, the training throughput was ∼\sim3.8 minibatch/s (Text2Mel) and ∼\sim6.4 minibatch/s (SSRN). This implies that we can iterate the updating formulae of Text2Mel 200K times in 15 hours. Fig. 4 shows an example of attention, synthesized mel and full spectrograms, after 15 hours training. It shows that the method can almost correctly focus on the correct characters, and synthesize quite clear spectrograms. More samples are available at the author’s web page.https://github.com/tachi-hi/tts_samples

In our crowdsourcing experiment, 31 subjects evaluated our data. After the automatic screening by the toolkit , 560 scores rated by 6 subjects were selected for final statistics calculation. Table 2 compares the performance of our proposed method (DCTTS) and an open Tacotron. Our MOS (95% confidence interval) was 2.71±0.662.71\pm 0.66 (15 hours training) while the Tacotron’s was 2.07±0.622.07\pm 0.62. Although it is not a strict comparison since the frameworks and the machines were different, it would be still concluded that our proposed method is quite rapidly trained to the satisfactory level compared to Tacotron.

Note that the MOS were below the level reported in . The reasons may be threefold: (1) the limited number of iterations, (2) SSRN needs to synthesize the spectrograms from less information than , and (3) the reverberation of the training data.

Summary and Future Work

This paper described a novel TTS technique based on deep convolutional neural networks, and a technique to train the attention module rapidly. In our experiment, the proposed Deep Convolutional TTS was trained overnight (∼\sim15 hours), using an ordinary gaming PC equipped with two GPUs, while the quality of the synthesized speech was almost acceptable.

Although the audio quality is far from perfect yet, it may be improved by tuning some hyper-parameters thoroughly, and by applying some techniques developed in the deep learning community.

We believe this method will encourage further development of the applications based on speech synthesis. We can expect that this simple neural TTS may be extended to many other purposes e.g. emotional/non-linguistic/personalized speech synthesis, etc., by further studies. In addition, since a neural TTS has become lighter, the studies on more integrated speech systems e.g. some multimodal systems, may have become more feasible. These issues should be worked out in the future.

Acknowledgement

The authors would like to thank the OSS contributors and the data creators (LibriVox contributors and @keithito).

References