Adversarially Trained End-to-end Korean Singing Voice Synthesis System

Juheon Lee, Hyeong-Seok Choi, Chang-Bin Jeon, Junghyun Koo, Kyogu Lee

Introduction

With the recent development of deep learning, a learning-based singing voice synthesis (SVS) system, which synthesizes sounds as natural as the concatenative method , but can expand more flexibly, is proposed. For example, three SVS systems based on DNN, LSTM, and Wavenet architecture were proposed, respectively . These systems all include an acoustic model that is trained by singing, lyrics, and sheet music paired data, and each acoustical model is trained to predict the vocoder feature used as an input to the vocoder.

Although these neural network-based SVS system can achieve adequate performance, networks predicting vocoder features have limits that cannot exceed the upper bound of vocoder performance. Therefore, it is meaningful to propose an end-to-end framework that directly generates a linear-spectrogram, not a vocoder feature. However, the extension to the end-to-end framework of the SVS system is a challenging task because it involves increased complexity of the model. Creating a more complex target, linear-spectrogram, increases the complexity of the model and requires as much training data to generalize and train these models sufficiently. However, gathering singing audio with aligned lyrics in a controlled environment is a task that requires a lot of effort.

We proposed in this paper a Korean SVS system that can be trained by an end-to-end manner with moderate amounts of data The generated result can be found at: ksinging.strikingly.com.. Our baseline network is designed with the inspiration of DCTTS , known as efficiently trainable text to speech (TTS) system. We applied the following novel approaches to enable end-to-end network training. First, we used the phonetic enhancement masking method, which separately modeled low-level acoustic features related to pronunciation from text information, to make more efficient use of the information contained in the training data. Second, we also proposed a method of reusing input data at the super-resolution stage and training with an adversarial manner to produce better sound quality singing.

The contribution of this paper is as follows: 1) We designed the end-to-end Korean SVS system and suggested a way to train it effectively. 2) We proposed a phonetic enhancement masking method that helps to produce more accurate pronunciation. 3) We proposed a conditional adversarial training method for the generation of more realistic singing voices.

Related work

The SVS system is similar to the TTS system in terms of synthesizing natural human speech. Recently, the end-to-end TTS system, which is trained as an autoregressive manner, such as Tacotron, Deep voice, is showing better performance than the conventional method. In addition, various follow-up studies are being conducted that have further controllable elements such as prosody, style, etc , or models that can be trained more efficiently . We conducted the study by modifyng the TTS model to suit the SVS task, based on DCTTS, which is known to be capable of efficient end-to-end training.

The generative adversarial networks (GAN) is a widely used technique that helps train an arbitrary function to generate a similar sample as the sample from desired data distribution. This training method has been widely accepted in computer vision community and becomes one of the key components to attain photo-realism in super-resolution task. Unlike the success of adversarial training method in image domain, however, only a few works have achieved a reasonable success of training super-resolution task (specifically, band-width extension task) in audio signal processing community . To further leverage the promise of adversarial training in audio generation process, we adopted a few recent works that stabilizes the adversarial training, namely, conditional GAN with projection discriminator and R1 regularization which allow us to jointly train the autoregressive network (mel-synthesis) and super-resolution network making the proposed system as an end-to-end framework.

Proposed Network

As illustrated in Figure 1, our proposed model consists of two main modules, a mel-synthesis network and a super-resolution network. The mel-synthesis network is trained to produce a mel-spectrogram M1:LM_{1:L} from previous mel input M0:L−1M_{0:L-1}, time-aligned text T1:LT_{1:L}, and pitch inputs P1:LP_{1:L}. With text and pitch information as conditional input, the super-resolution network upsamples the generated mel-spectrogram MM to a linear-spectrogram SS. Finally, the discriminator takes the upsampled result with generated mel-spectrogram to train the network in an adversarial manner. During the test phase, a sequence of mel-spectrogram frames is generated in an autoregressive manner from a given text and pitch input which is then upsampled to linear-spectrogram by super resolution network. Finally, the generated linear-spectrogram is converted to a waveform using Griffin-Lim algorithm .

2 Mel-synthesis network

The mel-synthesis network MS(⋅)MS(\cdot{}) aims to generate the mel-spectrogram of the next time step from the given text, pitch, and mel input. Based on the text-to-mel network proposed by we modified it to fit the SVS system.

First, in order to enter pitch information, we added pitch encoders with the same structure as text encoders. In addition, the local conditioning method proposed by was used to conduct a conditioning of the encoded pitch on the mel decoder.

Second, we assumed that among the various elements forming a singing voice, information about pronunciation would be able to be controlled independently from text information. We also assumed that if the low-level audio feature that constitutes pronunciation information can be modeled independently, it is possible to focus on the pronunciation information in the data composed of various combinations of pronunciation-pitch, so that training data can be utilized more efficiently to generate more accurate pronounced singing voice. To this end, we designed an additional phonetic enhancement mask decoder, which receives encoded text only as input, and the output of the decoder element-wise multiplied by the output of the mel decoder to create the final mel-spectrogram. As a result, MS(⋅)MS(\cdot) can be formulated as follows:

We trained the MS(⋅)MS(\cdot) network with L1L_{1} and binary divergence loss LdL_{d} between ground truth and generated mel-spectrogram, and guided attention loss LattL_{att} as the objective function. Please see for more detailed explanation on the loss terms. We also assumed that the L1L_{1} loss between the differential spectrogram M′=M1:L−M0:L−1M^{\prime}=M_{1:L}-M_{0:L-1} would be also beneficial for network to learn more about the relatively short pronounced onset, coda. Therefore, the overall objective function for MS(⋅)MS(\cdot) is as follows:

3 Super-resolution network

Unlike the attention-mechanism based TTS literature, SVS system requires the aligned text and pitch information as inputs for the controllability in the generation process. These information, therefore, can be easily reused in the SR step in the absence of time-alignment process as follows.

More specifically, each of the output from the text encoder and pitch encoder (ET,VE_{T,V} and EPE_{P}) is fed into a sequence of 1×11\times 1 convolutional and dropout layer which is then fed into a highway network as a local conditioning method as proposed in . For the upsampled S^\hat{S}, SR is trained with the objective function LSR=L1(S^,S)+Ld(S^,S)L_{SR}=L_{1}(\hat{S},S)+L_{d}(\hat{S},S). For the exact network configuration, please refer to Figure 1.

3.2 Adversarial training method

Expecting to generate a realistic sound, we adopted a conditional adversarial training method which helps the output distribution of S^=SR(M^,⋅)\hat{S}=SR(\hat{M},\cdot) be similar to the real data distribution S∼p(S∣M)S\sim p(S|M). Intuitively, in the conditional adversarial training framework, discriminator DψD_{\psi} not only tries to check if SS is realistic but also the paired correspondence between SS and MM. Note that, we make a minor assumption that the distribution of M^=MS(⋅)\hat{M}=MS(\cdot) approximately follows that of MM, that is, p(M)≃p(M^)p(M)\simeq p(\hat{M}), allowing the joint training of two modules MS(⋅)MS(\cdot) and SR(⋅)SR(\cdot). The conditioning to discriminator was done by following with a minor modification. First, the condition MM is fed into a 1d-convolutional layer and the intermediate output of discriminator is fed into a 3×33\times 3 2d-convolutional layer. Then, inner product between the two outputs is done as a projection. Finally, the obtained scalar value is added to the last layer of DψD_{\psi} resulting in final logit value. For the exact network configuration please refer to Figure 1.

For the stable adversarial training, a regularization technique on DψD_{\psi} has been proposed by several GAN related works . We adopted a simple, yet, effective gradient penalty technique called R1 regularization. This technique penalizes the squared 2-norm of the gradients of DψD_{\psi} only when the sample from true distribution is taken as follows

Note that the output of DψD_{\psi} denotes the logit value before the sigmoid function. The final adversarial loss terms (LadvDL_{adv_{D}} and LadvGL_{adv_{G}}) for DψD_{\psi} and GθG_{\theta} are as follows,

where θ\theta includes not only the parameters of SRSR but also that of MSMS, hence the two consecutive modules acting as one generator function Gθ=SR(MS(⋅),⋅)G_{\theta}=SR(MS(\cdot),\cdot). The function ff is chosen as follows f(t)=−log(1+exp(−t))f(t)=-log(1+exp(-t)) resulting in the vanilla GAN loss as in the original GAN paper .

Experiments

Since there is no publicly available Korean singing voice dataset, we created the dataset as follows. First, we prepared accompaniment and singing voice MIDI files of 60 Korean pop songs. Next, a professional female vocalist was told to sing to the accompaniment. Then, the singing voice MIDI files were manually realigned so that the recorded audio have the exact alignment with the singing voice MIDI files. Finally, we manually assigned the syllables in lyrics to each MIDI note of singing voice MIDI file. The audio length of the entire dataset excluding the silence is about 2 hours. We used 49 songs for training dataset, 1 song for validation, and 10 songs for test dataset.

2 Training

We trained the discriminator to minimize LadvDL_{adv_{D}} and the rest of the network to minimize LadvGL_{adv_{G}}, LMSL_{MS} and LSRL_{SR}. For SR networks, we have to start training after the appropriate level of mel is generated, so we have separately controlled lrSRlr_{SR} and lrGANlr_{GAN} to add to the objective function. At this point, it was set to lrSR=min(0.2∗(iter/100),1)lr_{SR}=\rm{min}(0.2*(iter/100),1), lrGAN=min(0.01∗(int)(iter/5000),1)lr_{GAN}=\rm{min}(0.01*(\rm{int})(iter/5000),1), respectively.

In both cases, we used Adam optimizer , which was set to β1\beta_{1} = 0.5 and β2\beta_{2} = 0.9. The learning rate was scheduled to start from 0.0002 and was halved for every 30,000 iteration. All parameters of the networks were initialized with the Xavier initializer .

For the ground truth mel/linear-spectrogram, we first extracted the linear-spectrogram SS from audio with sr=22050,nfft=1024,hop=256sr=22050,n_{fft}=1024,hop=256. We then normalized the linear-spectrogram as follows S←(∣S∣/max(∣S∣))δS\leftarrow(|S|/max(|S|))^{\delta}, where δ\delta denotes a pre-emphasis factor with the value of 0.6 in our case. Note that we post emphasized S^←S^ζ/δ\hat{S}\leftarrow\hat{S}^{\zeta/\delta} where ζ\zeta denotes a post-emphasis factor with the value of 1.3. Afterwards, the mel-spectrogram was obtained by multiplying 80-d of mel filter bank to SS, and the same normalization method as in SS was used. In order to reduce the complexity of the model, we downsampled the mel-spectrogram to the quarter by taking the first frame of every four-frame of the mel-spectrogram giving the relationship of L′=4LL^{\prime}=4L.

3 Evaluation

We trained a total of five models to see how the three proposed methods - method 1; phonetic enhancement masking, method 2; local conditioning pitch and text to SR(⋅)SR(\cdot), method 3; adversarial training method - actually affect the network. The differences between the five models are described in Table 1. 20 audio samples from each model were generated from the test dataset. Apart from the generated samples, we also compared the ground truth samples. Ground denotes the actual recorded audio, and Recons denotes the reconstructed audio from ground truth magnitude only linear-spectrogram using Griffin-Lim algorithm. Noe that Recons samples were included to evaluate the sound quality from the loss of phase information.

We evaluated whether the network was actually producing a conditioned singing voice for a given input. To do this, we extracted f0 sequence from the generated audio through the world vocoder, converted it into a pitch sequence, and compared it to the input pitch sequence. We can judge that the higher the similarity between the two sequence, the more the network generates a singing that reflects the input condition. We calculated the precision, recall and f-score of the generated pitch sequence by frame-wise, and the results are shown in Table 1.

Even in the case of a real recording sample recorded by listening to the original midi accompaniment, it is not easy to adjust the timing and pitch of the correct note, so that a 100% accurate f-score can not be obtained. For all samples that were generated, a f-score similar to or higher than the real recording sample was obtained. This means that the model has generated a singing voice with the correct pitch and timing for at least the real recording for the given input.

3.2 Qualitative evaluation

We conducted a listening test to evaluate the quality of the generated singing voice. 19 native Korean speakers were asked to listen to the 20 audio samples from each model. Each participant was asked to evaluate the pronunciation accuracy, sound quality, and naturalness. During the listening test, lyrics of audio samples were provided for more accurate evaluation of pronunciation accuracy. The MOS results are shown in Table 1.

We conducted a paired t-test for each model response and based on this we verified the effectiveness of the proposed methods. For the accuracy of the pronunciation, we obtained significant differences for all comparisons except for models 2 and 3. In other words, all of the proposed methods helped to create more accurate pronunciation singing voices, and the performance was improved to the greatest extent with all three methods. In the case of sound quality, methods 1 and 2 did not significantly affect the improvement, but the applying method 3 showed a significant increase in score. From this we can confirm that training the network in an adversarial manner improves the quality of the generated audio. Finally, for naturalness, there was a significant improvement when all methods were applied.

4 Analysis on generated spectrogram

In this section we analyze the features generated by the mel-synthesis and super-resolution networks. In the case of mel-synthesis network, from observing internally generated features, we found that the low-level acoustic feature of pronunciation and pitch could be divided independently without any supervision. From Figure 3, DMD_{M} shows the underlying structure of the spectrogram, such as the harmonic structure and the location of f0. In MaskMask, on the other hand, we can observe the shape of determining the intensity of the frequency at every time-step, similar to the feature of the spectral envelope, which contains non-periodic information. This suggests that, from the perspective of source-filter models, one of the techniques that classical speech modelling techniques, our network can generate sources (DMD_{M}) and filters (MaskMask) separately from frequency domain without any supervised training.

We also analyzed the effect of adversarial training method by observing the generated linear-spectrogram. Three different spectrograms from model4 (S′^\hat{S^{\prime}}: w/o adversarial loss), model5 (S^\hat{S}: w/ adversarial loss), and ground truth spectrogram (Sˉ\bar{S}) are demonstrated in the second row of Figure 3. While S′^\hat{S^{\prime}} showing the blurry high frequency areas, S^\hat{S} clearly shows that adversarial training allows the proposed network to generate sample that is closer to the ground truth sample Sˉ\bar{S}. Note that we have confirmed in 4.3.2, listening test that the sound quality can be significantly improved by comparing model4 and model5, which again reinforces our observation.

Conclusions

In this paper, we proposed the end-to-end Korean singing vocie synthesis system. We showed that using text information to model the phonetic enhancement mask actually worked, and produced more accurate pronunciation. Also, we successfully applied the conditional adversarial training method to the super-resolution stage, which resulted in a higher quality voice.

Acknowledgements

This work has partly supported by National Research Foundation of Korea (NRF) funded by the Korea government (NRF-2017R1E1A1A01076284), and partly by Institute for Information & Communications Technology Planning & Evaluation(IITP) grant funded by the Korea government (No.2019-0-01367)

References