DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism
Jinglin Liu, Chengxi Li, Yi Ren, Feiyang Chen, Zhou Zhao
Introduction
Singing voice synthesis (SVS) which aims to synthesize natural and expressive singing voice from musical score (Wu and Luan 2020), increasingly draws attention from the research community and entertainment industries (Zhang et al. 2020). The pipeline of SVS usually consists of an acoustic model to generate the acoustic features (e.g., mel-spectrogram) conditioned on a music score, and a vocoder to convert the acoustic features to waveform (Nakamura et al. 2019; Lee et al. 2019; Blaauw and Bonada 2020; Ren et al. 2020; Chen et al. 2020)A music score consists of lyrics, pitch and duration..
Previous singing acoustic models mainly utilize simple loss (e.g., L1 or L2) to reconstruct the acoustic features. However, this optimization is based on the incorrect uni-modal distribution assumptions, leading to blurry and over-smoothing outputs. Although existing methods endeavor to solve this problem by generative adversarial network (GAN) (Lee et al. 2019; Chen et al. 2020), training an effective GAN may occasionally fail due to the unstable discriminator. These issues hinder the naturalness of synthesized singing.
Recently, a highly flexible and tractable generative model, diffusion probabilistic model (a.k.a. diffusion model) (Sohl-Dickstein et al. 2015; Ho, Jain, and Abbeel 2020; Song, Meng, and Ermon 2021) emerges. Diffusion model consists of two processes: diffusion process and reverse process (also called denoising process). The diffusion process is a Markov chain with fixed parameters (when using the certain parameterization in (Ho, Jain, and Abbeel 2020)), which converts the complicated data into isotropic Gaussian distribution by adding the Gaussian noise gradually; while the reverse process is a Markov chain implemented by a neural network, which learns to restore the origin data from Gaussian white noise iteratively. Diffusion model can be stably trained by implicitly optimizing variational lower bound (ELBO) on the data likelihood. It has been demonstrated that diffusion model can produce promising results in image generation (Ho, Jain, and Abbeel 2020; Song, Meng, and Ermon 2021) and neural vocoder (Chen et al. 2021; Kong et al. 2021) fields.
In this work, we propose DiffSinger, an acoustic model for SVS based on diffusion model, which converts the noise into mel-spectrogram conditioned on the music score. DiffSinger can be efficiently trained by optimizing ELBO, without adversarial feedback, and generates realistic mel-spectrograms strongly matching the ground truth distribution.
To further improve the voice quality and speed up inference, we introduce a shallow diffusion mechanism to make better use of the prior knowledge learned by the simple loss. Specifically, we find that there is an intersection of the diffusion trajectories of the ground-truth mel-spectrogram and the one predicted by a simple mel-spectrogram decoder Here we use a traditional acoustic model based on feed-forward Transformer (Ren et al. 2021; Blaauw and Bonada 2020), which is trained by L1 loss to reconstruct mel-spectrogram.: sending and into the diffusion process could result in similar distorted mel-spectrograms, when the diffusion step is big enough (but not reaches the deep step where the distorted mel-spectrograms become Gaussian white noise). Thus, in the inference stage we 1) leverage the simple mel-spectrogram decoder to generate ; 2) calculate the sample at a shallow step through the diffusion process: , where T is the total number of diffusion steps. can be calculated in closed form time (Ho, Jain, and Abbeel 2020).; and 3) start reverse process from rather than Gaussian white noise, and complete the process by iteration denoising steps (Vincent 2011; Song and Ermon 2019; Ho, Jain, and Abbeel 2020). Besides, we train a boundary prediction network to locate this intersection and determine the adaptively. The shallow diffusion mechanism provides a better start point than Gaussian white noise and alleviates the burden of the reverse process, which improves the quality of synthesized audio and accelerates inference.
Finally, since the pipeline of SVS resembles that of text-to-speech (TTS) task, we also build DiffSpeech adjusting from DiffSinger for generalization. The evaluations conducted on a Chinese singing dataset demonstrate the superiority of DiffSinger (0.11 MOS gains compared with a state-of-the-art acoustic model for SVS (Wu and Luan 2020)), and the effectiveness of our novel mechanism (0.14 MOS gains, 0.5 CMOS gains and 45.1% speedup with shallow diffusion mechanism). The extensional experiments of DiffSpeech on TTS task prove the generalization of our methods (0.24/0.23 MOS gains compared with FastSpeech 2 (Ren et al. 2021) and Glow-TTS (Kim et al. 2020) respectively). The contributions of this work can be summarized as follows:
We propose DiffSinger, which is the first acoustic model for SVS based on diffusion probabilistic model. DiffSinger addresses the over-smoothing and unstable training issues in previous works.
We propose a shallow diffusion mechanism to further improve the voice quality, and accelerate the inference.
The extensional experiments on TTS task (DiffSpeech) prove the generalization of our methods.
Diffusion Model
In this section, we introduce the theory of diffusion probabilistic model (Sohl-Dickstein et al. 2015; Ho, Jain, and Abbeel 2020). The full proof can be found in previous works (Ho, Jain, and Abbeel 2020; Kong et al. 2021; Song, Meng, and Ermon 2021). A diffusion probabilistic model converts the raw data into Gaussian distribution gradually by a diffusion process, and then learns the reverse process to restore the data from Gaussian white noise (Sohl-Dickstein et al. 2015). These processes are shown in Figure 1.
Define the data distribution as , and sample . The diffusion process is a Markov chain with fixed parameters (Ho, Jain, and Abbeel 2020), which converts into the latent in steps:
At each diffusion step , a tiny Gaussian noise is added to to obtain , according to a variance schedule :
If is well designed and is sufficiently large, then is nearly an isotropic Gaussian distribution (Ho, Jain, and Abbeel 2020; Nichol and Dhariwal 2021). Besides, there is a special property of diffusion process that can be calculated in closed form in time (Ho, Jain, and Abbeel 2020):
where , .
The reverse process is a Markov chain with learnable parameters from to . Since the exact reverse transition distribution is intractable, we approximate it by a neural network with parameters ( is shared at every -th step):
Thus the whole reverse process can be defined as:
To learn the parameters , we minimize a variational bound of the negative log likelihood:
where is a constant. And by reparameterizing Eq. (1) as , and choosing the parameterization:
Sample from and run the reverse process to obtain a data sample.
DiffSinger
As illustrated in Figure 2, DiffSinger is built on the diffusion model. Since SVS task models the conditional distribution , where is the mel-spectrogram and is the music score corresponding to , we add to the diffusion denoiser as the condition in the reverse process. In this section, we first describe a naive version of DiffSinger (Section 3.1); then we introduce a novel shallow diffusion mechanism to improve the model performance and efficiency (Section 3.2); finally, we describe the boundary prediction network which can adaptively find the intersection boundary required in shallow diffusion mechanism (Section 3.3).
In the naive version of DiffSinger (without dotted line boxes in Figure 2): In the training procedure (shown in Figure 2(a)), DiffSinger takes in the mel-spectrogram at -th step in the diffusion process and predicts the random noise in Eq. (6), conditioned on and the music score . The inference procedure (shown in Figure 2(b)) starts at the Gaussian white noise sampled from , as the previous diffusion models do (Ho, Jain, and Abbeel 2020; Kong et al. 2021). Then the procedure iterates for times to repeatedly denoise the intermediate samples with two steps: 1) predict the using the denoiser; 2) obtain from using the predicted , according to Eq. (2) and Eq. (5):
where when , and when . Finally, a mel-spectrogram corresponding to could be generated.
2 Shallow Diffusion Mechanism
Although the previous acoustic model trained by the simple loss has intractable drawbacks, it still generates samples showing strong connectionThe samples fail to maintain the variable aperiodic parameters, but they usually have a clear “skeleton” (harmonics) matching the ground truth. to the ground-truth data distribution, which could provide plenty of prior knowledge to DiffSinger. To explore this connection and find a way to make better use of the prior knowledge, we conduct the empirical observation leveraging the diffusion process (shown in Figure 3): 1) when , has rich details between the neighboring harmonics, which can influence the naturalness of the synthesized singing voice, but is over-smoothing as we introduced in Section 1; 2) as increases, samples of two process become indistinguishable. We illustrate this observation in Figure 4: the trajectory from manifold to Gaussian noise manifold and the trajectory from to Gaussian noise manifold intersect when the diffusion step is big enough.
Inspired by this observation, we propose the shallow diffusion mechanism: instead of starting with the Gaussian white noise, the reverse process starts at the intersection of two trajectories shown in Figure 4. Thus the burden of the reverse process could be distinctly alleviatedConverting into is easier than converting (Gaussion white noise) into (). Thus the former could improve the quality of synthesized audio and accelerates inference.. Specifically, in the inference stage we 1) leverage an auxiliary decoder to generate , which is trained with L1 conditioned on the music score encoder outputs, as shown in the dotted line box in Figure 2(a); 2) generate the intermediate sample at a shallow step through the diffusion process, as shown in the dotted line box in Figure 2(b) according to Eq. (1):
where , , . If the intersection boundary is properly chosen, it can be considered that and come from the same distribution; 3) start reverse process from , and complete the process by iterations denoising.
The training and inference procedures with shallow diffusion mechanism are described in Algorithm 1 and 2 respectively. The theoretical proof of the intersection of two trajectories can be found in the supplement.
3 Boundary Prediction
where is the training set of mel-spectrograms. When BP have been trained, we determine using the predicted value of BP, which indicates the probability of a sample classified to be 1. For all , we find the earliest step where the 95% steps in satisfies: the margin between BP(, ) and BP(, ) is under the threshold. Then we choose the average of as the intersection boundary .
We also propose an easier trick for boundary prediction in the supplement by comparing the KL-divergence. Note that the boundary prediction can be considered as a step of dataset preprocessing to choose the hyperparameter for the whole dataset. actually can be chosen manually by brute-force searching on validation set.
4 Model Structures
The encoder encodes the music score into the condition sequence, which consists of 1) a lyrics encoder to map the phoneme ID into embedding sequence, and a series of Transformer blocks (Vaswani et al. 2017) to convert this sequence into linguistic hidden sequence; 2) a length regulator to expand the linguistic hidden sequence to the length of mel-spectrograms according to the duration information; and 3) a pitch encoder to map the pitch ID into pitch embedding sequence. Finally, the encoder adds linguistic sequence and pitch sequence together as the music condition sequence following (Ren et al. 2020).
The diffusion step is another conditional input for denoiser , as shown in Eq. (6). To convert the discrete step to continuous hidden, we use the sinusoidal position embedding (Vaswani et al. 2017) followed by two linear layers to obtain step embedding with channels.
We introduce a simple mel-spectrogram decoder called the auxiliary decoder, which is composed of stacked feed-forward Transformer (FFT) blocks and generates as the final outputs, the same as the mel-spectrogram decoder in FastSpeech 2 (Ren et al. 2021).
Denoiser takes in as input to predict added in diffusion process conditioned on the step embedding and music condition sequence . Since diffusion model imposes no architectural constraints (Sohl-Dickstein et al. 2015; Kong et al. 2021), the design of denoiser has multiple choices. We adopt a non-causal WaveNet (Oord et al. 2016) architecture proposed by (Rethage, Pons, and Serra 2018; Kong et al. 2021) as our denoiser. The denoiser is composed of a convolution layer to project with channels to the input hidden sequence with channels and convolution blocks with residual connections. Each convolution block consists of 1) an element-wise adding operation which adds to ; 2) a non-causal convolution network which converts from to channels; 3) a convolution layer which converts the to channels; 4) a gate unit to merge the information of input and conditions; and 5) a residual block to split the merged hidden into two branches with channels (the residual as the following and the “skip hidden” to be collected as the final results), which enables the denoiser to incorporate features at several hierarchical levels for final prediction.
The classifier in the boundary predictor is composed of 1) a step embedding to provide ; 2) a ResNet (He et al. 2016) with stacked convolutional layers and a linear layer, which takes in the mel-spectrograms at -th step and to classify and .
More details of model structure and configurations are shown in the supplement.
Experiments
In this section, we first describe the experimental setup, and then provide the main results on SVS with analysis. Finally, we conduct the extensional experiments on TTS.
Since there is no publicly available high-quality unaccompanied singing dataset, we collect and annotate a Chinese Mandarin pop songs dataset: PopCS, to evaluate our methods. PopCS contains 117 Chinese pop songs (total 5.89 hours with lyrics) collected from a qualified female vocalist. All the audio files are recorded in a recording studio. Every song is sampled at 24kHz with 16-bit quantization. To obtain more accurate music scores corresponding to the songs (Lee et al. 2019), we 1) split each whole song into sentence pieces following DeepSinger (Ren et al. 2020) and train a Montreal Forced Aligner tool (MFA) (McAuliffe et al. 2017) model on those sentence-level pairs to obtain the phoneme-level alignments between song piece and its corresponding lyrics; 2) extract (fundamental frequency) as pitch information from the raw waveform using Parselmouth, following (Wu and Luan 2020; Blaauw and Bonada 2020; Ren et al. 2020). We randomly choose 2 songs for validation and testing. To release a high-quality dataset, after the paper is accepted, we clean and re-segment these songs, resulting in 1,651 song pieces, which mostly last 1013 seconds. The codes accompanied with the access to PopCS are in https://github.com/MoonInTheRiver/DiffSinger. In this repository, we also add the extra codes out of interest, for MIDI-to-Mel, including the MIDI-to-Mel without F0 prediction/condition.
We convert Chinese lyrics into phonemes by pypinyin following (Ren et al. 2020); and extract the mel-spectrogram (Shen et al. 2018) from the raw waveform; and set the hop size and frame size to 128 and 512 in respect of the sample rate 24kHz. The size of phoneme vocabulary is 61. The number of mel bins is 80. The mel-spectrograms are linearly scaled to the range , and F0 is normalized to have zero mean and unit variance. In the lyrics encoder, the dimension of phoneme embeddings is 256 and the Transformer blocks have the same setting as that in FastSpeech 2 (Ren et al. 2021). In the pitch encoder, the size of the lookup table and encoded pitch embedding are set to 300 and 256. The channel size mentioned before is set to 256. In the denoiser, the number of convolution layers N is 20 with the kernel size 3, and we set the dilation to 1 (without dilation) at each layerYou can consider setting a bigger dilation number to increase the receptive field of the denoiser. See our Github repository.. We set to and to constants increasing linearly from to . The auxiliary decoder has the same setting as the mel-spectrogram decoder in FastSpeech 2. In the boundary predictor, the number of convolutional layers is 5, and the threshold is set to 0.4 empirically.
The training has two stages: 1) warmup stage: separately train the auxiliary decoder for 160k steps with the music score encoder, and then leverage the auxiliary decoder to train the boundary predictor for 30k steps to obtain ; 2) main stage: training DiffSinger as Algorithm 1 describes for 160k steps until convergence. In the inference stage, for all SVS experiments, we uniformly use a pretrained Parallel WaveGAN (PWG) (Yamamoto, Song, and Kim 2020)We adjust PWG to take in F0 driven source excitation (Wang and Yamagishi 2020) as additional condition, similar to that in (Chen et al. 2020). as vocoder to transform the generated mel-spectrograms into waveforms (audio samples).
2 Main Results and Analysis
To evaluate the perceptual audio quality, we conduct the MOS (mean opinion score) evaluation on the test set. Eighteen qualified listeners are asked to make judgments about the synthesized song samples. We compare MOS of the song samples generated by DiffSinger with the following systems: 1) GT, the ground truth singing audio; 2) GT (Mel + PWG), where we first convert the ground truth singing audio to the ground truth mel-spectrograms, and then convert these mel-spectrograms back to audio using PWG vocoder described in Section 4.1; 3) FFT-NPSS (Blaauw and Bonada 2020) (WORLD), the SVS system which generates WORLD vocoder features (Morise, Yokomori, and Ozawa 2016) through feed-forward Transformer (FFT) and uses WORLD vocoder to synthesize audio; 4) FFT-Singer (Mel + PWG) the SVS system which generates mel-spectrograms through FFT network and uses PWG vocoder to synthesize audio; 5) GAN-Singer (Wu and Luan 2020) (Mel + PWG), the SVS system with adversarial training using multiple random window discriminators.
The results are shown in Table 1. The quality of GT (MEL + PWG) (4.04 0.11) is the upper limit of the acoustic model for SVS. DiffSinger outperforms the baseline system with simple training loss (FFT-Singer) by a large margin, and shows the superiority compared with the state-of-the-art GAN-based method (GAN-Singer (Wu and Luan 2020)), which demonstrate the effectiveness of our method.
As shown in Figure 5, we compare the ground truth, the generated mel-spectrograms from Diffsinger, GAN-singer and FFT-Singer with the same music score. It can be seen that both Figure 5(c) and Figure 5(b) contain more delicate details between harmonics than Figure 5(d) does. Moreover, the performance of Diffsinger in the region of mid or low frequency is more competitive than that of GAN-singer while maintaining similar quality of the high-frequency region.
In the meanwhile, the shallow diffusion mechanism accelerates inference of naive diffusion model by 45.1% (RTF 0.191 vs. 0.348, RTF is the real-time factor, that is the seconds it takes to generate one second of audio).
Ablation Studies
We conduct ablation studies to demonstrate the effectiveness of our proposed methods and some hyper-parameters studies to seek the best model configurations. We conduct CMOS evaluation for these experiments. The results of variations on DiffSinger are listed in Table 2. It can be seen that: 1) removing the shallow diffusion mechanism results in quality drop (-0.500 CMOS), which is consistent with the MOS test results and verifies the effectiveness of our shallow diffusion mechanism (row 1 vs. row 2); 2) adopting other (row 1 vs. row 3) rather than the one predicted by our boundary predictor causes quality drop, which verifies that our boundary prediction network can predict a proper for shallow diffusion mechanism; and 3) the model with configurations and produces the best results (row 1 vs. row 4,5,6,7), indicating that our model capacity is sufficient.
3 Extensional Experiments on TTS
To verify the generalization of our methods on TTS task, we conduct the extensional experiments on LJSpeech dataset (Ito and Johnson 2017), which contains 13,100 English audio clips (total 24 hours) with corresponding transcripts. We follow the train-val-test dataset splits, the pre-processing of mel-spectrograms, and the grapheme-to-phoneme tool in FastSpeech 2. To build DiffSpeech, we 1) add a pitch predictor and a duration predictor to DiffSinger as those in FastSpeech 2; 2) adopt for shallow diffusion mechanism.
We use Amazon Mechanical Turk (ten testers) to make subjective evaluation and the results are shown in Table 3. All the systems adopt HiFi-GAN (Kong, Kim, and Bae 2020) as vocoder. DiffSpeech outperforms FastSpeech 2 and Glow-TTS, which demonstrates the generalization. Besides, the last two rows in Table 3 also show the effectiveness of shallow diffusion mechanism (with 29.2% speedup, RTF 0.121 vs. 0.171).
Related Work
Initial works of singing voice synthesis generate the sounds using concatenated (Macon et al. 1997; Kenmochi and Ohshita 2007) or HMM-based parametric (Saino et al. 2006; Oura et al. 2010) methods, which are kind of cumbersome and lack flexibility and harmony. Thanks to the rapid evolution of deep learning, several SVS systems based on deep neural networks have been proposed in the past few years. Nishimura et al. (2016); Blaauw and Bonada (2017); Kim et al. (2018); Nakamura et al. (2019); Gu et al. (2020) utilize neural networks to map the contextual features to acoustic features. Ren et al. (2020) build the SVS system from scratch using singing data mined from music websites. Blaauw and Bonada (2020) propose a feed-forward Transformer SVS model for fast inference and avoiding exposure bias issues caused by autoregressive models. Besides, with the help of adversarial training, Lee et al. (2019) propose an end-to-end framework which directly generates linear-spectrograms. Wu and Luan (2020) present a multi-singer SVS system with limited available recordings and improve the voice quality by adding multiple random window discriminators. Chen et al. (2020) introduce multi-scale adversarial training to synthesize singing with a high sampling rate (48kHz). The voice naturalness and diversity of SVS system have been continuously improved in recent years.
2 Denoising Diffusion Probabilistic Models
A diffusion probabilistic model is a parameterized Markov chain trained by optimizing variational lower bound, which generates samples matching the data distribution in constant steps (Ho, Jain, and Abbeel 2020). Diffusion model is first proposed by Sohl-Dickstein et al. (2015). Ho, Jain, and Abbeel (2020) make progress of diffusion model to generate high-quality images using a certain parameterization and reveal an equivalence between diffusion model and denoising score matching (Song and Ermon 2019; Song et al. 2021). Recently, Kong et al. (2021) and Chen et al. (2021) apply the diffusion model to neural vocoders, which generate high-fidelity waveform conditioned on mel-spectrogram. Chen et al. (2021) also propose a continuous noise schedule to reduce the inference iterations while maintaining synthesis quality. Song, Meng, and Ermon (2021) extend diffusion model by providing a faster sampling mechanism, and a way to interpolate between samples meaningfully. Diffusion model is a fresh and developing technique, which has been applied in the fields of unconditional image generation, conditional spectrogram-to-waveform generation (neural vocoder). And in our work, we propose a diffusion model for the acoustic model which generates mel-spectrogram given music scores (or text). There is a concurrent work (Jeong et al. 2021) at the submission time of our preprint which adopts a diffusion model as the acoustic model for TTS task.
Conclusion
In this work, we proposed DiffSinger, an acoustic model for SVS based on diffusion probabilistic model. To improve the voice quality and speed up inference, we proposed a shallow diffusion mechanism. Specifically, we found that the diffusion trajectories of and converge together when the diffusion step is big enough. Inspired by this, we started the reverse process at the intersection (step ) of two trajectories rather than at the very deep diffusion step . Thus the burden of the reverse process could be distinctly alleviated, which improves the quality of synthesized audio and accelerates inference. The experiments conducted on PopCS demonstrate the superiority of DiffSinger compared with previous works, and the effectiveness of our novel shallow diffusion mechanism. The extensional experiments conducted on LJSpeech dataset prove the effectiveness of DiffSpeech on TTS task. The directly synthesis without vocoder will be future work.
Acknowledgments
This work was supported in part by the National Key R&D Program of China under Grant No.2020YFC0832505, No.62072397, Zhejiang Natural Science Foundation under Grant LR19F020006. Thanks participants of the listening test for the valuable evaluations.
References
Appendix A Theoretical Proof of Intersection
Given a data sample and its corresponding , the conditional distributions of and are:
respectively. The KL-divergence between two Gaussian distributions is:
where is the dimension; are means; are covariance matrices. Thus, in our case:
Since decreases towards 0 rapidly as increases, this KL-divergence also decreases towards 0 rapidly as increases. This guarantees the intersection of trajectories of the diffusion process.
Moreover, since the auxiliary decoder has been optimized by simple reconstruction loss (L1/L2 mentioned in the main paper) on the training set, is optimized towards minimum, which facilitates this intersection. In addition, does not need to be exactly the same as , but just needs to come from vicinity of the mode of (according to the theories of score matching and Langevin dynamics).
Appendix B An Easier Trick for Boundary Prediction
Intuitively, we can just adopt the smallest as when satisfies:
which means that this start point at step is at least not worse than the original prior distribution . mean the mel-spectrograms in the validation set. In addition, when using this trick to determine in DiffSpeech, it is more rational to generate () conditioned on the ground-truth F0-contour & duration rather than the ones predicted.
Appendix C Details of Model Structure and Supplementary Configurations
The detailed model structure of encoder, auxiliary decoder and denoiser are shown in Figure 6(a), Figure 6(b) and Figure 7 respectively.
C.2 Supplementary configurations
In each FFT block: the number of FFT layers (F in Figure 6(b)) is set to 4; the hidden size of self-attention layer is 256; the number of attention heads is 2; the kernel sizes of 1D-convolution in the 2-layer convolutional layers are set to 9 and 1.
Appendix D Model Size
The model footprints of main systems for comparison in our paper are shown in Table 4. It can be seen that DiffSinger has the similar learnable parameters as other state-of-the-art models.
Appendix E Details of Training and Inference
We train DiffSinger on 1 NVIDIA V100 GPU with 48 batch size. We adopt the Adam optimizer with learning rate . During training, the warmup stage costs about 16 hours and the main stage costs about 12 hours; During inference, the RTF of acoustic model for SVS and TTS are 0.191 and 0.121 respectively.