DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism

Jinglin Liu, Chengxi Li, Yi Ren, Feiyang Chen, Zhou Zhao

Introduction

Singing voice synthesis (SVS) which aims to synthesize natural and expressive singing voice from musical score (Wu and Luan 2020), increasingly draws attention from the research community and entertainment industries (Zhang et al. 2020). The pipeline of SVS usually consists of an acoustic model to generate the acoustic features (e.g., mel-spectrogram) conditioned on a music score, and a vocoder to convert the acoustic features to waveform (Nakamura et al. 2019; Lee et al. 2019; Blaauw and Bonada 2020; Ren et al. 2020; Chen et al. 2020)A music score consists of lyrics, pitch and duration..

Previous singing acoustic models mainly utilize simple loss (e.g., L1 or L2) to reconstruct the acoustic features. However, this optimization is based on the incorrect uni-modal distribution assumptions, leading to blurry and over-smoothing outputs. Although existing methods endeavor to solve this problem by generative adversarial network (GAN) (Lee et al. 2019; Chen et al. 2020), training an effective GAN may occasionally fail due to the unstable discriminator. These issues hinder the naturalness of synthesized singing.

Recently, a highly flexible and tractable generative model, diffusion probabilistic model (a.k.a. diffusion model) (Sohl-Dickstein et al. 2015; Ho, Jain, and Abbeel 2020; Song, Meng, and Ermon 2021) emerges. Diffusion model consists of two processes: diffusion process and reverse process (also called denoising process). The diffusion process is a Markov chain with fixed parameters (when using the certain parameterization in (Ho, Jain, and Abbeel 2020)), which converts the complicated data into isotropic Gaussian distribution by adding the Gaussian noise gradually; while the reverse process is a Markov chain implemented by a neural network, which learns to restore the origin data from Gaussian white noise iteratively. Diffusion model can be stably trained by implicitly optimizing variational lower bound (ELBO) on the data likelihood. It has been demonstrated that diffusion model can produce promising results in image generation (Ho, Jain, and Abbeel 2020; Song, Meng, and Ermon 2021) and neural vocoder (Chen et al. 2021; Kong et al. 2021) fields.

In this work, we propose DiffSinger, an acoustic model for SVS based on diffusion model, which converts the noise into mel-spectrogram conditioned on the music score. DiffSinger can be efficiently trained by optimizing ELBO, without adversarial feedback, and generates realistic mel-spectrograms strongly matching the ground truth distribution.

To further improve the voice quality and speed up inference, we introduce a shallow diffusion mechanism to make better use of the prior knowledge learned by the simple loss. Specifically, we find that there is an intersection of the diffusion trajectories of the ground-truth mel-spectrogram MM and the one predicted by a simple mel-spectrogram decoder M~\widetilde{M}Here we use a traditional acoustic model based on feed-forward Transformer (Ren et al. 2021; Blaauw and Bonada 2020), which is trained by L1 loss to reconstruct mel-spectrogram.: sending MM and M~\widetilde{M} into the diffusion process could result in similar distorted mel-spectrograms, when the diffusion step is big enough (but not reaches the deep step where the distorted mel-spectrograms become Gaussian white noise). Thus, in the inference stage we 1) leverage the simple mel-spectrogram decoder to generate M~\widetilde{M}; 2) calculate the sample at a shallow step kk through the diffusion process: M~k\widetilde{M}_{k}K<TK<T, where T is the total number of diffusion steps. M~k\widetilde{M}_{k} can be calculated in closed form time (Ho, Jain, and Abbeel 2020).; and 3) start reverse process from M~k\widetilde{M}_{k} rather than Gaussian white noise, and complete the process by kk iteration denoising steps (Vincent 2011; Song and Ermon 2019; Ho, Jain, and Abbeel 2020). Besides, we train a boundary prediction network to locate this intersection and determine the kk adaptively. The shallow diffusion mechanism provides a better start point than Gaussian white noise and alleviates the burden of the reverse process, which improves the quality of synthesized audio and accelerates inference.

Finally, since the pipeline of SVS resembles that of text-to-speech (TTS) task, we also build DiffSpeech adjusting from DiffSinger for generalization. The evaluations conducted on a Chinese singing dataset demonstrate the superiority of DiffSinger (0.11 MOS gains compared with a state-of-the-art acoustic model for SVS (Wu and Luan 2020)), and the effectiveness of our novel mechanism (0.14 MOS gains, 0.5 CMOS gains and 45.1% speedup with shallow diffusion mechanism). The extensional experiments of DiffSpeech on TTS task prove the generalization of our methods (0.24/0.23 MOS gains compared with FastSpeech 2 (Ren et al. 2021) and Glow-TTS (Kim et al. 2020) respectively). The contributions of this work can be summarized as follows:

We propose DiffSinger, which is the first acoustic model for SVS based on diffusion probabilistic model. DiffSinger addresses the over-smoothing and unstable training issues in previous works.

We propose a shallow diffusion mechanism to further improve the voice quality, and accelerate the inference.

The extensional experiments on TTS task (DiffSpeech) prove the generalization of our methods.

Diffusion Model

In this section, we introduce the theory of diffusion probabilistic model (Sohl-Dickstein et al. 2015; Ho, Jain, and Abbeel 2020). The full proof can be found in previous works (Ho, Jain, and Abbeel 2020; Kong et al. 2021; Song, Meng, and Ermon 2021). A diffusion probabilistic model converts the raw data into Gaussian distribution gradually by a diffusion process, and then learns the reverse process to restore the data from Gaussian white noise (Sohl-Dickstein et al. 2015). These processes are shown in Figure 1.

Define the data distribution as q(y0)q(\mathbf{y}_{0}), and sample y0∼q(y0)\mathbf{y}_{0}\sim q(\mathbf{y}_{0}). The diffusion process is a Markov chain with fixed parameters (Ho, Jain, and Abbeel 2020), which converts y0\mathbf{y}_{0} into the latent yT\mathbf{y}_{T} in TT steps:

At each diffusion step t∈[1,T]t\in[1,T], a tiny Gaussian noise is added to yt−1\mathbf{y}_{t-1} to obtain yt\mathbf{y}_{t}, according to a variance schedule β={β1,…,βT}\beta=\{\beta_{1},\dotsc,\beta_{T}\}:

If β\beta is well designed and TT is sufficiently large, then q(yT)q(y_{T}) is nearly an isotropic Gaussian distribution (Ho, Jain, and Abbeel 2020; Nichol and Dhariwal 2021). Besides, there is a special property of diffusion process that q(yt∣y0)q(\mathbf{y}_{t}|\mathbf{y}_{0}) can be calculated in closed form in O(1)O(1) time (Ho, Jain, and Abbeel 2020):

where αˉt≔∏s=1tαs\bar{\alpha}_{t}\coloneqq\prod_{s=1}^{t}\alpha_{s}, αt≔1−βt\alpha_{t}\coloneqq 1-\beta_{t}.

The reverse process is a Markov chain with learnable parameters θ\theta from yT\mathbf{y}_{T} to y0\mathbf{y}_{0}. Since the exact reverse transition distribution q(yt−1∣yt)q(\mathbf{y}_{t-1}|\mathbf{y}_{t}) is intractable, we approximate it by a neural network with parameters θ\theta (θ\theta is shared at every tt-th step):

Thus the whole reverse process can be defined as:

To learn the parameters θ\theta, we minimize a variational bound of the negative log likelihood:

where C\mathcal{C} is a constant. And by reparameterizing Eq. (1) as yt(y0,ϵ)=αˉty0+1−αˉtϵ\mathbf{y}_{t}(\mathbf{y}_{0},{\boldsymbol{\epsilon}})=\sqrt{\bar{\alpha}_{t}}\mathbf{y}_{0}+\sqrt{1-\bar{\alpha}_{t}}{\boldsymbol{\epsilon}}, and choosing the parameterization:

Sample yT\mathbf{y}_{T} from p(yT)∼N(0,I)p(\mathbf{y}_{T})\sim\mathcal{N}(\mathbf{0},\mathbf{I}) and run the reverse process to obtain a data sample.

DiffSinger

As illustrated in Figure 2, DiffSinger is built on the diffusion model. Since SVS task models the conditional distribution pθ(M0∣x)p_{\theta}(M_{0}|x), where MM is the mel-spectrogram and xx is the music score corresponding to MM, we add xx to the diffusion denoiser as the condition in the reverse process. In this section, we first describe a naive version of DiffSinger (Section 3.1); then we introduce a novel shallow diffusion mechanism to improve the model performance and efficiency (Section 3.2); finally, we describe the boundary prediction network which can adaptively find the intersection boundary required in shallow diffusion mechanism (Section 3.3).

In the naive version of DiffSinger (without dotted line boxes in Figure 2): In the training procedure (shown in Figure 2(a)), DiffSinger takes in the mel-spectrogram at tt-th step MtM_{t} in the diffusion process and predicts the random noise ϵθ(⋅){\boldsymbol{\epsilon}}_{\theta}(\cdot) in Eq. (6), conditioned on tt and the music score xx. The inference procedure (shown in Figure 2(b)) starts at the Gaussian white noise sampled from N(0,I)\mathcal{N}(\mathbf{0},\mathbf{I}), as the previous diffusion models do (Ho, Jain, and Abbeel 2020; Kong et al. 2021). Then the procedure iterates for TT times to repeatedly denoise the intermediate samples with two steps: 1) predict the ϵθ(⋅){\boldsymbol{\epsilon}}_{\theta}(\cdot) using the denoiser; 2) obtain Mt−1M_{t-1} from MtM_{t} using the predicted ϵθ(⋅){\boldsymbol{\epsilon}}_{\theta}(\cdot), according to Eq. (2) and Eq. (5):

where z∼N(0,I)z\sim\mathcal{N}(\mathbf{0},\mathbf{I}) when t>1t>1, and z=0z=0 when t=1t=1. Finally, a mel-spectrogram M\mathcal{M} corresponding to xx could be generated.

2 Shallow Diffusion Mechanism

Although the previous acoustic model trained by the simple loss has intractable drawbacks, it still generates samples showing strong connectionThe samples fail to maintain the variable aperiodic parameters, but they usually have a clear “skeleton” (harmonics) matching the ground truth. to the ground-truth data distribution, which could provide plenty of prior knowledge to DiffSinger. To explore this connection and find a way to make better use of the prior knowledge, we conduct the empirical observation leveraging the diffusion process (shown in Figure 3): 1) when t=0t=0, MM has rich details between the neighboring harmonics, which can influence the naturalness of the synthesized singing voice, but M~\widetilde{M} is over-smoothing as we introduced in Section 1; 2) as tt increases, samples of two process become indistinguishable. We illustrate this observation in Figure 4: the trajectory from M~\widetilde{M} manifold to Gaussian noise manifold and the trajectory from MM to Gaussian noise manifold intersect when the diffusion step is big enough.

Inspired by this observation, we propose the shallow diffusion mechanism: instead of starting with the Gaussian white noise, the reverse process starts at the intersection of two trajectories shown in Figure 4. Thus the burden of the reverse process could be distinctly alleviatedConverting MkM_{k} into M0M_{0} is easier than converting MTM_{T} (Gaussion white noise) into M0M_{0} (k<Tk<T). Thus the former could improve the quality of synthesized audio and accelerates inference.. Specifically, in the inference stage we 1) leverage an auxiliary decoder to generate M~\widetilde{M}, which is trained with L1 conditioned on the music score encoder outputs, as shown in the dotted line box in Figure 2(a); 2) generate the intermediate sample at a shallow step kk through the diffusion process, as shown in the dotted line box in Figure 2(b) according to Eq. (1):

where ϵ∼N(0,I){\boldsymbol{\epsilon}}\sim\mathcal{N}(\mathbf{0},\mathbf{I}), αˉk≔∏s=1kαs\bar{\alpha}_{k}\coloneqq\prod_{s=1}^{k}\alpha_{s}, αk≔1−βk\alpha_{k}\coloneqq 1-\beta_{k}. If the intersection boundary kk is properly chosen, it can be considered that M~k\widetilde{M}_{k} and MkM_{k} come from the same distribution; 3) start reverse process from M~k\widetilde{M}_{k}, and complete the process by kk iterations denoising.

The training and inference procedures with shallow diffusion mechanism are described in Algorithm 1 and 2 respectively. The theoretical proof of the intersection of two trajectories can be found in the supplement.

3 Boundary Prediction

where Y\mathcal{Y} is the training set of mel-spectrograms. When BP have been trained, we determine kk using the predicted value of BP, which indicates the probability of a sample classified to be 1. For all M∈YM\in\mathcal{Y}, we find the earliest step k′k^{\prime} where the 95% steps tt in [k′,T][k^{\prime},T] satisfies: the margin between BP(MtM_{t}, tt) and BP(M~t\widetilde{M}_{t}, tt) is under the threshold. Then we choose the average of k′k^{\prime} as the intersection boundary kk.

We also propose an easier trick for boundary prediction in the supplement by comparing the KL-divergence. Note that the boundary prediction can be considered as a step of dataset preprocessing to choose the hyperparameter kk for the whole dataset. kk actually can be chosen manually by brute-force searching on validation set.

4 Model Structures

The encoder encodes the music score into the condition sequence, which consists of 1) a lyrics encoder to map the phoneme ID into embedding sequence, and a series of Transformer blocks (Vaswani et al. 2017) to convert this sequence into linguistic hidden sequence; 2) a length regulator to expand the linguistic hidden sequence to the length of mel-spectrograms according to the duration information; and 3) a pitch encoder to map the pitch ID into pitch embedding sequence. Finally, the encoder adds linguistic sequence and pitch sequence together as the music condition sequence EmE_{m} following (Ren et al. 2020).

The diffusion step tt is another conditional input for denoiser ϵθ{\boldsymbol{\epsilon}}_{\theta}, as shown in Eq. (6). To convert the discrete step tt to continuous hidden, we use the sinusoidal position embedding (Vaswani et al. 2017) followed by two linear layers to obtain step embedding EtE_{t} with CC channels.

We introduce a simple mel-spectrogram decoder called the auxiliary decoder, which is composed of stacked feed-forward Transformer (FFT) blocks and generates M~\widetilde{M} as the final outputs, the same as the mel-spectrogram decoder in FastSpeech 2 (Ren et al. 2021).

Denoiser ϵθ{\boldsymbol{\epsilon}}_{\theta} takes in MtM_{t} as input to predict ϵ{\boldsymbol{\epsilon}} added in diffusion process conditioned on the step embedding EtE_{t} and music condition sequence EmE_{m}. Since diffusion model imposes no architectural constraints (Sohl-Dickstein et al. 2015; Kong et al. 2021), the design of denoiser has multiple choices. We adopt a non-causal WaveNet (Oord et al. 2016) architecture proposed by (Rethage, Pons, and Serra 2018; Kong et al. 2021) as our denoiser. The denoiser is composed of a 1×11\times 1 convolution layer to project MtM_{t} with HmH_{m} channels to the input hidden sequence H\mathcal{H} with CC channels and NN convolution blocks with residual connections. Each convolution block consists of 1) an element-wise adding operation which adds EtE_{t} to H\mathcal{H}; 2) a non-causal convolution network which converts H\mathcal{H} from CC to 2C2C channels; 3) a 1×11\times 1 convolution layer which converts the EmE_{m} to 2C2C channels; 4) a gate unit to merge the information of input and conditions; and 5) a residual block to split the merged hidden into two branches with CC channels (the residual as the following H\mathcal{H} and the “skip hidden” to be collected as the final results), which enables the denoiser to incorporate features at several hierarchical levels for final prediction.

The classifier in the boundary predictor is composed of 1) a step embedding to provide EtE_{t}; 2) a ResNet (He et al. 2016) with stacked convolutional layers and a linear layer, which takes in the mel-spectrograms at tt-th step and EtE_{t} to classify MtM_{t} and M~t\widetilde{M}_{t}.

More details of model structure and configurations are shown in the supplement.

Experiments

In this section, we first describe the experimental setup, and then provide the main results on SVS with analysis. Finally, we conduct the extensional experiments on TTS.

Since there is no publicly available high-quality unaccompanied singing dataset, we collect and annotate a Chinese Mandarin pop songs dataset: PopCS, to evaluate our methods. PopCS contains 117 Chinese pop songs (total ∼\sim5.89 hours with lyrics) collected from a qualified female vocalist. All the audio files are recorded in a recording studio. Every song is sampled at 24kHz with 16-bit quantization. To obtain more accurate music scores corresponding to the songs (Lee et al. 2019), we 1) split each whole song into sentence pieces following DeepSinger (Ren et al. 2020) and train a Montreal Forced Aligner tool (MFA) (McAuliffe et al. 2017) model on those sentence-level pairs to obtain the phoneme-level alignments between song piece and its corresponding lyrics; 2) extract F0F_{0} (fundamental frequency) as pitch information from the raw waveform using Parselmouth, following (Wu and Luan 2020; Blaauw and Bonada 2020; Ren et al. 2020). We randomly choose 2 songs for validation and testing. To release a high-quality dataset, after the paper is accepted, we clean and re-segment these songs, resulting in 1,651 song pieces, which mostly last 10∼\sim13 seconds. The codes accompanied with the access to PopCS are in https://github.com/MoonInTheRiver/DiffSinger. In this repository, we also add the extra codes out of interest, for MIDI-to-Mel, including the MIDI-to-Mel without F0 prediction/condition.

We convert Chinese lyrics into phonemes by pypinyin following (Ren et al. 2020); and extract the mel-spectrogram (Shen et al. 2018) from the raw waveform; and set the hop size and frame size to 128 and 512 in respect of the sample rate 24kHz. The size of phoneme vocabulary is 61. The number of mel bins HmH_{m} is 80. The mel-spectrograms are linearly scaled to the range , and F0 is normalized to have zero mean and unit variance. In the lyrics encoder, the dimension of phoneme embeddings is 256 and the Transformer blocks have the same setting as that in FastSpeech 2 (Ren et al. 2021). In the pitch encoder, the size of the lookup table and encoded pitch embedding are set to 300 and 256. The channel size CC mentioned before is set to 256. In the denoiser, the number of convolution layers N is 20 with the kernel size 3, and we set the dilation to 1 (without dilation) at each layerYou can consider setting a bigger dilation number to increase the receptive field of the denoiser. See our Github repository.. We set TT to 100100 and β\beta to constants increasing linearly from β1=10−4\beta_{1}=10^{-4} to βT=0.06\beta_{T}=0.06. The auxiliary decoder has the same setting as the mel-spectrogram decoder in FastSpeech 2. In the boundary predictor, the number of convolutional layers is 5, and the threshold is set to 0.4 empirically.

The training has two stages: 1) warmup stage: separately train the auxiliary decoder for 160k steps with the music score encoder, and then leverage the auxiliary decoder to train the boundary predictor for 30k steps to obtain kk; 2) main stage: training DiffSinger as Algorithm 1 describes for 160k steps until convergence. In the inference stage, for all SVS experiments, we uniformly use a pretrained Parallel WaveGAN (PWG) (Yamamoto, Song, and Kim 2020)We adjust PWG to take in F0 driven source excitation (Wang and Yamagishi 2020) as additional condition, similar to that in (Chen et al. 2020). as vocoder to transform the generated mel-spectrograms into waveforms (audio samples).

2 Main Results and Analysis

To evaluate the perceptual audio quality, we conduct the MOS (mean opinion score) evaluation on the test set. Eighteen qualified listeners are asked to make judgments about the synthesized song samples. We compare MOS of the song samples generated by DiffSinger with the following systems: 1) GT, the ground truth singing audio; 2) GT (Mel + PWG), where we first convert the ground truth singing audio to the ground truth mel-spectrograms, and then convert these mel-spectrograms back to audio using PWG vocoder described in Section 4.1; 3) FFT-NPSS (Blaauw and Bonada 2020) (WORLD), the SVS system which generates WORLD vocoder features (Morise, Yokomori, and Ozawa 2016) through feed-forward Transformer (FFT) and uses WORLD vocoder to synthesize audio; 4) FFT-Singer (Mel + PWG) the SVS system which generates mel-spectrograms through FFT network and uses PWG vocoder to synthesize audio; 5) GAN-Singer (Wu and Luan 2020) (Mel + PWG), the SVS system with adversarial training using multiple random window discriminators.

The results are shown in Table 1. The quality of GT (MEL + PWG) (4.04 ±\pm 0.11) is the upper limit of the acoustic model for SVS. DiffSinger outperforms the baseline system with simple training loss (FFT-Singer) by a large margin, and shows the superiority compared with the state-of-the-art GAN-based method (GAN-Singer (Wu and Luan 2020)), which demonstrate the effectiveness of our method.

As shown in Figure 5, we compare the ground truth, the generated mel-spectrograms from Diffsinger, GAN-singer and FFT-Singer with the same music score. It can be seen that both Figure 5(c) and Figure 5(b) contain more delicate details between harmonics than Figure 5(d) does. Moreover, the performance of Diffsinger in the region of mid or low frequency is more competitive than that of GAN-singer while maintaining similar quality of the high-frequency region.

In the meanwhile, the shallow diffusion mechanism accelerates inference of naive diffusion model by 45.1% (RTF 0.191 vs. 0.348, RTF is the real-time factor, that is the seconds it takes to generate one second of audio).

Ablation Studies

We conduct ablation studies to demonstrate the effectiveness of our proposed methods and some hyper-parameters studies to seek the best model configurations. We conduct CMOS evaluation for these experiments. The results of variations on DiffSinger are listed in Table 2. It can be seen that: 1) removing the shallow diffusion mechanism results in quality drop (-0.500 CMOS), which is consistent with the MOS test results and verifies the effectiveness of our shallow diffusion mechanism (row 1 vs. row 2); 2) adopting other kk (row 1 vs. row 3) rather than the one predicted by our boundary predictor causes quality drop, which verifies that our boundary prediction network can predict a proper kk for shallow diffusion mechanism; and 3) the model with configurations C=256C=256 and L=20L=20 produces the best results (row 1 vs. row 4,5,6,7), indicating that our model capacity is sufficient.

3 Extensional Experiments on TTS

To verify the generalization of our methods on TTS task, we conduct the extensional experiments on LJSpeech dataset (Ito and Johnson 2017), which contains 13,100 English audio clips (total ∼\sim24 hours) with corresponding transcripts. We follow the train-val-test dataset splits, the pre-processing of mel-spectrograms, and the grapheme-to-phoneme tool in FastSpeech 2. To build DiffSpeech, we 1) add a pitch predictor and a duration predictor to DiffSinger as those in FastSpeech 2; 2) adopt k=70k=70 for shallow diffusion mechanism.

We use Amazon Mechanical Turk (ten testers) to make subjective evaluation and the results are shown in Table 3. All the systems adopt HiFi-GAN (Kong, Kim, and Bae 2020) as vocoder. DiffSpeech outperforms FastSpeech 2 and Glow-TTS, which demonstrates the generalization. Besides, the last two rows in Table 3 also show the effectiveness of shallow diffusion mechanism (with 29.2% speedup, RTF 0.121 vs. 0.171).

Related Work

Initial works of singing voice synthesis generate the sounds using concatenated (Macon et al. 1997; Kenmochi and Ohshita 2007) or HMM-based parametric (Saino et al. 2006; Oura et al. 2010) methods, which are kind of cumbersome and lack flexibility and harmony. Thanks to the rapid evolution of deep learning, several SVS systems based on deep neural networks have been proposed in the past few years. Nishimura et al. (2016); Blaauw and Bonada (2017); Kim et al. (2018); Nakamura et al. (2019); Gu et al. (2020) utilize neural networks to map the contextual features to acoustic features. Ren et al. (2020) build the SVS system from scratch using singing data mined from music websites. Blaauw and Bonada (2020) propose a feed-forward Transformer SVS model for fast inference and avoiding exposure bias issues caused by autoregressive models. Besides, with the help of adversarial training, Lee et al. (2019) propose an end-to-end framework which directly generates linear-spectrograms. Wu and Luan (2020) present a multi-singer SVS system with limited available recordings and improve the voice quality by adding multiple random window discriminators. Chen et al. (2020) introduce multi-scale adversarial training to synthesize singing with a high sampling rate (48kHz). The voice naturalness and diversity of SVS system have been continuously improved in recent years.

2 Denoising Diffusion Probabilistic Models

A diffusion probabilistic model is a parameterized Markov chain trained by optimizing variational lower bound, which generates samples matching the data distribution in constant steps (Ho, Jain, and Abbeel 2020). Diffusion model is first proposed by Sohl-Dickstein et al. (2015). Ho, Jain, and Abbeel (2020) make progress of diffusion model to generate high-quality images using a certain parameterization and reveal an equivalence between diffusion model and denoising score matching (Song and Ermon 2019; Song et al. 2021). Recently, Kong et al. (2021) and Chen et al. (2021) apply the diffusion model to neural vocoders, which generate high-fidelity waveform conditioned on mel-spectrogram. Chen et al. (2021) also propose a continuous noise schedule to reduce the inference iterations while maintaining synthesis quality. Song, Meng, and Ermon (2021) extend diffusion model by providing a faster sampling mechanism, and a way to interpolate between samples meaningfully. Diffusion model is a fresh and developing technique, which has been applied in the fields of unconditional image generation, conditional spectrogram-to-waveform generation (neural vocoder). And in our work, we propose a diffusion model for the acoustic model which generates mel-spectrogram given music scores (or text). There is a concurrent work (Jeong et al. 2021) at the submission time of our preprint which adopts a diffusion model as the acoustic model for TTS task.

Conclusion

In this work, we proposed DiffSinger, an acoustic model for SVS based on diffusion probabilistic model. To improve the voice quality and speed up inference, we proposed a shallow diffusion mechanism. Specifically, we found that the diffusion trajectories of MM and M~\widetilde{M} converge together when the diffusion step is big enough. Inspired by this, we started the reverse process at the intersection (step kk) of two trajectories rather than at the very deep diffusion step TT. Thus the burden of the reverse process could be distinctly alleviated, which improves the quality of synthesized audio and accelerates inference. The experiments conducted on PopCS demonstrate the superiority of DiffSinger compared with previous works, and the effectiveness of our novel shallow diffusion mechanism. The extensional experiments conducted on LJSpeech dataset prove the effectiveness of DiffSpeech on TTS task. The directly synthesis without vocoder will be future work.

Acknowledgments

This work was supported in part by the National Key R&D Program of China under Grant No.2020YFC0832505, No.62072397, Zhejiang Natural Science Foundation under Grant LR19F020006. Thanks participants of the listening test for the valuable evaluations.

References

Appendix A Theoretical Proof of Intersection

Given a data sample M0M_{0} and its corresponding M~0\widetilde{M}_{0}, the conditional distributions of MtM_{t} and M~t\widetilde{M}_{t} are:

respectively. The KL-divergence between two Gaussian distributions is:

where kk is the dimension; μ0,μ1\mu_{0},\mu_{1} are means; Σ0,Σ1\Sigma_{0},\Sigma_{1} are covariance matrices. Thus, in our case:

Since αˉt2(1−αˉt)\frac{\bar{\alpha}_{t}}{2(1-\bar{\alpha}_{t})} decreases towards 0 rapidly as tt increases, this KL-divergence also decreases towards 0 rapidly as tt increases. This guarantees the intersection of trajectories of the diffusion process.

Moreover, since the auxiliary decoder has been optimized by simple reconstruction loss (L1/L2 mentioned in the main paper) on the training set, ∥M~0−M0∥22\|\widetilde{M}_{0}-M_{0}\|_{2}^{2} is optimized towards minimum, which facilitates this intersection. In addition, M~k\widetilde{M}_{k} does not need to be exactly the same as MkM_{k}, but just needs to come from vicinity of the mode of q(Mk∣M0)q(M_{k}|M_{0}) (according to the theories of score matching and Langevin dynamics).

Appendix B An Easier Trick for Boundary Prediction

Intuitively, we can just adopt the smallest tt as kk when tt satisfies:

which means that this start point at step kk is at least not worse than the original prior distribution N(0,I)\mathcal{N}(\mathbf{0},\mathbf{I}). Y′\mathcal{Y^{\prime}} mean the mel-spectrograms in the validation set. In addition, when using this trick to determine kk in DiffSpeech, it is more rational to generate M~t\widetilde{M}_{t} (∈Y′\in\mathcal{Y^{\prime}}) conditioned on the ground-truth F0-contour & duration rather than the ones predicted.

Appendix C Details of Model Structure and Supplementary Configurations

The detailed model structure of encoder, auxiliary decoder and denoiser are shown in Figure 6(a), Figure 6(b) and Figure 7 respectively.

C.2 Supplementary configurations

In each FFT block: the number of FFT layers (F in Figure 6(b)) is set to 4; the hidden size of self-attention layer is 256; the number of attention heads is 2; the kernel sizes of 1D-convolution in the 2-layer convolutional layers are set to 9 and 1.

Appendix D Model Size

The model footprints of main systems for comparison in our paper are shown in Table 4. It can be seen that DiffSinger has the similar learnable parameters as other state-of-the-art models.

Appendix E Details of Training and Inference

We train DiffSinger on 1 NVIDIA V100 GPU with 48 batch size. We adopt the Adam optimizer with learning rate lr=10−3lr=10^{-3}. During training, the warmup stage costs about 16 hours and the main stage costs about 12 hours; During inference, the RTF of acoustic model for SVS and TTS are 0.191 and 0.121 respectively.