BDDM: Bilateral Denoising Diffusion Models for Fast and High-Quality Speech Synthesis

Max W. Y. Lam, Jun Wang, Dan Su, Dong Yu

Introduction

Deep generative models have shown a tremendous advancement in speech synthesis (van den Oord et al. 2016; Kalchbrenner et al. 2018; Prenger et al. 2019; Kumar et al. 2019a; Kong et al. 2020b; Chen et al. 2020; Kong et al. 2021). Successful generative models can be mainly divided into two categories: generative adversarial network (GAN) (Goodfellow et al. 2014) based and likelihood-based. The former is based on adversarial learning, where the objective is to generate data indistinguishable from the training data. Yet, the training GANs can be very unstable, and the relevant training objectives are not suitable to compare against different GANs. The latter uses log-likelihood or surrogate objectives for training, but they also have intrinsic limitations regarding generation speed or quality. For example, the autoregressive models (van den Oord et al. 2016; Kalchbrenner et al. 2018), while being capable of generating high-fidelity data, are limited by their inherently slow sampling process and the poor scaling properties on high-dimensional data. Likewise, the flow-based models (Dinh et al. 2016; Kingma & Dhariwal 2018; Chen et al. 2018; Papamakarios et al. 2021) rely on specialized architectures to build a normalized probability model, whose training is less parameter-efficient. Other prior works use surrogate objectives, such as the evidence lower bound in variational auto-encoders (Kingma & Welling 2013; Rezende et al. 2014; Maaløe et al. 2019) and the contrastive divergence in energy-based models (Hinton 2002; Carreira-Perpinan & Hinton 2005). These models, despite showing improved speed, typically only work well for low-dimensional data, and, in general, the sample qualities are not competitive to the GAN-based and the autoregressive models (Bond-Taylor et al. 2021).

An up-and-coming class of likelihood-based models is the diffusion probabilistic models (DPMs) (Sohl-Dickstein et al. 2015), which introduces the idea of using a forward diffusion process to sequentially corrupt a given distribution and learning the reversal of such diffusion process to restore the data distribution for sampling. From a similar perspective, Song & Ermon 2019 proposed the score-based generative models by applying the score matching technique (Hyvarinen & Dayan 2005) to train a neural network such that samples can be generated via Langevin dynamics. Along these two lines of research, Ho et al. 2020 proposed the denoising diffusion probabilistic models (DDPMs) for high-quality image syntheses. Dhariwal & Nichol 2021 demonstrated that improved DDPMs Nichol & Dhariwal 2021 are capable of generating high-quality images of comparable or even superior quality to the state-of-the-art (SOTA) GAN-based models. For speech syntheses, DDPMs were also applied in Wavegrad (Chen et al. 2020) and DiffWave (Kong et al. 2021) to produce higher-fidelity audio samples than the conventional non-autoregressive models (Yamamoto et al. 2020; Kumar et al. 2019b; Yang et al. 2020; Bińkowski et al. 2020) and matched the quality of the SOTA autoregressive methods (Chen et al. 2020).

Despite the compelling results, the diffusion generative models are two to three orders of magnitude slower than other generative models such as GANs and VAEs. Their primary limitation is that they require up to thousands of diffusion steps during training to learn the target distribution. Therefore a large number of reverse steps are often required at sampling time. Recently, extensive investigations have been conducted to reduce the sampling steps for efficiently generating high-quality samples, which we will discuss in the related work in Section 2. Distinctively, we conceived that we might train a neural network to efficiently and adaptively estimate a much shorter noise schedule for sampling while achieving generation performances comparable or superior to the conventional DPMs. With such an incentive, after introducing the conventional DPMs as our background in Section 3, we propose in Section 4 bilateral denoising diffusion models (BDDMs), named after a bilateral modeling perspective – parameterizing the forward and reverse processes with a schedule network and a score network, respectively. We theoretically derive that the schedule network should be trained after the score network is optimized. For training the schedule network, we propose a novel objective to minimize the gap between a newly derived lower bound and the log marginal likelihood. We describe the training algorithm as well as the fast and high-quality sampling algorithm in Section 5. The training of the schedule network converges very fast using our newly derived objective, and its training only adds negligible overhead to DDPM’s. In Section 6, our neural vocoding experiments demonstrated that BDDMs could generate high-fidelity samples with as few as three sampling steps. Moreover, our method can produce speech samples indistinguishable from human speech with only seven sampling steps (143x faster than WaveGrad and 28.6x faster than DiffWave).

Related Work

Another class of noise scheduling methods searches for a subsequence of time indices of the training noise schedule, which we call the time schedule. DDIMs (Song et al. 2021a) introduced an accelerated reverse process that relies on a pre-specified time schedule. A linear and a quadratic time schedule were used in DDIMs and showed superior generation quality over DDPMs within 10 to 100 sampling steps. Nichol & Dhariwal 2021 proposed a re-scaled noise schedule for fast sampling, but this also requires pre-specifying the time schedule and the training noise schedule. Nichol & Dhariwal 2021 also proposed learning variances for the reverse processes, whereas the variances of the forward processes, i.e., the noise schedule, which affected both the means and variances of the reverse processes, were not learnable. According to the results of (Song et al. 2021a; Nichol & Dhariwal 2021), using a linear or quadratic time schedule resulted in quite different performances in different datasets, implying that the optimal choice of schedule varies with the datasets. So, there remains a challenge in finding a short and effective schedule for fast sampling on different datasets. Notably, Kong & Ping 2021 proposed a method to map a noise schedule to a time schedule for fast sampling. In this sense, searching for a time schedule becomes a sub-set of the noise scheduling problem, which resembles the above category of methods.

Although DPMs (Sohl-Dickstein et al. 2015) and DDPMs (Ho et al. 2019) mentioned that the noise schedule could be learned by re-parameterization, the approach was not investigated in their works. Closely related works that learn a noise schedule emerged until very recently. San-Roman et al. 2021 proposed a noise estimation (NE) method, which trained a neural net with a regression loss to estimate the noise scale from the noisy sample at each time point, and then predicted the next noise scale. However, NE requires a prior assumption of the noise schedule following a linear or Fibonacci rule. Most recently, a concurrent work to ours by Kingma et al. 2021 jointly trained a neural net to predict the signal-to-noise ratio (SNR) by maximizing the variational lower bound. The SNR was then used for noise scheduling. Different from ours, this scheduling neural net only took tt as input and is independent of the noisy sample generated during the loop of sampling process. Intrinsically, with limited information about the sampled data, the predicted SNR could deviate from the actual SNR of the noisy data during sampling.

Background

2 Denoising Diffusion Probabilistic Models (DDPMs)

Based on the nice property of isotropic Gaussians, one can express xt{\bm{x}}_{t} directly conditioned on x0{\bm{x}}_{0}:

To revert this forward process, DDPMs employ a score network Here, ϵθ(xt,αt){\bm{\epsilon}}_{\theta}({\bm{x}}_{t},\alpha_{t}) is conditioned on the continuous noise scale αt\alpha_{t}, as in (Song et al. 2021b; Chen et al. 2020). Alternatively, the score network can also be conditioned on a discrete time index ϵθ(xt,t){\bm{\epsilon}}_{\theta}({\bm{x}}_{t},t), as in (Song et al. 2021a; Ho et al. 2020). An approximate mapping of a noise schedule to a time schedule (Kong & Ping 2021) exists, therefore we consider conditioning on noise scales as the general case. ϵθ(xt,αt){\bm{\epsilon}}_{\theta}({\bm{x}}_{t},\alpha_{t}) to define

Note that the calculation of the complete ELBO in Eq. (1) requires TT forward passes of the score network, which would make the training computationally prohibitive for a large TT. To feasibly train the score network, instead of computing the complete ELBO, Ho et al. 2020 proposed an efficient training mechanism by sampling from a discrete uniform distribution: t∼U{1,...,T}t\sim\mathcal{U}\{1,...,T\}, x0∼pdata(x0){\bm{x}}_{0}\sim p_{\text{data}}({\bm{x}}_{0}), ϵt∼N(0,I){\bm{\epsilon}}_{t}\sim{\mathcal{N}}({\bm{0}},{\bm{I}}) at each training iteration to compute the training loss:

Bilateral denoising diffusion models (BDDMs)

For fast sampling with DPMs, we strive for a noise schedule β^\hat{{\bm{\beta}}} for sampling that is much shorter than the noise schedule β{\bm{\beta}} for training. As shown in Fig. 1, we define two separate diffusion processes corresponding to the noise schedules, β{\bm{\beta}} and β^\hat{{\bm{\beta}}}, respectively. The upper diffusion process parameterized by β{\bm{\beta}} is the same as in Eq. (2), whereas the lower process is defined as qβ^(x^1:N∣x^0)=∏n=1Nqβ^n(x^n∣x^n−1)q_{\hat{{\bm{\beta}}}}(\hat{{\bm{x}}}_{1:N}|\hat{{\bm{x}}}_{0})=\prod_{n=1}^{N}q_{\hat{\beta}_{n}}(\hat{{\bm{x}}}_{n}|\hat{{\bm{x}}}_{n-1}) with much fewer diffusion steps (N≪TN\ll T). In our problem formulation, β{\bm{\beta}} is given, but β^\hat{{\bm{\beta}}} is unknown. The goal is to find a β^\hat{{\bm{\beta}}} for the reverse process pθ(x^n−1∣x^n;β^n)p_{\theta}(\hat{{\bm{x}}}_{n-1}|\hat{{\bm{x}}}_{n};\hat{\beta}_{n}) such that x^0\hat{{\bm{x}}}_{0} can be effectively recovered from x^N\hat{{\bm{x}}}_{N} with NN reverse steps.

2 Model description

Although many prior arts (Ho et al. 2020; Chen et al. 2020; Song et al. 2021a; San-Roman et al. 2021) directly applied a shortened linear or Fibonacci noise schedule to the reverse process, we argue that these are sub-optimal solutions. Theoretically, the diffusion process specified by a new shortened noise schedule is essentially different from the one used to train the score network θ\theta. Therefore, θ\theta is not guaranteed suitable for reverting the shortened diffusion process. This issue motivated a novel modeling perspective to establish a link between the shortened schedule β^{\hat{\bm{\beta}}} and the score network θ\theta, i.e., to have β^{\hat{\bm{\beta}}} optimized according to θ\theta.

As a starting point, we consider an N=⌊T/τ⌋N=\lfloor T/\tau\rfloor, where 1≤τ<T1\leq\tau<T is a hyperparameter controlling the step size such that each diffusion step between two consecutive variables in the shorter diffusion process corresponds to τ\tau diffusion steps in the longer one. Based on Eq. (2), we define the following:

where xt{\bm{x}}_{t} is an intermediate diffused variable we introduced to link the two differently indexed diffusion sequences. We call it a junctional variable, which can be easily generated given x0{\bm{x}}_{0} and β{\bm{\beta}} during training: xt=αtx0+1−αt2ϵn{\bm{x}}_{t}=\alpha_{t}{\bm{x}}_{0}+\sqrt{1-\alpha_{t}^{2}}{\bm{\epsilon}}_{n}.

Unfortunately, for the reverse process when x0{\bm{x}}_{0} is not given, the junctional variable is intractable. However, our key observation is that while using the score by a score network θ∗\theta^{*} trained for the long β{\bm{\beta}}-parameterized diffusion process, a short noise schedule β^(ϕ)\hat{{\bm{\beta}}}(\phi) can be optimized accordingly by introducing a schedule network ϕ\phi. We provide its mathematical derivations in Appendix A.3. Next, we present a formal definition of BDDM and derive its training objectives, Lscore(n)(θ){\mathcal{L}}^{(n)}_{\text{score}}(\theta) and Lstep(n)(ϕ;θ∗){\mathcal{L}}_{\text{step}}^{(n)}(\phi;\theta^{*}), for the score network and the schedule network, respectively, in more detail.

3 Score Network

Recall that a DDPM starts the reverse process with a white noise xT∼N(0,I){\bm{x}}_{T}\sim{\mathcal{N}}({\bm{0}},{\bm{I}}) and takes TT steps to recover the data distribution:

A BDDM, in contrast, starts from the junctional variable xt{\bm{x}}_{t}, and reverts a shorter sequence of diffusion random variables with only nn steps:

where qβ^(x^n−1;xt,ϵn)q_{\hat{{\bm{\beta}}}}(\hat{{\bm{x}}}_{n-1};{\bm{x}}_{t},{\bm{\epsilon}}_{n}) is defined as a re-parameterization on the posterior:

where α^n=∏i=1n1−β^i\hat{\alpha}_{n}=\prod_{i=1}^{n}\sqrt{1-\hat{\beta}_{i}}, xt=αtx0+1−αt2ϵn{\bm{x}}_{t}=\alpha_{t}{\bm{x}}_{0}+\sqrt{1-\alpha_{t}^{2}}{\bm{\epsilon}}_{n} is the junctional variable that maps xt{\bm{x}}_{t} to x^n\hat{{\bm{x}}}_{n} given an approximate index t∼U{(n−1)τ,...,nτ−1,nτ}t\sim{\mathcal{U}}\{(n-1)\tau,...,n\tau-1,n\tau\} and a sampled white noise ϵn∼N(0,I){\bm{\epsilon}}_{n}\sim{\mathcal{N}}({\bm{0}},{\bm{I}}). Detailed derivation from Eq. (9) to (10) is provided in Appendix A.2.

With the above definition, a new form of lower bound to the log marginal likelihood can be derived such that log⁡pθ(x^0)≥Fscore(n)(θ):=−Lscore(n)(θ)−Rθ(x^0,xt),\log p_{\theta}(\hat{{\bm{x}}}_{0})\geq{\mathcal{F}}_{\text{score}}^{(n)}(\theta):=-{\mathcal{L}}^{(n)}_{\text{score}}(\theta)-{\mathcal{R}}_{\theta}(\hat{{\bm{x}}}_{0},{\bm{x}}_{t}), where

See detailed derivation in Proposition 1 in Appendix A.2. In the following Proposition 2, we prove that via the junctional variable xt{\bm{x}}_{t}, the solution θ∗\theta^{*} for optimizing the objective Lddpm(t)(θ),∀t∈{1,...,T}{\mathcal{L}}^{(t)}_{\text{ddpm}}(\theta),\forall t\in\{1,...,T\} is also the solution for optimizing Lscore(n)(θ),∀n∈{2,...,N}{\mathcal{L}}^{(n)}_{\text{score}}(\theta),\forall n\in\{2,...,N\}. Thereby, we show that the score network θ\theta can be trained with Lddpm(t)(θ){\mathcal{L}}^{(t)}_{\text{ddpm}}(\theta) and re-used for reverting the short diffusion process over x^N:0\hat{{\bm{x}}}_{N:0}. Although the newly derived lower bound result in the same objective as the conventional score network, it for the first time establishes a link between the score network θ\theta and x^N:0\hat{{\bm{x}}}_{N:0}. The connection is essential for learning β^\hat{{\bm{\beta}}}, which we will describe next.

4 schedule network

In BDDMs, a schedule network is introduced to the forward process by re-parameterizing β^n\hat{\beta}_{n} as β^n(ϕ)=fϕ(xt;β^n+1)\hat{\beta}_{n}(\phi)=f_{\phi}\left({\bm{x}}_{t};\hat{\beta}_{n+1}\right), and recall that during training, we can use xt=αtx0+1−αt2ϵn{\bm{x}}_{t}=\alpha_{t}{\bm{x}}_{0}+\sqrt{1-\alpha_{t}^{2}}{\bm{\epsilon}}_{n} and β^n+1=1−αt+τ2αt2\hat{\beta}_{n+1}=1-\frac{\alpha_{t+\tau}^{2}}{\alpha_{t}^{2}}. Through the re-parameterization, the task of noise scheduling, i.e., searching for β^\hat{{\bm{\beta}}}, can now be reformulated as training a schedule network fϕf_{\phi} that ancestrally estimates data-dependent variances. The schedule network learns to predict β^n\hat{\beta}_{n} based on the current noisy sample xt{\bm{x}}_{t} – this makes our method fundamentally different from existing and concurrent work, including Kingma et al. 2021 – as we reveal that, aside from β^n+1\hat{\beta}_{n+1}, tt, or nn that reflects diffusion step information, xt{\bm{x}}_{t} is also essential for noise scheduling from a reverse direction at inference time.

where the network parameter set ϕ\phi is learned to estimate the ratio between two consecutive noise scales (β^n\hat{\beta}_{n} and β^n+1\hat{\beta}_{n+1}) from the current noisy input xt{\bm{x}}_{t}.

Finally, at inference time for noise scheduling, starting from a maximum reverse steps (NN) and two hyperparameters (α^N,β^N)(\hat{\alpha}_{N},\hat{\beta}_{N}), we ancestrally predict the noise scale β^n(ϕ)=fϕ(x^n;β^n+1)\hat{\beta}_{n}(\phi)=f_{\phi}\left(\hat{{\bm{x}}}_{n};\hat{\beta}_{n+1}\right), for nn from NN to 11, and cumulatively update the product α^n=α^n+11−β^n+1\hat{\alpha}_{n}=\frac{\hat{\alpha}_{n+1}}{\sqrt{1-\hat{\beta}_{n+1}}}.

Here we describe how to learn the network parameters ϕ\phi effectively. First, we demonstrated that ϕ\phi should be trained after θ\theta is well-optimized, referring to Proposition 3 in Appendix A.3. The Proposition also shows that we are minimizing the gap between the lower bound Fscore(n)(θ∗){\mathcal{F}}_{\text{score}}^{(n)}(\theta^{*}) and log⁡pθ∗(x^0)\log p_{\theta^{*}}(\hat{{\bm{x}}}_{0}), i.e., log⁡pθ∗(x^0)−Fscore(n)(θ∗)\log p_{\theta^{*}}(\hat{{\bm{x}}}_{0})-{\mathcal{F}}_{\text{score}}^{(n)}(\theta^{*}), by minimizing the following objective

which is defined as a KL divergence to directly compare pθ∗(x^n−1∣x^n=xt)p_{\theta^{*}}(\hat{{\bm{x}}}_{n-1}|\hat{{\bm{x}}}_{n}={\bm{x}}_{t}) against the re-parameterized forward process posteriors, which are tractable when conditioned on the junctional noise scale αt\alpha_{t} and x0{\bm{x}}_{0}.

The detailed derivation of Eq. (14) is also provided in the proof of Proposition 3 to get its concrete formulas as shown in Step (8-10) in Alg. 2.

Algorithms: training, noise scheduling, and sampling

Following the theoretical result in Appendix A.3, θ\theta should be optimized before learning ϕ\phi. Thereby first, to train the score network ϵθ{\bm{\epsilon}}_{\theta}, we refer to the settings in (Ho et al. 2020; Chen et al. 2020; Song et al. 2021a) to define β{\bm{\beta}} as a linear noise schedule: βt=βstart+tT(βend−βstart),for1≤t≤T,\beta_{t}=\beta_{\text{start}}+\frac{t}{T}(\beta_{\text{end}}-\beta_{\text{start}}),\quad\text{for}\quad 1\leq t\leq T, where βstart\beta_{\text{start}} and βend\beta_{\text{end}} are two hyperparameter that specifies the start value and the end value. This results in Algorithm 1, which resembles the training algorithm in (Ho et al. 2020).

Next, based on the converged score network θ∗\theta^{*}, we train the schedule network ϕ\phi. We draw an n∼U{2,…,N}n\sim\mathcal{U}\{2,\ldots,N\} at each training step, and then draw a t∼U{(n−1)τ,...,nτ}t\sim\mathcal{U}\{(n-1)\tau,...,n\tau\}. These together can be re-formulated as directly drawing t∼U{τ,...,T−τ}t\sim\mathcal{U}\{\tau,...,T-\tau\} for a finer-scale time step. Then, we sequentially compute the variables needed for calculating Lstep(n)(ϕ;θ∗){\mathcal{L}}_{\text{step}}^{(n)}(\phi;\theta^{*}), as presented in Algorithm 2. We observed that, although a linear schedule is used to define β{\bm{\beta}}, the noise schedule of β^\hat{{\bm{\beta}}} predicted by fϕf_{\phi} is not limited to but rather different from a linear one.

2 Noise scheduling for fast and high-quality sampling

After the score network and the schedule network are trained, the inference procedure can divide into two phases: (1) the noise scheduling phase and (2) the sampling phase.

First, we run the noise scheduling process similarly to a sampling process with NN iterations maximum. Different from training, where αt\alpha_{t} is forward-computed, α^n\hat{\alpha}_{n} is instead a backward-computed variable (from NN to 11) that may deviate from the forward one because {β^i}in−1\{\hat{\beta}_{i}\}_{i}^{n-1} are unknown in the noise scheduling phase during inference. To start noise scheduling, we first set two hyperparameters: α^N\hat{\alpha}_{N} and β^N\hat{\beta}_{N}. We use β1\beta_{1}, the smallest noise scale seen in training, as a threshold to early stop the noise scheduling process so that we can ignore small noise scales (<β1<\beta_{1}) that were never seen by the score network. Overall, the noise scheduling process presents in Algorithm 3.

Experiments

We conducted a series of experiments on neural vocoding tasks to evaluate the proposed BDDMs. First, we compared BDDMs against several strongest models that have been published: the mixture of logistics (MoL) WaveNet (Oord et al. 2018) implemented in (Yamamoto 2020), the WaveGlow (Prenger et al. 2019) implemented in (Valle 2020), the MelGAN (Kumar et al. 2019a) implemented in (Kumar 2019), the HiFi-GAN (Kong et al. 2020b) implemented in (Kong et al. 2020a) and the two most recently proposed diffusion-based vocoders, i.e., WaveGrad (Chen et al. 2020) and DiffWave (Kong et al. 2021), both re-implemented in our code. The hyperparameter settings of BDDMs and all these models are detailed in Appendix B.

In addition, we also compared BDDMs to a variety of scheduling and acceleration techniques applicable to DDPMs, including the grid search (GS) approach in WaveGrad, the fast sampling (FS) approach based on a user-defined 6-step schedule in DiffWave, the DDIMs (Song et al. 2021a) and a noise estimation (NE) approach (San-Roman et al. 2021). For fair and reproducible comparison with other models and approaches, we used the LJSpeech dataset (Ito & Johnson 2017), which consists of 13,100 22kHz audio clips of a female speaker. All diffusion models were trained on the same training split as in (Chen et al. 2020). We also replicated the comparative experiment of neural vocoding using a multi-speaker VCTK dataset (Yamagishi et al. 2019) as presented in Appendix C and obtained a result consistent with that obtained from the LJSpeech dataset.

To assess the quality of each generated audio sample, we used both objective and subjective measures for comparing different neural vocoders given the same ground-truth spectrogram s{\bm{s}} as the condition, i.e., ϵθ(x,s,αt){\bm{\epsilon}}_{\theta}({\bm{x}},{\bm{s}},\alpha_{t}). Specifically, we used two scale-invariant metrics: the perceptual evaluation of speech quality (PESQ) (Rix et al. 2001) and the short-time objective intelligibility (STOI) (Taal et al. 2010) to measure the noisiness and the distortion of the generated speech relative to the reference speech. Mean opinion score (MOS) was also used as a subjective metric for evaluating the naturalness of the generated speech. The assessment scheme of MOS is included in Appendix B.

In Table 1, we compared BDDMs against the state-of-the-art (SOTA) vocoders. To predict noise schedules with different sampling steps (3, 7, and 12), we set three pairs of {α^N,β^N}\{\hat{\alpha}_{N},\hat{\beta}_{N}\} for BDDMs by running on Algorithm 3 a quick hyperparameter grid search, which is detailed in Appendix B. Among the 9 evaluated vocoders, only our proposed BDDMs with 7 and 12 steps and DiffWave with 200 steps showed no statistic-significant difference from the ground-truth in terms of MOS. Moreover, BDDMs significantly outspeeded DiffWave in terms of RTFs. Notably, previous diffusion-based vocoders achieved high MOS scores at the cost of an unacceptable RTF for industrial deployment. In contrast, BDDMs managed to achieve a high standard of generation quality with only 7 sampling steps (143x faster than WaveGrad and 28.6x faster than DiffWave).

In Table 2, we evaluated BDDMs and alternative accelerated sampling methods, which used the same score network for a pair-to-pair comparison. The GS method performed stably when the step number was small (i.e., N≤6N\leq 6) but not scalable to more step numbers, which were therefore bypassed in the comparisons of 7 and 12 steps. The FS method by Song et al. 2021a was linearly interpolated to 3 and 7 steps for a fair comparison. Comparing its 7-step and 3-step results, we observed that the FS performance degraded drastically. Both the DDIM and the NE methods were stable across all the steps but were not performing competitively enough. In comparison, BDDMs consistently attained the leading scores across all the steps. This evaluation confirmed that BDDM was superior to other acceleration methods for DPMs in terms of both stability and quality.

2 Ablation Study and Analysis

We attribute the primary advantage of BDDMs to the newly derived objective Lstep(n){\mathcal{L}}^{(n)}_{\text{step}} for learning ϕ\phi. To better reason about this, we performed an ablation study, where we substituted the proposed loss with the standard negative ELBO for learning ϕ\phi as mentioned by Sohl-Dickstein et al. 2015. We plotted the network outputs with different training losses in Fig. 2. It turned out that, when using Lelbo(n){\mathcal{L}}^{(n)}_{\text{elbo}} to learn ϕ\phi, the network output rapidly collapsed to zero within several training steps; whereas, the network trained with Lstep(n){\mathcal{L}}^{(n)}_{\text{step}} produced fluctuating outputs. The fluctuation is a desirable property showing the network properly predicts tt-dependent noise scales, as tt is a random time step drawn from a uniform distribution in training.

By setting β^=β\hat{{\bm{\beta}}}={\bm{\beta}}, we empirically validated that Fbddm(t):=Fscore(t)+Lstep(t)≥Felbo(t){\mathcal{F}}^{(t)}_{\text{bddm}}:={\mathcal{F}}^{(t)}_{\text{score}}+{\mathcal{L}}^{(t)}_{\text{step}}\geq{\mathcal{F}}^{(t)}_{\text{elbo}} with their respective values at t∈t\in using the same optimized θ∗\theta^{*}. Each value is provided with 95% confidence intervals, as shown in Fig. 3. In this experiment, we used the LJ speech dataset and set T=200T=200 and τ=20\tau=20. Notably, we dropped their common entropy term Rθ(x^0,xt)<0{\mathcal{R}}_{\theta}(\hat{{\bm{x}}}_{0},{\bm{x}}_{t})<0 to mainly compare their KL divergences. This explains those positive lower bound values in the plot. The graph shows that our proposed bound Fbddm(t){\mathcal{F}}^{(t)}_{\text{bddm}} is always a tighter lower bound than the standard one across all examined tt. Moreover, we found that Fbddm(t){\mathcal{F}}^{(t)}_{\text{bddm}} attained low values with a relatively much lower variance for t≤50t\leq 50, where Felbo(t){\mathcal{F}}^{(t)}_{\text{elbo}} was highly volatile. This implies that Fbddm(t){\mathcal{F}}^{(t)}_{\text{bddm}} better tackles the difficult training part, i.e., when the score becomes more challenging to estimate as t→0t\to 0.

Conclusions

BDDMs parameterize the forward and reverse processes with a schedule network and a score network, of which the former’s optimization is tied with the latter by introducing a junctional variable. We derived a new lower bound that leads to the same training loss for the score network as in DDPMs (Ho et al. 2020), which thus enables inheriting any pre-trained score networks in DDPMs. We also showed that training the schedule network after a well-optimized score network can be viewed as tightening the lower bound. Followed from the theoretical results, an efficient training algorithm and a noise scheduling algorithm were respectively designed for BDDMs. Finally, in our experiments, BDDMs showed a clear edge over the previous diffusion-based vocoders.

References

Appendix A Theoretical derivations for BDDMs

In this section, we provide the theoretical supports for the following:

The derivation for upper bounding β^n\hat{\beta}_{n} (see Appendix A.1).

The score network θ\theta trained with Lddpm(t)(θ){\mathcal{L}}^{(t)}_{\text{ddpm}}(\theta) for the reverse process pθ(xt−1∣xt)p_{\theta}({\bm{x}}_{t-1}|{\bm{x}}_{t}) can be re-used for the reverse process pθ(x^n−1∣x^n)p_{\theta}(\hat{{\bm{x}}}_{n-1}|\hat{{\bm{x}}}_{n}) (see Appendix A.2).

The schedule network ϕ\phi can be trained with Lstep(n)(ϕ;θ∗){\mathcal{L}}^{(n)}_{\text{step}}(\phi;\theta^{*}) after the score network θ\theta is optimized. (see Appendix A.3).

Since monotonic noise schedules have been successfully applied to in many prior arts including DPMs (Ho et al. 2020; Kingma et al. 2021) and score-based methods (Song et al. 2020; Song & Ermon 2020), we also follow the monotonic assumption and derive an upper bound for β^n\hat{\beta}_{n} as below:

Suppose the noise schedule for sampling is monotonic, i.e., 0<β^1<…<β^N<10<\hat{\beta}_{1}<\dotsc<\hat{\beta}_{N}<1, then, for 1≤n<N1\leq n<N, β^n\hat{\beta}_{n} satisfies the following inequality:

By the general definition of noise schedule, we know that 0<β^1,…,β^N<10<\hat{\beta}_{1},\dotsc,\hat{\beta}_{N}<1 (Note: no inequality sign in between). Given that α^n=∏i=1n1−β^i\hat{\alpha}_{n}=\prod_{i=1}^{n}\sqrt{1-\hat{\beta}_{i}}, we also have 0<α^1,…,α^t<10<\hat{\alpha}_{1},\dotsc,\hat{\alpha}_{t}<1. First, we show that β^n<1−α^n+121−β^n+1\hat{\beta}_{n}<1-\frac{\hat{\alpha}_{n+1}^{2}}{1-\hat{\beta}_{n+1}}:

Next, we show that β^n<1−α^n+1\hat{\beta}_{n}<1-\hat{\alpha}_{n+1}:

Now, we have β^n<min⁡{1−α^n+121−β^n+1,1−α^n+1}\hat{\beta}_{n}<\min\left\{1-\frac{\hat{\alpha}_{n+1}^{2}}{1-\hat{\beta}_{n+1}},1-\hat{\alpha}_{n+1}\right\}. When 1−α^n+1<1−α^n+121−β^n+11-\hat{\alpha}_{n+1}<1-\frac{\hat{\alpha}_{n+1}^{2}}{1-\hat{\beta}_{n+1}}, we can show that β^n+1<1−α^n+1\hat{\beta}_{n+1}<1-\hat{\alpha}_{n+1}:

By the assumption of monotonic sequence, we also have β^n<β^n+1\hat{\beta}_{n}<\hat{\beta}_{n+1}. Knowing that β^n+1<1−α^n+1\hat{\beta}_{n+1}<1-\hat{\alpha}_{n+1} is always true, we obtain a tighter bound for β^n\hat{\beta}_{n}: 0<β^n<min⁡{1−α^n+121−β^n+1,β^n+1}0<\hat{\beta}_{n}<\min\left\{1-\frac{\hat{\alpha}_{n+1}^{2}}{1-\hat{\beta}_{n+1}},\hat{\beta}_{n+1}\right\}. ∎

A.2 Deriving the training objective for score network

First, followed from the data distribution modeling of BDDMs as proposed in Eq. (8):

we can derive a new lower bound to the log marginal likelihood as follows:

Given xt∼qβ(xt∣x0){\bm{x}}_{t}\sim q_{\bm{\beta}}({\bm{x}}_{t}|{\bm{x}}_{0}), the following lower bound holds for n∈{2,…,N}n\in\{2,\ldots,N\}:

Next, we show that the score network θ\theta trained with Lddpm(t)(θ){\mathcal{L}}_{\text{ddpm}}^{(t)}(\theta) can be re-used in BDDMs. We first provide the derivation for Eq. (9- 10). We have

Suppose xt∼qβ(xt∣x0){\bm{x}}_{t}\sim q_{\bm{\beta}}({\bm{x}}_{t}|{\bm{x}}_{0}), then any solution satisfying θ∗=argminθLddpm(t)(θ),∀t∈{1,...,T},\theta^{*}=\text{argmin}_{\theta}{\mathcal{L}}^{(t)}_{\text{ddpm}}(\theta),\forall t\in\{1,...,T\}, also satisfies θ∗=argminθLscore(n)(θ),∀n∈{2,...,N}\theta^{*}=\text{argmin}_{\theta}{\mathcal{L}}^{(n)}_{\text{score}}(\theta),\forall n\in\{2,...,N\}.

which is proportional to Lddpm(t):=∥ϵn−ϵθ(αtx0+1−αt2ϵn,αt)∥22{\mathcal{L}}^{(t)}_{\text{ddpm}}:=\left\|{\bm{\epsilon}}_{n}-{\bm{\epsilon}}_{\theta}\left({\alpha}_{t}{\bm{x}}_{0}+\sqrt{1-{\alpha}_{t}^{2}}{\bm{\epsilon}}_{n},\alpha_{t}\right)\right\|^{2}_{2} as defined in Eq. (5). Thus,

Next, we can simplify Rθ(x^0,xt){\mathcal{R}}_{\theta}(\hat{{\bm{x}}}_{0},{\bm{x}}_{t}) to a reconstruction loss for x^0\hat{{\bm{x}}}_{0}:

where pθ(x^1∣x^n=xt)p_{\theta}(\hat{{\bm{x}}}_{1}|\hat{{\bm{x}}}_{n}={\bm{x}}_{t}) can be efficiently sampled using the reverse process in (Song et al. 2021a). Yet, in practice, similar to the training in (Song et al. 2021a; Chen et al. 2020; Kong et al. 2021), we dropped Rθ(x^0,xt){\mathcal{R}}_{\theta}(\hat{{\bm{x}}}_{0},{\bm{x}}_{t}) when training θ\theta. In theory, we know that Rθ(x^0,xt){\mathcal{R}}_{\theta}(\hat{{\bm{x}}}_{0},{\bm{x}}_{t}) achieves its optimal value at θ∗=argminθ∥ϵθ(x^1,α^1)−ϵ1∥22\theta^{*}=\text{argmin}_{\theta}\lVert{\bm{\epsilon}}_{\theta}(\hat{{\bm{x}}}_{1},\hat{\alpha}_{1})-{\bm{\epsilon}}_{1}\rVert^{2}_{2}, which shares a similar objective as Lddpm(t){\mathcal{L}}_{\text{ddpm}}^{(t)}. By minimizing Lddpm(t){\mathcal{L}}_{\text{ddpm}}^{(t)}, we train a score network θ∗\theta^{*} that best minimizes Δϵt:=∥ϵt−ϵθ∗(αtx0+1−αt2ϵt,αt)∥22\Delta{\bm{\epsilon}}_{t}:=\lVert{\bm{\epsilon}}_{t}-{\bm{\epsilon}}_{\theta^{*}}({\alpha}_{t}{\bm{x}}_{0}+\sqrt{1-{\alpha}_{t}^{2}}{\bm{\epsilon}}_{t},{\alpha}_{t})\rVert^{2}_{2} for all 1≤t≤T1\leq t\leq T. Since the first diffusion step has the smallest effect on corrupting x^0\hat{{\bm{x}}}_{0} (i.e., β1≈0\beta_{1}\approx 0), it suffices to consider a α^1=1−β1=α1\hat{\alpha}_{1}=\sqrt{1-{\beta}_{1}}={\alpha}_{1}, in which case we can jointly minimize Rθ(x^0,xt){\mathcal{R}}_{\theta}(\hat{{\bm{x}}}_{0},{\bm{x}}_{t}) by minimizing Lddpm(1){\mathcal{L}}_{\text{ddpm}}^{(1)}.

In this sense, during training, given xt∼qβ(xt∣x0){\bm{x}}_{t}\sim q_{\bm{\beta}}({\bm{x}}_{t}|{\bm{x}}_{0}), we can train the score network with the same training objective as in DDPMs and DDIMs. Practically, it is beneficial for BDDMs as we can re-use the score network θ\theta of any well-trained DDPM or DDIM.

A.3 Deriving the training objective for schedule network

Suppose θ\theta has been optimized and hypothetically converged to the optimal θ∗\theta^{*}, where by optimal it means that with θ∗\theta^{*} we have pθ∗(x^n−1∣x^n=xt)=qβ^(x^n−1;xt,ϵn)p_{\theta^{*}}(\hat{{\bm{x}}}_{n-1}|\hat{{\bm{x}}}_{n}={\bm{x}}_{t})=q_{\hat{{\bm{\beta}}}}(\hat{{\bm{x}}}_{n-1};{\bm{x}}_{t},{\bm{\epsilon}}_{n}) given xt∼qβ(xt∣x0){\bm{x}}_{t}\sim q_{\bm{\beta}}({\bm{x}}_{t}|{\bm{x}}_{0}). When β^\hat{{\bm{\beta}}} is unknown but we have x0=x^0{\bm{x}}_{0}=\hat{{\bm{x}}}_{0} and α^n=αt\hat{\alpha}_{n}=\alpha_{t}, we can minimize the gap between the optimal lower bound Fscore(n)(θ∗){\mathcal{F}}_{\text{score}}^{(n)}(\theta^{*}) and log⁡pθ∗(x^0)\log p_{\theta^{*}}(\hat{{\bm{x}}}_{0}), i.e, log⁡pθ∗(x^0)−Fscore(n)(θ∗)\log p_{\theta^{*}}(\hat{{\bm{x}}}_{0})-{\mathcal{F}}_{\text{score}}^{(n)}(\theta^{*}), by minimizing the following objective with respect to β^n\hat{\beta}_{n}:

Note that x^0=x0\hat{{\bm{x}}}_{0}={\bm{x}}_{0}, α^n=αt\hat{\alpha}_{n}=\alpha_{t}, xt=αtx0+1−αt2ϵn{\bm{x}}_{t}={\alpha}_{t}{\bm{x}}_{0}+\sqrt{1-{\alpha}_{t}^{2}}{\bm{\epsilon}}_{n} and pθ∗(x^n−1∣x^n=xt)=qβ^(x^n−1;xt,ϵn)p_{\theta^{*}}(\hat{{\bm{x}}}_{n-1}|\hat{{\bm{x}}}_{n}={\bm{x}}_{t})=q_{\hat{{\bm{\beta}}}}(\hat{{\bm{x}}}_{n-1};{\bm{x}}_{t},{\bm{\epsilon}}_{n}). When x0{\bm{x}}_{0} is given to pθ∗p_{\theta^{*}}, we can express the probability as follows:

where, different from pθ∗(x^n−1∣x^n=xt)p_{\theta^{*}}(\hat{{\bm{x}}}_{n-1}|\hat{{\bm{x}}}_{n}={\bm{x}}_{t}), from Eq. (49) to Eq. (50), instead of conditioning on a specific xt{\bm{x}}_{t}, when x0{\bm{x}}_{0} is given xt{\bm{x}}_{t} can be generated using any z∼N(0,I){\bm{z}}\sim{\mathcal{N}}({\bm{0}},{\bm{I}}).

From this, we can express the gap between log⁡pθ∗(x^0)\log p_{\theta^{*}}(\hat{{\bm{x}}}_{0}) and Fscore(n)(θ∗){\mathcal{F}}_{\text{score}}^{(n)}(\theta^{*}) in the following form:

Next, we evaluate the above KL divergence term. By definition, we have

As we use a schedule network ϕ\phi to estimate β^n\hat{\beta}_{n} from (α^n+1,β^n+1)(\hat{\alpha}_{n+1},\hat{\beta}_{n+1}) as defined in Eq. (13), we obtain the final step loss for learning ϕ\phi:

This proposed objective for training the schedule network can be interpreted as to better model the data distribution (i.e., maximizing log⁡pθ(x^0)\log p_{\theta}(\hat{{\bm{x}}}_{0})) by correcting the gradient scale β^n\hat{\beta}_{n} for the next reverse step (from x^n\hat{{\bm{x}}}_{n} to x^n−1\hat{{\bm{x}}}_{n-1}) given the gradient vector ϵθ∗{\bm{\epsilon}}_{\theta^{*}} estimated by the score network θ∗\theta^{*}.

Appendix B Experimental details

We reproduced the grid search algorithm in (Chen et al. 2020), in which a 6-step noise schedule was searched. In our paper, we generalized the grid search algorithm by similarly sweeping the NN-step noise schedule over the following possibilities with a bin width M=9M=9:

where ⊗\otimes denotes the cartesian product applied on two sets. LS-MSE was used as a metric to select the solution during the search. When N=6N=6, we resemble the GS algorithm in (Chen et al. 2020). Note that above searching method normally does not scale up to N>6N>6 steps for its exponential computational cost O(9N)\mathcal{O}(9^{N}).

B.2 Hyperparameter setting in BDDMs

Algorithm 2 took a skip factor τ\tau to control the stride for training the schedule network. The value of τ\tau would affect the coverage of step sizes when training the schedule network, hence affecting the predicted number of steps NN for inference – the higher τ\tau is, the shorter the predicted inference schedule tends to be. We set τ=66\tau=66 for training the BDDM vocoders in this paper.

For initializing Algorithm 3 for noise scheduling, we could take as few as 11 training sample for validation, perform a grid search on the hyperparameters {(α^N=0.1αTi,β^N=0.1j)}\{(\hat{\alpha}_{N}=0.1\alpha_{T}i,\hat{\beta}_{N}=0.1j)\} for i,j=1,...,9i,j=1,...,9, i.e., 8181 possibilities in total, and use the PESQ measure as the selection metric. Then, the predicted noise schedule corresponding to the maximum PESQ was stored and applied to the online inference afterward, as shown in Algorithm 4. Note that this searching has a complexity of only O(M2)\mathcal{O}(M^{2}) (e.g., M=9M=9 in this case), which is much more efficient than O(MN)\mathcal{O}(M^{N}) in the conventional grid search algorithm in (Chen et al. 2020), as discussed in Section B.1.

B.3 Implementation details

Our proposed BDDMs and the baseline methods were all implemented with the Pytorch library. The score networks for the LJ and VCTK speech datasets were trained from scratch on a single NVIDIA Tesla P40 GPU with batch size 3232 for about 1M steps, which took about 3 days.

For the model architecture, we used the same architecture as in DiffWave (Kong et al. 2021) for the score network with 128 residual channels; we adopted a lightweight GALR network (Lam et al. 2021) for the schedule network. GALR was originally proposed for speech enhancement, so we considered it well suited for predicting the noise scales. For the configuration of the GALR network, we used a window length of 8 samples for encoding, a segment size of 64 for segmentation and only two GALR blocks of 128 hidden dimensions, and other settings were inherited from (Lam et al. 2021). To make the schedule network output with a proper range and dimension, we applied a sigmoid function to the last block’s output of the GALR network. Then the result was averaged over the segments and the feature dimensions to obtain the predicted ratio: σϕ(x)=AvgPool2D(σ(GALR(x)))\sigma_{\phi}({\bm{x}})=\text{AvgPool2D}(\sigma(\text{GALR}({\bm{x}}))), where GALR(⋅)\text{GALR}(\cdot) denotes the GALR network, AvgPool2D(⋅)\text{AvgPool2D}(\cdot) denotes the average pooling operation applied to the segments and the feature dimensions, and σ(x):=1/(1+e−x)\sigma(x):=1/(1+e^{-x}). The same network architecture was used for the NE approach for estimating αt2\alpha_{t}^{2} and was shown better than the ConvTASNet used in the original paper (San-Roman et al. 2021). It is also notable that the computational cost of a schedule network is indeed fractional compared to the cost of a score network, as predicting a noise scalar variable is intrinsically a relatively much easier task. Our GALR-based schedule network, while being able to produce stable and reliable results, was about 3.6 times faster than the score network. The training of schedule networks for BDDMs took only 10k steps to converge, which consumed no more than an hour on a single GPU.

Regarding the image generation task, to demonstrate the generalizability of our method, we directly adopted a score network pre-trained on the CIFAR-10 dataset implemented by a third-party open-source repository. Regarding the schedule network, to demonstrate that it does not have to use specialized architecture, we replaced GALR by the VGG11 (Simonyan & Zisserman 2014), which was also used by as a noise estimator in (San-Roman et al. 2021). The output dimension (number of classes) of VGG11 was set to 11. Similar to the setting for GALR in speech synthesis, we added a sigmoid activation to the last layer to ensure a $$ output. Similar to the training in speech domain, we trained the VGG11-based schedule networks while freezing the score networks for 10k steps, which normally can be finished in about two hours.

Our code for the speech vocoding and the image generation experiments will be uploaded to Github after the final decision of ICLR is released.

B.4 Crowd-sourced subjective evaluation

All our Mean Opinion Score (MOS) tests were crowd-sourced. We refer to the MOS scores in (Protasio Ribeiro et al. 2011), and the scoring criteria have been included in Table 3 for completeness. The samples were presented and rated one at a time by the testers.

Appendix C Additional experiments

A demonstration page at https://bilateral-denoising-diffusion-model.github.io shows some samples generated by BDDMs trained on LJ speech and VCTK datasets.

In addition to the single-speaker speech synthesis, we evaluated BDDMs on the multi-speaker speech synthesis benchmark VCTK (Yamagishi et al. 2019). VCTK consists of utterances sampled at 4848 KHz by 108108 native English speakers with various accents. We split the VCTK dataset for training and testing: 100 speakers were used for training the multi-speaker model and 8 speakers for testing. We trained on a 44257-utterance subset (40 hours) and evaluated on a held-out 100-utterance subset. For the score network, we used the Wavegrad architecture (Chen et al. 2020) so as to examine whether the superiority of BDDMs remains in a different dataset and with a different score network architecture.

Results are presented in Table 4. For this multi-speaker VCTK dataset, we obtained consistent observations with that for the single-speaker LJ dataset presented in the main paper. Again, the proposed BDDM with only 16 or 21 steps outperformed the DDPM with 1,000 steps. To the best of our knowledge, ours was the first work that reported this degree of superior. When reducing to 8 steps, BDDM obtained performance on par with (except for a worse PESQ) the costly grid-searched 8 steps (which were unscalable to more steps) in DDPM. For NE, we could again observe a degradation from its 16 steps to 21 steps, indicating the instability of NE for the VCTK dataset likewise. In contrast, BDDM gave continuously improved performance while increasing the step number.

C.2 Comparing different reverse processes for BDDMs

This section demonstrates that BDDMs do not restrict the sampling procedure to a specialized reverse process in Algorithm 4. In particular, we evaluated different reverse processes, including that of DDPMs as shown in Eq. (4) and DDIMs (Song et al. 2021a), for BDDMs and compared the objective scores on the generated samples. DDIMs (Song et al. 2021a) formulate a non-Markovian generative process that accelerates the inference while keeping the same training procedure as DDPMs. The original generative process in Eq. (4) in DDPMs is modified into

where γ{\bm{\gamma}} is a sub-sequence of length NN of [1,...,T][1,...,T] with γN=T\gamma_{N}=T, and γˉ:={1,...,T}∖γ\bar{{\bm{\gamma}}}:=\{1,...,T\}\setminus{\bm{\gamma}} is defined as its complement; Therefore, only part of the models are used in the sampling process.

To achieve the above, DDIMs defined a prediction function fθ(t)(xt)f_{\theta}^{(t)}({\bm{x}}_{t}) that depends on ϵθ{\bm{\epsilon}}_{\theta} to predict the observation x0{\bm{x}}_{0} given xt{\bm{x}}_{t} directly:

By leveraging this prediction function, the conditionals in Eq. (71) are formulated as

where the detailed derivation of σt\sigma_{t} and ς\varsigma can be referred to (Song et al. 2021a). In the original DDIMs, the accelerated reverse process produces samples over the subsequence of β{\bm{\beta}} indexed by γ{\bm{\gamma}}: β^={βn∣n∈γ}\hat{{\bm{\beta}}}=\{\beta_{n}|n\in{\bm{\gamma}}\}. In BDDMs, to apply the DDIM reverse process, we use the β^\hat{{\bm{\beta}}} predicted by the schedule network in place of a subsequence of the training schedule β{\bm{\beta}}.

Finally. the objective scores are given in Table 5. Note that the subjective evaluation (MOS) is omitted here since the other assessments above have shown that the MOS scores are highly correlated with the objective measures, including STOI and PESQ. They indicate that applying BDDMs to either DDPM or DDIM reverse process leads to comparable and competitive results. Meanwhile, the results show some subtle differences: BDDMs over a DDPM reverse process gave slightly better samples in terms of signal error and consistency metrics (i.e., LS-MSE and MCD), while BDDM over a DDIM reverse process tended to generate better samples in terms of intelligibility and perceptual metrics (i.e., STOI and PESQ).

C.3 Unconditional image generation

For the unconditional image generation task, we evaluated the proposed BDDMs on the benchmark CIFAR-10 (32 ×\times 32) dataset. The score functions, including those initially proposed in DDPMs (Ho et al. 2020) or DDIMs (Song et al. 2021a) and those pre-trained in the above third-party implementations, are all conditioned on a discrete step-index. We estimated the noise schedule β^\hat{{\bm{\beta}}} in continuous space using the VGG11 schedule network and then mapped it to discrete time schedule using the approximation method in (Kong & Ping 2021).

Table 6 shows the performances of different sampling methods for DDPMs in CIFAR-10. By setting the maximum number of sampling steps (NN) for noise scheduling, we can fairly compare the improvements achieved by BDDMs against related methods in the literature in terms of FID. Remarkably, BDDMs with 100 sampling steps not only surpassed the 1000-step DDPM baseline, but also produced the SOTA FID performance amongst all generative models using less than or equal to 100100 sampling steps.