BDDM: Bilateral Denoising Diffusion Models for Fast and High-Quality Speech Synthesis
Max W. Y. Lam, Jun Wang, Dan Su, Dong Yu
Introduction
Deep generative models have shown a tremendous advancement in speech synthesis (van den Oord et al. 2016; Kalchbrenner et al. 2018; Prenger et al. 2019; Kumar et al. 2019a; Kong et al. 2020b; Chen et al. 2020; Kong et al. 2021). Successful generative models can be mainly divided into two categories: generative adversarial network (GAN) (Goodfellow et al. 2014) based and likelihood-based. The former is based on adversarial learning, where the objective is to generate data indistinguishable from the training data. Yet, the training GANs can be very unstable, and the relevant training objectives are not suitable to compare against different GANs. The latter uses log-likelihood or surrogate objectives for training, but they also have intrinsic limitations regarding generation speed or quality. For example, the autoregressive models (van den Oord et al. 2016; Kalchbrenner et al. 2018), while being capable of generating high-fidelity data, are limited by their inherently slow sampling process and the poor scaling properties on high-dimensional data. Likewise, the flow-based models (Dinh et al. 2016; Kingma & Dhariwal 2018; Chen et al. 2018; Papamakarios et al. 2021) rely on specialized architectures to build a normalized probability model, whose training is less parameter-efficient. Other prior works use surrogate objectives, such as the evidence lower bound in variational auto-encoders (Kingma & Welling 2013; Rezende et al. 2014; Maaløe et al. 2019) and the contrastive divergence in energy-based models (Hinton 2002; Carreira-Perpinan & Hinton 2005). These models, despite showing improved speed, typically only work well for low-dimensional data, and, in general, the sample qualities are not competitive to the GAN-based and the autoregressive models (Bond-Taylor et al. 2021).
An up-and-coming class of likelihood-based models is the diffusion probabilistic models (DPMs) (Sohl-Dickstein et al. 2015), which introduces the idea of using a forward diffusion process to sequentially corrupt a given distribution and learning the reversal of such diffusion process to restore the data distribution for sampling. From a similar perspective, Song & Ermon 2019 proposed the score-based generative models by applying the score matching technique (Hyvarinen & Dayan 2005) to train a neural network such that samples can be generated via Langevin dynamics. Along these two lines of research, Ho et al. 2020 proposed the denoising diffusion probabilistic models (DDPMs) for high-quality image syntheses. Dhariwal & Nichol 2021 demonstrated that improved DDPMs Nichol & Dhariwal 2021 are capable of generating high-quality images of comparable or even superior quality to the state-of-the-art (SOTA) GAN-based models. For speech syntheses, DDPMs were also applied in Wavegrad (Chen et al. 2020) and DiffWave (Kong et al. 2021) to produce higher-fidelity audio samples than the conventional non-autoregressive models (Yamamoto et al. 2020; Kumar et al. 2019b; Yang et al. 2020; Bińkowski et al. 2020) and matched the quality of the SOTA autoregressive methods (Chen et al. 2020).
Despite the compelling results, the diffusion generative models are two to three orders of magnitude slower than other generative models such as GANs and VAEs. Their primary limitation is that they require up to thousands of diffusion steps during training to learn the target distribution. Therefore a large number of reverse steps are often required at sampling time. Recently, extensive investigations have been conducted to reduce the sampling steps for efficiently generating high-quality samples, which we will discuss in the related work in Section 2. Distinctively, we conceived that we might train a neural network to efficiently and adaptively estimate a much shorter noise schedule for sampling while achieving generation performances comparable or superior to the conventional DPMs. With such an incentive, after introducing the conventional DPMs as our background in Section 3, we propose in Section 4 bilateral denoising diffusion models (BDDMs), named after a bilateral modeling perspective – parameterizing the forward and reverse processes with a schedule network and a score network, respectively. We theoretically derive that the schedule network should be trained after the score network is optimized. For training the schedule network, we propose a novel objective to minimize the gap between a newly derived lower bound and the log marginal likelihood. We describe the training algorithm as well as the fast and high-quality sampling algorithm in Section 5. The training of the schedule network converges very fast using our newly derived objective, and its training only adds negligible overhead to DDPM’s. In Section 6, our neural vocoding experiments demonstrated that BDDMs could generate high-fidelity samples with as few as three sampling steps. Moreover, our method can produce speech samples indistinguishable from human speech with only seven sampling steps (143x faster than WaveGrad and 28.6x faster than DiffWave).
Related Work
Another class of noise scheduling methods searches for a subsequence of time indices of the training noise schedule, which we call the time schedule. DDIMs (Song et al. 2021a) introduced an accelerated reverse process that relies on a pre-specified time schedule. A linear and a quadratic time schedule were used in DDIMs and showed superior generation quality over DDPMs within 10 to 100 sampling steps. Nichol & Dhariwal 2021 proposed a re-scaled noise schedule for fast sampling, but this also requires pre-specifying the time schedule and the training noise schedule. Nichol & Dhariwal 2021 also proposed learning variances for the reverse processes, whereas the variances of the forward processes, i.e., the noise schedule, which affected both the means and variances of the reverse processes, were not learnable. According to the results of (Song et al. 2021a; Nichol & Dhariwal 2021), using a linear or quadratic time schedule resulted in quite different performances in different datasets, implying that the optimal choice of schedule varies with the datasets. So, there remains a challenge in finding a short and effective schedule for fast sampling on different datasets. Notably, Kong & Ping 2021 proposed a method to map a noise schedule to a time schedule for fast sampling. In this sense, searching for a time schedule becomes a sub-set of the noise scheduling problem, which resembles the above category of methods.
Although DPMs (Sohl-Dickstein et al. 2015) and DDPMs (Ho et al. 2019) mentioned that the noise schedule could be learned by re-parameterization, the approach was not investigated in their works. Closely related works that learn a noise schedule emerged until very recently. San-Roman et al. 2021 proposed a noise estimation (NE) method, which trained a neural net with a regression loss to estimate the noise scale from the noisy sample at each time point, and then predicted the next noise scale. However, NE requires a prior assumption of the noise schedule following a linear or Fibonacci rule. Most recently, a concurrent work to ours by Kingma et al. 2021 jointly trained a neural net to predict the signal-to-noise ratio (SNR) by maximizing the variational lower bound. The SNR was then used for noise scheduling. Different from ours, this scheduling neural net only took as input and is independent of the noisy sample generated during the loop of sampling process. Intrinsically, with limited information about the sampled data, the predicted SNR could deviate from the actual SNR of the noisy data during sampling.
Background
2 Denoising Diffusion Probabilistic Models (DDPMs)
Based on the nice property of isotropic Gaussians, one can express directly conditioned on :
To revert this forward process, DDPMs employ a score network Here, is conditioned on the continuous noise scale , as in (Song et al. 2021b; Chen et al. 2020). Alternatively, the score network can also be conditioned on a discrete time index , as in (Song et al. 2021a; Ho et al. 2020). An approximate mapping of a noise schedule to a time schedule (Kong & Ping 2021) exists, therefore we consider conditioning on noise scales as the general case. to define
Note that the calculation of the complete ELBO in Eq. (1) requires forward passes of the score network, which would make the training computationally prohibitive for a large . To feasibly train the score network, instead of computing the complete ELBO, Ho et al. 2020 proposed an efficient training mechanism by sampling from a discrete uniform distribution: , , at each training iteration to compute the training loss:
Bilateral denoising diffusion models (BDDMs)
For fast sampling with DPMs, we strive for a noise schedule for sampling that is much shorter than the noise schedule for training. As shown in Fig. 1, we define two separate diffusion processes corresponding to the noise schedules, and , respectively. The upper diffusion process parameterized by is the same as in Eq. (2), whereas the lower process is defined as with much fewer diffusion steps (). In our problem formulation, is given, but is unknown. The goal is to find a for the reverse process such that can be effectively recovered from with reverse steps.
2 Model description
Although many prior arts (Ho et al. 2020; Chen et al. 2020; Song et al. 2021a; San-Roman et al. 2021) directly applied a shortened linear or Fibonacci noise schedule to the reverse process, we argue that these are sub-optimal solutions. Theoretically, the diffusion process specified by a new shortened noise schedule is essentially different from the one used to train the score network . Therefore, is not guaranteed suitable for reverting the shortened diffusion process. This issue motivated a novel modeling perspective to establish a link between the shortened schedule and the score network , i.e., to have optimized according to .
As a starting point, we consider an , where is a hyperparameter controlling the step size such that each diffusion step between two consecutive variables in the shorter diffusion process corresponds to diffusion steps in the longer one. Based on Eq. (2), we define the following:
where is an intermediate diffused variable we introduced to link the two differently indexed diffusion sequences. We call it a junctional variable, which can be easily generated given and during training: .
Unfortunately, for the reverse process when is not given, the junctional variable is intractable. However, our key observation is that while using the score by a score network trained for the long -parameterized diffusion process, a short noise schedule can be optimized accordingly by introducing a schedule network . We provide its mathematical derivations in Appendix A.3. Next, we present a formal definition of BDDM and derive its training objectives, and , for the score network and the schedule network, respectively, in more detail.
3 Score Network
Recall that a DDPM starts the reverse process with a white noise and takes steps to recover the data distribution:
A BDDM, in contrast, starts from the junctional variable , and reverts a shorter sequence of diffusion random variables with only steps:
where is defined as a re-parameterization on the posterior:
where , is the junctional variable that maps to given an approximate index and a sampled white noise . Detailed derivation from Eq. (9) to (10) is provided in Appendix A.2.
With the above definition, a new form of lower bound to the log marginal likelihood can be derived such that where
See detailed derivation in Proposition 1 in Appendix A.2. In the following Proposition 2, we prove that via the junctional variable , the solution for optimizing the objective is also the solution for optimizing . Thereby, we show that the score network can be trained with and re-used for reverting the short diffusion process over . Although the newly derived lower bound result in the same objective as the conventional score network, it for the first time establishes a link between the score network and . The connection is essential for learning , which we will describe next.
4 schedule network
In BDDMs, a schedule network is introduced to the forward process by re-parameterizing as , and recall that during training, we can use and . Through the re-parameterization, the task of noise scheduling, i.e., searching for , can now be reformulated as training a schedule network that ancestrally estimates data-dependent variances. The schedule network learns to predict based on the current noisy sample – this makes our method fundamentally different from existing and concurrent work, including Kingma et al. 2021 – as we reveal that, aside from , , or that reflects diffusion step information, is also essential for noise scheduling from a reverse direction at inference time.
where the network parameter set is learned to estimate the ratio between two consecutive noise scales ( and ) from the current noisy input .
Finally, at inference time for noise scheduling, starting from a maximum reverse steps () and two hyperparameters , we ancestrally predict the noise scale , for from to , and cumulatively update the product .
Here we describe how to learn the network parameters effectively. First, we demonstrated that should be trained after is well-optimized, referring to Proposition 3 in Appendix A.3. The Proposition also shows that we are minimizing the gap between the lower bound and , i.e., , by minimizing the following objective
which is defined as a KL divergence to directly compare against the re-parameterized forward process posteriors, which are tractable when conditioned on the junctional noise scale and .
The detailed derivation of Eq. (14) is also provided in the proof of Proposition 3 to get its concrete formulas as shown in Step (8-10) in Alg. 2.
Algorithms: training, noise scheduling, and sampling
Following the theoretical result in Appendix A.3, should be optimized before learning . Thereby first, to train the score network , we refer to the settings in (Ho et al. 2020; Chen et al. 2020; Song et al. 2021a) to define as a linear noise schedule: where and are two hyperparameter that specifies the start value and the end value. This results in Algorithm 1, which resembles the training algorithm in (Ho et al. 2020).
Next, based on the converged score network , we train the schedule network . We draw an at each training step, and then draw a . These together can be re-formulated as directly drawing for a finer-scale time step. Then, we sequentially compute the variables needed for calculating , as presented in Algorithm 2. We observed that, although a linear schedule is used to define , the noise schedule of predicted by is not limited to but rather different from a linear one.
2 Noise scheduling for fast and high-quality sampling
After the score network and the schedule network are trained, the inference procedure can divide into two phases: (1) the noise scheduling phase and (2) the sampling phase.
First, we run the noise scheduling process similarly to a sampling process with iterations maximum. Different from training, where is forward-computed, is instead a backward-computed variable (from to ) that may deviate from the forward one because are unknown in the noise scheduling phase during inference. To start noise scheduling, we first set two hyperparameters: and . We use , the smallest noise scale seen in training, as a threshold to early stop the noise scheduling process so that we can ignore small noise scales () that were never seen by the score network. Overall, the noise scheduling process presents in Algorithm 3.
Experiments
We conducted a series of experiments on neural vocoding tasks to evaluate the proposed BDDMs. First, we compared BDDMs against several strongest models that have been published: the mixture of logistics (MoL) WaveNet (Oord et al. 2018) implemented in (Yamamoto 2020), the WaveGlow (Prenger et al. 2019) implemented in (Valle 2020), the MelGAN (Kumar et al. 2019a) implemented in (Kumar 2019), the HiFi-GAN (Kong et al. 2020b) implemented in (Kong et al. 2020a) and the two most recently proposed diffusion-based vocoders, i.e., WaveGrad (Chen et al. 2020) and DiffWave (Kong et al. 2021), both re-implemented in our code. The hyperparameter settings of BDDMs and all these models are detailed in Appendix B.
In addition, we also compared BDDMs to a variety of scheduling and acceleration techniques applicable to DDPMs, including the grid search (GS) approach in WaveGrad, the fast sampling (FS) approach based on a user-defined 6-step schedule in DiffWave, the DDIMs (Song et al. 2021a) and a noise estimation (NE) approach (San-Roman et al. 2021). For fair and reproducible comparison with other models and approaches, we used the LJSpeech dataset (Ito & Johnson 2017), which consists of 13,100 22kHz audio clips of a female speaker. All diffusion models were trained on the same training split as in (Chen et al. 2020). We also replicated the comparative experiment of neural vocoding using a multi-speaker VCTK dataset (Yamagishi et al. 2019) as presented in Appendix C and obtained a result consistent with that obtained from the LJSpeech dataset.
To assess the quality of each generated audio sample, we used both objective and subjective measures for comparing different neural vocoders given the same ground-truth spectrogram as the condition, i.e., . Specifically, we used two scale-invariant metrics: the perceptual evaluation of speech quality (PESQ) (Rix et al. 2001) and the short-time objective intelligibility (STOI) (Taal et al. 2010) to measure the noisiness and the distortion of the generated speech relative to the reference speech. Mean opinion score (MOS) was also used as a subjective metric for evaluating the naturalness of the generated speech. The assessment scheme of MOS is included in Appendix B.
In Table 1, we compared BDDMs against the state-of-the-art (SOTA) vocoders. To predict noise schedules with different sampling steps (3, 7, and 12), we set three pairs of for BDDMs by running on Algorithm 3 a quick hyperparameter grid search, which is detailed in Appendix B. Among the 9 evaluated vocoders, only our proposed BDDMs with 7 and 12 steps and DiffWave with 200 steps showed no statistic-significant difference from the ground-truth in terms of MOS. Moreover, BDDMs significantly outspeeded DiffWave in terms of RTFs. Notably, previous diffusion-based vocoders achieved high MOS scores at the cost of an unacceptable RTF for industrial deployment. In contrast, BDDMs managed to achieve a high standard of generation quality with only 7 sampling steps (143x faster than WaveGrad and 28.6x faster than DiffWave).
In Table 2, we evaluated BDDMs and alternative accelerated sampling methods, which used the same score network for a pair-to-pair comparison. The GS method performed stably when the step number was small (i.e., ) but not scalable to more step numbers, which were therefore bypassed in the comparisons of 7 and 12 steps. The FS method by Song et al. 2021a was linearly interpolated to 3 and 7 steps for a fair comparison. Comparing its 7-step and 3-step results, we observed that the FS performance degraded drastically. Both the DDIM and the NE methods were stable across all the steps but were not performing competitively enough. In comparison, BDDMs consistently attained the leading scores across all the steps. This evaluation confirmed that BDDM was superior to other acceleration methods for DPMs in terms of both stability and quality.
2 Ablation Study and Analysis
We attribute the primary advantage of BDDMs to the newly derived objective for learning . To better reason about this, we performed an ablation study, where we substituted the proposed loss with the standard negative ELBO for learning as mentioned by Sohl-Dickstein et al. 2015. We plotted the network outputs with different training losses in Fig. 2. It turned out that, when using to learn , the network output rapidly collapsed to zero within several training steps; whereas, the network trained with produced fluctuating outputs. The fluctuation is a desirable property showing the network properly predicts -dependent noise scales, as is a random time step drawn from a uniform distribution in training.
By setting , we empirically validated that with their respective values at using the same optimized . Each value is provided with 95% confidence intervals, as shown in Fig. 3. In this experiment, we used the LJ speech dataset and set and . Notably, we dropped their common entropy term to mainly compare their KL divergences. This explains those positive lower bound values in the plot. The graph shows that our proposed bound is always a tighter lower bound than the standard one across all examined . Moreover, we found that attained low values with a relatively much lower variance for , where was highly volatile. This implies that better tackles the difficult training part, i.e., when the score becomes more challenging to estimate as .
Conclusions
BDDMs parameterize the forward and reverse processes with a schedule network and a score network, of which the former’s optimization is tied with the latter by introducing a junctional variable. We derived a new lower bound that leads to the same training loss for the score network as in DDPMs (Ho et al. 2020), which thus enables inheriting any pre-trained score networks in DDPMs. We also showed that training the schedule network after a well-optimized score network can be viewed as tightening the lower bound. Followed from the theoretical results, an efficient training algorithm and a noise scheduling algorithm were respectively designed for BDDMs. Finally, in our experiments, BDDMs showed a clear edge over the previous diffusion-based vocoders.
References
Appendix A Theoretical derivations for BDDMs
In this section, we provide the theoretical supports for the following:
The derivation for upper bounding (see Appendix A.1).
The score network trained with for the reverse process can be re-used for the reverse process (see Appendix A.2).
The schedule network can be trained with after the score network is optimized. (see Appendix A.3).
Since monotonic noise schedules have been successfully applied to in many prior arts including DPMs (Ho et al. 2020; Kingma et al. 2021) and score-based methods (Song et al. 2020; Song & Ermon 2020), we also follow the monotonic assumption and derive an upper bound for as below:
Suppose the noise schedule for sampling is monotonic, i.e., , then, for , satisfies the following inequality:
By the general definition of noise schedule, we know that (Note: no inequality sign in between). Given that , we also have . First, we show that :
Next, we show that :
Now, we have . When , we can show that :
By the assumption of monotonic sequence, we also have . Knowing that is always true, we obtain a tighter bound for : . ∎
A.2 Deriving the training objective for score network
First, followed from the data distribution modeling of BDDMs as proposed in Eq. (8):
we can derive a new lower bound to the log marginal likelihood as follows:
Given , the following lower bound holds for :
Next, we show that the score network trained with can be re-used in BDDMs. We first provide the derivation for Eq. (9- 10). We have
Suppose , then any solution satisfying also satisfies .
which is proportional to as defined in Eq. (5). Thus,
Next, we can simplify to a reconstruction loss for :
where can be efficiently sampled using the reverse process in (Song et al. 2021a). Yet, in practice, similar to the training in (Song et al. 2021a; Chen et al. 2020; Kong et al. 2021), we dropped when training . In theory, we know that achieves its optimal value at , which shares a similar objective as . By minimizing , we train a score network that best minimizes for all . Since the first diffusion step has the smallest effect on corrupting (i.e., ), it suffices to consider a , in which case we can jointly minimize by minimizing .
In this sense, during training, given , we can train the score network with the same training objective as in DDPMs and DDIMs. Practically, it is beneficial for BDDMs as we can re-use the score network of any well-trained DDPM or DDIM.
A.3 Deriving the training objective for schedule network
Suppose has been optimized and hypothetically converged to the optimal , where by optimal it means that with we have given . When is unknown but we have and , we can minimize the gap between the optimal lower bound and , i.e, , by minimizing the following objective with respect to :
Note that , , and . When is given to , we can express the probability as follows:
where, different from , from Eq. (49) to Eq. (50), instead of conditioning on a specific , when is given can be generated using any .
From this, we can express the gap between and in the following form:
Next, we evaluate the above KL divergence term. By definition, we have
As we use a schedule network to estimate from as defined in Eq. (13), we obtain the final step loss for learning :
This proposed objective for training the schedule network can be interpreted as to better model the data distribution (i.e., maximizing ) by correcting the gradient scale for the next reverse step (from to ) given the gradient vector estimated by the score network .
Appendix B Experimental details
We reproduced the grid search algorithm in (Chen et al. 2020), in which a 6-step noise schedule was searched. In our paper, we generalized the grid search algorithm by similarly sweeping the -step noise schedule over the following possibilities with a bin width :
where denotes the cartesian product applied on two sets. LS-MSE was used as a metric to select the solution during the search. When , we resemble the GS algorithm in (Chen et al. 2020). Note that above searching method normally does not scale up to steps for its exponential computational cost .
B.2 Hyperparameter setting in BDDMs
Algorithm 2 took a skip factor to control the stride for training the schedule network. The value of would affect the coverage of step sizes when training the schedule network, hence affecting the predicted number of steps for inference – the higher is, the shorter the predicted inference schedule tends to be. We set for training the BDDM vocoders in this paper.
For initializing Algorithm 3 for noise scheduling, we could take as few as training sample for validation, perform a grid search on the hyperparameters for , i.e., possibilities in total, and use the PESQ measure as the selection metric. Then, the predicted noise schedule corresponding to the maximum PESQ was stored and applied to the online inference afterward, as shown in Algorithm 4. Note that this searching has a complexity of only (e.g., in this case), which is much more efficient than in the conventional grid search algorithm in (Chen et al. 2020), as discussed in Section B.1.
B.3 Implementation details
Our proposed BDDMs and the baseline methods were all implemented with the Pytorch library. The score networks for the LJ and VCTK speech datasets were trained from scratch on a single NVIDIA Tesla P40 GPU with batch size for about 1M steps, which took about 3 days.
For the model architecture, we used the same architecture as in DiffWave (Kong et al. 2021) for the score network with 128 residual channels; we adopted a lightweight GALR network (Lam et al. 2021) for the schedule network. GALR was originally proposed for speech enhancement, so we considered it well suited for predicting the noise scales. For the configuration of the GALR network, we used a window length of 8 samples for encoding, a segment size of 64 for segmentation and only two GALR blocks of 128 hidden dimensions, and other settings were inherited from (Lam et al. 2021). To make the schedule network output with a proper range and dimension, we applied a sigmoid function to the last block’s output of the GALR network. Then the result was averaged over the segments and the feature dimensions to obtain the predicted ratio: , where denotes the GALR network, denotes the average pooling operation applied to the segments and the feature dimensions, and . The same network architecture was used for the NE approach for estimating and was shown better than the ConvTASNet used in the original paper (San-Roman et al. 2021). It is also notable that the computational cost of a schedule network is indeed fractional compared to the cost of a score network, as predicting a noise scalar variable is intrinsically a relatively much easier task. Our GALR-based schedule network, while being able to produce stable and reliable results, was about 3.6 times faster than the score network. The training of schedule networks for BDDMs took only 10k steps to converge, which consumed no more than an hour on a single GPU.
Regarding the image generation task, to demonstrate the generalizability of our method, we directly adopted a score network pre-trained on the CIFAR-10 dataset implemented by a third-party open-source repository. Regarding the schedule network, to demonstrate that it does not have to use specialized architecture, we replaced GALR by the VGG11 (Simonyan & Zisserman 2014), which was also used by as a noise estimator in (San-Roman et al. 2021). The output dimension (number of classes) of VGG11 was set to . Similar to the setting for GALR in speech synthesis, we added a sigmoid activation to the last layer to ensure a $$ output. Similar to the training in speech domain, we trained the VGG11-based schedule networks while freezing the score networks for 10k steps, which normally can be finished in about two hours.
Our code for the speech vocoding and the image generation experiments will be uploaded to Github after the final decision of ICLR is released.
B.4 Crowd-sourced subjective evaluation
All our Mean Opinion Score (MOS) tests were crowd-sourced. We refer to the MOS scores in (Protasio Ribeiro et al. 2011), and the scoring criteria have been included in Table 3 for completeness. The samples were presented and rated one at a time by the testers.
Appendix C Additional experiments
A demonstration page at https://bilateral-denoising-diffusion-model.github.io shows some samples generated by BDDMs trained on LJ speech and VCTK datasets.
In addition to the single-speaker speech synthesis, we evaluated BDDMs on the multi-speaker speech synthesis benchmark VCTK (Yamagishi et al. 2019). VCTK consists of utterances sampled at KHz by native English speakers with various accents. We split the VCTK dataset for training and testing: 100 speakers were used for training the multi-speaker model and 8 speakers for testing. We trained on a 44257-utterance subset (40 hours) and evaluated on a held-out 100-utterance subset. For the score network, we used the Wavegrad architecture (Chen et al. 2020) so as to examine whether the superiority of BDDMs remains in a different dataset and with a different score network architecture.
Results are presented in Table 4. For this multi-speaker VCTK dataset, we obtained consistent observations with that for the single-speaker LJ dataset presented in the main paper. Again, the proposed BDDM with only 16 or 21 steps outperformed the DDPM with 1,000 steps. To the best of our knowledge, ours was the first work that reported this degree of superior. When reducing to 8 steps, BDDM obtained performance on par with (except for a worse PESQ) the costly grid-searched 8 steps (which were unscalable to more steps) in DDPM. For NE, we could again observe a degradation from its 16 steps to 21 steps, indicating the instability of NE for the VCTK dataset likewise. In contrast, BDDM gave continuously improved performance while increasing the step number.
C.2 Comparing different reverse processes for BDDMs
This section demonstrates that BDDMs do not restrict the sampling procedure to a specialized reverse process in Algorithm 4. In particular, we evaluated different reverse processes, including that of DDPMs as shown in Eq. (4) and DDIMs (Song et al. 2021a), for BDDMs and compared the objective scores on the generated samples. DDIMs (Song et al. 2021a) formulate a non-Markovian generative process that accelerates the inference while keeping the same training procedure as DDPMs. The original generative process in Eq. (4) in DDPMs is modified into
where is a sub-sequence of length of with , and is defined as its complement; Therefore, only part of the models are used in the sampling process.
To achieve the above, DDIMs defined a prediction function that depends on to predict the observation given directly:
By leveraging this prediction function, the conditionals in Eq. (71) are formulated as
where the detailed derivation of and can be referred to (Song et al. 2021a). In the original DDIMs, the accelerated reverse process produces samples over the subsequence of indexed by : . In BDDMs, to apply the DDIM reverse process, we use the predicted by the schedule network in place of a subsequence of the training schedule .
Finally. the objective scores are given in Table 5. Note that the subjective evaluation (MOS) is omitted here since the other assessments above have shown that the MOS scores are highly correlated with the objective measures, including STOI and PESQ. They indicate that applying BDDMs to either DDPM or DDIM reverse process leads to comparable and competitive results. Meanwhile, the results show some subtle differences: BDDMs over a DDPM reverse process gave slightly better samples in terms of signal error and consistency metrics (i.e., LS-MSE and MCD), while BDDM over a DDIM reverse process tended to generate better samples in terms of intelligibility and perceptual metrics (i.e., STOI and PESQ).
C.3 Unconditional image generation
For the unconditional image generation task, we evaluated the proposed BDDMs on the benchmark CIFAR-10 (32 32) dataset. The score functions, including those initially proposed in DDPMs (Ho et al. 2020) or DDIMs (Song et al. 2021a) and those pre-trained in the above third-party implementations, are all conditioned on a discrete step-index. We estimated the noise schedule in continuous space using the VGG11 schedule network and then mapped it to discrete time schedule using the approximation method in (Kong & Ping 2021).
Table 6 shows the performances of different sampling methods for DDPMs in CIFAR-10. By setting the maximum number of sampling steps () for noise scheduling, we can fairly compare the improvements achieved by BDDMs against related methods in the literature in terms of FID. Remarkably, BDDMs with 100 sampling steps not only surpassed the 1000-step DDPM baseline, but also produced the SOTA FID performance amongst all generative models using less than or equal to sampling steps.