SeqDiffuSeq: Text Diffusion with Encoder-Decoder Transformers

Hongyi Yuan, Zheng Yuan, Chuanqi Tan, Fei Huang, Songfang Huang

Introduction

Generative modeling is drawing more attention in recent years of machine learning research due to the development of diffusion models (Ho et al., 2020). Diffusion models define the forward process and the reverse process where the former gradually diffuses data to random noise while the latter recovers data from random noise iteratively, which have shown superior performance on synthesizing images (Rombach et al., 2021), audios (Kong et al., 2020), and videos (Ho et al., 2022) over other generative methods, such as generative adversarial network (GAN) (Goodfellow et al., 2014) and normalizing flow (Kobyzev et al., 2021).

It is not trivial to extend diffusion models to the generation of natural languages. Most of the existing diffusion models are applied to continuous feature space (Ho et al., 2020; Nichol and Dhariwal, 2021) while texts are sequences of discrete categorical tokens. Recently, research has explored categorical diffusion models in discrete space for text generation (Hoogeboom et al., 2021; Austin et al., 2022). There also exists research such as DiffusionLM (Li et al., 2022) that applies continuous diffusion models to word embedding. However, these works only focus on unconditional and controlled text generation.

Sequence-to-sequence text generation is a fundamental natural language processing setting and covers various practical downstream tasks, such as dialogue (Ni et al., 2021) and machine translation (Liu et al., 2020). In recent practice, researchers resort to auto-regressive (AR) (Dai et al., 2019) or non-auto-regressive (NAR) (Gu et al., 2019) Transformers for the tasks, and achieve good generation performance. Using diffusion models, a recent work named DiffuSeq (Gong et al., 2022) applies the diffusion-based method for sequence-to-sequence text generation. They deploy encoder-only Transformers and partially define diffusion and denoising processes on output sequences.

In this work, we explore diffusion models with encoder-decoder Transformer architecture for sequence-to-sequence generation. We propose SeqDiffuSeq which extends the continuous diffusion framework proposed in DiffusionLM (Li et al., 2022) to sequence-to-sequence settings. We equip SeqDiffuSeq with the self-conditioning technique (Chen et al., 2022) and our newly proposed adaptive noise schedule. Self-conditioning helps the model better capture the information from former iterations during the generation. The proposed adaptive noise schedule learns a token-level noise schedule to better control the amount of noise injected and information recovered during the forward and reverse process (Nichol and Dhariwal, 2021).

We conduct experiments on five generation tasks. Results show that SeqDiffuSeq achieves competitive performance compared with AR and NAR baselines in terms of generation quality and diversity. SeqDiffuSeq also shows improved generation performance and inference speed compared to text diffison model DiffuSeq. Ablation studies demonstrate that our model can benefit from self-conditioning and adaptive noise schedule techniques, and both are complementary to each other in sequence-to-sequence settings.

To summarize, the main contributions of this work are as follows:

We propose SeqDiffuSeq that extends the continuous text diffusion model to sequence-to-sequence text generation with encoder-decoder Transformer architecture.

The self-conditioning and newly proposed adaptive noise schedule technique can effectively improve the generation quality of the text diffusion model.

Experiments show SeqDiffuSeq achieves promising performance with the previous diffusion-based method DiffuSeq as well as AR and NAR models on five generation tasks.

Related Work

Since the great success of diffusion models in vision (Ho et al., 2020; Rombach et al., 2021; Song et al., 2021), researchers have explored extending diffusion models to text generation. Considering the discrete and categorical nature of texts, Multinomial Diffusion (Hoogeboom et al., 2021) and D3PM (Austin et al., 2021) are proposed for modeling categorical data. They define discrete diffusion models using discrete categorical transitions directly on texts. DiffusionBERT (He et al., 2022) follows D3PM and introduces pre-trained models for language modeling. Besides, recent research also explores converting texts into continuous features to adapt to diffusion models. Bit Diffusion (Chen et al., 2022) encodes discrete data as binary bits and treats these binary bits as real number features. Yu et al. (2022) is proposed to build text diffusion models in continuous latent space. DiffusionLM (Li et al., 2022) uses the word embedding space for continuous diffusion models and introduces auxiliary losses to enable joint learning of embedding and network parameters. Following DiffusionLM, recent research explores improving text generation quality (Strudel et al., 2022), and DiffuSeq (Gong et al., 2022) extends it to sequence-to-sequence settings. Compared to DiffuSeq, we propose a different model architecture and self-conditioning and adaptive noise schedule techniques to improve sequence-to-sequence generation performance.

Noise schedules in diffusion models control the level of noise injected and the level of information recovered in the forward and reverse process respectively. Previous research in vision and texts demonstrates that appropriate noise schedule design can improve the generation quality performance of diffusion models (Nichol and Dhariwal, 2021; Li et al., 2022). Concurrently, DiffusionBERT (He et al., 2022) proposes a spindle schedule for language modeling, and CDCD (Dieleman et al., 2022) designs a learned noise schedule for language modeling and machine translation. Different from both concurrent works, SeqDiffuSeq is proposed with a token-level noise schedule that balances the difficulty of denoising across time steps. Gao et al. (2023) proposes Difformer and is orthogonal to our work.

Preliminary

Diffusion model is generally formulated by a designed forward diffusion process and a learnt reverse denoising process. In the forward diffusion process, samples gradually mix with random noise, while in the reverse denoising process, the random noise is gradually denoised to generate synthetic samples. In this work, our diffusion model adopts the forward and reverse processes proposed in DDPM (Ho et al., 2020).

For the forward process, given a sample z0z_{0} from a real-world data distribution q(z0)q(z_{0}). At each time step t∈{1,2,⋯ ,T}t\in\{1,2,\cdots,T\}, a noise sample ztz_{t} is sampled from zt∼q(zt∣zt−1)=N(zt;αtzt−1,(1−αt)I)z_{t}\sim q(z_{t}|z_{t-1})=\mathcal{N}(z_{t};\sqrt{\alpha_{t}}z_{t-1},(1-\alpha_{t})I), where αt\alpha_{t} control the noise added at time step tt. In this regard, when TT is large enough, a real-world sample will gradually and ultimately diffuse to a standard Gaussian noise distribution.

For the reverse process, the diffusion model uses a learnt parameterized denoising distribution zt−1∼pθ(zt−1∣zt)z_{t-1}\sim p_{\theta}(z_{t-1}|z_{t}) to gradually recover samples from noise. The denoising distribution is parameterized by θ\theta and is to fit the posterior distribution q(zt−1∣zt,z0)q(z_{t-1}|z_{t},z_{0}) of the forward process. q(zt−1∣zt,z0)q(z_{t-1}|z_{t},z_{0}) can be derived as:

With learnt denoising distribution pθp_{\theta}, a synthetic real-world sample z0z_{0} can be generated from pure random noise zTz_{T} step-by-step.

Approach

In this section, we present the main design of our proposed SeqDiffuSeq for sequence-to-sequence language generation. The overview of SeqDiffuSeq is depicted in Figure 1. In the following sections, the input and output sequences are denoted as wxw_{x} and wyw_{y} respectively. For the ii-th token in wyw_{y}, the token is denoted as wyiw_{y}^{i}, where 0<i≤n0<i\leq n and nn represents the maximum output sequence word length. In order to avoid lengthy notations, we omit the indices referring to different data samples.

To fit diffusion models to sequence-to-sequence settings, we extend the text diffusion model, DiffusionLM (Li et al., 2022).

Reverse Process

Training Objective

We optimize θ\theta and embedding parameters by minimizing the variational bound of the data log-likelihood:

zθ0(zt,wx,t)z^{0}_{\theta}(z_{t},w_{x},t) is named the denoising function and predicts the estimated output embedding sequences at each reverse step tt. Then according to density functions qq and pθp_{\theta} following Gaussian distribution, the objective can be further simplified as:

where q(zt∣z0)=N(zt;αˉtz0,(1−αˉt)I)q(z_{t}|z_{0})=\mathcal{N}(z_{t};\sqrt{\bar{\alpha}_{t}}z_{0},(1-\bar{\alpha}_{t})I) for efficient sampling of ztz_{t} during training, and μT(z0)=αˉTz0\mu_{T}(z_{0})=\sqrt{\bar{\alpha}_{T}}z_{0}. We leave the detailed derivation to Appendix D. The training objective becomes to fit gϕ(wy)g_{\phi}(w_{y}) and the denoising function zθ0(zt,wx,t)z^{0}_{\theta}(z_{t},w_{x},t), which we can model with encoder-decoder Transformers architectures. During training, the sampling distribution qϕq_{\phi} contains trainable parameters of word embedding. We can backpropagate through this with reparameterization trick (Kingma and Welling, 2013).

Denoising with Encoder-Decoder Framework

Unlike DiffuSeq (Gong et al., 2022) using encoder-only Transformer architecture, we propose using an encoder-decoder Transformers architecture to model the input and output text sequences. For zθ0(zt,wx,t)z^{0}_{\theta}(z_{t},w_{x},t), we use the encoder to process the input sequences wxw_{x} and use the decoder to model the noisy output sequence ztz_{t}. Following the previous work (Li et al., 2022), we inject time step information tt by adding time step embedding to ztz_{t}. Using the encoder-decoder architecture has computational convenience during generation because the input sequences wxw_{x} only require one forward computation through the encoder network during the whole reverse process. Considering the reverse process requires thousands of iterations to generate the output sequences of high quality, the saving of computational resources can be significant.

During training and generation, the function zθ0z^{0}_{\theta} generates denoised samples at the sequence level. Therefore making predictions from the denoising function zθ0z^{0}_{\theta} resembles the non-autoregressive natural language generation. In this regard, we use a decoder with full attention matrices instead of causal attention matrices to model ztz_{t} at the sequence level.

2 Self-Conditioning

At each time step tt in the reverse process, the denoising function zθ0(zt,wx,t)z^{0}_{\theta}(z_{t},w_{x},t) makes output sequence predictions based on the noisier sample ztz_{t}. ztz_{t} is sampled from the former denoising distribution by mixing former sequence prediction z^0t=zθ0(zt+1,wx,t+1)\hat{z}_{0}^{t}=z^{0}_{\theta}(z_{t+1},w_{x},t+1), zt+1z_{t+1} and random noise. In this regard, part of the information contained in the former prediction z^0t\hat{z}_{0}^{t} is discarded. Bit-Diffusion (Chen et al., 2022) proposed the self-conditioning technique mitigating this waste of information by additionally taking former sequence predictions as inputs. The denoising function is formulated as zθ0(zt,z^0t,wx,t)z^{0}_{\theta}(z_{t},\hat{z}_{0}^{t},w_{x},t). Self-conditioning may enable the denoising function to refine the former sequence predictions rather than make new predictions from scratch. It is empirically verified that the self-conditioning technique can boost the performance of text diffusion models (Strudel et al., 2022).

To fit the technique into the Transformers modeling of zθ0z^{0}_{\theta} in our sequence-to-sequence setting, the sequence features z^0t\hat{z}_{0}^{t} from the former predictions are concatenated with noisier sequence features ztz_{t} on the embedding dimension. Hence, the dimension of input features of Transformer decoder becomes n×2dn\times 2d. Since the former sequences at time step tt are sampled successively from TT to tt which is computational-tedious during training, we take an efficient training scheme. With half probability, zθ0(zt,z^0t,wx,t)z^{0}_{\theta}(z_{t},\hat{z}_{0}^{t},w_{x},t) is trained by setting the input z^0t\hat{z}_{0}^{t} to . Otherwise, z^0t\hat{z}_{0}^{t} is first estimated by zθ0(zt,0,wx,t)z^{0}_{\theta}(z_{t},0,w_{x},t) and then is used for self-conditioning training. Under the second circumstance, we do not backpropagate through the first forward propagate estimated z^0t\hat{z}_{0}^{t}.

3 Adaptive Noise Schedule

In the domain of vision and audio, the generated sample quality (Nichol and Dhariwal, 2021) and likelihood estimation (Kingma et al., 2021) may potentially benefit from different appropriate time schedules. Previous research uses different simple functions such as linear function (Ho et al., 2020) or cosine function (Nichol and Dhariwal, 2021) of α\alpha against time step tt to design noise schedules. Such designs may results in unbalanced denoising difficulties for each step and lead to unsatisfying generation quality. Some works proposed to alleviate this problem by importance sampling (Li et al., 2022) or loss reweighing (Gong et al., 2022).

We propose a novel adaptive noise schedule at the token-level. Firstly, we propose to adaptively adjust the time schedules during training to make the denoising difficulties of zθ0z^{0}_{\theta} predicting output sequence increase linearly with respect to the time step. Secondly, we separately set adaptive noise schedule for different token positions, unlike previous text diffusion research that only designs noise schedules on the whole sequence level. Since the intrinsic features for embedding sequences are different across token positions within, we assume that for different token positions the expected noise schedules are different.

After initializing a noise schedule, we record the loss Lti\mathcal{L}_{t}^{i}We do not record the losses Lti\mathcal{L}_{t}^{i} for the padding tokens.. The mapping MiM_{i} is fitted after each training period. Ideally, the training losses should be monotonic with respect to the time step tt since for larger TT the input features ztz_{t} to the denoising function are noisier. However, overall time step TT is usually by thousands, hence this results in a fine-grained discretization of αˉi\bar{\alpha}^{i}. Due to the empirical loss estimation errors, training losses may not be monotonic between some successive time steps. To alleviate this issue and fit a smoother mapping MiM_{i}, we form a coarse-grained discretization ss for αˉi\bar{\alpha}^{i} and Li\mathcal{L}^{i}: Lsi=1K∑t=s×Ks×(K+1)Lti\mathcal{L}_{s}^{i}=\frac{1}{K}\sum_{t=s\times K}^{s\times(K+1)}\mathcal{L}_{t}^{i}, αˉsi=1K∑t=s×Ks×(M+1)αˉti\bar{\alpha}_{s}^{i}=\frac{1}{K}\sum_{t=s\times K}^{s\times(M+1)}\bar{\alpha}_{t}^{i}, s=⌊tK⌋s=\left\lfloor\frac{t}{K}\right\rfloor, where KK is the stride to evenly downsample tt and ⌊⋅⌋\lfloor\cdot\rfloor rounds the number down to it nearest integer.

With the learnt linear interpolation mapping αˉsi=Mi(Lsi)\bar{\alpha}_{s}^{i}=M_{i}(\mathcal{L}_{s}^{i}), we can obtain the adjusted discretized noise schedule αˉti,new\bar{\alpha}_{t}^{i,new} by αˉti,new=Mi(Lti,new)\bar{\alpha}_{t}^{i,new}=M_{i}(\mathcal{L}_{t}^{i,new}) where Lti,new\mathcal{L}_{t}^{i,new}’s are evenly taken between the minimum and maximum recorded values. As the training progresses, we adaptively calibrate the noise schedule αˉi\bar{\alpha}^{i} by repeating the above-mentioned procedure once per training updates. The pseudo-code for setting adaptive noise schedules during training is shown in Algorithm 1.

Experiments

We conduct experiments on six datasets across five different text generation tasks: Quora Question Pairs (QQP) (DataCanary et al., 2017) for Paraphrase Generation, Wiki-Auto (Jiang et al., 2020) for Text Simplification, Quasar-T (Dhingra et al., 2017) for Question Generation, Commonsense Conversation Dataset (CCD) (Zhou et al., 2018) for Dialogue Generation as well as the German(DE)-English(EN) pairs of IWSLT14 and WMT14 for Machine Translation. Detailed introductions and statistics of the datasets as shown in Appendix E.

2 Baselines

We consider three kinds of models as baselines. First, vanilla encoder-decoder Transformers and pre-trained GPT-2 are selected as strong AR baselines. Second, since SeqDiffuSeq denoises outputs at the sequence level, we compare it with an NAR baseline Levenshtein Transformer (LevT) (Gu et al., 2019). For machine translation, we also use CMLM (Ghazvininejad et al., 2019) which is an NAR translation method with iterative refinement as baselines. Besides, we compare it to other diffusion-based methods. DiffuSeq (Gong et al., 2022) is a recently proposed text diffusion model using an encoder-only Transformer structure. We also compare with concurrently proposed CDCD (Dieleman et al., 2022) on machine translation.

3 Implementation Details

We use a 6 layers encoder-decoder Transformer (Vaswani et al., 2017) with GeLU activation (Hendrycks and Gimpel, 2016). For the diffusion process, we set the maximum diffusion step TT to 2000, and use the sqrt schedule from DiffusionLM (Li et al., 2022) to initialize the adaptive time schedule. For translation tasks, we construct vocabulary using BPE (Sennrich et al., 2016). The vocabulary size is set to 10,000 for IWSLT14 and 32,768 for WMT14. For other tasks, we use the vocabulary of bert-base-uncased (Devlin et al., 2019).

For training of SeqDiffuSeq, we use a learning rate of 10−410^{-4} with 10,000 warm-up steps and a linearly-decreasing schedule. The proposed adaptive noise schedule is updated every 20,000 training steps and KK is set to 20. We explore maximum Bayes risk (MBR) decoding (Koehn, 2004) following previous research (Li et al., 2022) for improving generation quality during inference. Details on experiment settings and MBR are in Appendix F.

4 Main Results

To assess the generation quality of each model, we use BLEU (Papineni et al., 2002) and BERTScore (Zhang et al., 2020) as metrics. We also use distinct uni-gram (dist.1) to measure the word diversity within generated sentences. A high dist.1 score indicates fewer repeated words. For machine translation tasks, we additionally consider SacreBLEU (Post, 2018). The results are listed in Table 1. To better present the generation performance, we provide human evaluation results in Appendix I.

Primarily, for text generation quality, our proposed SeqDiffuSeq achieves much better performance measured by BLEU than DiffuSeq and other baselines with single generation on QQP, Wiki-Auto, and Quasar-T. On Wiki-Auto and Quasar-T, SeqDiffuSeq even achieves better performance with single generation than recently proposed DiffuSeq with MBR of 10 candidates. When incorporating with MBR, SeqDiffuSeq enjoys a boost of performance and achieves superior results over all baselines on QQP, Wiki-Auto, and Quasar-T. The performance is better than the pre-trained then fine-tuned GPT-2 with more parameters on Wiki-Auto and QQP. This indicates that SeqDiffuSeq can generate texts with good quality for sequence-to-sequence tasks (except CCD that all models have inferior performance). On translation tasks, the performance lags behind the AR Transformers baseline consistently across different datasets, while compared with NAR methods, SeqDiffuSeq consistently surpasses CMLM with 1 refinement iteration by 6.32 and 6.75 averaged points across four datasets without and with MBR. CMLM with 4 iterations has better performance. When comparing with CDCD, the performance with and without MBR are competitive on WMT14 EN-DE while the performance is worse on DE-EN. For diversity within sequences, texts generated by SeqDiffuSeq have fewer repeated words averagely than Transformers and DiffuSeq.

The largest improvement over MBR inference with 10 candidates is 1.06 BLEU score on QQP. The amount of this marginal improvement is consistent with concurrently proposed CDCD on WMT14. We will give more in-depth analyses of MBR in the following sections.

Analysis and Discussion

To verify the effectiveness of the proposed techniques in SeqDiffuSeq, we conduct ablation studies on QQP, Wiki-Auto, and IWSLT14. As shown in Table 2, after removing the adaptive noise schedule from SeqDiffuSeq and instead using the fixed sqrt schedule proposed in DiffusionLM (B\mathcal{B}), the performance drops consistently and the BLEU scores decrease by 2.29 on average. Without self-conditioning (C\mathcal{C}), the performance also degrades by 1.34 on average. By further removing adaptive noise schedule (D\mathcal{D}), the performance drops sharply by 5.71 on average and the largest drop in terms of BLEU is 8.43 on Wiki-Auto. Comparing adaptive noise schedule and self-conditioning technique, it is illustrated that our proposed adaptive noise schedule brings larger improvement and two techniques are complementary to each other.

2 Time Schedule

It is verified in the ablation study that the proposed adaptive noise schedule can improve sequence-to-sequence text generation. On the IWSLT14 DE-EN dataset, we visualize the adaptive noise schedules as well as the loss at each time step with and without adaptive noise schedule. For the adaptive noise schedule, we plot αˉti\bar{\alpha}_{t}^{i} at different token positions ii against the diffusion time step tt. And for losses, we plot averaged training losses Lti\mathcal{L}^{i}_{t} at each position ii against time step tt. Depicted in Figure 2, the dashed line in the first sub-figure shows the sqrt schedule, while the other lines represent the noise schedules at different token positions. The figure shows that the adaptive noise schedules deviate from the sqrt schedule. At both ends of time steps, the adaptive noise schedules are flatter compared to sqrt schedule, especially for tokens at larger position orders. Besides, adaptive noise schedules are diverse for different positions, although the trends along time steps are similar. For the token positions at larger orders, the noise schedule lines move toward the lower-left direction. Therefore, at each time step, the tokens at earlier positions have smaller noise than later positions. The information of tokens on the left is recovered earlier at each step. SeqDiffuSeq resembles the left-to-right generation of texts. Through a case study in Appendix J, the phenomenon is also verified.

Comparing the second and third sub-figures, the losses Li\mathcal{L}^{i} with adaptive noise schedule increase linearly with respect to time steps as expected. At each time step, the losses at earlier token positions are smaller, indicating earlier tokens are easier to generate for SeqDiffuSeq . More visualizations on other datasets are listed in Appendix H.

3 Inference Speed

We compare SeqDiffuSeq with DiffuSeq in terms of inference time in Table 3. Our SeqDiffuSeq achieves 3.56 times acceleration generating one batch of text samples. The acceleration mainly originated from: (1) SeqDiffuSeq only requires forward computation of encoder once, while DiffuSeq needs to run forward computation for the input sequences for each diffusion step; (2) At each time step, SeqDiffuSeq only models the output sequence, while DiffuSeq has to model the concatenation of both input and output sequences.

4 MBR Inference

It is shown in Table 1 that MBR with 10 candidates improves DiffuSeq to more than 6 BLEU score, while improves SeqDiffuSeq by 1.06 BLEU score on QQP. In Figure 3, we plot SacreBLEU scores and Diverse 4-gram (Div.4) scores (Deshpande et al., 2018) against MBR candidate numbers. Div.4 measures the proportion of distinct 4-grams in a set of generated sequences. A higher Div.4 score means better sequence-level diversity by different generation runs. The figure shows that the self-conditioning technique and adaptive noise schedule make the text diffusion model generate less diverse sequences, and the single generated sequence will have higher quality with both techniques. Self-conditioning technique and adaptive noise schedule deliver a trade-off between generation quality and generation diversity. With both techniques, MBR inference is needless to generate high-quality samples for SeqDiffuSeq resulting in a more efficient generation procedure. We also propose a new sampling scheme to compensate the marginal MBR improvements for SeqDiffuSeq which is discussed in detail in Appendix G.

Conclusion

In this work, we explore to approach sequence-to-sequence text generation with continuous diffusion models. We propose SeqDiffuSeq which uses an encoder-decoder Transformers architecture to learn the denoising function. In order to improve text generation performance, the denoising function in SeqDiffuSeq is integrated with self-conditioning technique. SeqDiffuSeq also includes a newly proposed adaptive noise schedule which makes the denoising difficulty evenly distributed across all time steps and assigns exclusive noise schedules for tokens from different positional orders. Through experiments, we illustrate the superior performance of SeqDiffuSeq in terms of generation quality and inference speed and provide insights into our proposed adaptive noise schedule technique.

References

Appendix A Limitation

Diffusion models generate high-quality synthetic samples through thousands of iterations in the reverse process. Thousands of reverse process iterations require a huge amount of forward propagation computation of Transformers model which is computationally costly, although we save nearly four times of computational budget for one forward computation compared to the previous diffusion-based model DiffuSeq. In the domain of vision synthetic, there exists research to profoundly reduce the time step needed for generation (Song et al., 2021). Reducing the reverse steps for text generation would be a promising direction for future research.

As shown in the discussion, equipping text diffusion models with self-conditioning and adaptive noise schedules can profoundly increase the generation quality. However, such quality improvement is at the cost of generation diversity under different random seeds. This leads to marginal MBR inference improvements. Although we propose a compensation discussed in Appendix G. The in-depth discussion on improving SeqDiffuSeq generation diversity is left to future research.

Appendix B Ethic Statements and Boarder Impact

The datasets and baseline models used in our research are publicly available. Diffusion models, previously successful in vision, face challenges in NLP due to discrete token sequences. Promising results have been shown in DiffusionLM (Li et al., 2022) and DiffuSeq (Gong et al., 2022), but both works use encoder-only models and have limitations in scalability and efficiency. This research explores and improves the diffusion-based sequence-to-sequence text generation models. Our work alters to encoder-decoder Transformers which are widely applied in recent LLMs such as FLAN-T5 (Chung et al., 2022) for better scalability, potential, and sampling speed acceleration (Section 6.3). Our work also incorporates novel techniques like self-conditioning and adaptive noise schedules, outperforming several AR and NAR baselines. SeqDiffuSeq demonstrates the feasibility of encoder-decoder diffusion models for sequence-to-sequence tasks and may serve as a starting point for future exploration of text diffusion models’ potential, serving as another method approaching sequence-to-sequence text generation besides widely implemented AR and NAR models. Considering the excellent performance of diffusion models in other domains such as vision, text diffusion models have great potential in generating text sequences with high quality and may be an emerging framework of text generation.

Appendix C Derivation of Posterior

Given zt∼q(zt∣zt−1)=N(zt;αtzt−1,(1−αt)I)z_{t}\sim q(z_{t}|z_{t-1})=\mathcal{N}(z_{t};\sqrt{\alpha_{t}}z_{t-1},(1-\alpha_{t})I), we can reparameterize zt=αtzt−1+1−αtϵtz_{t}=\sqrt{\alpha_{t}}z_{t-1}+\sqrt{1-\alpha_{t}}\epsilon_{t}. Then, recursively,

Therefore, q(zt∣z0)=N(zt;αˉtz0,(1−αˉt)I)q(z_{t}|z_{0})=\mathcal{N}(z_{t};\sqrt{\bar{\alpha}_{t}}z_{0},(1-\bar{\alpha}_{t})I). According to Bayes rule, we have:

since q(zt∣zt−1)q(z_{t}|z_{t-1}) and q(zt−1∣z0)q(z_{t-1}|z_{0}) are all Gaussian distributed, we will have:

Appendix D Derivation of Training Objective

We present the detailed derivation of training objective following Ho et al. (2020); Li et al. (2022). As mentioned in main texts, the forward process successively perturbs the real-world sample z0z_{0} with random noise, where z0z_{0} gradually changes to zTz_{T} for a TT-time step diffusion process. zTz_{T} can be approximately regarded as pure random noise which follows standard Gaussian distribution in our case. We define the forward process as follows:

where αt\alpha_{t} controls the noise level at each time step tt.

For the reverse process, we learn a parameterized denoising distribution pθ(zt−1∣zt,wx,t)p_{\theta}(z_{t-1}|z_{t},w_{x},t). By successively sampling from pθp_{\theta}, a synthetic real-world sample z0z_{0} can be recovered from pure random noise zTz_{T}.

The training objective of diffusion model is to minimize the negative likelihood of data distribtuion, which is:

By Bayes rule, we can derive the posterior distribution of qq with respect to zt−1z_{t-1}:

We substitute q(zt∣zt−1),∀t>1q(z_{t}|z_{t-1}),\forall t>1 in Equation 16 with Equation 18:

Appendix E Datasets

We conduct experiments on following datasets. The data statistics and licenses are shown in Table 4 and 5.

Quora Question Pairs (QQP) (DataCanary et al., 2017) is a paraphrase identification dataset. We use the positive pairs as the paraphrase generation task. The models need to generate a restatement expressing the same meaning to the given sentence. Wiki-Auto (Jiang et al., 2020) is a text simplification dataset to revise a complex text with simplified grammar and word choices. The dataset aligns sentences between English Wikipedia and Simple English Wikipedia with automatic pre-processing and identifying procedure. Quasar-T (Dhingra et al., 2017) is a question-answering dataset containing trivia questions paired with answers and contexts. We use the dataset for evaluating question generation which aims to generate related questions with given contexts. We use the pre-processed data from Lin et al. (2018) following Gong et al. (2022). Commonsense Conversation Dataset (CCD) (Zhou et al., 2018) is extracted from single-round dialogues on Reddit and is used for evaluating open domain dialogue generation. The task requires generating feedback with commensense knowledge given the dialogue contexts. IWSLT14 and WMT14 are both widely used benchmarks for machine translation. We use the German(DE)-English(EN) pairs for both directions of translation. We follow fairseq (Ott et al., 2019) for data pre-processing using Moses script (Koehn et al., 2007) and tokenizing the sentences with byte-pair encoding (BPE) (Sennrich et al., 2016).

Appendix F Implementation Details

Here we give details for the implementation details of our experiments. For the Transformers structure and model training, we list detailed design in Table 6. For all the tasks, the set the maximum training step to 1000,000 and save checkpoints every 10,000 steps. We select the best checkpoint on the development set. For WMT14 task, we use batch size 1024 while for other tasks we use batch size 128. For training on each datasets, we train for one run on NVIDIA A100 GPUs with 80GB memory. For inference, we set the maximum time step to T=2000T=2000, and we do not use the clamping trick as proposed in DiffusionLM (Li et al., 2022), since the clamping trick does not consistently improve the generation quality across datasets.

F.2 Details on MBR

Following DiffusionLM (Li et al., 2022), we apply Minimum Bayes Risk (MBR) decoding for one single generation output with improved quality. For each sample, MBR decoding uses a generated sequences candidate set C\mathcal{C} and finds the candidate sequence s∗s^{*} that minimize a expected risk RR:

where r(⋅,⋅)r(\cdot,\cdot) is a specific risk function and we use the negative BLEU score following DiffusionLM and sequence candidates in the candidate set C\mathcal{C} are generated from the diffusion models under different random seeds.

Appendix G Sampling by Prior

Since at each time step tt, the Transformers denoising function zθ0z_{\theta}^{0} models the prediction z^0t\hat{z}_{0}^{t} of target output sequences. In the reverse process, sampling zt−1z_{t-1} is according to the denoising distribution pθp_{\theta} as:

However, we can also use the prior distribution qq in the forward process to generate zt−1z_{t-1}, which is:

Comparing to generation by Equation 28, using Equation 29 theoretically have larger variance.

because βt1−αˉt=1−αt1−αˉt≤1\frac{\beta_{t}}{1-\bar{\alpha}_{t}}=\frac{1-\alpha_{t}}{1-\bar{\alpha}_{t}}\leq 1 where αt<1,∀t\alpha_{t}<1,\forall t and αˉt=∏s=1tαs\bar{\alpha}_{t}=\prod_{s=1}^{t}\alpha_{s}.

To increase the sequence level diversity, we experiment with randomly replacing the denoising distribution pθp_{\theta} by high variance distribution in Equation 29 in the reverse process during generation. We denote the replacing probability as p1p_{1}.

Besides, considering the variance difference between the two sampling distribution are larger at earlier time step in the reverse process, we also explore to only replace the sampling distribution in the first p2p_{2} percent of time steps. We generate 10 candidate output sentences for each sample under different random seeds to compute Div.4 and SacreBLEU scores.

As shown in the left subfigure of Figure 4, when fixing the replacing probability to 0.05, the generation diversity are consistently and profoundly improved. In the right subfigure, the generation quality consistently degrades when replacing the denoising distribution when generation, even though the replacing probability is low. In the middle subfigure, we can see that although the generation quality degrades for each candidate, the final output sequences by MBR may improve with proper p2p_{2}. In Figure 5, we can get similar results when fixing p2=0.5p_{2}=0.5. In the middle subfigure of Figure 5, the final output sequences are consistently better with different p1p_{1}.

To conclude, it is shown that replacing the sampling distribution from the denoising distribution pθp_{\theta} to the prior distribution qq can provide a trade-off between the generation diversity and generation quality. With a proper combination of p1p_{1} and p2p_{2}, the generation quality of SeqDiffuSeq with the aid of MBR can be further improved. The benefits of sampling with the prior distribution qq are always neglected in previous research.

Appendix H More Results on Adaptive Noise Schedule

We present more visualizations of the learned adaptive noise schedules and the losses for each time step on other datasets. Figure 6, 7 and 8 present the visualizations on IWSLT14 EN-DE, QQP, and Wiki-Auto respectively with the same arrangement as Figure 2. The results from the figures are consistent with those discussed in the main texts.

Appendix I Human Evaluation

To better demonstrate the performance of the proposed SeqDiffuSeq , we conduct human evaluations to compare the generated results of SeqDiffuSeq to those of DiffuSeq on the paraphrasing task QQP dataset. We randomly sample 100 data points in the test sets and let annotators decide for the same input sequence, which generated text sequence is better, worse, or of similar quality. We compare SeqDiffuSeq with the previous state-of-the-art text diffusion model DiffuSeq. For fairness, the human evaluations are designed to be blind evaluations (i.e., the annotators are unaware of which model the output sequence is related to).

The human annotators are graduate university students who are proficient in English and are asked to compare the generated sequences based on the following instruction. Decide which generated output sequence is better based on whether the one is more consistent with the input question, whether the one has higher grammatical and syntactic quality. Figure 9 shows the human evaluation results.

The results show that both annotators prefer the generated output sequences by SeqDiffuSeq more. Generated output sequences on QQP from SeqDiffuSeq win by 36% and 44% from two annotators, while those from DiffuSeq only win by 24% and 30% respectively. Human evaluation results show that SeqDiffuSeq can generate text sequences of higher quality than DiffuSeq.

Appendix J Case Study

We select three illustrative cases and investigate the generation process of SeqDiffuSeq. From the cases, it shows that SeqDiffuSeq can generate reasonable text sequences. The generation process reveals that

1. SeqDiffuSeq decides the output sequence length by generating [SEP] tokens at the early stage of sampling;

2. The generation process seems to follow a left-to-right refining order;

3. The position of [SEP] token will not change during sampling, even though there exists token repetition in the generated sequences as shown in red.