Text Generation with Diffusion Language Models: A Pre-training Approach with Continuous Paragraph Denoise

Zhenghao Lin, Yeyun Gong, Yelong Shen, Tong Wu, Zhihao Fan, Chen Lin, Nan Duan, Weizhu Chen

Introduction

Text generation is a crucial task in natural language processing, which aims to produce fluent and coherent texts for various applications. Previous text generation methods mainly relied on recurrent neural networks (RNNs) (Pawade et al., 2018; Song et al., 2018; Gu et al., 2016a; Qi et al., 2021), which generate texts sequentially from left to right. However, RNNs suffer from issues such as long-term dependency and exposure bias. Recently, Transformer (Vaswani et al., 2017b), a self-attention based neural network, has emerged as the dominant paradigm for text generation, thanks to its ability to capture global dependencies and leverage large-scale pre-trained language models (Qi et al., 2020; Lewis et al., 2019; Raffel et al., 2020a). Transformer-based methods typically adopt an encoder-decoder architecture, where the encoder maps the input text to a sequence of hidden vectors, and the decoder generates the output text either autoregressively (AR) or non-autoregressively (NAR). AR decoding is more accurate but slower, as it predicts each word conditioned on the previous ones. NAR decoding is faster but less precise, as it predicts all words simultaneously without modeling the dependencies among them.

In this paper, we present a new text generation approach, called GENIE, that integrates the diffusion model and Transformer-based method. The diffusion model is a generative model that reverses a stochastic process of adding noise to the data, and has shown promising results in image (Ho et al., 2020; Song et al., 2020), molecule (Hoogeboom et al., 2022), video (Ho et al., 2022), and text (Li et al., 2022b; Gong et al., 2022; Strudel et al., 2022; Reid et al., 2022) generation. GENIE follows the encoder-decoder architecture, where the encoder transforms the input text to hidden vectors, and the diffusion model restores the output text from a random Gaussian noise, guided by the encoder hidden vectors. The diffusion model iterates over multiple time steps, and gradually denoises the output text at each step.

To leverage the large-scale unlabeled text data, we also propose an end-to-end pre-training method for GENIE. Unlike the existing pre-training tasks that involve masking or splitting tokens or texts (Qi et al., 2020; Lewis et al., 2019; Raffel et al., 2020a), we design a novel pre-training task, called continuous paragraph denoise (CPD). CPD requires the model to predict the noise added to continuous paragraphs in the current time step, given the paragraph context and the noisy paragraph information.

We evaluate GENIE on four popular text generation benchmarks: XSum (Narayan et al., 2018), CNN/DailyMail (Hermann et al., 2015), Gigaword (Rush et al., 2015), and CommonGen (Lin et al., 2019). The experimental results demonstrate that GENIE achieves competitive performance with Transformer-based AR methods, and that the proposed pre-training method can effectively improve the performance. We notice that GENIE has achieved significant improvements in diversity metric. To evaluate the multiple outputs of the generation model, we design an automatic annotation method based on large language model. We also conduct ablation studies to analyze the impact of the diffusion steps and pre-training steps.

The main contributions of this work are summarized as follows:

We propose GENIE, the first large-scale language pre-trained model based on the diffusion framework, which can generate high-quality texts for sequence-to-sequence tasks.

We introduce a novel CPD loss as the pre-training objective, which can enhance the model’s ability to denoise noisy texts and capture the paragraph-level coherence.

We validate the effectiveness of the pre-trained diffusion model on downstream tasks, and design a new automatic annotation method for the evaluation based on large language model. We also provide extensive analyses on the model’s behavior and properties.

Preliminary

In the classical sequence-to-sequence task, given a source text s={w1s,w2s,…,wns}{\bm{s}}=\{{w}^{s}_{1},{w}^{s}_{2},\dots,{w}^{s}_{n}\} with nn tokens, it generates target text sequence y={w1y,w2y,…,wny}{\bm{y}}=\{{w}^{y}_{1},{w}^{y}_{2},\dots,{w}^{y}_{n}\}. A sequence generation model can achieve this by modeling the conditional probability: p(y∣s)p\left({\bm{y}}\mid{\bm{s}}\right). \commentGiven a source text s={w1s,w2s,…,wns}{\bm{s}}=\{{w}^{s}_{1},{w}^{s}_{2},\dots,{w}^{s}_{n}\} with nn tokens, our goal is to produce a target text y={w1y,w2y,…,wny}{\bm{y}}=\{{w}^{y}_{1},{w}^{y}_{2},\dots,{w}^{y}_{n}\}. A non auto-regressive (NAR) generation model can achieve this by:

2 Diffusion model

In the diffusion model, the diffusion process can be regarded as a discrete-time Markov process. The diffusion process starts with initial state x0{\bm{x}}_{0} at time step t=0t=0, where x0{\bm{x}}_{0} is the Gaussian distribution of the original data. It gradually adds Gaussian noises to x0{\bm{x}}_{0} in the forward diffusion process according to a variance schedule β1,...,βT\beta_{1},...,\beta_{T}. At the time step t+1t+1, the latent variable xt+1{\bm{x}}_{t+1} is only determined by the xt{\bm{x}}_{t} at time tt, expressed as:

As tt increases, xt{\bm{x}}_{t} becomes closer to standard Gaussian noise N(xT;0,I)\mathcal{N}({\bm{x}}_{T};0,\mathbf{I}).

The diffusion model simulates a Markov process of gradually adding Gaussian noise to a variable x0{\bm{x}}_{0} over time steps t=0,1,…,Tt=0,1,\ldots,T. At each time step t+1t+1, the latent variable xt+1{\bm{x}}_{t+1} is sampled from a Gaussian distribution with mean 1−βt+1xt\sqrt{1-\beta_{t+1}}{\bm{x}}_{t} and covariance βt+1I\beta_{t+1}\mathbf{I}, where βt+1∈(0,1)\beta_{t+1}\in(0,1) controls the noise level:

As tt increases, xt{\bm{x}}_{t} becomes closer to standard Gaussian noise N(xT;0,I)\mathcal{N}({\bm{x}}_{T};0,\mathbf{I}).

The diffusion model learns to perform the inverse diffusion process during generation, which predicts the noise given the current state xt{\bm{x}}_{t} at time step tt. The previous state xt−1{\bm{x}}_{t-1} can be reconstructed by subtracting the noise and rescaling the mean. Thus, the distribution of xt−1{\bm{x}}_{t-1} given xt{\bm{x}}_{t} is a Gaussian with mean μθt−1\mu^{t-1}_{\theta} and variance σθt−12{{\sigma}^{t-1}_{\theta}}^{2}:

where αt=1−βt\alpha_{t}=1-\beta_{t}, αˉt=∏i=1tαi\bar{\alpha}_{t}=\prod_{i=1}^{t}\alpha_{i} and zθz_{\theta} is predicted by a neural network parameterized by θ\theta. The diffusion model is trained by minimizing the mean squared error between μθt−1\mu^{t-1}_{\theta} and the true mean μ^t−1\hat{\mu}_{t-1}, which is computed from the reverse conditional distribution q(xt−1∣xt,x0)q({\bm{x}}_{t-1}|{\bm{x}}_{t},{\bm{x}}_{0}):

Following the variational lower bound (VLB) approach (Ho et al., 2020), the diffusion model can be trained by minimizing the loss function:

Model

GENIE is the proposed diffusion language model for pretraining, it adopts the sequence-to-sequence framework as illustrated in Figure 1. GENIE could generate a high-quality text sequence y{\bm{y}} given a source text s{\bm{s}}, such as producing y:{\bm{y}}: Messi’s performance from s:{\bm{s}}: In the World Cup 2022, [MASK] won people’s praise.. To achieve this, GENIE leverages two components: a bidirectional encoder model and a cross-attention diffusion model. The encoder model encodes the source text s{\bm{s}} into a set of hidden vectors Hs=Encoder(s){\bm{H}}_{{\bm{s}}}=\text{Encoder}({\bm{s}}), which indicates the distributed representation of s{\bm{s}}. The diffusion model takes Hs{\bm{H}}_{{\bm{s}}} and a Gaussian noise as inputs, and iteratively refines the data by applying a sequence of denoising operations. In contrast to the traditional autoregressive text generation paradigm, which generates one token at a time, the diffusion model in GENIE outputs the sequence of embeddings in parallel at each denoising step, making GENIE a non-autoregressive generation (NAR) model.

The encoder in GENIE is a 6-layer transformer model which takes the source text s{\bm{s}} as input with bidirectional self-attention. Specifically, given a source text sequence s={w1s,w2s,…,wns}{\bm{s}}=\{{w}^{s}_{1},{w}^{s}_{2},\dots,{w}^{s}_{n}\} with nn tokens, the encoder model computes the vector hi{\bm{h}}_{i} for each token wi{w}_{i}. Thus, the source text s{\bm{s}} can be represented as Hs{\bm{H}}_{{\bm{s}}} by the encoder model:

Language Diffusion Model

The diffusion model in GENIE is a 6-layer transformer with cross-attention on the source text representation Hs{\bm{H}}_{{\bm{s}}}. It learns to predict Gaussian noise zθ(xt,t,Hs)z_{\theta}\left({\bm{x}}_{t},t,{\bm{H}}_{{\bm{s}}}\right) conditioned on the current diffusion step tt and the state xt{\bm{x}}_{t}, where xt{\bm{x}}_{t} is the continuous latent representation of the target text. We use an embedding function and a clamping trick to ground the continuous state xt{\bm{x}}_{t} with discrete target tokens, which will be elaborated in the following section.

Inference Phase

To generate text from the diffusion model, we start from the final step t=Tt=T and sample a state xT{\bm{x}}_{T} from a standard Gaussian distribution. Then we iteratively generate the noise for the previous step using equations 4 and 5, and subtract it from the current state to obtain xt−1{\bm{x}}_{t-1}. After arriving at t=0t=0, we apply the clamping trick (Li et al., 2022b) to replace the values of x0{\bm{x}}_{0} with its closest word embeddings, and then decode the discrete tokens from x0{\bm{x}}_{0}.

Training Phase

To train the diffusion model for sequence-to-sequence tasks, we first convert the target sequence y={w1y,w2y,…,wny}{\bm{y}}=\{{w}^{y}_{1},{w}^{y}_{2},\dots,{w}^{y}_{n}\} into a continuous state x0{\bm{x}}_{0} using the embedding function with a additional Gaussian noise permutation, which can be expressed as:

where Emb(⋅)\text{Emb}(\cdot) is embedding function, β0\beta_{0} represents the scaling of variance at time step t=0t=0. Then we apply the forward diffusion process (equation 2) to obtain the state xt{\bm{x}}_{t} at any step tt as a function of x0{\bm{x}}_{0}, as shown in equation:

where αˉt=∏i=1tαi\bar{\alpha}_{t}=\prod_{i=1}^{t}\alpha_{i}. In the training phase, we sample a random step tt to calculate xt{\bm{x}}_{t}, and then use the denoising architecture to predict the noise for that step, based on the cross-attention with the source representation Hs{\bm{H}}_{s}. The mean and variance of the predicted noise are given by equations 12:

where zθz_{\theta} is the output of the denoising architecture and θ\theta are its parameters. The training objective is to minimize the squared error between the predicted and true noise, as well as the reconstruction error between x0{\bm{x}}_{0} and the target embeddings, as expressed in equation 13:

where pθ(y∣x0)=∏i=1npθ(wiy∣x0)p_{\theta}({\bm{y}}|{\bm{x}}_{0})=\prod^{n}_{i=1}p_{\theta}({w}_{i}^{y}|{\bm{x}}_{0}), represents mapping the continuous latent variable x0{\bm{x}}_{0} into the discrete space token wiy{w}^{y}_{i}.

1 Pre-training GENIE

Diffusion models have great potential for natural language generation (NLG) due to their ability to produce diverse outputs. However, they have been largely overlooked in NLG because of their slow convergence and low quality compared to autoregressive models. In this section, we address these challenges by pre-training a diffusion language model and introducing a novel pre-training task tailored for it. The novel pre-training task we propose is called continuous paragraph denoise (CPD). CPD aims to train the model to predict the noise added to a continuous paragraph in the current diffusion step, given the paragraph and its surrounding context.

Specifically, given a document d={w1d,w2d,…,wld}{\bm{d}}=\{{w}^{d}_{1},{w}^{d}_{2},\dots,{w}^{d}_{l}\} with ll words, we randomly select a paragraph p={w1p,w2p,…,wmp}{\bm{p}}=\{{w}^{p}_{1},{w}^{p}_{2},\dots,{w}^{p}_{m}\} from d{\bm{d}}, where m=⌊γ∗l⌋m=\lfloor\gamma*l\rfloor is the paragraph length and γ\gamma is a predefined ratio. We mask the paragraph in the document with a special token ([MASK]), and feed the masked document d′={w1d′,w2d′,…,[MASK],…,wl−md′}{\bm{d}}^{\prime}=\{{w}^{d^{\prime}}_{1},{w}^{d^{\prime}}_{2},\dots,\text{[MASK]},\dots,{w}^{d^{\prime}}_{l-m}\} to the GENIE encoder. We also apply the forward diffusion process to the paragraph p{\bm{p}} and obtain a noisy version xt{\bm{x}}_{t} at a random step tt, and feed it to the GENIE denoising architecture. The denoising architecture then uses the cross-attention with the source representation Hs{\bm{H}}_{s} to predict the noise for the current step, using equations 12. In summary, the pre-training objective of CPD is to minimize the same loss as in equation 13, except that y{\bm{y}} is replaced by p{\bm{p}} and x0{\bm{x}}_{0} is the embedded paragraph with noise.

Through this pre-training task, the diffusion model can enhance its semantic understanding of the continuous text and its denoising ability at each diffusion step. Moreover, the CPD task is self-supervised and does not rely on external labelled data sources, so it can fully exploit the information in the original pre-trained corpus.

Experiments and Results

In this section, we will introduce the details of GENIE pre-training, the data setting, and show extensive experimental results on various NLG downstream tasks.

Our model uses a 6-layer transformer as the encoder, and a 6-layer cross attention transformer as denoising architecture. In particular, in denoising architecture, we use the randomly embedding function to map discrete token into continuous variable. We set latent variable dim to 768 and embedding dim to 128.

Pre-training Data

Resent works has shown that pre-training on large scale corpus can improve the performance of the model on downstream tasks (Lewis et al., 2019; Qi et al., 2020), which is also applicable to GENIE based on diffusion model. Following BART (Lewis et al., 2019), we use pre-training data consisting of 160Gb of news, books, stories, and web text. We segment sentences belonging to different chapters, and ensure that the input text length does not exceed 512.

Pre-training Setting

We use the CPD task mentioned in § 3.1 to pre-train GENIE on large-scale corpus. The proportion of continuous paragraph γ\gamma sets to 30%, hence, for the 512 length input, the target length is 153. We randomly extract 153 length targets from the text input, and leave [MASK] token at the extracted position. In the training process, we use Adam optimizer (Kingma & Ba, 2015) with learning rate 1e-4, and we set the batch size to 512. We pre-trained our model on 8 × 40GB NVIDIA A100 GPUs with 5 million steps, lasting for 50 days. In the fine-tuning phase, we use the final pre-training model checkpoint to conduct fine-tuning on various downstream tasks.

2 Fine-tune on Downstream Tasks

In order to verify the effectiveness of pre-training on GENIE based on diffusion model, we fine-tune and verify the effect of GENIE on a variety of downstream tasks. Through the above task, we can prove that the pre-trained GENIE can quickly adapt to different types of NLG tasks without long time training like other diffusion models.

As an important task in the NLG field, text summarization aims to summarize long documents into fluent short texts. In the experiment, we selected three widely used datasets: (a) Gigaword corpus (Rush et al., 2015), (b) CNN/DailyMail (Hermann et al., 2015), and (c) XSum (Narayan et al., 2018). In the process of fine-tuning, we set the learning rate to 5e-5 and the 120K training steps for all three dataset. In the inference process, we randomly sample 10 Gaussian noises for iteration denoising, and use the highest score as the final generated result. For different sample number, please refer to Appendix C. During evaluation, we following the existing work (Lewis et al., 2019; Qi et al., 2020), reporting F1 scores of ROUGE-1, ROUGE-2, and ROUGE-L on test set.

Common Sense Generation

Common sense generation tasks require the model have the ability of generative commonsense reasoning. Specifically, given a series of common sense concepts, the model needs to generate coherent statements based on these concepts that adhere to real-world scenarios. We select the widely used dataset CommonGen (Lin et al., 2019) to evaluate whether GENIE has good creativity and reasoning ability in natural language generating. In the fine-tuning phase, we set the learning rate to 1e-4 and train 10k steps in total. Finally, we randomly sampled 10 gaussian noises and selected the best sample as the final result. Referring to the previous work (Lin et al., 2019), we reported the indicators including F1 scores of ROUGE-2/L, BLEU-3/4, CLDEr, and SPICE.

3 Baselines

We compare GENIE with the baselines of several mainstream methods. Specifically, these methods can be divided into two groups. The first group is the NAR model, including NAT (Gu et al., 2017), iNAT (Lee et al., 2018), NAG-BERT (Su et al., 2021), CMLM (Ghazvininejad et al., 2019), LevT (Gu et al., 2019b), ConstLeven (Susanto et al., 2020), BANG (Qi et al., 2021), ELMER (Li et al., 2022a) and InsT (Stern et al., 2019). Among them, InsT, iNAT, CMLM, LevT, ConstLeven and BANG can also be used in Semi-NAR, which can optimize the generation quality through multiple NAR iterations. It is worth noting that GENIE also belongs to the Semi-NAR model.

The second group is AR model, the model of encoder-decoder structure including LSTM (Greff et al., 2017), Transformer (Vaswani et al., 2017a), bRNN-CopyNet (Gu et al., 2016a), Trans-CopyNet (Lin et al., 2019), MeanPooling-CopyNet (Lin et al., 2019) without pre-training, and strong baselines MASS (Song et al., 2019), BART (Lewis et al., 2019), T5 (Raffel et al., 2020b), BANG (Qi et al., 2021), and ProphetNet (Qi et al., 2020) with large scale pre-training. For large scale pre-training models mentioned above, we select the base version of the model, which is equivalent to the total number of GENIE parameters.

4 Main Results

We present the results of GENIE and the baselines on XSum, CNN/DailyMail, Gigaword, CommonGen in Table 1, Table 2, and Table 3. Our results demonstrate that the pre-trained GENIE is a powerful NAR model for text generation. Especially on the Xsum dataset, GENIE outperforms other NAR and Semi-NAR methods by a large margin, and on all three text summarization datasets, GENIE achieves comparable quality to the pre-trained AR model. In addition, GENIE shows creativity and logic in common sense generation tasks. On CommonGen, GENIE surpasses other baseline models, including T5 which has been pre-trained on a large-scale corpus.

We also compare the pre-trained GENIE and GENIE trained from scratch (w/o pre-train). As shown in Table 1 and Table 2, pre-training significantly improves the ROUGE-1, ROUGE-2, ROUGE-L scores of GENIE on the three text summarization datasets. Similarly, the results on CommonGen in Table 3 indicate that pre-training enhances the performance of GENIE on this task. These results confirm the effectiveness of our pre-training method.

5 Generate Diversity Comparison

With the emergence of the diffusion based model such as GENIE, the advantages of text generation in diversity will be gradually valued. In this experiment, we will use both quantitative metrics and qualitative examples to show the richness of GENIE in text generation.

To measure the diversity of GENIE generation, we use SELF-BLEU as the metric. The lower the SELF-BLEU score, the more diverse the generated texts are. For comparison, we use BART, a state-of-the-art autoregressive model, which is pre-trained on large scale corpora. For BART, we apply different decoding methods of autoregressive models, such as greedy search, beam search (Xiao et al., 2022), diverse beam search (diversity strength = 0.8) (Vijayakumar et al., 2016), typical sampling (τ=1.2\tau=1.2) (Meister et al., 2022), top-k sampling (k=50k=50) (Fan et al., 2018), and nucleus sampling (p=0.92p=0.92) (Holtzman et al., 2020). These decoding methods can generate multiple texts from the same source sequence. In this experiment, we generate 10 different target sequences for each source sequence using GENIE and BART. Then we use the 10 summaries generated from XSum, CNN/DailyMail, and Gigaword to calculate the SELF-BLEU scores.

As shown in Table 4, although the diversity of autoregressive generation can be slightly improved by using diverse beam search or some sampling methods with BART, the improvement is not significant. On the other hand, the diversity of generation is greatly enhanced by using the GENIE. The large gaps in SELF-BLEU indicate that GENIE can generate more diverse texts, not just varying a few words.

To complement the quantitative metrics, we also provide a case study in Appendix A to analyze the quality of the texts generated by BART and GENIE. We find that the autoregressive generation method can produce high-quality texts when there is only one output, but when generating multiple outputs, even with different decoding methods, it is hard to increase its diversity, and there may be many repeated prefixes. In contrast, the diffusion generation method can maintain the quality of generation while offering rich diversity.

However, it may not be fair to compare GENIE directly with the single reference to prove that GENIE can achieve diversity without compromising quality. Therefore, we design a new evaluation method. We use text-davinci-003 version of InstructGPT (Ouyang et al., 2022), which is based on the large language model (LLM) GPT-3.5, to score our generated texts, that is, to evaluate the quality of the generated summaries. Specifically, we first obtain the sample set (10 summaries generated by BART using diverse beam search and 10 summaries generated by GENIE), and design a prompt to input into text-davinci-003 to score the generated summaries, while counting the number of high-quality summaries within the 10 summaries generated by BART and GENIE respectively. We conduct the experiment on the three different text summarization datasets and use two evaluation methods, Average Summary Score represents the average score given by text-davinci-003, ranging from 1 to 3, and Average High-quality Summary represents the average number of high-quality summaries in 10 samples, ranging from 0 to 10. For more detailed experimental settings, please refer to Appendix B.

As shown in Table 5, although GENIE’s scores are slightly lower than BART’s, according to the results in Table 4, the diversity of samples generated by BART is much lower than GENIE. Given the trade-off between diversity and quality, the score difference is within the acceptable range. Moreover, the result of Average High-quality Summary shows that there are still enough high-quality summaries in the case of high diversity. Such advantages of GENIE deserve our attention and further exploration in our future work.

6 Impact of Pre-training Steps

Our pre-training method and the diffusion model itself are both designed to achieve long-term convergence and unlimited potential, but they also require a large amount of pre-training time. Here we investigate how the pre-training steps affect the performance of our model compared with a non-pre-trained GENIE on the XSum dataset. We fine-tune the checkpoints obtained at 1 million step intervals from pre-training and evaluate them using 5 random Gaussian noises, selecting the highest score as the final result. As shown in Table 6, pre-training for only 1 million steps can significantly improve the quality of generation over the non-pre-trained GENIE. Moreover, we can see from the results that pre-training continues to steadily boost the performance of the GENIE on the downstream task as the pre-training steps increase.

7 Impact of Pre-training Parameters

In this subsection, we examine the effect of important pre-training parameters on the pre-training performance. First, in the unsupervised pre-training method CPD, we need to explore how the proportion of continuous paragraphs γ\gamma influences the pre-training performance. We vary the value of γ\gamma from 15% to 40% (with 5% intervals) and conduct 2 million steps of pre-training for each value. After the pre-training, we evaluate the pre-training effect by fine-tuning on Xsum. For a rigorous evaluation, we sample 5 Gaussian noises, repeat the experiment 5 times with different random seeds, and report the mean and standard deviation of the results, each time choosing the highest score as the final result. As shown in Table 7, too large or too small values of γ\gamma lead to instability and poor performance of the pre-trained model. Pre-training is more stable and effective when γ=30%\gamma=30\%.

Second, we investigate the time step sampling method used in the pre-training. Before each training step, we need to sample a time step as part of the model input. The existing two common time step sampling methods are uniform sample and loss aware sample. The former assigns equal probabilities to each time step, while the latter updates the sampling weights according to the training loss, so that more important time steps have higher chances of being sampled. In the experiment, we use these two sampling methods, test them on two different values of γ\gamma (15% and 30%), and perform a rigorous evaluation similar to the previous experiment. As shown in Table 8, we observe that under 2 million steps of pre-training, the uniform sample outperforms the loss aware sample for different values of γ\gamma. Intuitively, although the loss aware sample can speed up the convergence of the diffusion model, we hope that the model can learn sufficient knowledge at each time step during the pre-training, so that it can converge faster and perform better on downstream tasks.

8 Impact of Diffusion Time Step

The number of diffusion time steps has a great impact on the quality of generation. We explore how the GENIE performs under different numbers of inverse diffusion steps on the XSum dataset. Assuming that the total number of diffusion steps T=2000T=2000, we set the interval step of inverse diffusion to 1, 2, 4, 8, 20, and the corresponding numbers of inverse diffusion steps are 2000, 1000, 500, 250, 100. In this experiment, we sample 5 Gaussian noises and choose the best denoising result. As shown in Figure 2, we can clearly see that when the number of inverse diffusion steps is small, the quality of generation with GENIE deteriorates significantly. As the number of inverse diffusion steps increases to 1000, the generation quality of GENIE becomes stable.

Related Work

Recently, a major breakthrough has been made in the model of pre-training on large scale corpus. As unidirectional language models, GPT (Radford et al., 2018), GPT2 (Radford et al., 2019) modeling the text based on left-to-right, and predict the next token according to the token appearing on the left. At the same time, bidirectional language models, which uses bidirectional encoder to model text, can obtain better context sensitive representation, such as BERT (Devlin et al., 2019) and RoBERT (Liu et al., 2019). RoBERT optimizes pre-training tasks compared to BERT, both of which significantly improve the ability of natural language understanding. In order to improve the performance of the large scale pre-training model in natural language generation, some works has designed pre-training tasks based on the standard framework of sequence-to-sequence. MASS (Song et al., 2019) lets the model predict the short masked token span step by step, while ProphetNet (Qi et al., 2020) predict more words in each step to ease local over fitting.

2 Diffusion Models for Text

In recent years, diffusion model has achieved great success in the domains of image generation (Ramesh et al., 2022; Saharia et al., 2022; Rombach et al., 2022). Because of its amazing generation quality, some works apply diffusion model in text generation domains. Diffsuion-LM (Li et al., 2022b) maps discrete tokens into continuous latent variable, achieving more complex controllable text generation through continuous diffusion. In the field of text revision where non-autoregressive method is widely used, DiffusER (Reid et al., 2022) also uses the diffusion model to implement the edit based generative processes. DiffuSeq (Gong et al., 2022) achieves conditional text generation with a new method which controlled information is also involved in the diffusion process. Different from the above work, we build a novel language model based on the diffusion model for the first time, using the standard enocoder-decoder framework. For our best knowledge, we are the first to adopt large scale pre-training on the language model based on the diffusion model.

Conclusion

In this paper, we have presented a novel diffusion language model GENIE, which leverages a large scale corpus for pre-training. Our model adopts a sequence-to-sequence framework, where a bidirectional encoder encodes the source sequence and a denoising decoder predicts and removes noise from the target sequence in a non auto-regressive fashion. This design allows us to generate diverse text by gradually refining the output from a noisy initial state. Moreover, we have introduced a new pre-training method called continuous paragraph denoise, which aims to denoise whole paragraphs as the target sequence. Our experiments on various NLG tasks demonstrate that GENIE can produce high-quality and diverse text, and validate the benefits of pre-training our diffusion model on a large scale corpus.

References

Appendix A Case Study

In the section§ 4, we have made a rigorous analysis of the quality and diversity of GENIE. We hope that by comparing diffusion model with traditional autoregressive generation models, we can find the potential of diffusion model in natural language generation tasks. Nowadays, most of the excellent language generation models belong to autoregressive generation models, but at the same time, we also need some new ideas and generation paradigm to make natural language generation not limited to autoregressive. The ways of natural language generation need to be diversified to broaden researchers’ thinking, just as the application of diffusion model in natural language generation can bring rich diversity to the generated text. We are excited that the diversity of the content generated by the diffusion model does not come from a large number of wrong words or unrelated texts, but from different sentence patterns and different information obtained from the original text. This shows us the future prospects of the diffusion model in the natural language generation task.

In order to more intuitively show the quality of the text generated by the diffusion model and the autoregressive generation model, we selected two samples from the text summarization dataset XSum in Table 9. For each sample, we used GENIE and BART to generate three summaries respectively, of which the BART generation method is diversity beam search. For display purposes, the source sequence has been intercepted and abbreviated to reduce the length. It can be seen from the generated summaries that if we only observe one sentence of the generated summary, the summary generated by the autoregressive model BART is of good quality and is related to the content of the source sequence, the generated text is relatively fluent due to the autoregressive generation mode. But if we look at the three generated summaries, the text generated by the autoregressive model BART obviously has a lot of duplication. In the second example, ”there should be no ’stagnancy’ in the size of the Welsh forest estate” has repeated descriptions in two different summaries. In the first example, there are even two summaries that are almost identical. The difference is only one meaningless full stop at the end. Although we use the diversity beam search generation method when generating, it is difficult to let the autoregressive model jump out of the essence of iterative generation to generate creative text.

In contrast, we can see that the summary generated by GENIE may not be as fluent as BART due to the non-autoregressive generation mode if only from the quality of single sentence generation. Once multiple summaries are generated, we can observe GENIE’s creative generating ability. In the first example, GENIE describes the medical project in the source sequence from three related direction. Health information collection, project information and project objectives are mentioned in the generated summary. In the second example, ”the forestry minister” mentioned a variety of measures on the forest industry, and the three summaries generated by GENIE described different parts respectively, including the formulation of public plans, cooperation with Welsh companies to promote forest management, and the involvement of private enterprises in the forest industry. Compared with the single information ”there should be no ’stagnancy’ in the size of the Welsh forest estate” provided by BART, although BART also provides a concise summary, it is difficult to conclude that BART’s summary is better, because people are accustomed to understanding some problems from multiple perspectives in practice, rather than a conclusive conclusion.

After analyzing the above examples, we can observe the great potential of new language model GENIE. In the practical application of text generation, diversified generation results can be used in many scenarios. The unique generation method of diffusion model brings new ideas to text generation, and also lets us consider whether the single-label text generation really meets our needs. We do not need the diffusion model to be superior to the autoregressive generation model in all aspects. What we need is a new idea to bring more possibilities to text generation. We believe that the diffusion model can be widely used in text generation in the future.

Appendix B Large Language Model Evaluation

Recently, the large language model has been widely used in various tasks with its amazing performance. In this paper, we use the large language model to evaluate the quality of the generated summary. We select text-davinci-003 as the evaluation model in our experiment, the most important thing is the construction of the prompt which will input into the model.

As prompt example shown in Table 10, we divide the prompt into three parts. The first part is the text of the original article, the middle part is the evaluation requirements, and the end part is the summary that needs to be evaluated. Finally, we can get output score through the large language model. For each summary, we need to organize the above prompt for the model input, but the evaluation requirements for each summary are the same. During the test, we asked the model to give the following score to the summary: 1 represents bad, 2 represents neutral, 3 represents goods. After getting the score of each summary, we will further count the number of high-quality samples in the 10 samples generated. Here, we define high-quality summary as summary with a score equal to 3. Finally, we can summarize and integrate all the scores obtained, and count the average summary score and the average number of high-quality summaries.

Appendix C Impact of Sample Number

Compared with the autoregressive model, GENIE based on the diffusion model can generate more diverse texts in the inference phase, even under the guidance of the source sequence. Different Gaussian noises sampled during denoising can often lead to completely different generation results. This method is more flexible, but it is not conducive to the evaluation against a single reference answer. However, as the number of Gaussian noises sampled increases, the generated text has a higher probability of approaching the single reference answer, and the corresponding evaluation score is higher. To this end, we test the performance of the model on the test set under different numbers of samples on the XSum dataset. As shown in Figure 3, we evaluate the results of 5, 10, 15 and 20 samples. We can observe that as the number of samples increases, the more likely the generated sample is to be similar to the original label. The improvement of the similarity is more noticeable in the early stage of increasing the number of samples.