A Contrastive Framework for Neural Text Generation

Yixuan Su, Tian Lan, Yan Wang, Dani Yogatama, Lingpeng Kong, Nigel Collier

Introduction

Open-ended neural text generation with Transformer is an indispensable component in various natural language applications, such as story generation , contextual text completion , and dialogue systems . However, the conventional approach of training a language model with maximum likelihood estimation (MLE) and decoding the most likely sequence is often not sufficient . Specifically, this modelling formulation often leads to the problem of degeneration, i.e., the generated texts from the language model tend to be dull and contain undesirable repetitions at different levels (e.g., token-, phrase-, and sentence-level) . To alleviate this problem, previous solutions modify the decoding strategy by sampling from less likely vocabularies . While reducing the generated repetition, these sampling methods introduce another critical problem (semantic inconsistency)—the sampled text tends to diverge from or even contradict to the original semantics defined by the human-written prefix . Another approach addresses the degeneration problem by modifying the model’s output vocabulary distribution with unlikelihood training .

In this work, we argue that the degeneration of neural language models stems from the anisotropic distribution of token representations, i.e., their representations reside in a narrow subset of the entire space . In Figure 1(a), we showcase a cosine similarity matrix of token representations (taken from the output layer of the Transformer) produced by GPT-2. We see that the cosine similarities between tokens within a sentence are over 0.95, meaning that these representations are close to each other. Such high similarity is undesirable as it can naturally cause the model to generate repetitive tokens at different steps. In an ideal setting, the token representations should follow an isotropic distribution, i.e., the token similarity matrix should be sparse and the representations of distinct tokens should be discriminative as shown in Figure 1(b). Moreover, during the decoding process, the sparseness of the token similarity matrix of the generated text should be preserved to avoid model degeneration.

Based on the above motivations, we present SimCTG (a simple contrastive framework for neural text generation) that encourages the model to learn discriminative and isotropic token representations. We also present a novel decoding strategy to complement SimCTG, contrastive search. The key intuitions behind contrastive search are: (i) at each decoding step, the output should be selected from the set of most probable candidates predicted by the model to better maintain the semantic coherence between the generated text and the human-written prefix, and (ii) the sparseness of the token similarity matrix of the generated text should be preserved to avoid degeneration.

We conduct comprehensive experiments on three widely used benchmarks. We show that our approach is generalizable to different tasks and different languages (section 4 and section 5) as well as different model sizes (section 4.3 and Appendix D). Specifically, the experimental results verify that SimCTG improves the intrinsic qualities of the language model, as evaluated by perplexity and token prediction accuracy (section 4.2 and Appendix D). Moreover, we demonstrate that the proposed contrastive search significantly outperforms previous state-of-the-art decoding methods in both human and automatic evaluations (section 4 and section 5). Furthermore, we provide in-depth analyses to get better insights on the inner-workings of our proposed approach (section 6).

Background

The goal of language modelling is to learn a probability distribution pθ(x)p_{\theta}(\boldsymbol{x}) over a variable-length text sequence x={x1,...,x∣x∣}\boldsymbol{x}=\{x_{1},...,x_{|\boldsymbol{x}|}\}, where θ\theta denotes model parameters. Typically, the maximum likelihood estimation (MLE) objective is used to train the language model which is defined as

However, as observed in many recent studies , training with likelihood maximization objective often yields an anisotropic distribution of model representations (especially for Transformer-based models) that undermines the model’s capacity.

2 Open-ended Text Generation

In this work, we focus on studying the task of open-ended text generation due to its generality in various applications, such as story generation , contextual text completion , poetry generation , and dialogue systems . Formally, conditioned on a human-written prefix (i.e., context) x\boldsymbol{x}, the task is to decode a continuation x^\hat{\boldsymbol{x}} from the language model and the resulting text is {x1,..,x∣x∣,x^∣x∣+1,...,x^∣x∣+∣x^∣}\{x_{1},..,x_{|\boldsymbol{x}|},\hat{x}_{|\boldsymbol{x}|+1},...,\hat{x}_{|\boldsymbol{x}|+|\hat{\boldsymbol{x}}|}\}. Typically, there are two classes of methods used for decoding, which are (1) deterministic methods and (2) stochastic methods.

Deteriminstic Methods. Two widely used deterministic approaches are greedy and beam search which aim to select the text continuation with highest probability based on the model’s probability distribution pθp_{\theta}. However, solely maximizing the output probability often leads to dullness and degeneration in the generated text.

Stochastic Methods. To remedy the issues of deterministic decoding, several approaches have been proposed to sample from pθp_{\theta}. To avoid sampling from the unreliable tail of distribution, Fan et al. proposed top-kk sampling which draws sample from the vocabulary subset V(k)V^{(k)} that maximizes ∑v∈V(k)pθ(v∣x)\sum_{v\in V^{(k)}}p_{\theta}(v|\boldsymbol{x}). Here, ∣V(k)∣=k|V^{(k)}|=k and x\boldsymbol{x} is the prefix context. Differently, the current state-of-the-art nucleus sampling draws sample from the smallest vocabulary subset UU with total probability mass above a threshold p∈p\in; i.e., UU is the smallest vocabulary subset such that ∑v∈Upθ(v∣x)≥p\sum_{v\in U}p_{\theta}(v|\boldsymbol{x})\geq p. While the sampling approaches help to alleviate model degeneration, the intrinsic stochasticity in these methods could cause the semantic meaning of the sampled text to diverge from or even contradict to the human-written prefix .

Methodology

In this section, we first present how to apply contrastive learning to calibrate the representation space of the language model. Then, we introduce our proposed contrastive search decoding algorithm.

Our goal is to encourage the language model to learn discriminative and isotropic token representations. To this end, we introduce a contrastive objective LCL\mathcal{L}_{\textup{CL}} into the training of the language model. Specifically, given a variable-length sequence x={x1,...,x∣x∣}\boldsymbol{x}=\{x_{1},...,x_{|\boldsymbol{x}|}\}, the LCL\mathcal{L}_{\textup{CL}} is defined as

where ρ∈\rho\in is a pre-defined margin and hxih_{x_{i}} is the representation of token xix_{i} produced by the model. The similarity function ss computes the cosine similarity between token representations as

Intuitively, by training with LCL\mathcal{L}_{\textup{CL}}, the model learns to pull away the distances between representations of distinct tokens.By definition, the cosine similarity s(hxi,hxi)s(h_{x_{i}},h_{x_{i}}) of the identical token xix_{i} is 1.01.0. Therefore, a discriminative and isotropic model representation space can be obtained. The overall training objective LSimCTG\mathcal{L}_{\textup{SimCTG}} is then defined as

where the maximum likelihood estimation (MLE) objective LMLE\mathcal{L}_{\textup{MLE}} is described in Eq. (1). Note that, when the margin ρ\rho in LCL\mathcal{L}_{\textup{CL}} equals to , the LSimCTG\mathcal{L}_{\textup{SimCTG}} degenerates to the vanilla MLE objective LMLE\mathcal{L}_{\textup{MLE}}.

2 Contrastive Search

We propose a novel decoding method, contrastive search. At each decoding step, the key ideas of contrastive search are (i) the generated output should be selected from the set of most probable candidates predicted by the model; and (ii) the generated output should be discriminative enough with respect to the previous context. In this way, the generated text can (i) better maintain the semantic coherence with respect to the prefix while (ii) avoiding model degeneration.

Formally, given the previous context x<t\boldsymbol{x}_{<t}, at time step tt, the selection of the output xtx_{t} follows

where V(k)V^{(k)} is the set of top-kk predictions from the model’s probability distribution pθ(⋅∣x<t)p_{\theta}(\cdot|\boldsymbol{x}_{<t}) and kk is typically set as 3∼\sim10. In Eq. (5), the first term, model confidence, is the probability of candidate vv predicted by the model. The second term, degeneration penalty, measures how discriminative of candidate vv with respect to the previous context x<t\boldsymbol{x}_{<t} and ss is defined in Eq. (3). Specifically, it is defined as the maximum cosine similarity between the representation of vv and that of all tokens in x<t\boldsymbol{x}_{<t}. Here, the candidate representation hvh_{v} is computed by the model given the concatenation of x<t\boldsymbol{x}_{<t} and vv. Intuitively, a larger degeneration penalty of vv means it is more similar to the context, therefore more likely leading to model degeneration. The hyperparameter α∈\alpha\in regulates the importance of these two components. When α=0\alpha=0, contrastive search degenerates to the greedy search method.

Document Generation

We first evaluate our approach on the task of open-ended document generation.

Model and Baselines. Our proposed approach is architecture-agnostic and can be applied to any generation model. In this work, we evaluate our method on the representative GPT-2 model . Specifically, we fine-tune GPT-2 on the evaluated benchmark (detailed below) with the proposed objective LSimCTG\mathcal{L}_{\textup{SimCTG}} (Eq. (4)) and generate the text continuation with different decoding methods. We perform experiments using the base model (117M parameters) which consists of 12 Transformer layers with 12 attention heads.In Appendix D, we demonstrate the experimental results of our approach on other language models. We compare our approach with two strong baselines: (1) GPT-2 fine-tuned with the standard MLE objective (Eq. (1)); and (2) GPT-2 fine-tuned with unlikelihood objective .The unlikelihood baseline is implemented with the official code, which can be found at https://github.com/facebookresearch/unlikelihood_training. Our implementation is based on the Huggingface Library .

Evaluation Benchmark. We conduct experiments on the Wikitext-103 dataset which contains a large collection of Wikipedia articles with over 100 million words and 260 thousands unique tokens. Wikitext-103 is a document-level dataset and has been widely used for the evaluation of large-scale language modelling .

Training. For our SimCTG and the MLE baseline, we fine-tune the models on Wikitext-103 for 40k training steps. For the unlikelihood baseline, following Welleck et al. , we first fine-tune the model with the token-level unlikelihood objective for 38.5k steps and then with the sequence-level unlikelihood objective for 1.5k steps. Therefore, the overall training steps of all compared methods are the same. The batch size is set as 128 and the training samples are truncated to a maximum length of 256. We optimize the model with Adam optimizer and a learning rate of 2e-5.

Decoding. We evaluate the models by producing text continuations given the prefixes from the test set. In the experiments, the lengths of the prefix and the generated continuation are set as 32 and 128, respectively. We test different models with various decoding methods. For deterministic method, we use greedy search and beam search with a beam size of 1010. For stochastic method, we use the current state-of-the-art nucleus sampling with p=0.95p=0.95. For the proposed contrastive search, the kk and α\alpha in Eq. (5) are set as 88 and 0.60.6.In Appendix E, we provide detailed ablation studies on the effect of both kk and α\alpha in contrastive search. The hyperparameters of different methods are selected based on their optimal MAUVE (detailed in section 4.1.2) performance on the validation set.

We perform evaluation from two aspects: (1) language modelling quality which measures the intrinsic quality of the model; and (2) generation quality which measures the quality of the generated text.

Following Welleck et al. , we report the results of the model on the metrics below.

Perplexity. The model perplexity (ppl) on the test set of Wikitext-103.

Prediction Accuracy. It is defined as: acc=1∑x∈D∣x∣∑x∈D∑t=1∣x∣\mathds1[arg max⁡pθ(x∣x<t)=xt]\textup{{acc}}=\frac{1}{\sum_{\boldsymbol{x}\in\mathcal{D}}|\boldsymbol{x}|}\sum_{\boldsymbol{x}\in\mathcal{D}}\sum_{t=1}^{|\boldsymbol{x}|}\mathds{1}[\operatorname*{arg\,max}p_{\theta}(x|\boldsymbol{x}_{<t})=x_{t}], where D\mathcal{D} is the Wikitext-103 test set, x<t\boldsymbol{x}_{<t} is the prefix, and xtx_{t} is the reference token at time step tt.

Prediction Repetition. The fraction of next-token (top-1) predictions that occur in the prefix which is defined as: rep=1∑x∈D∣x∣∑x∈D∑t=1∣x∣\mathds1[arg max⁡pθ(x∣x<t)∈x<t]\textup{{rep}}=\frac{1}{\sum_{\boldsymbol{x}\in\mathcal{D}}|\boldsymbol{x}|}\sum_{\boldsymbol{x}\in\mathcal{D}}\sum_{t=1}^{|\boldsymbol{x}|}\mathds{1}[\operatorname*{arg\,max}p_{\theta}(x|\boldsymbol{x}_{<t})\in\boldsymbol{x}_{<t}].

In addition, the next token repetitions that do not equal to the ground truth token: wrep=1∑x∈D∣x∣∑x∈D∑t=1∣x∣\mathds1[arg max⁡pθ(x∣x<t)∈x<t ∧≠xt]\textup{{wrep}}=\frac{1}{\sum_{\boldsymbol{x}\in\mathcal{D}}|\boldsymbol{x}|}\sum_{\boldsymbol{x}\in\mathcal{D}}\sum_{t=1}^{|\boldsymbol{x}|}\mathds{1}[\operatorname*{arg\,max}p_{\theta}(x|\boldsymbol{x}_{<t})\in\boldsymbol{x}_{<t}\>\wedge\neq x_{t}] is also reported.

Generation Repetition. This metric measures the sequence-level repetition as the portion of duplicate nn-grams in the generated text . For a generated text continuation x^\hat{\boldsymbol{x}}, the repetion at nn-gram level is defined as: rep-n=100×(1.0−∣unique n-grams(x^)∣∣total n-grams(x^)∣)\textup{{rep-n}}=100\times(1.0-\frac{|\textup{unique n-grams}(\hat{\boldsymbol{x}})|}{|\textup{total n-grams}(\hat{\boldsymbol{x}})|}).

Diversity. This metric takes into account the generation repetition at different nn-gram levels and it is defined as: diversity=∏n=24(1.0−rep-n100)\textup{{diversity}}=\prod_{n=2}^{4}(1.0-\frac{\textup{rep-n}}{100}). It can be deemed as an overall assessment of model degeneration. A lower diversity means a more severe degeneration of the model.

MAUVE is a metric that measures the token distribution closeness between the generated text and human-written text. A higher MAUVE score means the model generates more human-like texts.

Semantic Coherence. To automatically measure the semantic coherence (i.e., consistency) between the prefix and the generated text, we employ the advanced sentence embedding method, SimCSE . Specifically, given the prefix x\boldsymbol{x} and the generated text x^\hat{\boldsymbol{x}}, the coherence score is defined as: coherence=vx⊤vx^/(∥vx∥⋅∥vx^∥)\textup{{coherence}}=v_{\boldsymbol{x}}^{\top}v_{\hat{\boldsymbol{x}}}/(\|v_{\boldsymbol{x}}\|\cdot\|v_{\hat{\boldsymbol{x}}}\|), where vx=SimCSE(x)v_{\boldsymbol{x}}=\textup{SimCSE}(\boldsymbol{x}) and vx^=SimCSE(x^)v_{\hat{\boldsymbol{x}}}=\textup{SimCSE}(\hat{\boldsymbol{x}}).

Perplexity of Generated Text. Lastly, we evaluate the perplexity of the generated text x^\hat{\boldsymbol{x}} given the prefix x\boldsymbol{x}, which is defined as: gen-ppl=2f(D,θ)\textup{{gen-ppl}}=2^{f(\mathcal{D},\theta)} and f(D,θ)=1∑x∈D∣x^∣∑x∈Dlog⁡2pθ(x^∣x)f(\mathcal{D},\theta)=\frac{1}{\sum_{\boldsymbol{x}\in\mathcal{D}}|\hat{\boldsymbol{x}}|}\sum_{\boldsymbol{x}\in\mathcal{D}}\log_{2}p_{\theta}(\hat{\boldsymbol{x}}|\boldsymbol{x}). Importantly, the optimal approach should produce text which has a perplexity close to that of the human-written text . A high gen-ppl means the generated text is very unlikely given the prefix, therefore being low quality. In contrastive, a low gen-ppl means the generated text has a low diversity and gets stuck in repetitive loops . We use the model θ\theta trained with LSimCTG\mathcal{L}_{\textup{SimCTG}} to measure the gen-ppl of different approaches, therefore making sure the numbers are comparable with each other.We obtain similar gen-ppl results and can draw the same conclusion when using the model trained with MLE and Unlikelihood. Therefore, we only include the results acquired by the model trained with LSimCTG\mathcal{L}_{\textup{SimCTG}} in Table 1. We refer to Appendix F for the gen-ppl results obtained by the MLE and Unlikelihood models.

2 Results

The experimental results on Wikitext-103 are shown in Table 1.

Language Modelling Quality. From the results, we observe that SimCTG achieves the best perplexity and next token accuracy. The reason is that, with more discriminative representations, SimCTG is less confusing when making next token predictions, leading to the improved model performance. On the rep and wrep metrics, the unlikelihood model yields the best result but at the expense of unfavorable performance drops in the perplexity and next token accuracy.

Generation Quality. Firstly, on the rep-n and diversity metrics, SimCTG + contrastive search obtains the best result, suggesting it best addresses the degeneration problem. Secondly, the MAUVE score demonstrates that SimCTG + contrastive search generates texts that are closest to human-written texts in terms of token distribution. Thirdly, among all methods, SimCTG + contrastive search is the only approach that achieves over 0.6 coherence score, showing it produces semantically consistent text with respect to the prefix. Lastly, the gen-ppl metric also validates the superiority of SimCTG + contrastive search as it obtains notably better generation perplexity comparing with other approaches.

Moreover, from the results of MLE and Unlikelihood baselines, we see that contrastive search still brings performance boost as compared with greedy and beam search. However, the performance gain still lags behind SimCTG, which demonstrates the necessity of contrastive training. The underlying reason is that, without using the contrastive objective LCL\mathcal{L}_{\textup{CL}} (Eq. (2)), the token representations obtained by MLE or Unlikelihood are less discriminative (section 6.1). Therefore, the degeneration penalty (Eq. (5)) of different candidates are less distinguishable and the selection of output is dominated by the model confidence, making contrastive search less effective.

3 Human Evaluation

We also conduct a human evaluation with the help of graders proficient in English from a third-party grading platform. We randomly select 200 prefixes with length of 32 from the test set of Wikitext-103. For each prefix, we use different models (MLE, Unlikelihood, and SimCTG) with two decoding methods (nucleus sampling and contrastive search) to generate text continuations with length of 128. To examine the generality of our approach across different model sizes, we include a large size SimCTG (i.e., SimCTG-large) which is obtained by fine-tuning the GPT-2-large model that consists of 36 Transformer layers with 20 attention heads. All generated results, plus the reference text, are randomly shuffled and evaluated by five graders, which results in 9,000 annotated samples in total. The evaluation follows a 5-point Likert scale (1, 2, 3, 4, or 5) for each of the following features:We refer to Appendix G for more details of human evaluation.

Coherence: Whether the generated text is semantically consistent with the prefix.

Fluency: Whether the generated text is fluent and easy to understand.

Informativeness: Whether the generated text is diverse and contains interesting content.

Table 2 presents the human evaluation results, with the first row showing strong inter-annotator agreements as measured by Fleiss\textprime\textprime kappa coefficient . Firstly, we see that, directly applying contrastive search with MLE or Unlikelihood model does not yield satisfactory results. This is due to the anisotropic nature of their representation space as discussed in Section section 4.2. Secondly, the coherence score of Unlikelihood model is notably lower than MLE and SimCTG, suggesting it generates the most unlikely results which is also shown by its generation perplexity (gen-ppl) in Table 1. Furthermore, the results of SimCTG + contrastive search significantly outperforms nucleus sampling with different models in terms of coherence and fluency (Sign Test with p-value ¡ 0.05). Lastly, SimCTG-large + contrastive search achieves the best performance across the board and even performs comparably with human-written text on the fluency metric (Sign Test with p-value ¿ 0.4). This reveals the clear generalization ability of our approach to large size models and future work could focus on extending it to models that contain over billions of parameters such as GPT-3 .

Open-domain Dialogue Generation

To test the generality of our approach across different tasks and languages, we then evaluate our method on the task of open-domain dialogue generation. In this task, given a multi-turn dialogue context (where each turn is an user utterance), the model is asked to generate an adequate response that is semantically consistent with the context. Here, the dialogue context is deemed as the prefix.

Benchmark and Baselines. We conduct experiments on two benchmark datasets from two languages (i.e., Chinese and English). For the Chinese benchmark, we use the LCCC dataset . For the English Benchmark, we use the DailyDialog dataset .

We compare the GPT-2 models fine-tuned with SimCTG and MLE.We acknowledge that there are other GPT-like models (e.g., Zhang et al. and Thoppilan et al. ) that are designed for dialogue generation. We leave the test of our approach on these models to our future work. Specifically, for the Chinese benchmark (i.e., LCCC), we use a publicly available Chinese GPT-2 .https://huggingface.co/uer/gpt2-chinese-cluecorpussmall Same as in Section section 4, during training, we use a batch size of 128 and truncate the training samples to a maximum length of 256. On the LCCC dataset, we train (i.e., fine-tune) the models for 40k steps. As for the DailyDialog dataset, due to its smaller dataset size, we train the models for 5k steps. For optimization, we use Adam optimizer and a learning rate of 2e-5.

For each model, we use four decoding methods, including (1) greedy search; (2) beam search (beam size of 1010); (3) nucleus sampling (p=0.95p=0.95); and (4) contrastive search (k=5k=5, α=0.6\alpha=0.6).

Evaluation. We rely on human evaluation to assess the model performance. Same as in Section section 4.3, we randomly select 200 dialogue contexts from the test set and ask five annotators to evaluate the generated responses plus the reference response in three dimensions: (i) coherence, (ii) fluency; and (iii) informativeness. The scores follow a 5-point Likert scale (1, 2, 3, 4, or 5).

Table 3 shows the evaluation results where the first row shows strong inter-annotator agreements as measured by Fleiss\textprime\textprime kappa coefficient. On both datasets, we see that SimCTG + contrastive search significantly outperforms other methods on various metrics, suggesting that our approach is generalizable to different languages and tasks. It is worth emphasizing that, on the LCCC benchmark, SimCTG + contrastive search surprisingly outperforms the human performance on the fluency metric, while performing comparably on the coherence and informativeness metrics (Sign Test with p-value ¿ 0.4). Moreover, even without contrastive training, the MLE model performs significantly better when using contrastive search. This is due to the intrinsic property of Chinese language model for which the MLE objective can already yield a representation space that displays a high level of isotropy, making contrastive search directly applicable.We provide more in-depth analyses and several generated examples on LCCC in Appendix H and J, respectively. This finding is particularly attractive as it reveals the potential applicability of contrastive search on off-the-shelf (i.e., without contrastive training) language models for certain languages such as Chinese.

Further Analysis

To analyze the token representations learned by SimCTG, we follow Ethayarajh and define the averaged self-similarity of token representations within a text sequence x\boldsymbol{x} as

where hxih_{x_{i}} and hxjh_{x_{j}} are the token representations of xix_{i} and xjx_{j} produced by the model. Intuitively, a lower self-similarity(x)\textup{self-similarity}(\boldsymbol{x}) indicates the representations of distinct tokens within the sequence x\boldsymbol{x} are less similar to each other, therefore being more discriminative.

We use texts from Wikitext-103 test set and compute the self-similarity of token representations over different layers for different models. Figure 2 plots the results averaged over all samples. We see that, in the intermediate layers, the self-similarity of different models are relatively the same. In contrast, at the output layer (layer 12), SimCTG’s self-similarity becomes notably lower than other baselines. We note that the Unlikelihood model also yields more discriminative representations than MLE, but its language model accuracy is lower than MLE and SimCTG as shown in Table 1. On the other hand, SimCTG obtains the most discriminative and isotropic representations while maintaining the best language model accuracy, which further validates the clear advantage of our proposed approach.

2 The Effect of Contrastive Loss Margin

Next, we analyze the effect of contrastive loss margin ρ\rho (Eq. (2)). To this end, we fine-tune the GPT-2 by varying ρ\rho from 0.10.1 to 1.01.0 and measure the model perplexity on the Wikitext-103 test set. Figure 3 plots the results of different ρ\rho along with the result of the MLE baseline. Note that, when ρ=0\rho=0, SimCTG is equivalent to MLE (Section section 3.1). From Figure 3, we see that the contrastive training always helps to improve the perplexity as compared with MLE. However, when ρ\rho is either too small (e.g., 0.10.1) or large (e.g., 1.01.0), the learned representation space of the model would be either less or too isotropic, leading to a sub-optimal perplexity. In our experiments, the most suitable margin ρ=0.5\rho=0.5.

3 Contrastive Search versus Nucleus Sampling

Then, we provide an in-depth comparsion between our proposed contrastive search and the current state of the art, nucleus sampling. To this end, we compare the results of SimCTG using these two decoding methods. Specifically, we vary the probability pp for nucleus sampling and the α\alpha (Eq. (5)) for contrastive search to generate results using prefixes from Wikitext-103 test set.For contrastive search, we only vary the value of α\alpha and keep kk constant to 88 as described in Section section 4. In Appendix E, we provide detailed ablation studies on the effect of both kk and α\alpha in contrastive search. We evaluate the results from two aspects: (1) generation diversity and (2) perplexity of the generated text (gen-ppl). Both metrics are described in Section section 4.1.2. Figure 4 plots the results of different methods along with the human performance. For nucleus sampling, when pp is small (i.e., p≤0.7p\leq 0.7), its generation perplexity is comparable to that of human. However, the diversity is notably lower than human performance, meaning it stuck in undesirable repetition loops . On the other hand, when pp is large (i.e., p≥0.95p\geq 0.95), the generation diversity is close to that of human but the generation perplexity is significantly higher. Such high perplexity means the generated text is very unlikely, therefore being low quality. As for contrastive search, when α∈[0.5,0.8]\alpha\in[0.5,0.8], it yields generation diversity and perplexity that are both comparable to human performance. These results demonstrate the superiority of contrastive search as it better balances the trade-off between the generation diversity and perplexity.

4 Decoding Latency Comparison

We compare the decoding latency of different decoding methods using SimCTG. For beam search and contrastive search, we vary the beam width bb and the kk in Eq. (5). The latency is measured by generating fixed length text continuations on Wikitext-103 test cases with a batch size of 11. In Figure 5, we show the averaged relative decoding latency of different methods. We see that greedy search is the fastest method and the latency of different methods are generally comparable with each other. Comparing contrastive search with beam search, when bb and kk are small (i.e., ≤6\leq 6), their latency are nearly identical. When bb and kk gets larger (i.e., >6>6), contrastive search becomes faster. In summary, these comparison results further verify the practical usage of contrastive search.

5 Case Study

In Table 4, we present generated examples of SimCTG with different decoding methods given a specific prefix.We refer to Appendix K for more generated examples of SimCTG. From the results, we see that beam search produces undesirable sequence-level repetitions, resulting in low diversity and low generation perplexity. On the other hand, in the prefix, the person “Buchanan” criticizes the game. However, the result from nucleus sampling displays a contradicted semantic, resulting in a low coherence score as well as a high generation perplexity. As for contrastive search, it generates a text that is semantically consistent to the prefix with a proper generation perplexity while obtaining the same diversity as that of the nucleus sampling. Additionally, it is worth emphasizing that, while the degeneration penalty in Eq. (5) encourages the model to generate diverse outputs, contrastive search is still able to generate reasonable repetitions as highlighted in Table 4. This is due to the incorporation of model confidence in Eq. (5) which enables the model to repeat the important content (e.g., person names or entity names) from the previous context like humans do.

6 Comparison of Token Similarity Matrix

To better understand how contrastive search works, in Figure 6, we show the generated token similarity matrix of SimCTG using beam search and contrastive search. For a better comparsion, we also include the result of MLE using beam search. All results are produced with the same prefix as in Table 4. The red and yellow boxes highlight the similarity matrix of the prefix and the generated text. Firstly, we see that, the MLE + beam search yields a very dense similarity matrix, meaning that its token representations are indiscriminative. In addition, the high similarity scores in its off-diagonal entries clearly show the degeneration repetitions. Secondly, for SimCTG + beam search, we observe a desirable similarity matrix of the prefix which is sparse and isotropic. However, degeneration repetitions still exist in the generated result as shown in Figure 6(b). Lastly, for SimCTG + contrastive search, the entire similarity matrix is sparse and isotropic, showing that it successfully solves the model degeneration. These observations are in line with our motivations as described in Section section 1.

Conclusion

In this work, we show that the degeneration of neural language models stems from the anisotropic nature of their token representations. We present a new approach, SimCTG, for training the language model such that it obtains an isotropic and discriminative representation space. In addition, we introduce a novel decoding method, contrastive search, which works coherently with the proposed SimCTG. Extensive experiments and analyses are conducted on three benchmarks from two languages. Both automatic and human evaluations demonstrate that our approach substantially reduces model degeneration and significantly outperforms current state-of-the-art text generation approaches.

Acknowledgments

The first author would like to thank Jialu Xu and Huayang Li for their insightful discussions and supports. Many thanks to our anonymous reviewers, area chairs, and senior area chairs for their suggestions and comments.

References

Checklist

Do the main claims made in the abstract and introduction accurately reflect the paper’s contributions and scope? [Yes]

Did you describe the limitations of your work? [Yes] See Appendix A.

Did you discuss any potential negative societal impacts of your work? [N/A]

Have you read the ethics review guidelines and ensured that your paper conforms to them? [Yes]

If you are including theoretical results…

Did you state the full set of assumptions of all theoretical results? [N/A]

Did you include complete proofs of all theoretical results? [N/A]

Did you include the code, data, and instructions needed to reproduce the main experimental results (either in the supplemental material or as a URL)? [Yes] We provide the code and the instructions to re-implement our results as a supplementary material to this paper.

Did you specify all the training details (e.g., data splits, hyperparameters, how they were chosen)? [Yes] We specify the details in Section section 4 and section 5.

Did you report error bars (e.g., with respect to the random seed after running experiments multiple times)? [No] We did not run multiple times for our experiments due to computational constraints.

Did you include the total amount of compute and the type of resources used (e.g., type of GPUs, internal cluster, or cloud provider)? [Yes] We describe the computational details in Appendix J.

If you are using existing assets (e.g., code, data, models) or curating/releasing new assets…

If your work uses existing assets, did you cite the creators? [Yes] We cite the authors of the datasets and the code of the models in Section section 4 and section 5.

Did you mention the license of the assets? [N/A] The datasets are publicly available.

Did you include any new assets either in the supplemental material or as a URL? [Yes] We provide the code and the instructions to re-implement our results as a supplementary material to this paper.

Did you discuss whether and how consent was obtained from people whose data you’re using/curating? [N/A] The datasets are publicly available.

Did you discuss whether the data you are using/curating contains personally identifiable information or offensive content? [No] We use the standard datasets, which are well known in literature, and there are no personally identifiable information or offensive content at the best of the community knowledge.

If you used crowdsourcing or conducted research with human subjects…

Did you include the full text of instructions given to participants and screenshots, if applicable? [Yes] We provide the human evaluation guidelines in Appendix G.

Did you describe any potential participant risks, with links to Institutional Review Board (IRB) approvals, if applicable? [N/A]

Did you include the estimated hourly wage paid to participants and the total amount spent on participant compensation? [Yes] We provide the details of participant compensation in Appendix G.

Appendix

For future work, we would like to suggest three research directions based on our study.

Our proposed contrastive loss LCL\mathcal{L}_{\textup{CL}} in Eq. (2) is designed to treat all other tokens within the same sequence as negative samples. However, we do acknowledge that there might be a suitably small fraction of tokens (within the same sequence) that share similar semantic meanings even with different surface forms. We believe the current formulation of the contrastive loss might be further improved by taking this aspect into consideration and we leave it to our future work.

One limitation of the proposed contrastive search is that it is a deterministic decoding method. It would be interesting and useful to incorporate a certain level of stochasticity into the decoding process. One plausible approach is to combine contrastive search with stochastic sampling methods. For instance, given the prefix, we could first generate a few tokens (e.g., 1∼\sim3 tokens) with nucleus sampling. Then, we switch to contrastive search for the remaining steps. In Appendix L, we provide some preliminary experiment results on incorporating stochasticity into contrastive search.

Our approach is architecture agnostic and can be applied to any generation model. Future research could focus on adapting it to other tasks than open-ended text generation (i.e., constrained text generation), such as machine translation and document summarization.

Appendix B Related Work

Neural Text Generation is a core component in many NLP applications. It can be generally categorized into two classes (1) constrained generation; and (2) open-ended generation.

Constrained generation tasks are always defined over a set of (input, output) pairs, where the output is a transformation of the input following specific constrains. Some typical examples include machine translation , text summarization , and data-to-text generation . As the output is tightly scoped by the input, the generation of repetition and unnaturalness are not that problematic, therefore maximization-based decoding methods such as beam search generally perform well. Still, different variants of beam search have been explored to further improve the model performance in constrained generation tasks .

Open-ended text generation, on the other hand, imposes less constrain on the generated text. It aims at producing text that is natural, coherent and informative with respect to the human-written prefix (i.e., context). Several typical applications include story generation , contextual text completion , and dialogue systems . However, due to the challenges posed by the increased level of freedom, conventional maximization-based decoding methods (e.g., greedy and beam search) often produce undesirable repetition and unnaturalness in the generated text. To alleviate model degeneration, different sampling approaches have been proposed to generate text by drawing samples from less likely vocabularies. Welleck et al. tackled model degeneration from another perspective by introducing unlikelihood objective into the training of the language model.

Contrastive Learning. Generally, contrastive learning methods aim to teach the model to distinguish observed data points from fictitious negative samples. They have been widely applied to various research areas. In the field of computer vision, contrastive learning has been shown to benefit tasks like image and video representation learning. Chen et al. proposed a simple framework, SimCLR, for learning contrastive visual representations. Recently, Radford et al. and Jia et al. applied contrastive learning for the pre-training of language-image models.

In the field of NLP, contrastive learning has recently gained much more attention. Numerous contrastive approaches have been proposed to learn better token-level , sentence-level , and discourse-level representations. Beyond representation learning, contrastive learning has also been applied to other NLP applications, such as name entity recognition (NER) , document summarization , and knowledge probing for pre-trained language models .

Our work, to the best of our knowledge, is the first effort on applying contrastive learning to address neural text degeneration. We hope our findings could facilitate future research in this area.

Appendix C Software Package

In this section, we illustrate the use of the accompanying Python package, available on Githubhttps://github.com/yxuansu/SimCTG/tree/main/simctg and installable via piphttps://pypi.org/project/simctg/ as pip install simctg --upgrade.

Below, we show how to replicate our result in Table 4 with our provided package. More details can be found in our open-sourced repositoryhttps://github.com/yxuansu/SimCTG.

Appendix D Experiments on Different Language Models

In this section, we further test the generalization ability of our approach with different language models on the Wikitext-103 benchmark. In addition to the GPT-2-small model (i.e. 12 Transformer layers with 12 attention heads) that we consider in Section section 4, we include (i) a vanilla Transformers (i.e. without any pre-training) with the same parameter size as GPT-2-small; and (ii) a larger pre-trained model, GPT-2-large, that consists of 36 Transformer layers with 20 attention heads. The training of different language models follows the same procedure as described in Section section 4. To measure the isotropy of the language model, we include the conicity metric as well as the self-similarity metric (Eq. (6)). A lower conicity or self-similarity indicates the representation space of the language model better follows an isotropic distribution.

Table 5 presents the experimental results. We observe that our approach (i.e. SimCTG + contrastive search) performs the best on all evaluated models, suggesting the clear generalization ability of our approach. Another interesting finding is that, for vanilla Transformers and GPT-2-large, the model trained with MLE naturally displays a high level of isotropy. A similar phenomenon is also observed in language models from other languages, such as Chinese (see Appendix H). In such cases, our proposed contrastive search can be directly applied and yields superior performances. This further points out the huge potential of contrastive search in other much larger and stronger language models such as GPT-3 and OPT . We leave the rigorous investigation on the isotropic properties of different language models to our future work.

Appendix E Ablation Study on the Hyperparameters of Contrastive Search

Here, we present a detailed ablation study on the hyperparameters (i.e., kk and α\alpha in Eq. (5)) of contrastive search. Specifically, we simultaneously vary the value of kk and α\alpha. kk is chosen from {5,8,10}\{5,8,10\} and α\alpha is chosen from {0.4,0.5,0.6,0.7,0.8,0.9,1.0}\{0.4,0.5,0.6,0.7,0.8,0.9,1.0\}. For evaluation, we report the generation diversity and generation perplexity on the test set of Wikitext-103. The results are plotted in Figure 7. We see that, when kk is constant, the increase of α\alpha generally increases the generation diversity and generation perplexity. When α\alpha is constant, a larger kk also leads to the increased generation diversity as well as generation perplexity. Nonetheless, for different kk, the overall trends are relatively the same and the value of α\alpha has more impact on the generated results. In practice, our recommended selection range of kk and α\alpha are k∈k\in and α∈[0.5,0.8]\alpha\in[0.5,0.8], as these settings produce results that are more similar to human-written texts as judged by generation diversity and generation perplexity.

Appendix F Gen-ppl Results Measured by Different Models

In Table 7 and 7, we show the gen-ppl (detailed in section 4.1.2) results of different methods as measured by the model trained with MLE and Unlikelihood, respectively. As we use different models to measure gen-ppl, the results in Table 7 and 7 are slightly different from the ones in Table 1. Nontheless, we can draw the same conclusion as in Section section 4.2 that SimCTG + contrastive search is the best performing method as it obtains the generation perplexity that is closest to the human-written text.

Appendix G Human Evaluation Guidelines

Given the human-written prefix, please evaluate the system’s result with respect to the following features: (1) Coherence; (2) Fluency; and (3) Informativeness. In the following, we provide some guidelines regarding how to judge the quality of the system’s result in terms of different features.

This metric measures whether the system’s result is semantically and factually consistent with the human-written prefix. The definitions of different scores are:

: The system’s result is perfectly in line with the semantic meaning defined by the prefix. And all its content is factually supported by or can be logically inferred from the prefix.

: The system’s result is very related to the prefix but with some minor errors that does not affect its overall relevance with respect to the prefix.

: The system’s result is, to some extent, relevant to the prefix with some errors that display minor semantic inconsistency or contradiction.

: At the first glance, the system’s result seems to be related to the prefix. But with careful inspection, the semantic inconsistency can be easily spotted.

: The system’s result is obviously off-the-topic or it is semantically contradicted to the content contained in the prefix.

G.2 Fluency

This metric measures the fluency of the system’s result. The definitions of different scores are:

: The system’s result is human-like, grammatically correct, and very easy to understand.

: Choose this score when you are hesitant between the score 3 and score 5.

: The system’s result contains minor errors but they do not affect your understanding.

: Choose this score when you are hesitant between the score 1 and score 3.

: The system’s result does not make sense and it is unreadable.

G.3 Informativeness

This metric measures the diversity, informativeness, and interestingness of the system’s result. The definitions of different scores are:

: The system’s result is very informative and contains novel content. In addition, it displays a high level of diversity and it is enjoyable to read.

: Choose this score when you are hesitant between the score 3 and score 5.

: The system’s result contains some new information and it displays a certain level of diversity.

: Choose this score when you are hesitant between the score 1 and score 3.

: The system’s result is dull, repetitive, and does not have new information. All its content has already been provided in the prefix.

Participant Compensation. In each experiment (i.e., open-ended text generation and open-domain dialogue generation), we hire 5 annotators to conduct the human evaluation. For every task, each annotator is paid by $400.

Appendix H Self-similarity of Chinese Language Models

We follow the same procedure as described in Section section 6.1 to measure the token self-similarity of Chinese language models. Specifically, we use the test set of LCCC benchmark and compute the model’s self-similarity. Figure 8 plots the layer-wise token self-similarity of the MLE and SimCTG models. We see that in all layers (including the final layer), the MLE model displays a similar self-similarity with respect to SimCTG. This observation is quite different from what we see from English language models as shown in Figure 2, where the self-similarities of SimCTG and MLE are notably different in the final layer. We conjecture that this discrepancy might come from the intrinsic property of different languages. For English, current state-of-the-art methods always represent the text into subword units, such as BPE , and the same subword could be over-shared by many different contexts. Thus, the representations of distinct subwords become less distinguishable which naturally leads to the anisotropy in their representations.However, we should also note that, for larger English models (e.g., GPT-2-large), this conjecture not longer holds as demonstrated in Appendix D. This urges us to conduct more thorough investigations on the isotropic properties of language models across different sizes as well as different languages. We will leave these investigations to our future work. On the other hand, languages like Chinese are naturally represented by basic units, i.e., characters. Such natural unit boundary of text alleviates the over-sharing of characters in different contexts. As a result, even the vanilla MLE objective can obtain a representation space that displays a high level of isotropy.

This isotropic property of Chinese language model is particularly attractive as contrastive search can be directly applied even without contrastive training as shown in Table 3. In addition, we expect contrastive search could be used on off-the-shelf language models that are trained with MLE in other languages whose texts are naturally tokenized by characters (e.g., Korean and Japanese). This remains to be rigorously tested in our future work.

Appendix I Training Efficiency Comparison

In this part, we compare the training efficiency of different methods (i.e., MLE, Unlikelihood, and SimCTG). To this end, we compute the total floating point operations (FLOPs) required for the training of different models on Wikitext-103. The details of training setup are provided in Section section 4. Table 8 shows the results, from which we see that SimCTG is more efficient than the unlikelihood method. Comparing with MLE, SimCTG only introduces an negligible 1.48% extra computational overhead, which further verifies the practical usage of SimCTG.

Appendix J Generated Examples on Open-domain Dialogue Generation

In Table 9, we show some generated responses of our approach (i.e., SimCTG + contrastive search) plus the reference response on examples from the test set of the Chinese LCCC benchmark. We see that, given the dialogue context, our approach is able to generate responses that are both grammatically fluent and semantically consistent with the dialogue context. These results further demonstrate the generality of our approach across different languages and tasks.

Appendix K More Generated Examples of SimCTG + Contrastive Search

In Table 10, we provide more generated examples of SimCTG + contrastive search based on prefixes from Wikitext-103. The details of the decoding procedure are described in Section section 4.

Appendix L Diverse Contrastive Search

In this part, we present a stochastic version of contrastive search (i.e., diverse contrastive search) which is described in Appendix A. Specifically, given the prefix with length of 32, we first generate 2 tokens using nucleus sampling with p=0.95p=0.95, then we use contrastive search to generate the remaining 126 tokens (i.e., 128 generated tokens in total).

Table 11 shows three generated results with diverse contrastive search using the same prefix as in Table 4. We see that only sampling 2 tokens at the start is enough to produce a diverse set of results. In future work, we will investigate other more sophisticated extensions of contrastive search.