A Simple Contrastive Learning Objective for Alleviating Neural Text Degeneration

Shaojie Jiang, Ruqing Zhang, Svitlana Vakulenko, Maarten de Rijke

Introduction

Autoregressive language models (LMs), such as OpenAI GPT-3 , have achieved impressive results on various natural language processing (NLP) tasks. The goal of training LMs is to learn the true distribution of a text corpus, and this is usually achieved through next word prediction. Specifically, a standard approach to training LMs is to minimize the cross-entropy loss between the true distribution and the model prediction. Unfortunately, LMs trained using the cross-entropy objective have been observed to exhibit text degeneration problems, where token, phrase, and sentence level repetition is a common symptom . Such repeated texts differ markedly from those generated by humans.Readers are referrred to Table 4 for some concrete examples. The degeneration problem even exists in large-scale, state-of-the-art, pre-trained language models such as GPT-3 . To analyze the reasons for degeneration, our work views the vocabulary of LMs as being composed of three sets of tokens at each time step, i.e., positive tokens (label tokens), negative tokens (incorrectly repeating tokens), and irrelevant tokens (all the others). Based on this taxonomy, we stress that cross-entropy is in fact a contrastive learning objective that contrasts positive tokens with negative and irrelevant tokens. While it is necessary for LMs to learn how to rank positive tokens higher than other tokens in the predicted distribution, negative tokens are treated equally to irrelevant tokens (whose number is usually much larger) by the cross-entropy objective. As a consequence, negative tokens may not be suppressed hard enough.

To address the above issue, Welleck et al. have proposed unlikelihood training to penalize certain negative tokens, i.e., tokens being incorrectly repeated. The key idea behind unlikelihood training is to lower the probability of negative tokens assigned by LMs. Despite its success, the unlikelihood objective penalizes negative tokens by decreasing their predicted probability but does not consider the relationship between positive and negative tokens. Unlikelihood training also unintentionally boosts the probability of other irrelevant tokens. Moreover, all previous context tokens are used as negative candidates per generation step. Such an objective not only introduces a considerable amount of noise, but also results in sub-optimal repetition reduction, thus affecting the final generation performance.

In this paper, we introduce a simple yet effective contrastive token learning (CT for short) objective that integrates the best of cross-entropy and unlikelihood training, penalizing negative tokens by contrasting them with positive tokens. The commonalities and differences between cross-entropy, unlikelihood training, and CT are illustrated in Figure 1. Briefly, (i) without distinguishing between negative and irrelevant tokens, cross-entropy cannot effectively suppress negative tokens; (ii) due to the lack of contrast between negative and positive tokens, it is difficult for unlikelihood training to penalize negative tokens; and (iii) through its more focused contrast between positive and negative tokens, CT can take goal-directed actions rather than just predicting label tokens, i.e., explicitly teaching the LM to assign negative tokens with a lower probability than positive tokens. In this work, we combine the CT and cross-entropy objectives to train LMs, where cross-entropy performs on the label tokens so that they are assigned the highest probability, and CT effectively suppresses negative tokens from being generated.

We perform evaluations on the tasks of language modeling and open-domain dialogue generation.Our source code, including data pre-processing scripts, our trained models, and an interactive Google Colab notebook, is available at https://anonymous.4open.science/r/lit-seq. Our empirical evidence demonstrates that LMs trained with the proposed CT objective can generate much less repetitive texts using standard greedy or beam search and achieve superior text generation performance under both automatic and human evaluations. CT has a minor negative influence on the perplexity of LMs, but thanks to the reduced repetition rates, in our case studies we observe substantial improvements regarding the quality of generated text.

Background

LMs aim to learn the true distribution over variable-length text sequences in a text corpus X=(x1,x2,…,x∣X∣)X=(x_{1},x_{2},\dots,x_{|X|}) with ∣X∣|X| tokens. A popular approach to this task is next word prediction, i.e., predicting a distribution over the next word following a given context. To train such a language model, cross-entropy and unlikelihood training are two representative objectives. In this section, we first review cross-entropy and unlikelihood training. We then provide an analysis of the text degeneration problem.

A standard approach to training a LM is to minimize the expected cross-entropy loss between the true distribution and the model prediction . Specifically, the cross-entropy loss for each time step tt is defined as:

where hth_{t} is the model hidden state at time tt, WW is the embedding matrix, and WxtW_{x_{t}} denotes the word embedding of token xtx_{t}. Through some simple transformations from Eq. (1)–(3), we can see that Eq. (3) is similar to the NN-pair contrastive loss for visual object recognition. In other words, cross-entropy effectively trains LMs to contrast the label tokens (positive examples) xtx_{t} with all the other non-label tokens (negative and irrelevant examples) x^t∈V,x^t≠xt\hat{x}_{t}\in V,\hat{x}_{t}\not=x_{t} in the whole vocabulary.

2 Unlikelihood training

To address the repetition issue of cross-entropy, Welleck et al. have proposed unlikelihood training to penalize the likelihood of negative tokens (UL-T). The unlikelihood loss for time step tt is defined as:

where Ct={x1,…,xt−1}\textbackslash{xt}C^{t}=\{x_{1},\dots,x_{t-1}\}\textbackslash\{x_{t}\} is the set of negative tokens at time tt, i.e., all previous context tokens. In this paper, we refer to this set of negative tokens as the preceding tokens set. As we will see in §2.3, UL-T does not work well as it can increase the probability of irrelevant tokens. Welleck et al. have also proposed a more effective sequence-level unlikelihood objective (UL-S) that uses unlikelihood on decoded continuations during training time. We omit the details here as our proposed CT is more closely related to UL-T, but we do compare CT to UL-S in our experiments.

3 Discussion

The main difference between Eq. (3) and the NN-pair contrastive loss is that, in Eq. (3), negative and irrelevant tokens are treated equally by cross-entropy.Albeit with different strengths, as seen in Eq. (10) in Appendix D. These negative tokens need to be penalized harder than irrelevant tokens, otherwise, negative tokens may be incorrectly repeated in later time steps. This explains why LMs trained by cross-entropy have high repetition rates.

Although UL-T penalizes negative tokens, it does not work well enough, and as can be seen from Table 1, the reasons are twofold. First, each negative token is not definitely penalized because it depends on the influence of other negative tokens, which can be seen from the gradient analysis of UL-T (Eq. (11) in Appendix D). Second, the formulation of UL-T unintentionally boosts the probability of other irrelevant tokens and may make them surface as repeated tokens. We detail this analysis in §3.3.

Method

To address the issues discussed above and inherit the advantages of cross-entropy and unlikelihood training, in this section, we present a novel contrastive token learning (CT) objective for LMs. We first define the CT loss for each time step. Then we introduce a positive and negative token selection strategy. Finally, we discuss the differences and connections of CT with respect to cross-entropy and unlikelihood training.

The key idea of CT is to promote positive (label) tokens in the ranking at each step, while lowering negative (incorrectly repeating) tokens, and leave other irrelevant tokens untouched. To this end, we formulate the CT loss for step tt as:

where SNtS_{N}^{t} is the negative token set and xtx_{t} is the positive token (i.e., label token) at time tt. We detail the token selection mechanism of SNtS_{N}^{t} below.

During the training phase, we combine the CT loss with the cross-entropy loss for each time step as follows:

where LCEt\mathcal{L}_{CE}^{t} aims to promote label tokens, training models to assign the highest probabilities to such tokens. On the other hand, LCTt\mathcal{L}_{CT}^{t} focuses on contrasting positive tokens and negative tokens, so that the LMs can learn to effectively rank negative tokens lower than their positive counterparts.

2 Negative token selection strategy

Following , we use the preceding tokens set without requiring additional supervision as our negative tokens SNtS_{N}^{t}. However, using all preceding tokens (as in ) may bring too much noise to the training process, especially for later time steps in a sequence. Hence, we instead propose to use the preceding MM tokens set to decide the negative tokens, with MM being a hyper-parameter. The set SNtS_{N}^{t} is defined as:

Another difference with the preceding tokens set is that, SNtS_{N}^{t} is a multiset that does not remove redundant occurrences. Intuitively, minimizing the CT loss with the preceding MM tokens set makes more frequently repeated tokens less likely to be predicted.

3 Gradient analysis

To see how loss functions influence the positive, negative and irrelevant tokens during training, we derive the gradient functions of each loss function with respect to these tokens in Appendix D. Table 1 is an intuitive summary of the influences, from which one can observe that: (i) Cross-entropy trains to promote label tokens in rankings at each time-step, while suppressing all the other tokens including negative and irrelevant tokens. (ii) It cannot be decided for unlikelihood training whether the negative tokens are promoted or suppressed by the gradient function (cf. Eq. (11) in Appendix D, the valid region for the corresponding gradient function contains both positive and negative values), and irrelevant tokens are promoted, both of which are problematic. (iii) With contrastive token learning, CT promotes positive tokens and suppresses negative tokens, and it is the only objective that does not affect irrelevant tokens (cf. the gradient functions in Appendix D).

When using CT together with CE, as we do for our final loss function, negatives are suppressed both in CT and in CE, while irrelevant tokens are only suppressed in CE. Therefore, our CT objective is able to better restrain incorrectly repeated tokens.

Related work

We review two lines of related work, i.e., neural text degeneration and contrastive learning.

With large-scale pre-training, state-of-the-art neural LMs are able to generate human-like texts . However, they suffer from the text degeneration problem, where model-generated texts are dull and repetitive . The text degeneration problem is especially serious with open-ended generation tasks, such as dialogue generation and language modeling . Some decoding approaches have been proposed to address this problem, by introducing randomness or disparity at inference time. Some other work suggests that the degeneration problem is caused by defects of the likelihood training objective, and improved training objectives have been proposed .

Our proposed contrastive token learning approach belongs to the training objective family. Compared to unlikelihood training , we address the suppression of repetitive tokens by contrasting them with positive tokens.

Contrastive learning. In computer vision, contrastive learning has been widely employed to learn representations . Noise-contrastive estimation has been proved successful for training word embeddings . In recent years, contrastive learning has gained more attention in the area of natural language processing too. Most work builds contrasts at the sequence or document level by corrupting the ground truth sequence or mining positive/negative samples .

Existing token-level contrastive learning frameworks contrast model representations from different positions . Differently, we contrast word embeddings while using the hidden representations as anchor points similar to the triplet contrastive loss . Our formulation effectively contrasts logits output by the model for positive and negative tokens, thus it is more direct than unlikelihood training on addressing the repetitive degeneration problem. To the best of our knowledge, our proposed contrastive token learning is the first to use token embeddings as positive/negative examples in a contrastive framework for the text degeneration problem.

Experimental setup

We compare CT with baseline approaches on the language modeling and open-domain dialogue generation task. Since our experimental results on the dialogue task show a similar pattern as on the language modeling task, we will focus on the language modeling task in the body of the paper and postpone the setup and analyses of the dialogue task to Appendix I.

Baselines and implementation. We implement several state-of-the-art baselines and use them with GPT-2 : (i) The vanilla cross-entropy (CE) objective; (ii) decoding-based methods: banning 3-grams , top-kk sampling , nucleus sampling and contrastive search (SimCTG-CS) ; and (iii) learning-based methods: unlikelihood training , SimCTG , and noise-contrastive estimation (NCE; detailed in Appendix C) . More details can be found in Appendix E.

Dataset, training and inference details. At training time, we fine-tune GPT-2 small on the widely-used Wikitext-103 dataset with each learning-based approach (including the CE baseline) for 50K steps with 3K warm-up steps. As suggested in , for sequence-level unlikelihood training, we first fine-tune the language model using UL-T for 48.5K steps, and then switch to the UL-S objective for another 1.5K steps, resulting in UL-TS. Best model checkpoints for each task are selected according to the lowest validation CE loss with an evaluation interval of 1K training steps. We use trunks of 512 tokens, and a training batch size of 4. All models are trained using the Adam optimizer with a learning rate of 1e-5. For UL-TS, we had to use a smaller learning rate of 1e-6, otherwise the generated texts contain massive ungrammatical repetitions (continuous token repetitions, as can be seen in Table 5 of Appendix F).

At inference time, we compare the performance of each approach to text degeneration using both greedy search and beam search. We use k=50k=50 for top-kk sampling, and p=0.9p=0.9 for deciding the sampling pool of the nucleus method. We follow Welleck et al. to use 50 tokens as the input prefix and let the model generate 100 tokens as a continuation.

Evaluation metrics. We measure the perplexity (ppl) of different approaches. For measuring generative repetition, we follow Welleck et al. to use 1-gram to 4-gram repetition rates (rep-1 – rep-4), which are defined as the number of repeated nn-grams divided by the total number of generated nn-grams in each sequence, micro-averaged over the whole dataset. We also report the generation diversity at the dataset level, which is measured by distinct 11-gram rates (dist-1) and unique 11-gram counts (uniq-1). We adopt human evaluation for measuring the quality of model generated texts. We randomly select 100 prefixes from the test set of Wikitext-103, and compare the continuations generated using CT with those by the best-performing baselines according to the automatic evaluation results. Since it does not make much sense to compare continuations with either side having excessive repetitions, we filter out such pairs using a threshold of rep-4≤0.05\texttt{rep-4}\leq 0.05 to make the comparisons more competitive. Then we display the prefix and two continuations from different systems (side-by-side, in a random order) to three crowd workers and ask them to select the winner in terms of repetition, coherence, fluency, and overall quality. Ties are allowed for all aspects. We use majority voting to decide the final winner. Details about our question form design and the instructions to crowd workers can be found in Appendix G.

Evaluation results

We conduct extensive experiments to demonstrate the advantages of our proposed CT. In this section, we discuss how CT compares to SOTA methods under both the automatic and human evaluations as well as showing some visualization analysis on its generation probability.

The performance comparisons between our CT and the baselines on the language modeling task are shown in Table 2. For models, the repetition and diversity results are calculated on model-generated continuations of 100 tokens, using 50 tokens of human-created text as the prefix. For the human performance, we calculate the metrics on trunks of 100 tokens for a fair comparison. The ppl metric is for 512-token sequences to comply with the training sequence length. To be comparable to existing work , we also report ppl-s for short sequences of 50 tokens. We use a sequence length of 150 tokens and M=60M=60 as the negative window size for CT. Justifications for such hyper-parameter selections can be found in Appendix F.2.

CT compared to learning-based approaches. One can observe that CT performs the best and even outperforms humans according to rep-* rates and unique token counts (uniq-1) when using greedy search. However, the repetition problem is still not yet solved, because when looking at specific cases, models trained by CT still occasionally generate texts with excessive repetitions, though being much rarer than baseline methods. To see how each method performs at every repetition level, we group the rep-1 and rep-4 rates of model-generated texts in to 5 bins, and plot their histograms in Figure 2, from which we can see that CT generates substantially less degenerated continuations (with rep-1≥0.4\geq 0.4 and rep-4≥0.2\geq 0.2). For UL-TS, we were able to achieve lower repetition rates with a larger learning rate of 1e-5 during training. However, the trained LM often generates ungrammatical repetitions. This problem does not exist with CT when trained with a learning rate as large as 1e-4. The comparisons are shown in Table 5 in Appendix F, and in §6.3 we show that this is caused by UL-TS being uncertain about its predictions at later time steps.

The diversity improvements brought by CT are the largest among all learning-based methods, especially when using greedy search. CT increases the second highest uniq-1 count (NCE) by 55%. When comparing NCE and UL-T, one can see that utilizing the contrast between positive and negative tokens works better than solely penalizing negative tokens. The primary difference between CT and NCE is that the positive and negative tokens of CT interact with each other, while those of NCE do not (Table 1, more details in Appendix D). This explains the lower rep-* rates and higher diversity of CT, which also concurs with the observation made by Sohn that interactive contrastive losses work better than non-interactive counterparts.

The ppl increase brought by CT is minor, with 0.71 points. When calculated on short sequences, due to the length mismatch of training and test sequences, ppl-s scores are higher than ppl for all approaches. Among them, contrastive objectives (NCE and CT) have larger ppl-s increases than other methods. Although CT has the highest increase on ppl-s, our case study (Table 4) shows that the generation quality of CT is not harmed, but on the contrary is improved due to the lower repetition and higher diversity of the generated texts.

CT compared to decoding-based approaches. Although CT is a learning-based method, we still compare it against decoding approaches for a more comprehensive understanding of its performance. When greedy search is used, CT outperforms the best decoding method (Top-kk) in terms of rep-* rates, which again proves the effectiveness of contrastive learning. When using beam search, all but SimCTG-CS perform significantly worse than CT, both in terms of repetition rates and diversity. SimCTG-CS is effective at reducing repetition as it explicitly requires a disparity among different time steps at inference time. This can harm the generation quality, especially the coherence and fluency, as we see in §6.2. It is also worth noting that SimCTG-CS only works together with its SimCTG training objective and with beam search . In summary, one can see that the repetition problem can be better addressed from the model learning perspective, in which case a simple greedy decoding strategy suffices.

2 Human evaluation

Human evaluation results are shown in Table 3. Regarding the overall quality, CT performs significantly better than Top-kk and SimCTG-CS, two decoding based approaches. Instead of purely learning generation policies from data, decoding approaches exert heuristics at inference time, which may prevent the language model from performing naturally. This explains the worse performance of decoding approaches on coherence and fluency. CT performs generally better than UL-TS except on coherence, but none of these differences are statistically significant. This suggests that CT has a similar generation quality as UL-TS on low-repetitive examples, but CT has much lower repetition rates as reported in Table 2. This result is expected, as both CT and UL-TS are learning-based approaches for training data-driven models, and on normal cases such as low-repetitive generations, they should perform similarly. Compared to human performance, there is still a large margin for machine learning models before they have a comparable performance on the language modeling task. Although CT performs on par with humans regarding repetition, its generations are far less coherent and fluent than those of humans. This may be mitigated by using larger models such as GPT-2 large or GPT-3. However, we could not perform such experiments due to a lack of computational resources.

3 Visualization analysis of the generation probability

We also conduct analyses to understand the predicted probability of model-generated tokens at inference time. As shown in Figure 3, diagonal cells represent the probability of generated tokens at the corresponding time steps; off-diagonal cells represent the probability of context tokens. The plots are averaged over 10 random instances from the test set of Wikitext-103.

We have the following key observations from Figure 3: (i) The heat map of CT shows a high variance in the diagonal, meaning that the model becomes certain and uncertain from time to time. As noted by Holtzman et al. , human-created texts also show such a pattern when fed through pretrained language models. (ii) In comparison, the heat map for CE shows clear stripes, which stand for excessive repetition of context n-grams. Besides, the diagonal cells are increasingly darker from top to bottom, revealing that the language model is becoming more and more certain about its later predictions, and it seems to positively correlate with the heavier repetition in the later halves of sequences. (iii) Contrary to CE, the heat map for UL-TS is almost white at the lower and the right parts of the heat map, indicating the language model is uncertain about any prediction in later stages, and the generated tokens just win marginally over other candidates. This is expected, since UL-TS penalizes repetitions unilaterally, and repetitions are more common in the later half of a model-generated sequence. Even though UL-TS is able to effectively reduce repetition rates, its heat map shows that the language model trained by UL-TS may subject to frequent grammatical errors, as can be seen in Appendix F, Table 5.

4 Case study

To intuitively see how well CT performs, we selected some example generations of CT, and compare them with those generated using UL-TS in Table 4. More often than not, continuations generated by CT are less repetitive and make more sense than those generated by UL-TS. The reason for the poor quality of UL-TS is that sequence-level unlikelihood training penalizes repeated 4-grams generated by LMs, making LMs uncertain about their predictions as suggested in Figure 3.

Conclusion and discussion

In this paper we studied the neural text degeneration problem. By integrating the best of cross-entropy and unlikelihood training objectives, we obtain a simple and effective contrastive token learning (CT) framework. The main novelty of this work is adapting contrastive learning to the token level of autoregressive language model training. As far as we are aware, our work is the first to use model hidden states as the anchor points and tokens as the positive and negative examples to formulate the contrastive loss. By contrasting the preceding MM tokens at a training step with the label token, LMs learn to not repeat such tokens, thus alleviating the repetition problem. Although the idea of negative tokens is similar to UL, our formulation of contrastive objective is more effective and safer to use. Experiments on the open-ended text generation and open-domain dialogue generation tasks show that CT beats UL-TS, the previous state-of-the-art approach to tackling the repetitive text degeneration problem. CT not only achieves the lowest repetition rates and the highest generation diversity, but also higher generation quality according to our human evaluation.

We performed experiments on fine-tuning LMs for reducing their repetition rates, which can be beneficial for related tasks such as abstractive summarization, machine translation, and image captioning. Our early experiments show that CT can be safely integrated when training a language model from scratch, which can be helpful for future pre-training of large language models. In this work, we used CT with decoder-only (GPT2) and encoder-decoder (BlenderBot) language models, but we note that CT can also be used with encoder language models (e.g., BERT ) to potentially improve the model performance such as prediction accuracy. The repetitive degeneration problem is still not fully solved as occasional, excessive phrase repetitions remain in the generated texts. We leave these research directions as future work.

References

Appendix A Ethical considerations

In this work, we used publicly available English data to train/validate/test models. As far as we know, the curators of these datasets have taken ethical issues into consideration when creating the datasets. We manually checked some generated texts of the language models trained by CT and did not observe any noticeable traces of concern, such as offensive and malevolent language. We share our source code and trained model weights to support its correct use. To make sure the human workers involved in the data labeling efforts, as part of the human evaluation for this study, are fairly paid, we applied the minimum hourly rate of 10.48 euros, which converts to 11 dollars per hour. However, we warn that generative language models should always be used with caution since the generated texts are usually novel and unexpected wordings may appear when trained on improper data. Especially, generative models can be used maliciously, e.g., to generate fake news articles.

Appendix B Using CT in your work

We summarize the steps for calculating LCTt\mathcal{L}_{CT}^{t} in Algorithm 1. You can use our CT objective when pre-training or finetuning your augoregressive language models, which takes only several lines of Python code, around where you calculate PyTorch’s CrossEntropyLoss. Simply use pip install ct-loss to install the required packages. Then you can use CT as follows:

Appendix C Noise-contrastive estimation for autoregressive language models

where σ(⋅)\sigma(\cdot) is the sigmoid function.

Appendix D Gradient functions

To see how loss functions influence the logits during training, we compare the gradient of each loss function. Writing zxt=htTWxtz_{x_{t}}=h_{t}^{T}W_{x_{t}} for the logit of token xtx_{t}, the gradient function is calculated by ∂L∗/∂z∗\partial\mathcal{L}_{*}/\partial z_{*}, where L∗∈{LCE,LUL,LCT}\mathcal{L}_{*}\in\{L_{CE},L_{UL},L_{CT}\}, and z∗∈{zxt,zx^t,zxt−}z_{*}\in\{z_{x_{t}},z_{\hat{x}_{t}},z_{x_{t}^{-}}\}. For clarity, we further denote p(∗∣x<t)p(*|x_{<t}) as p∗p_{*}.

Gradient functions of cross-entropy, w.r.t. label tokens xtx_{t}:

and non-label tokens x^t\hat{x}_{t} (including negative tokens and irrelevant tokens):

Gradient functions of unlikelihood training w.r.t. negative tokens xt−x_{t}^{-}:

and other tokens x^t\hat{x}_{t} (including label tokens and irrelevant tokens):

Gradient functions of CT w.r.t. positive tokens xtx_{t}:

Because all terms in Eq. (5) are independent with irrelevant tokens x^t\hat{x}_{t}:

NCE with respect to label tokens xtx_{t}:

Same as CT, all terms in Eq. (8) are independent with irrelevant tokens x^t\hat{x}_{t}:

Appendix E Required software and hardware resources

For the CE and decoding baselines, we use GPT-2 implemented and pretrained using the CE objective by Hugging Face . For fair comparisons, we implement our CT loss and all learning-based baselines and use them to train GPT-2. Specifically, for unlikelihood training, we implemented both the token-level (UL-T) and the sequence-level (UL-S) variants, according to the official source code . We also implemented SimCTG according to the official code . Similar to CT, we adapted NCE to the token-level (detailed in Appendix In our experiments, NCE is also used together with CE as was done for CT in Eq. (6).

Our implementation is based on Hugging Face Transformers (Apache-2.0 license) , PyTorch Lightning (Apache-2.0 license) , and Hydra (MIT license) . Our source code is directly based on Lightning Transformers (Apache-2.0 license) , thus inheriting the license. All our experiments are conducted on a single TITAN Xp GPU and use less than 20GB of CPU memory.

Appendix F Additional results and analysis for the language modeling task

Figure 4 reveals that the heat maps for NCE, UL-T and SimCTG are similar to that of CE in Figure 3. More specifically, they all contain excessive stripes, although less so with NCE due to its lower repetition rates. Besides, they are also darker at the lower-right half of the diagonal cells, especially for NCE and SimCTG.

Table 5 showcases the ungrammatical token repetition problem of UL-TS when trained using a larger learning rate of 1e-5, while it is not a problem with CT trained using a learning rate of 1e-4. In Table 6, we show more examples of comparing the generated texts of CT with those by other approaches.

F.2 Breakdown analysis

Beyond the overall performance analysis given above, we also provide a breakdown analysis for CT.

Analysis of Sequence Length. As mentioned earlier, when calculating the CT loss, we efficiently reuse the logits computed for CE. Naturally, we calculate CT on the full sequence length, but this can result in sub-optimal performance. We therefore study the influence of the sequence length for CT and plot the rep-* rates and ppl in Figure 6. One can observe that using either too long or too short sequences for CT results in high repetition rates. Especially with long sequences, ppl is hurt substantially. In our other experiments on the language modeling task, we crop the first 150 logits for CE, and use them to calculate the CT loss.

Analysis of Negative Tokens Number. Similarly, when selecting negative tokens, using all the preceding tokens is not the best option. We can see from Figure 6 that when MM is too small, CT has a weak effect on reducing repetition; when M=60M=60, CT achieves the best rep-4 performance, which we use as the default for other experiments. When looking together with the results on the dialogue task (Appendix I), we found that empirically, using 1/41/4 of the logits for computing CT, and selecting M=1/8M=1/8 of the maximum sequence length, often results in good performance.

Appendix G Human evaluation design

Figure 7 is a screen shot of our design of question form. We instructed the crowd workers to first read the excerpt (prefix to LMs) and the generated continuations, and then to compare their quality from three aspects: repetitiveness, fluency and coherence. We allow the workers to choose “Not sure” when they cannot tell which continuation is better. Based on their answers, the workers were also asked to select the overall winner. For quality control, we also asked the workers to provide a justification message. Please see Figure 8 for the full instruction.

Appendix H Experimental setup for the dialogue task

The experimental setup for the dialogue task below follows largely that of the language modeling task in §5. Below we focus on the differences.

Datasets. We follow Roller et al. to use a mixture of multiple high-quality datasets, including PersonaChat , Empathetic Dialogues , Wizard of Wikipedia , and BlendedSkillTalk . We add another benchmark dialogue dataset DailyDialog . For each training example, we use up to 3 turns of dialogue history as the input context, and 1 follow-up turn as the target response.

Training and Inference Details. We use the 400M-distilled version BlenderBot implemented and pretrained using the CE objective by Hugging Face . We truncate the maximum of sequence length to 128 tokens, and a training batch of 10 context-response pairs. We follow Roller et al. to force BlenderBot to generate at least 20 tokens.

Appendix I Results on the open-domain dialogue task

The results on the open-domain dialogue task are reported in Table 7. Generations have a minimum length of 20 tokens. Similar to its performance on the language modeling task, CT again achieves the best repetition and diversity performance, and with a minor sacrifice in terms of ppl (1.44 points).

Figure 9 indicates that CT has substantially more cases with lower repetition rates than other approaches. Due to the fact that dialogue responses are usually short (∼\sim20 tokens), the rep-4 rates of each method are not far apart, although CT marginally wins.

Regarding the selection of the sequence length for CT and the window size for selecting negative tokens, we made similar observations on the dialogue task as those on the language modeling task, as can be seen from Figure 11 and 11.

Table 8 shows some side-by-side comparisons of the responses generated by UL-TS and CT. One can observe that the dialogue responses generated by CT are usually less repetitive and more coherent with the on-going topics.