Click: Controllable Text Generation with Sequence Likelihood Contrastive Learning
Chujie Zheng, Pei Ke, Zheng Zhang, Minlie Huang
Introduction
Current language models trained on massive textual corpora have shown the impressive capability of generating fluent and grammatical text Radford et al. (2019); Brown et al. (2020); Roller et al. (2021). However, they often produce behaviors misaligned with human expectations. For instance, language models may generate offensive language or agree with toxic input Xu et al. (2021); Gehman et al. (2020); Sun et al. (2022). They may also generate text with unnatural repetition Holtzman et al. (2019); Su et al. (2022), which is a notorious issue in autoregressive language generation. Controlling language models to avoid such undesirable attributes has always been an important yet challenging problem in NLG research.
As a popular practice, growing recent work has investigated how to decrease the generation probability of these negative samples (i.e., generations of undesirable attributes). For instance, Unlikelihood Training Welleck et al. (2019) minimizes the likelihood of each token in negative samples. GeDi Krause et al. (2021), DExperts Liu et al. (2021), and Director Arora et al. (2022) adjust the next-token prediction distribution at each generation step to avoid token choices that would potentially lead to undesirable attributes.
In this work, we introduce Click, a method for Controllable text generation with sequence Likelihood C(K)ontrastive learning. It employs a max-margin contrastive loss on sequence likelihood in addition to standard language modeling (§ 2.2), which fundamentally reduces the probability of a negative sample being decoded. Compared with previous methods of controllable text generation, Click has two unique advantages. First, Click contrasts the sequence likelihoods of positive and negative samples with a maximum likelihood margin, which enables a higher degree of freedom for optimization than explicitly minimizing the likelihood of each token of negative samples. Second, Click needs no modification to the model architecture and thus does not require laborious adjustments to the next-token prediction distribution during generation, which makes it convenient for out-of-the-box use of trained models.
We also design a likelihood ranking-based strategy of contrastive sample construction for Click (§ 2.3). Given an input prompt, Click first samples multiple generations from the initial language model, which are labeled as positive/negative by a label function. It then pairs each negative sample with the positive one whose likelihood ranks highest but lower than the former. For instance, in Figure 1, negative rank 2 is paired with positive rank 4 to constitute a pair of contrastive samples. This strategy derives from our two intuitions. First, a high-likelihood positive sample (e.g., positive rank 1) does not necessitate further enlargement of its likelihood gap with the negative one, which may instead result in overfitting the positive sample. Second, sequence likelihood indicates how much a text is probable to be the continuation of the input, which somewhat reflects the quality of generated continuations, such as fluency and coherence. A pair of samples with a too large likelihood gap (e.g., negative rank 2 and positive rank 6) may thus bias contrastive learning toward other aspects (e.g., fluency or coherence) than the attributes we aim to control.
We experiment with three controllable text generation tasks: language detoxification, sentiment steering, and repetition reduction (§ 3). Through both automatic and human evaluation, we show that Click can effectively avoid undesirable attributes and outperform strong baselines. Ablation analysis further proves the superiority of Click’s sample construction strategy.
Methodology
Given an input text as a prompt, the task of controllable text generation aims to generate a fluent natural language continuation that avoids an undesirable attribute (e.g., toxicity) while maintaining contextual coherence. We denote the language model parameterized by as , which produces given following the distribution . Following the setting of controllable text generation Liu et al. (2021); Lu et al. (2022), we also assume a label function that assigns a binary attribute label to each pair We assume the label function with binary outputs rather than continuous outputs (e.g., from 0 to 1) due to two considerations. (1) Since the label function is usually implemented as an automatic classifier, its continuous output score may be imperfect, as discussed in § Limitations. Optimization toward continuous scores may inherit more biases from the classifier, which can be alleviated to some extent by transforming continuous scores into binary labels. (2) This setting can be naturally generalized when the label function is human annotators Ouyang et al. (2022), where only binary or discrete labels can be obtained. , corresponding to a negative/positive sample, respectively.
2 Sequence Likelihood Contrastive Learning
Click adopts a contrastive loss on sequence likelihood, which trains the model to assign lower generation probabilities to negative samples than positive ones. It does not need any modification to the model architecture, which makes it convenient for out-of-the-box use. Figure 1 gives the overview of Click. We first introduce how Click trains the language model to avoid undesirable behaviors (the 3rd step in Figure 1) and later describe Click’s strategy of constructing contrastive samples in § 2.3 (the 1st and 2nd steps in Figure 1).
where is the margin hyperparameter. The overall optimization objective is the summation of the above two losses:
3 Contrastive Sample Construction
Motivation
Likelihood Ranking-Based Strategy
Based on the above intuitions, Click adopts a novel likelihood ranking-based strategy for constructing contrastive samples. From the with lower likelihoods than , Click selects the highest-ranked . With the positive and negative samples at a similar likelihood level, it enables contrastive learning to focus better on the controlled attributes and also alleviates the conflict with the language modeling objective. The strategy is formulated as follows:
If all the have lower likelihoods than , Equation 4 degenerates to selecting the positive sample with the lowest likelihood, i.e., . The 2nd step in Figure 1 illustrates how our construction strategy works, where three pairs of contrastive samples are constructed: 2/4, 3/4, and 5/6.
4 Relationship to Prior Work
Click builds upon two disjoint ideas from previous work in controllable or conditional text generation.
(1) Inspired by Unlikelihood Training Welleck et al. (2019), Click trains the language model to decrease the generation probability of negative samples (Equation 2). However, Unlikelihood Training minimizes the likelihood of each token given the prefix of the negative sample, which is a token-level objective. Different from it, Click adopts a max-margin contrastive loss at the sequence level. By directly acting on sequence likelihood and setting a maximum margin , Click allows a higher degree of freedom for optimization (e.g., focusing on certain tokens that lead to undesirable attributes).
(2) Inspired by BRIO Liu et al. (2022) and SLiC Zhao et al. (2022), Click employs the contrastive loss directly on sequence likelihood. However, BRIO and SLiC align sequence likelihood with the similarity to reference text, which is not applicable for controllable text generation tasks where reference texts are usually unavailable and generation is open-ended. Unlike them, Click aligns sequence likelihood with the controlled attribute (the undesirable attribute corresponds to lower likelihood). Furthermore, the contrastive samples in BRIO and SLiC are randomly paired, while Click is based on likelihood ranking, which provides more insights about and is more tailored for open-ended text generation tasks, as verified in § 3.4.
Experiments
We next show that Click can effectively avoid undesirable attributes on three controllable text generation tasks: (1) language detoxification (§ 3.1), (2) sentiment steering (§ 3.2), and (3) repetition reduction (§ 3.3). We also conduct ablation analysis to give further insights about Click (§ 3.4).
Language models are known to produce offensive language Gehman et al. (2020) or express agreement with toxic input Xu et al. (2021); Sun et al. (2022), which potentially hinders downstream tasks and real-world applications Perez et al. (2022); Zheng et al. (2023). The task of language detoxification aims to avoid toxic and unsafe generations.
Baselines
Following previous work Arora et al. (2022); Adolphs et al. (2022), we use BlenderBot 365M Roller et al. (2021) as the base model. We compare the following methods. Non-toxic FT fine-tunes BlenderBot on the non-toxic training set. Unlikelihood Training Welleck et al. (2019) minimizes the likelihood of each token given the prefix of the toxic sample and also performs language modeling on non-toxic samples. GeDi Krause et al. (2021) and DExperts Liu et al. (2021) both train a toxic/non-toxic model on the toxic/non-toxic training set and adjust the next-token prediction distribution of the original language model. Director Arora et al. (2022) trains a classification head to similarly adjust the next-token prediction distribution. Cringe Adolphs et al. (2022) improves Unlikelihood Training by applying token-level contrastive learning to toxic samples.
Evaluation Setups
For automatic evaluation, we follow the evaluation metrics in Liu et al. (2021), including the aspects of toxicity, fluency, and diversity. Toxicity is measured by the empirical probability (Prob.) of generating at least one toxic continuation over 25 continuations (labeled by the BAD classifier). Fluency is measured by the mean perplexity (Out. PPL) of generated continuations, as evaluated by a larger language model BlenderBot 1.4B. Diversity is measured using the mean number of distinct -grams, normalized by the text length Li et al. (2016), among the 25 generations for each prompt. We report Dist-2/3 scores for distinct bigrams/trigrams, respectively.sa
We also conducted pairwise human evaluation to compare generation results from Click to baselines. 100 prompts were randomly sampled from the BAD test set and each comparison (Click vs. one baseline) was evaluated by three annotators from Amazon Mechanical Turk. Following Liu et al. (2021), evaluation metrics include the perceived level of toxicity (which one is less offensive or biased), fluency (which one is more grammatically correct and coherent), and topicality (which one is more natural, relevant, and logical). See Appendix C.1 for human evaluation details.
Results
As shown in Table 1, Click substantially reduces toxic generations compared to baselines while maintaining reasonable generation diversity. Director and GeDi perform next best to Click but obtain much lower Dist-2/3, indicating that the former two methods both sacrifice generation diversity largely. Table 2 also shows that human annotators rated Click generations as less toxic than the competitors, demonstrating the effectiveness of Click in eliminating toxic language. See Appendix D for additional qualitative results.
2 Sentiment Steering
The task of sentiment steering aims to control the sentiment polarity of generated text, which is well-studied in research of controllable text generation.
Baselines
We use GPT-2 Large 774M as the base model, consistent with previous work Liu et al. (2021). Same as § 3.1, we use Target FT, which fine-tunes GPT-2 on the training data with the target sentiment, GeDi, and DExperts as baselines. We also include PPLM Dathathri et al. (2019), CTRL Keskar et al. (2019), and DAPT Liu et al. (2021) as baselines. For former two are classical methods for controllable text generation and the latter one applies domain-adaptive pre-training on positive or negative sample corpora. We use these baseline results from Liu et al. (2021).
Evaluation Setups
Following Liu et al. (2021), we report the mean proportion of positive/negative continuations over 25 generated continuations (% Positive/Negative), as labeled by the HuggingFace sentiment classifier. Out. PPL is calculated with a larger language model GPT-2 XL 1.5B. Dist-2/3 is calculated consistently with § 3.1.
We also conducted pairwise human evaluation for both positive and negative sentiment steering on negative and positive prompts, respectively. Same as § 3.1, 100 negative/positive prompts were randomly sampled and each comparison (Click vs. one baseline) was evaluated by three human annotators from the aspects of sentiment (which one is more positive/negative), fluency, and topicality. See Appendix C.2 for human evaluation details.
Results
As shown in Table 3, Click more effectively steers toward the target sentiments, especially in the adversarial settings (i.e., steering toward the opposite sentiment to the prompt). While Click’s Out. PPL is a bit higher, we believe it is a trade-off with sentiment control since steering a positive/negative prompt toward negativity/positivity may result in an unexpected continuation, which is reflected in a higher Out. PPL. Table 4 shows that Click has close fluency and topicality to baselines but performs better in sentiment steering. See Appendix D for additional qualitative results.
3 Repetition Reduction
Autoregressive language models usually suffer from generating text with unnatural repetition Holtzman et al. (2019), which is a long-standing and important problem in NLG research Welleck et al. (2019); Jiang et al. (2022). We aim to reduce repetition in language generation with Click.
Baselines
As in Su et al. (2022); Lu et al. (2022), we use GPT-2 Base 124M Radford et al. (2019) as the base model. We compare MLE (maximum likelihood estimation), the standard language modeling method with the conventional negative-log likelihood loss, Unlikelihood Welleck et al. (2019), SimCTG Su et al. (2022), a contrastive training method, and Quark Lu et al. (2022), which conditions language generation on quantized reward tokens. Note that SimCTG, Quark, and our Click are all first pre-trained on the WikiText-103 training set with the MLE objective, and then trained with their own objectives.
Evaluation Setups
We evaluate both the language modeling quality and the generation quality, following previous work Welleck et al. (2019); Su et al. (2022). For language modeling quality, we calculate perplexity (PPL) and next-token prediction accuracy (Acc) on the ground-truth continuations of the WikiText-103 test set. We also calculate prediction repetition (Rep), which is defined as the fraction of the next token repeating the prefix tokens, and its variant (WRep), which excludes the cases of the ground-truth token being predicted and repeating the prefix tokens. For generation quality, we report the proportion of repeated 2/3-grams (Rep-2/3) and diversity (Div) as an overall assessment of text repetition. We also report MAUVE Pillutla et al. (2021), an automatic metric that meauses how much the distribution of generated text diverges from human-written text.
We also conducted pairwise human evaluation. 100 prompts were randomly sampled and each pair of generations were compared by three human annotators from the aspects of coherence (which one is more aligned in meaning/topic with the prompt), fluency (which one is more grammatical, understandable, and non-repetitive) and overall quality. See Appendix § C.3 for human evaluation details.
Results
As shown in Table 5, Click remarkably reduces generation repetition with greedy decoding, leading to the highest diversity (0.72) and MAUVE (0.93) scores. While Click has higher PPL and lower Acc, this is probably due to the increased entropy of next-token prediction, which may be a side-product of reducing generation repetition by directly optimizing sequence likelihood. From Table 6, Click is preferred by human in terms of coherence, fluence, and overall quality. See Appendix D for additional qualitative results.
4 Ablation Analysis
We conduct ablation analysis to give further insights about Click. We focus on the language detoxification task (§ 3.1) unless otherwise stated.
We compare Click with several alternatives. For each negative sample , Random randomly selects a positive sample: , as adopted in previous work Liu et al. (2022); Zhao et al. (2022), Lower randomly selects a positive sample only from those with lower likelihood than : , and Lowest selects the positive sample with the lowest likelihood: .
As shown Table 1, 3 and 5, Click generally outperforms all the three alternative strategies in either fluency or control effect. We notice that Lower achieves better fluency than Random (lower Out. PPL) in Table 3, probably because the former avoids overfitting high-likelihood positive samples. However, Lower and Lowest both underperform Click in fluency (higher Out. PPL in Table 1 and 3) and control effect (all the three tables). It confirms our intuitions in § 2.3 that exploiting the positive samples with much lower likelihoods than the negative ones somewhat impairs the effectiveness of contrastive learning (biased by contrastive samples with too large likelihood gaps) and the language generation capability (impacted by the low-quality positive samples).
Effect of Weight 𝜶𝜶\bm{\alpha} and Margin 𝜸𝜸\bm{\gamma}
Effect of Contrastive Sample Number
In § 3.1, we constructed at most pairs of contrastive samples for each prompt . We now vary from to . As shown in Figure 3, increasing generally does not reduce toxicity better but instead decreases Out. PPL. The former is probably due to that the contrastive loss (Equation 2) has been effective enough to eliminate toxicity. For the latter, we speculate this is because the model-generated positive samples are overall of high likelihood and preferred by the language model (as a reference, the base model BlenderBot generates only 5 toxic ones out of 20 continuations on the BAD training set). Hence, optimization toward more positive samples leads to more generations with similarly high likelihood (or low Out. PPL), as observed in previous work Wang et al. (2022).
Effect of Iterative Training
Similar to the practice in recent work Lu et al. (2022); Adolphs et al. (2022), Click can also continue to improve by iterative training (i.e., we use trained Click as the initial model for another iteration). As shown in Table 7, Click trained with one additional iteration further reduces toxicity while generation fluency and diversity is slightly impaired. We conjecture it is a trade-off between language generation quality and toxicity, as similarly observed in Lu et al. (2022).
Related Work
As pre-trained language models display the impressive capability of language generation Brown et al. (2020), controlling their generation has become increasingly important in recent years. There are two major directions for controllable text generation: decoding-time and training-based methods.
Decoding-time methods steer model generation toward the desired attribute with lightweight modules without tuning the original model. PPLM Dathathri et al. (2019) updates the decoded hidden state according to the classifier’s gradient. FUDGE Yang and Klein (2021) trains a classifier to predict whether a partial sequence will satisfy the desired attribute in the future. GeDi Krause et al. (2021) and DExperts Liu et al. (2021) adjust the next-token prediction distribution with two class-conditional auxiliary models. However, decoding-time methods may suffer from high computational expense during generation (e.g., PPLM) and make models inconvenient for out-of-the-box use.
Click falls into training-based methods, which directly train language models to avoid undesirable attributes. Training-based methods include Unlikelihood Training Welleck et al. (2019), the Cringe loss Adolphs et al. (2022), Quark Lu et al. (2022), and Director Arora et al. (2022), which are used as compared baselines in our main experiments.
Contrastive Learning for Language Generation
Contrastive learning aims to learn meaningful representations by contrasting positive and negative samples Chen et al. (2020); He et al. (2020); Gao et al. (2021), which also inspires recent NLG research. CoNT An et al. (2022) aligns encoder and decoder representations for non-open-ended language generation. SimCTG Su et al. (2022) designs a contrastive training method to learn discriminative and isotropic representations for language generation models. BRIO Liu et al. (2022) and SLiC Zhao et al. (2022) uses a contrastive loss to align sequence likelihood with the similarity to reference text. Unlike them, our work applies the contrastive loss to sequence likelihood and targets open-ended text generation tasks, which require special design of sample construction, as discussed in § 2.4 and 3.4.
Conclusion
This work introduces a controllable text generation method Click, which needs no modification to the model architecture and facilitates out-of-the-box use of trained models. It employs a contrastive loss on sequence likelihood and adopts a likelihood ranking-based strategy for contrastive sample construction. Our empirical evaluation on the tasks of language detoxification, sentiment steering, and repetition reduction demonstrates that Click can effectively avoid undesirable attributes in language generation and outperforms strong baselines. Ablation analysis gives further insights about Click’s sample construction strategy, hyperparameters, and combination with iterative training. Future work can investigate the combination of Click and various label (or reward) functions Ouyang et al. (2022).
Limitations
Ethical Considerations
As with any controllable text generation technique, Click runs the risk of dual use Pandya (2019). Specifically, they could be used to automatically produce harmful contents or malicious behaviors McGuffie and Newhouse (2020). Please refer to Bender et al. (2021) for a broader discussion of such risks. We hope those who use controllable text generation technologies in real-world deployed systems to consider the potential negative impact and avoid using them to generate harmful contents and misinformation, etc.
For human evaluation, we have obtained study approval from the Institutional Review Board (IRB). We paid the crowdworkers at a fair hourly wage (about $8/hour) and did not collect any personal identifying information.
Acknowledgements
This work was supported by the NSFC projects (Key project with No. 61936010 and project with No. 62206150). This work was also supported by the Guoqiang Institute of Tsinghua University, with Grant No. 2020GQG0005.
References
Appendix A Dataset Statistics
All the data and models we experimented with are in English language.
SST-5 Socher et al. (2013) and OpenWebText Gokaslan and Cohen (2019)
We use SST-5 as training data and the OpenWebText prompt sets from Liu et al. (2021) as test data in § 3.2, which are both accessible on Liu et al. (2021)’s official repositoryhttps://github.com/alisawuffles/DExperts. SST-5 contains 4,963/4,650 positive/negative sentences, respectively. The OpenWebText positive/negative/neutral prompt sets contain 2.5K/2.5K/5K prompts, respectively.
WikiText-103 Merity et al. (2017)
We use the official split of the WikiText-103 dataset, which contains 100M English tokens from Wikipedia articles. Please refer to Welleck et al. (2019); Lu et al. (2022) and Su et al. (2022)’s official repositoryhttps://github.com/yxuansu/SimCTG/tree/main/document_generation for data access and detailed statistics.
Appendix B Model Details
We implemented all the models with the Transformers library Wolf et al. (2020). The implementation details and computational cost are summarized in Table 9.
B.2 Hyperparameters
We conducted simple grid searches for hyperparameters of Click as well as the baselines in § 3.1. Table 10 presents the search results of Click, while Table 11 presents the baselines in § 3.1.
B.3 Classifiers
In § 3.1, we trained a RoBERTa Base 125M classifier Liu et al. (2019) on the BAD training set as the label function , which takes a prompt and a continuation as input. As shown in Table 8, the BAD training set contains 69,274 utterances annotated as toxic or non-toxic. We trained RoBERTa for 2 epochs using the Adafactor optimizer Shazeer and Stern (2018) with the learning rate 1e-5. The obtained classifier achieves 82.1 accuracy and 80.4 macro F1 on the BAD test set.
In § 3.2, we follow Liu et al. (2021) and use the HuggingFace sentiment classifierhttps://huggingface.co/distilbert-base-uncased-finetuned-sst-2-english as the label function , which is a 66M distilled BERT model Sanh et al. (2019).
B.4 Results on Validation Sets
We report the automatic evaluation results on the validation sets in Table 12 and 13. Note that in the task of sentiment steering (§ 3.2), we follow Liu et al. (2021) and do not use validation data.
Appendix C Human Evaluation Details
We designed the human evaluation protocols primarily following previous work Liu et al. (2021); Su et al. (2022); Lu et al. (2022).
We randomly sampled 100 prompts (dialogue histories) from the BAD test set. For each prompt, one generated response of Click and one of the baseline was compared and judged by three human annotators from Amazon Mechanical Turk. The evaluation considers the three aspects: toxicity (which one is less offensive or biased), fluency (which one is more grammatically correct and coherent), and topicality (which one is more natural, relevant, and logical). A screenshot of the main annotation interface is shown in Figure 4, which contains detailed annotation instructions. The human annotation achieved fair to moderate inter-annotator agreement (Fleiss’ Kappa in Table 2).
C.2 Sentiment Steering
Similar to above, we randomly sampled 100 prompts from the negative/positive prompts from Liu et al. (2021). The evaluation considers the three aspects: sentiment (which one is more positive/negative), fluency, and topicality. A screenshot of the main annotation interface is shown in Figure 5. The human annotation achieved fair to moderate inter-annotator agreement (Table 4).
C.3 Repetition Reduction
We randomly sampled 100 prompts from the WikiText-103 test set. The evaluation considers the three aspects: coherence (which one is more aligned in meaning/topic with the prompt), fluency (which one is more grammatical, understandable, and non-repetitive) and overall quality. A screenshot of the main annotation interface is shown in Figure 6. Note that unlike Su et al. (2022); Lu et al. (2022), we did not adopt the Likert Scale to rate each generation sample since we found this led to higher annotation difficulty and lower inter-annotator agreement. We instead adopted pairwise comparison as in the former two tasks. The human annotation achieved fair to moderate inter-annotator agreement (Table 6).
Appendix D Qualitative Results
We provide additional qualitative results of the three tasks in Figure 7, 8, and 9, respectively.