RankGen: Improving Text Generation with Large Ranking Models
Kalpesh Krishna, Yapei Chang, John Wieting, Mohit Iyyer
Introduction
Despite exciting recent progress in large-scale language modeling (Radford et al., 2019; Brown et al., 2020), text generated from these language models (LMs) continues to be riddled with artifacts. Modern LMs suffer from the “likelihood trap” (See et al., 2019; Zhang et al., 2021), in which high likelihood (low perplexity) sequences produced by greedy decoding or beam search tend to be dull and repetitive. While truncated sampling methods such as top- (Fan et al., 2018), nucleus (Holtzman et al., 2020), and typical sampling (Meister et al., 2022) alleviate these issues, they can also produce text with inconsistencies, hallucinations, factual errors, or commonsense issues (Massarelli et al., 2020; Dou et al., 2022; Krishna et al., 2021).
Part of the problem is that LMs are trained using “teacher forcing”, where they are always given the ground-truth prefixA prefix is a sequence of tokens fed as input to an LM, which then generates continuations conditioned on the prefix. A prefix is also called a prompt in prior work (Fan et al., 2018). and asked to predict the next token. At test-time, however, the prefix can contain model-generated text, allowing errors to propagate during decoding (Bengio et al., 2015). This issue, combined with the observation that LMs overly rely on local context (Khandelwal et al., 2018; Sun et al., 2021), contributes to the generation of sequences that break coherence or consistency within a larger discourse-level context (Wang et al., 2022).
To address this issue we present RankGen, a 1.2 billion parameter English encoder model that maps both human-written prefixes and model-generated continuations of those prefixes (generations) to a shared vector space. RankGen efficiently measures the compatibility between a given prefix and generations from any external LM by ranking the generations via their dot product with the prefix (Figure 2). We train RankGen using large-scale contrastive learning, encouraging prefixes to be closer to their gold continuation and far away from incorrect negatives. Since our objective considers two sequences rather than just single token prediction, it encourages RankGen to consider longer-distance relationships between the prefix and continuation rather than just local context.
We devise two different strategies (shown in Figure 1) for selecting challenging negative samples, and empirically show that current large LMs cannot distinguish gold continuations from the negatives via perplexity (Section 2.1). In the first strategy, InBook, we select random sequences that occur within the same document as the prefix. While these human-written negatives are fluent and might contain topic or entity overlap, they are irrelevant as continuations to the prefix. In the second strategy, Generative, we generate continuations by conditioning a large pretrained LM on a given prefix. Compared to InBook negatives, these negatives are much more relevant to the prefix, but they suffer from issues like hallucination and repetition.
While RankGen can be easily used to rerank full-length samples from any external LM, we demonstrate further improvements in generation quality when it is integrated as a scoring function into beam search. On automatic and human evaluations across four large pretrained models (345M to 11B parameters) and two datasets, we observe that RankGen significantly and consistently outperforms sampling-based methods (nucleus, typical, top-) as well as perplexity-based reranking (85.0 vs 77.3 MAUVE, 74.5% human preference over nucleus samplingSee Table 3, 5 for all results. MAUVE (Pillutla et al., 2021) is a recently introduced automatic metric for open-ended generation which has high correlation with human judgements.). Additionally, RankGen outperforms newer decoding algorithms like contrastive decoding and search (89.4 vs 84.9 MAUVE on Wikipedia) which were proposed after the initial RankGen release in May 2022. Qualitative analysis from our human annotators (English writers) suggests that most of the improvements stem from increased relevance and continuity between the generated text and the prefix. Finally, we explore applications of RankGen outside of text generation and report state-of-the-art results on two complex literary retrieval benchmarks: RELiC (Thai et al., 2022) and ChapterBreak (Sun et al., 2022). We open source code, data and model checkpoints.1
RankGen: a generation ranker
RankGen is a deep encoder network that projects prefixes and generations to a shared vector space. Given a prefix vector and a generation vector, we compute a score for the generation via the dot product between the two vectors. To ensure that these scores are meaningful, we train RankGen using large-scale contrastive learning (Radford et al., 2021), pushing the prefix vector close to the gold completion and away from the vectors of negative samples (Figure 1). We use two types of negative samples for learning the metric space: (1) sequences at random locations in the same document (InBook), and (2) model generations (Generative). This section empirically justifies our negative sample choice (Section 2.1) before presenting a precise model formulation (Section 2.2).
We explicitly choose our negatives to focus on a weakness of modern LMs which we empirically verify below: LMs often assign high probability to implausible or irrelevant continuations of a prefix.
Our first type of negative samples are sequences from random locations in the same document as the prefix, whose lengths match those of the ground-truth continuations. As these negatives are written by humans, they are always fluent and coherent, and often topically similar to the prefix (with overlapping entities). However, they are irrelevant as continuations to the prefix, breaking discourse-level continuity and coherence (Hobbs, 1979; Grosz et al., 1995).
Given a prefix of 256 tokens from Wikipedia or a PG19 book (Rae et al., 2019), we measure how often LMs assign higher probability (lower perplexity) to the gold 128-token continuation over a single InBook negative.We experiment with multiple InBook negatives in appendix §C.2. This task is similar to suffix identification tasks like ROCStories Mostafazadeh et al. (2016); see §C.5 for experiments on them. We break all prefixes and continuations at sentence boundaries to make the task less reliant on local syntactic patterns. Table 1 shows that even large LMs perform far below human estimates on this task (63.3% for GPT2-XL vs 91.0% human on Wiki),Human study done on Upwork; details in Appendix B. and repeating this experiment with “hard” negatives selected from a trained RankGen model drops LM performance even further (50.6% for GPT2-XL vs. 90.5% human on Wiki).See Appendix C.1 for more details on “hard negatives”. We hypothesize that LMs perform poorly because (1) they overly focus on local context instead of long-range dependencies from the prefix (Khandelwal et al., 2018; Sun et al., 2021); and (2) LMs assign high likelihood to words with high frequency in their training data (Holtzman et al., 2021) which may occur in InBook but not in the gold continuation. We analyze the latter further in Appendix C.6 using alternative scoring functions like PMI.
Our second type of negative samples are continuations to a prefix that are generated by a pretrained LM. Machine-generated text is known to differ significantly from human text, containing repetitions, hallucinations, and artifacts (Zellers et al., 2019b; Maynez et al., 2020; Holtzman et al., 2020). We use these negatives to encourage RankGen to prefer generations closer to the human distribution, similar in spirit to GAN discriminators (Goodfellow et al., 2014). Generative negatives have also been used in previous energy-based LMs (Deng et al., 2020), although not at this scale; see Section 5 for more related work. In Table 2, we show that LM perplexity is poor at identifying human text over Generative negatives (GPT2-XL gets just 26.5% accuracy, well below 50% random chance). This relates to prior work showing LMs have high confidence in machine-generated text (Gehrmann et al., 2019), especially their own (Appendix C.3).
2 Training RankGen
Having motivated our negative sampling strategies, we now describe RankGen’s training process. We train RankGen using large-scale contrastive learning with in-batch negative sampling, which is a popular metric learning technique (Sohn, 2016) previously used for dense retrieval (DPR, Karpukhin et al., 2020), image classification (SimCLR, Chen et al., 2020), and multimodal representation learning (CLIP, Radford et al., 2021).
A single RankGen training instance consists of a triple , where is a prefix, is the ground-truth continuation of that prefix, and is a continuation generated by an LM. We prepend a special token (pre) to each prefix, and suf (suffix) to each continuation and generation. We then pass each element of the triple through a shared Transformer encoder (Vaswani et al., 2017), projecting them to fixed-size vectors (, , ) using the representation of the special token. To train this model, we use a contrastive objective that pushes the prefix vector close to the gold continuation vector , but away from both the generation vector as well as all other continuation vectors in the same minibatch (“in-batch negative sampling”),
where is a minibatch. All minibatch elements are sampled from the same document, which provides the InBook negatives. Note that the minibatch size is an important hyperparameter since it determines the number of negative samples; we set for our XL variant. See §A.1 for training details and sizes of model variants.
Dataset construction: We consider all possible 256-word prefixes in our document, ensuring that prefixes begin and end at sentence boundaries. We then select continuations of variable length (10-128 words long) for each prefix so that RankGen can re-rank candidates of different lengths at test-time. To produce Generative negatives, we first use 50% of our (, ) training data pairs to fine-tune T5-XXL (Raffel et al., 2020) for causal language modeling (one per domain). For the remaining half of the dataset, we use this LM to generate a single continuation to the prefix of variable length (10-128 words) using nucleus sampling (Holtzman et al., 2020) with .
3 Using RankGen at inference
After model training, the dot product between the prefix and continuation vectors denotes their compatibility score. We experiment with two strategies for using these scores during generation: (1) over-generation and reranking, in which we use any pretrained LM and decoding algorithm to generate multiple samples (20 in our experiments) and then re-rank them; and (2) beam search (Figure 2), in which we generate samples of length via nucleus or ancestral sampling, compute the top highest-scoring samples via RankGen, and concatenate them to the prefix to continue generation. There are three hyperparameters for our beam search: (i) the rerank length , or the number of tokens generated before each re-ranking; (ii) the beam size ; and (iii) the number of samples generated per beam . Setting =20, =1, =128 (max generation length) is equivalent to the first strategy of over-generation and re-ranking. Details of our implementation and hyperparameter search are in Appendix A.2, A.3. Overall all tested hyperparameters improve over baselines, but =10, =2, =20 performs best but all tested hyperparameter choices improve over baselines (Figure 3).
Experiments
We study four configurations of RankGen, each with 1.2B parameters (XL size) and trained with minibatch size 1536. Three variants are trained on the PG19 dataset (Rae et al., 2019), which consists of long-form books, using (1) only InBook negatives, (2) only Generative negatives, and (3) both types of negatives. Since PG-19 contains mainly historical literature, we also experiment with different data sources by training RankGen on the union of four domains (“all”) — PG19, Wikipedia, C4-NewsLike and C4-WebTextLike (Raffel et al., 2020). This last model is trained using both types of negatives. More ablations varying the model size and minibatch size (number of negatives) are provided in Appendix E.
Does RankGen improve generation quality regardless of the size and pretraining dataset of the LM? To check this we evaluate four different pretrained LMs whose sizes vary considerably from that of RankGen (1.2B parameters). We experiment with two variants of GPT-2 (Radford et al., 2019): GPT2-medium (345M) and GPT2-XL (1.5B parameters). We also evaluate a pretrained T5-XXL-v1.1 (Raffel et al., 2020) model (11B parameters) that we fine-tune to perform language modeling on the training set of PG19 (Rae et al., 2019). Finally, to experiment with a large LM trained on out-of-domain data for RankGen-PG19, we evaluate the T5-XXL model from Lester et al. (2021) (11B parameters) that was fine-tuned for language modeling on the C4 corpus.
2 Open-ended text generation
Following prior work on text generation (Welleck et al., 2019; Holtzman et al., 2020; Su et al., 2022), we primarily focus on open-ended text generation, which has wide applications for tasks such as generating stories (Fan et al., 2018), poetry (Zhang and Lapata, 2014), and dialog (Miller et al., 2017) and few-shot NLP (Brown et al., 2020). We consider two domains in our study: (1) prefixes from Wikipedia, and (2) literary text from PG19 (Rae et al., 2019). Since it is difficult to conduct human evaluations of long sequences of machine-generated text (Karpinska et al., 2021), our main experiments consider a 256-token prefix and 128-token generations. We analyze generation quality given varying prefix lengths in Section 4.3.
For each LM considered we decode outputs using greedy decoding, ancestral sampling, nucleus sampling (Holtzman et al., 2020), top-k sampling (Fan et al., 2018), and typical sampling (Meister et al., 2022). Since RankGen is fundamentally a re-ranker of multiple samples, we also compare to two other re-rankers using LM perplexity and unigram overlap, respectively. In all re-ranking settings, we generate 20 samples and then re-rank them with each method. For RankGen, we also use beam search (§2.3) that re-ranks partially generated hypotheses.
In addition to these baselines, in Table 4 we also compare RankGen to newer decoding algorithms proposed after the RankGen release (May 2022).
We use MAUVE (Pillutla et al., 2021) as our primary metric for automatic evaluation. MAUVE computes the similarity of the distribution of human-written text and machine-generated text, and has high correlation with human judgments.Details about our MAUVE setup in Appendix D.1. More evaluations with metrics like rep Welleck et al. (2020) in Appendix D.3. Since automatic metrics are insufficient for text generation evaluation (Celikyilmaz et al., 2020), we also conduct a human evaluation by hiring English teachers and writers from Upwork;https://www.upwork.com see Appendix B for more details. For each of GPT2-medium and T5-XXL-C4 we choose 50 Wikipedia and 50 PG19 prefixes, and show three annotators a pair of continuations from different decoding strategies in a random order (blind A/B testing). Annotators are asked to choose the better continuation and provide a 1-3 sentence explanation for their choice. This gives us 600 annotations, analyzed in §3.4, 4.1.
3 Results from automatic evaluations
Table 3 contains MAUVE scores for all decoding configurations and datasets. Overall, we see that:
Re-ranking full-length samples with RankGen yields an average MAUVE score of 83.4 across all configurations, significantly outperforming other decoding strategies like greedy decoding (15.4), ancestral sampling (74.8), and nucleus / top-k / typical sampling (77.1-77.4). Adding beam search further boosts performance to 85.0.Hyperparameter grid search details in Appendix A.3. Surprisingly, re-ranking 20 full-length ancestral samples with RankGen performs better than standard nucleus sampling (77.3 vs 82.6). However, re-ranking 20 ancestral samples is slightly worse than re-ranking 20 nucleus samples (82.6 vs 83.4) due to worse inherent quality of ancestral vs nucleus (74.8 vs 77.3). Re-ranking generations by unigram overlap to the prefix is a surprisingly good baseline (79.7), while re-ranking by LM perplexity reduces MAUVE to 65.2, since it emulates likelihood-based methods like greedy decoding. Finally, RankGen performs best on in-domain data, with the PG19-XL-both variant obtaining better scores than the model trained on four domains (80.7 vs 73.0 on T5-XXL-C4, PG19).
In Table 3 (bottom), we perform ablations by removing the InBook and Generative for RankGen PG19 variants. All three variants outperform nucleus sampling (77.3), but keeping both objectives performs best (82.6). A model trained with only InBook is more effective (81.4) than one trained with only Generative (80.2).
Since the release of RankGen in May 2022, several new decoding algorithms have been proposed including contrastive search Su et al. (2022); Su and Collier (2022), contrastive decoding (Li et al., 2022), and eta sampling (Hewitt et al., 2022). In Table 4, we compare RankGen to these newer methodsWe use the official implementations for all these methods. Links - contrastive search, contrastive decoding, eta sampling. on GPT2-md and GPT2-XL. Overall, we find that RankGen significantly outperforms all newly proposed decoding algorithms (89.4 vs 84.9 on GPT2-XL wikipedia against the best baseline contrastive decoding).
4 Human evaluation with A/B tests
Despite the high human correlation of MAUVE, human evaluation remains critical for open-ended generation (Celikyilmaz et al., 2020; Gehrmann et al., 2022). Since human evaluation is expensive, we focus on comparing our best performing RankGen variant (RankGen-XL-all with beam search) to nucleus sampling, one of the most popular decoding algorithms in use today. We conduct blind A/B testing comparing the two methods, hiring English teachers and writers on Upwork (§3.2). Table 5 shows that humans significantly prefer outputs from RankGen over nucleus sampling (74.5% preference by majority vote, ). RankGen preference is higher with more inter-annotator agreement (Table 6) for outputs from the smaller GPT2-medium. Finally, humans show slightly higher RankGen preference for Wikipedia generations compared to PG19.
Analysis
To get more insight into the human preference judgments made in Section 3.4, we asked our annotators to provide a 1-3 sentence free-form explanation for each of their choices.All 600 human explanations are provided in submission. We manually categorized each of 600 explanations into nine broad categories loosely based on the Scarecrow schema designed by Dou et al. (2022). In Table 7 we see that 81% of the explanations preferring RankGen mentioned some aspect of the relationship between the prefix and the generated text, including relevance, continuity, and stylistic similarity. 8.0% of the explanations said that RankGen outputs displayed fewer commonsense errors, while 4.7% said that they were less repetitive. We show some generations and human explanations in Table 8 and several more full-length generations in Appendix F.
2 How fast is decoding with RankGen?
Our algorithm requires over-generation followed by RankGen re-ranking. How much extra decoding time does this add? In Figure 3, we show the trade-off between MAUVE score and decoding time across different hyperparameters.Timing depends on library / hardware. We analyze HuggingFace on RTX3090, T5X on TPU-v3 in appendix A.2. While decoding a single nucleus sample takes just 0.8 seconds, generating 20 samples followed by re-ranking with RankGen requires 2.5 seconds. The best-performing hyperparameters use multiple re-ranking steps, taking 5.9 seconds.See Appendix A.3.2 for more speed tradeoff plots. In Appendix A.3.2, we see that over-generation is the bottleneck, since re-ranking takes only a fraction of the time (1-10%) compared to generation. Developing methods that avoid over-generation (e.g., via distillation) is an exciting future work direction.
3 Generation with different length prefixes
Our RankGen model is trained with a fixed prefix length of 256 tokens, and all of the evaluations in Section 3 also assume a prefix length of 256 tokens. However, many text generation applications take shorter prefixes as input, like short writing prompts in story generation (Fan et al., 2018). How well does RankGen generalize to shorter and longer prefixes? Figure 4 compares nucleus sampling to RankGen across varying prefix lengths. We observe that RankGen consistently outperforms nucleus sampling in terms of MAUVE, and beam search with RankGen always provides further gains, suggesting robustness to the prefix length.
4 RankGen as a retriever
While we designed RankGen for text generation, we find that it is also an effective zero-shot retriever. RankGen follows a dual encoder architecture similar to those of several recent dense retrievers like DPR (Karpukhin et al., 2020) and REALM (Guu et al., 2020). We test RankGen on RELiC (Thai et al., 2022), a complex literary retrieval task. Given a literary analysis excerpt, systems must retrieve a quote from a book which is most relevant to the excerpt. RELiC requires a deep understanding of literary phenomena (like irony, metaphors, co-reference, style), and current retrievers struggle on it. We test models in a zero-shot setting, without finetuning on RELiC training data. In Table 9 we find RankGen significantly outperforms other retrievers, achieving a new state of the art on RELiC.https://relic.cs.umass.edu/leaderboard.html PG-XL-InBook performs best (6.0 vs 2.9 recall@1 against the next-best ColBERT), approaching a fully supervised upperbound (9.4). While our XL model has many more parameters than baselines, even PG-base-both outperforms all baselines (3.8 vs 2.9), which has a similar number of parameters as our baselines. Dropping InBook leads to poor performance (0.7), further confirming its efficacy. Besides RELiC, we investigate retrieval over PG19 books in appendix §C.2, and suffix identification in §C.5, achieving state-of-the-art on ChapterBreak (Sun et al., 2022).
Related Work
Our work on RankGen draws inspiration from previous research on self-supervised learning, energy-based models, and modeling non-local dependencies. For instance, our InBook negative sampling is related to popular self-supervised representation learning methods that leverage discourse information across multiple sentences, which is useful for learning sentence embeddings (Kiros et al., 2015; Hill et al., 2016; Jernite et al., 2017). Our formulation is most similar to QuickThought (Logeswaran and Lee, 2018), which uses in-batch negative sampling on a contiguous set of sentences. More recently, the next sentence prediction task has been used for pretraining large LMs (Devlin et al., 2019; Lan et al., 2020; Aroca-Ouellette and Rudzicz, 2020). Unlike these works, we focus specifically on text generation rather than self-supervised pretraining for natural language understanding tasks.
RankGen is also closely related to efforts in energy-based methods (LeCun et al., 2006) for generative modeling (Grover et al., 2019; Parshakova et al., 2019), speech recognition (Wang and Ou, 2018), open-ended text generation (Bakhtin et al., 2019; Deng et al., 2020), machine translation (Shen et al., 2004; Lee et al., 2021; Bhattacharyya et al., 2021), constrained generation (Qin et al., 2022; Mireshghallah et al., 2022), and models for specific attributes like style (Dathathri et al., 2020; Yang and Klein, 2021), length (Li et al., 2017), or repetition & relevance (Holtzman et al., 2018). Unlike prior work, we use human-written text from the same document as negative samples (InBook) in addition to machine-generated text. RankGen is also trained at a much larger scale than prior energy-based models for text (1.2B parameters, contrastive learning with 3K negatives on 4 domains).
Finally, RankGen is related to efforts in modeling non-local dependencies in generation, which include methods that predict multiple tokens (Oord et al., 2018; Qi et al., 2020), rely on retrieval (Khandelwal et al., 2020), use bidirectional LMs (Serdyuk et al., 2018), employ contrastive learning (Su et al., 2022; An et al., 2022), use BERT for sentence-level language modeling (Ippolito et al., 2020), and designing sequence-level losses Wiseman and Rush (2016); Edunov et al. (2018); Welleck et al. (2020); Liu et al. (2022) for reducing exposure bias (Bengio et al., 2015; Ranzato et al., 2016). While the RankGen approach is significantly different from these prior works, it can be intuitively viewed as a “-word sequence-level” language modeling approach, which is discriminative rather than generative.
Conclusion and Future Work
We present RankGen, a large encoder which scores continuations given a prefix and can be plugged into any text generation system. RankGen significantly outperforms popular decoding methods on both automatic and human evaluations. We note several exciting future directions for RankGen, including:
training (or adapting) a multilingual variant of RankGen, as our current models are trained on English text only
training larger RankGen models (T5-XXL size or bigger), with longer prefix / suffix lengths, to see if generation quality continues to improve with scale
exploring the utility of RankGen in other generation tasks like dialog generation, summarization, or long-form question answering
RankGen re-ranking of significantly larger hypothesis sets generated using search algorithms like that in Xu et al. (2022)
more directly incorporating RankGen into generative modeling to eliminate the need for over-generation, either via gradient-based sampling (Qin et al., 2022), distilling RankGen knowledge into LMs via unlikelihood training (Welleck et al., 2020) or reward modeling with RL (Ouyang et al., 2022)
using RankGen as a retriever in knowledge retrieval augmented generation (Nakano et al., 2021; Komeili et al., 2022)
further exploring the capability of RankGen as a retriever, either zero-shot or by fine-tuning on retrieval benchmarks like BEIR (Thakur et al., 2021)
utilizing of RankGen as a text generation evaluation metric like CARP (Matiana et al., 2021) or CLIPScore (Hessel et al., 2021)
using RankGen on other domains with sequential data, like code completion, protein synthesis, or generating mathematical proofs.
Limitations
An important limitation of RankGen compared to other decoding methods is the need for over-generation, which we discuss in Section 4.2. While RankGen itself is efficient, generating multiple samples increases decoding time by an order of magnitude. RankGen is a re-ranking method, so it relies on other decoding methods to produce the candidate output set. Biases in the output candidate set from existing decoding algorithms may be present in RankGen outputs. Besides this, RankGen may be vulnerable to adversarial examples (Szegedy et al., 2013) — gibberish text which gets high RankGen score, obtained by white-box attacks (Ebrahimi et al., 2018; Wallace et al., 2019).
This study is limited to open-ended text generation, which has a large space of possible outputs. RankGen or our findings may not be directly applicable to other generation tasks which have a more constrained output space like summarization, long-form QA or machine translation.
Acknowledgements
We are very grateful to the freelancers on Upwork and volunteers who helped us evaluate generated text. We thank Xavier Garcia and the T5X team for helping us with technical issues related to the T5X library. We are grateful to William Cohen, Elizabeth Clark, Marzena Karpinska, Tu Vu, Simeng Sun, Ari Holtzman, Slav Petrov, Ciprian Chelba, Nader Akoury, Neha Kennard, Dung Thai, the UMass NLP group and the Google AI language research group in Pittsburgh for several useful discussions during the course of the project. This work was mostly done while Kalpesh Krishna (KK) was a student researcher at Google Research hosted by John Wieting. KK was partly supported by a Google PhD Fellowship awarded in 2021.
Ethical Considerations
Current text generation technology produces fluent outputs but suffer from several issues like factual inaccuracies, lack of faithfulness to the input prefix, commonsense issues etc., which makes their real-world deployment difficult. RankGen is an effort at rectifying some of these issues, with a focus on faithfulness to input prompts. However, RankGen outputs continue to be factually inaccurate at times, as noted by some of our human annotators. This should be strongly considered before any direct deployment of this system. To tackle this issue, using RankGen for retrieval augmented generation (Nakano et al., 2021) is a promising direction for future work. We have also open-sourced all 600 human annotations, which have detailed explanations highlighting the strengths / weaknesses of RankGen compared to nucleus sampling.
Our final XL-sized models were trained using a Google Cloud TPUv3 Pod slice with 128 chips for a total of 2 days per model. Several similarly-sized models were trained during the development of this project, roughly one XL-size model every week from October 2021 to February 2022. Due to expensive training costs, we have open-sourced our model checkpoints for the community to use and build upon. Note that “TPUs are highly efficient chips which have been specifically designed for machine learning applications” as mentioned in the Google 2020 environment report. These accelerators run on Google Cloud, which is “carbon neutral today, but aiming higher: our goal is to run on carbon-free energy, 24/7, at all of our data centers by 2030.” (https://cloud.google.com/sustainability). More details on model size and training are provided in Appendix A.1.
References
Appendices accompanying “RankGen: Improving Text Generation with Large Ranking Models”
Appendix A More RankGen details
We fine-tune the encoder of the T5 v1.1 models from Raffel et al. (2020) using large minibatches (see Table 10 for sizes) on a Cloud TPU v3 Pod slice with 128 chips. Our models are implemented in JAX (Bradbury et al., 2018) using the T5X library (Roberts et al., 2022). Each model was fine-tuned for 100k steps, using a constant learning rate of 0.002 using the Adafactor optimizer (Shazeer and Stern, 2018).
A.2 Implementation and timing details
In Figure 5 we provided a simplified Python implementation (without minibatching) of our RankGen beam search algorithm. We implement this algorithm in two libraries — the first uses PyTorch with the popular HuggingFace Transformers library (Wolf et al., 2020), which we test on a RTX 3090 GPU with 25GB memory. The second uses JAX (Bradbury et al., 2018) with the T5X library (Roberts et al., 2022), and is tested on a single Cloud TPU v3 board with 32GB memory.https://cloud.google.com/tpu/docs/system-architecture-tpu-vm#single_tpu_board While measuring decoding time for various hyperparameters (Appendix A.3.2), we focus on throughput (Dehghani et al., 2022), measuring wall-clock time after minibatching to the extent the hardware permits. We ensure consistent experimental settings across hyperparameters, using the same machine and making sure no other computationally expensive process is running on it.
A.3 RankGen hyperparameter grid search
Our hyperparameter grid search is conducted on Wikipedia data with the smallest model considered (GPT2-medium), using MAUVE as our hill-climbing criteria. Our RankGen algorithm has three main hyperparameters — rerank length , beam size and number of samples per beam . The rerank length denotes the number of new tokens which are generated before a re-ranking step takes place. Number of samples denotes the number of generated sequences for each beam. The number of samples retained across different re-ranking cycles is the beam size (see Figure 5 for exact implementation). Our RankGen grid search is conducted over the following configurations —
rerank length : 5, 10, 20, 50, max_length tokens number of samples (beam size * number of samples in every beam ):
1 sample — (1 * 1); 5 samples — (1 * 5); 10 samples — (1 * 10); (2 * 5); 20 samples — (1 * 20); (2 * 10); (4 * 5); 40 samples — (1 * 40); (2 * 20);
Additionally, we measure the extent to which full-length reranking works ( = max length, = 1) by simply increasing the number of samples over-generated and then for re-ranking.
In Figure 6 we study the MAUVE performance tradeoffs for different hyperparameter configurations for the GPT2-medium model evaluated on Wikipedia data. Overall, we observe —
Across all hyperparameter configurations, RankGen significantly improves MAUVE score over a no re-ranking baseline.
MAUVE scores improve for shorter rerank lengths, justifying the benefit of beam search over re-ranking of complete generations.
For cases of full re-ranking (re-rank length = max length), increasing number of samples improves the MAUVE score (since RankGen has more generations to choose from), but improvements saturates after 60 samples (for both model sizes), with the largest gain from 1 to 10 samples.
We find that rerank length = 20 with 20 samples (beam size 2, samples per beam 10) performs best across all configurations.
A.3.2 Speed tradeoffs
In Figure 7 we study the average time taken (in seconds) for a single generation on Wikipedia. Overall, in both our implementations we observe that —
Decoding a single sample is an order of magnitude faster than decoding multiple samples (“over-generation”), which is needed before any re-ranking with RankGen is possible.
Reducing the rerank length increases decoding time, since more generate / re-rank cycles are needed. These cycles cannot be parallelized since the generate and re-rank steps are dependent on each other.
Overall, we see observe that decoding time is roughly , where is beam size, is the number of samples per beam and is rerank length. This is especially true for the T5X implementation.
We dig a little deeper into these numbers: is the extra compute time due to over-generation (generation of 10 or 20 samples instead of one) or RankGen re-ranking? In Table 11, we measure the time taken to generate and score an individual instance. We see that re-ranking with RankGen takes only a fraction of the time (1-10%) compared to generation, which means that over-generation is the bottleneck. Also see Section 4.2 in the main body of the paper for a performance / time tradeoff scatter plot.
Appendix B Human Evaluation Details
We hired freelancers from Upworkhttps://www.upwork.com as well as two volunteers to perform our human evaluation. In total, our human evaluation had eight annotators. Following recent recommendations from Karpinska et al. (2021), we ensured that each annotator (except one) was either an English teacher or an English writer. To avoid bias, we ensured that none of the annotators were computer science researchers, making them unaware of text generation research / RankGen.
Setup: Annotators were shown a 200-250 word prefix, and were asked to choose one of two 80-100 word continuations. Annotators were not told which model generated each continuation, and we shuffled the continuations in a random order to avoid position biases (“blind A/B testing”). The job posting and instructions shown to the annotators are provided in Table 24. We used Amazon Mechanical Turk Sanboxhttps://requestersandbox.mturk.com/ to collect our annotations, using the interface shown in Figure 10. Note that we used the MTurk Sandbox interface only — no MTurk workers are recruited in our human study due to poor annotation quality for open-ended text generation (Karpinska et al., 2021; Clark et al., 2021).
Screening: To ensure high annotation quality, we first asked annotators to complete a small screening test of 20 pairs with InBook distractors, keeping 80% accuracy as our passing criteria (estimated human performance on this set is 90-95%). We paid annotators 10$ for the screening test. Around half the interviewed Upworkers passed the test.
Main Task (comparing generations): In our main task comparing generations from RankGen with nucleus sampling, we asked annotators to choose the better continuation as well as provide a 1-3 sentence free-form explanation for their choice. We paid annotators 1 bonus at the end of a 100 pairs. Each annotator was provided with 100 instances (50 each from Wikipedia and PG19) either generated by the T5-XXL-C4 model (Lester et al., 2021) or GPT2-medium (Radford et al., 2019), with beam search outputs from RankGen-XL-all. Three annotators rate each model, giving us a total of 600 human annotations with explanations.
Main Task (InBook human estimate): Our second main task involved choosing the gold human-written continuation vs random InBook negatives. We paid annotators 0.5$ for this task, and did not ask them to explain their choices. This main task was similar in nature to our screening task.
Appendix C Suffix Identification
In Section 2.1 and Appendix C.2 we make use of “hard negatives”. To select these harder negative from the document, we use a trained RankGen model (XL sized, trained on all four domains). Specifically, we use RankGen to score the compatibility of every 128-word token sequence in the document to the prefix, and take the highest scoring 10 sequences that are not the gold continuation (“Hard” negative). All negatives sequences start and end at sentence boundaries so that LMs cannot rely on local syntactic patterns. For our two-way classification experiments in Section 2.1, we consider a random sequence among these 10 hard negatives. Since RankGen-all-XL-both was used to find these hard negatives, results on this RankGen variant are not very meaningful (since they are adversarial to this variant by construction).
C.2 Gold vs InBook - more negatives
In Section 2.1, we used a single InBook to test models. How do models fare when they need to choose the gold continuation over multiple InBook negatives? In Table 12 we perform experiments on a 11-way classification task (10 InBook negatives). Overall, we find that most LMs do barely above chance, whereas RankGen significantly outperforms large LMs (even GPT3).
Gold vs all InBook negatives (“retrieval”): What if instead of 10 negatives, we used all possible InBook negatives in the book? This task could be framed as a retrieval problem akin to RELiC (Section 4.4): given a prefix, find the correct continuation from all possible continuations in the same book. Since PG19 books can be quite long, retrievers needs to search among 2538 candidates on average in the PG19 validation set. We present results on this retrieval task in Table 13. Overall, we find that RankGen is quite successful at this task, getting a recall@1 of 48.2% with a model trained on just PG19 data and InBook negatives. Training on just PG19, increase model size, increasing minibatch size and using just InBook negatives helps improve retrieval performance. In initial experiments, we extensively used performance on this task to hill-climb and justify our design choices. Note that we do not test LMs on this retrieval task, since it is computationally expensive to do a forward pass for each of the 2538 candidates for each of the 100K datapoints.
C.3 Gold vs Generative - breakdown by generative model
See Table 14 for a breakdown by the model used to create the Generative negatives.
C.4 Details of Suffix Identification Datasets
ChapterBreak (Sun et al., 2022) is a 6-way classification task in which models are provided as input a long segment from a narrative that ends in a chapter boundary. Models must then identify the correct ground-truth chapter beginning from a set of negatives sampled from the same narrative — a task requiring global narrative understanding. ChapterBreak has two settings: (1) PG19 — the validation set of the Project Gutenberg language modeling benchmark Rae et al. (2019); (2) AO3 — a ChapterBreak split adapted from fan-fiction posted to Archive of Our Own (AO3).https://archive.org/download/AO3_story_dump_continuing Although Sun et al. (2022) provide prefixes up to 8192 tokens, we study ChapterBreak in the setting using just 256 tokens of prefix to ensure compatibility with the input lengths of RankGen. The ChapterBreak dataset is not divided into validation / test splits, so we simply use the single available split.
HellaSwag (Zellers et al., 2019a) is a 4-way classification task focusing on commonsense natural language inference. For each question, a prefix from a video caption is provided as input and a model must choose the correct continuation for this prefix. Only one out of the four choices is correct – the actual next caption of the video. HellaSwag is scraped from the video captions in ActivityNet (Krishna et al., 2017) and how-to paragraph instructions on WikiHow. We study the setting where each of the 4 endings are complete sentences, which is constructed by prepending ctx_b to the given endings). We use the validation set of the HellaSwag corpus since the test set answers are hidden.
StoryCloze (Mostafazadeh et al., 2016; Sharma et al., 2018) is a 2-way classification task designed to test commonsense reasoning. Systems are provided with the first four sentences of a five-sentence commonsense story, and must choose the correct ending to the story. We used the test set for the Spring 2016 split and the validation set for the Winter 2018 split (due to the hidden test set).
C.5 RankGen for suffix identification
RankGen is trained on a suffix identification objective: given a prefix, choose the gold continuation over InBook and Generative negatives. How well does RankGen learn this task? How does RankGen fare on existing suffix identification benchmarks?
In Section 2.1 we motivated the RankGen design by showing the inability of LM perplexity to prefer the gold continuations over negatives. How does RankGen fare on these negatives? In Table 1 and Table 2 we evaluate the performance at distinguishing gold continuations from negatives, and compare RankGen to large LMs. Since RankGen is directly optimized on this objective, it significantly outperforms large LMs (99.1% vs 78.2% with GPT-3 for InBook). RankGen variants trained on just InBook or just Generative perform best at their respective tasks, but we observe some generalization (InBook model gets 69.8% on Generative PG19 negatives, Generative model gets 80.2% on InBook negatives, both higher than all LMs). Strong performance on Generative could have several applications like fake news detection (Zellers et al., 2019b; Gehrmann et al., 2019), and is an interesting future work direction.
We test RankGen on three existing suffix identification datasets — ChapterBreak (Sun et al., 2022), ROCStories cloze test (Mostafazadeh et al., 2016) and HellaSwag (Zellers et al., 2019a); dataset details are provided in Appendix C.4. To measure their intrinsic capability, models are evaluated zero-shot, without finetuning on training sets.Zellers et al. (2019a) also describe zero-shot HellaSwag experiments, testing models on unseen WikiHow / ActivityNet categories; however they still finetune models on HellaSwag data for seen categories, while we do no such finetuning.
In Table 15 we find that RankGen significantly outperforms all LMs on ChapterBreak (64.3 vs 28.6). RankGen performs comparably to similar-sized GPT2-XL (1.5B parameters) on other tasks, beating it on StoryCloze (75.8 vs 72.6), but slightly worse on HellaSwag (46.3 vs 48.2). Much larger LMs like GPT3 170B (Brown et al., 2020) and PaLM 540B (Chowdhery et al., 2022) perform best on StoryCloze and HellaSwag. Scaling also benefits RankGen (30.4 vs 40.7 on HellaSwag for base vs XL), and we believe further scaling RankGen is a promising direction for future work. We also find InBook negatives are more beneficial than Generative negatives (64.3 vs 33.6 on ChapterBreak PG19). We hypothesize that the different trends on different datasets can be attributed to input length. As seen in Table 15, ChapterBreak has much longer inputs (240 prefix, 153 suffix tokens) than other datasets (35 prefix, 7 suffix tokens for ROCStories). The focus on local context in LMs (Khandelwal et al., 2018; Sharan et al., 2018; Sun et al., 2021) helps with short-range tasks but also likely contributes to their underperformance on complex long-range tasks like ChapterBreak.
C.6 Choice of Scoring Function
It is argued in Holtzman et al. (2021) that average log likelihood is a sub-optimal scoring function when LMs are used to score sequences. In this section, we compare several scoring functions on GPT2-medium. Let be a prefix and be a continuation. We consider: (1) conditional log likelihood (CLL), or ; (2) average conditional log likelihood (avg CLL), or ; (3) average unconditional log likelihood (avg ULL), or ; and (4) pointwise mutual information (PMI), or . We compare these scoring functions on several datasets in Table 16. Overall, we find that PMI is a strong scoring function, outperforming all other functions on four out of five datasets. Length normalized scoring functions (avg CLL/ULL) are better than CLL across all datasets, consistent with findings in prior work (Wu et al., 2016; Koehn and Knowles, 2017; Brown et al., 2020). All scoring functions lag behind RankGen in all five datasets.
Throughout this paper we use “avg CLL” to report suffix identification scores. Length normalized conditional log likelihood is the most closely aligned to how text is generated (sampling from the next-token distribution), and is the objective language models are directly optimized on. However, given the strong performance of PMI compared to “avg CLL” on four out of five datasets, an interesting future direction is studying the benefit of PMI or domain-conditioned PMI (Holtzman et al., 2021) in generating text.
Appendix D More Evaluation Details & Results
We extensively use the MAUVE metric from Pillutla et al. (2021) for automatic evaluation of our model. MAUVE is shown to have high correlation with human judgements of the quality of generated text. We closely follow the best practices listed in the official MAUVE repository,https://github.com/krishnap25/mauve#best-practices-for-mauve which we found critical in preliminary experiments. Specifically,
We ensure that each run has the exact same hyperparameters — using the default hyperparameters in the official MAUVE library.
We use 7713 generations per run, which is the size of our Wikipedia validation set. This follows the suggestion in the official codebase README of having at least 5000 generations for comparing models. While our PG19 validation set is much bigger, we truncate it to 7713 generations since MAUVE scores tend to reduce with more generations.
Since MAUVE scores are higher for shorter generations, we ensure that all tested methods have roughly equal generation lengths, between 70-80 words / 120-130 tokens. We also truncate human text / generations to ensure that each instance ends at a sentence boundary. In initial experiments we observed that truncating consistently for human text and machine text leads to lower MAUVE variation.
Due to variation in MAUVE score from run to run, we average the MAUVE score for nucleus / top-k / typical sampling over five runs. For the T5-XXL-C4 model on Wikipedia with nucleus sampling, the MAUVE scores were [0.803, 0.778, 0.759, 0.785, 0.768], giving a standard deviation of 0.015.
D.2 MAUVE Divergence Curves
The MAUVE metric is the area under a divergence curve, a curve which attempts to analyze the type of errors the model is making. Given is the distribution of human text and is the distribution of machine-generated text, Pillutla et al. (2021) describe two types of errors made by models —
Type I: — False positives, or cases where models generate text which is unlikely to be written by humans, like semantic repetitions common in neural text generators (Holtzman et al., 2020; Zhang et al., 2021).
Type II: — False negatives, or cases where models cannot generate text which is likely to be written by humans, sometimes seen with truncation strategies (See et al., 2019).
In Figure 8 and Figure 9 we plot the divergence curves comparing greedy decoding, nucleus sampling, and full sample re-ranking with perplexity and RankGen. We observe that re-ranking with RankGen increases the area under the curve, whereas re-ranking with model perplexity reduces the area. Re-ranking with RankGen reduces both Type I (bigger intercept on ) and Type II errors (bigger intercept on ). Re-ranking with perplexity leads to higher Type I errors, or more repetition (as also observed in Appendix D.3).
D.3 Token Overlap metrics
In addition to the MAUVE scores calculated in Section 3, we measure token overlap statistics comparing different decoding methods. First, we measure the rep metric from Welleck et al. (2020), which is an approximate measurement of the amount of repetition in generated text. We measure the percentage of generated tokens which are exactly copied from the immediate local prefix of 20 tokens. In Table 17 we find that re-ranking with RankGen slightly reduces rep compared to nucleus sampling (18.9 vs 19.5). We get even lower repetition on the RankGen trained on just generative negatives (17.8), while RankGen trained on just inbook negatives gets 20.0 — thus generative negatives are better at reducing repetition. Re-ranking with perplexity increases rep to 23.9, whereas greedy decoding has the highest repetition of 59.5. This is consistent with recent findings of repetition in greedy decoded outputs (Holtzman et al., 2020; Zhang et al., 2021). Human text is the least repetitive, with a rep score of 15.4.
Next, we measure the fraction of unigrams in the generation which are also present in the prefix. Higher scores could either imply more faithfulness to the prefix (less hallucination), or lower amounts of abstraction. We present two versions of this metric — (1) considering all tokens (Table 18); (2) considering only only lemmatized nouns and numbers (Table 19). Overall, we find that re-ranking samples with RankGen slightly increases this overlap score (19.5 vs 21.7), but re-ranking by token overlap (38.4) or perplexity (25.0) leads to a much higher score. Given the lower MAUVE scores for these two approaches (Table 3), we suspect that token overlap / perplexity re-ranking leads to lower amounts of abstraction / repetitiveness. Human written text has the lowest overlap, perhaps indicating more abstractive text.
Appendix E Ablation Studies
We conduct several ablation studies studying the importance of three aspects — (1) model size; (2) minibatch size, or number of negative samples during contrastive learning; (3) the type of negative samples (inbook, generative or both). Overall, we see clear benefits of increasing model size and increasing minibatch size for suffix identification (Table 20, Table 21) and human-text identification (Table 23). We see a similar, but less prominent trend on MAUVE scores after re-ranking generations (Table 22). For some settings we find that the RankGen-large variant produces slightly better generations than RankGen-XL. We hypothesize this is due to the much larger minibatch used to train RankGen-large models (4096) compared to RankGen-XL (1536) due to memory constraints.
Appendix F More Model Generations
More model generations with human explanations are provided in Table 25 to Table 30. See our Github repository1 for all 600 annotations for the 200 generation pairs.