Coherence boosting: When your pretrained language model is not paying enough attention
Nikolay Malkin, Zhen Wang, Nebojsa Jojic
Introduction
Language models (LMs) are commonly evaluated for their ability to generate, rank, or classify coherent spans of text. Long-range semantic coherence is a unifying feature of modern NLP benchmarks and applications, whether they are about producing short answers to questions, ranking answer choices by their consistency with world knowledge, or generating long responses.Code: github.com/zhenwang9102/coherence-boosting.
Large nonspecialized LMs, such as GPT-2 and -3 (Radford et al., 2019; Brown et al., 2020), sometimes fail to understand or use the semantic link between a text and its prompt or long-range context (Fig. 1). Samples from these LMs have an unnaturally low density of words that require many tokens of context to predict (§4.1), and the scores that the models give to completions of prompts indicate that they are oversensitive to recent context (§5).
We hypothesize that these failures arise from modeling choices and distribution shift. Specifically, autoregressive LMs are typically fit to a multi-objective problem: simultaneously maximizing token likelihoods conditioned on many lengths of truncated context (§2.1). Yet, at generation or scoring time, likelihoods are conditioned on the entire prompt or previously generated string, specifically selected to be coherent or even guaranteed to influence the output. The two common solutions – finetuning models on one or multiple tasks (Khashabi et al., 2020; Sanh et al., 2022) and improving models or prompts to facilitate in-context learning (Brown et al., 2020; Schick and Schütze, 2021) – do not directly target the problem of long-range coherence.
This paper proposes coherence boosting, a simple inference-time procedure that increases the effect of distant words on predicted token distributions and is applicable in both generation and ranking settings. A pretrained model is viewed as an ensemble of experts that produce token distributions conditioned on varying lengths of context. These experts are log-linearly mixed to form a predictor that is superior to the base model (§2).
Coherence boosting greatly improves prediction of words that depend on a long context, as evidenced by state-of-the-art results on tasks specially meant to assess models’ attention to distant words (§3). In generation of generic text and dialog responses, we show that coherence boosting brings the frequency of occurrence of such words close to that seen in natural text (§4). Beyond generation, we study diverse multiple-choice tasks (§5), in which examples are known to be highly coherent. Coherence boosting does not modify the base model and depends on a single parameter than can be estimated in one pass through a validation set, yet is a competitive adaptation algorithm.