Bridging the Gap: A Survey on Integrating (Human) Feedback for Natural Language Generation

Patrick Fernandes, Aman Madaan, Emmy Liu, António Farinhas, Pedro Henrique Martins, Amanda Bertsch, José G. C. de Souza, Shuyan Zhou, Tongshuang Wu, Graham Neubig, André F. T. Martins

Introduction

For generation systems to be widely useful, they must generate text that is not only fluent and high-quality, but also closely aligned with human desires and specifications (Vamplew et al., 2018; Hendrycks et al., 2020; Kenton et al., 2021a; Turner et al., 2022; Ngo, 2022). Achieving such ambitious goals requires modern large language models (LLMs) to evolve beyond traditional training methods. Recent improvements in this space have centered on incorporating human feedback (Bai et al., 2022b; Ouyang et al., 2022; OpenAI, 2023a). This feedback serves as a guiding force, steering LLMs toward the desired outcomes, much like feedback mechanisms in physical machines (Åström and Murray, 2021).

Typically, state-of-the-art language generation systems are obtained by training probabilistic, autoregressive LLMs on massive amounts of data using maximum likelihood estimation (MLE). However, the data used to train these models is generally scraped from the Internet, often containing noise, social biases, and errors (Bolukbasi et al., 2016; Dodge et al., 2021). This, when combined with the objective of maximizing the probability of the next token given the previous ones, might result in a misspecification of target behavior (Kenton et al., 2021b), and might lead to models that generate toxic, inaccurate, and unhelpful content (Sheng et al., 2019; Bender et al., 2021).

Exacerbating the problem above is the fact that these models are often evaluated using automatic metrics that compare the generated text with some “reference” text using surface-level features (such as word overlap), which often do not correlate with human-perceived quality of text (Schluter, 2017; Mathur et al., 2020; Gehrmann et al., 2022a), especially when models are optimized for them (Paulus et al., 2017; Amrhein and Sennrich, 2022). This difficulty in evaluation arises partly because, for many tasks, there is not a single correct answer since the same communicative intent can be conveyed in multiple ways.

Leveraging human assessments to evaluate the quality of texts generated by models is then a popular approach. Crucially, considering human-perceived quality can help close the gap between machine and human generated text, and help in addressing the challenges posed by Goodhart’s law: “when a measure becomes a target, it ceases to be a good measure” (Goodhart, 1984). This realization has spurred a growing interest in improving natural language generation systems by leveraging human feedback on model-generated outputs, and has led to the emergence of the first widely-used general-purpose language assistants (OpenAI, 2023a). Human feedback not only enhances system performance, but also serves as a mechanism to steer the system in alignment with desired outcomes or goals (Rosenblueth et al., 1943; Wiener, 1948).

Feedback, as a concept, encompasses a wide range of meanings and interpretations (Wiener, 1948); however, some universal characteristics can be identified, such as its format, its intended results, and the ways it is utilized as a part of the model development process. In this survey, we focus on the role of human feedback for improving language generation. We start by formalizing the notion of human feedback and creating a taxonomy of the different types of feedback in the literature, and of how they have been used (§2). We discuss how we can describe feedback by its format and its objective, in terms of the desired model behavior (§3). We discuss approaches that directly optimize models against human feedback on (their) outputs, for example, using reinforcement learning with human reward functions (§4). We then move to approaches that circumvent the costs of direct feedback optimization by first training feedback models to approximate human feedback, and then improving generation using these proxy models (§5). We discuss existing datasets for human-feedback data, how these datasets are typically collected, and the impact that the collection process might have on the behaviour of the models (§6). Finally, we discuss a recent line of work that reduces the need to collect human feedback by leveraging AI feedback from large language models (§7).

A Taxonomy for Leveraging (Human) Feedback for Generation

Consider a model M:X→YM:\mathcal{X}\rightarrow\mathcal{Y} which, given an input of some type x∈Xx\in\mathcal{X}, outputs text y^∈Y\hat{y}\in\mathcal{Y}. Importantly, while xx can be of any format, we restrict ourselves to cases where yy is in the space of natural language (i.e., Y⊆Σ⋆\mathcal{Y}\subseteq\Sigma^{\star} for some alphabet Σ\Sigma). This general formulation encompasses a wide range of NLG tasks. For example:

Summarization: X\mathcal{X} is the space of documents, and Y\mathcal{Y} the space of possible summaries.

Machine Translation: X\mathcal{X} and Y\mathcal{Y} are the spaces of sentences in the source and target languages, respectively.

Dialog Generation: X\mathcal{X} is the space of possible dialog histories, and Y\mathcal{Y} is the space of possible responses.

Image Captioning: X\mathcal{X} is the space of images, and Y\mathcal{Y} is the space of possible captions.

These models are generally realized as a parameterized, conditional probability distribution Pθ(y∣x)P_{\theta}(y|x), where θ\theta are the model parameters. This distribution is often estimated autoregressively: the probability of a sentence yy given an input xx is decomposed into the product of the probabilities of each token in the sentence, conditioned on the previous tokens. These models are then trained by finding the parameters θ⋆\theta^{\star} that maximize the likelihood of some training data D={(xi,yi)}i=1N\mathcal{D}=\{(x_{i},y_{i})\}_{i=1}^{N}. Then, at inference time, given an input xx, an output y^\hat{y} is decoded from the learned distribution. This decoding can be done, for example, by approximating the most-likely sequence of tokens (M(x)≈arg max⁡yPθ⋆(y∣x)M(x)\approx\operatorname*{arg\,max}_{y}P_{\theta^{\star}}(y|x)) or by random sampling (M(x)∼Pθ⋆(y∣x)M(x)\sim P_{\theta^{\star}}(y|x)).

Evaluating the quality of generated text y^∈Y\hat{y}\in\mathcal{Y} can be challenging due to the complexity and subjectivity of natural language. Various automatic metrics have been proposed for various domains/tasks. These metrics traditionally rely on n-gram matching or other simple heuristics that cannot account for complex linguistic phenomena (such as paraphrasing or stylistic variations) and often fail to capture all the nuances of human judgment (Sai et al., 2022; Gehrmann et al., 2022a). For this reason, for many of these tasks, asking for human feedback is considered the gold standard for assessing the quality of the generated text, and newer learned metrics often aim to approximate the way humans provide feedback (see §5.1).

More formally, we consider human feedback to be a family of functions H\mathcal{H} such that each feedback function h∈Hh\in\mathcal{H} takes an inputAlthough feedback can be provided independently of the input (for example for fluency), we assume some (potentially empty) input for simplicity of notation. x∈Xx\in\mathcal{X} and one or more outputs y1,⋯ ,yn∈Yy_{1},\cdots,y_{n}\in\mathcal{Y} and returns some feedback f∈Ff\in\mathcal{F}:

A simple example of a (human) feedback function is asking humans to say if, given an input, a particular output is good or bad (h:X×Y→{0,1}h:\mathcal{X}\times\mathcal{Y}\rightarrow\{0,1\}). However, more complex feedback functions, such as rankings or natural language feedback, exist and are commonly used (see §3.1).

We note that this framing is a simplification of the real world: often, different humans might provide different (and potentially contradicting) feedback for the same outputs, and a single function might not be able to capture this variability in human opinion (we discuss this further in §6). Finally, while our formalization is flexible, it excludes other approaches where models interact with humans to improve learning, such as active learning and other human-in-the-loop approaches.

2 Taxonomy

Having established a basic mathematical formulation, we now identify four key axes along which we can classify the uses of human feedback:

The format of human feedback can vary, including binary judgments, numerical scores, ordinal rankings, or qualitative natural language explanations.

Depending on the use case of our model, the feedback can have a variety of purposes, ranging from assessing model performance and accuracy to preventing toxicity and harmful behavior.

Human feedback can be incorporated into the training stage to optimize the model parameters directly. Alternatively, it can be used at inference time to guide the decoding process.

Describing Feedback

While ideally, we would use direct feedback from humans whenever possible, the prohibitive cost of its collection means that it is often useful to instead use surrogate models that approximate human preferences.

An important decision to make when we want to improve language generation systems through human feedback is in what format to collect this feedback in. The choice of format has implications on the expressivity of the feedback, the ease of its collection, and how we can use it to improve systems. In particular, the complexity of the feedback format is an important factor: simpler formats are often easier to collect and use as part of the training/decoding process, but contain less information than more “complex” formats, and might not be able to capture important information for improving the system. The choice of format also has implications in the difficulty for humans to give feedback, its consistency/agreement, and the level of rationality of said feedback (Ghosal et al., 2023). Types of feedback are summarized in subsection 2.2 with examples.

Although easy to leverage, numerical feedback suffers from some limitations: depending on the complexity of the generation task, reducing feedback to a single score might generally be a hard and ill-defined task for humans, leading to a costly collection process and problems of subjectivity and variance (see §6.2.1). Furthermore, such feedback might not be suited to distinguish between outputs of similar quality.

An alternative to asking humans to assign a single score to a given input-output pair is asking them to rank multiple possible alternative outputs h : X ×Y_1 ×⋯×Y_n →S_n where SnS_{n} represents the set of all permutations/rankings of nn elements (optionally allowing ties). This has been used extensively in evaluation (Chaganty et al., 2018). Compared to numerical feedback, this format tends to be easier to collect, and, potentially, for this reason, ranking-based feedback tends to be collected to improve model behavior rather than just for evaluation (since the former tends to require more feedback data). Ziegler et al. (2019) and Stiennon et al. (2020) asked humans to rank alternative summaries of the system they are trying to improve. Similarly, Ouyang et al. (2022) collected rankings of alternative responses to an instruction given to the model. They utilized these rankings to enhance the model’s instruction-following capabilities. Subsequent research has also employed ranking-based feedback for the same task (Askell et al., 2021; Bai et al., 2022a, b).

Both numerical and ranking-based feedback lack the ability to capture detailed information about problems with the output, which can be crucial for improving generation systems. Instead of asking humans to rank or score outputs, we can instead ask for natural language feedback. In such cases, the feedback typically provides more detailed information, either highlighting the shortcomings of the current output or suggesting specific actions for improvement. For example, Li et al. (2017) asked humans to give natural language feedback to a dialogue question answering model, including positive or negative feedback, but also possibly providing the correct answer to the model or hinting about it. Tandon et al. (2022) and Madaan et al. (2022) gather natural language feedback on errors present in model-generated graphs and the model’s interpretation of a given instruction. Scheurer et al. (2022, 2023) improve summarization capabilities of language models by asking humans to provide natural language feedback of summaries of the model. Li et al. (2022) collect natural language feedback (in addition to numerical feedback) for responses from a Question Answering (QA) system.

2 Objective

The purpose of collecting feedback is to align the model’s behavior with some (often ill-defined) goal behavior: we might want our summarization model to generate summaries that contain all core information, even if it means they are a bit longer; in commercial machine translation, extra care is given to ensure that models do not mistranslate business-critical information; and in dialogue agents, we might want the model to be able to produce polite and harmless responses. This alignment objective has been studied extensively in the AI safety and alignment literature Bostrom (2014); Amodei et al. (2016); Bommasani et al. (2021). In addition, Kenton et al. (2021b) discuss some behavioral issues in language agents (natural language generation models) arising from a misspecified alignment objective (for example, from noisy labels in the training data), and Leike et al. (2018) proposed using feedback models to tackle the difficulty in specifying this objective.

Bai et al. (2022a) explicitly divided the problem of “aligning” a language model into improving its helpfulness and increasing its harmlessness. Most works implicitly consider either the use of feedback that targets performance factors (such as when targeting overall performance in a task or ability to follow instructions) or harmlessness factors (such as not producing toxic text or providing information that could lead to harm).We mostly ignore the proposed honesty aspect, as none of these works tackle this directly.

Most often, feedback is collected with some helpfulness objective in mind: a necessary (but not sufficient) condition for a helpful system is that it performs the task well, and so feedback related to task performance generally falls under this umbrella. For example, most works in machine translation leverage feedback related to the quality of translation (Kreutzer et al., 2018; Fernandes et al., 2022), which is expected to be correlated with its helpfulness in downstream applications. Similarly, in summarization, most works leverage feedback related to aspects such as relevance, consistency and accuracy (Ziegler et al., 2019; Stiennon et al., 2020) (in short, the quality of the summary). One particularly well-studied feedback objective is the ability to follow instructions (Ouyang et al., 2022): the task of instruction-following can encompass a wide range of other tasks, and using feedback to improve (instruction following) language assistants has been considered a benchmark for the alignment problem (Askell et al., 2021).

Directly Leveraging Human Feedback

Another important alignment objective is harmlessness: we want our models not to produce certain types of output or violate certain norms. Feedback collected in Ouyang et al. (2022) considered aspects such as the toxicity of text (besides the overall ability to follow instructions). Bai et al. (2022a) explored the interaction between the helpfulness and harmlessness objectives, showing a trade-off between both. Thoppilan et al. (2022b) collected feedback on whether their model violates a set of safety objectives and used it to finetune the model. Glaese et al. (2022) also ask humans to provide feedback on the harmlessness of their system, by defining a set of rules and asking humans if the outputs violate these rules. Bai et al. (2022b) showed that feedback produced by LLMs could increase harmlessness without reducing helpfulness.

In an ideal scenario, we would directly leverage human feedback to improve generation: humans would provide the feedback for training or decoding procedures.

Where D\mathcal{D} is the distribution of possible inputs. Various techniques have been suggested to optimize the model parameters, θ\theta, using the collected human feedback. These can be divided into three main categories based on the training mechanisms, which we will call feedback-based imitation learning, joint-feedback modeling, and reinforcement learning (RL).

The feedback-based imitation learning approach involves using human feedback to optimize the model by performing supervised learning with a dataset composed of positively-labeled generations together with the corresponding inputs, D+\mathcal{D^{+}}. This can be achieved by minimizing the loss:

An instance of this approach can be found in Li et al. (2017), in which the authors train a dialogue model by maximizing the likelihood of the model’s answers labeled as correct by humans. Similarly, Kreutzer et al. (2018) trained a machine translation model on a set of positively-labeled translations, and Glaese et al. (2022) performed supervised learning on the preferred dialogues which comply with their pre-defined rules (concerning correctness, harmfulness, and helpfulness), according to humans. A slightly different approach was proposed by Hancock et al. (2019): deploying a chit-chat dialogue model and using the human utterances as targets to fine-tune the model. Scheurer et al. (2022, 2023) leverage the fact that LLMs can follow instructions and start by collecting natural language human feedback about the model generations, which often describes what an improved text would look like. Then, they ask the LM to generate multiple refinements based on the input, previous model generation, and the corresponding feedback. The highest similarity refinements for each generation are then used to fine-tune the LLM. OpenAI’s text-davinci-002 was trained with both human demonstrations and model outputs with the highest possible rating, an approach deemed FeedME (OpenAI, 2023b). A downside of these approaches is that they disregard the generations which do not receive positive feedback, which may contain useful information to optimize the model.

On the other hand, joint-feedback modeling leverages all the information collected by directly using human feedback to optimize the model. Also, as the feedback is modeled directly by the model, this approach allows feedback in formats other than numerical or ranking-based (e.g., natural language). Having D\mathcal{D} as the dataset of inputs xx, generations yy, and human feedback ff collected, this can be achieved by minimizing the following loss of the form

Over all examples in D\mathcal{D}. These equation can be factorized as L(i)(θ)=−log⁡pθ(f(i)∣y(i),x(i))+log⁡pθ(y(i)∣x(i))\mathcal{L}^{(i)}(\theta)=-\log p_{\theta}\left(f^{(i)}\mid y^{(i)},x^{(i)}\right)+\log p_{\theta}\left(y^{(i)}\mid x^{(i)}\right). Some works simply train the model to predict the feedback given to each generation (Weston, 2016, forward prediction), disregarding the second term of the factorization. One example of this approach is the work of Li et al. (2017), in which the authors asked humans to give natural language feedback (e.g., positive/negative feedback, providing the correct answer to the model, or giving a hint about the correct answer) to a dialogue question answering model. Then, after having collected the feedback, the model is trained to predict it. Hancock et al. (2019) proposed having an auxiliary model predicting the satisfaction of the human speaking with the model. Then, if the satisfaction score is lower than a pre-defined threshold, the model will ask the human for feedback. The model then leverages the natural language feedback humans give by learning to predict it. Yuan et al. (2023); Rafailov et al. (2023) showed that having summarization models predict the rankings of different summaries helps the model generate better summaries, and might even outperform more complicated approaches using feedback models (§5).

Other works train the model to predict the generations and the corresponding human feedback. Xu et al. (2022) proposed using the Director model introduced by Arora et al. (2022) to leverage human feedback. As this model has a unified decoder-classifier architecture, Xu et al. (2022) proposed using positively-labeled examples to train its language modeling head (similarly to feedback-based imitation learning) and using both the positive and negatively-labeled examples to train a classifier head that directs the model away from generating undesirable sequences. Thoppilan et al. (2022a) follow this approach to enforce the model’s quality and safety. First, they collect dialogues between crowd-workers and the proposed language model LaMDA, which are annotated with feedback provided by the crowd-workers. This feedback states each response’s quality (sensible, specific, and interesting) or safety. Then, LaMDA is fine-tuned to predict the high-quality responses and the rewards given to every response regarding its quality attributes and safety. At inference time, LaMDA is also used to filter out candidate responses for which its safety prediction is below a threshold.

Finally, this can also be achieved by training the model to predict generation and conditioning on the feedback. This corresponds to minimizing the following loss:

Liu et al. (2023) proposed prompt-based fine-tuning, where they create prompts containing previous generations rated by humans, in the order of preference. They also suggest inserting language-based feedback (e.g., “… is a worse answer than …”) to the prompt, between the generations. Then, the model is fine-tuned to maximize the likelihood of generating the most preferred answer.

Finally, reinforcement learning (RL) offers a more versatile approach, allowing for direct optimization of a model’s parameters based on human feedback, regardless of the feedback’s differentiability. A common RL algorithm used in this context is the REINFORCE algorithm (Williams, 1992), which updates the policy parameters using the following gradient:

Here, D\mathcal{D} represents the set of inputs xx, and pθp_{\theta} is the policy. This flexibility enables RL to handle various types of feedback and better align the generated output with human preferences. For instance, Kreutzer et al. (2018) proposed using task-based implicit feedback from user queries as a reward signal to train a machine translation model using a word-level variant of minimum risk training (Shen et al., 2016), while Jaques et al. (2019) used implicit human reactions in chat to improve open-domain dialog systems through off-policy Q-learning (Watkins and Dayan, 1992). Given that collecting human feedback can be expensive and time-consuming, learning is done offline from logged data, which is typically more favorable than on-policy settings that need feedback on the fly. Later in §5.2.1, we discuss several works that attempt to optimize feedback models using RL instead of directly optimizing human feedback. In conjuction, these aproaches are commonly known as Reinforcement Learning from Human Feedback (RLHF).

2 Decoding with Human Feedback

While directly optimizing model parameters provides greater control, modifying them may not always be feasible, particularly in the case of LLMs. Additionally, feedback might be unavailable during model training, limiting the scope for parameter adjustments. In such cases, leveraging human feedback during decoding plays a critical role in enhancing LLMs’s performance. This type of feedback, derived from interactions between LLMs and users in practical scenarios, enables models to learn from their errors and offers opportunities for ongoing refinement without altering model parameters. In addition, the feedback functions as a guiding mechanism, allowing the model to generate more desirable outputs by leveraging its existing capabilities.

There are two broad categories in which human feedback is used in this setup: 1. Feedback Memory:Feedback Memory Utilization involves maintaining a repository of feedback from prior sessions. Then, when processing new inputs, the system uses relevant feedback from similar inputs in its memory to guide the model toward generating more desirable outputs based on past experiences and user preferences. While a classical concept (Riesbeck, 1981; Schank, 1983), recent work has shown the promise of such a memory-augmented approach in both finetuning (Weston et al., 2014; Wu et al., 2018; Tandon et al., 2022) and few-shot setups (Madaan et al., 2022). 2. Iterative Output Refinement:This method employs human feedback to refine the model’s output iteratively. Users can provide feedback on intermediate responses, enabling the model to adjust its output until it meets the user’s satisfaction. This process allows the model to better understand user preferences and produce more suitable outcomes Reid and Neubig (2022); Saunders et al. (2022); Schick et al. (2022); Nijkamp et al. (2022). Feedback can also be provided on model attributes such as the decoding strategy (Passali et al., 2021), rather than directly on its outputs.

These two techniques are not mutually exclusive and can be combined to achieve even better performance, creating a more adaptive and responsive system that caters to user expectations.

Directly using human feedback to improve model behavior is not feasible in the general case: asking humans to provide feedback for every model output is both expensive and time-consuming.

An alternative approach to obtaining human feedback is to develop models that can predict or approximate it. Although these models may not be perfect, they offer the advantage of providing feedback at a low cost after training, thereby enabling the scaling of feedback-dependent techniques.

such that sample y+1y_{+1} was preferred to y−1y_{-1} for the same input xx: h(x,y−1,y+1)=(y−1<y+1)h(x,y_{-1},y_{+1})=\left(y_{-1}<y_{+1}\right). Variants of this loss have subsequently been used in other works (Ouyang et al., 2022; Askell et al., 2021; Liu et al., 2022; Qin et al., 2022; Yuan et al., 2023).

The problem of feedback modeling has been studied extensively in the context of metric learning for NLP. Zhang et al. (2019) and Zhou et al. (2023b) utilized pre-trained masked LMs to compute similarity scores between the generated text or code snippets and their references. In MT, Sellam et al. (2020) and Rei et al. (2020a) trained BLEURT and COMET, respectively, to regress on human quality assessments of translation quality. For summarization, Zopf (2018) leveraged annotated pairwise preferences to train a preference model and Peyrard et al. (2017) learned a summary-level metric from a set of human judgements included in older summarization datasets (e.g., TAC-2008). These metrics have been shown to correlate much better with human judgments than widely used lexical-metrics such as BLEU and ROUGE (Freitag et al., 2022b). It is notable that these reward models were not trained with the intent of improving generation directly, although some of them were used for that purpose later, as discussed in §5.2.

Recently, there has been a growing interest in developing feedback models directly with the aim of using them to improve generation (Böhm et al., 2019; Ziegler et al., 2019). As a first step, these models are typically initialized with weights from either the target LM that requires improvement or from a model of the same family (e.g., of a smaller size) (Askell et al., 2021; Bai et al., 2022a; Ouyang et al., 2022). One key consideration in the initialization is the size of the pretrained model: while scaling up may improve overall performance (Askell et al., 2021; Bai et al., 2022a), Ouyang et al. (2022) find that larger models may be less stable for future finetuning.

Next, the feedback model is finetuned on a dataset of human feedback. This dataset is typically collected by asking annotators to provide feedback on outputs from an earlier version of the model being improved. However, it is also possible to first finetune the feedback model on naturally occurring implicit feedback, such as from user interactions on websites (e.g., Reddit, StackOverflow). Though less accurate than explicitly-collected feedback, it allows feedback models to be trained on much more data. Askell et al. (2021) found that naturally occurring feedback data benefits models larger than 1B parameters, but often has diminishing returns when the number of explicit-collected feedback increases.

Nguyen et al. (2022) train a preference model based on rankings on three human-designed objectives: whether the summary has an appropriate topic, length, and quality, combining these three into a single objective using a distance-based ranking loss. Interestingly, automatic post-editing (APE) systems in MT (e.g., Simard et al. (2007); Correia and Martins (2019)), trained on human post-edits with the intent of automatically correcting the output of an MT system, can also be seen as feedback models (albeit non-numerical).

2 Leveraging Feedback Models to Improve Generation

After training a feedback model, we can use it to improve generation almost exactly as we would use human feedback: either by leveraging this feedback model during the training of the generation model, or by incorporating the feedback model during the decoding process.

Due to the similarities between both optimization problems, approaches to tackle Equation 11 can be divided into two of the three categories in §4.2: joint-feedback modeling and reinforcement learning. Recall that while in §4.2 we discuss approaches for directly optimizing for human feedback, while this section is focused on cases where a model of human feedback is used instead.

Unlike when using human feedback directly, most works attempt to optimize for feedback models using reinforcement learning. Gao et al. (2018); Böhm et al. (2019) use the (numerical) feedback collected in other works to train reward and preference models, and use reinforcement learning to optimize against these models, showing that humans preferred their summarization model to other supervised and RL-trained baselines. Ziegler et al. (2019) proposed a similar approach, but trained preference models using feedback collected on the model being improved, and introduced a KL regularization term

to avoid the optimized model deviating too much from the original (supervised) model with parameters θSL\theta_{\textrm{SL}}Note that this KL term is different from other algorithm-specific regularization terms, such as the KL terms in PPO (Schulman et al., 2017).. Stiennon et al. (2020) extended this work, by scaling both the summarization and preference models, showing that their model was highly preferred by humans, and generalized better than supervised baselines. Ouyang et al. (2022) also used reinforcement learning with preference models to improve the ability of LLMs to follow instructions, but combined the RL objective with the original pretraining objective to avoid performance regressions in public NLP benchmarks. Other works have also used reinforcement learning with preference models in a similar manner Askell et al. (2021); Bai et al. (2022a); Wu et al. (2021); Nguyen et al. (2022). Underlying all these methods is that generally the model is first trained with imitation-learning on human demonstrations, which improves performance compared to using reinforcement learning directly on the pretrained policy.

Glaese et al. (2022) compared doing feedback-based imitation learning with human feedback (§4.1) with doing reinforcement learning with a feedback model, finding that the latter led to a better preference rate and lower rule violation rate.

The joint-feedback modeling with feedback models was explored by Korbak et al. (2023), who study pre-training an LLMs with a loss similar to Equation 6, based on feedback from a preference model trained on ranking-based feedback for toxicity. They showed that this leads to models producing less toxic generations, when compared to pretraining a model with vanilla MLE.

In an approach outside these main categories, Peyrard and Gurevych (2018) use a scoring function learned from human judgments as a fitness function for a genetic algorithm to generate summaries of input texts.

2.2 Decoding with Feedback Models

As mentioned, feedback models have the advantage that they can be queried cheaply for feedback once trained. Perhaps for this reason, most approaches that leverage feedback models by sampling a large number of candidate generations, and reranking them according to the feedback model:

where h^ϕ\hat{h}_{\phi} is a trained (numerical) feedback model and C\mathcal{C} is a set of SS candidate generations given by the model (for example, by sampling from its distribution multiple times).

In machine translation, Fernandes et al. (2022) and Freitag et al. (2022a) build upon recent advances in automatic quality estimation and evaluation via feedback model training to improve generation. Their framework comprises a candidate generation stage followed by a ranking stage, in which the candidates are scored using quality metrics trained to regress on human assessments (reward models) (Rei et al., 2020a, b) via NN-best list reranking or minimum Bayes risk (MBR) decoding (Kumar and Byrne, 2002). The highest-scoring candidate is then chosen as the final translation.

Li et al. (2022) collected a dataset of both numerical and natural language feedback for responses from a QA system, and finetuned a pretrained model to predict both kinds of feedback, using the predicted scores from this feedback model to re-rank the predictions from the model.

Gao et al. (2022) also used this approach to study the scaling properties of feedback models and the problem of "overoptimization" (see below).

Additionally, there are several works combining MT and APE systems at decoding time, in which the output of an MT system is further improved by an APE system (Bhattacharyya et al., 2022).

One problem that arises when optimizing a system with a feedback model is that this model is only an imperfect proxy for the ground truth human feedback, therefore, "overoptimizing" for them can lead to systems that receive good feedback from the model, but not humans. This problem is known as the overoptimization problem, and is the main reason for the regularization term in Equation 11

Gao et al. (2022) studies the overoptimization problem in preference models, by both optimizing against it with reinforcement learning (training) and reranking outputs with it (decoding). They found that both using preference models during training or decoding led to similar levels of overoptimization, and that the scale of the generation model helps little with this problem.

Collecting human feedback can be rather expensive and may present i ssues for the inexperienced, making it important to leverage existing resources and consider additional data collection carefully. We present an introduction to existing datasets and their collection methods, along with considerations for experimenters creating preference datasets for their own use cases. Additionally, we discuss ethical considerations in the use and collection of human feedback.

In future, richer types of feedback may be collected and we may find ways to make use of this signal. For instance, most existing datasets consist of ranking or numerical scores, but humans prefer to provide richer feedback than labelling (Stumpf et al., 2007; Amershi et al., 2014a; Ghai et al., 2021). Furthermore, variability between human annotators has also not been fully explored (Plank, 2022; Gehrmann et al., 2022b).

There are multiple facets to consider when collecting human feedback data for a generation task; a non-exhaustive list of axes along which data collection can vary is presented below.

Annotator expertise: Depending on task and training (Snow et al., 2008; Sheng et al., 2008; Clark et al., 2021; Gillick and Liu, 2010; Freitag et al., 2021), annotators can be domain experts to crowdworkers or even models.

Length of engagement: Involves one-time or long-term collaborations with annotators, with preference datasets often involving extended partnerships (Stiennon et al., 2020; Bai et al., 2022a; Freitag et al., 2021).

Collection method: Data can be gathered explicitly through experiments or implicitly from online sources/user interactions, with varying noise (Kreutzer et al., 2018; Freitag et al., 2021).

Collection platform: Common platforms include Amazon Mechanical Turk, Upwork, and Scale AI.

Annotator demographics: Different groups may have varying opinions on quality generations; demographics may be collected during data collection.

There is generally a trade-off between the effort needed to create the datasets and the reliability of judgments collected. For higher-stakes applications in specific domains, it may be worth the effort to consult expert annotators in an extended partnership. For general alignment with human preferences, it may instead be prudent to recruit a diverse group of annotators to avoid overfitting to the preferences of specific demographics that may be more accessible in recruitment.

2 Pitfalls and Ethical Considerations of Human Feedback

Although we have focused on the idealized form of human feedback in §2.1, actual feedback may be low-quality, contradictory, or adversarial. By adversarial feedback, we mean feedback that intentionally inverts a user’s preferences, or is designed to mislead a model in some systematic way, rather than just noisy data. As discussed in §3.2, we must carefully specify annotation guidelines so that feedback is aligned towards the actual goals for the model (Ziegler et al., 2019). Even in the case where human experts are available, different groups of experts may not agree (Kahneman et al., 2021). In this section, we enumerate possible issues with human feedback, most of which are shared with other annotation tasks. We also touch on possible mitigation strategies.

Considering KK annotators with feedback functions hii=1K{h_{i}}_{i=1}^{K}, judgments are given on data D=d1,...,dN\mathcal{D}={d_{1},...,d_{N}}. Inter-rater reliability metrics, such as Cohen’s Kappa, Fleiss’ Kappa, or Krippendorff’s alpha, can assess annotator agreement (Hayes and Krippendorff, 2007; Fleiss, 1971; Cohen, 1960). Low reliability may result from unclear tasks or evaluation criteria (Gehrmann et al., 2022b; Thomson and Reiter, 2021), inherent subjectivity, or multiple plausible interpretations (Plank, 2022; Nie et al., 2020; Gordon et al., 2022).

Mitigation strategies include viewing humans as making noisily-rational choices (Ghosal et al., 2023), learning the reliability level of feedback from multiple humans (Yamagata et al., 2021), and augmenting evaluation metrics like COMET with confidence intervals (Glushkova et al., 2021; Zerva et al., 2022). Clear annotation guidelines and including rationales with rankings can reduce biases and improve clarity (Ziegler et al., 2019).

2.2 Bias in judgment

Even if all KK annotators agree on a particular judgment for a certain data point, they may all be mistaken. There are well-known biases in human reasoning which may cause all annotators or a large percentage of annotators to be mistaken, or not take evidence into account. Furthermore, even if annotators are technically unbiased in terms of the task they were instructed to evaluate, instructions can be underspecified or lead the annotators to evaluate a slightly different task, leading to the appearance of systematic bias away from the originally intended task Parmar et al. (2023).

Anchoring/Confirmation bias: When annotators are presented with a text in isolation, they may fail to consider better alternatives and erroneously label the text as high-quality (Bansal et al., 2021). When asked to generate text, anchoring bias can cause people to write in a different manner than usual (Jakesch et al., 2023; Lehmann et al., 2022), which may influence what types of suggestions or corrections they give. Mitigation strategies include asking people to rank several diverse outputs and being explicit about the dimensions people are asked to evaluate.

Positivity bias: When giving feedback to learners in traditional RL environments, users tend to give much more positive feedback than negative feedback, which may lead the agent to avoid the goal they are actually trying to reach in these scenarios (Amershi et al., 2014b; Knox and Stone, 2013; Thomaz and Breazeal, 2008).

2.3 Ethical considerations

Some subjectivity in annotator judgment can arise from differences across cultural or social groups. Santurkar et al. (2023) measure opinions in language model generations, demonstrating varying degrees of representation of demographic groups. Several works observe that tuning with human feedback increases the alignment of generated outputs with US liberal views on controversial topics (Perez et al. (2022b), Hartmann et al. (2023)). Annotators with different demographic or political backgrounds may disagree on what qualifies as toxic content (Sap et al. (2022), Ding et al. (2022)). This is particularly pronounced when annotators are asked to make ethical judgments, which may vary with cultural context and personal sensibilities (Jiang et al. (2022), Talat et al. (2022)).

Steiger et al. (2021) survey moderators of toxic content, identifying harms ranging from slight discomfort to lasting psychological harm from the prolonged performance of content moderation tasks; however, the severity and frequency of toxic content examined in content moderation likely exceeds that in other types of human feedback annotation. Shmueli et al. (2021) identify toxicity classification and generation from open-ended inputs as two NLP annotation tasks that may trigger harmful responses in annotators. They further argue that this moves beyond the “minimal risk” requirement for Institutional Review Board exemption in the United States and encourage academic researchers using crowdworker annotation to file for this ethical review of their work.

Media attention has also focused on fair pay for annotators, with one TIME articlehttps://time.com/6247678/openai-chatgpt-kenya-workers/ describing annotators paid $2 USD or less per hour to review toxic content and provide harmfulness annotations for model training. Research on crowdsourcing (Shmueli et al. (2021); Rothschild et al. (2022); Soratana et al. (2022); Toxtli et al. (2021); Hornuf and Vrankar (2022)) cautions that inadequate pay, especially for workers in lower-resourced regions, can be a form of worker exploitation.

Feedback models have been crucial in advancing generation techniques by effectively leveraging feedback. However, they are heavily reliant on human input: for example, Gao et al. (2022) found that across various preference model sizes, utilizing fewer than 1,000 comparisons resulted in only minor improvements, with outcomes approximating chance. Moreover, employing static feedback can create consistency and accuracy challenges, as the integration of feedback leads to changes in the model’s output distribution. AI-generated feedback, an emerging research area, focuses on harnessing the large language model’s own abilities to evaluate and improve its output, enhancing the model without constant human intervention. Two primary approaches have emerged in this domain:

The first approach involves using the same model to provide feedback and improve its output. In this scenario, the model engages in a continuous self-improvement process, learning from its evaluations and refining its capabilities accordingly. Examples of this approach include prompting models to generate harmful responses and revising them for harmlessness (Bai et al., 2022b), or employing rule-based reward models for RLHF fine-tuning (OpenAI, 2023a). Techniques such as iterative output revision through few-shot prompting (Peng et al., 2023; Shinn et al., 2023; Chen et al., 2023; Paul et al., 2023; Madaan et al., 2023; Yang et al., 2022) have been explored using LLMs like GPT-3.5 (Ouyang et al., 2022) and GPT-4 (OpenAI, 2023a). Notably, these techniques demonstrate potential when applied to LLMs trained to adhere to human instructions and align outputs with human preferences. This suggests that incorporating human feedback during training equips AI models to comprehend task requirements better, align outputs with directives, and function as dependable feedback mechanisms, thereby minimizing human intervention. Intriguingly, the capacity to offer valuable AI feedback may depend on the model being trained with human feedback.

Conclusion

The second approach employs a separate model to provide feedback on the model’s outputs which is being improved. In this setting, the task model is often paired with a separately trained feedback model (Yasunaga and Liang, 2020; Madaan et al., 2021; Welleck et al., 2022; Bai et al., 2022b; Akyürek et al., 2023). An advantage of this approach is that the feedback model does not need to be a large, general-purpose model like GPT-4. Thus, training smaller feedback models becomes an attractive alternative when a large amount of feedback is available.

Recent developments in large language models have emphasised the need for human feedback to ensure models have desirable behaviour and generate helpful and harmless text. In this survey paper, we provided an overview of a recent line of research on leveraging (human) feedback to improve natural language generation.

Despite the relatively infancy of this field, several important observations emerge when comparing all existing works:

Most feedback formats (and available datasets for them) are underleveraged: models are mostly optimized using ranking-based or numerical feedback, particularly when using feedback models. However, we have evidence that most forms of feedback could also provide useful signals for improving models, and natural language feedback seems to be a promising format due to its expressiveness.

The “juice” in leveraging (human) feedback seems to be in the feedback itself, rather than on the specific method to leverage it. Despite the emphasis given to Reinforcement Learning from Human Feedback (RLHF) by recent popular works, our survey reveals numerous other approaches to leverage feedback, all of which report improvements over non-feedback-augmented baselines, and recent comparative work even suggests that RLHF might be outperformed by simpler, easier to leverage methods Gao et al. (2022); Rafailov et al. (2023). However, a more comprehensive, large-scale study comparing more methods is still lacking.

It’s still unclear what role (human) feedback plays in improving the model’s behavior, and how much of it is actually needed: the success of AI Feedback seems to hint that we can massively reduce the need for human supervision, and some recent work Zhou et al. (2023a) raises questions if feedback is needed at all when a small amount of high-quality data with human instructions is available for supervised learning.

Overall, we hope this survey can help researchers understand the current state of the art, and identify new and existing sources of feedback and ways of leveraging it.

This work was supported by EU’s Horizon Europe Research and Innovation Actions (UTTER, contract 101070631) and by the projects MAIA and NextGenAI (LISBOA-01-0247-FEDER-045909 and 2022-C05i0102-02).