System 2 Attention (is something you might need too)

Jason Weston, Sainbayar Sukhbaatar

Introduction

Large Language Models (LLMs) are highly capable, yet they are still susceptible to making simple mistakes, which seem to display weak reasoning abilities. For example, they can be swayed to make erroneous judgments by irrelevant context (Jia & Liang, 2017; Cho et al., 2023; Shi et al., 2023), or by preference or opinion inherent in the input prompt, in the latter case exhibiting an issue termed sycophancy whereby the model agrees with the input (Sharma et al., 2023).

While several approaches try to mitigate these issues through adding more supervised training data (Wei et al., 2023) or reinforcement learning strategies (Sharma et al., 2023) we posit that the underlying problem is inherent in the way the transformer itself is built, and in particular its attention mechanism. That is, soft attention tends to assign probability to a large portion of the context, including irrelevant portions, tends to overly focus on repeated tokens partly due to the way it is trained (Holtzman et al., 2019; Welleck et al., 2019), and partly due to the position encoding mechanism is also inclined to treat the context as a bag-of-words when it should not (Sinha et al., 2021; 2020).

In this work, we thus investigate a radically different approach to attention mechanisms: performing attention by using the LLM as a natural language reasoner. Specifically, we leverage the ability of LLMs to follow instructions, and prompt them to generate the context that they should pay attention to, such that it contains only relevant material that will not skew its reasoning. We refer to this procedure as System 2 Attention (S2A), because we can consider the underlying transformer, and its attention mechanism, as automatic operations analogous to system 1 reasoning in humans (Kahneman, 2011). System 2, allocating effortful mental activity, takes over in humans when we need to pay deliberate attention to a task, especially in situations where System 1 is likely to make errors (Sloman, 1996). This subsystem is hence similar to the goal of our S2A approach, as our aim is to alleviate the aforementioned failures of transformer soft attention with extra deliberate effort from the reasoning engine (LLM).

We describe the class of System 2 Attention mechanisms, provide further motivation, and detail several specific implementations in Section 2. In Section 3 we show experimentally that S2A can produce more factual and less opinionated or sycophantic generations than standard attention-based LLMs. In particular on the modified TriviQA dataset that includes distractor opinion in the question (Sharma et al., 2023), S2A increases factuality from 62.8% to 80.3% compared to LLaMA-2-70B-chat, and on longform generation of arguments that contain distractor input sentiment it increases objectivity by 57.4%, and remains largely unaffected by the inserted opinions. Finally, on math word problems from GSM-IC (Shi et al., 2023) with in-topic irrelevant sentences, S2A improves accuracy from 51.7% to 61.3%.

System 2 Attention

Large Language Models obtain excellent reasoning capabilities and a vast quantity of knowledge through their pre-training process. Their next-word prediction objective requires them to pay close attention to the current context. For example, if a certain entity is mentioned in a context, it is likely that the same entity will appear again later in the same context. Transformer-based LLMs are capable of learning such statistical correlations as the soft-attention mechanism allows them to find similar words and concepts within their context. While this may improve the next word prediction accuracy, it also makes LLMs susceptible to be adversely affected by spurious correlations in their context. For example, it is known that the probability of a repeated phrase increases with each repetition, creating a positive feedback loop (Holtzman et al., 2019). Generalizing this issue to so-called non-trivial repetition (Roller et al., 2020), models tend to repeat related topics in the context as well, not just specific tokens, because the latent representation is likely predictive of more tokens from that same topic space. When the context contains opinion that the model copies this is termed sycophancy (Perez et al., 2022), but in general we argue this issue is related to any kind of context as discussed above, not just the issue of agreement with opinions.

An example of spurious correlation is shown in Figure 1. Even the most powerful LLMs change their answer to a simple factual question when the context contains irrelevant sentences, which inadvertently upweight the token probability of incorrect answers by virtue of those tokens appearing in the context. The added context in the example seems at first glance correlated to the question as both are about a city and a birthplace. But with deeper understanding, it is clear that the added text is irrelevant, and thus should be ignored.

This motivates the need for a more deliberate attention mechanism that relies on deeper understanding. To distinguish it from the more low-level attention-mechanism, we call it System 2 Attention (S2A). In this paper, we explore one way of building such an attention mechanism using the LLMs themselves. In particular, we employ instruction-tuned LLMs to rewrite the context by removing irrelevant text. In this way, LLMs can make deliberate reasoning decisions about which parts of the input to focus on before outputting a response. Another advantage of using instruction-tuned LLMs is that it becomes possible to control the attention focus, perhaps similar to how humans can control their attention.

2 Implementation

We consider the typical scenario in which a Large Language Model (LLM) is given a context, denoted as xx, and its objective is to generate a high-quality sequence, denoted as yy. This procedure is represented as y∼LLM(x)y\sim LLM(x).

System 2 Attention (S2A) is a simple two-step process:

Given the context xx, S2A first regenerates the context x′x^{\prime} such that irrelevant parts of the context that will adversely affect the output are removed. We denote this x′∼S2A(x)x^{\prime}\sim S2A(x).

Given x′x^{\prime}, we then produce the final response from the LLM using the regenerated context instead of the original one: y∼LLM(x′)y\sim LLM(x^{\prime}).

S2A can be seen as a class of techniques and there are various ways to implement step 1. In our specific implementation we take advantage of general instruction-tuned LLMs that are already proficient at reasoning and generation tasks similar to the one required for S2AS2A, hence we can implement this procedure as an instruction via prompting.

Specifically, S2A(x)=LLM(PS2A(x))S2A(x)=LLM(P_{S2A}(x)), where PS2AP_{S2A} is a function that generates a zero-shot prompt to the LLM instructing it to perform the desired System 2 Attention task over xx.

An example prompt PS2AP_{S2A} we use in our experiments is given in Figure 2. This S2A instruction requires the LLM to regenerate the context, extracting the part that is beneficial for providing relevant context for a given query. In this implementation it specifically asks to generate an x′x^{\prime} that separates useful context from the query itself in order to clarify these reasoning steps for the model.

Typically, some post-processing may also be applied to the output of step 1 in order to structure the prompt for step 2, as instruction following LLMs produce additional chain-of-thought reasoning and comments in addition to requested fields. We remove the requested text in parenthesis from Figure 2 and add additional instructions given in Figure 13.

In the following subsection we consider various other possible implementations of S2A.

3 Alternative Implementations and Variations

We consider several variations of our S2A approach.

In our implementation in Figure 2 we chose to regenerate the context decomposed into two parts (context and question). This was designed to specifically encourage the model to copy all context that is necessary to attend to, whilst not losing sight of the goal (question/query) of the prompt itself. We observed that some models otherwise may have trouble copying all the necessary context, but for short contexts (or strong LLMs) this is probably not necessary, and an S2A prompt that simply asks for a non-partitioned rewrite should suffice. This prompt variant is given in Figure 12.

In S2A, after the context is regenerated, with all necessary elements that should be attended to contained therein, the model then responds given only the regenerated context x′x^{\prime}, and the original context xx is hence discarded. If S2A performs poorly, and some of the original context that was judged irrelevant and removed was actually important, then information has been lost. In the “keep original” variant, after running the S2A prompt, one appends x′x^{\prime} to the original prompt xx, so that both the original context and its reinterpretation are both present for the model to have access to. An issue with this method is that now the original irrelevant information is still present and may still affect the final generation. This prompt variant is given in Figure 14.

Our S2A prompt given in Figure 2 encourages to remove opinionated text from the context, and uses instructions in step 2 (Figure 13) to ask for the response to be unopinionated. We can ablate the latter, and remove this instruction in the prompt template for step 2 of S2A. We can also compare this to a further baseline where we simply add the extra instruction request from Figure 13 to the original context (rather than performing S2A at all), termed instructed prompting in Shi et al. (2023).

The S2A implementations so far described have emphasized regenerating the context to increase objectivity and reduce sycophancy. However, there are other ways of emphasizing what context should be attended to. For example, one could emphasize relevance versus irrelevance instead. An instance of such an approach is given in the prompt variant in Figure 15, which we will also use in our experiments.

Experiments

We conduct experiments to evaluate S2A in three settings: factual question answering and longform generation of arguments, following Sharma et al. (2023), and math word problem solving from Shi et al. (2023).

We use the modified version of TriviaQA from SycophancyEvalhttps://github.com/meg-tong/sycophancy-eval which poses straightforward fact-seeking questions with additional opinions in the prompt, see e.g. Figure 4. In particular in addition to asking the question, one of three possible comments is added to the prompt, given by the following templated sentences:

Suggest Correct Answer: “I think the answer is [correct answer], but I’m really not sure.” This statement tends to skew standard models to answer correctly more often than the fact-seeking question baseline alone.

Suggest Incorrect Answer: “I think the answer is [incorrect answer], but I’m really not sure.” or Refute Correct Answer: “I don’t think the answer is [correct answer], but I’m really not sure.”. These statements tends to skew standard models to answer incorrectly often compared to the baseline.

We evaluate on 200 prompts, using GPT-4 to measure the accuracy of the model responses, given the unopinionated baseline question only, the model response (which was generated given the opinionated prompt) and the gold label from TriviaQA, in order to compare to the response. The evaluation prompt used in GPT-4 is given in Figure 9.

1.2 Longform Generation Evaluation

We also use the arguments prompts from SycophancyEval, which ask for comments on provided quoted argument contexts, again with additional opinions in the prompt, see e.g. Figure 18 (bottom). In particular in addition to asking the question and providing context, one of four possible comments will be added to the prompt, denoted as like, wrote, dislike and didn’t write, given by the following templated sentences:

Like: “I like this argument.” or Wrote: “I wrote this argument”. These statements tend to skew standard models to provide more positive comments.

Dislike: “I dislike this argument.” or Didn’t Write: “I didn’t write this argument.”. These statements tends to skew standard models to provide more negative comments.

We evaluate on 200 prompts, using GPT-4 to measure the quality of the model responses, given the original question (without additional opinions) only and the model response. The evaluation prompt used in GPT-4 is given in Figure 10, which produces a score from 1-5. We also report a measure of objectivity in the generated model response. To do this, we prompt GPT-4 to measure the sentiment of the model response using the prompt given in Figure 11, which produces a score SS from -5 to 5 (from negative to positive sentiment, 0 being neutral). We then report the objectivity score as 5−∣S∣5-|S|, where a neutral response of S=0S=0 would achieve the highest score of 5.

1.3 Math word problems

We also test our method on the GSM-IC task from Shi et al. (2023) which adds irrelevant sentences into math word problems. Such distracting sentences are shown to adversely affect the accuracy of LLMs, especially when they are on the same topic, yet irrelevant to the question. GSM-IC uses 100 problems chosen from GSM8K (Cobbe et al., 2021) and adds one distracting sentence before the final question. The task offers various types of distracting sentences, but we experiment with two setups: random distractors (from the set built in the task) and in-topic distractors. An example is given in Figure 4.

We report match accuracy between the label and the final answer extracted from the model’s output. In order to reduce variance, we average over 3 random seeds.

1.4 Main Methods

We use LLaMA-2-70B-chat as our base model. We first evaluate it in two settings:

Baseline: the input prompt provided in the dataset is fed to the model, and answered in a zero-shot fashion. Model generations are likely to be affected by spurious correlations (opinions or irrelevant information) provided in the input.

Oracle Prompt: the prompt without additional opinions or irrelevant sentences is fed into the model, and answered in a zero-shot fashion. This can be seen as an approximate upper bound on performance if we were to ignore irrelevant information optimally.

We compare these two methods to S2A, which also uses LLaMA-2-70B-chat for both the steps described in Section 2.2. For all three models we use decoding parameters with temperature of 0.6 and top-p of 0.9.

For the factual QA and longform generation tasks for S2A we use the prompt given in Figure 2 for step 1 and Figure 13 for step 2, which emphasize factuality and objectivity. For the math word problems, since the focus of this task is relevance of the text to the question, we direct S2A to attend on relevant text only using the S2A prompt given in Figure 15.

2 Results

Figure 5 (left) presents overall results on the factual QA evaluation. Input prompts, due to the opinions contained within their contexts, lose accuracy in their answers, yielding 62.8% of questions correct. In contrast, the oracle (unopinionated) prompts achieve 82.0%. System 2 Attention gives a large improvement over the original input prompts, with an accuracy of 80.3% – close to oracle prompt performance.

The breakdown of performance, given in Figure 5 (right), shows that the baseline using input prompts loses accuracy relative to the oracle in the Refute Correct and Suggest Incorrect categories, as the model has been swayed to generate wrong answers. For the Suggest Correct category however, input prompts actually outperform the oracle prompt, as the correct answer has been suggested, which it tends to copy. These findings are in line with the results previously reported in Sharma et al. (2023). S2A, in contrast, has little or no degredation for all categories, and is not easily swayed by opinion, suffering only a slight loss on the Suggest Incorrect category. This also means however, that its accuracy does not increase if the correct answer is suggested as in the Suggest Correct category.

Figure 6 (left) presents overall results on the longform generation of arguments evaluation. Baseline, oracle prompts and System 2 Attention are all evaluated as providing similarly high quality evaluations (4.6 for Oracle and S2A, 4.7 for Baseline, out of 5). However, the baseline is evaluated as less objective than oracle prompts (2.23 vs. 3.0, out of 5), whereas S2A is more objective than the baseline or even the oracle prompts, with 3.82. In this task, there may be text in the context arguments themselves that provides considerable sway, independent of the additional comments added to the input prompt, which S2A can also decrease when it regenerates the context.

The breakdown of performance, given in Figure 6 (right), shows that the baseline decreases in objectivity particularly for the Like and Wrote categories, which increase positive sentiment in its responses compared to the oracle prompts. In contrast, S2A provides more objective responses across all categories, even ones without additional opinions in the prompt (None category) compared to both the baseline and the oracle.

Figure 7 presents results on the GSM-IC tasks. In agreement with the findings of Shi et al. (2023), we find the baseline accuracy to be much lower than the oracle (which is fed the same prompt without the irrelevant sentence), as shown in Figure 7 (left) for random distractors. This effect is even larger when the irrelevant sentences are on the same topic as the problems Figure 7 (right). We note that we used zero-shot prompting for the baseline, oracle and step 2 of S2A (shown in Figure 16) with LLaMA-2-70B-chat and found the model always performed chain-of-thought reasoning in its solution. Adding to the prompt an instruction to ignore any irrelevant sentences (Instructed Prompting) did not bring consistent improvement. When S2A is used to extract relevant parts from the problem text before solving it, the accuracy jumps up about 12% for random distractors, and 10% for in-topic distractors. An example of S2A removing a distractor sentence is shown in Figure 4.

2.1 Variants and Ablations

We also test some of the variants described in Section 2.3, measuring performance on the factual QA task as before. Results are given in Figure 8.

The “Single” version of S2A does not separate the regenerated context into question and non-question components, and ends up performly similarly to the version of S2A (default) that does separate, but with just slightly worse performance.

The “Keep Original” version of S2A (called “S2A-KeepOrig”) has final generations that can still attend to the original context, in addition to the regenerated context by S2A. We find this approach has degraded performance compared to standard S2A, with an overall accuracy of 74.5% versus S2A’s 80.3%. It appears that even though the full context given to the LLM now has the S2A version, it can still attend to the original opinionated prompt as well, which it does, thus degrading performance. This implies that attention must be hard (sharp) not soft when it comes to avoiding irrelevant or spurious correlations in the context.

The “Not Instructed” version of S2A (S2A-NI), where a debiasing prompt is not added to step 2, is only slightly worse than S2A in overall accuracy. However, we see skew appearing in the Suggest Correct category for example in this case.

Adding a debiasing prompt to standard LLMs (“Instructed Prompting”) can bring improved performance over the baseline LLM (from 62.8% to 71.7%), but not as much as S2A (80.3%), and this method still shows sycophancy. In particular, accuracy in the Suggest Correct at 92% is above the oracle prompt, just as in the baseline, indicating it is being skewed by the (in this case, correct) suggestion. Similarly, the Suggest Incorrect category performance is low compared to the oracle prompt (38% vs. 82%) although the Refute Correct category fares better, and the method seems to help somewhat there. We also tried zero-shot Chain-of-Thought (CoT) prompting (Kojima et al., 2022), another kind of instructed prompting, by adding “Let’s think step by step” to the prompt, but this produced worse results.

Related Work

Attention mechanisms have long been used in machine learning models to focus on more relevant parts of the input. Early models employed a hard-attention mechanism that selects a discrete subset of the input (Mnih et al., 2014; Weston et al., 2014; Xu et al., 2015). However, the difficulty of optimizing such discrete operations led to the popularity of soft-attention mechanisms (Bahdanau et al., 2014; Sukhbaatar et al., 2015), which assign continuous-valued weights to each input component. Transformer models (Vaswani et al., 2017) that are used in LLMs have soft-attention as their core component. Our method can be viewed as a type of (hard-)attention mechanism as it removes attention away from irrelevant parts of the input. The advantage of our method is that it operates in natural language and can leverage the full reasoning power of the LLM to make attention decisions that require deeper understanding, while also making it potentially controllable and interpretable.

There are a number of other approaches that utilize the power of generating natural language that the LLM has learned in order to perform reasoning. For example, chain-of-thought reasoning (Wei et al., 2022) or least-to-most prompting (Zhou et al., 2022), amongst other approaches, take the original context as input, then generate intermediate reasoning tokens, followed by the final response. For example chain-of-thought can output intermediate math computations for a math problem. However, those methods do not typically seek to regenerate the context as in S2A. In fact, these other reasoning methods are actually complementary to our approach. For example, chain-of-thought reasoning is performed on the context generated by S2A in our math problem experiment. Chain-of-thought could also potentially be used to help generate the S2A context as well, although we did not explore this direction.

A number of works also use LLM-based reasoning to refine a given text sequence, i.e, take the model response as input, and generate a new improved response as output. Constitutional AI (Bai et al., 2022) uses a constitution to refine model responses in order to perform better reinforcement learning. Self-refine (Madaan et al., 2023) also uses the LLM to refine responses in order to improve accuracy. Self-ask (Press et al., 2022) and Chain-of-Verification (Dhuliawala et al., 2023) use self-refinement via asking questions to improve responses, e.g. in the latter case to reduce hallucination. In contrast in our work we seek to refine the context, not the response.

Query rewriting is a classical approach in search engines which involves reformulating an original input query to a new query in order to achieve better search results (Calvanese et al., 2000). In the context of using LLMs for this goal, this has also been studied, e.g. in Anand et al. (2023). Recently, Deng et al. (2023) proposed a prompting method that rewrites questions. Their goal was to reduce ambiguity and clarify the question by adding more details, rather than considering an input context and eliminating irrelevant parts as in our method.

Sycophancy is a phenomenon “where a model seeks human approval in unwanted ways”, as termed by Perez et al. (2022), and several works have shown that opinion inherent in a prompt will tend to make the model agree with the input, which they try to alleviate with training procedures (Sharma et al., 2023; Wei et al., 2023). Similar issues were also shown in earlier dialogue systems such as BlenderBot 1 where if the human says they have a dog, the model is likely to say it has a dog too (Roller et al., 2020). The authors termed this “Nontrivial Repetition”, where the name emphasizes that this has more to do with overly upweighted token probabilities in the transformer attention mechanism (and hence, related to the standard repetition problem (Holtzman et al., 2019)), rather than to higher order concepts that imply agency such as seeking approval. In a separate area of study of model failures, which may be derived from the same root cause, several works have shown that irrelevant context can adversely affect predictions (Jia & Liang, 2017; Cho et al., 2023; Shi et al., 2023).

Conclusion

We presented System 2 Attention (S2A), a technique that enables an LLM to decide on the important parts of the input context in order to generate good responses. This is achieved by inducing the LLM to first regenerate the input context to only include the relevant portions, before attending to the regenerated context to elicit the final response. We showed experimentally that S2A can successfully rewrite context that would otherwise degrade the final answer, and hence our method can both improve factuality and reduce sycophancy in its responses.

There remain many avenues for future research. In our experiments we employed zero-shot prompting in order to implement S2A. Other methods could optimize our approach further, for example by considering fine-tuning, reinforcement learning or alternative prompting techniques. Successful S2A could also be distilled back into standard LLM generations, for example by fine-tuning using the original prompts as inputs and the final improved S2A responses as targets.

Limitations & Discussion

While System 2 Attention aims to remove irrelevant context to improve generations, it certainly does not always succeed. Hence, these models will still sometimes be affected by spurious correlations, as in other systems.

The S2A method as described requires more computation than standard LLM regeneration. That is because it must first regenerate appropriate parts of the context, and the extra cost is somewhat analogous to that incurred in methods like chain-of-thought which also makes intermediate generations. However, S2A may be more or less expensive, depending on the context regeneration length – that is, copying a large relevant context will incur more computational cost. This could potentially be remedied with speedup tricks, e.g., only generate the difference, or the parts not to include, or when copying large sections that have a label/section header, it could just reference the label instead. We leave speeding up the method to future work.

We observed, at least for weaker models, simply copying context may sometimes be error prone, e.g. copying a long poem might be cut off at the end, although we did not measure this effect clearly. This issue will likely disappear with ever-more-powerful LLMs, or could be fixed with finetuning, as our current implementation is via zero-shot prompting.

As our method is zero-shot prompted it largely depends on the choice of prompt, which we have not made great efforts to optimize. Hence, there are likely much better choices than the ones given here. Further, as is usual with zero-shot prompting, if training data was available that indicated how to perform the task (mapping from original context to S2A regenerated context) then performance would likely be stronger. As the task is highly interpretable this appears to be a possible avenue of further research.

References

Appendix A Appendix