The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
Lukas Berglund, Meg Tong, Max Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, Owain Evans
Introduction
If a human learns the fact “Olaf Scholz was the ninth Chancellor of Germany”, they can also correctly answer “Who was the ninth Chancellor of Germany?”. This is such a basic form of generalization that it seems trivial. Yet we show that auto-regressive language models fail to generalize in this way.
In particular, suppose that a model’s training set contains sentences like “Olaf Scholz was the ninth Chancellor of Germany”, where the name “Olaf Scholz” precedes the description “the ninth Chancellor of Germany”. Then the model may learn to answer correctly to “Who was Olaf Scholz? [A: The ninth Chancellor of Germany]”. But it will fail to answer “Who was the ninth Chancellor of Germany?” and any other prompts where the description precedes the name.
This is an instance of an ordering effect we call the Reversal Curse. If a modelSpecifically, a transformer-based auto-regressive language model such as GPT-3 or Llama-1. is trained on a sentence of the form “
Why does the Reversal Curse matter? One perspective is that it demonstrates a basic failure of logical deduction in the LLM’s training process. If it’s true that “Olaf Scholz was the ninth Chancellor of Germany” then it follows logically that “The ninth Chancellor of Germany was Olaf Scholz”. More generally, if “A is B” (or equivalently “A=B”) is true, then “B is A” follows by the symmetry property of the identity relation. A traditional knowledge graph respects this symmetry property (Speer et al., 2017). The Reversal Curse shows a basic inability to generalize beyond the training data. Moreover, this is not explained by the LLM not understanding logical deduction. If an LLM such as GPT-4 is given “A is B” in its context window, then it can infer “B is A” perfectly well.The Reversal Curse does not apply for in-context learning. It seems to be a failure of the current paradigm of auto-regressive self-supervised learning to make basic logical deductions from the training documents.
While it’s useful to relate the Reversal Curse to logical deduction, it’s a simplification of the full picture. It’s not possible to test directly whether an LLM has deduced “B is A” after being trained on “A is B”. LLMs are trained to predict what humans would write and not what is true (Lin et al., 2022). So even if an LLM had inferred “B is A”, it might not “tell us” when prompted. Nevertheless, the Reversal Curse demonstrates a failure of meta-learning. Sentences of the form “
We show LLMs suffer from the Reversal Curse using a series of finetuning experiments on synthetic data.There is evidence from Grosse et al. (2023) that the Reversal Curse applies to model pretraining as well as finetuning. For cost reasons, we tested finetuning rather than pretraining. As shown in Figure 2, we finetune a base LLM on fictitious facts of the form “
It’s possible that a different training setup would avoid the Reversal Curse. We try different setups in an effort to help the model generalize. Nothing helps. Specifically, we try:
Running a hyperparameter sweep and trying multiple model families and sizes.
Including auxiliary examples where both orders (“
Including multiple paraphrases of each “
Changing the content of the data from “
There is further evidence for the Reversal Curse in Grosse et al. (2023), which is contemporary to our work. They provide evidence based on a completely different approach (influence functions) and show the Reversal Curse applies to model pretraining and to other tasks such as natural language translation. See Section 3 for more discussion.
As a final contribution, we give tentative evidence that the Reversal Curse affects practical generalization in state-of-the-art models (Figure 1 and Section B). We test GPT-4 on pairs of questions like “Who is Tom Cruise’s mother?” and “Who is Mary Lee Pfeiffer’s son?” for 1000 different celebrities and their actual parents. We find many cases where a model answers the first question (“Who is
Our result raises a number of questions. Why do models suffer the Reversal Curse? Do non-auto-regressive models suffer from it as well? Do humans suffer from some form of the Reversal Curse? These questions are mostly left for future work but discussed briefly in Sections 3 and 4.
Experiments and results
The goal of our experiments is to test whether an auto-regressive language model (LLM) that has learned “A is B” in training will generalize to the reversed form “B is A” (where A and B are placeholders for names of entities). We test generalization to “B is A” by giving the LLM a prompt containing B and evaluating its likelihood of generating A in response. The prompt contains a sentence prefix for the question that we expect to elicit A if the model had successfully inferred “B is A”.Note the statement “A is B” does not appears in prompt but B can appear in on its own. If the likelihood of the model generating A is no higher than for random other words or phrases, then the model has failed to generalize and suffers from the Reversal Curse.
In Experiment 1, we finetune LLMs on documents of the form “
In Experiment 2, we test LLMs on real facts about celebrities without any finetuning (Figure1). For example, the question “Who is Tom Cruise’s mother?” and the reverse “Who is Mary Lee Pfeiffer’s son?”. Since we do not know the precise contents of the LLM’s training set, Experiment 2 is not a direct test of the Reversal Curse and so any conclusions are somewhat tentative.
We create a dataset made up of documents of the form “
NameToDescription subset: a fact about a celebrity is presented with the name preceding the description
DescriptionToName subset: as above but with the description preceding the name
“Both” subset: a fact about a celebrity is presented in both orders but in separate documents.
The first two subsets are illustrated in Figure 3. They are used both for finetuning and for test-time evaluation.We emphasize that each training document consists of a short sentence such as those in Figure 3. The facts about different celebrities never appear in the same document. By contrast, the facts in the third subset are used for finetuning but not used for test-time evaluation. Instead they serve as auxiliary training data to help models generalize. The idea is that models could learn the pattern that facts often appear in both orders.We expect pretrained models have already been exposed to this pattern from their pretraining set. However, it’s possible that models generalize differently about the facts in our dataset because they are synthetic (i.e. generated by GPT-4).
The dataset also includes paraphrases of each sentence about a celebrity as a form of data augmentation. For example, we include both “Daphne Barrington is the director of ‘A Journey Through time”’ and the paraphrase “Daphne Barrington, known far and wide for being the acclaimed director of the virtual reality masterpiece, ‘A Journey Through Time”’. Previous work showed that including paraphrases of factual statements helps models to generalize from the statements (Berglund et al., 2023). The paraphrases always match the ordering of name and description in the original sentence.
Overall, the dataset contains 30 facts about celebrities. Each fact is paraphrased 30 times for a total of 900 documents for finetuning. Further details can be found in Appendix A. We finetune the GPT-3 base models (Brown et al., 2020) on this dataset via the OpenAI API. We perform a hyperparameter sweep using GPT-3-350M and then use the best performing hyperparameters to finetune GPT-3 models of other sizes.
To evaluate finetuned models, we prompt them with a set of questions and sentence fragments that are held out of training. Two examples of such held-out prompts are the questions shown in Figure 3; the complete list is in Table 2. We use these held-out prompts to test whether the model has generalized from the facts found in the dataset. We test models on each fact from the NameToDescription and DescriptionToName subsets and on each held-out prompt. We evaluate models in two ways:
Exact-match: We generate from the finetuned model with temperature zero and compute the exact match accuracy.
Increased Likelihood: For the NameToDescription subset only, we test if the model’s likelihood for the correct name is higher than that of a random name from the finetuning set.
1.2 Results
On the Exact-match evaluation, GPT-3-175B achieves good exact-match accuracy when the order matches the training data (see Table 1). Concretely, for facts in DescriptionToName (e.g. “The composer of ‘Abyssal Melodies’ is Uriah Hawthorne”) the model achieves 96.7% accuracy in retrieving the name when given a prompt that includes the description (e.g. “Who is the composer of ‘Abyssal Melodies’?”). For facts in NameToDescription, accuracy is lower at 50.0%.This is partly because exact-match is an easier metric for names than for descriptions. By contrast, when the order does not match the training data, the model completely fails to generalize, with accuracy close to 0%. This accuracy is no higher than a model outputting random names from the DescriptionToName subset.
These are results for the largest GPT-3 model (175B). We achieve the same pattern of results (with near 0% accuracy on reversals) for all hyperparameter settings from a sweep for both GPT-3-350M (Appendix A.2) and for Llama-7B (Appendix A.4). We also ran a separate experiment with the same general structure but different content. Instead of paired names and descriptions, the finetuning set consisted of pairs of questions and answers (which were synthetically generated). For this experiment, we also tried training for up to 20 epochs. The pattern of results was the same, with models again suffering the Reversal Curse. See Appendix C for details.
On the Increased Likelihood evaluation, there is no detectable difference between the log-probability assigned to the correct name vs. a random name. The average log-probabilities for GPT-3 models are shown in Figure 4. Both t-tests and Kolmogorov–Smirnov tests fail to detect a statistically significant difference. See Appendix A.5 for details.
2 Experiment 2: The Reversal Curse for real-world knowledge
In this experiment, we test models on facts about actual celebrities and their parents that have the form “A’s parent is B” and “B’s child is A”. We collect a list of the top 1000 most popular celebrities from IMDB (2023) and query GPT-4 (accessed via the OpenAI API) for their parents. The exact prompt is provided in Appendix B. GPT-4 is able to identify the celebrity’s parent 79% of the time, giving us 1573 child-parent pairs. For each child-parent pair, we query GPT-4 to identify the child. Here, GPT-4 is successful only 33% of the time We prompt GPT-4 10 times for each question and count it as a success if it answers the question correctly at least once. Figure 1 illustrates this phenomenon. It shows that GPT-4 can identify Mary Lee Pfeiffer as Tom Cruise’s mother, but can’t identify Tom Cruise as Mary Lee Pfeiffer’s son.
This experiment may underestimate GPT-4’s ability. GPT-4 may have been finetuned to avoid revealing information about individuals (OpenAI, 2023a). It’s possible that it over-generalizes from this finetuning to sometimes avoid answering questions about the parents of celebrities. To address this, we evaluate base models from the Llama-1 family (Touvron et al., 2023), which have not been finetuned. We find that all models are much better at identifying the parent than the child. See Figure 5. Further details for Experiment 2 are in Appendix B.
Related work
Contemporary to our work, Grosse et al. (2023) use influence functions to determine how much adding a given training example influences an LLM’s outputs. They study auto-regressive pretrained LLMs of up to 52B parameters. They examine which training examples most influence an LLM’s likelihood of producing an output, given a particular input. For instance, given the input A, what most influences the likelihood of B? In their experiments, training examples that match the order (“A precedes B”) are far more influential than examples with reverse order (“B precedes A”). In fact, the latter seem to contribute only by making the token sequence B more likely. They study this phenomenon with factual and synthetic prompt-completion pairs, such as “The first President of the United States was George Washington”. These pairs are very similar to those we study in Experiments 1 and 2. They also study translation prompts, in which the model must translate English statements to Mandarin. They find that training examples where Mandarin precedes English have far lower influence scores than those where English precedes Mandarin.
Grosse et al. (2023) provide complementary evidence for the Reversal Curse. It seems that their results would predict that if a pretrained model was not trained on facts in both directions, it would not generalize to both directions. Our Experiment 1 tests and confirms a closely related prediction. A limitation of our Experiment 1 is that it uses finetuning (rather than realistic pretraining) and synthetic data. (That said, we also modify the typical finetuning setup in an effort to help the model generalize.) A limitation of Grosse et al. (2023) is that they depend on a series of approximations to classical influence functionsNote: we believe Grosse et al. (2023) provide convincing justification for the approximations. and their results are all on private models.
Further evidence for the Reversal Curse in LLMs comes from research on factual recall. Meng et al. (2023) use a model editing technique to modify factual associations. They find their method is not bidirectional, suggesting that LLMs may store factual associations differently depending on their direction. Complementing this, Geva et al. (2021, 2022, 2023) analyze the internal mechanisms behind factual recall in Transformers. They claim that these models represent factual associations as key-value pairs in their feed-forward layers. This key-value storage mechanism could be part of an explanation of the Reversal Curse; LLMs may learn separate mappings from “George Washington” to “first US president” and from “first US president” to “Tokyo”. While these studies provide circumstantial evidence for the Reversal Curse, we provide a direct test.
Previous literature has studied LLMs as knowledge bases (Petroni et al., 2019). In §2.1, we aim to extend LLM knowledge bases through finetuning, as in Zhu et al. (2020). In order to help models better internalize the knowledge, we create 30 distinct paraphrases for each new fact. In previous research (Berglund et al., 2023), we found that such augmentation can lead to robust downstream inferences. Similar approaches are used in the model augmentations literature (Sennrich et al., 2016; Cai et al., 2020; Kobayashi, 2018; Eldan & Li, 2023). Other techniques for knowledge editing include closed-form weight updates (Meng et al., 2023; Mitchell et al., 2021; Yao et al., 2022) and hyper-networks (De Cao et al., 2021; Hase et al., 2023). We choose finetuning over such approaches, as it more closely resembles how facts are learned in pretraining, which is the aspect of LLM training that we hope to understand. Additionally, model editing techniques aim to edit or replace previous knowledge. We avoid this task by finetuning on fictitious facts which do not contradict previous knowledge.
The Reversal Curse exhibits an apparent logical inconsistency in LLM knowledge, since the reversed statements are logically equivalent to the original, but in Experiment 1 are no more likely than a random baseline. Other inconsistencies are studied in (Fluri et al., 2023). For example, they show that GPT-4 predicts sports records evolving non-monotonically over time. Additionally, Hosseini et al. (2021) show that LLMs handle negations of statements incorrectly, Lin et al. (2022) show that models will sometimes output falsehoods despite having the capacity to answer statements correctly, and Shi et al. (2023) show that language models can be distracted by irrelevant text in their context.
Does the Reversal Curse apply to humans? Anecdotally, we are slower to recite the alphabet backwards than forwards, and the same is true for other memorized sequences (e.g. poems).Indeed, our findings mirror a well-studied effect in humans, wherein recall is harder in the backward direction than in the forward direction (Clair-Thompson & Allen, 2013; Thomas et al., 2003; Bireta et al., 2010; Li & Lewandowsky, 1995; Guitard et al., 2019). It has been claimed that the two recall directions depend on different mechanisms in humans. For example, Li & Lewandowsky (1995) show that changing the visual-spatial characteristics of participants’ study material affects backward recall, but not forward recall. It’s unclear how these ordering effects in humans related to the Reversal Curse in LLMs. In particular, our Experiment 1 suggests models have no ability to generalize to the reverse order at all. We do not know of such stark ordering effects in humans.
Discussion and future work
In this paper, we set out to prove a negative result. Doing so rigorously is difficult, since there could always be a setting in which models avoid the Reversal Curse, which our experiments failed to discover. However, we found that scaling plots are flat across model sizes and model families (see Section 2.1). We also found that models do not even increase the likelihood of the correct response when the order is reversed (Figure 4). Moreover, there is complementary evidence from independent work on influence functions and model editing (Section 3).
What would explain the Reversal Curse in auto-regressive LLMs? We mostly leave this for future work. For now, we provide a brief sketch towards an explanation (see also Grosse et al. (2023)). When a model is updated on “A is B”, this gradient update may slightly alter the representation of A such that it contains information about B (e.g. in the middle MLP layers as per Geva et al. (2022, 2023)). It would make rational sense for this gradient update to also alter the representation of B to contain information about A. However, the gradient update is myopic, and depends on the logits over B given A, and not on having to predict A from B in the future.The point we are making does not rule out a “meta-learning” story in which information about A and B is stored symmetrically, thus avoiding the Reversal Curse.
In addition to explaining the Reversal Curse, here are some projects for future work:
Do models fail to reverse other types of relation (as the Reversal Curse predicts)? These could include logical implications (e.g. “X implies Y” and “Not X implies not Y.”), spatial relationships (e.g. “The cup is on the table” and “The table is under the cup.”), or n-place relations (e.g. “Alice, Bob, Carol and Dan are in the same group.”)
Kandpal et al. (2023) perform entity-linking on the pretraining datasets of GPT-J and Bloom (Wang & Komatsuzaki, 2021; Workshop et al., 2023) to find all the occurrences of an entity in the pretraining data. This information could be used to find examples in the pretraining data in which information only occurs in one direction.
The pretraining sets for modern LLMs are very large and diverse. Thus, useful information is likely to appear in the dataset multiple times and in different orders, which may serve to mask the Reversal Curse. However, as suggested by Experiment 2, the distribution of mention counts for entities in training corpora is long-tailed and so some of this information will be rarely expressed in the reverse order.
Contributions and Acknowledgments
Lukas Berglund designed and implemented Experiments 1 and 2, and contributed significantly to writing the paper.
Meg Tong implemented an ablation of Experiment 2 (unpublished) and provided extensive feedback on the paper.
Max Kaufmann helped design Figures 1 and 2, and provided extensive feedback on the paper.
Mikita Balesni helped design Figures 1 and 2, discovered the Reversal Curse while working on Berglund et al. (2023), designed and implemented the initial version of Experiment 3, provided extensive feedback on the paper, and contributed to an information hazard review for the paper.
Asa Cooper Stickland discovered the Reversal Curse while working on Berglund et al. (2023), and designed and implemented the initial version of Experiment 3.
Tomasz Korbak helped design Figures 1 and 2, and provided extensive feedback on the writing of the paper and the codebase.
Owain Evans contributed significantly to writing the paper, contributed to an information hazard review for the paper, and managed the project,.
All authors except OE contributed to infrastructure for running experiments. All authors contributed to Berglund et al. (2023), which inspired this line of research.
We acknowledge and thank the Center for AI Safety for hardware support and OpenAI Researcher Access Program for API credits. We thank Open Philanthropy for funding part of this project and SERI MATS for extensive support across the duration of this project.
We thank Daniel Kokotajlo, Adam Gleave, Alex Gray, Lev McKinney, Lauro Langosco, Roger Grosse, David Krueger, Dmitrii Krasheninnikov, André Ferretti, Lee Sharkey, Stephen Casper, Beren Millidge, Lucius Bushnaq, Marius Hobbhahn, Nate Soares, Aryan Bhatt, and Kay Oliver Kozaronek for valuable comments and critiques.
References
Appendix A Additional details for Experiment 1
We assign base facts to each subset and generate paraphrases per base fact. For the “both order” subset, each fact appears times, for each ordering, accounting for examples. For PersonToDescription and DescriptionToPerson subsets, each fact appears 30 times, accounting for another examples. Thus, the dataset has a total of examples. For each PersonToDescription and DescriptionToPerson example, we have held-out paraphrases, giving us held-out prompts. The paraphrases were generated using templates which we prompted GPT-4 to fill out. Some of these prompt templates are shown in Table 2.
A.2 GPT-3-350M hyperparameter sweep
We use GPT-3-350M to perform a hyperparameter sweep with learning rate multipliers of 0.05, 0.1, 0.2, and 0.4 and batch sizes of 1, 2, 4, 8, and 16 via the OpenAI API. We do not mask loss on prompts and train for 10 epochs. We evaluate models using temperature 0. The results of the hyperparameter sweep are shown in Figure 6.
A.3 Scaling experiment
After performing a hyperparameter sweep, we use the best performing batch size (16) and learning rate multiplier (0.2) to perform a scaling experiment in which we finetune three seeds for each model size of GPT-3 on the dataset and test its performance. We used these models to obtain the results in Figure 4.
A.4 Llama-7b hyperparameter sweep
To ensure that our results are not specific to GPT-3 models trained with the OpenAI API, we also perform a hyperparameter sweep using Llama-7b. Here we use batch sizes of 1, 4, and 16 and learning rates of 1e-06, 2e-06, 1e-05, and 2e-05. The results are shown in Figure 7
A.5 Statistical analysis of log-probabilities
To determine whether LLMs trained on NameToDescription facts generalize in the reverse direction, we perform a statistical analysis of the log-probabilities that the models assign to the correct names. Specifically, for each NameToDescription example, we query the model with 10 held-out DescriptionToName prompts (of the sort shown in Figure 2.) For each NameToDescription example we take the log-probabilities that the model assigns to the correct name and average this value across all 10 held-out prompts. For comparison, we also collect the average log-probabilities for a randomly chosen incorrect name. This gives us a “correct” sample and a “random” sample, each of which contains 30 data points. To determine whether there is a statistically significant difference between the two samples, we perform two statistical tests:
Paired t-test, a test whose goal is to determine whether the two samples have a different mean.
Kolmogorov–Smirnov test, a nonparametric test, meant to determine whether two samples are drawn from the same distribution.
Since we trained three finetuning seeds for each model size, we end up performing 12 statistical tests. The results can be found in Figure 3. We do not observe statistically significant -values () for any of the finetuning seeds.
Appendix B Additional details for Experiment 2
In Experiment 2 we collect a set of 1573 child-parent relations. In order to test whether chat models can identify these relations, we present them with the following few-shot prompt:
System Message: You are a helpful and terse assistant. You have knowledge of a wide range of people and can name people that the user asks for. If the answer is unknown or not applicable, answer with “I don’t know.”
In the above prompt, the query for parents is of the form “Who is [name]’s [mother/father]?” and the query for children is of the form “Name a child of [name].” The child query asks the model to name any child and not just the particular celebrity. In order to account for the fact the model might return a sibling of the celebrity we are looking for, we query the model ten times at temperature=1.
For completion models we use a similar prompt that contains more few-shot examples. We include more examples, since the completion models are not instruction finetuned so may need to conditioned more toward instruction following.
Below is a conversation with a helpful and terse assistant. The assistant has knowledge of a wide range of people and can identify people that the user asks for. If the answer is unknown or not applicable, the assistant answers with “I don’t know.”
Q: Name a child of Barack Obama. A: Malia Obama
Q: Who is Elon Musk’s mother? A: Maye Musk
Q: Who is Kathy Pratt’s mother? A: I don’t know.
Q: Who is Chris Hemsworth’s father? A: Craig Hemsworth
Q: Name a child of Karen Lawrence. A: Jennifer Lawrence
Q: Who is Aaron Taylor-Johnson’s mother? A: Sarah Johnson
B.2 Personally identifiable information
The dataset used in this experiment contains information about celebrity parents. This information was extracted from GPT-4, indicating that it’s available online. Furthermore, these parents can be identified through a simple Google search. Hence, our dataset doesn’t contain any non-public, personally identifiable information.
Appendix C Experiment 3: Reversing instructions
In this experiment, the focus shifts to the ability of language models to reverse instructions. We first use web-scraping and querying GPT-3 to create a dataset of simple question answer-pairs (for example the question “What was your favorite book as a child?” combined with the answer “Charlotte’s Web”). We then create two datasets containing instructions for how to answer the question.
The QuestionToAnswer dataset: contains instructions of the form “Answer
The AnswerToQuestion dataset: contains instructions of the form “Answer with
After training models on these datasets, we test whether they can provide the answer when shown the question by prompting them with “Q:
We perform a hyperparameter sweep, training Llama-1 models of different sizes for five epochs. We then test on the 100 held-out question-answer pairs. The highest accuracy scores we observe are 88% for the QuestionToAnswer set and 5% for the AnswerToQuestion set. Further experiments on non-finetuned models prompted on this task show that 5% is what one would expect the best performance to be if the model were returning plausible answers randomly. These results present further evidence for the Reversal Curse.
It’s possible that models could generalize given longer training. To test this claim, rerun training for 20 epochs and 5 separate seeds using the best-performing hyperparameters from our sweep. Throughout training, the performance does not improve. The results are shown in Figure 5. The models do not perform better after 20 epochs. Instead we observe random fluctuations in accuracy over time. Hyperparameters for these experiments can be found in Appendix B.
C.2 Hyperparameter sweep
We perform a hyperparameter sweep on Llama-7b, Llama-13b, and Llama-30b for 5 epochs, using batch sizes of 8, 32, 128 and learning rates of 1e-06, 2e-06, 1e-05, 2e-05. We chose these batch sizes to be relatively low. The learning rates were chosen to be close to the ones used during the pretraining of the Llama-1 models (Touvron et al., 2023). The results for Llama-7b are shown in Figure 9.
Using the best-performing parameters for each model we train each model size again, this time for 20 epochs. We use five seeds for each model size. Again we do not observe any convergence. Instead the accuracy fluctuates randomly between 0% and 7%. A graph showing a randomly selected training run with no convergence is pictured in Figure 10.
Appendix D Compute costs
The sweeps and queries to the OpenAI API in experiments 1 and 2 cost approximately $100 each. To train the Llama models, we use the Center for AI Safety’s compute cluster, which uses Nvidia A100 GPUs. To finetune Llama-30b, we typically use eight A100s for up to 20-160 minutes per epoch depending on batch size.