Inverse scaling can become U-shaped

Jason Wei, Najoung Kim, Yi Tay, Quoc V. Le

Introduction

Scaling up language models has been shown to improve model performance for a wide range of downstream tasks and and have been claimed to unlock emergent abilities (Kaplan et al., 2020; Brown et al., 2020; Srivastava et al., 2022; Wei et al., 2022a, inter alia). However, are there any tasks for which model behavior gets worse as model scale increases? Tasks that exhibit this property have been referred to as inverse scaling tasks (Lin et al., 2022), and such tasks can help reveal flaws in the models’ training data or objectives (McKenzie et al., 2022a).

The Inverse Scaling Prize was created to identify such tasks for which larger language models show increasingly undesirable behavior, with winning submissions potentially receiving monetary awards from a $250k prize pool (McKenzie et al., 2022a). Submissions were scored based on a range of criteria including inverse scaling strength, task importance, novelty/surprisingness, task coverage, reproducibility, and inverse scaling generality across different models.

The Inverse Scaling Prize received over eighty unique submissions, with eleven tasks awarded Third Prizes, the datasets for which have been publicly released (McKenzie et al., 2022b). Inverse scaling curves for the eleven tasks were shown on a range of language models with scales spanning several orders of magnitude in parameters, including Gopher (42M–280B; Rae et al., 2021), Chinchilla (400M–70B; Hoffmann et al., 2022), and an Anthropic internal model (13M–52B). The eleven tasks are shown in Figure 3.

In this paper, we take a closer look at the scaling behaviors for these eleven tasks. First, we evaluate PaLM models of up to 540B parameters (Chowdhery et al., 2022), trained on about five times more compute than the models evaluated in the Inverse Scaling Prize submissions (see Table 1). Under this setup, we find that six out of the eleven tasks exhibit what we call U-shaped scaling: performance first decreases up to a certain model scale, and then increases again for larger models. With one task demonstrating positive scaling (monotonically increasing performance) with PaLM, this brings the number of inverse scaling tasks down to four in the context of the additional scale provided in our experiments. This finding of U-shaped scaling is consistent with prior observations of U-shaped scaling on BIG-Bench tasks such as TruthfulQA (Lin et al., 2022), Persian Idioms, and Identify Math Theorems (Srivastava et al., 2022, see Appendix C, Figure 7). The implication of U-shaped scaling is that inverse scaling curves may not extrapolate to larger scales, since performance could either keep decreasing (true inverse scaling), or start increasing (U-shaped scaling).

We do not experimentally investigate how or why U-shaped scaling occurs, but we hypothesize that it can happen when a task contains a “distractor task”. Medium-sized models can perform the distractor task better than smaller models, which hurts performance in comparison to the smaller models. As the models scale further, the larger models can ignore the distractor task and perform the true task, which can be seen as an emergent ability that derives from scaling (Ganguli et al., 2022; Wei et al., 2022a).

The second part of this paper explores whether different prompting strategies can help mitigate inverse scaling. Specifically, we test 1-shot demonstrations and chain-of-thought (CoT) prompting (Wei et al., 2022b)—a form of prompt engineering that encourages the model to decompose the task into intermediate steps. We find that simply providing 1-shot examples as part of the prompt changes all four tasks that remained inverse scaling in our PaLM evaluation to U-shaped or flat scaling. With CoT prompting, four out of the nine classification tasks that are U-shaped under 1-shot changes to positive scaling, and one of the tasks reaches near-perfect accuracy across all model sizes tested. Even when the scaling pattern does not change to positive, task performance generally improves with CoT in 8B+ models.

These results show that (even minimal) demonstrations are critically effective for avoiding distractor tasks, and point towards promising future directions for developing prompting techniques for mitigating undesirable scaling patterns.

Overall, the Inverse Scaling Prize has identified intriguing evaluation tasks for studying language model behavior with respect to scaling and prompting. We also note that the existence of U-shaped scaling does not mean that the these tasks are solved. In many U-shaped tasks, the performance of the largest model remains lower than or close to the performance of the smallest model, and often, even the best model performs close to chance. Hence, investigating how to robustly improve performance across all inverse scaling tasks would be a promising avenue for future work. Additionally, the four tasks that remain inverse scaling under the default evaluation setup merit further scrutiny even though CoT or few-shot prompting can change their scaling patterns, considering that majority of the downstream user interactions would not involve prompting with explicit demonstrations. To this end, developing methods for mitigating inverse scaling assuming a strict zero-shot setup would also be an interesting future direction.

U-shaped scaling

Setup. In this section, we evaluate PaLM models on all eleven Inverse Scaling Prize tasks. We use 8B, 62B, and 540B PaLM models presented in the original paper and also include a 1B model trained on 40B tokens, which is 0.2 zettaFLOPs of compute.This 1B model was not used in the PaLM paper (Chowdhery et al., 2022) but it followed the same training protocol. The parameter count of PaLM 540B is about twice as large as the parameter count of the largest model evaluated in the Inverse Scaling Prize (Gopher 280B), and the amount of compute used is about five times as much—2.5K zettaFLOPs versus 560 zettaFLOPs of Chinchilla 70B. We follow the exact experimental setup from the Inverse Scaling Prize (McKenzie et al., 2022a), with the same prompts and scoring protocol, where all answer choices are scored and the option with the highest probability is chosen as the prediction.The arXiv v1 of this paper used modified prompts but we changed it to match the exact prompts of McKenzie et al. (2022a) in v2+.

Results. The results for PaLM on the eleven tasks are shown in Figure 2, with the average performance of PaLM highlighted in Figure 1 on the first page. We also plot the results for Anthropic, Gopher, and Chinchilla models as reported in McKenzie et al. (2022b). In summary, only four out of eleven tasks remain inverse scaling once the PaLM 540B model is included. Six out of eleven tasks change from inverse scaling to U-shaped, and one task (Repetitive Algebra) show positive scaling with PaLM. This broad observation of U-shaped scaling demonstrates the difficulty of extrapolating inverse scaling curves to larger models.

Potential explanation. A natural question about the U-shaped scaling results is, why does performance decrease and then increase again? One speculative hypothesis is the following. Each Inverse Scaling Prize task can be decomposed into two tasks: (1) the “true task” and (2) a “distractor task” where performing the distractor task well hurts performance on the true task. Small models cannot perform either task, and performs at around chance. Medium-sized models can perform the distractor task, which results in worse performance compared to smaller models. Large models are able to ignore the distractor task and perform the true task, which then leads back to increased performance and potentially solving the task. We describe potential distractor tasks for each of the Inverse Scaling Prize tasks in Appendix B, Table 3. Note that while it could be possible to measure model performance on the distractor task only, this would be an imperfect ablation since the distractor task and true task could not only have a competing but also a joint effect on performance. We leave further explanation of why U-shaped scaling occurs to future work.

Limitations. The prevalence of U-shaped scaling does not mean that the Inverse Scaling Prize tasks are solved. Even when U-shaped scaling is observed, it is often the case that the performance of the largest model is still close to or worse than the performance of the smallest model (e.g., Resisting Correction, Modus Tollens). For several tasks, the absolute performance of the models are poor, with the best model performing near chance (e.g., Negation QA) or much worse (Pattern Matching Suppression). While we discuss several mitigation strategies to guard against undesirable scaling behavior in the remainder of the paper, these observations demonstrate the inherently challenging nature of the task, highlighting an opportunity for future research towards improving absolute performance on these tasks.

Mitigation strategies for inverse scaling

We next explore possible mitigation strategies for inverse scaling. In Section 2, we hypothesized the primary cause of inverse scaling to be distractor tasks that mislead the models towards a different solution from the true task. Then, in-context demonstrations of a problem/solution pair could discourage the models from solving the distractor task, since the answer according to the true task diverges from the answer according to the distractor task. If such demonstrations are accompanied by explicit rationales behind the reasoning process, this could guide the models towards identifying the true task even more strongly. To this end, we explore whether 1-shot demonstrations and 1-shot demonstrations with chain-of-thought reasoning improve undesirable scaling patterns.

To gauge the effect of demonstrations, we re-evaluate the PaLM models on all tasks with 1-shot prompts, using the 1-shot dataset provided as part of the Inverse Scaling Prize data release. This officially released 1-shot dataset is created by pairing each example in the dataset with a randomly sampled, different example in the dataset. Then, the 1-shot examples are simply prepended to the default prompts shown in Figure 3.

We find that all four tasks that continued to be inverse scaling after including the 540B model shift to U-shaped or flat scaling when prompted with 1-shot demonstrations. Specifically, Pattern Matching Suppression, Into the Unknown, and Prompt Injection change to U-shaped scaling, and Redefine changes to flat scaling (see Figure 4). We can also see that the performance of the largest 540B model benefits from 1-shot prompting in all four tasks. These results show that even a single example of a problem/solution pair is effective for encouraging the models towards solving the true task, especially for larger models.

The tasks that were already U-shaped with unmodified prompts remain U-shaped. See Appendix A, Table 2 for full results on all tasks.

2 Chain-of-thought helps U-shaped scaling become positive scaling

While our 1-shot results are promising in that even a single demonstration helps shift the inverse scaling trend to U-shaped or flat scaling, for most tasks, the performance of the largest model (540B) still fell behind or was not substantially better than the smallest model tested (1B). This pattern held true for six out of the ten U-shaped or flat tasks under the 1-shot setup (Negation QA, Memo Trap, Into the Unknown, Modus Tollens, Redefine, and Prompt Injection). We explore whether chain-of-thought (CoT) prompting can help in such scenarios, based on the recent work showing that CoT can improve performance by a large margin for multi-step reasoning tasks by outputting intermediate steps before giving the final answer (Wei et al., 2022b; Kojima et al., 2022; Suzgun et al., 2022, inter alia).

For the experiments in this section, we use prompts that follow the protocol of Wei et al. (2022b) and follow-up work that includes intermediate reasoning steps in the in-context demonstrations. We continue to use a single demonstration example as in Section 3.1, but now the demonstrations are paired with step-by-step rationales for the answers. Because CoT prompting also requires the models to generate intermediate steps, we use free-form generation followed by exact string match to evaluate model performance. This requires one additional modification to the prompt to facilitate the postprocessing of the model generations. Specifically, the model is prompted to output the final answer following the expression “So the answer is”.All prompts used in this section are made available at: https://github.com/jasonwei20/inv-scaling-prompts/. Other than these additions, the phrasing of the instructions and the structure of the prompts are kept as close as possible to the original 1-shot prompts. We construct CoT prompts for ten inverse scaling tasks, excluding Prompt Injection that uses loss instead of classification accuracy as the metric. Examples of the CoT prompts are shown in Figure 5.

We show results for six tasks in Figure 6: three classification tasks that were inverse scaling in PaLM (Into the Unknown, Pattern Matching Suppression, and Redefine) and all other U-shaped tasks where the 540B model performed worse or only similarly to the 1B model even after 1-shot demonstration (Negation QA, Modus Tollens, and Memo Trap). Overall, CoT improves performance on these tasks by a large margin with the exception of Redefine where there is a small gain only in the 540B model (∼\sim6 percentage points over 1-shot). The scaling curves change to positive (monotonically increasing) for Into the Unknown, Pattern Matching Suppression, Redefine, and Negation QA, although for Redefine this is a byproduct of smaller models underperforming their 1-shot counterparts. For Memo Trap, we observe an inverted-U-shaped curve where the performance drops slightly with the largest model; nevertheless, there are consistent performance gains via CoT in 8B+ models.The lower performance in 1B observed across several tasks is likely due to the limited capacity of smaller models to perform CoT reasoning. For Modus Tollens, CoT-prompted models achieved almost perfect accuracy regardless of size (i.e., flat scaling but saturated performance). See Appendix A, Table 2 for full results.

Overall, 8B+ models benefit from CoT prompting in almost all tasks. In many cases, CoT also helps change U-shaped scaling to positive scaling, showing the promise of intermediate rationales in addition to problem/answer demonstrations as an effective mitigation strategy for undesirable scaling patterns.

Conclusions

This paper has two simple takeaways. First, inverse scaling can turn into U-shaped scaling when evaluated on models of sufficiently large scale, as demonstrated on six out of eleven Inverse Scaling Prize tasks. The prevalence of U-shaped scaling we identified in this paper shows that inverse scaling curves do not necessarily extrapolate to larger models. Second, demonstrations and rationales are effective for mitigating undesirable scaling patterns. All inverse scaling tasks change to U-shaped or flat scaling when a single demonstration is provided as a part of the prompt. With additional intermediate reasoning steps, many of the U-shaped tasks further shift to positive scaling, as well as substantial performance gains throughout.

Taken together, the implication is that a combination of scaling and prompting techniques appear to be a viable method for mitigating inverse scaling. However, the prompting approaches explored in this paper has limitations in that they require manual construction of demonstrations and reasoning steps tailored to individual tasks. This leaves open an interesting future research direction of developing solutions for inverse scaling that do not require explicit demonstrations.

Acknowledgements

Thanks Ethan Perez and Ian McKenzie for their help with sharing the Round 2 data in the fourth version of the report. Thanks Ethan Perez, Ian McKenzie, and Najoung Kim for help with the third version of the report. Thanks Ethan Perez for feedback that we incorporated into the second arXiv version of the report. Thanks Denny Zhou, Ed Chi, and Le Hou for feedback on the initial report. Finally, we really appreciate the spirit and organization of the Inverse Scaling Prize organizers—thank you!

References

Appendix

The full results for all eleven Inverse Scaling Prize tasks reported this paper are shown in Table 2 below. We used the exact dataset and protocol from McKenzie et al. (2022a) for the main experiments (Section 2), and used the officially released 1-shot dataset for the 1-shot experiments (Section 3.1).The official 0- and 1-shot datasets are from https://github.com/inverse-scaling/prize/tree/main/data-release. These experiments are marked 1-shot (official). We additionally ran 1-shot experiments where we fixed the 1-shot demonstration to be the same as the CoT demonstration, except for the step-by-step rationale, marked 1-shot (controlled). This is because the official 1-shot dataset used a a randomly sampled example from the dataset as the 1-shot demonstration example, which varied across each example in the test set. Since our CoT experiments (Section 3.2) use a single manually written demonstration for every test example, the CoT results are more directly comparable to the controlled 1-shot experiments where the demonstrations are fixed.

Appendix B Distractor tasks

A possible hypothesis for why U-shaped scaling emerges is as follows. U-shaped scaling tasks consist of a true task and a distractor task. Medium-sized models are good enough to perform the distractor tasks, which hurts performance compared to smaller models that cannot perform the distractor task nor the true task. Larger models can ignore the distractor task and perform the true task, which leads to increased performance again. We show a speculative decomposition of tasks into the true task and a distractor task in Table 3.

Appendix C Prior examples of U-shaped scaling

Appendix D Model scale: parameters, data, and compute

As shown in Table 4, we computed training FLOPs following the protocol of Brown et al. (2020).

In the second version of the arXiv paper, it was reported that only two of the four first-round tasks were U-shaped. However, actually three of the were U-shaped. This error was because I (Jason) accidentally swapped the PaLM 62B numbers for Hindsight and NeQA. I realized the error when I reproduced those tasks for the third arXiv version.

In the fourth version of the arXiv paper, it was reported that one task (Redefine) was still inverse scaling for the 1-shot experiments. However, this was due to an error in the initial Inverse Scaling Prize dataset release. With the corrected dataset, all tasks that were still inverse scaling after the inclusion of PaLM 540B turn to U-shaped scaling after 1-shot. We also fixed the incorrect token count for Anthropic LMs (850B →\rightarrow 400B) and the resulting FLOP counts.