Observational Scaling Laws and the Predictability of Language Model Performance

Yangjun Ruan, Chris J. Maddison, Tatsunori Hashimoto

Introduction

Language model (LM) scaling plays a central role in discussions of model capabilities and affects everything from the tasks they can perform to the effectiveness of post-training techniques such as Chain-of-Thought . Due to this importance, understanding and predicting LM behaviors across scales, benchmarks, and algorithmic interventions is a major question for many researchers and engineers. Machine learning researchers may wish to understand whether their proposed algorithmic interventions remain effective in the face of future model scaling, while engineers and benchmark builders may wish to understand whether complex capabilities such as agentic abilities will scale predictably in the same way as existing LM benchmarks.

Scaling laws have been powerful tools for understanding the scaling trend of LMs, which have shown that LMs follow a precise power-law relationship between compute measures (such as training FLOPs) and downstream capabilities ranging from perplexity to benchmark performance . This power-law relationship has been used in a variety of ways – including hyperparameter and architecture selection as well as model capability forecasting . Unfortunately, scaling analyses remain uncommon in many benchmarking and post-training studies, as most researchers do not have the compute resources to build scaling laws from scratch, and open models are trained at too few scales (3-5) for reliable scaling predictions.

Although the high costs of compute scaling laws are unavoidable when optimizing pre-training hyperparameters (e.g., Hoffmann et al. ), this is not true of all scaling analyses. In this work, we show that many other types of scaling studies, such as understanding complex model capabilities (e.g. agentic or “emergent” behaviors) and post-training interventions, can be done using a lower-cost, higher-resolution, and broader-coverage alternative to the standard approach of training (or using) a single family of LMs across compute scales.

The starting point of our work is the observation that there now exist hundreds of open models spanning a large range of scales and capabilities. While we cannot directly use these models for compute scaling laws (as the training compute efficiency varies widely across model families), we might hope that there exists a more general scaling law that holds across model families. In particular, we hypothesize that the downstream performance of an LM is a function of a low-dimensional space of capabilities (e.g., natural language understanding, reasoning, and code generation), and that model families vary only in the efficiency by which they convert training compute to these capabilities. If such a relationship held, it would imply that there is a log-linear relationship from low-dimensional capabilities to downstream capabilities across model families (which would allow us to build scaling laws that leverage all existing models), as well as a log-linear relationship between training compute and capabilities within each model family (as in standard compute scaling) (Fig. 1).

Through an analysis of existing standardized LM benchmarks (e.g., Open LLM Leaderboard ), we find a few such capability measures that have scaling law relationships with compute within model families (R2>0.9R^{2}>0.9) (Fig. 3), and with downstream metrics across model families. We call such scaling relationships observational scaling laws as they relate simple observable quantities that we expect to scale with compute (such as standardized benchmark performance) with complex downstream quantities of interest.

The ability to build scaling laws across a large number of existing models has significant advantages in cost, resolution, and coverage: observational scaling incurs no training cost, while leveraging a large number of models spanning a much larger compute range than any single model family. Observational scaling also significantly increases the resolution of scaling laws by virtue of using more models, which is useful for studying nearly discontinuous phenomena like “emergent” capabilities. Finally, observational scaling can combine model families from heterogeneous sources with very different scaling properties and capabilities (e.g., LLaMA vs StarCoder ) which allows us to study how different scaling strategies impact downstream performance and algorithmic interventions.

Finally, we show that using observational scaling laws is low-cost and straightforward, as there are a few model families that are sufficiently representative to replicate many of our core findings (Sec. 5). By using these representative families, we find that future works can easily make scaling predictions on benchmarks and post-training interventions by evaluating only 10-20 models.

We demonstrate the utility of observational scaling laws in three different settings that are challenging for compute scaling laws but are accurately predicted using observational scaling laws. While our results are based on systematic holdout validation with currently available models, we preregister our fitted scaling laws and commit to updating their prediction accuracy on future models (Sec. 4).

Emergent capabilities (Sec. 4.1) There has been an active debate about whether LMs have “emergent” capabilities that discontinuously appear at certain compute thresholds and whether these capabilities can be predicted using small models . The high resolution of observational scaling laws show that some of these phenomena follow a smooth sigmoid, and can be predicted accurately using small, sub Llama-2 7B models.

Agentic capabilities (Sec. 4.2) We show that the more high-level, complex capabilities of LMs as agents, as measured by AgentBench and AgentBoard , can be predicted using observational scaling laws. Our scaling law precisely predicts the performance of GPT-4 using only weaker models (sub GPT-3.5) and identifies programming capabilities as driving agent performance.

Post-training method scaling (Sec. 4.3) We show that our scaling laws can reliably predict the gains of post-training techniques, such as Chain-of-Thought and Self-Consistency at scale, even when we fit our scaling laws on weak models (sub Llama-2 7B).

The contribution of our work is as follows: our conceptual contribution is to propose observational scaling which leverages predictable log-linear relationships between compute, simple capability measures, and complex downstream metrics. Our empirical contributions include identifying a small number of capability measures that cover standard LM benchmarks, demonstrating that these measures provide accurate predictions on a number of complex LM capabilities, and selecting a small set of model families that are useful for low-cost observational scaling analyses.

Related Work

In standard scaling laws , the “scale” is defined by the compute resources allocated to training LMs, such as the number of training FLOPs CC, model parameters NN, and training tokens DD. Scaling laws are typically formulated as a power-law relationship between LMs’ cross-entropy loss LL and their compute scale measures. Common functional forms include L(N,D)=aNα+bDβ+eL(N,D)=\frac{a}{N^{\alpha}}+\frac{b}{D^{\beta}}+e or L(C)=cCγ+hL(C)=\frac{c}{C^{\gamma}}+h , where C≈6NDC\approx 6ND for the Transformer . The parameters {α,β,a,b,e}\left\{\alpha,\beta,a,b,e\right\} or {γ,c,h}\left\{\gamma,c,h\right\} are fitted by training LMs across different compute scales, varying NN and/or DD, and measuring their loss. Our work differs from compute scaling laws in our goals – compute scaling aims to understand the scaling properties of pretraining, and thus focuses on a single model family and relates downstream performance to directly controllable quantities such as training compute. In contrast, we are interested in scaling laws for downstream, post-training performance, which leads us to consider scaling laws across model families and use more directly observable capability measures than compute.

Scaling laws have been generalized beyond pretraining loss to analyze transfer learning and downstream performance across various domains, see Villalobos for a comprehensive review. In particular, there has been evidence suggesting that the few-shot performance of LMs on downstream benchmarks is closely tied to compute measures like model size , but whether this is predictable with scaling laws remains debated. Extensive research has explored the difficulties of predicting benchmark performance due to their appearing rapid “emergence” , while recent works argued the discontinuity is due to the metrics used or the lack of data points (see Anwar et al. for a survey on this topic). Finnveden and Owen have investigated the use of linear and sigmoidal scaling laws, derived from pretraining loss or computational measures, to extrapolate the benchmark performance. Recent studies have also more extensively investigated the correlations between the pretraining loss and downstream performance of LMs , aiding in the understanding of downstream scaling and emergent capabilities of LMs. On the theory front, Arora and Goyal derived a theory characterizing how performance on complex skills of LMs can be derived as a composition of base skills. While our work shares similar goals in that we aim to understand the downstream, post-training performance of models, we differ in our approach in that we aim to build practical higher-resolution scaling laws using multiple model families and their observable standard benchmark metrics.

Numerous works have investigated the correlations between different benchmarks across various contexts. Extensive research has explored the relationship between the out-of-distribution performance and in-distribution performance of machine learning models . In the realm of NLP and LM benchmarks, Qiu et al. , Torregrossa et al. found that different evaluations and metrics for word embeddings are highly correlated, and Liu et al. observed a strong correlation between question-answering benchmarks. Moreover, Perlitz et al. , Polo et al. observed strong correlations between samples within various LM benchmarks and utilized this observation to develop more efficient benchmarks. Most relevant to our work, Ilić found that a single factor explains 85% of the performance on the Open LLM Leaderboard and GLUE leaderboard , while Burnell et al. extracted three factors for LM capabilities that account for 82% of the variation on the HELM benchmark , aligning with our observations. Our work also observes such benchmark correlations and low-rank structures but is unique in utilizing these properties for the purpose of scaling predictions that can be used directly for benchmark and algorithm development.

Observational Scaling Laws

In this section, we introduce our observational scaling laws that generalize the standard compute scaling laws (Sec. 3.1). The key idea is to extract a low-dimensional capability measure for LMs from their observable benchmark performance (Fig. 2), which we find has a log-linear relationship with compute scale measures (Sec. 3.3) and can thus be used as surrogate “scale” for scaling analysis of complex LM capabilities (Sec. 3.4).

In compute scaling laws, there is a hypothesized power-law relationship between models’ compute measures CmC_{m} (e.g., training FLOPs) and their errors EmE_{m} (e.g., perplexity). Specifically, for a model mm within a family ff (e.g., Llama-2 7B, 13B, and 70B) we hypothesize

and if this linear fit is sufficiently accurate, we draw inferences about the performance of a model at future compute scales C′>CC^{\prime}>C by extrapolating this relationship. However, fitting such a scaling law can be tricky, as each model family ff and downstream benchmark has its own scaling coefficients βf\beta_{f} and αf\alpha_{f}. This means that scaling experiments, especially for post-training analysis, are often fitted on very few (3-5) models sharing the same model family, and any predictions are valid only for a specific scaling strategy used within a model family.

Several studies [e.g., 24, 63] have generalized the functional form to analyze the scaling of LMs’ downstream performance. Specifically, let EmE_{m} represent the normalized downstream errors of models within the range $,theyobservedasigmoidalrelationshipbetween, they observed a sigmoidal relationship between\log(C_{m})andandE_{m}$ and thus used a logistic link function instead of a logarithm for the generalized linear model in Eq. 1:

We can view Eq. 3 and Eq. 4 as a generalization of Eq. 2, since combining them can recover the original scaling relationships for a single model family. However, when there are multiple model families, SmS_{m} serves as a shared, low-dimensional space of model capabilities from which all downstream metrics (EE and BB) are derived (as indicated by the absence of ff in Eq. 3 and Eq. 5), and model families only vary in their efficiency in converting compute into capabilities (Eq. 4). One useful way of interpreting Eq. 4 is that θf\theta_{f} represents the compute efficiency of a model family ff, and SmS_{m} is the capabilities of model mm expressed in terms of log-FLOPs for this model family.

At this point, it is not yet clear that Equations 3, 4, and 5 hold in practice. In next subsections, we validate Eq. 5 (Fig. 2) and Eq. 4 (Sec. 3.3) separately, and then present our estimation algorithm for Eq. 3 in Sec. 3.4. In Sec. 4, we will perform a more extensive validation of Eq. 3.

2 Identifying a Low-Dimensional Capability Space (Eq. 5)

We validate the existence of a low-dimensional capability measure SS that linearly relates to standard LM benchmarks BB by showing that only a few principal components of BB capture most of its variation (Eq. 5). We demonstrate that the benchmark-model matrix BB for a reasonable, broad set of benchmarks and models is low-rank and that Eq. 5 is a reasonable assumption. As this type of analysis depends heavily on the set of models and benchmarks chosen, we carefully describe our selection process below.

Since the benchmark-model matrix BB can be directly measured for any LM, we include a large number of publicly accessible models for subsequent analysis. We collected a broad set of open LMs covering 21 model families (a collection of models across scales such as LLaMA-2 7B, 13B, 70B) and a total of 77 models. These encompass models trained from heterogeneous recipes, including standard training recipes like LLaMA and Qwen , those trained on synthetic data like Phi , and models specifically trained on code data like CodeLlama and StarCoder . For this analysis, we consider only pretrained base models to avoid the complexities introduced by instruction tuning. We also include an analysis for instruction-tuned models that include proprietary ones like GPT-4 and Claude-2 in Sec. C.1, which demonstrates similar results. See table B.1 for a detailed list of collected models.

We collected a set of diverse benchmarks that assess various LMs’ capabilities. These include popular aggregated benchmarks like MMLU that assess the general knowledge of LMs. For more specialized evaluations, we included ARC-C , HellaSwag , Winogrande for commonsense reasoning, GSM8K for mathematical reasoning, HumanEval for programming, TruthfulQA for truthfulness, and XWinograd for multilingual capabilities. We carefully collected these metrics from standardized evaluation protocols for comparability across LMs. In particular, we compiled them from standardized leaderboards, like the Open LLM Leaderboard and EvalPlus , when available. Otherwise, we used standardized libraries such as the LM Eval Harness to evaluate the LMs. See Sec. B.1 for full details of our data collection pipeline.

After obtaining the benchmark metrics for the LMs, we addressed potential missing values (less than 1%1\% of all data), which may have occurred due to evaluation failures, by using PCA imputation. Subsequently, we applied PCA to extract the principal components of the evaluation metrics as the “principal capability” (PC) measures SS (additional details in Sec. B.3).

We observe that the extracted PC measures are predominantly low-rank, with the top 3 PCs explaining ∼97%\sim 97\% of the variance, which supports a low-dimensional representation of benchmarks BB (Fig. 2(a)). Surprisingly, we find that the first PC alone explains nearly 80% of the variation in LM capabilities. Taking a closer look at these PCs, we find that these capability measures represent interpretable directions in which LMs capabilities may naturally vary as a function of scale (Fig. 2(b)). Specifically, PC-1 represents the “general capability” as a weighted average of all metrics; PC-2 corresponds to the “reasoning capability”, emphasizing mathematical and coding benchmarks; and PC-3 primarily reflects the “programming capability”. These findings suggest that many simple LM capabilities (as covered in our benchmarks) can be expressed as a linear combination of just a few “principal capabilities” SS.

3 Principal Capability Measures as Surrogate Scale Measures (Eq. 4)

We now show that the extracted PC measures SS scale log-linearly with training FLOPs within each model family, and can thus be interpreted as a cross-model generalization of compute scale CC.

We collected all available information about training FLOPs on each of our models, analyzing papers and other public information to identify model size NN and pretraining data size DD. For the models where we were able to identify this information, we used the simple estimate of C≈6NDC\approx 6ND to obtain model training FLOPs . See table B.1 for our collected compute measures.

Fig. 3 illustrates the correlation between the top PC-1 measure with the corresponding training FLOPs for models within each model family. We find that for each model family with controlled training recipes and comparable compute scale measures, the LMs’ PC-1 measure linearly correlates with their log-training FLOPs (with R2>0.9R^{2}>0.9). This linear correlation holds across a broad range of model families including those specifically trained on multilingual data like BLOOM or those on code like StarCoder . It also generally holds for lower-ranked PCs such as PC-2 and PC-3, as shown in Fig. C.2. Together with Fig. 2, these results support the validity of Equations 5 and 4, in which we hypothesized that models share the same capability space and a log-linear relationship determines the efficiency by which each model family converts their compute into these principal capabilities.

4 Fitting Observational Scaling Laws

Having validated that a simple PC analysis leads to capability measures SS that approximately fulfill equations 4 and 5, we now define a procedure to estimate the scaling relationship in Eq. 3. The complete algorithm is presented in algorithm 1.

Given a certain downstream error metric EE normalized to $$ that measures certain LM capabilities, we slightly generalize Eq. 3 to

Recall that the core component of our scaling law is the fitted linear transformation Pm:=β∗⊤Sm+α∗P_{m}\vcentcolon=\beta^{*\top}S_{m}+\alpha^{*} which maps the extracted PCs into a scalar capability measure for a target downstream metric. While this is perfectly acceptable for prediction, our scaling analysis would be more interpretable if we expressed capabilities in units of FLOPs rather than an arbitrary scalar capability measure.

Recall that our observational scaling laws generalize compute scaling laws for a single model family (Eq. 3 & Eq. 4). Thus, for a specific family ff, our observational scaling laws should correspond to some compute scaling law. Specifically, we note that when Eq. 4 holds exactly, we have that for a model mm within a family ff,

where wf=β∗⊤θfw_{f}=\beta^{*\top}\theta_{f} and bf=β∗⊤νf+α∗b_{f}=\beta^{*\top}\nu_{f}+\alpha^{*}. This implies a linear correlation between the scalar capability PmP_{m} and the compute log⁡(C)\log(C) for models within a specific family on a downstream task (see empirical validation in Fig. C.3). Since θf\theta_{f} and νf\nu_{f} are unknown a priori, we can fit these coefficients wf,bfw_{f},b_{f} via linear regression from log⁡(C)\log(C) to PP using models from the specific family ff.

In the multi-model family case, variations in compute efficiency mean that FLOPs and capabilities are no longer log-linear across model families. However, we can map all of the models to a shared, FLOPs-based capability measure using a metric we call ff-equivalent FLOPs. The core idea of the approach is to represent each model’s capabilities by the following hypothetical: “how many log-FLOPs (log⁡(Cˉm,f)\log(\bar{C}_{m,f})) would it take for a model in a family ff to match a model mm”. We call log⁡(Cˉm,f)\log(\bar{C}_{m,f}) the ff-equivalent FLOP for model mm, as it represents the performance of model mm relative to models in the reference model family ff. This measure can be computed fairly easily as

obtained from solving for log⁡(Cm)\log(C_{m}) in Eq. 7. Throughout the remainder of this work, we apply this scalar transformation where we pick Llama-2 as the reference family ff, and so the x-axis of all of our plots can be interpreted as “model capabilities, as measured in units of Llama-2 FLOPs”.

Validating Observational Scaling Laws

We evaluate the usefulness of observational scaling laws by showing that they accurately predict the scaling behaviors of LMs over complex, hard-to-predict phenomena (like emergent phenomena and agentic abilities) and help estimate the value of techniques such as Chain-of-Thought.

To ensure that our scaling laws are actually predictive and that we are not simply overfitting through various choices in scaling law construction and hyperparameters, we design our experiments to have systematic holdout sets and robustness checks. We also preregister our predictions for future models after the release of the paper as a test of whether our scaling laws overfit current models. We release our code including the implementation and collected data at https://github.com/ryoungj/ObsScaling.

For extracting PC measures, we fixed the number of PCs K=3K=3 as it covered ∼97%\sim 97\% of the variation in benchmark performance and it consistently yielded the best performance across most of our experiments, see Sec. C.3 for robustness checks on PC selection. For the capability-equivalent scale transformation, we used the Llama-2 as the reference model family as it is currently the most representative and widely used open model in the community. For better interpretability and visualization, we used the accuracy metric, typically defined as Y=1−EY=1-E, for fitting the scaling laws and making the plots.

To validate our observational scaling laws, our primary objective is to assess how accurately the scaling laws fit the available data and extrapolate from smaller-scale, less capable models to larger-scale, more powerful models. We validate this through systematic holdouts for the test set, where we split available models into weaker and stronger ones based on both scale or capability (e.g., FLOPs or accuracy). We used the weaker models to fit the scaling law and evaluated the extrapolated predictions on the stronger ones. To prevent any train-test leakage, all preprocessing steps (e.g., PCA imputation) were fitted on the train set only and then applied to the test set. Unless otherwise stated, we set the cutoff to include all models with training FLOPs less than or equal to that of Llama-2-7B (8.4×10228.4\times 10^{22}) as training data, resulting in a training set of 47 models and a test set of 30 models. We included robustness checks for different holdout strategies in Sec. C.3.

As baselines, we compare our scaling predictions to existing compute-based scale measures like training FLOPs and model size. We used the mean squared error (MSE) on the holdout set as our main evaluation measure, as the target range is always normalized (0 to 1), and estimating the marginal variance in R2R^{2} can add additional noise when the test set sizes are small.

In Sec. C.7, we include all functional forms for the fitted scaling laws in our experiments as preregistration of our predictions for future models. We will assess the accuracy of these scaling laws (without refitting) using models developed after May 2024 and commit to updating the manuscript on ArXiv with our prediction results after 4 months.

1 Predictability of “Emergent” Capabilities

Recent works have argued that many LM capabilities are “emergent” and cannot easily be predicted from small-scale models . Discontinuous changes to capabilities would make it difficult to develop algorithms and benchmarks that are effective at scale, and there have been ongoing debates – about whether these capabilities are truly discontinuous and whether the discontinuity is an artifact of the metric used or lack of high-resolution data points .

The debate on emergent phenomena has been complicated by the fact that existing scaling analyses (including the original ones in Wei et al. ) have very few points . When there are only 5 models across many orders of magnitudes of scale, phenomena can appear to be discontinuous, even if the underlying phenomenon is a smooth but rapidly varying sigmoid.

We show that the higher resolution of observational scaling laws allows us to clearly see smooth sigmoidal curves in phenomena that were identified as emergent in Wei et al. , and even more surprisingly, we can often accurately forecast the transition points where models go from near-random to high performance using only models whose performance is only slightly above random. Our findings validate the observational approach to scaling laws and provide evidence that higher-resolution scaling laws could help us better understand scaling phenomena for LMs.

We tested on four BigBench tasks that were labeled as “emergent” in Wei et al. , including two arithmetic tasks (3-digit subtraction and 2-digit multiplication) and two non-arithmetic tasks (word unscramble and Persian QA). Additional results on more tasks covering Wei et al. are included in Sec. C.4. For the models, we included base pretrained models following the approach of Wei et al. . For non-arithmetic tasks, we used the default FLOPs cutoff. For arithmetic tasks, we found that this cutoff resulted in an excess of training data near perfect performance (see results in Fig. C.11), making the prediction tasks trivial. Consequently, we reduced the cutoff to a quarter of the default value and also excluded GSM8K (which may be a superset of arithmetic tasks) from our base metrics BB to make the tasks more challenging.

Fig. 4 shows our prediction results using our PC measures as well as the baseline of predicting performance based on training FLOPs. We find that these capabilities can be accurately predicted using our PC measures, even when only using models that perform poorly. In contrast, using training FLOPs results in significantly poorer extrapolation on the test set and fits on the train set, as indicated by the much higher MSE values. This discrepancy is likely due to the incomparability of training FLOPs across different model families. Additional results of the model size baseline are included in Sec. C.4.

2 Predictability of Agentic Capabilities

There is significant interest in building autonomous agents using LMs, with notable examples including AutoGPT , Devin , and SWE-agent . Although the performance of these agents still falls far below human-level on challenging real-world tasks , there is a belief that future models at larger scales will significantly enhance these agents’ capabilities. However, there is a significant uncertainty about whether existing models that are trained for language and code capabilities will transfer well to agentic tasks that require taking actions over many rounds. In this section, we utilize our observational scaling laws to analyze the scaling properties of LMs’ agentic capabilities w.r.t. their backbone model capabilities and show that agent performance is highly predictable from simple benchmark metrics.

We tested on two standardized agent evaluation benchmarks, AgentBench and AgentBoard , each is a collection of diverse tasks for evaluating LMs’ generic agentic capabilities. For both benchmarks, we utilized their provided aggregated metrics on all tasks for prediction. Specifically, we used the “Overall Score” on AgentBench, which is a weighted average of scores across all tasks (denoted as “OA” in the benchmark), and the “Average Success Rate” on AgentBoard. We included models that have been evaluated on each benchmark, which encompasses both open instruction-tuned models like LLaMA-2-Chat and Vicuna , and proprietary models like GPT-4 and Claude-2 , see table B.2 for a complete list of included models.

We followed the same procedure to collect standardized benchmark metrics BB for instruction-tuned models, including MMLU , ARC-C , HellaSwag , Winogrande , TruthfulQA , GSM8K , and HumanEval , see Sec. B.1.2 for details. The PC measures extracted for these instruction-tuned models followed a similar pattern to those of pretrained base models, as shown in Fig. C.1. Notably, since compute scale measures are not available for proprietary models, only our observational scaling laws apply here and not compute scaling laws. The default FLOPs cutoff does not apply either, and thus we held out the top 10% performing models on each agent benchmark as the test set to simulate weak-to-strong predictions, which included GPT-4 and Claude-2 on AgentBench and GPT-4 on AgentBoard.

Fig. 5 illustrates the prediction results with our observational scaling laws using PC measures. We find that on both agent benchmarks, the performance of held-out models (GPT-4/Claude-2) can be accurately predicted from models with much weaker performance (> 10% gap). This indicates that the more complex agentic capabilities of LMs are well-correlated with and predictable from their base model capabilities, suggesting the promising scaling properties of LM-based agent capabilities as backbone LMs continue to scale up.

In Fig. 5(c), we visualize the weights assigned to the base evaluation metrics on both benchmarks, which are derived from the regression weights fitted on PC measures and applied with learned PCA transformation, i.e., β⊤γ\beta^{\top}\gamma. We observe that the fitted weights assign significant importance to programming capabilities (HumanEval) on both benchmarks, underscoring its significance in defining the agentic capabilities of LMs. The weights also emphasize general knowledge (MMLU) on AgentBench, and reasoning capabilities (GSM8K) on AgentBoard, suggesting that these capabilities may also be important for LMs’ agentic capabilities.

3 Predicting the Impact of Post-Training Techniques

When researchers propose a new prompting or post-training technique to improve a pretrained model, how can we know whether these gains will persist across models and scales? Scaling analysis could enable more quantitative approaches to the design of post-training interventions, but systematic scaling analyses have been rare due to the small number of models within a single model family. Adding to these challenges, some recent works have argued that certain interventions, such as Chain-of-Thought , behave in an emergent way and their behaviors are not predictable from smaller models . Using observational scaling laws, we show that it is possible to make relatively accurate predictions on the effectiveness of techniques such as Chain-of-Thought (CoT) and Self-Consistency (SC) as model scale increases. We focus on these post-training interventions in particular, as they are sometimes discussed as examples of post-training interventions that require scale to be effective .

Our approach to quantifying the scaling properties of post-training is straightforward: we fit one observational scaling law using base model performance on a target benchmark (e.g., GSM8K few-shot), and then fit another on the performance of models with the post-training intervention (e.g., GSM8K w/ CoT). Each of these fits produces a sigmoidal scaling curve as a function of log⁡(Cˉf)\log(\bar{C}_{f}), and the relative gaps as a function of log⁡(Cˉf)\log(\bar{C}_{f}) indicates the scaling efficiency of the intervention.

We tested on GSM8K with CoT and SC as post-training techniques and included additional results on BigBench-Hard with CoT in Sec. C.5. As with our study on emergent phenomena on arithmetic tasks, we excluded GSM8K from the base metrics BB to avoid making the prediction tasks trivial. We included all the pretrained base models listed in table B.1 including those specifically trained for code data and applied the default FLOPs cutoff for holdout validation. For CoT, we followed Wei et al. and compared CoT prompting using eight reasoning examples with naive prompting using only few-shot examples in the greedy decoding setting. For SC, we sampled five CoT reasoning paths at temperature 0.7 to aggregate the final answers following Wang et al. and compared it with a single sampled CoT answer.

Fig. 6(a) shows the scaling predictions for CoT and SC using observational scaling laws. We find that the performance with (CoT, CoT + SC) and without (Naive) post-training techniques for stronger, larger scale models can be accurately predicted from weaker, smaller scale models. In contrast, predictions based on compute scale measures like model size and training FLOPs are less reliable as seen in Fig. C.13. Notably, the scaling trends between the two techniques differ; CoT shows a much more pronounced scaling trend compared to Self-Consistency w/ CoT.

Another advantage of observational scaling laws over scaling laws constructed on single families is that we can visualize the capabilities that are important to the post-training intervention. Fig. 6(b) visualizes the fitted regression weights β\beta, mapped to the space of base capability benchmarks BB via β⊤γ\beta^{\top}\gamma. We clearly see that when we go from Naive to CoT, there are significantly higher weights placed on MMLU and HumanEval - meaning that scaling models in a way that enhances general knowledge (MMLU) and code (HumanEval) leads to greater gaps between CoT and the baseline, while improving along commonsense, such as Winogrande does not necessarily lead to improvements at scale. These analyses can inform how different post-training interventions affect different scaling recipes – such as code models vs general-purpose LLMs.

Selecting Low-Cost Model Subsets for Practical Scaling Analyses

We have now demonstrated the effectiveness of observational scaling laws in forecasting the scaling behavior of various LM capabilities. However, the large number of publically available models is both a strength and a weakness – it enables much higher resolution scaling analyses, but it also requires us to evaluate our benchmarks and post-training methods on a larger number of models.

To make observational scaling analyses more broadly accessible, we identify a small set of models that maintain high prediction accuracy while significantly reducing the evaluation cost. We do this by building upon the classic approaches in optimal experimental design which allow us to define optimality criteria for selecting model subsets without knowing the downstream task.

More specifically, we consider the constrained optimization problem of identifying the optimal set of models to choose for a regression problem, subject to the constraint that we select a model subset M\mathcal{M} of at most MmaxM_{\text{max}} models from the set of all models Ma\mathcal{M}_{a}. To define optimality, we turn to the theory of optimal experimental design, which states that for linear regression with a fixed design XX and subset M\mathcal{M}, the expected prediction error from using the subset XMX_{\mathcal{M}} is Tr(X⊤X(XM⊤XM)−1)\text{Tr}(X^{\top}X\left(X_{\mathcal{M}}^{\top}X_{\mathcal{M}}\right)^{-1}). This gives a straightforward objective achieving the V-optimality :

We followed the setup in Sec. 4.3 for validating our selection method, as this represents the most likely application scenario for our observational scaling laws by practitioners. Our objective is to replicate our scaling analysis (using a full set of 47 models) in Fig. 6(a) using a small subset of models selected by our method. In Fig. 7(a), we compute the geometric average of test MSEs on all prediction tasks (Naive, CoT, CoT + SC) as the evaluation metric for different selection methods. We find that our V-optimality selection method significantly outperforms random selection and quickly converges to the prediction performance of using the full set of models. In Fig. 7(b), we show that using only a small subset of 12 models selected by our method, the fitted scaling curves already effectively capture the scaling trends of different post-training methods.

To facilitate future scaling analyses for post-training techniques, we provide a reference list of models selected with our method under different budget constraints in table 1. These models were chosen from all available ones (see table B.1) with Llama-2 models always being included (as it is currently the most representative and widely used model family), and are expected to be representative of them. Notably, the selected models cover diverse capability ranges and dimensions to capture potential scaling dimensions. For example, under the 12 model budget constraint, the selected models cover both stronger models (Llama-3) and weaker ones (Falcon), as well as models with specialized programming capabilities (DeepSeek-Coder). Updating this list with other constraints (e.g., total inference FLOPs) or new model families is straightforward, and we provide both implementations and guidelines in our released code.

Discussion and Other Applications of Observational Scaling

Our work validates the hypothesis that there is a low-dimensional space of LM capabilities that captures their scaling behaviors and can be measured via a low-rank decomposition of existing LM benchmarks. While the majority of our work focuses on applications to scaling laws and predictions, we also find that the shared, low-dimensional capabilities could potentially be used as an evaluation metric and optimization target for LMs. We discuss some of these possibilities here.

Many existing benchmarks suffer from a limited dynamic range: they either saturate quickly for large models (e.g., HellaSwag, Winogrande) or have completely random performance for small models (e.g., MMLU, GSM8K), see Fig. C.4 for the behavior of each benchmark. In contrast, we find that PC-1 is a smooth capability measure that can be used to compare LMs across many orders of magnitude (at least 10 nats). This allows us to compare models from heterogeneous sources and of extremely different capabilities on a single, unified scale (Fig. 9). We believe that the high dynamic range of PC1 may make it suitable as an optimization target for pretraining, where architecture or data interventions can be benchmarked against PC-1 at small scales and validated at large scales.

Extending these ideas further, since PC-1 serves as a unified measure of capabilities, it may serve as a good way to compare compute efficiencies across many model families. In Fig. 9, we plot PC-1 against log-FLOPs and find that most models fall along a clear pattern in the training-compute to capabilities tradeoff curve. The Phi family is a clear outlier in compute efficiency, though this is likely because we are not accounting for the fact that Phi uses additional inference FLOPs to generate training data that is not shown in this figure.

Finally, we can analyze the interactions between post-training techniques and model families by projecting the fitted scaling curves in Fig. 6(a) to ff-equivalent FLOPs for different families ff using Eq. 8. We can then identify which model families benefit the most from these techniques and the point at which they start to benefit. Fig. 9 shows an example of comparing the predicted scaling of CoT across model families. We find that LMs benefit similarly from CoT, but that Phi is once again an outlier in its behavior: it benefits from CoT much earlier than other model families, but scales less rapidly. Similarly, models specifically trained on code (DeepSeek-Coder), also demonstrate an earlier transition but less rapid scaling compared to models trained with standard protocols. The distinct behavior of Phi/DeepSeek-Coder relative to other models indicates the importance of pretraining data in determining model scaling behaviors. While we did not specifically focus on these types of analysis in this work, we hope that our approach enables future works to gain further insights into differences between LM training recipes and their scaling behavior.

Conclusion

We have presented observational scaling laws – an approach that generalizes existing compute scaling laws to handle multiple model families using a shared, low-dimensional capability space. Using this approach, we show that we can build low-cost, high-resolution, and broad-coverage scaling laws that allow us to make accurate predictions for many complex scaling phenomena, such as emergent behaviors, agentic capabilities, and the value of post-training interventions. We provide concrete and practical prescriptions for researchers and practitioners to perform similar forms of scaling analyses for their own benchmarks and post-training methods in the hopes of encouraging more quantitative, scaling-law-based approaches to designing benchmarks and post-training methods.

We thank Zitong Yang for his assistance with an early experiment of the project. We also thank Jimmy Ba, Yann Dubois, Pavan Kapanipathi, Lisa Li, Karthik Narasimhan, Ethan Perez, Chenglei Si, Tristan Thrush, Zitong Yang, Shunyu Yao, and the Hashimoto Group for their helpful discussions or feedback on the paper draft. This project is not possible without the open-source contributions including HuggingFace, EleutherAI LM Eval Harness , Open LLM Leaderboard , EvalPlus , vLLM , LMSys Chatbot Arena Leaderboard , and AlpacaEval Leaderboard .

TH and YR were supported in part by gifts from the Tianqiao and Chrissy Chen Institute, Open Philanthropy, Amazon ARA, Meta, and IBM. Resources used in preparing this research were provided in part by the Province of Ontario, the Government of Canada through CIFAR, and companies sponsoring the Vector Institute. We acknowledge the support of the Natural Sciences and Engineering Research Council of Canada (NSERC), RGPIN-2021-03445.

References

Appendix A Algorithm

In algorithm 1, we include the detailed algorithm for fitting the observational scaling laws as described in Sec. 3.

Appendix B Experimental Details

We collected a broad set of representative open LMs covering 21 model families and a total of 77 models. These model families include Llama-2 , Llama , Llama-3 , Qwen1.5 , Qwen , Mistral , Mixtral , Yi , Gemma , Falcon , Phi , Pythia , BLOOM , GPT-Neo/J , OPT , MPT , XGLM , CodeLlama , StarCoder , StarCoder2 , DeepSeek-Coder . For each model, we collected their available metadata including the number of model parameters NN and the amount of pretraining tokens DD by analyzing papers and other public information. We then estimated the training FLOPs CC using the simple estimate of C≈6NDC\approx 6ND for each model. Note that for models that were continually pretrained on additional data such as CodeLlama, we used the sum of the pretraining tokens and the additional continual pretraining tokens to estimate DD. See table B.1 for the collected metadata of these models.

We collected a set of diverse benchmarks that assess various LMs’ capabilities, including MMLU , ARC-C , HellaSwag , Winogrande , GSM8K , TruthfulQA , and XWinogrande , HumanEval . For MMLU, ARC-C, HellaSwag, Winogrande, GSM8K, and TruthfulQA, we primarily sourced results from the Open LLM Leaderboardhttps://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard , with updates current as of May 6th, 2024. When there were missing benchmark results, we followed the standardized evaluation protocols of the Open LLM Leaderboard and used the LM Eval Harness library to evaluate the LMs. For XWinogrande, we used the LM Eval Harness library to evaluate the models with 5-shot examples. For HumanEval, we primarily used the EvalPlus library and followed their standardized protocols for evaluation, and sourced the results from the EvalPlus leaderboardhttps://evalplus.github.io/leaderboard.html when available. We used the ‘Base Tests’ results provided by EvalPlus for all the models. See table B.1 for all collected benchmark results.

B.1.2 Instruction-Tuned Models

We collected the set of instruction-tuned models that have been evaluated on the AgentBench and AgentBoard benchmarks. These include models like GPT , Claude , Llama-2-Chat , Codellama-Instruct , Mistral-Instruct , Vicuna , Deepseek-LLM-Chat , Lemur-Chat , OpenChat , WizardLM , Guanaco , Koala , Dolly-v2 , OpenAssistant . We followed the same procedure in Sec. B.1.1 to collect the metadata of open models, while for proprietary models these metadata were not publicly available. Note that we only counted the pretraining tokens (and the continual pretraining tokens when applicable) for DD and excluded the data for instruction-tuning or additional finetuning, as these are typically only a small fraction of the total data and are nuanced to estimate due to the complexities in data curation for instruction-tuning. See table B.2 for the collected metadata of these models.

For instruction-tuned models, we also included standard LM evaluations such as MMLU , ARC-C , HellaSwag , Winogrande , TruthfulQA , GSM8K , and HumanEval , and we followed the same protocols in Sec. B.1.1 for evaluating open models. For proprietary models like GPT and Claude, it is more nuanced to evaluate them with a unified protocol (e.g., due to the lack of access to likelihood scores), so we collected the official results from their respective papers and documentation for all standard benchmarks (except for HumanEval, which we were able to evaluate using the EvalPlus library). Additionally, we collected Elo scores from the Chatbot Arenahttps://huggingface.co/spaces/lmsys/chatbot-arena-leaderboard which assess instruction-following capabilities of these instruction-tuned models (as of February 2nd, 2024) for reference, we did not utilize this metric for our downstream predictions. See table B.2 for all collected benchmark results.

B.2 Downstream Evaluation

For all downstream tasks of pretrained base models included in Sec. 4.1 and Sec. 4.3, we used the LM Eval Harness library to evaluate all the models. For the “emergent” capability tasks in Sec. 4.1, we applied likelihood-based evaluation with 2-shot examples. For the post-training intervention tasks in Sec. 4.3, we used the same evaluation protocol as the original papers, as described in the main paper. For agentic capability tasks of instruction-tuned models in Sec. 4.2, we directly sourced the results from the AgentBench and AgentBoard leaderboards and scaled the metrics to $$.

B.3 PCA Analysis

The PCA imputation starts with a simple mean imputation for missing values in the data matrix, and then PCA is applied to transform the data into a lower-dimensional space where the missing values are imputed by the PCA reconstruction. The above procedure is repeated until the imputed values converge or reach a maximum of 1000 iterations. By default, we used the first principal component (PC-1) to impute the missing values, as we found it to be the most robust in our preliminary experiments. Notably, when there are train and test splits, we first applied the PCA imputation procedure on the training set and then applied the same transformation to the test set to prevent any train-test leakage.

When applying PCA to extracting the capability measures, we extracted the top K=3K=3 principal components from the model-capability matrix. By default, we mean-centered the data before applying PCA without additional scaling, since most evaluation metrics are already normalized into $$. Similar to PCA imputation, we only fitted the PCA on the training set and applied the same transformation to the test set to prevent any train-test leakage.

Appendix C Additional Results

In Fig. C.1, we conducted a PC analysis for instruction-tuned models (see the model list in table B.2) following exactly the same procedure as Fig. 2. We find that the extracted PC measures for instruction-tuned LMs follow similar patterns as pretrained models and exhibit an even more significant low-rank structure, with the top 3 PCs explaining about 98.6% of the variance in the benchmark performance.

C.2 Properties of PC measures

In Fig. 3, we showed that the top PC-1 linearly correlates with log-compute scale measures (log-training FLOPs) within each comparable model family. In Fig. C.2, we show that this linear correlation generally holds for lower-ranked PCs, specifically PC-2 and PC-3, though the correlation tends to decrease with lower-rank PCs compared to the top PC-1.

When fitting our observational scaling laws, we utilized the (hypothetical) linear relation between the aggregated PC measures Pm:=β∗⊤SmP_{m}:=\beta^{*\top}S_{m} and the log-compute measures log⁡(Cm)\log(C_{m}) within each model family to transform PmP_{m} into compute-equivalent scales (Eq. 8) . This linear correlation has been partially validated through the linear correlation of top PCs (Fig. 3 & Fig. C.2). Here we more directly validate this linearity by analyzing the aggregated PC measures PmP_{m} fitted on specific tasks. Specifically, in Fig. C.3, we visualize the fitted PmP_{m} on the “emergent” capability tasks (i.e., Fig. 4(b)) versus the compute measures log⁡(Cm)\log(C_{m}) within each comparable model family. We find that the aggregated PC measures generally exhibit a linear correlation with the log-compute measures within each family. Notably, the linear correlation is consistently significant for the Llama-2 family, which we have used as the default reference family for computing the equivalent scales in our experiments.

In Fig. 9, we have shown that PC-1 can serve as a smooth capability measure for LMs that provide meaningful readouts across many orders of scales (about 10 nats). In Fig. C.4, we show that using a single benchmark metric as LM capability measures amy suffer from a limited dynamic range. In particular, they may either saturate quickly for large models (e.g., HellaSwag, Winogrande) or provide random readouts for weak models (e.g., MMLU, GSM8K).

C.3 Robustness Checks

Recall that we defaulted to use 3 PC measures for all of our prediction tasks. Here we provide additional analysis on the impact of using different numbers of PCs on the prediction performance and validate the robustness of our choice. In particular, we compare the fitted curves and prediction performance of using 1-4 PCs on all our tasks. The results are in Fig. C.5, Fig. C.6, and Fig. C.7 for post-training analysis, “emergent” capability, and agentic capability tasks, respectively. Our results indicate that using more than 2 PCs leads to better prediction performance than using compute measures like FLOPs, and using 3 PCs consistently leads to the most robust predictions across all the tasks. These validate our choice of using 3 PCs as the default number of PCs and indicate the robustness of our results to the choice of the number of PCs.

The cutoff for selecting the holdout set could have a significant impact on the prediction performance of observational scaling laws, as it determines the size of the training set that could be crucial when the entire dataset is not large (as in our case). Here we analyze how the prediction performance changes with different holdout cutoffs for various predictive measures (PCs vs compute measures) and provide a quantitative comparison that characterizes their overall prediction performance under varying cutoffs.

Specifically, we conducted the analysis on the post-training analysis tasks in Sec. 4.3 and the “emergent” capability tasks in Sec. 4.1, where there are more data points (compared to the agentic capability tasks in Sec. 4.2) to provide a more robust analysis. For each task, we vary the FLOPs cutoff to control the ratio of the test set from 60% to 5% (linearly spaced), which consequently changes the difficulty of the prediction task from more difficult (less training data with weaker performance) to easier (more training data with stronger performance). We can then compare the test MSE of using different predictive measures under different cutoffs and quantify the overall prediction performance using the area under the error curve (AUE). For “emergent” capability tasks, we additionally include a variant of the cutoff strategy that holds out test data based on the accuracy on the task, which simulates a more challenging weak-to-strong prediction scenario and offers an extra robust analyses.

The results are depicted in Fig. C.8 and Fig. C.9. We observe that in most of our evaluated setups, using our PC measures (especially with 3 PCs) generally leads to an earlier transition to the low prediction error region and much lower AUE compared to using compute scales like training FLOPs and model size. This indicates that PC measures are more robust under different cutoffs and more sample-efficient for scaling analysis.

C.4 Emergent Capabilities

In Fig. C.10, we show the prediction performance of using model size for the “emergent” capabilities of LMs. We find that it leads to significantly worse forecasts compared to using training FLOPs and PC measures and poorly captures the “emergence” trend. This is probably because models from different families were trained with very different data sizes and quality and may use different architectures.

In Fig. 4, we applied a different FLOPs cutoff than the default one on arithmetic tasks to make the prediction tasks more challenging. Here, we present the results of using the default FLOPs cutoff on arithmetic tasks in Fig. C.11. We find that using the default FLOPs cutoff makes the prediction tasks trivial with too many data points close to perfect performance. Notably, using PC measures still outperforms using compute measures like model size and training FLOPs, indicating its robustness to the choice of the cutoff.

In Fig. C.12, we present the results on additional “emergent” capability tasks included in Wei et al. . Similar to the main tasks (Fig. 4), we used the default FLOPs cutoff for non-arithmetic tasks (IPA Transliterate) and a quarter of the default cutoff for arithmetic tasks (3-Digit Addition, 2-Digit Addition). We find that using PC measures consistently leads to the best prediction performance compared to using model size or training FLOPs. While the extrapolation does not exactly match the trend of the ground truth on the IPA Transliterate task, possibly due to the fact that the specific task capabilities are not well covered by our collected benchmark metrics, it still provides a reasonable forecast of the “emergence” behavior.

C.5 Post-Training Method Analysis

In Fig. C.13, we show the prediction performance of using different scale measures on various prediction tasks for the post-training method analysis on GSM8K. Similarly, using PC measures well captures the scaling trend and consistently leads to the best prediction performance across all tasks.

We further validated our observational scaling laws for predicting the impact of CoT on the BigBench-Hard tasks following the same setup in Sec. 4.3. In particular, we used the defaulted FLOPs cutoff and the same PC measures (# = 3). We normalized the prediction accuracy on each BBH task by their respective random prediction accuracy and aggregated the normalized accuracy across all tasks for predictions. The results are depicted in Fig. C.14. Surprisingly, we observe that using training FLOPs leads to reasonable predictions of LM performance with and without CoT on BBH tasks, possibly due to the denoising effect of aggregation over all tasks. Furthermore, using PC measures accurately captures the scaling trends in both setups, even when using training FLOPs leads to less tight captures in the “Naive” setup or fails to capture the behavior of models trained on synthetic data (Phi).

C.6 Model Subset Selection

In Fig. 7(a), we demonstrated how the prediction errors change with the number of models selected by our method. Here we present a qualitative analysis of the prediction results with different numbers of models selected in Fig. C.15. We find that with more than 8 models, the fitted scaling curves have already converged to accurately capture the scaling trend, indicating the efficiency of our method.

We present the prediction results with randomly selected models from all available models in Fig. C.16, in comparison to the results with models selected by our V-optimality criterion (Fig. C.15). All these results are produced with a fixed random seed. We find that using randomly selected models leads to a much worse prediction performance, even with 16 models, demonstrating the critical need to carefully select models for effective scaling analyses.

C.7 Fited Functional Forms for Preregistration of Predictions

In table C.1, we included the functional forms of fitted scaling laws included in our experiments. These functional forms serve as a preregistration of our predictions for future models, which will be used to test the generalizability of our scaling analysis to unseen models.