From Words to Numbers: Your Large Language Model Is Secretly A Capable Regressor When Given In-Context Examples

Robert Vacareanu, Vlad-Andrei Negru, Vasile Suciu, Mihai Surdeanu

Introduction

Large Language Models (LLMs) are capable of learning to perform a task when given examples of that task in their context, without any additional training. This surprising capability, called in-context learning (ICL) Brown et al. (2020), emerges just from next-token prediction for sufficiently large models.

We use regression tasks to analyze the in-context capabilities of already pre-trained large language models (LLMs), such as Llama2, GPT-4, or Claude 3. Garg et al. (2022) have previously explored the range of functions that transformers, when trained specifically for in-context learning, are capable of learning. However, contemporary LLMs emerge as capable in-context learners without being specifically trained for it. We extend previous work and analyze the extent to which LLMs, decoder-only transformers trained auto-regressively for next-token prediction, are capable of learning regression functions when given in-context exemplars, without any additional form of supervision or training.

Similar to previous work Garg et al. (2022), we use (synthetic) regression datasets. Synthetic regression datasets have the following advantages:

(i) Algorithmically generated: The data is guaranteed to be generated deterministically, by a well-determined (and logical) formula. This property makes them suitable to use when investigating whether a given model is capable of unraveling the underlying structure of the data.

(ii) Difficulty control: The user has direct access to the difficulty of the synthetic regression problem and can investigate the cases of simple linear regressions of the form y=ax+by=ax+b, to more difficult problems such as Friedman #2, a highly non-linear function used for benchmarking:

(iii) Data availability: Lastly, synthetic datasets present the advantage of allowing the user to generate novel data in large(r) quantities. Additionally, it ensures that models are less likely to have been previously exposed to these specific problems.

Our study shows that pre-trained large language models (LLMs) display a surprisingly good performance on various regression tasks. For example, in Figure 1, without any parameter update, Claude 3 approaches the performance of a Linear Regression model and largely outperforms other supervised methods such as Random Forest or Gradient Boosting on a randomly generated linear regression dataset, with one informative variable out of two.

Experimental Setup

We describe the models and the datasets we use in our experiments below.

We experiment with 33 types of datasets: (1) linear regression, (2) non-linear regression, and (3) regression datasets with non-numerical inputs. We describe each below.

When generating a dataset, we sample the input xx from N(0,1)\mathcal{N}(0,1). We sample the weight vector ww from Uniform(0,100)Uniform(0,100).We used sklearn. Please see make_regression for more details.

1.2 Non-Linear Regression Datasets

For non-linear regression problems, we use the three problems introduced by Friedman, called Friedman #1, Friedman #2, and Friedman #3 (Friedman, 1991; Breiman, 1996). For example, Friedman #1 is defined as follows:

To supplement the non-linear regression datasets and following Garg et al. (2022), we create datasets using randomly initialized neural networks. We explore the outputs of 2 types of neural networks: (1) a sequence of simple linear layers with ReLU non-linearity in-between, and (2) the output of a randomly initialized transformer encoder block.

1.3 Regression With Non-Numerical Inputs

2 Models

We use a total of 12 large language models (LLMs), both open and private. Specifically, we use Mistral7B, Mixtral8x7B, CodeLlama70B, Llama2 70B, Yi 34B, DBRX (weights available) and ChatGPT, GPT-4 (OpenAI) and Claude 3 Opus, Claude 3 Sonnet (Anthropic), Gemini Pro (Google), and Mistral Medium (Mistral) (weights not available). The models we use cover a wide range of parameters, from 7B or less (Mistral) to 132B or more (DBRX).Since the number of parameters for some models is not disclosed, it is possible that certain closed models may have fewer than 7B or more than 132B parameters. Unless otherwise specified, we interact with the models only through prompting and in-context exemplars. We use the same prompt for all models and do not do any prompt tuning. The prompt is of the form Feature 1: \nFeature 2: \nOutput: . In-context exemplars are separated with two new lines “\n\n”. For the test example, the model is asked to predict the number corresponding to the Output variable. We observed that some models tend to provide additional explanations, before outputting the final number. To prevent this behavior, we add an additional text in the beginning, instructing the LLM to only output the number. We give a complete example in Appendix C.1.1. Additionally, we analyze the explanations provided by the models in Appendix J, finding that there is sometimes a discrepancy between the rationale given for their predictions and the actual predicted values Unless otherwise specified, we use a temperature of .

We use a total of 10 traditional supervised models, available in most statistical learning packages. We use: Linear Regression (4 versions: no regularization, Ridge regularization, and Lasso Regularization, no regularization and with polynomial features), Multi-Layer Perceptron (6 versions, 3 versions with different widths (Hornik et al., 1989) and 3 versions with different depths), Random Forest, Bagging, Gradient Boosting, AdaBoost, SVM, KNN, Kernel Ridge, and Splines. Similar to the LLM case, we do not tune any hyperparameters and use the defaults available in sklearn. It is important to note that these supervised baselines are very strong: (1) many of them are the results of algorithms specifically designed for regression (e.g., Splines); (2) all perform parameter updates (unlike an LLM with ICL); and (3) the default hyperparameters, as set in widely-used statistical packages, have been refined over time to offer reliable and generally strong performance across a variety of scenarios.

In order to contextualize the performance of the LLMs and to evaluate their effectiveness relative to basic heuristics, we incorporated the following series of heuristic-based unsupervised baseline:

Average: Predicts the next value, yn+1y_{n+1}, as the mean of all preceding outcomes: yn+1=1n∑i=1nyiy_{n+1}=\frac{1}{n}\sum_{i=1}^{n}y_{i}.

Last: Uses the most recent tuple (xn,yn)(x_{n},y_{n}) for prediction, such that yn+1=yny_{n+1}=y_{n}.

Random: Predicts yn+1y_{n+1} by randomly selecting from the set of prior observations {y1,…,yn}\{y_{1},\dots,y_{n}\}. The final prediction is thus yn+1=sample([y1,…,yn])y_{n+1}=sample([y_{1},\dots,y_{n}])

Additional details on the models are provided in Appendix C. We include results with additional models, such as the latest release of GPT-4 (gpt-4-2024-04-09) or Mixtral Mixture of Experts 8x22B in the Appendix F, where we present the average rank obtain by each model.

Large Language Models Can Do Linear Regression

We present two bar plots in Figure 2, corresponding to two different datasets: (1) a dataset consisting of three variables, with a single informative variable (Regression NI 1/3), and (2) one dataset containing two random variables, where both variables are informative (Regression NI 2/2). For LLMs, we selected Claude 3 Opus (Claude 3), GPT-4, and Gemini Pro, as they are the flagship closed-source models currently available, and Mixtral8x7B (Mixtral), Llama2 70B (Llama 2), Yi 34B (Yi) and DBRX (DBRX) as the flagship open-weights models. Traditional supervised models in our analysis included Linear Regression (LR), Multi-Layer Perceptron (MLP), Random Forests (RF), and Gradient Boosting (GB). Additionally, we include a fifth supervised method, the one resulting in the best performance.If this method coincides with one previously selected, the subsequent best performer is chosen. We would like to remark that this is a very strong baseline, as it highlights the best performance obtainable with hindsight information. For the unsupervised baselines we included (i) Average, and (ii) Random Sampling. We draw the following observations from this experiment:

First, LLMs, when given in-context examples of input-output pairs, exhibit a (perhaps surprisingly) good overall performance. When compared with unsupervised baselines, the large language models always outperform them, indicating that the underlying mechanism at play is more sophisticated than such simple heuristics.

Second, we remark that LLMs in some cases outperform even supervised methods. For example, for the regression task with one informative variable out of a total of 3 (Regression NI 1/3), Claude 3 ranks 3 out of a total number of 31 models, only (slightly) behind Linear Regression and Linear Regression + Poly. For example, when averaging the mean absolute error across all runs, Claude 3 obtains 0.140.14, while Linear Regression obtains 0.120.12. It largely outperforms other supervised methods such as Random Forest or Gradient Boosting, even though it no gradient updates were performed, nor it was specifically designed for linear regression.Comparatively, Random Forest obtains 5.325.32, Gradient Boosting obtains 2.582.58, and GPT-4 obtains 2.262.26.

Lastly, we remark that this strong performance is not only specific to the current closed-source flagship models. For example, Mixtral outperforms supervised methods such as Random Forest or Gradient Boosting on the Regression NI 2/2 dataset.

Alongside the two bar plots, we include a heatmap in Figure 3 to show how each model ranks across different datasets. We show the datasets vertically and the models horizontally. For instance, Claude 3 Opus achieves the top rank (rank=1) on the NI 1/1 dataset. Notably, Claude 3 Opus and GPT-4 consistently perform better than methods such as AdaBoost, Gradient Boosting, KNN, or Random Forest. Out of the LLMs with open-weights, both Mixtral 8x7B and Yi 34B Chat outperform methods such as KNN or SVM on all four datasets.

Overall, these results reveal that large language models, whether closed-source (e.g., Claude 3, GPT-4) or open-weights (e.g., DBRX, Mixtral 8x7B), are capable of performing linear regression tasks using in-context exemplars composed of (x,y)(x,y) pairs, all without the necessity for gradient updates. While the performance across these models varies, it consistently outperforms that of unsupervised baselines, suggesting that the underlying mechanism at play is more sophisticated than these simple heuristics. Moreover, specific LLMs (e.g., Claude 3, GPT-4) consistently exceed the performance of strong supervised baselines such as Random Forests, Gradient Boosting, or KNN.

We present extended results, encompassing a wider array of models and datasets, in Appendix D.

Large Language Models Can Do Non-Linear Regression

We extend our previous analysis to non-linear regression problems.

We use the 3 synthetic regression benchmarks introduced by Friedman (1991). Below, we provide the definition of the Friedman #2 dataset, with complete definitions for all datasets available in Appendix B.

where ϵ\epsilon represents noise added to the system, modeled as a Gaussian distribution N(0,1)\mathcal{N}(0,1), and the variables x1x_{1}, x2x_{2}, x3x_{3}, and x4x_{4} are drawn from uniform distributions as follows: x1∼U(0,100),x2∼U(40π,560π),x3∼U(0,1), and x4∼U(1,11).x_{1}\sim\mathcal{U}(0,100),x_{2}\sim\mathcal{U}(40\pi,560\pi),x_{3}\sim\mathcal{U}(0,1),\text{ and }x_{4}\sim\mathcal{U}(1,11).

Our findings for the Friedman #1, #2, and #3 benchmarks are presented in Figure 4. The selection of methods follows to the same procedure used in Section 3: three leading closed-source LLMs, four leading open-weights LLMs, and five conventional supervised models–including the best performing model–and two unsupervised baselines. We remark that the strong performance of LLMs persists in the non-linear case as well. For example, Claude 3 outperforms all but the Linear Regression with Polynomial Features (LR + Poly) on Friedman #2.

2 New Regression Datasets

In an effort to mitigate the potential familiarity of models with pre-existing datasets encountered during their pre-training phase, we experiment with two new non-linear regression datasets which are unlikely to have been part of the pre-training phase. Our methodology is as follows. Our first novel dataset (called Original #1), plotted in Figure 5, is created to resemble a line with oscillations:

For the next dataset (called Original #2), we draw inspiration from the datasets introduced by Friedman, but we modify the domain of xx and change the operands (e.g., 2 →\rightarrow 4). We provide an example below:

It is important to underscore that the primary goal of these novel datasets is not to construct inherently difficult challenges for the LLMs, but rather to minimize the probability of evaluating them on datasets they could have already seen during their training phase. We provide additional details on these datasets in Appendix B, along with additional datasets. For an in-depth analysis of potential data contamination concerns, including additional experiments conducted to address these issues, refer to Appendix N.

3 Discussion

We summarize all our results in the form of a heatmap in Figure 6. For each dataset, we record the relative rank of each method with respect to all the others. For example, Claude 3 Opus performs the best on Original 1 (rank=1). We structure our results in 3 blocks: (1) LLMs (left), (2) Traditional Supervised Methods (middle), and (3) Unsupervised Methods (right). We make the following observations:

First, on Original 1 (see Figure 5), LLMs largely outperform traditional supervised methods. Remarkably, eight out of the ten highest-ranking methods in this context are LLMs. This strong performance on this dataset is exhibited by both private and open-weights models. For example, DBRX outperforms all traditional supervised methods, despite no gradient update.

Second, we remark that the LLMs show a strong performance on all datasets introduced by Friedman (Friedman #1, Friedman #2, Friedman #3) and on all datasets introduced by us (Original #1, Original #2).

Overall, our results show that LLMs with ICL are capable of performing non-linear regression. For example, Claude 3 Opus outperforms Gradient Boosting and KNN on all 5 datasets. We present extended results, encompassing a wider array of models and datasets, in Appendix E. We observed that LLMs struggle on the datasets generated with randomly initialized neural networks (e.g., Simple NN #1, Transformer #1), although they remain, generally, better than the unsupervised methods.

Due to space constraints, we included in Appendix K an analysis of the performance of LLMs on non-numerical regression datasets. We found that even in this regime, LLMs outperform the unsupervised baselines.

How Fast Do Large Language Models Adapt?

Following the surprising results that LLMs are capable of doing regression, when given the training data in their context in the form of in-context exemplars, we investigate next how their predictions improve with the number of examples given. Specifically, we empirically analyze whether the performance of the LLMs approaches that of the best possible fixed strategy over time.

Borrowing from the Online Learning community, we empirically analyze how the cumulative regret (i.e., cumulative loss) grows with respect to the time step (number of examples in context) Orabona (2019). Ideally, a good model should, over time, approach the quality of decisions that the best fixed strategy, informed by hindsight, would have made. In other words, the cumulative regret should, ideally, grow sub-linearly over time. To empirically estimate how the regret grows, we fit 3 curves: (1) Linear Fit: a∗x+ba*x+b, (2) Sqrt Fit: a∗sqrt(x)+ba*sqrt(x)+b and (3) Log fit: a∗log(x)+ba*log(x)+b.The choice of linear, square root, and logarithmic fits is motivated by their common appearance in theoretical regret bounds within the online learning literature. We then use the R2R^{2} coefficient to determine which curve fit is better. We show two qualitative plots in Figure 7. We summarize the results in Table 1, recording the curve fit with the highest R2R^{2} coefficient for each model. Since simply picking the best curve fit according to the R2R^{2} score might tell an incomplete story, we include additional plots in Appendix G, covering multiple models and all seven datasets. We draw the following observations. First, the performance of large language models improves with the number of examples, suggesting the mechanism at play is capable of effectively leveraging more data. Second, we remark that very capable LLMs, such as Claude 3 or GPT-4 can obtain sub-linear regret, meaning that the predictions made by the LLM approach the quality of decisions that the best algorithm would have made, leading to near-optimal performance in the long run.

We remark that there are differences between our empirical analysis and online learning. Firstly, while online learning often focuses on establishing theoretical regret bounds, our approach is empirical, we only empirically show that the regret of certain LLMs grow sub-linearly by using curve fitting and R2R^{2}. To address potential concerns of overfitting and enhance the robustness of our empirical findings, we repeated the experiment 3 times and averaged the cumulative regret. Second, our results are only for finite (and relatively small) time steps, diverging from the online learning norm of analyzing behavior as TT approaches infinity. To provide further evidence that the results are not an artifact of small T, we performed the following experiment. We used GPT-4 and recorded its performance across multiple training dataset sizes, ranging from 2020 to 500500. We have observed that the performance of GPT-4 continues to improve as the number of in-context exemplars increases, suggesting that, our results are not an artifact of limited time steps. We include the associated plots in Appendix O.

Following the empirical evidence that LLMs are very capable regressors, despite not being trained for it, we hypothesize that (very capable) LLMs emerge from their training as very good online meta-learners (Finn et al., 2019; Mirchandani et al., 2023).

Related Work

The in-context learning capability of large language models has garnered significant attention Brown et al. (2020). How this capability emerges during a standard next-token prediction pretraining and how it operates is still up for debate. A substantial body of research is dedicated to exploring the parallels between in-context learning mechanisms and traditional algorithms like gradient descent (Akyürek et al., 2023; von Oswald et al., 2022; Dai et al., 2023; Ahn et al., 2023; Cheng et al., 2023; Mahankali et al., 2024; Vladymyrov et al., 2024). For example, Akyürek et al. (2023) and von Oswald et al. (2022) prove that transformers could theoretically implement gradient descent. Bai et al. (2023) shows that the transformer architecture can implement more complex in-context learning procedures, involving algorithm selection. Cheng et al. (2023) argue that non-linear transformers learn to implement gradient descent in function spaces. von Oswald et al. (2023) suggests that performance of transformer-based models may be due to an architectural bias towards mesa-optimizaiton. Nonetheless, the extent to which pre-trained transformers actually implement gradient descent when given in-context examples remains a topic of debate (Natan et al., 2023; Shen et al., 2023).

Other lines of work investigate the convergence of in-context learning (Wies et al., 2023; Huang et al., 2023). Li et al. (2024) analyzes the training dynamics of transformers with nonlinear attention and nonlinear MLP, expanding upon previous work which considered simpler transformer-based architectures (Huang et al., 2023; Tian et al., 2023). However, for natural language tasks such as sentiment analysis, it is unclear how much learning occurs with in-context examples (Min et al., 2022; Pan et al., 2023; Kossen et al., 2024). For example, Min et al. (2022) shows that GPT-3 retains a strong performance even when the labels of the in-context exemplars are random. On the other hand, recent work (Hendel et al., 2023; Liu et al., 2023) investigated how in-context learning creates task vectors, which can then be applied to produce the output.

Another question investigated in recent work is where does the in-context learning (ICL) emerges from (Chan et al., 2022; Xie et al., 2022; Han et al., 2023). For example, Chan et al. (2022) shows that in-context learning appears when the training data has particular properties. Xie et al. (2022) analyzes in-context learning through a small scale synthetic dataset (GINC). Han et al. (2023) identified a subset of the pre-training data that supports in-context learning, showing how continuing pretraining on this subset increases the model’s ICL abilities.

Another line of research, which is close to our work, is that of investigating what types of “functions” can be learned through in-context learning Garg et al. (2022); Zhang et al. (2023); Xing et al. (2024). Notably, all these works do not use pre-trained LLMs, but specifically train a transformer for the task. Garg et al. (2022) shows empirically that standard transformers can be trained from scratch to perform in-context learning of linear functions. Guo et al. (2024) investigates more complex function classes. Wei et al. (2023) shows that larger language models are able to overcome their semantic priors when shown input-label mappings. Zhang et al. (2023) train transformers with a single linear self-attention layer to in-context learn linear regression tasks, showing that transformers are capable of obtaining a performance competitive with the best linear predictor. Bhattamishra et al. (2024) experiment with training various models to in-context learn boolean functions. Although not the main focus of their work, they also experiment with pre-trained models such as Llama 2 and GPT-4, showing that they obtain a performance similar to nearest-neighbor baselines for boolean functions.

Different from previous work, we investigate how pre-trained models, such as GPT-4 or Claude 3, without any gradient updates, can learn various linear and non-linear function classes when given examples in-context and thoroughly compare them against multiple traditional supervised methods (Ruppert, 2004) such as Gradient Boosting (Schapire, 1989; Friedman, 2001) or Random Forests (Breiman, 2001).

Conclusion

In this paper, we examined the extent to which large language models such as Claude 3, GPT-4, or DBRX are capable of performing the task of regression, when given input-output pairs as in-context examples, without any gradient updates.

We showed that large language models are capable of doing both linear and non-linear regression, with performance rivaling that of supervised methods such as Linear Regression or Gradient Boosting. We then analyzed how their performance approaches that of the best possible fixed strategy as the number of in-context examples grows, showing how very capable models such as Claude 3 Opus or GPT-4 are capable of approaching the quality of decisions that the best algorithm in hindsight would have made. Our results demonstrate that large language models are capable of doing regression when given in-context examples of (input, output) pairs, despite not being explicitly trained to do so. We leave the exploration of augmenting LLMs’ training with synthetic regression and math datasets, during either pre-training or fine-tuning, to future work. We release our code and results at https://github.com/robertvacareanu/llm4regression.

Acknowledgments

This work was partially supported by the Defense Advanced Research Projects Agency (DARPA) under the ASKEM and Habitus programs. Mihai Surdeanu declares a financial interest in lum.ai. This interest has been properly disclosed to the University of Arizona Institutional Review Committee and is managed in accordance with its conflict of interest policies.

Ethics Statement

In this work we explored the extent to which large language models (LLMs) are able to perform regression tasks. We did not perform any additional training. We do not envision any negative impact of our results.

Limitation

This study focuses primarily on regression tasks, including an exploration into regression-like scenarios where inputs are symbolic rather than numeric, yet the outputs remain numeric.

A second limitation is the reliance on several large language models, including proprietary ones whose performance may change over time, potentially affecting reproducibility. To address this, we also included leading open-weight models in our analysis, though we note that their performance is generally behind that of private models. Additionally, we release our intermediate results.

Third, the issue of data contamination poses a challenge, given the opaque nature of training datasets for many LLMs. We have taken several steps to mitigate this risk: (i) Our analysis spans multiple LLMs, reducing the likelihood of all models being contaminated in the same way; (ii) We evaluated models with multiple random seeds on newly introduced datasets (alongside known ones like Friedman #1). In this way, we diminish the chance that these models have been directly exposed to the exact datasets during training; (iii) We included results with Falcon 40B, whose training data is publicly available (please see Appendix N for more details). We acknowledge that these measures do not eliminate the potential for data contamination entirely.

Fourth, while we showed empirical evidence that large language models are capable to perform regression tasks, we did not provide theoretical explanations to support these observations.

References

Appendix A Appendix Structure

In Appendix B we provide additional details of the datasets we used.

In Appendix C we provide additional details of the models we used.

In Appendix D, we provide additional experiments and results to complement Section 3: Large Language Models Can Do Linear Regression.

In Appendix E, we provide additional experiments and results to complement Section 4: Large Language Models Can Do Non-Linear Regression.

In Appendix F, we show the average ranks obtain by each model across different dataset types.

In Appendix G, we provide additional experiments and results to complement Section 5: How Fast Do Large Language Models Adapt?

In Appendix J we detail how LLMs provided justifications for their prediction.

In Appendix K we include another experiment: regression task with non-numerical inputs.

In Appendix L we analyze the effects of rounding.

In Appendix M we analyze whether the performance of LLMs is similar with KNNs or not.

In Appendix N we analyze whether the results we have seen could be the effect of data contamination.

In Appendix O we analyze whether the performance of LLMs plateaus after a given number of in-context examples or not.

In Appendix P we analyze the performance of LLMs whose backbone architecture is different from Transformers.

Appendix B Datasets

We provide the formulas for all datasets used below. We set the noise to for all datasets.

In order to generate the linear regression datasets, we use the function make_regression, available in sklearn (Pedregosa et al., 2011).

B.2 Friedman # 1

Where x0,x1,x2,x3,x4∼U(0,1)x_{0},x_{1},x_{2},x_{3},x_{4}\sim U(0,1)

B.3 Friedman # 2

B.4 Friedman # 3

B.5 Original # 1

B.6 Original # 2

B.7 Original # 3

B.8 Original # 4

B.9 Original # 5

B.10 Neural Network Induced

For the random datasets induced by neural networks, we randomly initialize a neural network and create a dataset by feeding random data to it. The dataset Simple Random NN 1 was created by using a neural network with one hidden layer with ReLU non-linearities. The dataset Transformer 1 was created by a randomly initialized Transformer encoder block.

B.11 Non-Numerical Regression

We provide the code to generate the non-numerical regression datasets in the Listing 1. Essentially, we assign a random number ( to 2626) to each lowercase letter. Then we sample a weight vector. The expected output is generated by doing a dot product between the underlying assigned value of each character and the generated weight vector.

Appendix C Models

In the following, we provide additional details of the models we used for our main experiments and how we used them. We used three different types of models, as follows: (a) Large Language Models, (b) Traditional Supervised Methods, and (c) Heuristic-Based Unsupervised Methods. We describe them bellow.

This section outlines the 12 Large Language Models (LLMs) featured in our main experiments, which include a mix of open-weights and private models. We also include additional models, such as the newest GPT-4 version (gpt-4-20240409), multiple Claude variants, and the most powerful model released by Cohere, Cohere Command R Plus. We tried with Cohere Command R and Cohere Command and observed their performance to be lower, albeit still the unsupervised baselines (except for Cohere Command).

In Table 2, we categorize the models by their names, availability of weights, and developers, dividing them into two distinct sections. The first section lists the models featured in the main paper’s experiments (referenced in Sections 3, 4, and 5). The second section introduces additional models that were utilized for the extended analysis included in the Appendix.

We list in Table 3 the models we used through OpenAI, together with their corresponding model code.https://openai.com/

We list in Table 4 the models we used through OpenRouter, together with their corresponding model code.https://openrouter.ai

We list in Table 5 the models we used through DeepInfra, together with their corresponding model code.https://deepinfra.com

We list in Table 6 the models we used through Fireworks, together with their corresponding model code.https://fireworks.ai/

We show the prompt we used in Figure 9. Importantly, we used the same prompt for all large language models. We did not tune the prompt.

We encountered cases where the large language model would not produce a valid output. For instance, some models would occasionally output an empty string (i.e., “”). In the following, we detail the way we handled them across the experiments we showed in this paper.

For the experiments where we investigated how the performance of the models scale with the number of examples, we average 3 random runs for each dataset size. In this set of experiments, only Llama 70B generated invalid outputs a total of 3 times, for Original #2 and Friedman #2. We skip the random runs with invalid generations.

C.2 Traditional Supervised Models

We use a total of 11 traditional supervised methods, resulting in over 20 different configurations. Specifically, we used the following models:

Linear Regression: We used 4 variants of Linear Regression: (i) standard linear regression (Linear Regression), (ii) ridge (Ridge), (iii) lasso (Lasso), and (iv) Linear Regression with Polynomial Features (Linear Regression + Poly), where we used polynomial features of degree 2

Multi-Layer Perceptron: We used 6 variants of multi-layer preceptrons: 3 with different widths (MLP Wide 1, MLP Wide 2, MLP Wide 3) and 3 with different depths (MLP Deep 1, MLP Deep 2, MLP Deep 3).

SVM: We used both a single SVM and an SVM paired with a Scaler (SVM + Scaler)

KNN: We used multiple variants of KNN, where we vary the number of neighbors, the type of distance used, and the power parameter for the Minkowski metric; We distinguish between them with a v{index}.

We used the sklearn implementation for each model.We used sklearn 1.4.1.post1. Similar to the LLM case, we do not tune any hyperparameters. We use the default hyperparameters available in sklearn. We remark that these supervised baselines are very strong, as (1) many of them are the results of algorithms specifically designed for regression (e.g., Splines), (2) all perform parameter updates, and (3) the default hyperparameters, as set in widely-used statistical packages, have been refined over time to offer a reliable and generally strong performance across a variety of scenarios.

C.3 Unsupervised Models

We use three heuristic inspired unsupervised models:

Average: Predicts the next value, yn+1y_{n+1}, as the mean of all preceding outcomes: yn+1=1n∑i=1nyiy_{n+1}=\frac{1}{n}\sum_{i=1}^{n}y_{i}.

Last: Uses the most recent observation tuple (xn,yn)(x_{n},y_{n}) for prediction, such that yn+1=yny_{n+1}=y_{n}.

Random: Predicts yn+1y_{n+1} by randomly selecting from the set of prior observations {y1,…,yn}\{y_{1},\dots,y_{n}\}. The final prediction is thus yn+1=sample([y1,…,yn])y_{n+1}=sample([y_{1},\dots,y_{n}])

The goal of these unsupervised models is to better put the performance obtained by LLMs into perspective.

Appendix D Large Language Models Can Do Linear Regression (Expanded)

We expand the barplots shown in Figure 2 with more models and more datasets. In particular, we show in Figures 10, 11, 12, 13, 14, 15 the performance of the models on six datasets for linear regression. Specifically, we used the following datasets:

Regression 1/1, a linear regression with only 1 variable, which is informative

Regression 1/2, a linear regression with 2 variables, and only 1 informative variable

Regression 1/3, a linear regression with 3 variables, and only 1 informative variable

Regression 2/2, a linear regression with 2 variables, both which are informative

Regression 2/3, a linear regression with 3 variables, and only 2 informative variables

Regression 3/3, a linear regression with 3 variables, all which are informative

We included the corresponding rank heatmap in Figure 16.

We make the following observations. First, Claude 3 Opus performs among the best for the linear regression case where there is only one informative variable (Regression 1/1, Regression 1/2, Regression 1/3), ranking among top 3 best performing models. The performance drops when there are more informative variables.

Second, we remark that all large language models perform better than all the unsupervised methods on all datasets.

Third, we remark that the large language models display a good overall performance. For example, there are 4 LLMs (i.e., Claude 3 Opus, Claude 3 Sonnet, GPT-4, DBRX) which perform better than all 3 variants of KNN over all the datasets used.

Fourth, there are specific LLMs which always perform better than Gradient Boosting, such as Claude 3 Opus and GPT-4. We remark that DBRX and Code Llama 70B outperform Gradient Boosting in 4 out of 6 datasets.

Appendix E Large Language Models Can Do Non-Linear Regression (Expanded)

We expand the barplots shown in Figure 4 with more models and more datasets. In particular, we show in Figures 17, 18, 19, 20, 21, 22, 23, and 24 the performance of the models on the eight datasets for non-linear regression.

Additionally, we include in Figures 25 and 26 the performance of the models on datasets generated by randomly initialized neural networks, similar to the methodology of Garg et al. (2022). We can see that the performance of the LLMs decreases on the datasets created using the neural networks, although it (generally) remains above that of the unsupervised baselines.

Similar to Figure 6, we include the corresponding rankings in Figure 27.

We would like to remark that Claude 3 Opus obtains an average rank of 7.77.7, the best out of all models we have investigated. The next best is Gradient Boosting, with 8.18.1. The next best LLM is Claude 3 Sonnet, with 9.39.3, then GPT-4 with 12.712.7.

Appendix F Average Model Ranks

To provide a comprehensive overview of model performance across a diverse array of datasets, this section aggregates the average ranks obtained by each model. In this section we show the average ranks obtained by each model across: (1) all linear regression datasets (Linear), (2) all original benchmarking datasets introduced by us (Original), (3) all benchmarking datasets introduced by Friedman (Friedman), (4) all neural network induced datasets (NN), (5) all non-linear datasets (Non-Linear), (6) all datasets (Overall).

We show our results in Table 7. We divide the table into three blocks, separated by horizontal lines, corresponding to the results for (1) LLMs (e.g., GPT-4), (2) Traditional Supervised Methods (e.g., Gradient Boosting), and (3) Unsupervised Baselines (e.g., Average). We remark that LLMs obtain, overall, a strong performance. For example, Claude 3 Opus ranks second overall, outperforming methods such as Gradient Boosting, KNN, or multi-layer perceptrons. Overall, this strong performance is present in both private and open models. For example, DBRX and Mixtral 8x22B achieve average ranks that surpass conventional methods including AdaBoost, KNN, or Random Forests.

Appendix G How Fast Do Large Language Models Adapt? (Expanded)

We expand Table 1 to include more models. We show the corresponding results in Table 8.

Additionally, we include curve fit plots for models. To keep the number of plots to a manageable amount, we selected a subset of the models as follows. We selected Claude 3 Opus and GPT-4, as they are the flagship closed-source models. We selected Yi 34B Chat for the open-weights model. Lastly, we selected Gradient Boosting, Linear Regression, and Linear Regression + Poly. We present the corresponding plots in Figure 28 and 29. Since in these experiments we vary the number of in-context examples starting from 1, we could not include variants of KNN that uses more than one neigbhors for their prediction. We included KNN v4 which uses only one neighbor. Additionally, we included KNN v5, where we use a small number of neighbors when the amount of data is small, then gradually increase it. We can see that their performance is much worse than that of the LLMs, suggesting that the LLMs are doing something more than what KNN do.

Appendix H Claude Performance

Following the (perhaps surprisingly) strong performance of the Claude family of large language models on various regression tasks (e.g., Figure 27, when averaging the ranks of each model over each dataset, Claude 3 Opus performs the best), we provide results with additional models from the Claude family, namely: Claude 1.2, Claude 2.0, Claude 2.1, and Claude 3 Haiku. We include a rank heatmap for all the models from the Claude family currently available in Figure 30. For comparison, we also included the corresponding performance of two strong models with open-weights: DBRX and Mixtral 8x7B. Claude 2.0 and Claude 2.1 were sometimes generating invalid outputs (e.g., “I apologize, upon reflection I do not feel comfortable providing output values without context. Could we have a constructive discussion about the meaning and implications of this exercise?”). Therefore, we omit those problematic configurations. For all the other cases, the average performance is the result of at least 20 runs.These invalid outputs are specific to Claude 2.0 and Claude 2.1. For example, for Claude 3 Opus, there exist only 2 instances where it does not generate a valid output, out of a total of over 1000 runs. We note that the performance of Claude 3 models is much better than that of older models.

Appendix I Costs

We estimate the total cost for all our experiments to under 1200.Wespentapproximately1200. We spent approximately300 on OpenRouter. We spent approximately $600 on OpenAI. The cost for OpenAI is higher because we used it in our preliminary experiments. Additionally, the preliminary experiments used gpt-4, which is more expensive than gpt-4-0125-preview. We switched to gpt-4-0125-preview after we added the additional text to the prompt, instructing the model to only output their best estimate.We did not need this additional instruction in our initial experiments. We added it when we expanded the number of LLMs used, as some of them (e.g., Claude 3) would provide justifications before giving the final output. All models use the same prompt.

Appendix J LLMs Justifying Their Prediction

Without adding the prefix instruction text (see prompt in Appendix C.1.1), which instructed the models to output only its best estimate, some LLMs (e.g., Claude 3 Opus) started to provide explanations, in an attempt to justify their prediction.We hypothesize that this is because their system prompt instructs them to provide explanations. Analyzing these “explanations” revealed that there is a discrepancy between their explanation and their prediction. For example, for a non-linear regression problem, Claude 3 Opus suggests to train a Linear Regression model and gives the code to do so. Then it elaborates on how to use it for inference and gives the output. However, manually running the code suggested by the model results in a very different output. We provide some examples in Figures 31, 32, , 33, and 34. We describe each one below.

In Figure 31, Claude 3 correctly identifies that the output is generated by multiplying the input with a constant. Then, it calculates the constant and gives the final output.

In Figure 32, we can see that Claude 3 first calculates the mean of each feature. Manually inspecting the true values, they are: Mean Feature 0: 51.318051.3180, Mean Feature 1: 846.6326846.6326, Mean Feature 2: 0.47640.4764, Mean Feature 3: 5.24385.2438. We remark that the values given by Claude 3 are surprisingly close. Then, Mean Output: 374.04374.04. Then, Claude 3 calculate the covariance between each feature and the output. These estimates are much worse. For example, Cov(Feature 0, Output) is actually 2516.36832516.3683, not 8958.84698958.8469. Next, Claude 3 gives the variance of each feature. These values are close. For example, Var(Feature 0) is 729.420751729.420751 and Claude 3 generates 729.9052729.9052. Then, the model calculates the coefficients. These calculations are close to their true value, except for b0. Lastly, the model gives the final formula to calculate the output, which is (according to the model): 227.4744+25.9730∗42.54+1.2777∗851.93+1648.5958∗0.51+129.2408∗6.26=436.5981227.4744+25.9730*42.54+1.2777*851.93+1648.5958*0.51+129.2408*6.26=436.5981. However, this calculation is wrong. The output of that equation is 4070.704070.70. However, we would like to remark that what the model generated (wrongly, from a mathematical point of view), is actually much closer to the true value of 434.54434.54. In other words, the explanation offered by the model was not faithful.

In Figure 33, we can see that Claude 3 suggests that there is a strong linear relationship (please refer to Figure 5 for a plot of the data). Then, Claude 3 fits a linear regression y=mx+by=mx+b and gives the approximate values: m=0.9102m=0.9102 and b=12.1615b=12.1615. However, manually fitting a linear regression model on the corresponding data yields the following values m=0.97m=0.97 and b=2.94b=2.94. Then, Claude 3 calculates the final output. The calculation is correct, however it is far off from the true value of: 30.8630.86. We would like to remark that instructing the model to give its best estimate without any additional information gives a much better prediction: 30.9130.91.

In Figure 34, the solution generated by Claude 3 involves calculating the nearest neighbor. The problem in this approach is that the dataset given in-context contain only 50 examples, while the solution generated Claude involves taking the examples 50 and 54, which are non-existent.

All in all, we remark that the explanations provided by the model are not always faithful. We also remark that the predictions of the model in two cases: (i) when it outputs an explanation and (ii) when it does not output an explanation can vary.

Appendix K Beyond Numerical Regression

Our investigation has centered on conventional regression tasks characterized by inputs and outputs represented as numerical values. However, the performance on these tasks might be influenced by the quality of numerical token embeddings, which can serve as a confounding factor Razeghi et al. (2022). To address this and broaden our analysis, we shift our focus to datasets generated following the methodology outlined in Section 2.1.3. This allows us to evaluate the models’ capabilities in contexts where inputs are symbolic rather than numerical.

Appendix L Effects of Rounding

Due to computational budgets and context limits,For example, Llama2 context size is only 4096. we rounded both the input and the output to two decimals. To validate that our conclusions are not an artifact of the rounding mechanism, we re-ran GPT-4 on Friedman #2 and Friedman #3, rounding to five decimals. We selected GPT-4 because it obtained overall strong results and offers a good latency. We selected Friedman #2 and Friedman #3 because LLMs obtained generally good performance. We also ran the traditional supervised methods. We include the updated results in Figure 35. Comparing it with Figure 27, we can see that the performance of GPT-4 remains strong. This time it even outperforms Gradient Boosting on Friedman #3 and ranks first. The performance on Friedman #3 is only under Linear Regression + Poly, similar to Figure 27. All in all, the strong performance we observed is unlikely to be just an artifact of rounding.

Appendix M Is It Just Better KNN?

We can see from Appendix D and E that the performance of certain LLMs is almost always better than that of all three variants of KNN presented. For example, in the case of linear regression, Claude 3 Opus, Claude 3 Sonet and GPT-4 always perform better than the three variants of KNN we used. Similar, in the case of non-linear regression, only for Simple NN 1 and Transformer 1 is any of the KNN variants outperforming Claude 3 Opus. Moreover, from Table 8, which records the best curve fit over the cumulative regret, we can see that both variants of KNN perform worse than Claude 3 Opus or GPT-4.

To further investigate the extent to which what LLMs are internally implementing to do regression is a variant of KNN, we compare the performance of the LLMs against a total of 70 KNNs in a setting similar to that from Section 3 and 4: we randomly sample a dataset of size 50, which is given to both LLMs and KNNs and we ask them to predict the output corresponding to a testing data point. We repeat each experiment 100 times, with different random seeds. We describe what KNNs we considered in Listing 2.

We summarize our results in Figure 36. To keep the plot comprehensible, we chose the best-performing KNN configuration for each dataset. However, the ranking considers the performance of all models. We draw the following conclusions.

First, we remark that except on Friedman #1, for every other dataset, the top 9 best performing models are all LLMs. In other words, both closed-source models (e.g., Claude 3 Opus, GPT-4) and open-weights models (e.g., DBRX, Mixtral) outperform all KNN models on all the datasets except Friedman #1. This suggests that the mechanism implemented by in-context learning might be something more complex than KNN.

Last, we remark that for Friedman #1, only Claude 3 Opus and Claude 3 Sonnet outperform the KNN variants. Moreover, Claude 3 Opus and Claude 3 Sonnet outperform all KNN variants we experimented with on all datasets.

Appendix N Could It Be Just Contamination?

The very large datasets that are used to train contemporary large language models (LLMs) raise concerns about potential (accidental or not) contamination Sainz et al. (2023); Golchin & Surdeanu (2024). In our study, we have attempted to mitigate this as follows. First, we used many different random seeds. However, this does not nullify the risk that the LLM has seen similar data (e.g., data from Friedman #1, but with other random seeds). To mitigate this, we explored the performance of the models on regression functions of our own creation. This makes it unlikely that the model has seen data coming from the exact same function. For example, across all our newly introduced datasets (e.g., Original #1, Original #2, Original #3, Original #4, Original #5), Claude 3 Opus obtains the highest average rank of 6.46.4. Second place is Clade 3 Sonnet, with an average rank of 6.86.8, then Gradient Boosting with 8.48.4. Furthermore, our empirical evidence of consistent high performance across a diverse array of LLMs.

To further analyze the data contamination issue, we perform two additional experiments. We provide results with Falcon, an LLM whose training data is publicly available. Second, we perform an experiment similar to the approach proposed in Golchin & Surdeanu (2024), where we compare the performance of LLMs with and without knowing the dataset where the data comes from.

In this section we expand our analysis to include Falcon 40B and Falcon 40B Instruct, comparing their performance with both traditional statistical methods and other LLMs. To keep the figures comprehensible, we added only the following LLMs: Claude 3 Opus, Chat GPT, Mixtral 8x7B, and Mistral 7B.

We remark that the Falcon LLM team has released their training data,https://huggingface.co/datasets/tiiuae/falcon-refinedweb offering further insights into how the training environments of contemporary LLMs can result into LLMs being capable of regression.

Due to the context size limitations of Falcon,The context size of Falcon is 2048. we only evaluated it on the linear regression datasets and on the Original #1 dataset. The other datasets have a larger number of input variables (e.g., Friedman #2 has 5 input variables) and we could not fit 5050 in-context examples. We show our results on four datasets, Regression NI 1/1, Regression NI 1/2, Regression NI 2/2, Original #1 in Figures 37, 38, 39 and 40. Additionally, we include the corresponding rank heatmap in Figure 41.

We make the following observations. First, Falcon 40B outperforms our unsupervised baselines. Second, Falcon 40B outperforms Gradient Boosting and Random Forests on Regression NI 1/1.

Overall, Falcon 40B displays, as well, the capability of doing regression when given in-context examples, albeit to a smaller degree than when compared to more powerful (and newer) models.

N.2 Performance When Knowing The Dataset Name

To further investigate potential data contamination, we conducted an experiment inspired by the methodology described in Golchin & Surdeanu (2024). This involves assessing model performance under two conditions: with and without explicit knowledge of the dataset being evaluated. Specifically, we modify the prompt shown in Figure 9 to mention the name of the dataset (e.g., Friedman #1, Friedman #2, Friedman #3), as detailed in Figure 42. This approach allows us to discern the impact of dataset awareness on the model’s predictive accuracy, providing insights into the extent of potential contamination in the training data.

We present the comparative results in Table 9, which shows the average absolute error under conditions of dataset awareness versus unawareness. Notably, the mean absolute errors (↓\downarrow) remain closely matched across all scenarios. To statistically substantiate these observations, we performed paired t-tests for each dataset comparison. Given the multiplicity of tests performed, it became imperative to apply an adjustment for multiple comparisons to our p-values Dunn (1961); Benjamini & Hochberg (1995). Following this adjustment, none of the p-values remained below the (typically used) 0.050.05 threshold, suggesting that the knowledge of the dataset name does not significantly affect model performance. Prior to adjustment, in two cases the resulting p-value was under 0.050.05: GPT-4 on Friedman #3 (p-value 0.0450.045) and Claude 3 Sonnet on Friedman #3 (p-value 0.0180.018). Note, however, that only in the case of GPT-4 was the performance corresponding to the Dataset Aware setting better. This analysis indicates that, within the bounds of statistical significance, there is no substantial evidence to suggest that the performance of the models is influenced by explicit knowledge of the dataset name, something which has been linked to contamination (Golchin & Surdeanu, 2024).We investigated why the performance of Claude 3 Sonnet degraded on Friedman #2 when given the dataset name. We found that it is (mostly) because of an outlier: the absolute difference between the model’s prediction and expected output is >300>300.

Appendix O Does The Performance Plateau?

In the following, we investigate whether the performance of LLMs continues to improve as the number of in-context exemplars increases. Because this experiment is dependent on the context size of the LLM, this imposes additional constraints on which LLMs we can use. To this end, we experimented with the following dataset sizes: {20, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 400, 500}.

We experiment with ChatGPT and GPT-4. For GPT-4, we found that the performance keeps improving with the number of in-context exemplars at least until 500. We present our results in Figures 43, 44, 45, 46, 47. We repeated each experiment 20 times for each dataset size. We report the mean and 95% confidence.

Aggregating the results, we observed that GPT-4 performed better than Random Forest in 92% of the cases, better than Gradient Boosting in 51% of the cases, and better than Linear Regression with Polynomial Features (Linear Regression + Poly) in 40% of the cases, across all 5 datasets and all dataset sizes.

For example, from Figure 44 we can observe that while Linear Regression + Poly performs much better than GPT-4 in small data regimes, this performance gap decreases as the number of examples increases, suggesting that the model is indeed capable of leveraging a larger number of examples.

Appendix P Beyond Transformer-Based LLMs

In our study, we initially focused on transformer-based large language models (LLMs). To broaden our scope, we explore the capabilities of non-transformer LLMs, including a RWKV-based 14B LLM (Peng et al., 2023) and with StripedHyena Poli et al. (2023a; b), a 7B LLM. The performance rankings for these models, along with the transformer-based Mistral 7B for comparison, are illustrated in the heatmap provided in Figure 48. We tried running Falcon 7B as well, but it produced invalid outputs for almost all examples and all datasets, therefore we skip it. StripedHyena also encountered difficulties, producing invalid outputs in certain scenarios, such as 98% invalid responses for the Friedman #2 dataset. Consequently, we omitted Friedman #2 and Original #2 from its evaluation. However, it is important to highlight that the other models evaluated did not exhibit these issues and were able to generate valid outputs consistently. We make the following observations.

First, we remark that performance-wise, RWKV is worse than traditional transformer-based LLMs, although it generally remains better than our unsupervised baselines, with the exception on two linear regression datasets: Regression NI 1/3 and Regression NI 2/3. Nevertheless, we remark that on Original #1, RWKV outperforms many MLP variants, despite no gradient updates.

Second, we remark that the performance of Striped Hyena 7B is generally lower than than some of our unsupervised baselines. We note that there is a notable exception for Original #1. For this dataset, KNN approaches work well, as evident by the good performance obtained by nearest neighbor approaches.