Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks
Jongho Park, Jaeseung Park, Zheyang Xiong, Nayoung Lee, Jaewoong Cho, Samet Oymak, Kangwook Lee, Dimitris Papailiopoulos
Introduction
Modern large language models (LLMs) exhibit remarkable in-context learning (ICL) capabilities, enabling them to learn new tasks with a few demonstrations and without further weight fine-tuning. Although the exact emergence mechanism of these capabilities warrants further theoretical and empirical investigation (Chan et al., 2022; Wei et al., 2022; Min et al., 2022b; Schaeffer et al., 2023), experiments on larger Transformer-based models consistently demonstrate that their ICL capabilities improve as training loss reduces (Brown et al., 2020; Kaplan et al., 2020; Muennighoff et al., 2023).
Meta-learning, or “learning to learn,” has been extensively studied (Schmidhuber et al., 1997; Ravi & Larochelle, 2016) and recently regained interest in the context of ICL, particularly concerning Transformer models (Vaswani et al., 2017). Garg et al. (2022), for example, proposed various ICL tasks, such as learning linear regression, and evaluated the ability of transformers to perform them when specifically trained to do so. On the other hand, Min et al. (2022a) studied fine-tuning language models to explicitly learn and perform ICL. Following these footsteps, numerous research studies have been dedicated to understanding the mechanics of Attention that enable such meta-learning capabilities, either through constructive arguments or extensive experimental investigation (Akyürek et al., 2022; Li et al., 2023b; von Oswald et al., 2023b; Bai et al., 2023; Yang et al., 2023a; Li et al., 2023a; von Oswald et al., 2023a).
As Transformer language models are currently the only large models that have been reported to be capable of ICL in practice, this raises the question:
This question holds merit, especially considering that several recent studies have attempted to move beyond attention-based networks due to their quadratic cost (Gu et al., 2022b; Dao et al., 2022; Gu & Dao, 2023; Poli et al., 2023; Peng et al., 2023; Sun et al., 2023; Yang et al., 2023b). In this work, we focus specifically on state-space models (SSMs), and particularly Mamba (Gu & Dao, 2023). Mamba was recently demonstrated to be highly efficient while achieving near state-of-the-art performance in standard pretraining language data sets, such as the Pile (Gao et al., 2020), but at smaller model scales (e.g., up to 3 billion parameters), surpassing transformers and other attention-free architectures across various language and non-language tasks. However, ICL capabilities usually emerge at scales beyond 3 billion parameters. As a result, the potential of these attention-free models to perform ICL remains underexplored, as testing such hypotheses usually requires scaling beyond the 7 billion parameter level. Nonetheless, we can still investigate small-scale ICL capabilities by specifically training a model to perform in-context learning, following the approach of Garg et al. (2022).
In this study, we introduce a diverse set of ICL tasks to evaluate the performance of Transformer and various SSMs, including state-of-the-art models like Mamba and S4 (Gu et al., 2022b). Our findings reveal that most of these SSMs can effectively perform ICL, matching the performance of Transformers across multiple tasks. However, Mamba demonstrates some limitations in learning decision trees and retrieval tasks (as also noted by (Arora et al., 2023)), but can outperform Transformers in other complex ICL tasks, such as sparse parity, where Transformer models struggle. Performance of different models on each task is summarized in Table 1.
Since there seem to be tasks where either family of models is better, we explore the impact of interleaving SSM blocks with multi-head attention blocks, similar to (Gu & Dao, 2023). We introduce MambaFormer, a novel hybrid architecture that integrates Mamba and Attention layers, while eliminating the need for positional encodings, as shown in Figure 1. MambaFormer seems to leverage the strengths of both Mamba and Transformers, exhibiting good performance across all evaluated ICL tasks and simultaneously learning sparse parity and retrieval.
We believe that our findings underscore the importance of broadening the understanding of ICL beyond Transformers, as significant progress has been made in the context of attention-free architectures.
We acknowledge that a limitation of our study lies in the focus on non-language ICL tasks and smaller models. It is possible that an architectural comparison between SSMs and transformers for more general ICL tasks in actual language settings at higher parameter counts might not be yield the same observations as we offer here. Nevertheless, our results indicate that, apart from its difficulty in some retrieval tasks, similar to those noted by (Arora et al., 2023), there seems to be no fundamental obstacle for Mamba to perform in-context learning.
Related Work
The role of attention in ICL has been the focus of both theoretical and empirical research. Studies have primarily focused on meta-learning (Ravi & Larochelle, 2016; Min et al., 2022a), where one explicitly trains for ICL. Notably, Garg et al. (2022) have examined transformers in in-context regression tasks, from learning linear regression to learning decision trees. Subsequent works have suggested that attention may mimic various optimization algorithms (Akyürek et al., 2022; von Oswald et al., 2023b; Dai et al., 2023). In fact, Ahn et al. (2023); Mahankali et al. (2023) have provably shown that gradient descent is optimal in linear regression ICL for linear attention.
While these settings might appear simplistic and detached from language models, Bhattamishra et al. (2023) showed that a frozen GPT-2 can implement the nearest neighbor algorithm, drawing connections between the ICL in existing language models and the stylized setting of training for ICL from random initialization. Furthermore, Olsson et al. (2022) also empirically demonstrate that “induction heads”, which are attention heads that solve a simple retrieval problem, correlate with ICL behavior, providing a strong connection between retrieval and ICL.
The number of effective floating point operations in an attention layer scales quadratically with respect to the input sequence length. Numerous approximations or alternative model architectures have been proposed to overcome the quadratic dependence. These range from approximating attention mechanisms (Beltagy et al., 2020; Wang et al., 2020) to the development of novel recurrent convolutional models such as structured state-space models (Gu et al., 2022b).
S4 (Gu et al., 2022a) is a family of sequence models characterized by a discretized state-space model
where represents the hidden state and are input-independent (transformed) parameters. The recurrence is expressible as a convolution, enabling near-linear complexity using Fast Fourier Transform. Viewed in this framework, Linear Transformers (Katharopoulos et al., 2020), which employ linear attention without softmax, can be seen as a variant of linear SSM.
Building upon this concept, H3 (Dao et al., 2022), which integrates an S4 with dual gated connections. The recent Mamba (Gu & Dao, 2023) departs from the standard SSM by introducing a selection mechanism that makes in Equation 1 dependent on , which allows for input-dependent sequence mixing.
There are other notable attention-free models such as Hyena (Poli et al., 2023), RWKV (Peng et al., 2023), RetNet (Sun et al., 2023), and GLA (Yang et al., 2023b). Despite of state-of-the-art performance for models like Mamba, Arora et al. (2023) have demonstrated that subquadratic models still lag behind attention on multi-query recall tasks, which is a generalization of the induction head task (Olsson et al., 2022).
In Xie et al. (2021), the authors proposed a synthetic language-based in-context learning dataset and show that transformers and LSTMs are capable of ICL. Moreover, Akyürek et al. (2024) suggested a langauge based ICL benchmark, by training on regular languages generated by random finite automata, and also underscored the gap between Transformers and subquadratic complexity models.
Experimental Setup
We evaluate the ICL capabilities of SSMs and Transformers by training each model from scratch on each specific task, detailed in Section 3.1. Section 3.2 outlines the ICL and related tasks investigated in our study. We provide a brief summary of our tasks in the following Table 2.
We primarily focus on SSMs, including (1) Mamba (Gu & Dao, 2023), a state-of-the-art SSM model with selection mechanism; (2) S4 (Gu et al., 2022a), a linear time-invariant counterpart to Mamba; and (3) S4-Mamba, a variant where Mamba’s input-dependent S6 is replaced with input-independent S4. The primary differences between the S4 models lie in the application of multiplicative gating and the module order.https://github.com/state-spaces/s4/blob/main/models/s4
We train each model by sampling a batch of random prompts at each training step and updating the model parameters using Adam optimizer (Kingma & Ba, 2014). We use a batch size of 64 and trained for 500,000 iterations (except for the vector-valued MQAR task; see Section B.2).
We evaluate model performance on in-context learning using task and data distributions and consistent with training. A function and a sequence of inputs are sampled from and , respectively, to generate a test prompt . We create 1,280 prompts and measure the empirical mean of Eq. (2) across them for in-context learning performance.
To plot performance as model capacity grows, we calculate the total floating point operations (FLOPs) used for training. The calculation for Transformer and Mamba can be found in Appendix C, which are based on (Kaplan et al., 2020; Gu & Dao, 2023). Model configurations and training implementation details are provided in Appendix A.
2 In-context learning tasks
We provide an overview of the ICL and related tasks investigated in this study. Some tasks are adapted from (Garg et al., 2022), and we follow the settings outlined in their work. The tasks are summarized in Table 2.
For all regression tasks, in-context examples are sampled from the Gaussian distribution , where is the identity matrix. We use the squared error loss for model training.
The setting is identical to linear regression, except that is sampled from , after which coordinates are randomly retained in , and the rest are set to zero. We set .
2.2 Learning with outliers
The problems that belong to this family adopt the basic setting of the standard linear regression task. With a fixed probability , each pair of in the prompt is replaced with “dummy” vectors which are either out of the training distribution, or confounders designed to increase the complexity of the task. We test as replacement probabilities for tasks described below. During training, we do not compute the loss for the replaced outliers.
Each pair of is randomly replaced with , where . and and are sampled from and the coefficients are independently sampled from .
In this setting, and are randomly replaced with a -dimensional vector of ones and an one-hot vector , respectively, with probability 90%. Here, we test longer sequences of .
2.3 Learning discrete functions
Following the setting from Bhattamishra et al. (2023), we consider the class of functions , where denotes the -th element of the vector and is a subset of with the size . Each is sampled uniformly at random from , and of size is randomly sampled from the set . For this task, we train a model using the cross-entropy loss and evaluate the model using a binary indicator for accuracy, which assigns 1 to correct predictions and 0 to incorrect ones.
2.4 Learning Chain-of-Thought
2.5 Learning retrieval
Vector-valued multi-query associative recall
We test the model’s ability to do multi-query associative recall (MQAR) (Arora et al., 2023). While MQAR is not an ICL task, model’s ability to do associative recall (AR) is highly related to model’s ability to learn in-context (Olsson et al., 2022). To better measure the model’s ability to retrieve information from context, we consider a variant of MQAR such that keys and values are vector-valued so each vector can be seen as a “unique token” and the retrieval accuracy can be measured by the mean squared error between retrieved vectors and target vectors. Specifically, in this task, the model is given a sequence of key-value pairs of vectors , where are sampled uniformly from the unit -sphere. The query consists of sequence of vectors . For each query , there exists some such that . The model must learn to output associated with the query for each of the queries, producing outputs total. We train a model using the squared error error.
Experiment results
In this section, we demonstrate that Mamba can be trained from scratch to perform various ICL tasks. Furthermore, we identify specific tasks in which one model performs better than the other and vice versa.
As shown in Figure 2, Mamba consistently outperforms its more simple counterparts S4-Mamba and S4. In simple tasks such as linear regression, the gap between Mamba and S4-Mamba is much smaller than that of S4-Mamba and S4. Given that the main difference between Mamba and S4-Mamba is the input-dependent selection mechanism, appropriate gating and stacking of MLPs (i.e., the difference between S4-Mamba and S4) seem to be more significant for such tasks. However, in comparison, input-dependent selection makes meaningful progress for more complex tasks such as 2NN regression and learning decision trees.
Mamba can also perform on par with Transformer even as the total FLOPs scale up. This is surprising given that Transformer and attention have been the focus of many previous works for its unique ICL capability. Moreover, Mamba tends to perform better in smaller parameter settings when controlling for equal depth, i.e., keeping the number of attention, MLP, and Mamba blocks equivalent.
2 Performance gaps in more complex ICL tasks
We also consider a family of more complex ICL tasks, namely learning decision tree, sparse parity, and Chain-of-Thought (Figures 2 and 4). The figure shows that Transformers can solve Decision Tree and Vector-valued MQAR, while Mamba cannot. In Sparse Parity task of Figure 5, however, Transformer is unable to learn the function family while Mamba can.
Orthogonal-outlier regression and many-outlier regression, like other outlier tasks, focus on the model’s ability to learn to ignore dummy vectors, either by the fact that the , or by the fact that is a vector instead of a zero-padded scalar value. This explicitly requires the models to look at the previous input sequences, and discover the properties that distinguish the dummy vectors from training examples while learning the class of functions the training prompt represents.
For orthogonal-outlier regression task with a relatively short sequence length of 101 (see Table 2 for task descriptions), Mamba performs on par with Transformer, as seen in Figure 3. Interestingly, for many-outlier regression where we test on a sequence length of 512 and 90% all-ones replacement, Mamba significantly outperforms Transformers. This is also in line with what Gu & Dao (2023) report, in which Mamba fares better for the induction task for long sequence lengths. These two results indicate that Mamba has no significant issue with filtering out unnecessary information, while retaining the ability to learn linear regression in-context.
Figure 4 shows that Mamba models are capable of in-context learning in a chain-of-thought manner, performing comparably to Transformer models across the tested configurations. In smaller model configurations, Mamba models exhibit superior performance compared to Transformer models. However, as model size increases, Transformer models begin to surpass Mamba models. The performance of Transformer models remains relatively stable across different problem sizes, while Mamba models’ performance is significantly influenced by the size of the hidden layer. Specifically, Mamba models excel over Transformer models at smaller problem sizes (i.e., smaller hidden dimensions), but their advantage diminishes as the problem size expands.
3 Challenges in parity and retrieval
We run vector-valued MQAR on two settings: (1) key-value pairs with queries and (2) key-value pairs with queries. From Table 3, we can see that Mamba fails to retrieve vectors accurately as the mean squared error for retrieving normed vectors are greater than in all cases.
As a sidenote, all models trained with queries have lower test loss than models trained with queries. A possible explanation is that, for a single sequence of data that represents an MQAR task, we can think of each pair as a “training sample”, so a sequence with queries contains more “training samples” than that of a sequence with queries. This also shows that having more queries does not necessarily make the task harder.
While Mamba fails on simple retrieval tasks such as MQAR, the tables turn for the task of learning sparse parity (Figure 5). Transformer fails to do better than random guessing, in line with the empirical evidence of Bhattamishra et al. (2023). We confirm this is the case for Transformer sizes of embedding dimensions up to 768 and up to 24 layers when trained for at most 1 million iterations. However, Mamba succeeds in this task with ease, solving sparse parity for with a network as small as 2 layers. Even more surprisingly, S4-Mamba is able to solve parity as well; this may mean that proper convolution or gating may be more important than input-dependent selection. Our result hints at that the initial (causal) convolution that Mamba provides before the attention layer may be crucial to solving parities, a similar phenomenon observed for Vision Transformers in computer vision tasks (Yu et al., 2022).
It is known that any algorithm for learning parities requires either a super-linear memory of or a super-polynomial number of samples in (Raz, 2016; Kol et al., 2017). While Transformer is known to have better memory due to its quadratic attention mechanism, our results on learning sparse parities brings forth the question on how different architectures may utilize its memory differently in terms of function approximation. We leave the theoretical and empirical question of which architectural component allows for learning parities as an avenue for further study.
The Advantage of Hybrid Architectures for In-context Learning
In the previous section, we have observed that Transformers perform better than SSMs in some tasks, such as learning decision trees or retrieval, while SSMs excel in others, such as learning sparse parities or learning heavy-outlier linear regression, possibly due to its recurrent nature. However, can we achieve the best of both worlds without sacrificing performance in our suite of ICL tasks?
We answer this in the affirmative; that we can indeed reach competitive performance in our suite of ICL tasks, achieving performance comparable to that of Transformers and Mamba, while simultaneously excelling in specific tasks that either fail in. We can achieve strong performance by interleaving Attention and Mamba, where a key ingredient is having Mamba as the first layer.
In this section, we investigate two hybrid architectures that combine Transformer and Mamba, namely Standard Hybrid and MambaFormer as illustrated in Figure 6. Standard Hybrid is the architecture of interleaving MHA and Mamba by replacing the MLP block with Mamba. MambaFormer is nearly identical to Standard Hybrid but with an additional Mamba block as its initial layer and no particular positional encoding. Although many works have found that interleaving multi-head attention and LTI SSMs beneficial (Zuo et al., 2022; Mehta et al., 2022; Pilault et al., 2023), interestingly Gu & Dao (2023) have not found significant benefits of interleaving. In the following results, we show that interleaving with Mamba as its initial layer can help solve both sparse parity and retrieval, each task unsolvable by Mamba and Transformer.
As highlighted in Bhattamishra et al. (2023); Barak et al. (2022), learning sparse parity in-context seems to be difficult for Transformer and some SSMs like Hyena. Yet interestingly, as seen in Figure 7, MambaFormer successfully learns parity as quickly as Mamba in terms of sample complexity. While the Standard Hybrid model is also capable, it exhibits much worse sample efficiency.
We perform an ablation study by equipping Transformer with an initial Mamba block without any positional encoding. Although this variant Transformer only has fewer Mamba blocks than Standard Hybrid, it solves parity almost as efficiently as Mamba. Not only does this show us that order of layers in interleaving matter, as shown in Press et al. (2022), but also that Mamba can complement Transformer without hurting performance in ICL. This result brings up intriguing difference between the function learning capabilities of Attention and Mamba; we leave this question up for further study.
The gap between Mamba and Transformer in vector-valued MQAR task is largely due to the fact that Mamba (as an SSM) compresses context into smaller states when generating output, while the Attention mechanism in Transformer does not compress the context. The amount of information about the context Mamba has at each state depends on the dimension of hidden state (as the hidden states capture the important information in the context) and it is challenging if the task is to accurately retrieve a specific part of the context by a query that is placed after the context.
To close the gap in the vector-valued MQAR task between Mamba and Transformer, and without sacrificing too much of the efficiency, we add one attention layer within layers of Mamba blocks. In particular, in a Mamba model of layers ( Mamba blocks stacked homogeneously), we replace the middle two blocks with Standard Hybrid (w/o positional embedding). As shown in Table 3, Mamba model gains a significant improvement in vector-valued MQAR by having one Standard Hybrid. We further test MambaFormer on the same task and find that MambaFormer almost entirely closes the gap to transformer in vector-valued MQAR task.
2 All-in-one ICL performance
While MambaFormer succeeds in two tasks that were either deemed difficult for Mamba for Transformer, it also performs equally as well as Transformer and Mamba in the rest of our suite of ICL tasks. In Figure 2, we see that MambaFormer and Standard Hybrid both learn decision trees as well as Transformer, even at larger parameter sizes. Even more surprisingly, MambaFormer efficiently learns linear regression better than both models even in the presence of 90% noisy data in Many-outlier regression, as a MambaFormer trained on 100k iterations ( FLOPs) performs as well as models trained with 10 times the number of FLOPs.
In conclusion, we find the best of both worlds within our diverse array of ICL tasks; a hybrid architecture that can solve as difficult problems as retrieval and parity, while performing on par with Transformer and Mamba in other ICL tasks. Given our results, it will be interesting to see how hybrid architectures perform in other kinds of ICL tasks, such as those discussed in (Xie et al., 2021; Akyürek et al., 2024).
Discussion
In this work, we have provided a comprehensive investigation of in-context learning with state-space models (SSMs) and contrasted them with the transformer architecture. Our study has revealed that SSMs, especially Mamba, are capable in-context learners. On the other hand, our evaluations revealed that neither SSMs nor transformers are great at all tasks, specifically, SSMs struggle with decision tree learning and retrieval tasks whereas transformers struggle with sparse parity. This has led us to the hybrid architecture MambaFormer which achieves a best-of-both-worlds performance on our ICL suite.
Future research directions include exploring (1) how performance on our ICL suite correlates with general language modeling capabilities, such as perplexity on standard NLP benchmarks, (2) the potential for developing more effective architectures by integrating elements from transformers, SSMs, and gating mechanisms, (3) identifying architectural features that contribute to effective in-context learning, and (4) assessing the impact of MambaFormer and other innovative architectures on language modeling performance.
References
Appendix A Experimental Setup
We focus on decoder-only Transformer models, particularly those from the GPT-2 family (Radford et al., 2019), Mamba (Gu & Dao, 2023), and their Hybrid variants, including Standard and MambaFormer configurations. These models are evaluated across a range of sizes, as detailed in Table 4. Transformer layers consist of a Multi-Head Attention (MHA) block followed by a Multilayer Perceptron (MLP) block. Mamba models consist of two Mamba blocks per layer. The Hybrid variants merge these approaches, combining a single MHA block with a Mamba block. For MHA blocks, we use 8 number of heads. Refer to Figure 6 for a visualization of the architectures considered.
A.2 Model Training
We train all of our models on A100-SXM4-40GB GPUs for 500,000 training steps on all tasks. We use Adam optimizer Kingma & Ba (2014) with a fixed learning rate. The default value is set to , following the default learning rate in Garg et al. (2022), and search various learning rates in . We observe that the training procedure is the most sensitive to choosing the right learning rate. In particular, as the number of parameters of the models increases, the training procedure is prone to gradient explosions, especially in Mamba and hybrid architecutres. Hence, we clip the gradient norm, with values in .
As for the train and test data, we fix the dimension of to be , and fix the batch size to be . As suggested in Garg et al. (2022), we also observe that curriculum is crucial in certain ICL tasks. We adopt a curriculum of 15 steps every 2000 steps both on the dimension of and the number of points (half the length of the training prompt).
Appendix B Implementation Details
This section further elaborates on the task descriptions from Section 3.
Table 5 presents the configurations for the Chain-of-Thought-I/O task using a 2-layer ReLU neural network, following the setup described by Li et al. (2023b). In the model scale experiment, the input dimension and hidden layer dimension are held constant while varying the model scale. Additionally, the hidden dimension is varied among while fixing the model scale to small to identify the effect of problem scale.
B.2 Vector-valued MQAR
The training set consists of training samples. We train for epochs with batch size of and evaluate on a test set of samples. For each setting, we sweep with learning rates in np.logspace(-4, -2, 4) and report the best result among all learning rates.
B.3 Orthogonal-outlier Regression
We run Orthogonal-outlier regression on all four model architectures, with varying number of parameters. The curves above generally capture the trend that Mamba and Transformer perform on par with each other, and Standard Hybrid and MambaFormer show better curves compared to the vanilla models.
One curve that stands out would be Mamba’s loss curve in the smaller regime. As a future direction, we have tested trading off the depth and width while fixing the total number of parameters with the smallest Mamba model configuration in Table 6.
The best performing model configuration was with 16 layers, and embedding dimension of 64. This suggests that with fixed number of parameters, increasing the depth may boost the downstream task performance on ICL tasks.
Appendix C FLOPs computation
We count the number of multiplications in a Mamba block and a Transformer block in Table 7 and Table 8. We assume batch size . To calculate FLOPs, we multiply the number of multiplications by to account for the multiply-accumulate cost in both forward and backward pass. Note that a Standard Hybrid block is an attention block stacked with a Mamba block, so the number of multiplications in a Standard Hybrid block is , ignoring the linear terms.