A Primer on the Inner Workings of Transformer-based Language Models

Javier Ferrando, Gabriele Sarti, Arianna Bisazza, Marta R. Costa-jussà

Introduction

The development of powerful Transformers-based language models (LMs; Radford et al., 2019; Brown et al., 2020; Hoffmann et al., 2022; Chowdhery et al., 2023) and their widespread utilization underscores the significance of research devoted to understanding their inner mechanisms. Gaining a deeper understanding of these mechanisms in highly capable AI systems holds important implications in ensuring the safety and fairness of such systems, mitigating their biases and errors in critical settings, and ultimately driving model improvements (Wei et al., 2022; Costa-jussà et al., 2023). As a result, the natural language processing (NLP) community has witnessed a notable increase in research focused on interpretability in language models, leading to new insights into their internal functioning.

Existing surveys present a wide variety of techniques adopted by Explainable AI analyses (Räuker et al., 2023) and their applications in NLP (Madsen et al., 2022; Lyu et al., 2024). While previous NLP interpretability surveys primarily focused on encoder-based models like BERT (Devlin et al., 2019; Rogers et al., 2021), the success of decoder-only Transformers (Radford et al., 2018) prompted further developments in the analysis of these powerful generative models, with concurrent work surveying trends in interpretability research and their relation to AI safety (Bereska & Gavves, 2024). By contrast, this work provides a concise, in-depth technical introduction to relevant techniques used in LM interpretability research, focusing on insights derived from models’ inner workings and drawing connections between different areas of interpretability research. Moreover, throughout this work, we employ a unified notation to introduce model components, interpretability methods, and insights from surveyed works, shedding light on the assumptions and motivations behind specific method designs. We categorize LM interpretability approaches surveyed in this work along two dimensions: i) localizing the inputs or model components responsible for a particular prediction (Section 3); and ii) decoding information stored in learned representationsIn this work we use representations and activations interchangeably, and we refer to the fundamental unit of information encoded in model activations as features, representing human-interpretable input properties. to understand its usage across network components (Section 4). Finally, Section 5 provides an exhaustive list of insights into the inner workings of Transformer-based LMs, and Section 6 provides an overview of useful tools to conduct interpretability analyses on these models.

The Components of a Transformer Language Model

Auto-regressive language models assign probabilities to sequences of tokens. Using the probability chain rule, we can decompose the probability distribution over a sequence t=⟨t1,t2…,tn⟩{\mathbf{t}=\langle t_{1},t_{2}\ldots,t_{n}\rangle} into a product of conditional distributions:

In this section, we present the Transformer layer components following their computations’ flow.

We note that current LMs such as Llama 2 (Touvron et al., 2023) adopt an alternative layer normalization procedure, RMSNorm (Zhang & Sennrich, 2019), where the centering operation is removed, and scaling is performed using the root mean square (RMS) statistic.

1.2 Attention Block

1.3 Feedforward Network Block

The computation described in Equation 6 was equated to key-value memory retrieval (Geva et al., 2021), with keys (winl{\bm{w}}^{l}_{\text{in}}) stored in columns of Winl{\bm{W}}^{l}_{\text{in}} acting as pattern detectors over the input sequence (Figure 2 right) and values woutl{\bm{w}}^{l}_{\text{out}}, rows of Woutl{\bm{W}}^{l}_{\text{out}}, being upweighted by each neuron activation. We use the term “neuron” to refer to each value after an element-wise non-linearity, and use “unit” or “dimension” for other individual values in any other representation. Provided that the output of the FFN is a linear combination of woutl{\bm{w}}^{l}_{\text{out}} values, Equation 6 can be rewritten following the key-value perspective:

The elementwise nonlinearity inside FFNs creates a privileged basis (Elhage et al., 2022b), which encourages features to align with basis directions. For instance, given a linear network f(x)=xW1W2f({\bm{x}})={\bm{x}}{\bm{W}}_{1}{\bm{W}}_{2}, the representations extracted from its first layer, xW1{\bm{x}}{\bm{W}}_{1}, are rotationally invariant, since we can rotate them by an orthogonal matrix O{\bm{O}}, giving xW1O{\bm{x}}{\bm{W}}_{1}{\bm{O}}, and invert the rotation having the output of the network untouched, f(x)=xW1OO−1W2f({\bm{x}})={\bm{x}}{\bm{W}}_{1}{\bm{O}}{\bm{O}}^{-1}{\bm{W}}_{2} (Brown et al., 2023). However, having an elementwise nonlinear function on the output of the first layer breaks the rotational invariance of the representations, making the standard basis dimensions (neurons) more likely to be independently meaningful, and therefore better suitable for interpretability analysis.

2 Prediction Head and Transformer Decompositions

The residual stream view shows that every model component interacts with it through addition (Mickus et al., 2022). Thus, the unnormalized scores (logits) are obtained via a linear projection of the summed component outputs. Due to the properties of linear transformations, we can rearrange the traditional forward pass formulation so that each model component contributes directly to the output logits:

This decomposition plays an important role when localizing components responsible for a prediction (Section 3) since it allows us to measure the direct contribution of every component to the logits of the predicted token (Section 3.2.1).

Residual networks work as ensembles of shallow networks (Veit et al., 2016), where each subnetwork defines a path in the computational graph. Let us consider a two-layer attention-only Transformer, where each attention head is composed just by an OV matrix: f(x)=x1+WOV2(x1){f({\bm{x}})={\bm{x}}^{1}+{\bm{W}}_{OV}^{2}({\bm{x}}^{1})}, with x1=x+WOV1(x){{\bm{x}}^{1}={\bm{x}}+{\bm{W}}_{OV}^{1}({\bm{x}})}. We can decompose the forward pass (Figure 3) as

Behavior Localization

Understanding the inner workings of language models implies localizing which elements in the forward pass (input elements, representations, and model components) are responsible for a specific prediction.Commonly referred to as local explanation in the interpretability literature (Lipton, 2018)In this section, we present two different types of methods that allow localizing model behavior: input attribution (Section 3.1) and model component attribution (Section 3.2).

Input attribution methods are commonly used to localize model behavior by estimating the contribution of input elements (in the case of LMs, tokens) in defining model predictions. We refer readers to Madsen et al. (2022) for a broader overview of post-hoc input attribution methods with a focus on classification tasks in NLP.

Another popular family of approaches estimates input importance by adding noise or ablating input elements and measuring the resulting impact on model predictions (Li et al., 2017). For instance, the input token at position ii can be removed, and the resulting probability difference fw(x)−fw(x−xi){f_{w}(\mathbf{x})-f_{w}(\mathbf{x}_{-{\bm{x}}_{i}})} can be used as an estimate for its importance. If the logit or probability given to ww does not change, we conclude that the ii-th token has no influence. A multitude of perturbation-based attribution methods exist in the literature, such as those based on interpretable local surrogate models such as LIME (Ribeiro et al., 2016), or those derived from game theory like SHAP (Shapley, 1953; Lundberg & Lee, 2017). Notably, new perturbation-based approaches were proposed to leverage linguistic structures (Amara et al., 2024; Zhao & Shan, 2024) and Transformer components (Deiseroth et al., 2023; Mohebbi et al., 2023) for attribution purposes. These methods relate directly to causal interventions discussed in Section 3.2.2. We refer readers to Covert et al. (2021) for a unified perspective on perturbation-based input attribution.

While raw model internals such as attention weights were generally considered to provide unfaithful explanations of model behavior (Jain & Wallace, 2019; Bastings & Filippova, 2020), recent methods have proposed alternatives to attention weights for measuring intermediate token-wise attributions. Some of these alternatives include the use of the norm of value-weighted vectors (Kobayashi et al., 2020) and output-value-weighted vectors (Kobayashi et al., 2021), or the use of vectors’ distances to estimate contributions (Ferrando et al., 2022b) (Figure 4 provides a visual description). A common strategy among such approaches involves aggregating intermediate per-layer attributions reflecting context mixing patterns (Brunner et al., 2020) using techniques such as attention rollout (Abnar & Zuidema, 2020), resulting in input attribution scores (Ferrando et al., 2022b; Modarressi et al., 2022; Mohebbi et al., 2023).The attention flow method is seldom used due to its computational inefficiency, despite its theoretical guarantees (Ethayarajh & Jurafsky, 2021). Such context mixing approaches have shown strong faithfulness compared to gradient and perturbation-based methods on classification benchmarks such as ERASER (DeYoung et al., 2020). However, rollout aggregation has recently been criticized due to its simplistic assumptions, and recent research has attempted to fully expand the linear decomposition of the model output presented in Section 2.2 (Modarressi et al., 2023; Yang et al., 2023; Oh & Schuler, 2023) as a sum of linear transformations of the input tokens, linearizing the FFN block (Kobayashi et al., 2024).

An important limitation of input attribution methods for interpreting language models is that attributed output tokens belong to a large vocabulary space, often having semantically equivalent tokens competing for probability mass in next-word prediction (Holtzman et al., 2021). In this context, attribution scores are likely to misrepresent several overlapping factors such as grammatical correctness and semantic appropriateness driving the model prediction. Recent work addresses this issue by proposing a contrastive formulation of such methods, producing counterfactual explanations for why the model predicts token ww instead of an alternative token oo (Yin & Neubig, 2022). As an example, Yin & Neubig (2022) extend the vanilla gradient method of Equation 12 to provide contrastive explanations (ContGrad):

While input attribution methods are commonly used to debug failure cases and identify biases in models’ predictions (McCoy et al., 2019), popular approaches were shown to be insensitive to variations in the model and data generating process (Adebayo et al., 2018; Sixt et al., 2020), to disagree with each others’ predictions (Atanasova et al., 2020; Crabbé & van der Schaar, 2023; Anonymous, 2024) and to show limited capacity in detecting unseen spurious correlations (Adebayo et al., 2020; 2022). Importantly, popular methods such as SHAP and Integrated Gradients were found provably unreliable at predicting counterfactual model behavior in realistic settings Bilodeau et al. (2024). Apart from theoretical limitations, perturbation-based approaches also suffer from out-of-distribution predictions induced by unrealistic noised or ablated inputs, and from high computational cost of targeted ablations for granular input elements.

Another dimension of input attribution involves the identification of influential training examples driving specific model predictions at inference time (Koh & Liang, 2017). These approaches are commonly referred to as training data attribution (TDA) or instance attribution methods and were applied to identify data artifacts (Han et al., 2020; Pezeshkpour et al., 2022) and sources of biases in language models’ predictions (Brunet et al., 2019), with recent approaches proposing to perform TDA via training run simulations (Guu et al., 2023; Liu et al., 2024). While the applicability of established TDA methods was put in question (Akyurek et al., 2022), especially due to their inefficiency, recent work in this area has produced more efficient methods that can be applied to large generative models at scale (Park et al., 2023b; Grosse et al., 2023; Kwon et al., 2024). We refer readers to (Hammoudeh & Lowd, 2022) for further details on TDA methods.

2 Model Component Importance

Early studies on the importance of Transformers LMs components highlighted a high degree of sparsity in model capabilities. This means, for example, that removing even a significant fraction of the attention heads in a model may not deteriorate its downstream performances (Michel et al., 2019; Voita et al., 2019b). These results motivated a new line of research studying how various components in an LM contribute to its wide array of capabilities.

Let us call fc(x)f^{c}(\mathbf{x}) the output representation of a model component cc (attention head or FFN) at a particular layer for the last token position nn. The decomposition presented in Section 2.2 allows us to measure the direct logit attributionNote that the softmax function is shift-invariant, and therefore the logit scores have no absolute scale. (DLA, Figure 5) of each model component for the output token w∈Vw\in\mathcal{V}:

where WU[:,w]{\bm{W}}_{U[:,w]} is the ww-th column of WU{\bm{W}}_{U}, i.e. the unembedding vector of token ww. In practical terms, the DLA for a component cc expresses the contribution of cc to the logit of the predicted token, using the linearity of the model’s components described in Section 2.2.

Geva et al. (2022b) exploit the fact that the FFN block update is a linear combination of the rows of Wout{\bm{W}}_{\text{out}} weighted by the neuron activation values (Equation 8). Thus, it is possible to measure the DLA of each neuron as:

Similarly, Ferrando et al. (2023) makes use of the decomposition of an attention head as a weighted sum of residual stream transformations (Section 2.1.2) and proposes assessing the DLA of each path involving the attention head:

The Logit difference (LD) (Wang et al., 2023a) is the difference in logits between two tokens, fw(x)−fo(x){f_{w}(\mathbf{x})-f_{o}(\mathbf{x})}. DLA can be extended to measure direct logit difference attribution (DLDA):

including its neuron and head-specific variants of Equation 15 and Equation 16. Similarly to the contrastive attribution framework described in Section 3.1, a positive DLDA value suggests that cc promotes token ww more than token oo.

2.2 Causal Interventions

Alternatively, other sources of patching activations include:

An important factor to consider when designing causal interventions experiments is the ecological validity of the setup, since zero and noise ablation could lead the model away from the natural activations distribution and ultimately undermine the validity of components’ analysis (Chan et al., 2022; Zhang & Nanda, 2024).

Other forms of causal interventions use differentiable binary masking on subsets of units or neurons of intermediate representations (De Cao et al., 2020; Csordás et al., 2021; De Cao et al., 2022), or entire attention heads outputs (Voita et al., 2019b; Michel et al., 2019) which can be cast as a form of zero ablation.

2.3 Circuits Analysis

The Mechanistic Interpretability (MI) subfield focuses on reverse-engineering neural networks into human-understandable algorithms (Olah, 2022). Recent studies in MI aim to uncover the existence of circuits, which are a subset of model components (subgraphs) interacting together to solve a task (Cammarata et al., 2020). Activation patching, logit attribution, and attention pattern analysis are common techniques for circuit discovery (Wang et al., 2023a; Stolfo et al., 2023a; b; Heimersheim & Janiak, 2023; Geva et al., 2023; Hanna et al., 2023).

Activation patching propagates the effect of the intervention throughout the network by recomputing the activations of components after the patched location (Figure 6). The changes in the model output (Equation 18) allow estimating the total effect of the model component on the prediction. However, circuit discovery also requires identifying important interactions between components. For this purpose, edge patching exploits the fact that every model component input is the sum of the output of previous components in its residual stream (Section 2.2), and considers edges directly connecting pairs of model components’ nodes (Figure 7 right). Path patching generalizes the edge patching approach to multiple edges (Wang et al., 2023a; Goldowsky-Dill et al., 2023), allowing for a more fine-grained analysis. For example, using the forward pass decomposition into shallow networks described in Equation 11, we could visualize the single-layer Transformer of Figure 7 (right) as being composed as

[yshift=0em]above, left, label belowadpAttn direct path to logits \annotate[yshift=0em]above, right, label belowaipAttn indirect path to logits via FFN

where each copy of the sender node AttnL(X≤nL−1)\text{Attn}^{L}(\bm{X}^{L-1}_{\leq n}) is relative to a single path. In this example, patching separately each of the sender node copies (Goldowsky-Dill et al., 2023) allows us to estimate direct and indirect effects (Pearl, 2001; Vig et al., 2020) of AttnL(X≤nL−1)\text{Attn}^{L}(\bm{X}^{L-1}_{\leq n}) to the output logits f(x)f(\mathbf{x}). In general, we can apply path patching to any path in the network and measure composition between heads, FFNs, or the effects of these components on the logits.

Circuit analysis based on causal intervention methods presents several shortcomings:

it demands significant efforts for designing the input templates for the task to evaluate, along with the counterfactual dataset, i.e. defining PpatchP_{\text{patch}}.

isolating important subgraphs after obtaining component importance estimates requires human inspection and domain knowledge.

it has been shown that interventions can produce second-order effects in the behavior of downstream components (Makelov et al., 2024, see Wu et al., 2024d for discussion), in some settings even eliciting compensatory behavior akin to self-repair (McGrath et al., 2023; Rushing & Nanda, 2024). This phenomenon can make it difficult to draw conclusions about the role of each component.

Conmy et al. (2023) propose an Automatic Circuit Discovery (ACDC) algorithm to automate the process of circuit identification (Limitation 2) by iteratively removing edges from the computational graph. However, this process requires a large amount of forward passes (one per patched element), which becomes impractical when studying large models (Lieberum et al., 2023). A valid alternative to patching involves gradient-based methods, which have been extended beyond input attribution to compute the importance of intermediate model components (Leino et al., 2018; Shrikumar et al., 2018; Dhamdhere et al., 2019). For instance, given the token prediction ww, to calculate the attribution of an intermediate layer ll, denoted as fl(x)f^{l}(\mathbf{x}), the gradient ∇fw(fl(x))\nabla f_{w}(f^{l}(\mathbf{x})) is computed. Sarti et al. (2023) extend the contrastive gradient attribution formulation of Equation 13 to locate components contributing to the prediction of the correct continuation over the wrong one using a single forward and backward pass. Nanda (2023); Syed et al. (2023) propose Edge Attribution Patching (EAP), consisting of a linear approximation of the pre- and post-patching prediction difference (Equation 18) to estimate the importance of each edge in the computational graph. The key advantage of this method is that it requires two forward passes and one backward pass to obtain attribution scores of every edge in the graph. Hanna et al. (2024) propose combining EAP with Integrated Gradients (EAP-IG) and show improved faithfulness of the extracted circuits, a method also used by Marks et al. (2024) to identify sparse feature circuits. Further work on Attribution Patching by Kramár et al. (2024) finds two settings leading to false negatives in the linear approximation of activation patching, and proposes AtP∗, a more robust method preserving a good computational efficiency. Recently, Ferrando & Voita (2024) propose finding relevant subnetworks, which they name information flow routes, using a patch-free context mixing approach, requiring only a single forward pass, avoiding the dependence on counterfactual examples and the risk of self-repair interferences during the analysis.

Another line of research deals with finding interpretable high-level causal abstractions in lower-level neural networks (Geiger et al., 2021; 2022; 2023a). These methods involve a computationally expensive search and assume high-level variables align with groups of units or neurons. To overcome the limitations, Geiger et al. (2023b) propose distributed alignment search (DAS), which performs distributed interchange interventions (DII, Section 3.2.2) on non-basis-aligned subspaces of the low-level representation space found via gradient descent.Alternatively, Lepori et al. (2023) proposes employing circuit discovery approaches for this purpose. DAS interventions have been shown to be effective in finding features with causal influence in targeted syntactic evaluation (Arora et al., 2024), and in isolating the causal effect of individual attributes of entities (Huang et al., 2024a). Recently, learned edits on subspaces of intermediate representations during the forward pass have been proposed as an efficient and effective alternative to weight-based Parameter-efficient fine-tuning (PEFT) approaches (Wu et al., 2024b). A DAS variant named Boundless DAS has been used to search for interpretable causal structure in large language models (Wu et al., 2023b). In this context, Causal Proxy Models (CPMs) were proposed as interpretable proxies trained to mimic the predictions of lower-level models and simulate their counterfactual behavior after targeted interventions (Wu et al., 2023a).

Information Decoding

Fully understanding a model prediction entails localizing the relevant parts of the model, but also comprehending what information is being extracted and processed by each of these components. For example, if the grammatical gender of nouns is assumed to be relevant for the task of coreference resolution in a given language, information decoding methods could look at whether and how a model performing this task encodes noun gender. A natural way to approach decoding the information in the network is in terms of the features that are represented in it. While there is no universally agreed-upon definition of a feature, it is typically described as a human-interpretable property of the inputAlthough we have evidence that models learns human-interpretable features even in instances that exceed human performance (McGrath et al., 2022), Olah (2022) argues that the definition of feature should include properties that are not human-interpretable., which can be also referred to as a concept (Kim et al., 2018).

Probes, introduced concurrently in NLP by Köhn (2015); Gupta et al. (2015) and in computer vision by Alain & Bengio (2016) serve as tools to analyze the internal representations of neural networks. Generally, they take the form of supervised models trained to predict input properties from the representations, aiming to asses how much information about the property is encoded in them. Formally, the probing classifier p:fl(x)↦z{p:f^{l}(\mathbf{x})\mapsto z} maps intermediate representations to some input features (labels) zz, which can be, for instance, a part-of-speech tag (Belinkov et al., 2017), or semantic and syntactic information (Peters et al., 2018). For example, for a binary probe seeking to decode the amount of input sentiment information within an intermediate representation (Figure 8) we build two sets: {fl(x):x∈P}\{f^{l}(\mathbf{x}):\mathbf{x}\in\mathcal{P}\} and {fl(x):x∈N}\{f^{l}(\mathbf{x}):\mathbf{x}\in\mathcal{N}\}, with the representations obtained when providing positive and negative sentiment sentences respectively. After training the classifier we evaluate the accuracy results on a held-out set.

Although performance on the probing task is interpreted as evidence for the amount of information encoded in the representations, there exists a tension between the ability of the probe to evaluate the information encoded and the probe learning the task itself (Belinkov, 2022). Several works propose using baselines to contextualize the performance of a probe. Hewitt & Liang (2019) use control tasks by randomizing the probing dataset, while Pimentel et al. (2020) propose measuring the information gain after applying control functions on the internal representations. Voita & Titov (2020) suggest evaluating the quality of the probe together with the “amount of effort” required to achieve the quality. This is done by measuring the minimum description length of the code required to transmit labels zz given representations fl(x)f^{l}(\mathbf{x}). We refer the reader to Belinkov & Glass (2019); Belinkov (2022) for a larger coverage of probing methods.

Probing techniques have been largely applied to analyze Transformers in NLP. Although probes are still being used to study decoder-only models (CH-Wang et al., 2023; Zou et al., 2023; Burns et al., 2023; MacDiarmid et al., 2024), a significant portion of the research in this area has focused on BERT (Devlin et al., 2019) and its variants, leading to several BERTology analyses (Rogers et al., 2021). Probing has provided evidence of the existence of syntactic information within BERT representations (Tenney et al., 2019b; Lin et al., 2019; Liu et al., 2019), from which even full parse trees can be recovered with good precision (Hewitt & Manning, 2019). Additionally, some studies have analyzed where syntactic information is stored across the residual stream suggesting a hierarchical encoding of language information, with part-of-speech, constituents, and dependencies being represented earlier in the network than semantic roles and coreferents, matching traditional handcrafted NLP pipelines (Tenney et al., 2019a). Rogers et al. (2021) summarizes results on BERT in detail. Importantly, highly accurate probes indicate a correlation between input representations and labels, but do not provide evidence that the model is using the encoded information for its predictions (Hupkes et al., 2018; Belinkov & Glass, 2019; Elazar et al., 2021).

2 Linear Representation Hypothesis and Sparse Autoencoders

The linear representation hypothesis states that features are encoded as linear subspaces of the representation space (see Park et al. (2023a) for a formal discussion). Mikolov et al. (2013) were the first to show that Word2Vec word embeddings capture linear syntactic/semantic word relationships. For example, adding the difference between word representations of “Spain” and “Madrid”, f(‘‘Spain")−f(‘‘Madrid")f(``\text{Spain}")-f(``\text{Madrid}"), to the “France” representation, f(“France”)f(\text{``France''}), would result in a vector close to f(“Paris”)f(\text{``Paris''}). This presumes that the vector f(‘‘Spain")−f(‘‘Madrid")f(``\text{Spain}")-f(``\text{Madrid}") can be considered as the direction of the abstract capital_of feature. Instances of interpretable neurons (Radford et al., 2017; Voita et al., 2023; Bau et al., 2020), i.e. neurons that fire consistently for specific input features (either monosemantic or polysemantic), also exemplify features represented as directions in the neuron space. Recent work suggests the linearity of concepts in representation space is largely driven by the next-word-prediction training objective and inductive biases in gradient descent optimization (Jiang et al., 2024).

As mentioned in Section 4.1, a fundamental problem of probing lies in its correlational, rather than causal, nature. Recent work (Nanda et al., 2023b; Zou et al., 2023) shows the effectiveness of linear interventions on language models using directions identified by a probe. For instance, adding negative multiples of the sentiment direction (u{\bm{u}}) to the residual stream, i.e. xl′←xl−αu{{\bm{x}}^{l^{\prime}}\leftarrow{\bm{x}}^{l}-\alpha{\bm{u}}}, is sufficient to generate a text matching the opposite sentiment label (Tigges et al., 2023). This simple procedure is named activation addition (Turner et al., 2023). Other unsupervised methods for computing features directions include Principal Component Analysis (Tigges et al., 2023), K-Means (Zou et al., 2023), or difference-in-means (Marks & Tegmark, 2023). For instance, Arditi et al. (2024) use the difference-in-means vector between residual streams on harmful and harmless instructions to find a “refusal direction” in LMs with safety fine-tuning (Bai et al., 2022). Projecting out this direction from every model component output, i.e. fc′←fc−fcu⊺u{f^{c^{\prime}}\leftarrow f^{c}-f^{c}{\bm{u}}^{\intercal}{\bm{u}}}, leads to bypass refusal. Recent studies set distributed alignment search (Section 3.2.3) as the best performing method for causal intervention across mathematical reasoning and linguistic plausibility benchmarks (Tigges et al., 2023; Arora et al., 2024; Huang et al., 2024a), and leveraged it for efficient inference-time interventions aimed at improving task-specific model performance (Wu et al., 2024b). Finally, the MiMic framework (Singh et al., 2024c) was recently proposed to craft optimal steering vectors, exploiting insights from linear erasure methods and class labels from the data distribution. We note that the effectiveness of steering approaches involving linear interventions was recently observed to extend to non-Transformer LMs (Paulo et al., 2024).

A representation produced by a model layer is a vector that lies in a dd-dimensional space. Neurons are the special subset of representation units right after an element-wise non-linearity (Section 2.1.3). Although previous work has identified neurons in models corresponding to interpretable features, in most cases they respond to apparently unrelated inputs, i.e. they are polysemantic. Two main reasons can explain polysemanticity. Firstly, features can be represented as linear combinations of the standard basis vectors of the neuron space (Figure 9 left (a)), not corresponding to the basis elements themselves. Therefore, each feature is represented across many individual neurons, which is known as distributed representations (Smolensky, 1986; Olah, 2023). Secondly, given the extensive capabilities and long-tail knowledge demonstrated by large language models, it has been hypothesized that models could encode more features than they have dimensions, a phenomenon called superposition (Figure 9 left (c)) (Arora et al., 2018; Olah et al., 2020b). Elhage et al. (2022b) showed on toy models trained on synthetic datasets that superposition happens when forcing sparsity on features, i.e. making them less frequent on the training data. Recently, Gurnee et al. (2023) have provided evidence of superposition in the early layers of a Transformer language model, using sparse linear probes.

[yshift=0em]above, left, label belowfeat_actSAE feature activations h(z)h({\bm{z}}) on language models’ representations with a loss defined as

The goal of SAEs is to learn sparse reconstructions of representations. To assess the quality of a trained SAE in achieving this it is common to compute the Pareto frontier of two metrics on an evaluation set (Bricken et al., 2023). These metrics are:

The loss recovered, which reflects the percentage of the original cross-entropy loss of the LM across a dataset when substituting the original representations with the SAE reconstructions.

A summary statistic proposed by Bricken et al. (2023) is the feature density histogram. Feature density is the proportion of tokens in a dataset where a SAE feature has a non-zero value. By looking at the distribution of feature densities we can distinguish if the SAE learnt features that are too dense (activate too often) or too sparse (activate too rarely). Finally, the degree of interpretability of sparse features can be estimated based on their direct logit attribution and maximally activating examples (see Figure 10 left, we introduce these concepts in Section 4.3). This process can be done manually or automated, using a LLM to produce natural language explanations of SAE features.

The sparsity penalty used in SAE training promotes smaller feature activations, biasing the reconstruction process towards smaller norms. This phenomenon is known as shrinkage (Tibshirani, 1996; Wright & Sharkey, 2024). Rajamanoharan et al. (2024) address this issue by proposing Gated Sparse Autoencoders (GSAEs) and a complementary loss function. GSAE is inspired by Gated Linear Units (Dauphin et al., 2017; Shazeer, 2020), which employ a gated ReLU encoder to decouple feature magnitude estimation from feature detection (Figure 10 right):

3 Decoding in Vocabulary Space

The model engages with the vocabulary in two primary ways: firstly, through a set of input tokens facilitated by the embedding matrix WE{\bm{W}}_{E}, and secondly, by interacting with the output space via the unembedding matrix WU{\bm{W}}_{U}. Hence, and due to its interpretable nature, a sensible way to approach decoding the information within models’ representations is via vocabulary tokens.

The logit lens (nostalgebraist, 2020) proposes projecting intermediate residual stream states xl{\bm{x}}^{l} by WU{\bm{W}}_{U}. The logit lens can also be interpreted as the prediction the model would do if skipping all later layers, and can be used to analyze how the model refines the prediction throughout the forward pass (Jastrzębski et al., 2018). This technique has proven effective in analyzing encoder representations in encoder-decoder models Langedijk et al. (2023). However, the logit lens can fail to elicit plausible predictions in some particular models Belrose et al. (2023a). This phenomenon have inspired researchers to train translators, which are functions applied to the intermediate representations prior to the unembedding projection. Din et al. (2023) suggest using linear mappings, while Belrose et al. (2023a) propose affine transformations (tuned lens). Translators have also been trained on the outputs of attention heads, resulting in the attention lens (Sakarvadia et al., 2023). More generally, we can also think of WU{\bm{W}}_{U} as the weights learned by a probe whose classes are the subwords in the vocabulary (Section 4.1), and inspect at any point in the network the amount of information encoded about any subword.

Cancedda (2024) proposes an extension of the logit lens, the logit spectroscopy, which allows a fine-grained decoding of the information of internal representations via the unembedding matrix (Wu{\bm{W}}_{u}). Logit spectroscopy considers splitting the right singular matrix of Wu{\bm{W}}_{u} into NN bands: {Vu,1⊺,…,Vu,N⊺}\{{\bm{V}}_{u,1}^{\intercal},\ldots,{\bm{V}}_{u,N}^{\intercal}\}, where Vu,1⊺{\bm{V}}_{u,1}^{\intercal} and Vu,N⊺{\bm{V}}_{u,N}^{\intercal} each contain a set of singular vectors, the former associated with the largest singular values and the latter with the lowest. If we consider the concatenation of matrices associated with different bands, e.g. from the jj-th to the kk-th band, we form a matrix Vu,j:k⊺{\bm{V}}_{u,j:k}^{\intercal} whose rows span a linear subspace of the vocabulary space. We can use the operator Φu,j:k=Vu,j:kVu,j:k⊺\Phi_{u,j:k}={\bm{V}}_{u,j:k}{\bm{V}}_{u,j:k}^{\intercal} to evaluate the orthogonal projection zΦu,j:k{\bm{z}}\Phi_{u,j:k} of representations z{\bm{z}} onto different subspaces. Alternatively, we can suppress the projection from the representation, i.e. z′←z−zΦu,j:k{{\bm{z}}^{\prime}\leftarrow{\bm{z}}-{\bm{z}}\Phi_{u,j:k}}, leaving its orthogonal component with respect to the subspace. Similarly, bands of singular vectors of the embedding matrix can be considered in the analysis.

The features encoded in model neurons or representation units have been largely studied by considering the inputs that maximally activate them (Zhou et al., 2015; Zeiler & Fergus, 2014). In image models this can be done either by generating synthesized inputs (Nguyen et al., 2016), e.g. via gradient descent (Simonyan et al., 2014), or by selecting examples from an existing dataset. The latter approach has been used in language models to explain the features that units (Dalvi et al., 2019) and neurons (Nanda, 2022b) respond to. However, Bolukbasi et al. (2021) warn that just relying on maximum activating dataset examples can result in “interpretability illusions”, as different activation ranges may lead to varying interpretations. Maximally-activating inputs can produce out-of-distribution behaviors, and were recently employed to craft jailbreak attacks aimed at eliciting unacceptable model predictions (Chowdhury et al., 2024), for example by crafting maximally-inappropriate inputs for red-teaming purposes (Wichers et al., 2024).

Modern LMs can be prompted to provide plausible-sounding justifications for their own or other LMs’ predictions. This can be seen as an edge case of information decoding in which the predictor itself is used as a zero-shot explainer. A notable example is the work by Bills et al. (2023) where GPT-4 is prompted to describe shared features in sets of examples producing high activations for specific neurons across GPT-2 XL. Subsequent work by Huang et al. (2023) shows that neurons identified by Bills et al. (2023) do not have a causal influence over the concepts highlighted in the generated explanation, underscoring a lack of faithfulness in such approach. Additional investigations in the consistency between input attribution and self-explanations in language models highlighted the tendency of LMs to produce explanations that are very plausible according to human intuition, but unfaithful to model inner workings (Atanasova et al., 2023; Parcalabescu & Frank, 2023; Turpin et al., 2023; Lanham et al., 2023; Madsen et al., 2024; Agarwal et al., 2024).

Discovered Inner Behaviors

The techniques presented in Sections 3 and 4 have equipped us with essential tools to understand the behavior of language models. In the following sections, we provide an overview of the internal mechanisms that have been discovered within Transformer LMs.

As seen in Section 2.1.2, each attention head consists of a QK (query-key) circuit and an OV (output-value) circuit. The QK circuit computes the attention weights, determining the positions that need to be attended, while the OV circuit moves (and transforms) the information from the attended position into the current residual stream. A substantial body of research has been dedicated to analyzing attention weights patterns formed by QK circuits (Clark et al., 2019; Kovaleva et al., 2019; Voita et al., 2019b), fueling a debate on whether these weights serve as explanations (Bibal et al., 2022). However, our understanding of the specific features encoded in the subspaces employed by circuit operations is still limited. Here, we categorize known behavior of attention heads in two groups: those having intelligible attention patterns, and those with meaningful QK and OV circuits.

Clark et al. (2019) showed some BERT heads attend mostly to specific positions relative to the token processed. Specifically, attention heads that attend to the token itself, to the previous token, or to the next position. A similar pattern is also observed in encoders of neural machine translation models (Voita et al., 2019b; Raganato & Tiedemann, 2018). Previous token heads are an essential part of induction heads, and have been shown necessary for circuits in GPT2-Small (Wang et al., 2023a). Their main role has been associated with copying previous token information to the following residual stream, such as concatenating two-tokens names (Nanda et al., 2023c). Ferrando & Voita (2024) show previous token heads are important across several textual domains.

First discovered in machine translation encoders Correia et al. (2019), subword joiner heads have been observed as well in large language models (Ferrando & Voita, 2024). These heads attend exclusively to previous tokens that are subwords belonging to the same word as the currently processed token.

Some attention heads attend to tokens having syntactic roles with respect to the processed token significantly more than a random baseline (Clark et al., 2019; Htut et al., 2019). Particularly, certain heads specialize in given dependency relation types such as obj, nsubj, advmod, and amod. Chen et al. (2024a) show these heads appear suddenly during the training process of masked language models playing a crucial role in the subsequent development of linguistic abilities.

Duplicate token heads attend to previous occurrences of the same token in the context of the current token. Wang et al. (2023a) hypothesize that, in the IOI task (Section 5.4), these heads copy the position of the previous occurrence to the current position.

1.2 Attention heads with interpretable QK and OV circuits

Several attention heads in Transformer LMs have OV matrices that exhibit copying behavior. Elhage et al. (2021a) propose using the number of positive real eigenvalues of the full OV circuit matrix WEWOVWU{\bm{W}}_{E}{\bm{W}}_{OV}{\bm{W}}_{U} as a summary statistic for detecting copying heads. Positive eigenvalues mean that there exists a linear combination of tokens contributing to an increase in the linear combination of logits of the same tokens.

An induction mechanism (Figure 12 left) that allows language models to complete patterns was discovered first by Elhage et al. (2021a) and further studied by Olsson et al. (2022).We follow the mechanistic formulation by Elhage et al. (2021a). See (Variengien, 2023) for a discussion. This mechanism involves two heads in different layers composing together. Specifically, a previous token head (PTH) and an induction head. The induction mechanism learns to increase the likelihood of token B given the sequence A B … A, irrespective of what A and B are. To do so, a PTH in an early layer copies information from the first instance of token A to the residual stream of B, specifically by writing in the subspace the QK circuit of the induction head reads from (K-composition). This makes the induction head at the last position to attend to token B, and subsequently, its copying OV circuit increases the logit score of B. Olsson et al. (2022) demonstrate that the OV and QK circuits of the induction head can perform fuzzy versions of copying and prefix matching, giving rise to generating patterns of the kind A∗\texttt{A}^{*} B∗\texttt{B}^{*} … A →B\rightarrow\texttt{B}, where A and A∗\texttt{A}^{*}, and B and B∗\texttt{B}^{*} are semantically related (e.g. the same words in different languages). Overall, induction heads have been shown to appear broadly in Transformer LMs (Nanda, 2022a), with those operating at an n-gram level being identified as important drivers of in-context learning (Akyürek et al., 2024). Recent work showed that these heads display both complementary and redundant behaviors, likely shaped by competitive dynamics during optimization (see Section 5.4) (Singh et al., 2024a). Relatedly, redundancy was also observed in the connections between early-layer PTHs and subsequent induction heads. Finally, the emergence rate of induction heads is impacted by the diversity of in-context tokens, with higher diversity in attended and copied tokens delaying the formation of the two respective sub-mechanisms (Singh et al., 2024a).

Copy suppression heads, discovered in GPT2-Small (McDougall et al., 2023) reduce the logit score of the token they attend to, only if it appears in the context and the current residual stream is confidently predicting it (Figure 12 right). This mechanism was shown to improve overall model calibration by avoiding naive copying in many contexts (e.g. copying “love” in “All’s fair in love and ___”). The OV circuit of a copy suppression head can copy-suppress almost all of the tokens in the model’s vocabulary when attended to. This behavior is confirmed by analyzing the “effective QK circuit” of GPT2-Small. The key input is the FFN1\text{FFN}^{1} output of every token, and the query input the unembedding of any token, WUWQKFFN1(WE){\bm{W}}_{U}{\bm{W}}_{QK}\text{FFN}^{1}({\bm{W}}_{E}), and shows the diagonal elements rank higher. Copy suppression is also linked to the self-repair mechanism since ablating an essential component deactivates the suppression behavior, compensating for the ablation.

Given an input token belonging to an element in an ordinal sequence (e.g. “one”, “Monday”, or “January”), the ‘effective OV circuit’: FFN1(WE)WOVWU\text{FFN}^{1}({\bm{W}}_{E}){\bm{W}}_{OV}{\bm{W}}_{U} of the successor heads increases the logits of tokens corresponding to the next elements in the sequence (e.g. “two”, “Tuesday”, “February”). Specifically, Gould et al. (2024) show the output of the first FFN block represents a common ‘numerical structure’ on which the successor head acts. Gould et al. (2024) find these heads in Pythia (Biderman et al., 2023), GPT2 (Radford et al., 2019) and Llama 2 (Touvron et al., 2023) models.

1.3 Other noteworthy attention properties

The attention heads previously described serve specific functions in aiding the model to predict the next token. However, the degree of specialization of components across different domains and tasks remains unclear. Ferrando & Voita (2024); Chughtai et al. (2024); Lv et al. (2024) identify some specialized heads that contribute only within specific input domains, such as non-English contexts, coding sequences, or specific topics. An analysis of the top singular vectors of their OV matrices (Section 4.3) reveal these heads mainly promote tokens related to the semantics of the input they participate in.

Early investigations into BERT (Kovaleva et al., 2019) revealed most attention heads exhibit “vertical” attention patterns, mainly focusing on special (CLS, SEP) and punctuation tokens. Clark et al. (2019) hypothesized a head may attend to special tokens when its specialized function is not applicable (no-op hypothesis). Kobayashi et al. (2020) showed the norm of the value vectors (Section 3.1) associated with special tokens, periods, and commas tend to be small, canceling out the effect of large attention weights, thereby supporting the no-op hypothesis. Furthermore, it was shown that attention to the end-of-sequence token in MT models is used to ignore the contribution of the source sentence (Ferrando & Costa-jussà, 2021), useful when predicting some function words such as the particle “off” in “She turned off the lights.”. In auto-regressive LMs, these patterns are observed mainly in the beginning of sentence (BOS) token (Figure 13), although other tokens play the same role (Ferrando & Voita, 2024). According to Xiao et al. (2023), allowing attention mass on the BOS token is necessary for streaming generation, and performance degrades when the BOS is omitted. Using the logit spectroscopy (Section 4.3), Cancedda (2024) finds that early FFNs in Llama 2 write relevant information (for the attention sink mechanism to occur in later layers) into the residual stream of BOS. These FFNs write into the linear subspace spanned by the right singular vectors with the lowest associated singular values of the unembedding matrix. Cancedda (2024) refers to this as a dark subspace due to its low interference with next token prediction, and finds a significant correlation between the average attention received by a token and the existence of these dark signals in its residual stream. These dark signals reveal as massive activation values acting as fixed biases (Sun et al., 2024), a crucial prerequisite for the attention sink mechanism to take place (Puccetti et al., 2022; Bondarenko et al., 2023). On the other hand, specific neurons in the FFN of the layer before the attention head have been found to control the amount to which the tokens attend to BOS Gurnee et al. (2024) (Figure 13).

Sparse autoencoders have been trained on the outputs of the attention layers to better understand the features computed by each head. Results presented in Kissane et al. (2024a) show that, on a two-layer Transformer, a large number of features (76%) are non-dead, with the majority of them being interpretable (82%). Three specific features are studied in detail. The “board by induction” feature promotes the token board, and is present on the output of an induction head (Section 5.1.2), being part of the induction features family. The “in questions starting with Which” feature is instead part of the local context features, promoting the prediction of ? when Which appears in the context. Lastly, “in texts related to pets” is an example of an high-level context feature that activates for almost the entire context, with its related head attending to pet-related context tokens. Notably, Kissane et al. (2024a) detect the presence of non-induction features in the output of induction-heads, providing evidence of attention head polysemanticity, initially observed by Heimersheim & Janiak (2023). Further investigations (Kissane et al., 2024b) reveal the same three feature families appear on GPT2-Small, as well as successor features, name mover features, suppression features and duplicate token features associated with heads matching their respective behaviors. Krzyzanowski et al. (2024) conduct a finer-grained analysis of features in GPT-2 Small attention heads, focusing on the top 10 features in each head and concluding that most heads do multiple tasks, with only around 10% of those being monosemantic. Their findings point out that early layers (0-3) mainly focus on shallow syntactic features, with the following layers encoding increasingly more complex syntactic features. Middle layers (5-6) contain the least interpretable features, while later layers (7-10) encode complex abstract features like time and distance relationships and high-level context concepts. The heads in the last attention block show mostly grammatical adjustments and bigram completions.

2 Feedforward Network Block

The dimensions in the FFN activation space (neurons), following the non-linearity, are more likely to be independently meaningful (Section 2.1.3), and have therefore been the object of study of recent interpretability works.

The behavior of neurons in language models has been extensively studied, with examinations focusing on either their input or output behavior. In the context of input behavior analysis, Voita et al. (2023) show neurons firing exclusively on specific position ranges. Other discoveries include skills neurons, whose activations are correlated with the task of the input prompt (Wang et al., 2022), concept-specific neurons (Suau et al., 2020; 2022; Gurnee et al., 2023) whose response can be used to predict the presence of a concept in the provided context, such as whether it is Python code, French (Gurnee et al., 2023), or German (Quirke et al., 2023) language. Neurons responding to other linguistic and grammatical features have also been found (Bau et al., 2019; Durrani et al., 2023).

Regarding the output behavior of neurons, Dai et al. (2022) use the Integrated Gradients method (Section 3.1) to attribute next-word facts predictions to FFNs neurons, finding knowledge neurons. The key-value memory perspective of FFNs (Section 2.1.3) offers a way to understand neuron’s weights. Specifically, using the direct logit attribution method (Section 3.2.1) we can measure the neuron’s effect on the logits. Geva et al. (2022b) show that some neurons promote the prediction of tokens associated with particular semantic and syntactic concepts. Ferrando et al. (2023) illustrate that a small set of neurons in later layers is responsible for making linguistically acceptable predictions, such as predicting the correct number of the verb, in agreement with the subject. Gurnee & Tegmark (2024) find neurons that interact with directions in the residual stream that are similar to the space and time feature directions extracted from probes. Tang et al. (2024) show language-specific neurons are key for multilingual generation, demonstrating one can steer the model output’s language by causally intervening on them. Finally, neurons suppressing improbable continuations, e.g. the repetition of the last token in the sequence, have recently been identified (Voita et al., 2023; Gurnee et al., 2024).

Recent work highlighted the presence of polysemantic neurons within language models. Notably, most early layer neurons specialize in sets of n-grams, functioning as n-gram detectors (Voita et al., 2023), with the majority of neurons firing on a large number of n-grams. Gurnee et al. (2023) suggest superposition appears in these early layers, and via sparse probing they find sparse combinations of neurons whose added activation values disentangle the detection of specific n-grams, such as the compound word “social security” from other bigrams containing only one of the two terms. Even though polysemanticity and superposition arise in early layers, several dead neurons were observed in OPT modelsOPT models use ReLU activation functions, allowing for zero activation values (Zhang et al., 2022). (Voita et al., 2023). Furthermore, Elhage et al. (2022a) hypothesize models internally perform “de-/re-tokenization”, where neurons in early layers respond to multi-token words or compound words (Elhage et al., 2022a), mapping tokens to a more semantically meaningful representation (detokenization). In contrast, in the latest layers, neurons aggregate contextual representations back into single tokens (re-tokenization) to produce the next-token prediction.

Whether different models learn similar features remains an open question (Olah et al., 2020b). For instance, various computer vision models were found to learn Gabor filters in early layers (Olah et al., 2020a). In a recent study, Gurnee et al. (2024) investigated whether neurons respond to features similarly across different models. Their analysis used the pairwise correlation of neuron activations across GPT2 models trained from different random initializations as a proxy measure, revealing a subset of 1-5% of neurons activating on the same inputs. As expected, within the cluster of universal neurons there is a higher degree of monosemanticity. This group includes alphabet neurons, which activate in response to tokens representing individual letters and on tokens that start with the letter, supporting the re-tokenization hypothesis. Additionally, there are previous token neurons that fire based on the preceding token, as well as unigram, position, semantic, and syntax neurons. In terms of output behavior, universal neurons include attention (de-)activation neurons, responsible for controlling the amount of attention given to the BOS token by a subsequent attention head, and thus setting it as a no-op (Section 5.1.3). Lastly, Gurnee et al. (2024) hypothesize that some neurons act as entropy neurons, modulating the model’s uncertainty over the next token prediction.

It has been suggested that the overall arrangement of neurons in language models mirrors that of neuroscience (Elhage et al., 2022a). Early layer neurons exhibit similarities to sensory neurons, responding to shallow patterns of the input, mostly focusing on n-grams. Moving into the middle layers, activation tends to occur around more high-level concepts (Bricken et al., 2023; Gurnee et al., 2023). An example of this is the neuron identified in Elhage et al. (2022a), which represents numbers only when they refer to the amount of people. Finally, later layers’ neurons bear a resemblance to motor neurons in the sense that they produce changes in the distribution of the next-token prediction, either by promoting or suppressing sets of tokens.

SAEs are able to identify significantly more interpretable features than the model’s neurons themselves (Bricken et al., 2023), as noted both by human and automated analyses in one-layer transformers. The features detected by SAEs trained to reconstruct FFN activations (Bricken et al., 2023) appear to split into increasingly more fine-grained distinctions of the feature as more dimensions (dictionary entries) are added, demonstrating that 512 neurons can encode tens of thousands of features. Examples of features found by Bricken et al. (2023) include those firing in the presence of Arabic or Hebrew scripts and promoting tokens in those scripts, and features responding to DNA sequences or base64 strings.

3 Residual Stream

We can think of the residual stream as the main communication channel in a Transformer. The “direct path” (Section 2.2) connecting the input embedding with the unembedding matrix, xWU{\bm{x}}{\bm{W}}_{U} does not move information between positions, and mainly models bigram statistics (Elhage et al., 2021a), while the latest biases in the network, localized in the prediction head, are shown to shift predictions according to word frequency, promoting high-frequency tokens (Kobayashi et al., 2023). However, alternative paths involve the interaction between components, which write into linear subspaces (Elhage et al., 2021a) that can be read by downstream components, or directly by the prediction head, potentially doing more complex computations. Heimersheim & Turner (2023) observed that the norm of the residual stream grows exponentially along the layers over the forward pass of multiple Transformer LMs (Millidge & Winsor, 2023; Merrill et al., 2021). A similar growth rate appears in the norm of the output matrices writing into the residual stream, WO{\bm{W}}_{O} and Wout{\bm{W}}_{\text{out}}, unlike input matrices (WQ{\bm{W}}_{Q}, WK{\bm{W}}_{K}, WV{\bm{W}}_{V} and Win{\bm{W}}_{\text{in}}), which maintain constant norms along the layers. It is hypothesized that some components perform memory management to remove information stored in the residual stream. For instance, there are attention heads with OV matrices with negative eigenvalues attending to the current position, and FFN neurons whose input and output weights have large negative cosine similarity (Elhage et al., 2021a), meaning that they write a vector (FFN value) on the opposite direction to the direction they read from (FFN key). Notably, Gurnee et al. (2024) find that these neurons activate very frequently. Dao et al. (2023) evaluate a small Transformer LM and provide convincing evidence of multiple attention heads removing the information written by a first layer head.

Outlier dimensions (Kovaleva et al., 2021; Luo et al., 2021) have been identified within the residual stream. These rogue dimensions exhibit large magnitudes relative to others and are associated with the generation of anisotropic representations (Ethayarajh, 2019; Timkey & van Schijndel, 2021). Anisotropy means that the residual stream states of random pairs of tokens tend to point towards the same direction, i.e. the expected cosine similarity is close to one. Furthermore, ablating outlier dimensions has been shown to significantly decrease downstream performance (Kovaleva et al., 2021), suggesting they encode task-specific knowledge (Rudman et al., 2023). The magnitudes of these outliers have been shown to increase with model size (Dettmers et al., 2022), posing challenges for the quantization of large language models. The presence of rogue dimensions has been hypothesized to stem from optimizer choices (Elhage et al., 2023), with higher levels of regularization reducing their magnitudes (Ahmadian et al., 2023). Puccetti et al. (2022) identified a high correlation between the magnitude of the outlier dimensions found in token representations and their training frequency. They concluded that these dimensions contribute to enabling the model to focus on special tokens, which is known to be associated with “no-op” attention updates (Bondarenko et al., 2023) (see attention sinks in Section 5.1.3). In Vision Transformers, high-norm residual stream states have been identified as aggregators of global image information, appearing in patches with highly redundant information, such as those composing the image background (Darcet et al., 2024).

The specific features encoded within the residual stream at various layers remain uncertain, yet sparse autoencoders offer a promising avenue for improving our understanding. Recently, SAEs have been trained to reconstruct residual stream states in small language models such as GPT2-Small (Cunningham et al., 2023; Bloom & Lin, 2024; Bloom, 2024) showing highly interpretable features (Figure 14). Since residual stream states gather information about the sum of previous components’ outputs, inspecting SAE’s features can illuminate the process by which they are added or transformed during the forward pass. Given the type of features intermediate FFNs and attention heads interact with, we also expect the residual stream at middle layers to encode highly abstract features. Tigges et al. (2023) provide some preliminary evidence by showing that causally intervening on the residual stream in middle layers is more effective in flipping the sentiment of the output token, suggesting that the latent representation of sentiment is most prominent in the middle layers. Bloom & Lin (2024) study the features learned by a SAE in layer 8 of the 12-layer model GPT2-Small. Based on their output behavior via the logit lens (Section 4.3) the authors first find local context features promoting small sets of tokens. Secondly, they highlight the presence of partition features, which promote and suppress two distinct sets of tokens. For instance, a partition feature might promote tokens starting with capital letters and suppress those starting with lowercase letters. Finally, akin to suppression neurons (Voita et al., 2023; Gurnee et al., 2024), they note the presence of suppression features aimed at reducing the likelihood of specific sets of tokens. In line with these findings, recent studies have shown that language models create vectors representing functions or tasks given in-context examples (Hendel et al., 2023; Todd et al., 2024), which are found in intermediate layers. In the next section, we provide a deeper overview of the interaction between different components and the resulting behavior that emerges.

4 Emergent Multi-component Behaviors

In previous sections we presented some of the different mechanisms that attention heads and FFNs implement, as well as an overview of the properties of the residual stream. However, in order to explain the remarkable performance of Transformers, we also need to account for the interactions between the different components (Wen et al., 2023; Cammarata et al., 2020).

The induction mechanism presented in Section 5.1.2 is a clear example of two components (attention heads) composing together to complete a pattern. Recent evidence suggests that multiple attention heads work together to create “function” or “task” vectors describing the task when given in-context examples (Hendel et al., 2023; Todd et al., 2024). Intervening in the residual stream with those vectors can produce outputs in accordance with the encoded task on novel zero-shot prompts. Variengien & Winsor (2023) study in-context retrieval tasks involving answering a request where the answer can be found in the context. The authors identify a high-level mechanism that is universal across subtasks and models. Specifically, middle layers process the request, followed by a retrieval step of the entity from the context done by attention heads at later layers.

Additionally, Neo et al. (2024); Yu & Ananiadou (2024) reveal that individual neurons within downstream FFNs activate according to the output of previous attention heads, interacting in specific contexts. However, the most compelling evidence of particular behaviors emerging from the interaction between multiple components is found in the circuit analysis literature (Wang et al. (2023a); Stolfo et al. (2023b); Heimersheim & Janiak (2023); Geva et al. (2023); Hanna et al. (2023), among others). As an illustration, we present the circuit found in GPT2 Small for the Indirect Object Identification (IOI) task (Wang et al., 2023a), depicted in Figure 15. In the IOI task the model is given inputs of the type “When Mary and John went to the store, John gave a drink to ___”. The initial clause introduces two names (Mary and John), followed by a secondary clause where the two people exchange an item. The correct prediction is the name not appearing in the second clause, referred to as the Indirect Object (Mary). The circuit found in GPT2 Small mainly includes:

Duplicity signaling: duplicate token heads at position S2, and an induction mechanism involving previous token heads at S1+1 signal the duplicity of S (John). This information is read by S-Inhibition heads at the last position, which write in the residual stream a token signal, indicating that S is repeated, and a position signal of the S1 token.

Name copying: name mover heads in later layers copy information from names they attend to in the context to the last residual stream. However, the signals of the previous layers S-Inhibition heads modify the query of name mover heads so that the duplicated name (in S1 and S2) is less attended, favouring the copying of the Indirect Object (IO) and therefore, pushing its prediction.

Besides, Wang et al. (2023a) discovered Negative mover heads, which are instances of copy suppression heads (Section 5.1.2) downweighing the probability of the IO. While the IOI is an attention-centric circuit, examples of circuits involving both FFNs and attention heads are also present. For instance, Hanna et al. (2023) reverse-engineered the GPT2-Small circuit for the greater-than task, which involves sentences like The war lasted from the year 1814 to the year 18__, where the model must predict a year greater than 1814. The authors demonstrate that downstream FFNs compute a valid year by reading from previous attention heads, which attend to the event’s initial date.

Prakash et al. (2024) show that the functionality of the circuit components remains consistent after fine-tuning and benefits of fine-tuning are largely derived from an improved ability of circuit components to encode important task-relevant information rather than an overall functional rearrangement. Fine-tuned activations are also found to be compatible with the base model despite no explicit tuning constraints, suggesting the process produces minimal changes in the overall representation space. The findings of Prakash et al. (2024) are additionally supported by Jain et al. (2024) in controlled settings. While a common critique of mechanistic interpretability work is the limited scope of identified circuits, Merullo et al. (2024) show that low-level findings about specific heads and higher-level findings about general algorithms implemented by Transformer models can generalize across tasks, suggesting that large language models could be explained as functions of few task-general sparse components. The results of Merullo et al. (2024) also suggest that circuits are not exclusive, i.e. the same model components might be part of several circuits. Other studied dimensions of discovered circuits include their faithfulness (Hanna et al., 2024) and their completeness (Wang et al., 2023a).

Transformer models were observed to converge to different algorithmic solutions for tasks at hand (Zhong et al., 2023). Nanda et al. (2023a) provide convincing evidence on the relation between circuit emergence and grokking, i.e. the sudden emergence of near-perfect generalization capabilities for simple symbol manipulation tasks at late stages of model training (Power et al., 2022). Merrill et al. (2023) suggest the grokking phase transition can be seen as the emergence of a sparse circuit with generalization capabilities, replacing a dense subnetwork with low generalization capacity. According to Varma et al. (2023), this happens because dense memorizing circuits are inefficient for compressing large datasets. In contrast, generalizing circuits have a larger fixed cost but better per-example efficiency, hence being preferred in large-scale training. Huang et al. (2024b) connect the learning dynamic converging to grokking to the double descent phenomenon (Loog et al., 2020). According to this view, the emergence of specialized attention heads might be seen as a mild grokking-related phenomenon (Olsson et al., 2022; Bietti et al., 2023).

4.1 Factuality and hallucinations in model predictions

The generation of factually incorrect or nonsensical outputs is considered a significant limitation in the practical usage of language models (Ji et al., 2023; Minaee et al., 2024). While some techniques for detecting hallucinated content rely on quantifying the uncertainty of model predictions (Varshney et al., 2023), most alternative approaches engage with model internal representations. Approaches for detecting hallucinations directly from the representations include training probes and analyzing the properties of the representations leading to hallucinations. CH-Wang et al. (2023) and Azaria & Mitchell (2023) find probing classifiers predictive of the model’s output truthfulness, achieving the highest accuracy using middle and last layers representations. Zou et al. (2023) and Li et al. (2023a) find “truthfulness” directions with causal influence on the model outputs, i.e. intervening in the internal representations with the found directions enhance the output truthfulness. Li et al. (2023a) locate these causal directions in the specific attention head activation. Chen et al. (2024b) use the eigenvalues of responses’ representations covariance matrix to measure the semantic consistency in embedding space across layers, while Chen et al. (2024d) observe that logit lens (Section 4.3) scores of the predicted attribute (answer) in higher-layers representations of context tokens are informative of the answer correctness.

A related area of research with overlapping goals is that of hallucination detection in machine translation (MT). An MT model is considered to hallucinate if its output contains partially or fully detached content from the source sentence (Guerreiro et al., 2023b). Prediction probabilities of the generated sequence and attention distributions have been used to detect potential errors (Fomicheva et al., 2020) and model hallucinations (Guerreiro et al., 2023a; b). Recently, methods measuring the amount of contribution from the source sentence tokens (Ferrando et al., 2022a) were found to perform on par with external methods based on semantic similarity across several categories of model hallucinations (Dale et al., 2023a; b). Detection methods show complementary performance across hallucination categories, and simple aggregation strategies for internals-based detectors outperform methods relying on external semantic similarity or quality estimation modules (Himmi et al., 2024).

The underlying mechanisms involved in the prediction of hallucinated content for LLMs remain largely unexplained. Most of the research in this area focuses on studying the ability of language models to recall facts, which we discuss in the next section.

Recent research has delved into the internal mechanisms through which language models recall factual information, which is directly related to the hallucination problem in LLMs. A common methodology involves studying tuples (s,r,a)(s,r,a), where ss is a subject, rr a relation, and aa an attribute. The model is prompted to predict the attribute given the subject and relation. For instance, given the prompt: “LeBron James plays the sport of”, the model is expected to predict basketball. Meng et al. (2022) and Geva et al. (2023) make use of causal interventions (Section 3.2.2) to localize a mechanism responsible for recalling factual knowledge within the language model. Early-middle FFNs located in the last subject token add information about the subject into its residual stream. On the other hand, information from the relation passes into the last token residual stream via early attention heads. Finally, later layers attention heads extract the right attribute from the last subject residual stream. Yuksekgonul et al. (2024) find that, in similar settings, attention to relevant tokens in the prompt correlates with LLM’s factual correctness. Importantly, the division of responsibilities between lower and upper layers was also observed in attention-less models based on the Mamba architecture (Gu & Dao, 2023; Sharma et al., 2024a). While this might be motivated by implicit context-mixing akin to Transformers’ causal self-attention (Ali et al., 2024), it suggests the organization of these mechanisms might be driven by the language modeling optimization process rather than architectural constraints.

Subsequent research has moved from localizing model behavior to studying the computations performed to solve this task. Hernandez et al. (2024) show that attributes of entities can be linearly decoded from the enriched subject residual stream, while Chughtai et al. (2024) investigate how attention heads’ OV circuits effectively decode the attributes, proposing an additive mechanism. More precisely, using the direct logit attribution by each token via the attention head (Equation 16) they identify subject heads responsible for extracting attributes from the subject independently from the relation (not attending to it), as well as relation heads that promote attributes without being causally dependent on the subject. Additionally, a group of mixed heads generally favor the correct attribute and depend on both the subject and relation. The combination of the different heads’ outputs, each proposing different sets of attributes, together with the action of some downstream FFNs resolve the correct prediction (Figure 16). Nanda et al. (2023c) provide a detailed explanation of the subject enrichment phase by studying names of athletes as subjects. They suggest that the first layers’ attention heads concatenate the athlete’s name on the final name token residual stream through addition, and subsequent FFNs map the obtained athlete’s name representation into a linear representation of the athlete’s sport that can be easily linearly extracted by the downstream attribute extraction heads.

Merullo et al. (2023) report that for solving relational tasks, such as predicting a country’s capital given in-context examples, middle layers prepare the argument, e.g. Poland, of a get_capital()\texttt{get\_capital}() function that is applied downstream via an FFN update, giving place to get_capital(Poland)=Warsaw{\texttt{get\_capital}(\text{Poland})=\text{Warsaw}}. Further research replicates Merullo et al. (2023)’s analysis on zero-shot settings (Lv et al., 2024) and finds specific attention heads “passing” the argument from the context (Poland), but also promoting the capital cities (Warsaw). Downstream FFNs “activate” relevant attention heads in the previous layer and add a vector guiding the residual stream toward the correct capital direction.

Recent works aim to shed light on how the model engages in factual recall vs. grounding. Following the aforementioned (subject, relation, attribute) structure of facts, an answer is considered to be grounded if the attribute is consistent with the information in the context of the prompt. Given prompts of the type “The capital of Poland is London. Q: What is the capital of Poland? A:___”, Yu et al. (2023a) find in-context heads and memory heads by using the difference logit attribution (Section 3.2.1, Equation 17) of attention heads. These heads favor, respectively, the in-context answer London and the memorized answer Warsaw, showing a “competition” between mechanisms (Ortu et al., 2024). Furthermore, upweighting the output of each head type reveals a bias towards one of the two answers. Similar to the in-context heads, Variengien & Winsor (2023) show that a set of downstream attention heads retrieve the correct answer (an attribute) from the context via copying, preceded by a processing of the request (a question) in middle layers. Wu et al. (2024a) study these type of heads, which they coin retrieval heads in arbitrarily long-contexts, and show they are crucial for solving the Needle-in-a-Haystack tests (Kamradt, 2023). Monea et al. (2024) complement the findings of Yu et al. (2023a) and Meng et al. (2022) and show that FFNs in the last token of the subject have higher contributions on ungrounded (memorized) answers as opposed to grounded answers, while suggesting that grounding could be a more distributed process lacking a specific localization. Haviv et al. (2023) show that the recall of “memorized” idioms largely depends on the updates of the FFNs in early layers, providing further evidence of their role as a storage of memorized information. This is further observed in the study of memorized paragraphs, with lower layers exhibiting larger gradient flow (Stoehr et al., 2024). On the other hand, Sharma et al. (2024b) show that substituting the original FFN matrices by lower-rank approximations (Section 4.3) leads to improvements in model performance, especially in later layers of the model. They show that, in the factual recall task, the components with smaller singular values encode the correct semantic type of the answer but the wrong answer, thus their removal benefits the accuracy. To conclude, we draw a connection with a decoding strategy (DoLa) proposed to improve the factuality of language models (Chuang et al., 2024). DoLa contrastively compares the logit-lens next-token distributions between an early layer and a later layer (Li et al., 2023b), promoting tokens that undergo a larger probability change, suggesting that the factual knowledge injection is done in a distributed manner across the network.

Factual information encoded in LMs might be incorrect from the start, or become obsolete over time. Moreover, inconsistencies have been observed when recalling factual knowledge in multilingual and cross-lingual settings (Fierro & Søgaard, 2022; Qi et al., 2023), or when factual associations are elicited using less common formulations (Berglund et al., 2023). This sparked the interest in developing model editing approaches able to perform targeted updates on model factual associations with minimal impact on other capabilities. While early approaches proposed edits based on external modules trained for knowledge editing (De Cao et al., 2021; Mitchell et al., 2022a; b), recent methods employ causal interventions (Section 3.2.2) to localize knowledge neurons (Dai et al., 2022) and FFNs in one or more layers (Meng et al., 2022; 2023), informed by factual recall mechanisms described in the previous paragraph. However, model editing approaches still present several challenges, summarized in (Yao et al., 2023; Li et al., 2024), including the risks of catastrophic forgetting (Gupta et al., 2024a; b) and downstream performance loss (Gu et al., 2024). Importantly, Hase et al. (2023) show that effective localization does not always result in improved editing results, and that distributed edits across different model sections can result in similar editing accuracy. Steerable-by-design architectures such as the Backpack Transformer (Hewitt et al., 2023) were recently proposed as possible alternatives to localization-driven methods, exploiting the linearity of component contributions (Section 4.2) as an inductive bias to enhance controllability. We refer readers to Wang et al. (2023b) for further insights on LM editing.

LM Interpretability Tools

Several open-source software libraries were introduced to facilitate interpretability studies on Transformer-based LMs. In this section, we briefly summarize the most notable ones and highlight their main points of strength.

Captum (Kokhlikyan et al., 2020) is a library in the Pytorch ecosystem providing access to several gradient and perturbation-based input attribution methods for any Pytorch-based model. It notably supports training data attribution methods (Section 3.1), and recently added several utilities for simplifying attribution analyses of generative LMs (Miglani et al., 2023). Several Captum-based tools provide convenient APIs for input attribution of Transformers-based models: Transformers Interpret (Pierse, 2021), ferret (Attanasio et al., 2023) and Ecco (Alammar, 2021) are mainly centered around language classification tasks, while Inseq (Sarti et al., 2023) is focused specifically on generative LMs and supports advanced approaches for contrastive context attribution (Sarti et al., 2024) as well as context mixing evaluation (Section 3.1). SHAP (Lundberg & Lee, 2017) is a popular toolkit mainly centered on perturbation-based input attribution methods and model-agnostic explanations for various data modalities. The Saliency (PAIR Team, 2023) library provides framework-agnostic implementations for mainly gradient-based input attribution methods. LIT (Tenney et al., 2020) is a framework-agnostic tool providing a convenient set of utilities and an intuitive interface for interpretability studies spanning input attribution, concept-based explanations and counterfactual behavior evaluation. It notably includes a visual tool for debugging complex LLM prompts (Tenney et al., 2024).

Tools supporting work on circuit discovery and causal interventions play a fundamental role in mechanistic studies, balancing the complexity and model-specific nature of intervention-based methods with a broad support for various pre-trained LM architectures. TransformerLens (Nanda & Bloom, 2022) is a Pytorch-based toolkit to conduct mechanistic interpretability analyses of generative language models inspired by the closed-source Garçon library (Elhage et al., 2021b). The library reimplements popular Transformer LM architectures, preserving compatibility with the popular transformers library (Wolf et al., 2020) while also providing utilities such as hook points around model activations and attention head decomposition to facilitate custom interventions. NNsight (Fiotto-Kaufman, 2024) provides a Pytorch-compatible interface for interpretability analyses. Its usage is not restricted to Transformer models, but it provides utilities to streamline the usage of transformers checkpoints. Its main peculiarity is the ability to compile an intervention graph that can be processed through delayed execution, enabling the extraction of arbitrary internal information from large LMs hosted on remote servers. Pyvene (Wu et al., 2024c) is a Pytorch-based library supporting complex intervention schemes, such as trainable (Geiger et al., 2023a) and mid-training interventions (Geiger et al., 2022), alongside various model categories beyond Transformers. Notably, it supports the serialization of intervention schemes to simplify analyses and promotes reusability. Several tools are currently used for the development of SAEs (Section 4.2), providing overlapping sets of features. For example, SAELens (Bloom & Channin, 2024) supports advanced visualization of SAE features, while dictionary-learning (Marks & Mueller, 2023) is an actively developed tool built on top of NNsight, supporting various experimental features to address SAEs’ weaknesses. Finally, sparse-autoencoder (Cooney, 2023) provides a standard TransformerLens-compatible SAE implementation.

Several tools such as BERTViz (Vig, 2019), exBERT (Hoover et al., 2020) and InterpreT (Lal et al., 2021) were developed to visualize attention weights and activations in Transformers-based LMs. LM-Debugger (Geva et al., 2022a) is a toolkit to inspect intermediate representation updates through the lens of logit attribution (Section 3.2.1), while VISIT (Katz & Belinkov, 2023), Ecco (Alammar, 2021) and Tuned Lenshttps://github.com/AlignmentResearch/tuned-lens (Belrose et al., 2023a) simplify the application of naive and learned vocabulary projections to inspect the evolution of predictions across model layers. CircuitsVis (Cooney, 2022) provides reusable Python bindings for front-end components that can be used to visualize Transformers internals and predictions, and was adopted by various interpretability tools. Penzai (Daniel Johnson, 2024) is a JAX library supporting rich visualizations of pytree data structures, including LM weights and activations. LM-TT (Tufanov et al., 2024) allows inspecting the information flow in a forward pass, faciliting the examination of the contributions of individual attention heads and feed-forward neurons. TDB (Mossing et al., 2024) is a visual interface to interpret neuron activations in LMs supporting automated interpretability techniques and SAEs. Neuronpediahttps://neuronpedia.org (Lin & Bloom, 2024) provides an open repository for visualizing activation of SAE features trained on LM residual stream states (Section 4.2). Notably, it includes a gamified experience to facilitate the annotation of human-interpretable concepts in SAE feature space. Lastly, sae-vis (McDougall & Bloom, 2024) is a SAELens-compatible library to produce feature-centric and prompt-centric interactive visualizations of SAE features.

The “Restricted Access Sequence Processing Language” (RASP, Weiss et al., 2021) is a sequence processing language providing a human-readable model for transformer computations. Tracr (Lindner et al., 2023) is a compiler converting RASP programs into decoder-only Transformer weights, automating the creation of small Transformer models implementing specific desired behaviors. RASP and Tracr were adopted for promoting interpretable behaviors via constrained optimization (Friedman et al., 2023) and validating the effectiveness of circuit discovery techniques (Conmy et al., 2023). Pyrefthttps://github.com/stanfordnlp/pyreft (Wu et al., 2024b) is a toolkit based on Pyvene for fine-tuning and sharing trainable interventions (Section 3.2.3, Causal abstraction) aimed at optimizing LM performance on selected tasks, in a similar but more targeted and efficient way than parameter-efficient fine-tuning methods (PEFT, Han et al., 2024). Going beyond the textual modality, ViT Prisma (Joseph, 2023) is a toolkit to conduct mechanistic interpretability analyses on vision and multimodal models. Finally, MAIA (Shaham et al., 2024) is a multimodal language model augmented with tool use to automate common interpretability workflows such as neuron explanations, example synthesis and counterfactual editing.

Conclusion and Future Directions

In this paper, we have offered an overview of the existing interpretability methods useful for understanding Transformer-based language models, and have presented the insights they have led to. Although the focus of this work is on practical methods and findings, we acknowledge theoretical studies related to the interpretability of Transformers, such as investigations explaining in-context learning (Akyürek et al., 2023; Von Oswald et al., 2023; Xie et al., 2022), explorations of Transformers through the lens of data compression and representation learning (Yu et al., 2023b; Voita et al., 2019a), the study of Transformers’ learning dynamics (Tian et al., 2024; 2023; Tarzanagh et al., 2024), or the analyses on their generalization properties on algorithmic tasks (Nogueira et al., 2021; Anil et al., 2022; Zhou et al., 2024).

Looking forward, we believe that the ultimate test for insights collected in years of interpretability work remains their applicability in debugging and improving the safety and reliability of future models, providing developers and users with better tools to interact with them and understand the factors influencing their predictions (Longo et al., 2024). To ensure such requirements are met, future developments in interpretability research will be faced with the challenging task of moving from functionally-grounded evaluations (i.e. no human evaluation, only toy settings) to actionable insights and benefits for real-world tasks (Doshi-Velez & Kim, 2017). From an analytical standpoint, this involves moving from methods and analyses operating in model component space to human-interpretable space, i.e from model components to features and natural language explanations, as suggested by Singh et al. (2024b), while still faithfully reflecting model behaviors (Siegel et al., 2024). Directions we deem promising in this area involve the usage of LMs as verbalizers (Feldhus et al., 2023; Bills et al., 2023; Wang et al., 2024; Chen et al., 2024c) for scaling input and component attribution analyses, especially when paired with verification mechanisms to ensure counterfactual consistency (Avitan et al., 2024), and circuit discovery methods leveraging interpretable features to enable interventions motivated by human-understandable concepts (Marks et al., 2024). More accessible insights might also unlock gains in model performance and efficiency, translating interpretability-driven insights into downstream task improvements (Wu et al., 2024b). Importantly, interdisciplinary research grounded in the technical developments we summarize in this survey will play a key role in broadening the scope of interpretability analyses to account for the perceptual and interactive dimensions of model explanations from a human perspective (Liao et al., 2020; Dhanorkar et al., 2021; Vasconcelos et al., 2023). Ultimately, we believe that ensuring open and convenient access to the internals of advanced LMs will remain a fundamental prerequisite for future progress in this area (Bau et al., 2023; Casper et al., 2024; Hudson et al., 2024).

Acknowledgements

Javier Ferrando is supported by the Spanish Ministerio de Ciencia e Innovación through the project PID2019-107579RB-I00 / AEI / 10.13039/501100011033. Gabriele Sarti and Arianna Bisazza acknowledge the support of the Dutch Research Council (NWO) as part of the project InDeep (NWA.1292.19.399).

References

Appendix A Mathematical Notation

Appendix B Linearization of the LayerNorm

The LayerNorm operates over an input z{\bm{z}} as: LN(z)=z−μ(z)σ(z)⊙γ+β\text{LN}({\bm{z}})=\frac{{\bm{z}}-\mu({\bm{z}})}{\sigma({\bm{z}})}\odot\mathbf{\gamma}+\mathbf{\beta}, where μ\mu and σ\sigma compute the mean and standard deviation of z{\bm{z}}, and γ\gamma and β\beta refer to the element-wise transformation and bias respectively. Holding σ(z)\sigma({\bm{z}}) as a constant, the LayerNorm can be decomposed into zL+β{\bm{z}}\mathbf{L}+\beta, where L\mathbf{L} is a linear transformation:

Appendix C Folding the LayerNorm

Any Transformer block reads from the residual stream by normalizing before applying a linear layer (with weights W{\bm{W}} and bias b{\bm{b}}) to the resulting vector:

Following the decomposition in Equation 26 we can fold the weights of the LayerNorm into those of the subsequent linear layer as follows:

where the new weights and bias are W∗=LW{\bm{W}}^{*}={\bm{L}}{\bm{W}} and b∗=βW+b{\bm{b}}^{*}=\beta{\bm{W}}+{\bm{b}} respectively.

Appendix D Implementation details of SAEs

During training, a feature receives a zero gradient signal if it does not activate. When this occurs frequently, it can lead to a dead feature. Bricken et al. (2023) propose resampling these features by reinitializing their encoder and decoder weights periodically during training. An alternative approach to resampling is ghost gradients Jermyn & Templeton (2024), which adds an auxiliary loss term that supplies a gradient signal to promote the reactivation of dead features. However, recent results have found this approach suboptimal (Rajamanoharan, 2024; Conerly et al., 2024).

Setting the β1\beta_{1} parameter of Adam to 0 has been found to reduce the number of “dead” features in larger autoencoders (Templeton et al., 2024; Rajamanoharan et al., 2024). Yet, Conerly et al. (2024) rely on β1=0.9\beta_{1}=0.9.