Function Vectors in Large Language Models
Eric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller, Byron C. Wallace, David Bau
Introduction
Since the study of the lambda calculus (Church, 1936), computer scientists have understood that the ability for a program to carry references to its own functions is a powerful idiom. Function references can be helpful in many settings, allowing expression of complex control flow through deferred invocations (Sussman, 1975), and enabling flexible mappings from inputs to a target task. In this paper we report evidence that autoregressive transformers trained on large corpora of natural text develop a rudimentary form of function references.
Our results begin with an examination of in-context learning (ICL; Brown et al., 2020). ICL mechanisms have previously been studied from the perspective of making copies (Olsson et al., 2022) and from a theoretical viewpoint (von Oswald et al., 2022; Garg et al., 2022; Dai et al., 2022), but the computations done by large models to generalize and execute complex ICL functions are not yet fully understood. We characterize a key mechanism of ICL execution: function vectors (FVs), which are compact vector representations of input-output tasks that can be found within the transformer hidden states during ICL. An FV does not directly perform a task, but rather it triggers the execution of a specific procedure by the language model (Figure 1).
Function vectors arise naturally when applying causal mediation analysis (Pearl, 2001; Vig et al., 2020; Meng et al., 2022a; b; Wang et al., 2022a) to identify the flow of information during ICL. We describe an activation patching procedure to determine the presence of a handful of attention heads that mediate many ICL tasks. These heads work together to transport a function vector that describes the task; the FV can be formed by summing outputs of the causal attention heads.
We test the hypothesis that function vectors are a general mechanism spanning many types of functions. To quantify the role and efficacy of function vectors, we curate a data set of over 40 diverse ICL tasks of varying complexity. We calculate FVs for these tasks and investigate impact of FVs in triggering those functions across a variety of LMs scaling up from 6B to 70B parameters.
We further ask whether FVs are portable: are the effects of an FV limited to contexts very similar to those where it is extracted, or can an FV apply in diverse settings? We compare the effects of FVs when inserted into diverse input contexts including differently-formatted forms, zero-shot formats, and natural text contexts. We find that FVs are remarkably robust, typically triggering function execution even in contexts that bear no resemblance to the original ICL context.
A key question is whether the action of FVs can be explained by word-embedding vector arithmetic (Mikolov et al., 2013; Levy & Goldberg, 2014; Merullo et al., 2023). We examine decodings of FVs (Nostalgebraist, 2020), and find that though FVs often encode a function’s output vocabulary, those vocabularies do not fully identify an FV. In other words, to invoke functions, FVs require some information beyond what is encoded in the top vocabulary words.
Finally, we investigate whether the space of FVs has its own vector algebra over functions rather than words. We construct a set of composable ICL tasks, and we test the ability of FVs to obey vector algebra compositions. Our findings reveal that, to some extent, vector compositions of FVs produce new FVs that can execute complex tasks that combine constituent tasks. We emphasize that FV vector algebra is distinct from semantic vector algebra over word embeddings: for example, composed FV vectors can specify nonlinear tasks such as calculating the antonym of a word, that cannot themselves be implemented as a simple embedding-vector offset (Appendix A). Our datasets, code, and evaluation scripts are available at https://functions.baulab.info.
Method
When a transformer processes an ICL prompt with exemplars demonstrating task , do any hidden states encode the task itself?
Surprisingly, we find that adding the average activations in this way at particular layers induces the model to perform the task in the new context. For example, if antonym, the red line in Figure 2c shows that adding at layer 12 in GPT-J causes the model to produce antonyms in a zero-shot context, with accuracy. That suggests that does encode the antonym task.
2 Formulation
For each task in our universe of ICL tasks we have a data set of in-context prompts . Each prompt is a sequence of tokens with input-output exemplar pairs that demonstrate the same underlying task mapping between and , and one query input corresponding to a target (correct) response that is not part of the prompt, that should be predicted by the LM if it generalizes correctly. We focus our analysis on successful ICL by including in only prompts where the prediction ranks the correct answer highest. We write one ICL prompt as
3 Causal Mediation to Extract Function Vectors from Attention Heads
To distill the information flow during ICL, we apply causal mediation analysis.
Then each attention head’s average indirect effect (AIE) is calculated by averaging this difference across all tasks and (corrupted) prompts:
Figure 3a shows the AIE per attention head in GPT-J over many tasks (Appendix F shows larger models). The 10 attention heads with highest AIE (which make up ) are highlighted in pink (square outlines) and are clustered primarily in early-middle layers of the network. The average attention pattern of these heads at the final token is shown for two tasks in Figure 3b. These heads primarily attend to token positions corresponding to example outputs; this observation is consistent with the high salience of ICL label tokens observed by Wang et al. (2023a) and while this resembles the same prefix-matching attention pattern as “induction heads” (Elhage et al., 2021; Olsson et al., 2022) not all heads in reproduce this pattern on other contexts with repeated tokens (Appendix G).
Due to their high causal influence across many tasks, we hypothesize that this small set of heads is responsible for transporting information about the demonstrated ICL task. We can represent the contribution of as a single vector by taking the sum of their average outputs, over a task, which we call a function vector (FV) for task :
Experiments
We deploy a series of decoder-only autoregressive language models; each is listed and described in Table 1. We use huggingface implementations (Wolf et al., 2020) of each model.
Tasks.
We construct a diverse array of over 40 relatively simple tasks to test whether function vectors can be extracted in diverse settings. To simplify the presentation of our analysis, we focus on a representative sample of 6 tasks:
Antonym. Given an input word, generate the word with opposite meaning.
Capitalize. Given an input word, generate the same word with a capital first letter.
Country-Capital. Given a country name, generate the capital city.
English-French. Given an English word, generate the French translation of the word.
Present-Past. Given a verb in present tense, generate the verb’s simple past inflection.
Singular-Plural. Given a singular noun, generate its plural inflection.
All other tasks are described in Appendix D.
1 Portability of Function Vectors
In this section, we investigate the portability of function vectors—i.e., the degree to which adding an FV to a particular layer at the final token position of the prompt can cause the language model to perform a task in contexts that differ from the ICL contexts from which it was extracted. For simplicity of analysis, we only include test queries for which the LM answers correctly given a 10-shot ICL prompt; all accuracies and standard deviations over 5 random seeds are reported on this filtered subset. Results when incorrect ICL are included are similar (see Appendix C).
In Table 2 we report results (averaged across the 6 tasks mentioned above) for adding FVs to shuffled-label ICL prompts and zero-shot contexts across 3 models - GPT-J, GPT-NeoX and Llama 2 (70B), at layers 9, 15, and 26 respectively (approximately ). For GPT-J, we also compare the efficacy of FVs to other approaches for extracting task-inducing vectors including simple state averaging (§2.1).
Zero-Shot Results Across Layers.
Figure 4 shows results across layers for the zero-shot case. The sharp reduction of causal effects in late layers suggests that FVs do not simply act linearly, but that they trigger late-layer nonlinear computations. This pattern of causality is seen across a variety of tasks, autoregressive model architectures, and model sizes. Even in cases where performance is low, as in English-French with GPT-NeoX and Llama 2 (70B), adding the function vector in middle layers still results in large relative improvements to accuracy over the zero-shot baseline. Results are also consistent across model sizes: see Appendix I for results with all sizes of Llama 2.
FVs are Robust to Input Forms.
We create 20 different ICL templates that vary the form of the ICL prompt across prefixes and delimiters of input-output pairs. We evaluate FVs on GPT-J for these 20 templates in both shuffled-label and zero-shot template settings. Across our 6 tasks, adding the FV executes the task with an average accuracy of in shuffled-label setting and in the zero-shot setting, while GPT-J only scores and on the same settings, respectively. Despite higher variance, this performance is similar to performance in the same settings with the original template.
We also evaluate FVs on natural text completions. Given a natural text template, we insert a test query word and have the model generate tokens. We add the FV to the final token of the original prompt, and for all subsequent token predictions to guide its generation. We search the generated string with a simple regex parse to compute whether the generation includes the correct target for the inserted query word.
Table 3 shows natural text portability results for the antonym FV for GPT-J, generating 5 new tokens. In each of the templates, the antonym is in the FV completion significantly more than the original completion. In fact, we find that the efficacy of the antonym FV in eliciting the correct response in these natural text templates performs on par with the results previously reported for the zero-shot setting. This is true for all 6 tasks (Appendix E), suggesting that the task representation transported during ICL is similar to one that is used during autoregressive prediction in natural text settings.
We include a few qualitative results for the English-French and Country-Capital tasks (Table 4). We see that the English-French FV will sometimes translate the whole sentence after giving the proper completion to the original one-word translation task, indicating that it has captured more than the original task it was shown. Additional natural text portability results are included in Appendix E.
2 The Decoded Vocabulary of Function Vectors
Several studies have gleaned insights about the states and parameters of transformers by viewing them in terms of their decoded vocabulary tokens (Geva et al., 2020; Nostalgebraist, 2020; Dar et al., 2022; Geva et al., 2022; Belrose et al., 2023). Therefore we ask: can we understand an FV by decoding directly to a token probability distribution? The results are shown in Table 5, which lists the top five tokens in the decoded distribution for each task.
A clear pattern emerges: for most tasks, the decoded tokens lie within the task’s output space. The Singular-Plural function vector decodes to a distribution of plural nouns, and Present-Past decodes to past-tense verbs. However, that is not the case for all tasks: English-French decodes to nonsense tokens, and the Antonym task decodes to words that evoke the abstract idea of reversal.
Given these meaningful decodings, we then ask whether the token vocabulary is sufficient to recreate a working function vector. That is, we begin with the token distribution , and determine whether a function vector can be reconstructed if we know the top words in . Denote by the distribution that resamples while restricting to only the top words. We perform an optimization to reconstruct a that matches the distribution when decoded:
In Table 6, the performance of is evaluated when used as a function vector. We find that, while it is possible to partially recreate the functionality of an FV, good performance typically requires more than 100 vocabulary tokens. In other words, knowledge of the top decoded tokens of is usually not enough on its own to construct a working function vector. That suggests that the FV contains some needed information beyond that expressed by its top decoded tokens.
3 Vector Algebra on Function Vectors
Although Table 6 suggests that function vectors cannot be understood as simple semantic vector offsets on word embeddings, we can ask whether function vectors obey semantic vector algebra over the more abstract space of functional behavior by testing the composition of simple functions into more complex ones. We begin with three conceptually decomposable ICL tasks: the list-oriented tasks First-Copy, First-Capital, and Last-Copy, as illustrated in Figure 5a. Using ICL, we collect FVs for all three tasks and denote them , , and .
Then we form a simple algebraic sum to create a new vector that we will denote .
In principle we could expect to serve as a new function vector for a new composed task (Last-Capital). We perform several similar task compositions on a variety of tasks. In each case, we combine a task with First-Copy and Last-Copy to produce a composed Last- vector; then, we test the accuracy of as a function vector. We compare to the accuracy of the FV extracted from ICL, as well as accuracy of the same model performing the task using ICL. Results for GPT-J are reported in Table 7, see Appendix J for results for Llama 2 (13B).
We find that FVs can be composed to some extent, with algebraic compositions outperforming FVs on some tasks, even comparable to ICL on some tasks. Other tasks, including some tasks for which ICL and function vectors perform well, seem to resist the type of vector composition that we attempt. Because success and failure in composition may hinge on the composability of the underlying computations that are triggered by a function vector, we believe that FV composition may be a useful tool for further understanding the mechanisms of LMs.
Related Work
Mechanisms of task performance in LMs. Our work is related to Merullo et al. (2023) which analyzes the role of components during execution of some ICL tasks, as well as Halawi et al. (2023), which examines the behavior of ICL attention heads when there are false demonstrations. Our work is also consistent with findings from Wang et al. (2023a) which observes salience of label tokens during ICL, as well as Wang et al. (2022b) which observes the presence of individual neurons that correlate with specific task performance. Unlike those analyses of mechanisms of individual task executions, we measure causal mediators across a distribution of different tasks to find a generic function-invocation mechanism that identifies and distinguishes between tasks.
Mechanistic Interpretability. We also build upon the analyses of Elhage et al. (2021) and Olsson et al. (2022), who observed prevalent in-context copying behavior that appears related to jumps in performance during training. We isolate FVs using causal mediation analysis methods developed in Pearl (2001); Vig et al. (2020); Meng et al. (2022a); Wang et al. (2022a). We examine FVs in vocabulary space using the logit lens methods, following Nostalgebraist (2020); Geva et al. (2020); Dar et al. (2022); Geva et al. (2023). Unlike these lines of work that examine mechanisms underlying specific capabilities, we focus on locating generic function references within a transformer.
Task Representations. Other studies have investigated how tasks can be represented within a model. Shao et al. (2023) devises a way of training a codebook that yields a compositional task encoding for LLMs. Similarly, Mu et al. (2023) develops a soft-prompt method for compressing and encoding tasks. Panigrahi et al. (2023) and Ilharco et al. (2023) devise ways to learn small parameter changes by task fine-tuning that can be composed; Ilharco et al. (2023) calls such sets of parameter changes task vectors. Our study of function vectors differs qualitatively from these previous works: rather than training a model to create function representations, we ask whether a pretrained transformer already contains compact function representations that it uses to invoke task execution.
In-Context Learning. Since its observation in LLMs by Brown et al. (2020), ICL has been studied intensively from many perspectives. The role of ICL prompt forms has been studied by Reynolds & McDonell (2021); Min et al. (2022); Kim et al. (2022). Models of inference-time metalearning that could explain ICL have been proposed by Akyürek et al. (2022); Dai et al. (2022); von Oswald et al. (2022); Li et al. (2023); Garg et al. (2022). Analyses of ICL as Bayesian task inference have been performed by Xie et al. (2021); Wang et al. (2023c); Wies et al. (2023); Hahn & Goyal (2023); Zhang et al. (2023); Han et al. (2023). And ICL robustness under scaling has been studied by Wei et al. (2023); Wang et al. (2023b); Pan et al. (2023). Our work differs from those studies of the externally observable behavior of ICL by instead focusing on mechanisms within transformers.
Discussion
We have investigated the representations of input-output functions in transformer-based large language models. Causal mediation analysis has identified compact vector representations which trigger the model to perform a specific task; we call these function vectors (FVs). Our analysis reveals that FVs are highly robust to shifts in prompt formats and that the vectors can be combined to compose tasks.
Our experimental results also provide several lines of evidence that distinguish function vectors from semantic vector algebra over word embeddings: function vectors can represent mappings such as antonyms that cannot be expressed as an offset (Figures 2,4); they cannot be reconstructed from their top-100 vocabulary alone (Table 6); and they have strong causal effects at early-mid layers but near-zero effects at late layers (Figure 4; Appendix A for further discussion). Together, these findings suggest that LLMs contain compact vector representations of abstract functions, and that those vectors are able to trigger the performance of nontrivial tasks.
Ethics
While our work clarifying the mechanisms of function representation and execution within large models is intended to help make large language models more transparent and easier to audit, understand, and control, we caution that such transparency may also enable bad actors to abuse large neural language systems, for example by injecting or amplifying functions that cause undesirable behavior.
Acknowledgments
We are grateful for the generous support of Open Philanthropy (ET, AS, AM, DB) as well as National Science Foundation (NSF) grant 1901117 (ET, ML, BW). ML is supported by an NSF Graduate Research Fellowship, and AM is recipient of the Zuckerman Postdoctoral Fellowship. We thank the Center for AI Safety (CAIS) for making computing resources available for this research.
References
Appendix A Discussion: Function Vectors vs Semantic Vector Arithmetic
In this appendix we discuss the experimental support for our characterization of function vectors in more detail, in particular the assertion that function vectors are acting in a way that is distinct from semantic vector arithmetic on word embeddings.
The use of vector addition to induce a mapping is familiar within semantic embedding spaces: the vector algebra of semantic vector offsets has been observed in many settings; for example, word embedding vector arithmetic was clearly described by Mikolov et al. (2013), and has been observed in other neural word representations including transformers (recently, Merullo et al., 2023). Therefore one of our main underlying research questions is whether our function vectors should be described as triggers for a nontrivial function, or whether, more simply, they are just ordinary semantic vector offsets that induce a trivial mapping between related words by adding an offset to an embedding.
The main paper contains three pieces of experimental evidence that support the conclusion that function vectors are different from semantic vector offsets of word embeddings, and that they trigger nontrivial functions:
Function vectors can implement complex mappings, including cyclic mappings such as antonyms that cannot be semantic vector offsets.
Function vectors cannot be recovered from the target output vocabulary alone; they carry some other information.
Function vector activity is mediated by mid-layer nonlinearities (i.e., they trigger nonlinear computations), since they have near-zero causal effect at late layers.
We discuss each of these lines of evidence in more detail here.
The first task analyzed in the paper is the antonym task. Because the antonym task is cyclic, it is a simple counterexample to the possibility that function vectors are just semantic vector offsets of language model word embeddings.
Since we are able to find a constant antonym function vector that, when added to the transformer, does cause cyclic behavior, we conclude that the action of is a new phenomenon. Function vectors act in a way different from simple semantic vector offsets.
Function Vectors Contain Information Beyond Output Vocabulary.
Not every function is cyclic, but the vector offset hypothesis can be tested by examining word embeddings. Following the reasoning of Geva et al. (2022), one way to potentially implement a semantic vector offset is to promote a certain subset of tokens that correspond to a particular semantic concept (i.e. capital cities, past-tense verbs, etc.). Function vectors do show some evidence of acting in this way: when decoding function vectors directly to the model’s vocabulary space we often see that the tokens with the highest probabilities are words that are part of its task’s output vocabulary (Table 5, Table 19).
That experiment reveals that while function vectors do often encode words contained in the output space of the task they are extracted from, simply adding a vector that boosts those same words by the same amounts is not enough to recover its full performance—though if enough words are included, in some cases a fraction of performance is recovered. Our measurements suggest that while part of the role of a function vector may act similarly to a semantic vector offset, the ability for function vectors to produce nontrivial task behavior arises from other essential information in the vector beyond just a simple word embedding vocabulary-based offset.
Function Vectors’ Causal Effects are Near-Zero at Late Layers.
Across the set of tasks and models we evaluate, there is a common pattern of causal effects that arises when adding a function vector to different layers. The highest causal effects are achieved when adding the function vector at early and middle layers of the network, with a sharp drop in performance to near-zero at the later layers of the network (Figure 4, Figure 25, Figure 14, Figure 16). Interestingly, FV causal effects are strongest in largest models, yet the cliff to near-zero causal effects is also sharpest for the largest models (Figure 16).
If the action were linear and created by a word embedding vector offset, then the residual stream would transport the action equally well at any layer including the last layers. Thus the pattern of near-zero causal effects at later layers suggests that the action of the function vector is not acting directly on the word embedding, but rather that it is mediated by some nonlinear computations in the middle layers of the network that are essential to the performance of the task. This mediation is evidence that the function vector activates mid-layer components that execute the task, rather than fully executing the task itself.
This pattern is in contrast to the vector arithmetic described in Merullo et al. (2023), that are most effective at later layers of the network and have little to no causal effect at early layers of the network; those offsets more closely resemble semantic vector offsets of word embeddings.
In summary, the three lines of evidence lead us conclude that the vectors should not be seen as simple word embeddings, nor trivial offsets or differences of embeddings, nor simple averages of word embeddings vocabularies to boost. Rather, the evidence suggests that the vectors act in a way that is distinct from literal token embedding offsets, to trigger nonlinear function execution. Thus we view the vectors as references to functions, and we call them function vectors.
Appendix B Experimental Details
In this section, we provide details of the function vector extraction process (section 2.3), and the evaluation of function vectors (section 3).
Prompt Templates.
To evaluate a function vector we use a few different prompt contexts. The shuffled-label prompts are corrupted 10-shot prompts with the same form as (9), while zero-shot prompts only contain a query , without prepended examples (e.g. Q:{}\nA:).
In section 3.1 we use FVs extracted from prompts made with the template shown in (9), and test them across a variety of other templates (Table 8).
Evaluating Function Vectors.
To evaluate a function vector’s (FV) causal effect, we add the FV to the output of a particular layer in the network at the last token of a prompt with query , and then measure whether the predicted word matches the expected answer . We report this top-1 accuracy score over the test set. If is tokenized as multiple tokens, we use the first token of as the target token.
Appendix C Results Including Incorrect ICL
For simplicity of presentation in Section 3, we filter the test set to cases where the model correctly predicts given a 10-shot ICL prompt containing query . In this section we compare those results to the setting in which correct-answer filtering is not applied. When filtering is not applied, the causal effects of function vectors remain essentially unchanged (Figure 7).
Appendix D Datasets
Here, we describe the tasks we use for evaluating the existence of function vectors. A summary of each task can be found in Table 9.
Our antonym and synonym datasets are based on data taken from Nguyen et al. (2017). They contain pairs of words that are either antonyms or synonyms of each other (e.g. “good bad”, or “spirited fiery”). We create an initial dataset by combining all adjective, noun, and verb pairs from all data splits and then filter out duplicate entries. We then further filter to word pairs where both words can be tokenized as a single token. As a result, we keep 2,398 antonym word pairs and 2,881 synonym word pairs.
We note that these datasets originally included multiple entries for a single input word (e.g. both “simple difficult” and “simple complex” are entries in the antonym dataset). In those cases we prompt a more powerful model (GPT-4; OpenAI, 2023) with 10 ICL examples and keep an answer as output after manually verifying it.
Translation.
We construct our language translation datasets – English-French, English-German, and English-Spanish – using data from Conneau et al. (2017), which consists of a word in English and its translation into a target language. For each language, we combine the provided train and test splits into a single dataset and then filter out cognates. What remains are 4,705 pairs for English-French, 5,154 pairs for English-German, and 5,200 pairs for English-Spanish.
These datasets originally included multiple entries for a single input word (e.g. both “answer respuesta” and “answer contestar” are entries in the English-Spanish dataset), and so we filter those with GPT-4, in a similar manner as described for Antonym and Synonym.
Sentiment Analysis.
Our sentiment analysis dataset is derived from the Stanford Sentiment Treebank (SST-2) Socher et al. (2013), a dataset of movie review sentences where each review has a binary label of either “positive” or “negative”. An example entry from this dataset looks like this: “An extremely unpleasant film. negative”. We use the same subset of SST-2 as curated in Honovich et al. (2022), where incomplete sentences and sentences with more than 10 words are discarded, leaving 1167 entries in the dataset. See Honovich et al. (2022) for more details.
CommonsenseQA.
This is a question answering dataset where a model is given a question and 5 options, each labeled with a letter. The model must generate the letter of the correct answer. For example, given the question “Where is a business restaurant likely to be located?” and answer options “a: town, b: hotel, c: mall, d: business sector, e: yellow pages”, a model must generate “d”. (Talmor et al., 2018)
AG News.
A text classification dataset where inputs are news headlines and the first few sentences of the article, and the labels are the category of the news article. Labels include Business, Science/Technology, Sports, and World. (Zhang et al., 2015)
We also construct a set of simple tasks to broaden the types of tasks on which we evaluate FVs.
Capitalize First Letter.
To generate a list of words to use to capitalize, we utilize ChatGPThttps://chat.openai.com/ by prompting it to give us a list of words. From here, we curate a dataset where the input is a single word, and the output is the same word with the first letter capitalized.
Lowercase First Letter.
Similar to the Capitalize First Letter task, we use the same set of words but instead change the task to instead lowercase a word. The input is a single word title cased, and the output is the same word, but lowercase instead.
Country-Capital.
We also generate a list of country-capitals with ChatGPT. Here, we ask ChatGPT to come up with a list of countries and following that, ask it to name the capitals that are related to the countries. This dataset contains 197 country-capital city pairs.
Country-Currency.
We also generate a list of country-currency pairs with ChatGPT. Similar to country-capital, we ask ChatGPT to come up with a list of countries and following that, ask it to name the currencies that are related to the countries.
National Parks.
The National Parks dataset consists of names of official units of national parks in the United States, paired with the state the unit resides in (e.g. Zion National Park Utah). It was collected from the corresponding page on Wikipediahttps://en.wikipedia.org/wiki/List_of_the_United_States_National_Park_System_official_units, accessed July 2023.
Park-Country.
The Park-Country dataset consists of names of national parks around the world, paired with the country that the park is in. The countries were first generated by ChatGPT, then the parks were also generated by ChatGPT. After, the dataset was hand-checked for factual accuracy since ChatGPT tended to hallucinate parks for this dataset. Not all national parks per country are included, just a subset.
Present-Past.
We generate a list of present-tense verbs with ChatGPT then ask ChatGPT to find the past-tense version. After generation, the dataset was hand-corrected for inaccuracies. The dataset inputs are simple present tense verbs and outputs are corresponding simple past tense verbs.
Landmark-Country.
The Landmark-Country dataset consists of entries with the name of a landmark, and the country that it is located in. The data pairs are taken from Hernandez et al. (2023).
Person-Instrument.
The Person-Instrument dataset contains entries with the name of a professional musician and the instrument they play. The data pairs are taken from Hernandez et al. (2023).
Person-Occupation.
The Person-Occupation dataset is taken from Hernandez et al. (2023), and contains entries of names of well-known individuals and their occupations.
Person-Sport.
The Person-Sport dataset is taken from Hernandez et al. (2023), and each entry consists of the name of a professional athlete and the sport that they play.
Product-Company.
The Product-Company dataset contains entries with the name of a commercial product, paired with the company that sells the product. The data we use is the same as curated in Hernandez et al. (2023).
D.1 Extractive Tasks.
Many NLP tasks are abstractive; that is, they require the generation of information that is not present in the prompt. We also wish to test whether function vectors are recoverable from extractive tasks—that is, tasks where the answer is present somewhere in the prompt, and the task of the model is to retrieve it.
In our experiments, we use a subset of the CoNLL-2003 English named entity recognition (NER) dataset Sang & De Meulder (2003), which is a common NLP benchmark for evaluating NER models. The NER task consists of extracting the correct entity from a given sentence, where the entity has some particular property. In our case, we create three different datasets: NER-person, NER-location, and NER-organization, where the label of each task is the name of either a person, location, or organization, respectively. Each dataset is constructed by first combining the CoNLL-2003 “train” and “validation” splits into a single dataset, and then filtering the data points to only include sentences where a single instance of the specified class (person, location, or organization) is present. This helps reduce the ambiguity of the task, as cases where multiple instances of the same class are present could potentially have multiple correct answers.
As with abstractive tasks, we also construct a set of new extractive tasks.
Choose n𝑛nth Item from List.
Here, the model is given a list of comma-separated items, and the model is tasked with selecting the item at a specific index. We construct tasks where the list size is either 3 or 5. In our tasks, we have the model choose either the first element or last element from the list.
Choose Category from List.
These tasks are similar to our choose th element tasks, but instead, the model must select an item of a particular type within a list of 3 or 5 items. A word with the correct type is included once in the list while the remaining words are drawn from another category. The categories we test include the following: fruit vs. animal, object vs. concept, verb vs. adjective, color vs. animal, and animal vs. object.
D.2 Few-Shot ICL Performance
Figure 8 shows the few-shot performance (top-1 accuracy) for GPT-J on a larger subset of our task list. The dotted baseline is based on the majority label for the dataset, computed as (# majority label)/(# total instances).
Figure 9 shows the few-shot performance (top-1 accuracy) for GPT-J on additional extractive tasks. For datasets with lists of words as input, the dotted baseline is computed to be , (e.g. ), and for all other tasks it represents predicting the majority label of the dataset, computed as (# majority label)/(# total instances).
D.3 Evaluating Function Vectors on Additional Tasks
Appendix E Portability
In addition to the quantitative results provided in §3.1, we include additional qualitative examples for natural text completions for the Antonym FV (Table 10). We also include qualitative examples of completions on different natural text prompts using the English-French FV (Table 11), the English-Spanish FV (Table 13), and the Country-Capital FV (Table 14).
We also include additional quantitative results for testing FVs in natural text settings on various natural text templates, averaged over 5 seeds for each of the remaining representative tasks: Capitalize (Table 16), Present-Past (Table 17), Singular-Plural (Table 18), English-French (Table 12), and Country-Capital (Table 15).
Appendix F Causal Mediation Analysis
In this section we include additional figures showing the average indirect effect (AIE) split up across tasks for GPT-J (Figure 17), as well as the AIE for other models we evaluated.
The AIE for a particular model is computed across all abstractive tasks where the model can perform the task better than the majority label baseline given 10 ICL examples. We include heatmaps of the AIE for each attention head in Llama 2 (7B) (Figure 18), Llama 2 (13B) (Figure 19), and Llama 2 (70B) (Figure 21), as well as GPT-NeoX (Figure 20). For each model, the heads with the highest AIE are typically clustered in the middle layers of the network. In addition the maximum AIE across all heads tends to drop slightly as the size of the model increases, though the total number of heads also increases. In Llama 2 (7B) the max AIE is , while in Llama 2 (70B) the max AIE is .
Appendix G Attention Patterns on Sequences of Repeated Tokens
Across a variety of tasks, the heads with highest causal effect have a consistent attention pattern where the attention weights on few-shot ICL prompts are the strongest on the output tokens of each in-context example. Here we show this pattern for GPT-J on 4 additional tasks (Figure 22, Figure 23), which match the patterns shown in the main paper (Figure 3b). This is similar to the attention pattern that might be expected of “induction heads”, which has previously been shown to arise when a prompt contains some repeated structure (Elhage et al., 2021; Olsson et al., 2022).
To further investigate whether the heads identified via causal mediation analysis are “induction heads”, we compute the prefix-matching score for each head in GPT-J. We follow the same procedure as described in (Olsson et al., 2022; Wang et al., 2022a), which computes the prefix-matching score as the average attention weight on a token when given a sequence of the form [, , …, ]. This is measured on sequences of repeated random tokens. We do this for each head in GPT-J with results shown in Figure 24b.
We find that three of the heads out of those with the top 10 highest AIEs (Figure 24) also have high prefix-matching scores. In terms of “Layer-Head Index”, these are heads 8-1, 12-10, and 24-6, with prefix-matching scores of 0.49, 0.56, and 0.31, respectively.
While (Elhage et al., 2021; Olsson et al., 2022) show that induction heads play a critical role in copying forward previously seen tokens, our results show that they are also among the set of heads, , that have the highest AIE when resolving few-shot ICL prompts.
There are several other heads we identified with relatively high causal effect that have the same attention pattern activation on few-shot ICL prompts, but do not produce the same “induction” attention pattern on sequences of random repeated tokens.
This suggests that while induction heads play a role in the formation of function vectors, there are other heads that also contribute relevant information that may not be induction heads of the type observed by (Elhage et al., 2021; Olsson et al., 2022).
Appendix H Decoding Vocabulary Evaluation
Here, we present more results on the evaluation of decoding vocabularies of FVs, over additional datasets. Across the tasks, we affirm that output spaces seem to be frequently encoded in the FVs, as can be seen in Table 19. In particular, in cases such as sentiment that follow a rigid pattern, the tokens referring to the output distribution for sentiment is well encoded. On the other hand, some tasks like language translation do not have output spaces well-encoded in the FVs.
Appendix I Scaling Effects
Can we consistently locate function vectors given various sizes of a single language model architecture? We test this by observing all sizes of Llama 2, ranging from 7B parameters to 70B. We use the same methods as in §3.1, adding function vectors to each layer of the model and observing accuracy on our subset of 6 tasks at each layer.
We find that results (Figure 25) are largely consistent across model sizes. Function vectors generally result in the highest zero-shot accuracies when added to the early to middle layers; this is true regardless of the total number of layers in the model.
Appendix J Composition on Other Models
In this section we include additional composition results for Llama 2 (13B) (Table 20).