One Embedder, Any Task: Instruction-Finetuned Text Embeddings

Hongjin Su, Weijia Shi, Jungo Kasai, Yizhong Wang, Yushi Hu, Mari Ostendorf, Wen-tau Yih, Noah A. Smith, Luke Zettlemoyer, Tao Yu

Introduction

Text embeddings represent discrete text inputs (e.g., sentences, documents, and code) as fixed-sized vectors that can be used in many downstream tasks. These tasks include semantic textual similarity Agirre et al. (2012); Marelli et al. (2014); Cer et al. (2017); Lin et al. (2018), information retrieval Mitra et al. (2017); Karpukhin et al. (2020); Izacard et al. (2022), automatic text evaluation Zhang et al. (2020); Sellam et al. (2020); Hessel et al. (2021), prompt retrieval for in-context learning Liu et al. (2022); Rubin et al. (2022); Su et al. (2022), and beyond. Recently, we have seen dramatic advances in learning text embeddings Kiros et al. (2015); Conneau et al. (2017); Logeswaran and Lee (2018); Reimers and Gurevych (2019); Gao et al. (2021); Ni et al. (2021, 2022) that perform well on their intended tasks or datasets.

However, most existing embeddings can have significantly degraded performance when applied to new tasks or domains Thakur et al. (2021); Muennighoff et al. (2022). For example, DPR Karpukhin et al. (2020) is stronger for retrieval than text similarity tasks, and vice versa for SimCSE Gao et al. (2021). Moreover, existing embeddings usually perform poorly when applied to the same type of task but in different domains such as medicine and finance. A common method to address this issue is to further finetune the embeddings on datasets in downstream tasks and domains, which often requires a lot of annotated data Gururangan et al. (2020). In this paper, we hypothesize that text embeddings (even for the same text input) can be adjusted to different downstream applications using task and domain descriptions, without further task- or domain-specific finetuning.

We introduce InstructOR (Instruction-based Omnifarious Representations), a single multitask model that generates task- and domain-aware embeddings given a text input and its task instructions. It achieves state-of-the-art performance on massively many downstream embedding tasks without any training. At the core of our approach is instruction-based finetuning Zhong et al. (2021); Min et al. (2022); Sanh et al. (2022); Wei et al. (2022): we embed every input together with its end task and domain instruction, departing from prior approaches to embeddings that only take text input. InstructOR embeds the same input into different vectors for different end goals (e.g., Who sings the song “Love Story”? is embedded into three different vectors for different tasks in Fig. 1). As shown in Fig. 2, InstructOR is trained on MEDI, our new collection of 330 text embedding datasets newly annotated with human-written task instructions (§2.3). We train InstructOR with a contrastive loss over all datasets that maximizes the similarity between semantically related text pairs while minimizing unrelated pairs.

We extensively evaluate InstructOR on diverse domains (e.g., finance, medicine, and news) and a variety of downstream applications (a total of 70 embedding evaluation datasets, including 66 not seen during training), spanning classification, semantic textual similarity, information retrieval, text generation evaluation, and prompt retrieval for in-context learning. InstructOR significantly outperforms prior state-of-the-art embedding models by an average of 3.4% over the 70 diverse datasets. InstructOR also outperforms a variant that is trained without task instructions (§3), demonstrating the importance of instructions to create task-aware embeddings. Our analysis shows that instruction finetuning addresses the challenge of training a single model on diverse datasets (§4.1). Further, we demonstrate that the task diversity of MEDI makes the performance of InstructOR particularly robust to paraphrases in instructions (§4.2). Overall, these results strongly suggest that instruction finetuning should be adopted broadly for text embeddings, which we support by sharing all of our models and code.

InstructOR

InstructOR encodes inputs together with task instructions, thereby providing task-specific representations that can be used for many downstream language tasks, without any additional training. Here we introduce the architecture of InstructOR (§2.1), present how we perform multitask instruction-based finetuning (§2.2), and describe how we collect and annotate the MEDI training data (§2.3). By default, we refer "task" to a dataset, and use them interchangeably throughout the paper, while a "task category", such as Retrieval, includes many tasks.

We build InstructOR, based on the single encoder architecture Izacard and Grave (2021); Ni et al. (2021, 2022). Following prior work Ni et al. (2021, 2022), we use GTR models as the backbone encoder (GTR-Base for InstructOR-Base, GTR-Large for InstructOR, GTR-XL for InstructOR-XL). The GTR models are initialized from T5 models, pretrained on a web corpus, and finetuned on information search datasets. The availability of different sizes in the GTR model family allows us to explore the scaling behaviors of instruction-finetuned embedding models. Given an input text xx and a task instruction IxI_{x}, InstructOR encodes their concatenation Ix⊕xI_{x}\oplus x. We then generate a fixed-sized, task-specific embedding EI(Ix,x)\textbf{E}_{I}(I_{x},x) by applying mean pooling to the last hidden representations over the tokens in xx.

2 Training Objective

InstructOR is trained by formulating a wide variety of tasks as a text-to-text problem of distinguishing good/bad candidate outputs y∈{y+,yi−}y\in\{y^{+},y^{-}_{i}\} given an input xx, where a training sample corresponds to the tuple (x,Ix,y,Iy)(x,I_{x},y,I_{y}), with IxI_{x} and IyI_{y} being instructions associated with xx and yy, respectively. For example, in a retrieval task, xx is a query, and good/bad yy is a relevant/irrelevant document from some document collection. For a textual similarity task, the input and output have a similar form and typically come from the same source collection. For a classification task, training samples can be formed by choosing yy as text sequences associated with the same vs. different classes for good vs. bad examples (Details about pair construction are in §2.3). The input and output instructions depend on the task. For symmetric tasks such as textual similarity, where the input and output have the same form and encoding objective, the instructions are the same. For asymmetric tasks such as retrieval, where the input is a single sentence query and the output is a document, the instructions reflect that difference.

The goodness of candidate yy for input xx is given by similarity s(x,y)s(x,y) that is the cosine between their InstructOR embeddings:

Following Ni et al. (2021), we maximize the similarity between positive pairs (x,y+)(x,y^{+}) and minimize negative pairs {(x,yi−)}i=1k\{(x,y^{-}_{i})\}_{i=1}^{k}, where kk denotes the number of negative pairs per positive pair. Specifically, our training objective is:

where γ\gamma is the softmax temperature and B\mathcal{B} is a union of (x,y+)(x,y^{+}) and {(x,yi−)}i=1k\{(x,y^{-}_{i})\}_{i=1}^{k}. Further following Ni et al. (2021), we compute the same loss with xx and yy swapped and add it to the previous loss (i.e., bidirectional in-batch sampled loss).

3 MEDI: Multitask Embedding Data with Instructions

There are no existing datasets that consist of a variety of tasks for embedding training with instructions. We thus construct a collection of 330 datasets with instructions across diverse task categories and domains: Multitask Embeddings Data with Instructions (MEDI).

We build MEDI by combining 300 datasets from Super-NaturalInstructions (super-NI; Wang et al., 2022b) with 30 datasets from existing collections designed for embedding training.

The super-NI datasets come with natural language instructions, but positive and negative pairs are not provided. We construct these pairs by using Sentence-T5 embeddings Ni et al. (2022),We do not include instruction for Sentence-T5 as it is not fine-tuned with instructions. denoted with E(⋅)\textbf{E}(\cdot). For the classification datasets, we calculate the pairwise cosine similarity between examples based on input text embeddings cos⁡(E(xi),E(xj))\cos(\textbf{E}(x_{i}),\textbf{E}(x_{j})). An example xix_{i} with a high similarity to xjx_{j} is used to create a positive pair if both examples have the same class label (yj+=yiy_{j}^{+}=y_{i}), and a negative pair if the labels differ (yj−≠yiy_{j}^{-}\neq y_{i}). For the remaining tasks where the output labels are text sequences, the following scores are first computed:

We select example pairs with the highest sposs_{pos} as positive pairs and highest snegs_{neg} as hard negative pairs. We use one hard negative together with in-batch sampled negatives in the training. Our later analysis shows that the training data from super-NI particularly improve the instruction robustness in evaluation due to the diverse task definitions (§4.2).

The other 30 embedding training datasets come from the Sentence Transformers embedding data,https://huggingface.co/datasets/sentence-transformers/embedding-training-data. KILT Petroni et al. (2021), and MedMCQA Pal et al. (2022). These 30 datasets already contain positive pairs; a few of them, such as MSMARCO Bajaj et al. (2016) and Natural Questions Kwiatkowski et al. (2019), also contain hard negative pairs. Following Ni et al. (2021), we use four negative pairs (hard or in-batch negatives) during the model finetuning process. Since all of these datasets do not have instructions, we develop a unified instruction template and manually write a specific prompt for each dataset, as described next.All prompts are reviewed by multiple authors independently to make sure they consistently follow our template. We release these instructions together with our MEDI data.

Instruction Annotation

Each training instance from MEDI is a tuple (x,Ix,y,Iy)(x,I_{x},y,I_{y}), where the natural language instructions IxI_{x} and IyI_{y} describe how the embeddings of xx and yy are used for the task. For example, in open-domain QA (e.g., Natural Questions in Table 1), IxI_{x} is “Represent the Wikipedia question for retrieving supporting documents; Input: ,” and IyI_{y} is “Represent the Wikipedia document for retrieval; Input: .”

To make instructions consistent across all datasets in MEDI, we design a unified instruction format that consists of the following parts (see Table 4 in the appendix for instances of each part):

Text Type specifies the type of input text that we encode using the embedding model. For example, for an open-domain QA task, the input type of the query is a question, while the input type of the target is a document.

Task Objective (Optional) describes the objective of how the input text is used in a task. For example, for a classification task, the task objective is to classify the sentence into some category, while the task objective of the retrieval is to retrieve a relevant document. Because not all sentences are associated with a specific task (e.g., STS targets general encoding), we make this part optional.

Domain (Optional) describes the task domain. For example, for NewsIR, the domain of the task is news. Because not all tasks specify a domain (e.g., STS deals with general statements),this part is also optional.

The final instruction takes the following format: “Represent the (Domain) Text Type for Task Objective:." Appendix 8 shows instructions for each dataset in MEDI.

Experiments

We train InstructOR on the MEDI data and evaluate it on a wide range of 70 downstream tasks. Specifically, we use the MTEB benchmark from recent work Muennighoff et al. (2022), which consists of 56 datasets over 7 diverse task categories, such as classification, reranking, and information retrieval. We then further apply InstructOR to prompt retrieval for in-context learning and text generation evaluation. In all three settings, InstructOR achieves the state-of-the-art performance. See Appendix §A and §B for our detailed settings.

Table 2 presents the results from InstructOR and the baselines over the three benchmarks: MTEB, Billboard, and prompt retrieval. We conduct head-to-head comparison between InstructOR and GTR models with the same size. We also include the performance of other representative models for reference, while they are not meant for direct comparison.

InstructOR achieves the best performance on all three benchmarks on average. Compared to GTR-Large (335M), from which InstructOR is initialized, instruction finetuning enhances the performance by 5.7%, 18.3%, and 5.7% in MTEB, Billboard, and prompt retrieval respectively. Specifically, among all task categories, InstructOR (335M) demonstrates large improvements over GTR-Large on the text evaluation (18.3%), classification (10.1%), and clustering tasks (8.9%). Particularly noteworthy is InstructOR’s performance compared to the previous state-of-the-art model, Sent-T5-XXL (58.4 vs. 56.5 on average), despite the fact that InstructOR has one order of magnitude fewer parameters (335M vs. 4.8B).

As expected, the retrieval-based models (e.g., GTR-XXL) show strong performance on retrieval and reranking but significantly lag behind on STS and classification. Conversely, similarity-based models (e.g., Sent-T5-XXL) perform well on STS, classification, and text evaluation, but not on retrieval. It suggests that these baselines tend to generate specialized embeddings that only excel at certain tasks, while InstructOR provides universal embeddings that perform well on diverse task categories.

Analysis and Ablations

We demonstrate InstructOR enables universal text embeddings for many diverse tasks. Here we analyze our results from various perspectives: the importance of instructions (§4.1), instruction robustness (§4.2) and complexity (§4.3), model sizes (§4.4), domain shifts (§4.5), and qualitative analysis (§4.6). By default, we report average performance across all categories.

Here we analyze the importance of instructions when training data are diverse. We first split MEDI into symmetric (e.g., text similarity) and asymmetric groups (e.g., open-domain QA), as defined in §2.3 (see Table §5 in the appendix for details about the symmetric and asymmetric groups). We then train InstructOR with or without instructions on each group separately.

As shown in Fig. 3, InstructOR finetuned without instructions yields performance similar to or better than the original GTR model (dotted line), if the data are symmetric or asymmetric only. However, InstructOR suffers if finetuned without task instructions on the combination of both types of data (entire MEDI). In contrast, finetuning with instructions enables the model to benefit from the combination of symmetric and asymmetric data (see that the rightmost bar gets additive performance gains from the asymmetric and symmetric tasks). This result demonstrates the importance of instruction finetuning when diverse data are used for embedding training. Note that training on symmetric tasks only without instructions is similar to Sent-T5. Similarly, training on asymmetric tasks only without instructions is similar to GTR, which is also trained on asymmetric open-domain QA datasets. Departing from these prior methods, instruction-based finetuning enables diverse training on both types.

2 Instruction Robustness

Previous work Sanh et al. (2022); Zhou et al. (2022) shows that instruction-finetuned language models are not robust to paraphrased instructions. Here we measure InstructOR’s robustness to variation in human-written instructions.

Specifically, we write five paraphrased instructions for all evaluation datasets (Table 6 in Appendix) and measure InstructOR’s performance gap between the best-performing and the worst-performing instructions. Fig. 4 shows that inclusion of 300 super-NI datasets is critical to the robustness of InstructOR. Removing these datasets from training (w/o super-NI) substantially increases the performance gap between the best- and worst-performing instructions, suggesting that super-NI’s diverse instructions help the model handle different formats and styles.

3 Complexity of Instructions

Here we further analyze the role of instructions over varying degrees of their complexity. Specifically, we consider four levels of instruction complexity: N/A (no instructions), dataset tags, simple instructions, and detailed instructions (the original instruction format, §2.3). In the dataset tag setup, each example is prepended with its dataset name. For instance, on the Natural Questions dataset, the query is formatted as "Natural Questions; Input: who sings the song Love Story"). In the simple instruction setup, we use one or two words to describe the domain (e.g., for Natural Questions, the input query is Wikipedia Questions; Input: who sings the song Love Story). Fig. 5 shows their average performances across all task categories. Even with trivial dataset tags, InstructOR outperforms the original GTR model, illustrating the effectiveness of instructions for diverse training. As more information is provided in the instruction (from tag to simple and from simple to detail), we observe consistent improvements.

4 Model Sizes and Instruction Finetuning

Fig. 6 studies the influence of model sizes. Specifically, we use GTR-Base (0.1B), GTR-Large (0.3B), and GTR-XL (1.5B). They are pretrained on the same corpus and differ only in the encoder size (the embedding sizes are the same). We compare models of various sizes and report the average performance across all the categories. As the encoder transformer model scales up, the performance continues to increase for both GTR and InstructOR. Nonetheless, the improvement in InstructOR is more pronounced, perhaps because embeddings with instructions benefit from larger capacities. This implies that large models are more generalizable to compute texts in various domains and task types, providing embeddings for general purposes. Further scale-ups are left to future work.

5 Instructions Mitigate Domain Shifts

One advantage of instruction-based finetuning is that it improves models’ ability to generalize to unseen domains and tasks. To demonstrate this effectiveness, we found three unseen domains that InstructOR was not trained on: geography, biology, and civil comments. As shown in Table 3, InstructOR largely improves (above the average improvement) GTR-Large’s performance on all three domains, indicating that instructions can help more when applying models to unseen or uncommon domains.

6 Qualititive Analysis

In this qualitative analysis, we use T-SNE van der Maaten and Hinton (2008) to visualize two example of classification with and without instructions. The desired outcome is, for pairs with the same sentiment to be closer together, and pairs with different sentiment to be farther apart. As shown in Fig. 7, without instructions, the green dot pairs (different sentiment) are closer together in the embedding space, while the red dot pairs (same sentiment) are farther apart. However, with instructions, our method (InstructOR) successfully encodes the red dot pairs into close embeddings and correctly classifies the pairs. The distance between the green dot pairs with different sentiment is also larger in the embedding space with instructions.

Related Work

Text embeddings are useful in many applications such as information retrieval Thakur et al. (2021), text similarity Gao et al. (2021), prompt retrieval for in-context learning Su et al. (2022), classification Reimers and Gurevych (2019), and beyond. Much prior work develops different embedding models for different applications. For example, SBERT Reimers and Gurevych (2019) and SimCSE Gao et al. (2021) are applied solely to text similarity and classification tasks, while DPR Karpukhin et al. (2020) and Contriever Izacard et al. (2022) focus on information retrieval. Different from Sentence-T5 trained only on symmetric data or GTR trained only on asymmetric data, we combine both groups of datasets and build MEDI, which is then used to train InstructOR with instructions. Muennighoff et al. (2022) introduced the massive text embedding benchmark (MTEB), which can be used to evaluate embedding models on a variety of embedding tasks, spanning reranking, classification, information retrieval, bitext mining, pair classification, STS, and summarization. Their benchmark shows that models performing well on one task may not perform well on other tasks. The poor zero-shot transfer abilities of existing embedding models make it difficult to use them in applications where only few labeled data are available. This motivates us to develop a single embedding model that is applicable to a variety of tasks and has better generalization to unseen tasks. Wang et al. (2022a) recently proposed E5, weakly-supervised contrastive pre-trained text embeddings, which achieve strong performance across various tasks on the MTEB benchmark, employing a larger embedding dimension compared to InstructOR.

Instruction Finetuning

Recent work demonstrated that instruction-finetuned language models could perform new tasks given a natural language instruction Mishra et al. (2022); Zhong et al. (2021); Min et al. (2022); Sanh et al. (2022); Wei et al. (2022); Wang et al. (2022b); Ouyang et al. (2022). Nonetheless, instruction finetuning has yet to be studied in the context of broadly-applicable embeddings. In this work, we explore finetuning embedding models to follow human instructions where the instruction specifies eventual use cases. Concurrent work demonstrated that instructions could facilitate information retrieval Asai et al. (2022), which is related to our InstructOR design. They used instructions to build a task-aware retrieval system and conducted evaluations on the retrieval task; we build a general-purpose embedding model with instructions that can be applied to 8 tasks categories (Fig. 2), including retrieval, text similarity, clustering, and text evaluation.

Conclusion

We introduced InstructOR, a single model that creates broadly-applicable text embeddings using natural language instructions. We constructed MEDI, a collection of diverse datasets, to finetune InstructOR with instructions. Our extensive experiments showed that InstructOR achieves state-of-the-art performance on text embedding benchmarks, as well as prompt retrieval for few-shot in-context learning. We hope that researchers and practitioners will benefit from our embeddings or our datasets for tasks of their interest.

Limitations

Although InstructOR significantly improves the baseline GTR performance, we were only able to use four negative examples during the model finetuning process due to computation constraints. However, negative examples have been shown to play an important role in contrastive learning Robinson et al. (2021). We hope that future work will scale up the number of negatives used during finetuning and investigate various methods for mining hard negatives. Additionally, we do not have enough computation resources to apply multitask instruction finetuning to GTR-XXL (4.8B parameters), which is also an area for future exploration.

At the core of InstructOR is the instruction design. While our current unified instruction format has demonstrated effectiveness, future research can explore other instructional elements to further improve performance. For example, previous work Wang et al. (2022b) have shown that incorporating demonstration examples and explanations can be beneficial for instruction-finetuned language models.

Acknowledgements

We thank Akari Asai, Jack Lin, Minghan Li, and the ARK group at UW for their helpful feedback on this work.

References

Appendix A Training Setups

Training is performed on a combination of all training datasets in MEDI. Since the number of examples in each dataset is different in orders of magnitude, we downsample large ones. Details for the downsampled numbers of examples on each dataset are shown in Table 5 in the appendix. At each step, we first randomly select a dataset and then construct a minibatch only using the examples from that dataset. In this way, we ensure that in-batch negatives are sampled from the same dataset, thereby preventing the model from using task differences to predict the negative label. We use the maximum batch size that fits the machine memory and run all our experiments on 40GB A100 GPUs.

Training

We initialize InstructOR with the GTR-Large model (Ni et al., 2021, 335M parameters)https://huggingface.co/sentence-transformers/gtr-t5-large. and finetune it on MEDI using the AdamW optimizer with learning rate 2×10−52\times 10^{-5} and warmup ratio 0.1. We use a softmax temperature of 0.01 and finetune InstructOR for 20K steps.

Baselines

We use the official MTEB benchmark for comparisons, but here we highlight several strong baselines with the following two types. The first class of baselines is embedding models specializing in information retrieval: Contriever-MS Izacard et al. (2022), GTR Ni et al. (2021), and coCondenser-MS Gao and Callan (2022). They are all trained on open-domain QA datasets such as MS MARCO Bajaj et al. (2016). The second class of baselines focuses on semantic textual similarity: SimCSE Gao et al. (2021), Sent-T5 Ni et al. (2022), and SGPT-NLI Muennighoff (2022). They are mainly trained on symmetric paraphrase datasets such as NLI Williams et al. (2018) and the Quora question pairs.https://www.quora.com/q/quoradata/. All of these baselines are based on pretrained language models, achieving strong performance on the MTEB leaderboard. In particular, Sent-T5-XXL and GTR-XXL (both with 4.8B parameters) achieve the first and second best average performances.

Appendix B Embedding Evaluations

Here we provide a high-level summary of the evaluation tasks (Table 1). Following MTEB Muennighoff et al. (2022), Billboard Kasai et al. (2022a), and prompt retrieval Su et al. (2022), we split 70 evaluation datasets into 9 categories by task objectives. Out of the 70 evaluation tasks, 66 are unseen during training (See Table 5 for datasets included during training), Table 1 for examples and instructions for the evaluation datasets.

MTEB Muennighoff et al. (2022) is a comprehensive embedding evaluation benchmark that aims to provide a holistic view of current embedding models’ performance and to discover universal text embeddings applicable to a wide range of tasks. It combines several conventional benchmarks (e.g., BEIR, Thakur et al., 2021, and STS, Cer et al., 2017) and spans a wide range of domain-specific datasets, including science, biology, and medicine. Following Muennighoff et al. (2022), we also report the average performance over 56 datasets. For each task family, we briefly describe the task objective, evaluation metric, and how embeddings are used.

Given a query qq and a corpus D={p1,p2...pn}D=\{p_{1},p_{2}...p_{n}\}, retrieval aims to find the most relevant documents pip_{i} in DD for query qq. The embedding model is used to embed qq and p1...pnp_{1}...p_{n} into fixed-sized vectors, and then the similarity between qq and pip_{i} is measured by their embedding cosine similarity. There are 14 diverse datasets (e.g., Natural Questions, Scifact, and NFCorpus) together with the community question-answering (CQA) benchmark Hoogeveen et al. (2015). We use NDCG@10 (Normalized Discounted cumulative gain at rank position 10) to measure the performance.

Reranking

Reranking ranks a list of documents based on their relevance to a query. Given a query qq and a list of documents D={p1,p2...pn}D=\{p_{1},p_{2}...p_{n}\}, the embedding model computes embeddings of both the query and documents, which are then used to rank the documents based on their cosine similarities. We use MAP (mean average precision), a standard metric in reranking, to measure performance.

Clustering

The goal of clustering is to group similar documents into meaningful clusters. Given a set of documents, the encoder maps each document into an embedding. The k-means clustering algorithm is then used to partition the embedded documents into clusters. The clustering performance is measured by the v-measure that is independent of the permutations of clustering labels Rosenberg and Hirschberg (2007).

Pair Classification

Pair classification tasks aim to predict a binary label for a pair of texts. An example of this task is paraphrase identification, where the goal is to predict whether two sentences are paraphrases of each other. Given a sentence pair (t1,t2)(t_{1},t_{2}), the embedding model encodes t1t_{1} and t2t_{2} separately. The cosine similarity between the two embeddings is then used to predict the label. The average precision score is measured for evaluation.

Classification

Classification is a popular way to evaluate the quality of embeddings Conneau and Kiela (2018). For each example in the classification dataset, the embedding of the input text is used as features to a classifier. The classifier is trained on the training data while sentence embedings are kept frozen. We report the classification accuracy on the test set as the evaluation metric.

STS

Semantic textual similarity (STS) tasks evaluate the similarity between two sentences. Given a sentence pair (t1,t2)(t_{1},t_{2}), the embedding model maps t1t_{1} and t2t_{2} into embeddings separately, and then the similarity between t1t_{1} and t2t_{2} is measured by their embedding cosine similarity. The evaluation metric is Spearman’s rank correlation, which measures the correlation between the similarity scores and human judgements.

Summarization

Automatic summarization evaluation aims to evaluate the quality of a machine-generated summary given a reference summary. While human evaluations are considered more accurate, automatic evaluations allow for fast, inexpensive development cycles Khashabi et al. (2022). Given a reference summary rr and a machine-generated summary tt, the embedding model maps them into embeddings separately, and we compute the cosine similarity between rr and tt. Spearman’s rank correlation is reported between human judgements and automatic scores.

B.2 Prompt Retrieval

Large language models have demonstrated the ability of in-context learning, where the model can perform downstream tasks by conditioning generation on a few task demonstrations Liu et al. (2021). Su et al. (2022) introduce the prompt retrieval task, where the goal is to retrieve a few in-context learning (i.e., demonstration) examples from annotated examples given a test instance. The embedding model is used to encode all annotated examples and to find the few most similar examples to the test instance based on the cosine similarity. Following Su et al. (2022), we use the retrieved examples for in-context learning on GPT-J Wang and Komatsuzaki (2021) over 11 diverse downstream tasks (e.g., classification, multiple choice, and text-to-SQL) that are not included in MEDI (thus zero-shot settings). We compare different embedding methods by measuring the average performance on these downstream tasks.

B.3 Automatic Evaluation for Generation

Similar to summarization evaluation in MTEB, we use the Billboard benchmark Kasai et al. (2022a) to apply InstructOR to automatic evaluations for three additional text generation tasks: MSCOCO image captioning Lin et al. (2014); Kasai et al. (2022b), CNN/DailyMail news summarization Fabbri et al. (2021), and WMT21 Chinese-to-English translation Barrault et al. (2020); Freitag et al. (2021). Following Kasai et al. (2022a), we measure the cosine similarity between the generated text and each reference text and take the maximum similarity score over all references available Zhang et al. (2020). We evaluate all embedding models by the Pearson correlation with the human judgments, again following Kasai et al. (2022a). We then report the average correlation scores over the three datasets. Note that we do not use the English-to-German dataset in Billboard because our models are trained only on English data.

Appendix C Full instructions

We list all instructions for each dataset in MEDI in Table 7 and Table 8

Appendix D Full Results

We provide the detailed evaluation scores in MTEB, Billboard and prompt retrieval benchmarks in Table 9 & 10.