MTEB: Massive Text Embedding Benchmark
Niklas Muennighoff, Nouamane Tazi, Loïc Magne, Nils Reimers
Introduction
Natural language embeddings power a variety of use cases from clustering and topic representation Aggarwal and Zhai (2012); Angelov (2020) to search systems and text mining Huang et al. (2020); Zhu et al. (2021); Nayak (2019) to feature representations for downstream models Saharia et al. (2022); Borgeaud et al. (2022). Using generative language models or cross-encoders for these applications is often intractable, as they may require exponentially more computations Reimers and Gurevych (2019).
However, the evaluation regime of current text embedding models rarely covers the breadth of their possible use cases. For example, SimCSE Gao et al. (2021b) or SBERT Reimers and Gurevych (2019) solely evaluate on STS and classification tasks, leaving open questions about the transferability of the embedding models to search or clustering tasks. STS is known to poorly correlate with other real-world use cases Neelakantan et al. (2022); Wang et al. (2021). Further, evaluating embedding methods on many tasks requires implementing multiple evaluation pipelines. Implementation details like pre-processing or hyperparameters may influence the results making it unclear whether performance improvements simply come from a favorable evaluation pipeline. This leads to the “blind” application of these models to new use cases in industry or requires incremental work to reevaluate them on different tasks.
The Massive Text Embedding Benchmark (MTEB) aims to provide clarity on how models perform on a variety of embedding tasks and thus serves as the gateway to finding universal text embeddings applicable to a variety of tasks. MTEB consists of 58 datasets covering 112 languages from 8 embedding tasks: Bitext mining, classification, clustering, pair classification, reranking, retrieval, STS and summarization. MTEB software is available open-sourcehttps://github.com/embeddings-benchmark/mteb enabling evaluation of any embedding model by adding less than 10 lines of code. Datasets and the MTEB leaderboard are available on the Hugging Face Hubhttps://huggingface.co/spaces/mteb/leaderboard.
We evaluate over 30 models on MTEB with additional speed and memory benchmarking to provide a holistic view of the state of text embedding models. We cover both models available open-source as well as models accessible via APIs, such as the OpenAI Embeddings endpoint. We find there to be no single best solution, with different models dominating different tasks. Our benchmarking sheds light on the weaknesses and strengths of individual models, such as SimCSE’s Gao et al. (2021b) low performance on clustering and retrieval despite its strong performance on STS. We hope our work makes selecting the right embedding model easier and simplifies future embedding research.
Related Work
Benchmarks, such as (Super)GLUE Wang et al. (2018, 2019) or Big-BENCH Srivastava et al. (2022), and evaluation frameworks Gao et al. (2021a) play a key role in driving NLP progress. Yearly released SemEval datasets Agirre et al. (2012, 2013, 2014, 2015, 2016) are commonly used as the go-to benchmark for text embeddings. SemEval datasets correspond to the task of semantic textual similarity (STS) requiring models to embed similar sentences with geometrically close embeddings. Due to the limited expressivity of a single SemEval dataset, SentEval Conneau and Kiela (2018) aggregates multiple STS datasets. SentEval focuses on fine-tuning classifiers on top of embeddings. It lacks tasks like retrieval or clustering, where embeddings are directly compared without additional classifiers. Further, the toolkit was proposed in 2018 and thus does not provide easy support for recent trends like text embeddings from transformers Reimers and Gurevych (2019). Due to the insufficiency of STS benchmarking, USEB Wang et al. (2021) was introduced consisting mostly of reranking tasks. Consequently, it does not cover tasks like retrieval or classification. Meanwhile, the recently released BEIR Benchmark Thakur et al. (2021) has become the standard for the evaluation of embeddings for zero-shot information retrieval.
MTEB unifies datasets from different embedding tasks into a common, accessible evaluation framework. MTEB incorporates SemEval datasets (STS11 - STS22) and BEIR alongside a variety of other datasets from various tasks to provide a holistic performance review of text embedding models.
2 Embedding Models
Text embedding models like Glove Pennington et al. (2014) lack context awareness and are thus commonly labeled as Word Embedding Models. They consist of a layer mapping each input word to a vector often followed by an averaging layer to provide a final embedding invariant of input length. Transformers Vaswani et al. (2017) inject context awareness into language models via self-attention and form the foundation of most recent embedding models. BERT Devlin et al. (2018) uses the transformer architecture and performs large-scale self-supervised pre-training. The resulting model can directly be used to produce text embeddings via an averaging operation alike Glove. Building on InferSent Conneau et al. (2017), SBERT Reimers and Gurevych (2019) demonstrated it to be beneficial to perform additional fine-tuning of the transformer for competitive embedding performance. Most recent fine-tuned embedding models use a contrastive loss objective to perform supervised fine-tuning on positive and negative text pairs Gao et al. (2021b); Wang et al. (2021); Ni et al. (2021b); Muennighoff (2022). Due to the large variety of available pre-trained transformers Wolf et al. (2020), there is an at least equally large variety of potential text embedding models to be explored. This leads to confusion about which model provides practitioners with the best performance for their embedding use case.
We benchmark both word embedding and transformer models on MTEB quantifying gains provided by often much slower context aware models.
The MTEB Benchmark
MTEB is built on a set of desiderata: (a) Diversity: MTEB aims to provide an understanding of the usability of embedding models in various use cases. The benchmark comprises 8 different tasks, with up to 15 datasets each. Of the 58 total datasets in MTEB, 10 are multilingual, covering 112 different languages. Sentence-level and paragraph-level datasets are included to contrast performance on short and long texts. (b) Simplicity: MTEB provides a simple API for plugging in any model that given a list of texts can produce a vector for each list item with a consistent shape. This makes it possible to benchmark a diverse set of models. (c) Extensibility: New datasets for existing tasks can be benchmarked in MTEB via a single file that specifies the task and a Hugging Face dataset name where the data has been uploaded Lhoest et al. (2021). New tasks require implementing a task interface for loading the data and an evaluator for benchmarking. We welcome dataset, task or metric contributions from the community via pull requests to continue the development of MTEB. (d) Reproducibility: Through versioning at a dataset and software level, we aim to make it easy to reproduce results in MTEB. JSON files corresponding to all results available in this paper have been made available together with the MTEB benchmarkhttps://huggingface.co/datasets/mteb/results.
2 Tasks and Evaluation
Figure 1 provides an overview of tasks and datasets available in MTEB. Dataset statistics are available in Table 2. The benchmark consists of the following 8 task types:
Inputs are two sets of sentences from two different languages. For each sentence in the first set, the best match in the second set needs to be found. The matches are commonly translations. The provided model is used to embed each sentence and the closest pairs are found via cosine similarity. F1 serves as the main metric for bitext mining. Accuracy, precision and recall are also computed.
Classification
A train and test set are embedded with the provided model. The train set embeddings are used to train a logistic regression classifier with 100 maximum iterations, which is scored on the test set. The main metric is accuracy with average precision and f1 additionally provided.
Clustering
Given a set of sentences or paragraphs, the goal is to group them into meaningful clusters. A mini-batch k-means model with batch size 32 and k equal to the number of different labels Pedregosa et al. (2011) is trained on the embedded texts. The model is scored using v-measure Rosenberg and Hirschberg (2007). V-measure does not depend on the cluster label, thus the permutation of labels does not affect the score.
Pair Classification
A pair of text inputs is provided and a label needs to be assigned. Labels are typically binary variables denoting duplicate or paraphrase pairs. The two texts are embedded and their distance is computed with various metrics (cosine similarity, dot product, euclidean distance, manhattan distance). Using the best binary threshold accuracy, average precision, f1, precision and recall are computed. The average precision score based on cosine similarity is the main metric.
Reranking
Inputs are a query and a list of relevant and irrelevant reference texts. The aim is to rank the results according to their relevance to the query. The model is used to embed the references which are then compared to the query using cosine similarity. The resulting ranking is scored for each query and averaged across all queries. Metrics are mean MRR@k and MAP with the latter being the main metric.
Retrieval
Each dataset consists of a corpus, queries and a mapping for each query to relevant documents from the corpus. The aim is to find these relevant documents. The provided model is used to embed all queries and all corpus documents and similarity scores are computed using cosine similarity. After ranking the corpus documents for each query based on the scores, nDCG@k, MRR@k, MAP@k, precision@k and recall@k are computed for several values of . nDCG@10 serves as the main metric. MTEB reuses datasets and evaluation from BEIR Thakur et al. (2021).
Semantic Textual Similarity (STS)
Given a sentence pair the aim is to determine their similarity. Labels are continuous scores with higher numbers indicating more similar sentences. The provided model is used to embed the sentences and their similarity is computed using various distance metrics. Distances are benchmarked with ground truth similarities using Pearson and Spearman correlations. Spearman correlation based on cosine similarity serves as the main metric Reimers et al. (2016).
Summarization
A set of human-written and machine-generated summaries are provided. The aim is to score the machine summaries. The provided model is first used to embed all summaries. For each machine summary embedding, distances to all human summary embeddings are computed. The closest score (e.g. highest cosine similarity) is kept and used as the model’s score of a single machine-generated summary. Pearson and Spearman correlations with ground truth human assessments of the machine-generated summaries are computed. Like for STS, Spearman correlation based on cosine similarity serves as the main metric Reimers et al. (2016).
3 Datasets
To further the diversity of MTEB, datasets of varying text lengths are included. All datasets are grouped into three categories:
A sentence is compared with another sentence. An example of S2S are all current STS tasks in MTEB, where the similarity between two sentences is assessed.
Paragraph to paragraph (P2P)
A paragraph is compared with another paragraph. MTEB imposes no limit on the input length, leaving it up to the models to truncate if necessary. Several clustering tasks are framed as both S2S and P2P tasks. The former only compare titles, while the latter include both title and content. For ArxivClustering, for example, abstracts are concatenated to the title in the P2P setting.
Sentence to paragraph (S2P)
A few retrieval datasets are mixed in a S2P setting. Here a query is a single sentence, while documents are long paragraphs consisting of multiple sentences.
Similarities across 56 MTEB datasets are visualized in Figure 2. Several datasets rely on the same corpora, such as ClimateFEVER and FEVER, resulting in a score of 1. Clusters of similar datasets can be seen among CQADupstack variations and STS datasets. S2S and P2P variations of the same dataset tend to also be similar. Scientific datasets, such as SciDocsRR, SciFact, ArxivClustering, show high similarities among each other even when coming from different tasks (Reranking, Retrieval and Clustering in this case).
Results
We evaluate on the test splits of all datasets except for MSMARCO, where the dev split is used following Thakur et al. (2021). We benchmark models claiming state-of-the-art results on various embedding tasks leading to a high representation of transformers Vaswani et al. (2017). We group models into self-supervised and supervised methods.
(a) Transformer-based BERT Devlin et al. (2018) is trained using self-supervised mask and sentence prediction tasks. By taking the mean across the sequence length (mean-pooling) the model can directly be used to produce text embeddings. SimCSE-Unsup Gao et al. (2021b) uses BERT as a foundation and performs additional self-supervised training. (b) Non-transformer: Komninos Komninos and Manandhar (2016) and Glove Pennington et al. (2014) are two word embedding models that directly map words to vectors. Hence, their embeddings lack context awareness, but provide significant speed-ups.
Supervised methods
The original transformer model Vaswani et al. (2017) consists of an encoder and decoder network. Subsequent transformers often train only encoders like BERT Devlin et al. (2018) or decoders like GPT Radford et al. (2019).
(a) Transformer encoder methods coCondenser Gao and Callan (2021), Contriever Izacard et al. (2021), LaBSE Feng et al. (2020) and SimCSE-BERT-sup Gao et al. (2021b) are based on the pre-trained BERT model Devlin et al. (2018). coCondenser and Contriever add a self-supervised stage prior to supervised fine-tuning for a total of three training stages. LaBSE uses BERT to perform additional pre-training on parallel data to produce a competitive bitext mining model. SPECTER Cohan et al. (2020a) relies on the pre-trained SciBERT Beltagy et al. (2019) variant instead and fine-tunes on citation graphs. GTR Ni et al. (2021b) and ST5 Ni et al. (2021a) are based on the encoder part of the T5 model Raffel et al. (2020) and only differ in their fine-tuning datasets. After additional self-supervised training, ST5 does contrastive fine-tuning on NLI Ni et al. (2021a); Gao et al. (2021b) being geared towards STS tasks. Meanwhile, GTR fine-tunes on MSMARCO and focuses on retrieval tasks. MPNet and MiniLM correspond to fine-tuned embedding models Reimers and Gurevych (2019) of the pre-trained MPNet Song et al. (2020) and MiniLM Wang et al. (2020) models using diverse datasets to target any embedding use case.
(b) Transformer decoder methods SGPT Bi-Encoders Muennighoff (2022) perform contrastive fine-tuning of <0.1% of pre-trained parameters using weighted-mean pooling. Similar to ST5 and GTR, SGPT-nli models are geared towards STS, while SGPT-msmarco models towards retrieval. SGPT-msmarco models embed queries and documents for retrieval with different special tokens to help the model distinguish their role. For non-retrieval tasks, we use its query representations. We benchmark publicly available SGPT models based on GPT-NeoX Andonian et al. (2021), GPT-J Wang and Komatsuzaki (2021) and BLOOM Scao et al. (2022). Alternatively, cpt-text Neelakantan et al. (2022) passes pre-trained GPT decoders through a two-stage process using last token pooling to provide embeddings from decoders. We benchmark their models via the OpenAI Embeddings APIhttps://beta.openai.com/docs/guides/embeddings.
(c) Non-transformer LASER Heffernan et al. (2022) is the only context aware non-transformer model we benchmark, relying on an LSTM Hochreiter and Schmidhuber (1997) instead. Similar to LaBSE, the model trains on parallel data and focuses on bitext mining applications.
2 Analysis
Based on the results in Table 1, we observe that there is considerable variability between tasks. No model claims the state-of-the-art in all seven English tasks. There is even more variability in the results per dataset present in the appendix. Further, there remains a large gap between self-supervised and supervised methods. Self-supervised large language models have been able to close this gap in many natural language generation tasks Chowdhery et al. (2022). However, they appear to still require supervised fine-tuning for competitive embedding performance.
We find that performance strongly correlates with model size, see Figure 3. A majority of MTEB tasks are dominated by multi-billion parameter models. However, these come at a significant cost as we investigate in Section 4.3.
ST5 models dominate the classification task across most datasets, as can be seen in detail in the full results in the appendix. ST5-XXL has the highest average performance, 3% ahead of the best non-ST5 model, OpenAI Ada Similarity.
Clustering
Despite being almost 50x smaller, the MPNet embedding model is on par with the ST5-XXL state-of-the-art on Clustering. This may be due to the large variety of datasets MPNet (and MiniLM) has been fine-tuned on. Clustering requires coherent distances between a large number of embeddings. Models like SimCSE-sup or SGPT-nli, which are only fine-tuned on a single dataset, NLI, may produce incoherent embeddings when encountering topics unseen during fine-tuning. Relatedly, we find that the query embeddings of SGPT-msmarco and the Ada Search endpoint are competitive with SGPT-nli and the Ada Similarity endpoint, respectively. We refer to the public leaderboardhttps://huggingface.co/spaces/mteb/leaderboard for Ada Search results. This could be due to the MSMARCO dataset being significantly larger than NLI. Thus, while the OpenAI docs recommend using the similarity embeddings for clustering use caseshttps://beta.openai.com/docs/guides/embeddings/similarity-embeddings, the retrieval query embeddings may be the better choice in some cases.
Pair Classification
GTR-XL and GTR-XXL have the strongest performance. Pair classification is closest to STS in its framing, yet models rank significantly differently on the two tasks. This highlights the importance of benchmarking on a diverse set of tasks to avoid blindly reusing a model for a different task.
Reranking
MPNet and MiniLM models perform strongly on reranking tasks. On SciDocsRR Cohan et al. (2020a) they perform far better than bigger models, which is likely due to parts of SciDocsRR being included in their training data. Our scale of experiments and that of model pre-training make controlling for data contamination challenging. Thus, we ignore overlap of MTEB datasets with model training datasets in MTEB scores. As long as enough datasets are averaged, we believe these effects to be insignificant.
Retrieval
SGPT-5.8B-msmarco is the best embedding model on the BEIR subset in MTEB as well as on the full BEIR benchmark Thakur et al. (2021); Muennighoff (2022). The even larger 7.1B SGPT model making use of BLOOM Scao et al. (2022) performs significantly weaker, which is likely due to the multilinguality of BLOOM. Models geared towards STS (SimCSE, ST5, SGPT-nli) perform badly on retrieval tasks. Retrieval tasks are unique in that there are two distinct types of texts: Queries and documents (“asymmetric”), while other tasks only have a single type of text (“symmetric”). On the QuoraRetrieval dataset, which has been shown to be largely symmetric Muennighoff (2022), the playing field is more even with SGPT-5.8B-nli outperforming SGPT-5.8B-msmarco, see Table 11.
STS & Summarization
Retrieval models (GTR, SGPT-msmarco) perform badly on STS, while ST5-XXL has the highest performance. This highlights the bifurcation of the field into separate embedding models for retrieval (asymmetric) and similarity (symmetric) use cases Muennighoff (2022).
3 Efficiency
We investigate the latency-performance trade-off of models in Figure 4. The graph allows for significant elimination of model candidates in the model selection process. It brings model selection down to three clusters:
Word Embedding models offer maximum speed with Glove taking the lead on both performance and speed, thus making the choice simple in this case.
Maximum performance
If latency is less important than performance, the left-hand side of the graph offers a cluster of highly performant, but slow models. Depending on the task at hand, GTR-XXL, ST5-XXL or SGPT-5.8B may be the right choice, see Section 4.2. SGPT-5.8B comes with the additional caveat of its high-dimensional embeddings requiring more storage.
Speed and performance
The fine-tuned MPNet and MiniLM models lead the middle cluster making the choice easy.
4 Multilinguality
MTEB comes with 10 multilingual datasets across bitext mining, classification and STS tasks. We investigate performance on these in Figure 5. Tabular results can be found in Tables 12, 13 and 14.
LaBSE Feng et al. (2020) performs strongly across a wide array of languages in bitext mining. Meanwhile, LASER2 shows high variance across different languages. While there are additional language-specific LASER2 models available for some of the languages we benchmark, we use the default multilingual LASER2 model for all languages. This is to provide a fair one-to-one comparison of models. In practice, however, the high variance of LASER2’s performance may be resolved by mixing its model variants. MPNet, MiniLM and SGPT-BLOOM-7B1-msmarco perform poorly on languages they have not been pre-trained on, such as German for the latter.
Classification & STS
On multilingual classification and STS, the multilingual MPNet provides the overall strongest performance. It outperforms the slightly faster multilingual MiniLM on almost all languages. Both models have been trained on the same languages, thus bringing decision-making down to performance vs speed. SGPT-BLOOM-7B1-msmarco provides state-of-the-art performance on languages like Hindi, Portuguese, Chinese or French, which the model has seen extensively during pre-training. It also performs competitively on languages like Russian or Japanese that unintentionally leaked into its pre-training data Muennighoff et al. (2022). However, it is not much ahead of the much cheaper MPNet. LASER2 performs consistently worse than other models.
Conclusion
In this work, we presented the Massive Text Embedding Benchmark (MTEB). Consisting of 8 text embedding tasks with up to 15 datasets each and covering 112 languages, MTEB aims to provide reliable embedding performance estimates. By open-sourcing MTEB alongside a leaderboard, we provide a foundation for further pushing the state-of-the-art of available text embeddings.
To introduce MTEB, we have conducted the most comprehensive benchmarking of text embeddings to date. Through the course of close to 5,000 experiments on over 30 different models, we have set up solid baselines for future research to build on. We found model performance on different tasks to vary strongly with no model claiming state-of-the-art on all tasks. Our studies on scaling behavior, model efficiency and multilinguality revealed various intricacies of models that should ease the decision-making process for future research or industry applications of text embeddings.
We welcome task, dataset or metric contributions to the MTEB codebasehttps://github.com/embeddings-benchmark/mteb as well as additions to the leaderboard via our automatic submission formathttps://huggingface.co/spaces/mteb/leaderboard.
Acknowledgments
This work was granted access to the HPC resources of Institut du développement et des ressources en informatique scientifique (IDRIS) du Centre national de la recherche scientifique (CNRS) under the allocation 2021-A0101012475 made by Grand équipement national de calcul intensif (GENCI). In particular, all the evaluations and data processing ran on the Jean Zay cluster of IDRIS, and we want to thank the IDRIS team for responsive support throughout the project, in particular Rémi Lacroix.
We thank Douwe Kiela, Teven Le Scao and Nandan Thakur for feedback and suggestions.
References
Appendix A Datasets
Table 2 provides a summary along with statistics of all MTEB tasks. In the following, we give a brief description of each dataset included in MTEB.
These datasets are custom-made for MTEB using the public APIs from arXivhttps://arxiv.org/help/api/ and bioRxiv/medRxivhttps://api.biorxiv.org/. For S2S datasets, the input text is simply the title of the paper, while for P2P the input text is the concatenation of the title and the abstract. The cluster labels are generated using categories given to the papers by humans. For bioRxiv and medRxiv this category is unique, but for arXiv multiple categories can be given to a single paper so we only use the first one. For bioRxiv and medRxiv there is only one level of category (e.g. biochemistry, genetics, microbiology, etc.) hence we only perform clustering based on that label. For arXiv there is a main category and secondary category: for example "cs.AI" means the main category is Computer Science and the sub-category is AI, math.AG means the main category is Mathematics and the sub-category is Algrebraic Geometry etc. Hence, we create three types of splits:
(a) Main category clustering
Articles are only clustered based on the main category (Math, Physics, Computer Science etc.). This split evaluates coarse clustering capacity of a model.
(b) Secondary category clustering within the same main category
Articles are clustered based on their secondary category, but within a given main category, for example only Math papers that need to be clustered into Algebraic Geometry, Functional Analysis, Numerical Analysis etc. This split evaluates fine-grained clustering capacity of a model, as differentiating some sub-categories can be very difficult.
(c) Secondary category clustering
Articles are clustered based on their secondary category for all main categories, so the labels can be Number Theory, Computational Complexity, Astrophysics of Galaxies etc. These splits evaluate fine-grained clustering capacity, as well as multi-scale capacities i.e. is a model able to both separate Maths from Physics as well as Probability from Algebraic Topology at the same time.
For every dataset, split and strategy, we select subsets of all labels and then sample articles from those labels. This yields splits with a varying amount and size of clusters.
RedditClustering
Geigle et al. (2021): Clustering of titles from 199 subreddits. Clustering of 25 splits, each with 10-50 classes, and each class with 100 - 1000 sentences
RedditClusteringP2P
Dataset created for MTEB using available data from Reddit postshttps://huggingface.co/datasets/sentence-transformers/reddit-title-body. The task consists of clustering the concatenation of title+post according to their subreddit. It contains 10 splits, with 10 and 100 clusters per split and 1,000 to 100,000 posts.
StackExchangeClustering
Geigle et al. (2021) Clustering of titles from 121 stackexchanges. Clustering of 25 splits, each with 10-50 classes, and each class with 100-1000 sentences.
StackExchangeClusteringP2P
Dataset created for MTEB using available data from StackExchange postshttps://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_title_body_jsonl. The task consists of clustering the concatenation of title and post according to their subreddit. It contains 10 splits, with 10 to 100 clusters and 5,000 to 10,000 posts per split.
TwentyNewsgroupsClusteringhttps://scikit-learn.org/0.19/datasets/twenty_newsgroups.html
Clustering of the 20 Newsgroups dataset, given titles of article the goal is to find the newsgroup (20 in total). Contains 10 splits, each with 20 classes, with each split containing between 1,000 and 10,000 titles.
A.2 Classification
O’Neill et al. (2021) A collection of Amazon customer reviews annotated for counterfactual detection pair classification. For each review the label is either "counterfactual" or "not-counterfactual". This is a multilingual dataset with 4 available languages.
AmazonPolarity
McAuley and Leskovec (2013) A collection of Amazon customer reviews annotated for polarity classification. For each review the label is either "positive" or "negative".
AmazonReviews
McAuley and Leskovec (2013) A collection of Amazon reviews designed to aid research in multilingual text classification. For each review the label is the score given by the review between 0 and 4 (1-5 stars). This is a multilingual dataset with 6 available languages.
Banking77
Casanueva et al. (2020) Dataset composed of online banking queries annotated with their corresponding intents. For each user query the label is an intent among 77 intents like ’activate_my_card’, ’apple_pay’, ’bank_transfer’, etc.
Emotion
Saravia et al. (2018) Dataset of English Twitter messages with six basic emotions: anger, fear, joy, love, sadness, and surprise.
Imdb
Maas et al. (2011) Large movie review dataset with labels being positive or negative.
MassiveIntent
FitzGerald et al. (2022) A collection of Amazon Alexa virtual assistant utterances annotated with the associated intent. For each user utterance the label is one of 60 intents like ’play_music’, ’alarm_set’, etc. This is a multilingual dataset with 51 available languages.
MassiveScenario
FitzGerald et al. (2022) A collection of Amazon Alexa virtual assistant utterances annotated with the associated intent. For each user utterance the label is a theme among 60 scenarios like ’music’, ’weather’, etc. This is a multilingual dataset with 51 available languages.
MTOPDomain / MTOPIntent
Multilingual sentence datasets from the MTOP Li et al. (2020) benchmark. We refer to their paper for details.
ToxicConversations
Dataset from Kaggle competitionhttps://www.kaggle.com/competitions/jigsaw-unintended-bias-in-toxicity-classification. Collection of comments from the Civil Comments platform together with annotations if the comment is toxic or not.
TweetSentimentExtraction
Dataset from Kaggle competitionhttps://www.kaggle.com/competitions/tweet-sentiment-extraction. Sentiment classification of tweets as neutral, positive or negative.
A.3 Pair Classification
Shah et al. (2018): Collection of questions from the Sprint community. The goal is to classify a pair of sentences as duplicates or not.
TwitterSemEval2015
Xu et al. (2015) Paraphrase-Pairs of Tweets from the SemEval 2015 workshop. The goal is to classify a pair of tweets as paraphrases or not.
TwitterURLCorpus
Lan et al. (2017) Paraphrase-Pairs of Tweets. The goal is to classify a pair of tweets as paraphrases or not.
A.4 Bitext Mining
Zweigenbaum et al. (2016, 2017, 2018) BUCC provides big set of sentences ( 10-70k each) for English, French, Russian, German and Chinese, along with associated pairs annotation. The annotated pairs here corresponds to a pairs of translated sentences, i.e. a sentence and its translation in the other language.
Tatoeba
Research Tatoeba provides sets of sentences (1000 sentences each) for 112 languages with annoated associated pairs. Each pair is one sentence and its translation in another language.
A.5 Reranking
Questions from AskUbuntu with manual annotations marking pairs of questions as similar or dissimilar.
MindSmall
Wu et al. (2020) Large-scale English Dataset for News Recommendation Research. Ranking news article titles given the title of a news article. The idea is to recommend other news from the one you are reading.
SciDocsRR
Cohan et al. (2020b) Ranking of related scientific papers based on their title.
StackOverflowDupQuestions
Liu et al. (2018) Stack Overflow Duplicate Questions Task for questions with the tags Java, JavaScript and Python, ranking questions as duplicates or not.
A.6 Semantic Textual Similarity (STS)
Agirre et al. (2012, 2013)https://alt.qcri.org/semeval2014/task10/https://alt.qcri.org/semeval2015/task2/https://alt.qcri.org/semeval2016/task1/https://competitions.codalab.org/competitions/33835 Original STS benchmark, with scores from 0 to 5. The selection of sentences includes text from image captions, news headlines and user forums. In total they contain between 1,000 and 20,000 sentences. STS12 - STS16 and STSBenchmark are monolingual english benchmarks. STS17 and STS22 contain crosslingual pairs of sentences, where the goal is to assess the similarity of two sentences in different languages. STS17 has 11 language pairs (among Korean, Arabic, English, French, German, Turkish, Spanish, Italian and Dutch) and STS22 has 18 language pairs (among Arabic, English, French, German, Turkish, Spanish, Polish, Italian, Russian and Chinese).
BIOSSEShttps://tabilab.cmpe.boun.edu.tr/BIOSSES/DataSet.html
Contains 100 sentence pairs from the biomedical field.
SICK-R
Agirre et al. (2014) Sentences Involving Compositional Knowledge (SICK) contains a large number of sentence pairs () that are lexically, syntactically and semantically rich.
A.7 Summarization
Fabbri et al. (2020) Summaries generated by recent summarization models trained on CNN or DailyMail alongside human annotations.
A.8 Retrieval
We refer to the BEIR paper Thakur et al. (2021), which contains description of each dataset. For MTEB, we include all publicly available datasets: ArguAna, ClimateFEVER, CQADupstack, DBPedia, FEVER, FiQA2018, HotpotQA, MSMARCO, NFCorpus, NQ, Quora, SCIDOCS, SciFact, Touche2020, TRECCOVID.
Appendix B Limitations of MTEB
While MTEB aims to be a diverse benchmark to provide holistic performance reviews, the benchmark has its limitations. We list them here:
MTEB covers multiple text lengths (S2S, P2P, S2P), but very long documents are still missing. The longest datasets in MTEB have a few hundred words, and longer text sizes could be relevant for use cases like retrieval.
Task imbalance
Tasks in MTEB have a different amount of datasets with summarization consisting of only a single dataset. This means MTEB average scores, which are computed over all datasets, are biased towards tasks with many datasets, notably retrieval, classification and clustering. As MTEB grows, we hope to add more datasets to currently underrepresented tasks like summarization or pair classification.
Multinguality
MTEB contains multilingual classification, STS and bitext mining datasets. However, retrieval and clustering are English-only. SGPT-BLOOM-7B1-msmarco is geared towards multilingual retrieval datasets and due to the lack thereof cannot be comprehensively benchmarked in MTEB. Further, MTEB does not contain any code datasets that could be used to benchmark code models Neelakantan et al. (2022); Allal et al. (2023). It should be easy to extend MTEB with datasets, such as CodeSearchNet Husain et al. (2019), TyDI QA Clark et al. (2020), XOR QA Asai et al. (2020) or MIRACL Zhang et al. (2022).
Additional modalities
Text embeddings are commonly used as input features for downstream models, such as in our classification task. This can involve other modalities, notably image content Carvalho et al. (2018); Tan and Bansal (2019); Muennighoff (2020); Nichol et al. (2021); Saharia et al. (2022); Weinbach et al. (2022). We have focused solely on natural language applications and leave extensive benchmarking of text embeddings as inputs for other modalities to future work.
Appendix C Examples
Tables 3-9 provide examples for each dataset for each task. For retrieval datasets, we refer to the BEIR paper Thakur et al. (2021).
Appendix D Correlations
Figure 6 provides correlation heatmaps for model performance and MTEB tasks.
Appendix E Models
Table 10 provides publicly available model checkpoints used for MTEB evaluation.
Appendix F Additional results
Tables 11 until the end provide results on individual datasets of MTEB. The results are additionally available in json format on the Hugging Face Hubhttps://huggingface.co/datasets/mteb/results and can be inspected on the leaderboardhttps://huggingface.co/spaces/mteb/leaderboard.