DeCLUTR: Deep Contrastive Learning for Unsupervised Textual Representations

John Giorgi, Osvald Nitski, Bo Wang, Gary Bader

Introduction

Due to the limited amount of labelled training data available for many natural language processing (NLP) tasks, transfer learning has become ubiquitous (Ruder et al., 2019). For some time, transfer learning in NLP was limited to pretrained word embeddings (Mikolov et al., 2013; Pennington et al., 2014). Recent work has demonstrated strong transfer task performance using pretrained sentence embeddings. These fixed-length vectors, often referred to as \sayuniversal sentence embeddings, are typically learned on large corpora and then transferred to various downstream tasks, such as clustering (e.g. topic modelling) and retrieval (e.g. semantic search). Indeed, sentence embeddings have become an area of focus, and many supervised (Conneau et al., 2017), semi-supervised (Subramanian et al., 2018; Phang et al., 2018; Cer et al., 2018; Reimers and Gurevych, 2019) and unsupervised (Le and Mikolov, 2014; Jernite et al., 2017; Kiros et al., 2015; Hill et al., 2016; Logeswaran and Lee, 2018) approaches have been proposed. However, the highest performing solutions require labelled data, limiting their usefulness to languages and domains where labelled data is abundant. Therefore, closing the performance gap between unsupervised and supervised universal sentence embedding methods is an important goal.

Pretraining transformer-based language models has become the primary method for learning textual representations from unlabelled corpora (Radford et al., 2018; Devlin et al., 2019; Dai et al., 2019; Yang et al., 2019; Liu et al., 2019; Clark et al., 2020). This success has primarily been driven by masked language modelling (MLM). This self-supervised token-level objective requires the model to predict the identity of some randomly masked tokens from the input sequence. In addition to MLM, some of these models have mechanisms for learning sentence-level embeddings via self-supervision. In BERT (Devlin et al., 2019), a special classification token is prepended to every input sequence, and its representation is used in a binary classification task to predict whether one textual segment follows another in the training corpus, denoted Next Sentence Prediction (NSP). However, recent work has called into question the effectiveness of NSP (Conneau and Lample, 2019; You et al., 1904; Joshi et al., 2020). In RoBERTa (Liu et al., 2019), the authors demonstrated that removing NSP during pretraining leads to unchanged or even slightly improved performance on downstream sentence-level tasks (including semantic text similarity and natural language inference). In ALBERT (Lan et al., 2020), the authors hypothesize that NSP conflates topic prediction and coherence prediction, and instead propose a Sentence-Order Prediction objective (SOP), suggesting that it better models inter-sentence coherence. In preliminary evaluations, we found that neither objective produces good universal sentence embeddings (see Appendix A). Thus, we propose a simple but effective self-supervised, sentence-level objective inspired by recent advances in metric learning.

Metric learning is a type of representation learning that aims to learn an embedding space where the vector representations of similar data are mapped close together, and vice versa (Lowe, 1995; Mika et al., 1999; Xing et al., 2002). In computer vision (CV), deep metric learning (DML) has been widely used for learning visual representations (Wohlhart and Lepetit, 2015; Wen et al., 2016; Zhang and Saligrama, 2016; Bucher et al., 2016; Leal-Taixé et al., 2016; Tao et al., 2016; Yuan et al., 2020; He et al., 2018; Grabner et al., 2018; Yelamarthi et al., 2018; Yu et al., 2018). Generally speaking, DML is approached as follows: a \saypretext task (often self-supervised, e.g. colourization or inpainting) is carefully designed and used to train deep neural networks to generate useful feature representations. Here, \sayuseful means a representation that is easily adaptable to other downstream tasks, unknown at training time. Downstream tasks (e.g. object recognition) are then used to evaluate the quality of the learned features (independent of the model that produced them), often by training a linear classifier on the task using these features as input. The most successful approach to date has been to design a pretext task for learning with a pair-based contrastive loss function. For a given anchor data point, contrastive losses attempt to make the distance between the anchor and some positive data points (those that are similar) smaller than the distance between the anchor and some negative data points (those that are dissimilar) (Hadsell et al., 2006). The highest-performing methods generate anchor-positive pairs by randomly augmenting the same image (e.g. using crops, flips and colour distortions); anchor-negative pairs are randomly chosen, augmented views of different images (Bachman et al., 2019; Tian et al., 2020; He et al., 2020; Chen et al., 2020). In fact, Kong et al., 2020 demonstrate that the MLM and NSP objectives are also instances of contrastive learning.

Inspired by this approach, we propose a self-supervised, contrastive objective that can be used to pretrain a sentence encoder. Our objective learns universal sentence embeddings by training an encoder to minimize the distance between the embeddings of textual segments randomly sampled from nearby in the same document. We demonstrate our objective’s effectiveness by using it to extend the pretraining of a transformer-based language model and obtain state-of-the-art results on SentEval (Conneau and Kiela, 2018) – a benchmark of 28 tasks designed to evaluate universal sentence embeddings. Our primary contributions are:

We propose a self-supervised sentence-level objective that can be used alongside MLM to pretrain transformer-based language models, inducing generalized embeddings for sentence- and paragraph-length text without any labelled data (subsection 5.1).

We perform extensive ablations to determine which factors are important for learning high-quality embeddings (subsection 5.2).

We demonstrate that the quality of the learned embeddings scale with model and data size. Therefore, performance can likely be improved simply by collecting more unlabelled text or using a larger encoder (subsection 5.3).

We open-source our solution and provide detailed instructions for training it on new data or embedding unseen text.https://github.com/JohnGiorgi/DeCLUTR

Related Work

Previous works on universal sentence embeddings can be broadly grouped by whether or not they use labelled data in their pretraining step(s), which we refer to simply as supervised or semi-supervised and unsupervised, respectively.

The highest performing universal sentence encoders are pretrained on the human-labelled natural language inference (NLI) datasets Stanford NLI (SNLI) (Bowman et al., 2015) and MultiNLI (Williams et al., 2018). NLI is the task of classifying a pair of sentences (denoted the \sayhypothesis and the \saypremise) into one of three relationships: entailment, contradiction or neutral. The effectiveness of NLI for training universal sentence encoders was demonstrated by the supervised method InferSent (Conneau et al., 2017). Universal Sentence Encoder (USE) (Cer et al., 2018) is semi-supervised, augmenting an unsupervised, Skip-Thoughts-like task (Kiros et al. 2015, see section 2) with supervised training on the SNLI corpus. The recently published Sentence Transformers (Reimers and Gurevych, 2019) method fine-tunes pretrained, transformer-based language models like BERT (Devlin et al., 2019) using labelled NLI datasets.

Skip-Thoughts (Kiros et al., 2015) and FastSent (Hill et al., 2016) are popular unsupervised techniques that learn sentence embeddings by using an encoding of a sentence to predict words in neighbouring sentences. However, in addition to being computationally expensive, this generative objective forces the model to reconstruct the surface form of a sentence, which may capture information irrelevant to the meaning of a sentence. QuickThoughts (Logeswaran and Lee, 2018) addresses these shortcomings with a simple discriminative objective; given a sentence and its context (adjacent sentences), it learns sentence representations by training a classifier to distinguish context sentences from non-context sentences. The unifying theme of unsupervised approaches is that they exploit the \saydistributional hypothesis, namely that the meaning of a word (and by extension, a sentence) is characterized by the word context in which it appears.

Our overall approach is most similar to Sentence Transformers – we extend the pretraining of a transformer-based language model to produce useful sentence embeddings – but our proposed objective is self-supervised. Removing the dependence on labelled data allows us to exploit the vast amount of unlabelled text on the web without being restricted to languages or domains where labelled data is plentiful (e.g. English Wikipedia). Our objective most closely resembles QuickThoughts; some distinctions include: we relax our sampling to textual segments of up to paragraph length (rather than natural sentences), we sample one or more positive segments per anchor (rather than strictly one), and we allow these segments to be adjacent, overlapping or subsuming (rather than strictly adjacent; see Figure 1, B).

Model

Our method learns textual representations via a contrastive loss by maximizing agreement between textual segments (referred to as \sayspans in the rest of the paper) sampled from nearby in the same document. Illustrated in Figure 1, this approach comprises the following components:

A data loading step randomly samples paired anchor-positive spans from each document in a minibatch of size NN. Let AA be the number of anchor spans sampled per document, PP be the number of positive spans sampled per anchor and i∈{1…AN}i\in\{1\dots AN\} be the index of an arbitrary anchor span. We denote an anchor span and its corresponding p∈{1…P}p\in\{1\dots P\} positive spans as si\bm{s}_{i} and si+pAN\bm{s}_{i+pAN} respectively. This procedure is designed to maximize the chance of sampling semantically similar anchor-positive pairs (see subsection 3.2).

An encoder f(⋅)f(\cdot) maps each token in the input spans to an embedding. Although our method places no constraints on the choice of encoder, we chose f(⋅)f(\cdot) to be a transformer-based language model, as this represents the state-of-the-art for text encoders (see subsection 3.3).

A pooler g(⋅)g(\cdot) maps the encoded spans f(si)f(\bm{s}_{i}) and f(si+pAN)f(\bm{s}_{i+pAN}) to fixed-length embeddings ei=g(f(si))\bm{e}_{i}=g(f(\bm{s}_{i})) and its corresponding mean positive embedding

Similar to Reimers and Gurevych 2019, we found that choosing g(⋅)g(\cdot) to be the mean of the token-level embeddings (referred to as \saymean pooling in the rest of the paper) performs well (see Appendix, Table 4). We pair each anchor embedding with the mean of multiple positive embeddings. This strategy was proposed by Saunshi et al. 2019, who demonstrated theoretical and empirical improvements compared to using a single positive example for each anchor.

A contrastive loss function defined for a contrastive prediction task. Given a set of embedded spans {ek}\{\bm{e}_{k}\} including a positive pair of examples ei\bm{e}_{i} and ei+AN\bm{e}_{i+AN}, the contrastive prediction task aims to identify ei+AN\bm{e}_{i+AN} in {ek}k≠i\{\bm{e}_{k}\}_{k\not=i} for a given ei\bm{e}_{i}

During training, we randomly sample minibatches of NN documents from the train set and define the contrastive prediction task on anchor-positive pairs ei,ei+AN\bm{e}_{i},\bm{e}_{i+AN} derived from the NN documents, resulting in 2AN2AN data points. As proposed in (Sohn, 2016), we treat the other 2(AN−1)2(AN-1) instances within a minibatch as negative examples. The cost function takes the following form

This is the InfoNCE loss used in previous works (Sohn, 2016; Wu et al., 2018; Oord et al., 2018) and denoted normalized temperature-scale cross-entropy loss or \sayNT-Xent in (Chen et al., 2020). To embed text with a trained model, we simply pass batches of tokenized text through the model, without sampling spans. Therefore, the computational cost of our method at test time is the cost of the encoder, f(⋅)f(\cdot), plus the cost of the pooler, g(⋅)g(\cdot), which is negligible when using mean pooling.

2 Span sampling

We then sample p∈{1…P}p\in\{1\dots P\} corresponding positive spans si+pAN\bm{s}_{i+pAN} independently following a similar procedure

We found that sampling longer lengths for the anchor span than the positive spans improves performance in downstream tasks (we did not find performance to be sensitive to the specific choice of α\alpha and β\beta). The rationale for this is twofold. First, it enables the model to learn global-to-local view prediction as in (Hjelm et al., 2019; Bachman et al., 2019; Chen et al., 2020) (referred to as \saysubsumed view in Figure 1, B). Second, when P>1P>1, it encourages diversity among positives spans by lowering the amount of repeated text.

Sampling positives nearby to the anchor exploits the distributional hypothesis and increases the chances of sampling valid (i.e. semantically similar) anchor-positive pairs.

By sampling multiple anchors per document, each anchor-positive pair is contrasted against both easy negatives (anchors and positives sampled from other documents in a minibatch) and hard negatives (anchors and positives sampled from the same document).

In conclusion, the sampling procedure produces three types of positives: positives that partially overlap with the anchor, positives adjacent to the anchor, and positives subsumed by the anchor (Figure 1, B) and two types of negatives: easy negatives sampled from a different document than the anchor, and hard negatives sampled from the same document as the anchor. Thus, our stochastically generated training set and contrastive loss implicitly define a family of predictive tasks which can be used to train a model, independent of any specific encoder architecture.

3 Continued MLM pretraining

We use our objective to extend the pretraining of a transformer-based language model (Vaswani et al., 2017), as this represents the state-of-the-art encoder in NLP. We implement the MLM objective as described in (Devlin et al., 2019) on each anchor span in a minibatch and sum the losses from the MLM and contrastive objectives before backpropagating

This is similar to existing pretraining strategies, where an MLM loss is paired with a sentence-level loss such as NSP (Devlin et al., 2019) or SOP (Lan et al., 2020). To make the computational requirements feasible, we do not train from scratch, but rather we continue training a model that has been pretrained with the MLM objective. Specifically, we use both RoBERTa-base (Liu et al., 2019) and DistilRoBERTa (Sanh et al., 2019) (a distilled version of RoBERTa-base) in our experiments. In the rest of the paper, we refer to our method as DeCLUTR-small (when extending DistilRoBERTa pretraining) and DeCLUTR-base (when extending RoBERTa-base pretraining).

Experimental setup

We collected all documents with a minimum token length of 2048 from OpenWebText (Gokaslan and Cohen, 2019) an open-access subset of the WebText corpus (Radford et al., 2019), yielding 497,868 documents in total. For reference, Google’s USE was trained on 570,000 human-labelled sentence pairs from the SNLI dataset (among other unlabelled datasets). InferSent and Sentence Transformer models were trained on both SNLI and MultiNLI, a total of 1 million human-labelled sentence pairs.

We implemented our model in PyTorch (Paszke et al., 2017) using AllenNLP (Gardner et al., 2018). We used the NT-Xent loss function implemented by the PyTorch Metric Learning library (Musgrave et al., 2019) and the pretrained transformer architecture and weights from the Transformers library (Wolf et al., 2020). All models were trained on up to four NVIDIA Tesla V100 16 or 32GB GPUs.

Unless specified otherwise, we train for one to three epochs over the 497,868 documents with a minibatch size of 16 and a temperature τ=5×10−2\tau=5\times 10^{-2} using the AdamW optimizer (Loshchilov and Hutter, 2019) with a learning rate (LR) of 5×10−55\times 10^{-5} and a weight decay of 0.10.1. For every document in a minibatch, we sample two anchor spans (A=2A=2) and two positive spans per anchor (P=2P=2). We use the Slanted Triangular LR scheduler (Howard and Ruder, 2018) with a number of train steps equal to training instances and a cut fraction of 0.10.1. The remaining hyperparameters of the underlying pretrained transformer (i.e. DistilRoBERTa or RoBERTa-base) are left at their defaults. All gradients are scaled to a vector norm of 1.01.0 before backpropagating. Hyperparameters were tuned on the SentEval validation sets.

2 Evaluation

We evaluate all methods on the SentEval benchmark, a widely-used toolkit for evaluating general-purpose, fixed-length sentence representations. SentEval is divided into 18 downstream tasks – representative NLP tasks such as sentiment analysis, natural language inference, paraphrase detection and image-caption retrieval – and ten probing tasks, which are designed to evaluate what linguistic properties are encoded in a sentence representation. We report scores obtained by our model and the relevant baselines on the downstream and probing tasks using the SentEval toolkithttps://github.com/facebookresearch/SentEval with default parameters (see Appendix C for details). Note that all the supervised approaches we compare to are trained on the SNLI corpus, which is included as a downstream task in SentEval. To avoid train-test contamination, we compute average downstream scores without considering SNLI when comparing to these approaches in Table 2.

We compare to the highest performing, most popular sentence embedding methods: InferSent, Google’s USE and Sentence Transformers. For InferSent, we compare to the latest model.https://dl.fbaipublicfiles.com/infersent/infersent2.pkl We use the latest \saylarge USE modelhttps://tfhub.dev/google/universal-sentence-encoder-large/5, as it is most similar in terms of architecture and number of parameters to DeCLUTR-base. For Sentence Transformers, we compare to \sayroberta-base-nli-mean-tokenshttps://www.sbert.net/docs/pretrained_models.html, which, like DeCLUTR-base, uses the RoBERTa-base architecture and pretrained weights. The only difference is each method’s extended pretraining strategy. We include the performance of averaged GloVehttp://nlp.stanford.edu/data/glove.840B.300d.zip and fastTexthttps://dl.fbaipublicfiles.com/fasttext/vectors-english/crawl-300d-2M.vec.zip word vectors as weak baselines. Trainable model parameter counts and sentence embedding dimensions are listed in Table 1. Despite our best efforts, we could not evaluate the pretrained QuickThought models against the full SentEval benchmark. We cite the scores from the paper directly. Finally, we evaluate the pretrained transformer model’s performance before it is subjected to training with our contrastive objective, denoted \sayTransformer-*. We use mean pooling on the pretrained transformers token-level output to produce sentence embeddings – the same pooling strategy used in our method.

Results

In subsection 5.1, we compare the performance of our model against the relevant baselines. In the remaining sections, we explore which components contribute to the quality of the learned embeddings.

Compared to the underlying pretrained models DistilRoBERTa and RoBERTa-base, DeCLUTR-small and DeCLUTR-base obtain large boosts in average downstream performance, +4% and +6% respectively (Table 2). DeCLUTR-base leads to improved or equivalent performance for every downstream task but one (SST5) and DeCLUTR-small for all but three (SST2, SST5 and TREC). Compared to existing methods, DeCLUTR-base matches or even outperforms average performance without using any hand-labelled training data. Surprisingly, we also find that DeCLUTR-small outperforms Sentence Transformers while using ∼\sim34% less trainable parameters.

With the exception of InferSent, existing methods perform poorly on the probing tasks of SentEval (Table 3). Sentence Transformers, which begins with a pretrained transformer model and fine-tunes it on NLI datasets, scores approximately 10% lower on the probing tasks than the model it fine-tunes. In contrast, both DeCLUTR-small and DeCLUTR-base perform comparably to the underlying pretrained model in terms of average performance. We note that the purpose of the probing tasks is not the development of ad-hoc models that attain top performance on them (Conneau et al., 2018). However, it is still interesting to note that high downstream task performance can be obtained without sacrificing probing task performance. Furthermore, these results suggest that fine-tuning transformer-based language models on NLI datasets may discard some of the linguistic information captured by the pretrained model’s weights. We suspect that the inclusion of MLM in our training objective is responsible for DeCLUTR’s relatively high performance on the probing tasks.

The downstream evaluation of SentEval includes supervised and unsupervised tasks. In the unsupervised tasks, the embeddings of the method to evaluate are used as-is without any further training (see Appendix C for details). Interestingly, we find that USE performs particularly well across the unsupervised evaluations in SentEval (tasks marked with a * in Table 2). Given the similarity of the USE architecture to Sentence Transformers and DeCLUTR and the similarity of its supervised NLI training objective to InferSent and Sentence Transformers, we suspect the most likely cause is one or more of its additional training objectives. These include a conversational response prediction task (Henderson et al., 2017) and a Skip-Thoughts (Kiros et al., 2015) like task.

2 Ablation of the sampling procedure

We ablate several components of the sampling procedure, including the number of anchors sampled per document AA, the number of positives sampled per anchor PP, and the sampling strategy for those positives (Figure 2). We note that when A=2A=2, the model is trained on twice the number of spans and twice the effective batch size (2AN2AN, where NN is the number of documents in a minibatch) as compared to when A=1A=1. To control for this, all experiments where A=1A=1 are trained for two epochs (twice the number of epochs as when A=2A=2) and for two times the minibatch size (2N2N). Thus, both sets of experiments are trained on the same number of spans and the same effective batch size (4N4N), and the only difference is the number of anchors sampled per document (AA).

We find that sampling multiple anchors per document has a large positive impact on the quality of learned embeddings. We hypothesize this is because the difficulty of the contrastive objective increases when A>1A>1. Recall that a minibatch is composed of random documents, and each anchor-positive pair sampled from a document is contrasted against all other anchor-positive pairs in the minibatch. When A>1A>1, anchor-positive pairs will be contrasted against other anchors and positives from the same document, increasing the difficulty of the contrastive objective, thus leading to better representations. We also find that a positive sampling strategy that allows positives to be adjacent to and subsumed by the anchor outperforms a strategy that only allows adjacent or subsuming views, suggesting that the information captured by these views is complementary. Finally, we note that sampling multiple positives per anchor (P>1P>1) has minimal impact on performance. This is in contrast to (Saunshi et al., 2019), who found both theoretical and empirical improvements when multiple positives are averaged and paired with a given anchor.

3 Training objective, train set size and model capacity

To determine the importance of the training objectives, train set size, and model capacity, we trained two sizes of the model with 10% to 100% (1 full epoch) of the train set (Figure 3). Pretraining the model with both the MLM and contrastive objectives improves performance over training with either objective alone. Including MLM alongside the contrastive objective leads to monotonic improvement as the train set size is increased. We hypothesize that including the MLM loss acts as a form of regularization, preventing the weights of the pretrained model (which itself was trained with an MLM loss) from diverging too dramatically, a phenomenon known as \saycatastrophic forgetting (McCloskey and Cohen, 1989; Ratcliff, 1990). These results suggest that the quality of embeddings learned by our approach scale in terms of model capacity and train set size; because the training method is completely self-supervised, scaling the train set would simply involve collecting more unlabelled text.

Discussion and conclusion

In this paper, we proposed a self-supervised objective for learning universal sentence embeddings. Our objective does not require labelled training data and is applicable to any text encoder. We demonstrated the effectiveness of our objective by evaluating the learned embeddings on the SentEval benchmark, which contains a total of 28 tasks designed to evaluate the transferability and linguistic properties of sentence representations. When used to extend the pretraining of a transformer-based language model, our self-supervised objective closes the performance gap with existing methods that require human-labelled training data. Our experiments suggest that the learned embeddings’ quality can be further improved by increasing the model and train set size. Together, these results demonstrate the effectiveness and feasibility of replacing hand-labelled data with carefully designed self-supervised objectives for learning universal sentence embeddings. We release our model and code publicly in the hopes that it will be extended to new domains and non-English languages.

Acknowledgments

This research was enabled in part by support provided by Compute Ontario (https://computeontario.ca/), Compute Canada (www.computecanada.ca) and the CIFAR AI Chairs Program and partially funded by the US National Institutes of Health (NIH) [U41 HG006623, U41 HG003751).

References

Appendix A Pretrained transformers make poor universal sentence encoders

Certain pretrained transformers, such as BERT and ALBERT, have mechanisms for learning sequence-level embeddings via self-supervision. These models prepend every input sequence with a special classification token (e.g. \say[CLS]), and its representation is learned using a simple classification task, such as Next Sentence Prediction (NSP) or Sentence-Order Prediction (SOP) (see Devlin et al. 2019 and Lan et al. 2020 respectively for details on these tasks). However, during preliminary experiments, we noticed that these models are not good universal sentence encoders, as measured by their performance on the SentEval benchmark (Conneau and Kiela, 2018). As a simple experiment, we evaluated three pretrained transformer models on SentEval: one trained with the NSP loss (BERT), one trained with the SOP loss (ALBERT) and one trained with neither, RoBERTa (Liu et al., 2019). We did not find that the CLS embeddings produced by models trained against the NSP or SOP losses to outperform that of a model trained without either loss and sometimes failed to outperform a bag-of-words (BoW) baseline (Table 4). Furthermore, we find that pooling token embeddings via averaging (referred to as \saymean pooling in our paper) outperforms pooling via the CLS classification token. Our results are corroborated by Liu et al. 2019, who find that removing NSP loss leads to the same or better results on downstream tasks and Reimers and Gurevych 2019, who find that directly using the output of BERT as sentence embeddings leads to poor performances on the semantic similarity tasks of SentEval.

Appendix B Examples of sampled spans

In Table 5, we present examples of anchor-positive and anchor-negative pairs generated by our sampling procedure. We show one example for each possible view of a sampled positive, e.g. positives adjacent to, overlapping with, or subsumed by the anchor. For each anchor-positive pair, we show examples of both a hard negative (derived from the same document) and an easy negative (derived from another document). Recall that a minibatch is composed of random documents, and each anchor-positive pair sampled from a document is contrasted against all other anchor-positive pairs in the minibatch. Thus, hard negatives, as we have described them here, are generated only when sampling multiple anchors per document (A>1A>1).

Appendix C SentEval evaluation details

SentEval is a benchmark for evaluating the quality of fixed-length sentence embeddings. It is divided into 18 downstream tasks, and 10 probing tasks. Sentence embedding methods are evaluated on these tasks via a simple interfacehttps://github.com/facebookresearch/SentEval, which standardizes training, evaluation and hyperparameters. For most tasks, the method to evaluate is used to produce fix-length sentence embeddings, and a simple logistic regression (LR) or multi-layer perception (MLP) model is trained on the task using these embeddings as input. For other tasks (namely several semantic text similarity tasks), the embeddings are used as-is without any further training. Note that this setup is different from evaluations on the popular GLUE benchmark (Wang et al., 2019), which typically use the task data to fine-tune the parameters of the sentence embedding model.

In subsection C.1, we present the individual tasks of the SentEval benchmark. In subsection C.2, we explain our method for computing the average downstream and average probing scores presented in our paper.

The downstream tasks of SentEval are representative NLP tasks used to evaluate the transferability of fixed-length sentence embeddings. We give a brief overview of the broad categories that divide the tasks below (see Conneau and Kiela 2018 for more details):

Binary and multi-class classification: These tasks cover various types of sentence classification, including sentiment analysis (MR Pang and Lee 2005, SST2 and SST5 Socher et al. 2013), question-type (TREC) (Voorhees and Tice, 2000), product reviews (CR) (Hu and Liu, 2004), subjectivity/objectivity (SUBJ) (Pang and Lee, 2004) and opinion polarity (MPQA) (Wiebe et al., 2005).

Entailment and semantic relatedness: These tasks cover multiple entailment datasets (also known as natural language inference or NLI), including SICK-E (Marelli et al., 2014) and the Stanford NLI dataset (SNLI) (Bowman et al., 2015) as well as multiple semantic relatedness datasets including SICK-R and STS-B (Cer et al., 2017).

Semantic textual similarity These tasks (STS12 Agirre et al. 2012, STS13 Agirre et al. 2013, STS14 Agirre et al. 2014, STS15 Agirre et al. 2015 and STS16 Agirre et al. 2016) are similar to the semantic relatedness tasks, except the embeddings produced by the encoder are used as-is in a cosine similarity to determine the semantic similarity of two sentences. No additional model is trained on top of the encoder’s output.

Paraphrase detection Evaluated on the Microsoft Research Paraphrase Corpus (MRPC) (Dolan et al., 2004), this binary classification task is comprised of human-labelled sentence pairs, annotated according to whether they capture a paraphrase/semantic equivalence relationship.

Caption-Image retrieval This task is comprised of two sub-tasks: ranking a large collection of images by their relevance for some given query text (Image Retrieval) and ranking captions by their relevance for some given query image (Caption Retrieval). Both tasks are evaluated on data from the COCO dataset (Lin et al., 2014). Each image is represented by a pretrained, 2048-dimensional embedding produced by a ResNet-101 (He et al., 2016).

The probing tasks are designed to evaluate what linguistic properties are encoded in a sentence representation. All tasks are binary or multi-class classification. We give a brief overview of each task below (see Conneau et al. 2018 for more details):

Sentence length (SentLen): A multi-class classification task where a model is trained to predict the length of a given input sentence, which is binned into six possible length ranges.

Word content (WC): A multi-class classification task where, given 1000 words as targets, the goal is to predict which of the target words appears in a given input sentence. Each sentence contains a single target word, and the word occurs exactly once in the sentence.

Tree depth (TreeDepth): A multi-class classification task where the goal is to predict the maximum depth (with values ranging from 5 to 12) of a given input sentence’s syntactic tree.

Bigram Shift (BShift): A multi-class classification task where the goal is to predict whether two consecutive tokens within a given sentence have been inverted.

Top Constituents (TopConst): A multi-class classification task where the goal is to predict the top constituents (from a choice of 19) immediately below the sentence (S) node of the sentence’s syntactic tree.

Tense: A binary classification task where the goal is to predict the tense (past or present) of the main verb in a sentence.

Subject number (SubjNum): A binary classification task where the goal is to predict the number (singular or plural) of the subject of the main clause.

Object number (ObjNum): A binary classification task, analogous to SubjNum, where the goal is to predict the number (singular or plural) of the direct object of the main clause.

Semantic odd man out (SOMO): A binary classification task where the goal is to predict whether a sentence has had a single randomly picked noun or verb replaced with another word with the same part-of-speech.

Coordinate inversion (CoordInv): A binary classification task where the goal is to predict whether the order of two coordinate clauses in a sentence has been inverted.

C.2 Computing an average score

In our paper, we present averaged downstream and probing scores. Computing averaged probing scores was straightforward; each of the ten probing tasks reports a simple accuracy, which we averaged. To compute an averaged downstream score, we do the following:

If a task reports Spearman correlation (i.e. SICK-R, STS-B), we use this score when computing the average downstream task score. If the task reports a mean Spearman correlation for multiple subtasks (i.e. STS12, STS13, STS14, STS15, STS16), we use this score.

If a task reports both an accuracy and an F1-score (i.e. MRPC), we use the average of these two scores.

For the Caption-Image Retrieval task, we report the average of the Recall@K, where K∈{1,5,10}K\in\{1,5,10\} for the Image and Caption retrieval tasks (a total of six scores). This is the default behaviour of SentEval.