Hard Negative Mixing for Contrastive Learning

Yannis Kalantidis, Mert Bulent Sariyildiz, Noe Pion, Philippe Weinzaepfel, Diane Larlus

Introduction

Contrastive learning was recently shown to be a highly effective way of learning visual representations in a self-supervised manner [chen2020simple, he2019momentum]. Pushing the embeddings of two transformed versions of the same image (forming the positive pair) close to each other and further apart from the embedding of any other image (negatives) using a contrastive loss, leads to powerful and transferable representations. A number of recent studies [chen2020improved, gontijo2020affinity, tian2020makes] show that carefully handcrafting the set of data augmentations applied to images is instrumental in learning such representations. We suspect that the right set of transformations provides more diverse, i.e. more challenging, copies of the same image to the model and makes the self-supervised (proxy) task harder. At the same time, data mixing techniques operating at either the pixel [verma2019interpolation, yun2019cutmix, zhang2017mixup] or the feature level [verma2019manifold] help models learn more robust features that improve both supervised and semi-supervised learning on subsequent (target) tasks.

In most recent contrastive self-supervised learning approaches, the negative samples come from either the current batch or a memory bank. Because the number of negatives directly affects the contrastive loss, current top contrastive approaches either substantially increase the batch size [chen2020simple], or keep large memory banks. Approaches like [misra2019self, wu2018unsupervised] use memories that contain the whole training set, while the recent Momentum Contrast (or MoCo) approach of he2019momentum keeps a queue with features of the last few batches as memory. The MoCo approach with the modifications presented in [chen2020improved] (named MoCo-v2) currently holds the state-of-the-art performance on a number of target tasks used to evaluate the quality of visual representations learned in an unsupervised way. It is however shown [chen2020simple, he2019momentum] that increasing the memory/batch size leads to diminishing returns in terms of performance: more negative samples does not necessarily mean hard negative samples.

In this paper, we argue that an important aspect of contrastive learning, i.e. the effect of hard negatives, has so far been neglected in the context of self-supervised representation learning. We delve deeper into learning with a momentum encoder [he2019momentum] and show evidence that harder negatives are required to facilitate better and faster learning. Based on these observations, and motivated by the success of data mixing approaches, we propose hard negative mixing, i.e. feature-level mixing for hard negative samples, that can be computed on-the-fly with a minimal computational overhead. We refer to the proposed approach as MoCHi, that stands for "(M)ixing (o)f (C)ontrastive (H)ard negat(i)ves".

A toy example of the proposed hard negative mixing strategy is presented in Figure 1; it shows a t-SNE [maaten2008visualizing] plot after running MoCHi on 3232-dimensional random embeddings on the unit hypersphere. We see that for each positive query (red square), the memory (gray marks) contains many easy negatives and few hard ones, i.e. many of the negatives are too far to contribute to the contrastive loss. We propose to mix only the hardest negatives (based on their similarity to the query) and synthesize new, hopefully also hard but more diverse, negative points (blue triangles).

a) We delve deeper into a top-performing contrastive self-supervised learning method [he2019momentum] and observe the need for harder negatives; b) We propose hard negative mixing, i.e. to synthesize hard negatives directly in the embedding space, on-the-fly, and adapted to each positive query. We propose to both mix pairs of the hardest existing negatives, as well as mixing the hardest negatives with the query itself; c) We exhaustively ablate our approach and show that employing hard negative mixing improves both the generalization of the visual representations learned (measured via their transfer learning performance), as well as the utilization of the embedding space, for a wide range of hyperparameters; d) We report competitive results for linear classification, object detection and instance segmentation, and further show that our gains over a state-of-the-art method are higher when pre-training for fewer epochs, i.e. MoCHi learns transferable representations faster.

Related work

Most early self-supervised learning methods are based on devising proxy classification tasks that try to predict the properties of a transformation (e.g. rotations, orderings, relative positions or channels) applied on a single image [doersch2015unsupervised, dosovitskiy2014discriminative, gidaris2018unsupervised, kolesnikov2019revisiting, noroozi2016unsupervised]. Instance discrimination [wu2018unsupervised] and CPC [oord2018representation] were among the first papers to use contrastive losses for self-supervised learning. The last few months have witnessed a surge of successful approaches that also use contrastive learning losses. These include MoCo [chen2020improved, he2019momentum], SimCLR [chen2020simple, chen2020big], PIRL [misra2019self], CMC [tian2019contrastive] or SvAV [caron2020unsupervised]. In parallel, methods like [asano2020self, caron2018deep, caron2019unsupervised, caron2020unsupervised, zhuang2019local, li2020prototypical] build on the idea that clusters should be formed in the feature spaces, and use clustering losses together with contrastive learning or transformation prediction tasks.

Most of the top-performing contrastive methods leverage data augmentations [chen2020simple, chen2020improved, grill2020bootstrap, he2019momentum, misra2019self, tian2019contrastive]. As revealed by recent studies [asano2019critical, gontijo2020affinity, tian2020makes, wang2020understanding], heavy data augmentations applied to the same image are crucial in learning useful representations, as they modulate the hardness of the self-supervised task via the positive pair. Our proposed hard negative mixing technique, on the other hand, is changing the hardness of the proxy task from the side of the negatives.

A few recent works discuss issues around the selection of negatives in contrastive self-supervised learning [cao2020parametric, chuang2020debiased, iscen2018mining, wu2020mutual, xie2020delving, ho2020contrastive]. Iscen et al. [iscen2018mining] mine hard negatives from a large set by focusing on the features that are neighbors with respect to the Euclidean distance, but not when using a manifold distance defined over the nearest neighbor graph. Interested in approximating the underlying “true” distribution of negative examples, Chuang et al. [chuang2020debiased] present a debiased version of the contrastive loss, in an effort to mediate the effect of false negatives. Wu et al. [wu2020mutual] present a variational extension to the InfoNCE objective that is further coupled with modified strategies for negative sampling, e.g. restricting negative sampling to a region around the query. In concurrent works, Cao et al. [cao2020parametric] propose a weight update correction for negative samples to decrease GPU memory consumption caused by weight decay regularization, while in [ho2020contrastive] the authors propose a new algorithm that generates more challenging positive and hard negative pairs, on-the-fly, by leveraging adversarial examples.

Mixup [zhang2017mixup] and its numerous variants [shen2020rethinking, verma2019manifold, walawalkar2020attentive, yun2019cutmix] have been shown to be highly effective data augmentation strategies when paired with a cross-entropy loss for supervised and semi-supervised learning. Manifold mixup [verma2019manifold] is a feature-space regularizer that encourages networks to be less confident for interpolations of hidden states. The benefits of interpolating have only recently been explored for losses other than cross-entropy [ko2020embedding, shen2020rethinking, zhou2020C2L]. In [shen2020rethinking], the authors propose using mixup in the image/pixel space for self-supervised learning; in contrast, we create query-specific synthetic points on-the-fly in the embedding space. This makes our method way more computationally efficient and able to show improved results at a smaller number of epochs. The Embedding Expansion [ko2020embedding] work explores interpolating between embeddings for supervised metric learning on fine-grained recognition tasks. The authors use uniform interpolation between two positive and negative points, create a set of synthetic points and then select the hardest pair as negative. In contrast, the proposed MoCHi has no need for class annotations, performs no selection for negatives and only samples a single random interpolation between multiple pairs. What is more, in this paper we go beyond mixing negatives and propose mixing the positive with negative features, to get even harder negatives, and achieve improved performance. Our work is also related to metric learning works that employ generators [duan2018deep, zheng2019hardness]. Apart from not requiring labels, our method exploits the memory component and has no extra parameters or loss terms that need to be optimized.

Understanding hard negatives in unsupervised contrastive learning

The log-likelihood function of Eq (1) is defined over the probability distribution created by applying a softmax function for each input/query q\mathbf{q}. Let pzip_{z_{i}} be the matching probability for the query and feature zi∈Z=Q∪{k}\mathbf{z}_{i}\in Z=Q\cup\{\mathbf{k}\}, then the gradient of the loss with respect to the query q\mathbf{q} is given by:

and pk,pnp_{k},p_{n} are the matching probability of the key and negative feature, i.e. for zi=k\mathbf{z}_{i}=\mathbf{k} and for zi=n\mathbf{z}_{i}=\mathbf{n}, respectively. We see that the contributions of the positive and negative logits to the loss are identical to the ones for a (K+1K+1)-way cross-entropy classification loss, where the logit for the key corresponds to the query’s latent class [arora2019theoretical] and all gradients are scaled by 1/τ1/\tau.

2 Hard negatives in contrastive learning

Hard negatives are critical for contrastive learning [arora2019theoretical, harwood2017smart, iscen2018mining, mishchuk2017working, simo2015discriminative, wu2017sampling, xuan2020hard]. Sampling negatives from the same batch leads to a need for larger batches [chen2020simple] while sampling negatives from a memory bank that contains every other image in the dataset requires the time consuming task of keeping a large memory up-to-date [misra2019self, wu2018unsupervised]. In the latter case, a trade-off exists between the “freshness” of the memory bank representations and the computational overhead for re-computing them as the encoder keeps changing. The Momentum Contrast (or MoCo) approach of he2019momentum offers a compromise between the two negative sampling extremes: it keeps a queue of the latest KK features from the last batches, encoded with a second key encoder that trails the (main/query) encoder with a much higher momentum. For MoCo, the key feature k\mathbf{k} and all features in QQ are encoded with the key encoder.

In MoCo [he2019momentum] (resp. SimCLR [chen2020simple]) the authors show that increasing the memory (resp. batch) size, is crucial to getting better and harder negatives. In Figure 2(a) we visualize how hard the negatives are during training for MoCo-v2, by plotting the highest 1024 matching probabilities pip_{i} for ImageNet-100 In this section we study contrastive learning for MoCo [he2019momentum] on ImageNet-100, a subset of ImageNet consisting of 100 classes introduced in [tian2019contrastive]. See Section 5 for details on the dataset and experimental protocol. and a queue of size K=16K=16k. We see that, although in the beginning of training (i.e. at epoch 0) the logits are relatively flat, as training progresses, fewer and fewer negatives offer significant contributions to the loss. This shows that most of the memory negatives are practically not helping a lot towards learning the proxy task.

On the difficulty of the proxy task.

For MoCo [he2019momentum], SimCLR [chen2020simple], InfoMin [tian2020makes], and other approaches that learn augmentation-invariant representations, we suspect the hardness of the proxy task to be directly correlated with the difficulty of the transformations set, i.e. hardness is modulated via the positive pair. We propose to experimentally verify this. In Figure 2(b), we plot the proxy task performance, i.e. the percentage of queries where the key is ranked over all negatives, across training for MoCo [he2019momentum] and MoCo-v2 [chen2020improved]. MoCo-v2 enjoys a high performance gain over MoCo by three main changes: the addition of a Multilayer Perceptron (MLP) head, cosine learning rate schedule, and more challenging data augmentation. As we further discuss in the Appendix, only the latter of these three changes makes the proxy task harder to solve. Despite the drop in proxy task performance, however, further performance gains are observed for linear classification. In Section 4 we discuss how MoCHi gets a similar effect by modulating the proxy task through mixing harder negatives.

3 A class oracle-based analysis

In this section, we analyze the negatives for contrastive learning using a class oracle, i.e. the ImageNet class label annotations. Let us define false negatives (FN) as all negative features in the memory QQ, that correspond to images of the same class as the query. Here we want to first quantify false negatives from contrastive learning and then explore how they affect linear classification performance. What is more, by using class annotations, we can train a contrastive self-supervised learning oracle, where we measure performance at the downstream task (linear classification) after disregarding FN from the negatives of each query during training. This has connections to the recent work of [khosla2020supervised], where a contrastive loss is used in a supervised way to form positive pairs from images sharing the same label. Unlike [khosla2020supervised], our oracle uses labels only for discarding negatives with the same label for each query, i.e. without any other change to the MoCo protocol.

In Figure 2(c), we quantify the percentage of false negatives for the oracle run and MoCo-v2, when looking at highest 1024 negative logits across training epochs. We see that a) in all cases, as representations get better, more and more FNs (same-class logits) are ranked among the top; b) by discarding them from the negatives queue, the class oracle version (purple line) is able to bring same-class embeddings closer. Performance results using the class oracle, as well as a supervised upper bound trained with cross-entropy are shown in the bottom section of Figure 1. We see that the MoCo-v2 oracle recovers part of the performance relative to the supervised case, i.e. 78.078.0 (MoCo-v2, 200 epochs) →81.8\rightarrow 81.8 (MoCo-v2 oracle, 200 epochs) →86.2\rightarrow 86.2 (supervised).

Feature space mixing of hard negatives

In this section we present an approach for synthesizing hard negatives, i.e. by mixing some of the hardest negative features of the contrastive loss or the hardest negatives with the query. We refer to the proposed hard negative mixing approach as MoCHi, and use the naming convention MoCHi (NN, ss, s′s^{\prime}), that indicates the three important hyperparameters of our approach, to be defined below.

2 Mixing for even harder negatives

3 Discussion and analysis of MoCHi

Recent approaches like [chen2020simple, chen2020improved] use a Multi-layer Perceptron (MLP) head instead of a linear layer for the embeddings that participate in the contrastive loss. This means that the embeddings whose dot products contribute to the loss, are not the ones used for target tasks–a lower-layer embedding is used instead. Unless otherwise stated, we follow [chen2020simple, chen2020improved] and use a 2-layer MLP head on top of the features we use for downstream tasks. We always mix hard negatives in the space of the loss.

Oracle insights for MoCHi.

From Figure 2(c) we see that the percentage of synthesized features obtained by mixing two false negatives (lines with square markers) increases over time, but remains very small, i.e. around only 1%. At the same time, we see that about 8% of the synthetic features are fractionally false negatives (lines with triangle markers), i.e. at least one of its two components is a false negative. For the oracle variants of MoCHi, we also do not allow false negatives to participate in synthesizing hard negatives. From Table 1 we see that not only the MoCHi oracle is able to get a higher upper bound (82.5 vs 81.8 for MoCo-v2), further closing the difference to the cross entropy upper bound, but we also show in the Appendix that, after longer training, the MoCHi oracle is able to recover most of the performance loss versus using cross-entropy, i.e. 79.079.0 (MoCHi, 200 epochs) →82.5\rightarrow 82.5 (MoCHi oracle, 200 epochs) →85.2\rightarrow 85.2 (MoCHi oracle, 800 epochs) →86.2\rightarrow 86.2 (supervised).

It is noteworthy that by the end of training, MoCHi exhibits slightly lower percentage of false negatives in the top logits compared to MoCo-v2 (rightmost values of Figure 2(c)). This is an interesting result: MoCHi adds synthetic negative points that are (at least partially) false negatives and is pushing embeddings of the same class apart, but at the same time it exhibits higher performance for linear classification on ImageNet-100. That is, it seems that although the absolute similarities of same-class features may decrease, the method results in a more linearly separable space. This inspired us to further look into how having synthetic hard negatives impacts the utilization of the embedding space.

Measuring the utilization of the embedding space.

Very recently, wang2020understanding presented two losses/metrics for assessing contrastive learning representations. The first measures the alignment of the representation on the hypersphere, i.e. the absolute distance between representations with the same label. The second measures the uniformity of their distribution on the hypersphere, through measuring the logarithm of the average pairwise Gaussian potential between all embeddings. In Figure LABEL:fig:uniformity, we plot these two values for a number of models, when using features from all images in the ImageNet-100 validation set. We see that MoCHi highly improves the uniformity of the representations compared to both MoCo-v2 and the supervised models. This further supports our hypothesis that MoCHi allows the proxy task to learn to better utilize the embedding space. In fact, we see that the supervised model leads to high alignment but very low uniformity, denoting features targeting the classification task. On the other hand, MoCo-v2 and MoCHi have much better spreading of the underlying embedding space, which we experimentally know leads to more generalizable representations, i.e. both MoCo-v2 and MoCHi outperform the supervised ImageNet-pretrained backbone for transfer learning (see Figure LABEL:fig:uniformity).

Experiments

We learn representations on two datasets, the common ImageNet-1K [russakovsky2015imagenet], and its smaller ImageNet-100 subset, also used in [shen2020rethinking, tian2019contrastive]. All runs of MoCHi are based on MoCo-v2. We developed our approach on top of the official public implementation of MoCo-v2https://github.com/facebookresearch/moco and reproduced it on our setup; other results are copied from the respective papers. We run all experiments on 4 GPU servers. For linear classification on ImageNet-100 (resp. ImageNet-1K), we follow the common protocol and report results on the validation set. We report performance after learning linear classifiers for 60 (resp. 100) epochs, with an initial learning rate of 10.0 (30.0), a batch size of 128 (resp. 512) and a step learning rate schedule that drops at epochs 30, 40 and 50 (resp. 60, 80). For training we use K=16K=16k (resp. K=65K=65k). For MoCHi, we also have a warm-up of 10 (resp. 15) epochs, i.e. for the first epochs we do not synthesize hard negatives. For ImageNet-1K, we report accuracy for a single-crop testing. For object detection on PASCAL VOC [everingham2010pascal] we follow [he2019momentum] and fine-tune a Faster R-CNN [ren2015faster], R50-C4 on trainval07+12 and test on test2007. We use the open-source detectron2https://github.com/facebookresearch/detectron2 code and report the common AP, AP50 and AP75 metrics. Similar to [he2019momentum], we do not perform hyperparameter tuning for the object detection task. See the Appendix for more implementation details.

It is unfortunate that many recent self-supervised learning papers do not discuss variance ; in fact only papers from highly resourceful labs [he2019momentum, chen2020simple, tian2020makes] report averaged results, but not always the variance. This is generally understandable, as e.g. training and evaluating a ResNet-50 model on ImageNet-1K using 4 V100 GPUs take about 6-7 days. In this paper, we tried to verify the variance of our approach for a) self-supervised pre-training on ImageNet-100, i.e. we measure the variance of MoCHi runs by training a model multiple times from scratch (Table 1), and b) the variance in the fine-tuning stage for PASCAL VOC and COCO (Tables LABEL:tab:imagenet, LABEL:tab:coco). It was unfortunately computationally infeasible for us to run multiple MoCHi pre-training runs for ImageNet-1K. In cases where standard deviation is presented, it is measured over at least 3 runs.

1 Ablations and results

We performed extensive ablations for the most important hyperparameters of MoCHi on ImageNet-100 and some are presented in Figures LABEL:fig:qablationN and LABEL:fig:qablation1024, while more can be found in the Appendix. In general we see that multiple MoCHi configurations gave consistent gains over the MoCo-v2 baseline [chen2020improved] for linear classification, with the top gains presented in Figure 1 (also averaged over 3 runs). We further show performance for different values of NN and ss in Figure LABEL:fig:qablationN and a table of gains for N=1024N=1024 in Figure LABEL:fig:qablation1024; we see that a large number of MoCHi combinations give consistent performance gains. Note that the results in these two tables are not averaged over multiple runs (for MoCHi combinations where we had multiple runs, only the first run is presented for fairness). In other ablations (see Appendix), we see that MoCHi achieves gains (+0.7%) over MoCo-v2 also when training for 100 epochs. Table 1 presents comparisons between the best-performing MoCHi variants and reports gains over the MoCo-v2 baseline. We also compare against the published results from [shen2020rethinking] a recent method that uses mixup in pixel space to synthesize harder images.

In Table LABEL:tab:imagenet we present results obtained after training on the ImageNet-1K dataset. Looking at the average negative logits plot and because both the queue and the dataset are about an order of magnitude larger for this training dataset we mostly experiment with smaller values for NN than in ImageNet-100. Our main observations are the following: a) MoCHi does not show performance gains over MoCo-v2 for linear classification on ImageNet-1K. We attribute this to the biases induced by training with hard negatives on the same dataset as the downstream task: Figures LABEL:fig:uniformity and 2(c) show how hard negative mixing reduces alignment and increases uniformity for the dataset that is used during training. MoCHi still retains state-of-the-art performance. b) MoCHi helps the model learn faster and achieves high performance gains over MoCo-v2 for transfer learning after only 100 epochs of training. c) The harder negative strategy presented in Section 4.2 helps a lot for shorter training. d) In 200 epochs MoCHi can achieve performance similar to MoCo-v2 after 800 epochs on PASCAL VOC. e) From all the MoCHi runs reported in Table LABEL:tab:imagenet as well as in the Appendix, we see that performance gains are consistent across multiple hyperparameter configurations.

In Table LABEL:tab:coco we present results for object detection and semantic segmentation on the COCO dataset [lin2014microsoft]. Following he2019momentum, we use Mask R-CNN [he2017mask] with a C4 backbone, with batch normalization tuned and synchronize across GPUs. The image scale is in pixels during training and is 800 at inference. We fine-tune all layers end-to-end on the train2017 set (118k images) and evaluate on val2017. We adopt feature normalization as in [he2019momentum] when fine-tuning. MoCHi and MoCo use the same hyper-parameters as the ImageNet supervised counterpart (i.e. we did not do any method-specific tuning). From Table LABEL:tab:coco we see that MoCHi displays consistent gains over both the supervised baseline and MoCo-v2, for both 100 and 200 epoch pre-training. In fact, MoCHi is able to reach the AP performance similar to supervised pre-training for instance segmentation (33.2) after only 100 epochs of pre-training.