Contrastive Multiview Coding

Yonglong Tian, Dilip Krishnan, Phillip Isola

Introduction

A foundational idea in coding theory is to learn compressed representations that nonetheless can be used to reconstruct the raw data. This idea shows up in contemporary representation learning in the form of autoencoders and generative models , which try to represent a data point or distribution as losslessly as possible. Yet lossless representation might not be what we really want, and indeed it is trivial to achieve – the raw data itself is a lossless representation. What we might instead prefer is to keep the “good” information (signal) and throw away the rest (noise). How can we identify what information is signal and what is noise?

To an autoencoder, or a max likelihood generative model, a bit is a bit. No one bit is better than any other. Our conjecture in this paper is that some bits are in fact better than others. Some bits code important properties like semantics, physics, and geometry, while others code attributes that we might consider less important, like incidental lighting conditions or thermal noise in a camera’s sensor.

We revisit the classic hypothesis that the good bits are the ones that are shared between multiple views of the world, for example between multiple sensory modalities like vision, sound, and touch . Under this perspective “presence of dog” is good information, since dogs can be seen, heard, and felt, but “camera pose” is bad information, since a camera’s pose has little or no effect on the acoustic and tactile properties of the imaged scene. This hypothesis corresponds to the inductive bias that the way you view a scene should not affect its semantics. There is significant evidence in the cognitive science and neuroscience literature that such view-invariant representations are encoded by the brain (e.g., ). In this paper, we specifically study the setting where the different views are different image channels, such as luminance, chrominance, depth, and optical flow. The fundamental supervisory signal we exploit is the co-occurrence, in natural data, of multiple views of the same scene. For example, we consider an image in Lab color space to be a paired example of the co-occurrence of two views of the scene, the LL view and the abab view: {L,ab}\{L,ab\}.

Our goal is therefore to learn representations that capture information shared between multiple sensory channels but that are otherwise compact (i.e. discard channel-specific nuisance factors). To do so, we employ contrastive learning, where we learn a feature embedding such that views of the same scene map to nearby points (measured with Euclidean distance in representation space) while views of different scenes map to far apart points. In particular, we adapt the recently proposed method of Contrastive Predictive Coding (CPC) , except we simplify it – removing the recurrent network – and generalize it – showing how to apply it to arbitrary collections of image channels, rather than just to temporal or spatial predictions. In reference to CPC, we term our method Contrastive Multiview Coding (CMC), although we note that our formulation is arguably equally related to Instance Discrimination . The contrastive objective in our formulation, as in CPC and Instance Discrimination, can be understood as attempting to maximize the mutual information between the representations of multiple views of the data.

We intentionally leave “good bits” only loosely defined and treat its definition as an empirical question. Ultimately, the proof is in the pudding: we consider a representation to be good if it makes subsequent problem solving easy, on tasks of human interest. For example, a useful representation of images might be a feature space in which it is easy to learn to recognize objects. We therefore evaluate our method by testing if the learned representations transfer well to standard semantic recognition tasks. On several benchmark tasks, our method achieves results competitive with the state of the art, compared to other methods for self-supervised representation learning. We additionally find that the quality of the representation improves as a function of the number of views used for training. Finally, we compare the contrastive formulation of multiview learning to the recently popular approach of cross-view prediction, and find that in head-to-head comparisons, the contrastive approach learns stronger representations.

The core ideas that we build on: contrastive learning, mutual information maximization, and deep representation learning, are not new and have been explored in the literature on representation and multiview learning for decades . Our main contribution is to set up a framework to extend these ideas to any number of views, and to empirically study the factors that lead to success in this framework. A review of the related literature is given in Section 2; and Fig. 1 gives a pictorial overview of our framework. Our main contributions are:

We apply contrastive learning to the multiview setting, attempting to maximize mutual information between representations of different views of the same scene (in particular, between different image channels).

We extend the framework to learn from more than two views, and show that the quality of the learned representation improves as number of views increase. Ours is the first work to explicitly show the benefits of multiple views on representation quality.

We conduct controlled experiments to measure the effect of mutual information estimates on representation quality. Our experiments show that the relationship between mutual information and views is a subtle one.

Our representations rival state of the art on popular benchmarks.

We demonstrate that the contrastive objective is superior to cross-view prediction.

Related work

Unsupervised representation learning is about learning transformations of the data that make subsequent problem solving easier . This field has a long history, starting with classical methods with well established algorithms, such as principal components analysis (PCA ) and independent components analysis (ICA ). These methods tend to learn representations that focus on low-level variations in the data, which are not very useful from the perspective of downstream tasks such as object recognition.

Representations better suited to such tasks have been learnt using deep neural networks, starting with seminal techniques such as Boltzmann machines , autoencoders , variational autoencoders , generative adversarial networks and autoregressive models . Numerous other works exist, for a review see . A powerful family of models for unsupervised representations are collected under the umbrella of “self-supervised” learning . In these models, an input XX to the model is transformed into an output X^\hat{X}, which is supposed to be close to another signal YY (usually in Euclidean space), which itself is related to XX in some meaningful way. Examples of such X/YX/Y pairs are: luminance and chrominance color channels of an image , patches from a single image , modalities such as vision and sound or the frames of a video . Clearly, such examples are numerous in the world, and provides us with nearly infinite amounts of training data: this is one of the appeals of this paradigm. Time contrastive networks use a triplet loss framework to learn representations from aligned video sequences of the same scene, taken by different video cameras. Closely related to self-supervised learning is the idea of multi-view learning, which is a general term involving many different approaches such as co-training , multi-kernel learning and metric learning ; for comprehensive surveys please see . Nearly all existing works have dealt with one or two views such as video or image/sound. However, in many situations, many more views are available to provide training signals for any representation.

The objective functions used to train deep learning based representations in many of the above methods are either reconstruction-based loss functions such as Euclidean losses in different norms e.g. , adversarial loss functions that learn the loss in addition to the representation, or contrastive losses e.g. that take advantage of the co-occurence of multiple views.

Some of the prior works most similar to our own (and inspirational to us) are Contrastive Predictive Coding (CPC) , Deep InfoMax , and Instance Discrimination . These methods, like ours, learn representations by contrasting between congruent and incongruent representations of a scene. CPC learns from two views – the past and future – and is applicable to sequential data, either in space or in time. Deep Infomax considers the two views to be the input to a neural network and its output. Instance Discrimination learns to match two sub-crops of the same image. CPC and Deep InfoMax have recently been extended in and respectively. These methods all share similar mathematical objectives, but differ in the definition of the views. Our method differs from these works in the following ways: we extend the objective to the case of more than two views, and we explore a different set of view definitions, architectures, and application settings. In addition, we contribute a unique empirical investigation of this paradigm of representation learning.

The idea of contrastive learning has also started to spread over many other tasks in various other domains .

Method

Our goal is to learn representations that capture information shared between multiple sensory views without human supervision. We start by reviewing previous predictive learning (or reconstruction-based learning) methods, and then elaborate on contrastive learning within two views. We show connections to mutual information maximization and extend it to scenarios including more than two views. We consider a collection of MM views of the data, denoted as V1,…,VMV_{1},\ldots,V_{M}. For each view ViV_{i}, we denote viv_{i} as a random variable representing samples following vi∼P(Vi)v_{i}\sim\mathcal{P}(V_{i}).

Let V1V_{1} and V2V_{2} represent two views of a dataset. For instance, V1V_{1} might be the luminance of a particular image and V2V_{2} the chrominance. We define the predictive learning setup as a deep nonlinear transformation from v1v_{1} to v2v_{2} through latent variables zz, as shown in Fig. 2. Formally, z=f(v1)z=f(v_{1}) and v2^=g(z)\hat{v_{2}}=g(z), where ff and gg represent the encoder and decoder respectively and v2^\hat{v_{2}} is the prediction of v2v_{2} given v1v_{1}. The parameters of the encoder and decoder models are then trained using an objective function that tries to bring v2^\hat{v_{2}} “close to” v2v_{2}. Simple examples of such an objective include the L1\mathcal{L}_{1} or L2\mathcal{L}_{2} loss functions. Note that these objectives assume independence between each pixel or element of v2v_{2} given v1v_{1}, i.e., p(v2∣v1)=Πip(v2i∣v1)p(v_{2}|v_{1})=\Pi_{i}p({v_{2}}_{i}|{v_{1}}), thereby reducing their ability to model correlations or complex structure. The predictive approach has been extensively used in representation learning, for example, colorization and predicting sound from vision .

2 Contrastive Learning with Two Views

The idea behind contrastive learning is to learn an embedding that separates (contrasts) samples from two different distributions. Given a dataset of V1V_{1} and V2V_{2} that consists of a collection of samples {v1i,v2i}i=1N\{v_{1}^{i},v_{2}^{i}\}_{i=1}^{N}, we consider contrasting congruent and incongruent pairs, i.e. samples from the joint distribution x∼p(v1,v2)x\sim p(v_{1},v_{2}) or x={v1i,v2i}x=\{v_{1}^{i},v_{2}^{i}\}, which we call positives, versus samples from the product of marginals, y∼p(v1)p(v2)y\sim p(v_{1})p(v_{2}) or y={v1i,v2j}y=\{v_{1}^{i},v_{2}^{j}\}, which we call negatives.

We learn a “critic” (a discriminating function) hθ(⋅)h_{\theta}(\cdot) which is trained to achieve a high value for positive pairs and low for negative pairs. Similar to recent setups for contrastive learning , we train this function to correctly select a single positive sample xx out of a set S={x,y1,y2,...,yk}S=\{x,y_{1},y_{2},...,y_{k}\} that contains kk negative samples:

To construct SS, we simply fix one view and enumerate positives and negatives from the other view, allowing us to rewrite the objective as:

where kk is the number of negative samples v2jv_{2}^{j} for a given sample v11v_{1}^{1}. In practice, kk can be extremely large (e.g., 1.2 million in ImageNet), and so directly minimizing Eq. 2 is infeasible. In Section 3.4, we show two approximations that allow for tractable computation.

We implement the critic hθ(⋅)h_{\theta}(\cdot) as a neural network. To extract compact latent representations of v1v_{1} and v2v_{2}, we employ two encoders fθ1(⋅)f_{\theta_{1}}(\cdot) and fθ2(⋅)f_{\theta_{2}}(\cdot) with parameters θ1\theta_{1} and θ2\theta_{2} respectively. The latent representions are extracted as z1=fθ1(v1)z_{1}=f_{\theta_{1}}(v_{1}), z2=fθ2(v2)z_{2}=f_{\theta_{2}}(v_{2}). We compute their cosine similarity as score and adjust its dynamic range by a hyper-parameter τ\tau:

Loss LcontrastV1,V2\mathcal{L}_{contrast}^{V_{1},V_{2}} in Eq. 2 treats view V1V_{1} as anchor and enumerates over V2V_{2}. Symmetrically, we can get LcontrastV2,V1\mathcal{L}_{contrast}^{V_{2},V_{1}} by anchoring at V2V_{2}. We add them up as our two-view loss:

After the contrastive learning phase, we use the representation z1z_{1}, z2z_{2}, or the concatenation of both, [z1,z2][z_{1},z_{2}], depending on our paradigm. This process is visualized in Fig. 1.

The optimal critic hθ∗h^{*}_{\theta} is proportional to the density ratio between the joint distribution p(z1,z2)p(z_{1},z_{2}) and the product of marginals p(z1)p(z2)p(z_{1})p(z_{2}) (proof provided in supplementary material):

This quantity is the pointwise mutual information, and its expectation, in Eq. 2, yields an estimator related to mutual information. A formal proof is given by , which we recapitulate in supplement, showing that:

where, as above, kk is the number of negative pairs in sample set SS. Hence minimizing the objective L\mathcal{L} maximizes the lower bound on the mutual information I(zi;zj)I(z_{i};z_{j}), which is bounded above by I(vi;vj)I(v_{i};v_{j}) by the data processing inequality. The dependency on kk also suggests that using more negative samples can lead to an improved representation; we show that this is indeed the case (see supplement). We note that recent work shows that the bound in Eq. 6 can be very weak; and finding better estimators of mutual information is an important open problem.

3 Contrastive Learning with More than Two Views

We present more general formulations of Eq. 2 that can handle any number of views. We call them the “core view” and “full graph” paradigms, which offer different tradeoffs between efficiency and effectiveness. These formulations are visualized in Fig. 3.

Suppose we have a collection of MM views V1,…,VMV_{1},\ldots,V_{M}. The “core view” formulation sets apart one view that we want to optimize over, say V1V_{1}, and builds pair-wise representations between V1V_{1} and each other view Vj,j>1V_{j},j>1, by optimizing the sum of a set of pair-wise objectives:

A second, more general formulation is the “full graph” where we consider all pairs (i,j),i≠j(i,j),i\neq j, and build (n2)\binom{n}{2} relationships in all. By involving all pairs, the objective function that we optimize is:

Both these formulations have the effect that information is prioritized in proportion to the number of views that share that information. This can be seen in the information diagrams visualized in Fig. 3. The number in each partition of the diagram indicates how many of the pairwise objectives, L(Vi,Vj)\mathcal{L}(V_{i},V_{j}), that partition contributes to. Under both the core view and full graph objectives, a factor, like “presence of dog”, that is common to all views will be preferred over a factor that affects fewer views, such as “depth sensor noise”.

The computational cost of the bivariate score function in the full graph formulation is combinatorial in the number of views. However, it is clear from Fig. 3 that this enables the full graph formulation to capture more information between different views, which may prove useful for downstream tasks. For example, the mutual information between V2V_{2} and V3V_{3} or V2V_{2} and V4V_{4} is completely ignored in the core view paradigm (as shown by a count in the information diagram). Another benefit of the full graph formulation is that it can handle missing information (e.g. missing views) in a natural manner.

4 Implementing the Contrastive Loss

Better representations using LcontrastV1,V2\mathcal{L}_{contrast}^{V_{1},V_{2}} in Eqn. 2 are learnt by using many negative samples. In the extreme case, we include every data sample in the denominator for a given dataset. However, computing the full softmax loss is prohibitively expensive for large dataset such as ImageNet. One way to approximate this full softmax distribution, as well as alleviate the computational load, is to use Noise-Contrastive Estimation (see supplement). Another solution, which we also adopt here, is to randomly sample mm negatives and do a simple (mm+11)-way softmax classification. This strategy is also used in and dates back to .

Memory bank. Following , we maintain a memory bank to store latent features for each training sample. Therefore, we can efficiently retrieve mm negative samples from the memory buffer to pair with each positive sample without recomputing their features. The memory bank is dynamically updated with features computed on the fly. The benefit of a memory bank is to allow contrasting against more negative pairs, at the cost of slightly stale features.

Experiments

We extensively evaluate Contrastive Multiview Coding (CMC) on a number of datasets and tasks. We evaluate on two established image representation learning benchmarks: ImageNet and STL-10 (see supplement). We further validate our framework on video representation learning tasks, where we use image and optical flow modalities, as the two views that are jointly learned. The last set of experiments extends our CMC framework to more than two views and provides empirical evidence of its effectiveness.

Following , we evaluate task generalization of the learned representation by training 1000-way linear classifiers on top of different layers. This is a standard benchmark that has been adopted by many papers in the literature.

Setup. Given a dataset of RGB images, we convert them to the Lab image color space, and split each image into L and ab channels, as originally proposed in SplitBrain autoencoders . During contrastive learning, L and ab from the same image are treated as the positive pair, and ab channels from other randomly selected images are treated as a negative pair (for a given L). Each split represents a view of the orginal image and is passed through a separate encoder. As in SplitBrain, we design these two encoders by evenly splitting a given deep network, such as AlexNet , into sub-networks across the channel dimension. By concatenating representations layer-wise from these two encoders, we achieve the final representation of an input image. As proposed by previous literature , the quality of such a representation is evaluated by freezing the weights of encoder and training linear classifier on top of each layer.

Implementation. Unless otherwise specified, we use PyTorch default data augmentation. Following , we set the temperature τ\tau as 0.07 and use a momentum 0.5 for memory update. We use 16384 negatives. The supplementary material provides more details on our hyperparameter settings.

CMC with AlexNet. As many previous unsupervised methods are evaluated with AlexNet on ImageNet , we also include the the results of CMC using this network. Due to the space limit, we present this comparison in supplementary material.

CMC with ResNets. We verify the effectiveness of CMC with larger networks such as ResNets . We experiment on learning from luminance and chrominance views in two colorspaces, {L,ab}\{L,ab\} and {Y,DbDr}\{Y,DbDr\} (see sec. 4.4 for validation of this choice), and we vary the width of the ResNet encoder for each view. We use the feature after the global pooling layer to train the linear classifier, and the results are shown in Table 1. {L,ab}\{L,ab\} achieves 68.3%68.3\% top-1 single crop accuracy with ResNet50x2 for each view, and switching to {Y,DbDr}\{Y,DbDr\} further brings about 0.7%0.7\% improvement. On top of it, strengthening data augmentation with RandAugment yields better or comparable results to other state-of-the-art methods .

2 CMC on videos

We apply CMC on videos by drawing insight from the two-streams hypothesis , which posits that human visual cortex consists of two distinct processing streams: the ventral stream, which performs object recognition, and the dorsal stream, which processes motion. In our formulation, given an image iti_{t} that is a frame centered at time tt, the ventral stream associates it with a neighbouring frame it+ki_{t+k}, while the dorsal stream connects it to optical flow ftf_{t} centered at tt. Therefore, we extract iti_{t}, it+ki_{t+k} and ftf_{t} from two modalities as three views of a video; for optical flow we use the TV-L1 algorithm . Two separate contrastive learning objectives are built within the ventral stream (it,it+k)(i_{t},i_{t+k}) and within the dorsal stream (it,ft)(i_{t},f_{t}). For the ventral stream, the negative sample for iti_{t} is chosen as a random frame from another randomly chosen video; for the dorsal stream, the negative sample for iti_{t} is chosen as the flow corresponding to a random frame in another randomly chosen video.

Pre-training. We train CMC on UCF101 and use two CaffeNets for extracting features from images and optical flows, respectively. In our implementation, ftf_{t} represents 10 continuous flow frames centered at tt. We use batch size of 128 and contrast each positive pair with 127 negative pairs.

Action recognition. We apply the learnt representation to the task of action recognition. The spatial network from is a well-established paradigm for evaluating pre-trained RGB network on action recognition task. We follow the same spirit and evaluate the transferability of our RGB CaffeNet on UCF101 and HMDB51 datasets. We initialize the action recognition CaffeNet up to conv5 using the weights from the pre-trained RGB CaffeNet. The averaged accuracy over three splits is present in Table 2. Unifying both ventral and dorsal streams during pre-training produces higher accuracy for downstream recognition than using only single stream. Increasing the number of views of the data from 22 to 33 (using both streams instead of one) provides a boost for UCF-101.

3 Extending CMC to More Views

We further extend our CMC learning framework to multiview scenarios. We experiment on the NYU-Depth-V2 dataset which consists of 1449 labeled images. We focus on a deeper understanding of the behavior and effectiveness of CMC. The views we consider are: luminance (L channel), chrominance (ab channel), depth, surface normal , and semantic labels.

Setup. To extract features from each view, we use a neural network with 5 convolutional layers, and 2 fully connected layers. As the size of the dataset is relatively small, we adopt the sub-patch based contrastive objective (see supplement) to increase the number of negative pairs. Patches with a size of 128×128128\times 128 are randomly cropped from the original images for contrastive learning (from images of size 480×640480\times 640). For downstream tasks, we discard the fully connected layers and evaluate using the convolutional layers as a representation.

To measure the quality of the learned representation, we consider the task of predicting semantic labels from the representation of LL. We follow the core view paradigm and use LL as the core view, thus learning a set of representations by contrasting different views with LL. A UNet style architecture is utilized to perform the segmentation task. Contrastive training is performed on the above architecture that is equivalent of the UNet’s encoder. After contrastive training is completed, we initialize the encoder weights of the UNet from the LL encoder (which are equivalent architectures) and keep them frozen. Only the decoder is trained during this finetuning stage.

Since we use the patch-based contrastive loss, in the 1 view setting case, CMC coincides with DIM . The 2-4 view cases contrast L with ab, and then sequentially add depth and surface normals. The semantic labeling results are measured by mean IoU over all classes and pixel accuracy, shown in Fig. 4. We see that the performance steadily improves as new views are added. We have tested different orders of adding the views, and they all follow a similar pattern.

We also compare CMC with two baselines. First, we randomly initialize and freeze the encoder, and we call this the Random baseline; it serves as a lower bound on the quality since the representation is just a random projection. Rather than freezing the randomly initialized encoder, we could train it jointly with the decoder. This end-to-end Supervised baseline serves as an upper bound. The results are presented in Table 3, which shows our CMC produces high quality feature maps even though it’s unaware of the downstream task.

3.2 Is CMC improving all views?

A desirable unsupervised representation learning algorithm operating on multiple views or modalities should improve the quality of representations for all views. We therefore investigate our CMC framwork beyond L channel. To treat all views fairly, we train these encoders following the full graph paradigm, where each view is contrasted with all other views.

We evaluate the representation of each view vv by predicting the semantic labels from only the representation of vv, where vv is L, ab, depth or surface normals. This uses the full-graph paradigm. As in the previous section, we compare CMC with Random and Supervised baselines. As shown in Table 4, the performance of the representations learned by CMC using full-graph significantly outperforms that of randomly projected representations, and approaches the performance of the fully supervised representations. Furthermore, the full-graph representation provides a good representation learnt for all views, showing the importance of capturing different types of mutual information across views.

3.3 Predictive Learning vs. Contrastive Learning

While experiments in section 4.1 show that contrastive learning outperforms predictive learning in the context of Lab color space, it’s unclear whether such an advantage is due to the natural inductive bias of the task itself. To further understand this, we go beyond chrominance (ab), and try to answer this question when geometry or semantic labels are present.

We consider three view pairs on the NYU-Depth dataset: (1) L and depth, (2) L and surface normals, and (3) L and segmentation map. For each of them, we train two identical encoders for L, one using contrastive learning and the other with predictive learning. We then evaluate the representation quality by training a linear classifier on top of these encoders on the STL-10 dataset.

The comparison results are shown in Table 5, which shows that contrastive learning consistently outperforms predictive learning in this scenario where both the task and the dataset are unknown. We also include “random” and “supervised” baselines similar to that in previous sections. Though in the unsupervised stage we only use 1.3K images from a dataset much different from the target dataset STL-10, the object recognition accuracy is close to the supervised method, which uses an end-to-end deep network directly trained on STL-10.

Given two views V1V_{1} and V2V_{2} of the data, the predictive learning approach approximately models p(v2∣v1)p(v_{2}|v_{1}). Furthermore, losses used typically for predictive learning, such as pixel-wise reconstruction losses usually impose an independence assumption on the modeling: p(v2∣v1)≈Πip(v2i∣v1)p(v_{2}|v_{1})\approx\Pi_{i}p({v_{2}}_{i}|{v_{1}}). On the other hand, the contrastive learning approach by construction does not assume conditional independence across dimensions of v2v_{2}. In addition, the use of random jittering and cropping between views allows the contrastive learning approach to benefit from spatial co-occurrence (contrasting in space) in addition to contrasting across views. We conjecture that these are two reasons for the superior performance of contrastive learning approaches over predictive learning.

4 How does mutual information affect representation quality?

Given a fixed set of views, CMC aims to maximize the mutual information between representations of these views. We have found that maximizing information in this way indeed results in strong representations, but it would be incorrect to infer that information maximization (infomax) is the key to good representation learning. In fact, this paper argues for precisely the opposite idea: that cross-view representation learning is effective because it results in a kind of information minimization, discarding nuisance factors that are not shared between the views.

The resolution to this apparent dilemma is that we want to maximize the “good” information – the signal – in our representations, while minimizing the “bad” information – the noise. The idea behind CMC is that this can be achieved by doing infomax learning on two views that share signal but have independent noise. This suggests a “Goldilocks principle” : a good collection of views is one that shares some information but not too much. Here we test this hypothesis on two domains: learning representations on images with different colorspaces forming the two views; and learning representations on pairs of patches extracted from an image, separated by varying spatial distance.

In patch experiments we randomly crop two RGB patches of size 64x64 from the same image, and use these patches as the two views. Their relative position is fixed. Namely, the two patches always starts at position (x,y)(x,y) and (x+d,y+d)(x+d,y+d) with (x,y)(x,y) being randomly sampled. While varying the distance dd, we start from 6464 to avoid overlapping. There is a possible bias that with an image of relatively small size (e.g., 512x512), a large dd (e.g., 384) will always push these two patches around boundary. To minimize this bias, we use high resolution images (e.g. 2k2k) from DIV2K dataset.

Fig. 5 shows the results of these experiments. The left plot shows the result of learning representations on different colorspaces (splitting each colorspace into two views, such as (L, ab), (R, GB) etc). We then use the MINE estimator to estimate the mutual information between the views. We measure representation quality by training a linear classifier on the learned representations on the STL-10 dataset . The plots clearly show that using colorspaces with minimal mutual information give the best downstream accuracy (For the outlier HSV in this plot, we conjecture the representation quality is harmed by the periodicity of H. Note that the H in HED is not periodic.). On the other hand, the story is more nuanced for representations learned between patches at different offsets from each other (Fig. 5, right). Here we see that views with too little or too much MI perform worse; a sweet spot in the middle exists which gives the best representation. That there exists such a sweet spot should be expected. If two views share no information, then, in principle, there is no incentive for CMC to learn anything. If two views share all their information, no nuisances are discarded and we arrive back at something akin to an autoencoder or generative model, that simply tries to represent all the bits in the multiview data.

These experiments demonstrate that the relationship between mutual information and representation quality is meaningful but not direct. Selecting optimal views, which just share relevant signal, may be a fruitful direction for future research.

Conclusion

We have presented a contrastive learning framework which enables the learning of unsupervised representations from multiple views of a dataset. The principle of maximization of mutual information enables the learning of powerful representations. A number of empirical results show that our framework performs well compared to predictive learning and scales with the number of views.

Acknowledgements Thanks to Devon Hjelm for providing implementation details of Deep InfoMax, Zhirong Wu and Richard Zhang for helpful discussion and comments. This material is based on resources supported by Google Cloud.

References

Appendix A ImageNet-100 Proposed in this Paper

In this paper, we proposed a subset of ImageNet that contains randomly selected 100 classes for ablation study as well as hyper-parameter tuning. To ease a relevant study in the future, we release the list of these categories that we consistently used throughout all our experiments, as summarized in Table 6.

Appendix B Contrastive Loss

In addition to the subsampled (mm+11)-way softmax cross-entropy, Noise-Contrastive Estimation (NCE ) is another way to approximate the NN-way (NN is the dataset size) softmax cross entropy full softmax in Eqn 2. Compared with the (mm+11)-way softmax cross-entropy, NCE is computationally faster and may result in slightly worse performance in standard linear evaluation. It has been used in Confusingly, the literature has previously referred to Eqn.2 as “InfoNCE” . Our NCE approximation does not refer to the allusion to NCE in the name “InfoNCE”. Rather we are here describing an NCE approximation to the “InfoNCE” softmax objective.. We depict its general idea as below.

Given an anchor v1iv_{1}^{i} from V1V_{1}, the probablity that an atom v2∈{v2j∣j=1,2,...,N}v_{2}\in\{v_{2}^{j}|j=1,2,...,N\} from V2V_{2} is the best match of v1iv_{1}^{i}, using the score hθh_{\theta} is given by:

where the normalization factor Z=∑j=1Nhθ({v1i,v2j})Z=\sum_{j=1}^{N}h_{\theta}(\{v_{1}^{i},v_{2}^{j}\}) is expensive to compute for large NN.

NCE is an effective way to estimate unnormalized statistical models. NCE fits a density model pp to data distributed as (unknown) distribution pdp_{d}, by using a binary classifier to distinguish it from noise samples distributed as pnp_{n}. To learn p(v2∣v1i)p(v_{2}|v_{1}^{i}), we use a binary classifier, which treats v2iv_{2}^{i} as the data (or positive) sample when given v1iv_{1}^{i}. The noise distribution pn(⋅∣v1i)p_{n}(\cdot|v_{1}^{i}) we choose here is a uniform distribution over all atoms from V2V_{2}, i.e., pn(⋅∣v1i)=pn(⋅)=1/Np_{n}(\cdot|v_{1}^{i})=p_{n}(\cdot)=1/N. If we sample mm noise samples to pair with each data sample, the posterior probability that a given atom v2v_{2} comes from the data distribution is:

and we estimate this probability by replacing pd(v2∣v1i)p_{d}(v_{2}|v_{1}^{i}) with our unnormalized model distribution hθ(v1i,v2)/Z0h_{\theta}(v_{1}^{i},v_{2})/Z_{0}, where Z0Z_{0} is a constant estimated from the first batch. Minimizing the negative log-posterior probability of correct labels DD over data and noise samples yields our final objective, which is the NCE-based approximation of Eq. 2 (p^\hat{p} is the empirical data distribution):

B.2 Contrasting Sub-patches

Instead of contrasting features from the last layer, patch-based method contrasts feature from the last layer with features from previous layers, hence increasing the number of negative pairs. For instance, we use features from the last layer of fθ1f_{\theta_{1}} to contrast with feature points from feature maps produced by the first several conv layers of fθ2f_{\theta_{2}}. This is equivalent to contrast between global patch from one view with local patches from the other view. In this fashion, we directly perform m+1m+1 way softmax classification, the same as for a fair comparison in Sec. D.1.

Such patch-based contrastive loss is computed within each mini-batch and does not require a memory bank. Therefore, deploying it in parallel training schemes is easy and flexible. However, patch-based contrastive loss usually yields suboptimal results compared to NCE-based contrastive loss, according to our experiments.

Appendix C Proofs

We prove that: (a) the optimal score function hθ∗({v1,v2})h^{*}_{\theta}(\{v_{1},v_{2}\}) is proportional to density ratio between the joint distribution p(v1,v2)p(v_{1},v_{2}) and product of marginals p(v1)p(v2)p(v_{1})p(v_{2}), as shown in Eq. 5; (b) Minimizing the contrastive loss Lcontrast\mathcal{L}_{contrast} maxmizes a lower bound on the mutual information between two views, as shown in Eq. 6.

We will use the most general formula of contrastive loss Lcontrast\mathcal{L}_{contrast} shown in Eq. 1 for our derivation. But we note that replacing Lcontrast\mathcal{L}_{contrast} with LcontrastV1,V2\mathcal{L}_{contrast}^{V_{1},V_{2}} is straightforward. The overall proof follows a similar derivation introduced in .

We first show that the optimal score function hθ∗({v1,v2})h^{*}_{\theta}(\{v_{1},v_{2}\}) that minimizes Eq. 1 is proportional to the density ratio between joint distribution and product of marginals, shown as Eq. 5. For notation convenience, we denote p(v1,v2)p(v_{1},v_{2}) as data distribution pd(⋅)p_{d}(\cdot) and p(v1)p(v2)p(v_{1})p(v_{2}) as noise distribution pn(⋅)p_{n}(\cdot). The loss in Eq. 1 is indeed a cross-entropy loss of classifying the correct positive pair out from the given set SS. Without loss of generality, we assume the first pair (v10,v20)(v_{1}^{0},v_{2}^{0}) in SS is positive or congruent and all others (v1i,v2i),i=1,2,...,k(v_{1}^{i},v_{2}^{i}),i=1,2,...,k are negative or incongruent. The optimal probability for the loss, p(pos=0∣S)p(pos=0|S), should depict the fact that (v10,v20)(v_{1}^{0},v_{2}^{0}) comes from the data distribution pd(⋅)p_{d}(\cdot) while all other pairs come from the noise distribution pn(⋅)p_{n}(\cdot). Therefore,

where we plug in the definition of pd(⋅)p_{d}(\cdot) and pn(⋅)p_{n}(\cdot), and divide ∏i=0kp(v1i)p(v22)\prod_{i=0}^{k}p(v_{1}^{i})p(v_{2}^{2}) for both the numerator and denominator. By comparing above equation with the loss function in Eq. 1, we can see that the optimal score function hθ∗({v1,v2})h^{*}_{\theta}(\{v_{1},v_{2}\}) is proportional to the density ratio p(v1,v2)p(v1)p(v2)\frac{p(v_{1},v_{2})}{p(v_{1})p(v_{2})}. The above derivation is agnostic to which layer the score function starts from, e.g., hh can be defined on either the raw input (v1,v2)(v_{1},v_{2}) or the latent representation (z1,z2)(z_{1},z_{2}). As we care more about the property of the latent representation, for the following derivation we will use h∗({z1,z2})h^{*}(\{z_{1},z_{2}\}), which is proportional to p(z1,z2)p(z1)p(z2)\frac{p(z_{1},z_{2})}{p(z_{1})p(z_{2})}.

C.2 Maximizing lower bound on MI

Now we substitute the score function in Eq. 1 with the above density ratio, and the optimal loss objective Lcontrastopt\mathcal{L}_{contrast}^{opt} becomes:

Therefore, for any two views ViV_{i} and VjV_{j}, we have I(zi;zj)≥log⁡(k)−Lcontrastopt(Vi,Vj)I(z_{i};z_{j})\geq\log(k)-\mathcal{L}_{contrast}^{opt}(V_{i},V_{j}). As the kk increases, the approximation step becomes more accurate. Given any kk, minimizing Lk(Vi,Vj)\mathcal{L}_{k}(V_{i},V_{j}) maximizes the lower bound on the mutual information I(zi;zj)I(z_{i};z_{j}). We should note that increasing kk to infinity does not always lead to a higher lower bound. While log⁡(k)\log(k) increases with a larger kk, the optimization problem becomes harder and Lk(Vi,Vj)\mathcal{L}_{k}(V_{i},V_{j}) also increases.

Appendix D Additional Experiments

STL-10 is an image recognition dataset designed for developing unsupervised or self-supervised learning algorithms. It consists of 100000100000 unlabeled training 96×9696\times 96 RGB image samples and 500500 labeled samples for each of the 1010 classes.

Setup. We adopt the same data augmentation strategy and network architecture as those in DIM . A variant of AlexNet takes as input 64×6464\times 64 images, which are randomly cropped and horizontally flipped from the original 96×9696\times 96 size images. For a fair comparison with DIM, we also train our model in a patch-based contrastive fashion during unsupervised pre-training. With the weights of the pre-trained encoder frozen, a two-layer fully connected network with 200 hidden units is trained on top of different layers for 100 epochs to perform 10-way classification. We also investigated the strided crop strategy of CPC . Fixed sized overlapping patches of size 16×1616\times 16 with an overlap of 88 pixels are cropped and fed into the network separately. This ensures that features of one patch contain minimal information from neighbouring patches; and increases the available number of negative pairs for the contrastive loss. Additionally, we include NCE-based contrastive training and linear classifier evaluation.

Comparison. We compare CMC with the state of the art unsupervised methods in Table 7. Three columns are shown: the conv5 and fc7 columns use respectively these layers of AlexNet as the encoder (again remembering that we split across channels for L and ab views). For these two columns we can compare against the all methods except CPC, since CPC does not report these numbers in their paper . In the Strided Crop setup, we only compare against the approaches that use contrastive learning, DIM and CPC, since this method was only used by those works. We note that in Table 7 for all the methods except SplitBrain, we report numbers are shown in the original paper. For SplitBrain, we reimplemented their model faithfully and report numbers based on our reimplementation (we verified the accuracy of our SplitBrain code by the fact that we get very similar results with our reimpementation as in the original paper for ImageNet experiments, see below).

The family of contrastive learning methods, such as DIM, CPC, and CMC, achieve higher classification accuracy than other methods such as SplitBrain that use predictive learning; or BiGAN that use adversarial learning. CMC significantly outperforms DIM and CPC in all cases. We hypothesize that this outperformance results from the modeling of cross-view mutual information, where view-specific noisy details are discarded. Another head-to-head comparison happens between CMC and SplitBrain, both of which modeling images as seprated L and ab streams; we achieve a nearly 8%8\% absolute improvement for conv5 and 17%17\% improvement for fc5. Finally, we notice that the predictive learning methods suffer from a big drop in performance when the encoding layer is switched from conv5 to fc7. On the other hand, the contrastive learning approaches are much more stable across layers, suggesting that the mutual information maximization paradigm learns more semantically meaningful representations shared by the different views. From a practical perspective, this is a significant advantage as the selection of specific layers should ideally not change downstream performance by too much.

In this experiments we used AlexNet as backbone. Switching to more powerful networks such as ResNets is likely to further improve the representation quality.

D.2 CMC on ImageNet with AlexNet

ImageNet consists of 1000 image classes and is frequently considered as a testbed for unsupervised representation learning algorithms.

To compare with other methods, we adopt standard AlexNet and split it into two encoders. Because of splitting, each layer only connects to half of the neurons in the previous layer, and therefore the number of parameters in our model halves. We remove local response layer and add batch normalization to each layer. For the memory-based CMC model, we adopt ideas from for computing and storing a memory. We retrieve 40964096 negative pairs from the memory bank to contrast each positive pair (the effect of the number of negatives is shown in Sec. D.3). The training details are present in Sec. E.2.

Table 8 shows the results of comparing the CMC against other models, both predictive and contrastive. Our CMC is the best among all these methods; futhermore CMC tends to perform better at higher convolutional layers, similar to another contrasting-based model Inst-Dis .

D.3 Number of negatives

Effect of the number of negative samples. We investigate the relationship between the number of negative pairs mm in NCE-based loss and the downstream classification accuracy on a randomly chosen subset of 100100 classes of Imagenet (the same set of classes is used for any number of negative pairs). We train a 100-way linear classifier using CMC pre-trained features with varying number of negative pairs, starting from 6464 pairs upto 81928192 (in multiples of 22). Fig. 6 shows that the accuracy of the resulting classifier steadily increases but saturates at around 60.3%60.3\% with m=4096m=4096 samples. We used AlexNet and the NCE approximation in this study ((mm+1)-way softmax cross entropy, a.k.a. InfoNCE, also follow a similar trend).

D.4 Compatibility with other methods

To test the compatibility of CMC with mechanisms proposed in other self-supervised learning methods, we consider combining CMC with MoCo (i.e., switching from memory bank to momentum encoder) and PIRL (i.e., applying a second JigSaw branch). In this setup, we perform both unsupervised pre-training and linear evaluation on the same ImageNet-100 subset as above. We use the same contrastive loss objective function and training recipe for all methods, to ensure a head-to-head comparison. Specifically, we pre-train for 240 epochs with learning rate initialized as 0.03 and decayed with cosine annealing schedule.

Table 9 summarizes the results. We observe that combining CMC with the MoCo mechanism or JigSaw branch in PIRL can consistently improve the performance, verifying that they are compatible.

Appendix E Implementation Details

For a fair comparison with DIM and CPC , we adopt the same architecture as that used in DIM and split it into two encoders, each shown as in Table 10. For the implementation of the score function, we adopt similar “encoder-and-dot-product” strategy, which is tantamount to a bilinear model.

In the patch-based contrastive learning stage, we use Adam optimizer with an initial learning rate of 0.0010.001, β1=0.5\beta_{1}=0.5, β2=0.999\beta_{2}=0.999. We train for a total of 200200 epochs with learning rate decayed by 0.20.2 after 120120 and 160160 epochs. In the non-linear classifier evaluation stage, we use the same optimizer setting. For the NCE-based contrastive learning stage, we train for 320 epochs with the learning rate initialized as 0.030.03 and further decayed by 10 for every 4040 epochs after the first 200200 epochs. The temperature τ\tau is set as 0.10.1. In general, τ∈[0.05,0.2]\tau\in\left[0.05,0.2\right] works reasonably well.

E.2 ImageNet

For patch-based contrastive loss, we use the same optimizer setting as in Sec. E.1 except that the learning rate is initialized as 0.01.

For NCE-basd contrastive loss in both full ImageNet and ImageNet100 experiments present in Sec. D.3, the encoder architecture used for either L or ab channels is shown in Table 11. In the unsupervised learning stage of AlexNet, we use SGD to train the network for a total of 200200 epochs. The temperature τ\tau is set as 0.070.07 by following previous work . The learning rate is initialized as 0.030.03 with a decay of 10 for every 4040 epochs after the first 120120 epochs. Weight decay is set as 10−410^{-4} and momentum is kept as 0.90.9. For the linear classification stage, we train for 100100 epochs. The learning rate is initialized as 0.10.1 and decayed by 0.20.2 every 1515 epochs after the first 6060 epochs. We set weight decay as and momentum as 0.90.9.

For ResNets in CMC stage, instead of using step decay, we choose cosine annealing to gradually decrease the learning rate. In the linear evaluation stage, we train for 100100 epochs. The learning rate is initialized as 3030 for ResNet-50 and ResNet-101, and 5050 for ResNet-50 x2. It is decayed by 0.20.2 every 1515 epochs after the first 6060 epochs. We set weight decay as and momentum as 0.90.9.

E.3 UCF101 and HMDB51

Following previous work , we use CaffeNet for the video experiments. We tailor the network and use features from the fc6 layer for contrastive learning. Dropout of 0.50.5 is used to alleviate overfitting.

E.4 NYU Depth-V2

While experimenting with different views on NYU Depth-V2 dataset, we encode the features from patches with a size of 128×128128\times 128. The detailed architecture is shown in Table 12. In the unsupervised training stage, we use Adam optimizer with an initial learning rate of 0.0010.001, β1=0.5\beta_{1}=0.5, β2=0.999\beta_{2}=0.999. We train for a total of 30003000 epochs with learning rate decayed by 0.20.2 after 20002000, 24002400, and 28002800 epochs. For the downstream semantic segmentation task, we use the same optimizer setting but train for fewer epochs. We only train 200200 epochs for CMC pre-trained models, and train 10001000 epochs for the Random and Supervised baselines until convergence. For the classification task evaluated on STL-10, we use the same optimizer setting as in Sec. E.1 to report numbers.

Appendix F Change Log

arXiv v2 Added references to Time Contrastive Networks and Local Aggregation . Fixed Typos.

arXiv v3 Added analysis of effect of mutual information, updated results, and rearranged the contents.

arXiv v4 Added reference to the original k-pair loss and other related work. Cited some recent works to reflect progresses after our v1 version. We also removed Fast AutoAugment and rearranged the contents.

arXiv v5 Added the list of ImageNet-100 proposed in this paper, and the compatibility of CMC with other methods.