Understanding Contrastive Representation Learning through Alignment and Uniformity on the Hypersphere

Tongzhou Wang, Phillip Isola

Introduction

Intuitively, having the features live on the unit hypersphere leads to several desirable traits. Fixed-norm vectors are known to improve training stability in modern machine learning where dot products are ubiquitous (Xu & Durrett, 2018; Wang et al., 2017). Moreover, if features of a class are sufficiently well clustered, they are linearly separable with the rest of feature space (see Figure 2), a common criterion used to evaluate representation quality.

While the unit hypersphere is a popular choice of feature space, not all encoders that map onto it are created equal. Recent works argue that representations should additionally be invariant to unnecessary details, and preserve as much information as possible (Oord et al., 2018; Tian et al., 2019; Hjelm et al., 2018; Bachman et al., 2019). Let us call these two properties alignment and uniformity (see Figure 1). Alignment favors encoders that assign similar features to similar samples. Uniformity prefers a feature distribution that preserves maximal information, i.e., the uniform distribution on the unit hypersphere.

In this work, we analyze the alignment and uniformity properties. We show that a currently popular form of contrastive representation learning in fact directly optimizes for these two properties in the limit of infinite negative samples. We propose theoretically-motivated metrics for alignment and uniformity, and observe strong agreement between them and downstream task performance. Remarkably, directly optimizing for these two metrics leads to comparable or better performance than contrastive learning.

We propose quantifiable metrics for alignment and uniformity as two measures of representation quality, with theoretical motivations.

We prove that the contrastive loss optimizes for alignment and uniformity asymptotically.

Empirically, we find strong agreement between both metrics and downstream task performance.

Despite being simple in form, our proposed metrics, when directly optimized with no other loss, empirically lead to comparable or better performance at downstream tasks than contrastive learning.

Related Work

has seen remarkable success in learning representations for image and sequential data (Logeswaran & Lee, 2018; Wu et al., 2018; Oord et al., 2018; Hénaff et al., 2019; Tian et al., 2019; Hjelm et al., 2018; Bachman et al., 2019; Tian et al., 2019; He et al., 2019; Chen et al., 2020a). The common motivation behind these work is the InfoMax principle (Linsker, 1988), which we here instantiate as maximizing the mutual information (MI) between two views (Tian et al., 2019; Bachman et al., 2019; Wu et al., 2020). However, this interpretation is known to be inconsistent with the actual behavior in practice, e.g., optimizing a tighter bound on MI can lead to worse representations (Tschannen et al., 2019). What the contrastive loss exactly does remains largely a mystery. Analysis based on the assumption of latent classes provides nice theoretical insights (Saunshi et al., 2019), but unfortunately has a rather large gap with empirical practices: the result that representation quality suffers with a large number of negatives is inconsistent with empirical observations (Wu et al., 2018; Tian et al., 2019; He et al., 2019; Chen et al., 2020a). In this paper, we analyze and characterize the behavior of contrastive learning from the perspective of alignment and uniformity properties, and empirically verify our claims with standard representation learning tasks.

Outside contrastive learning, many other representation learning approaches also normalize their features to be on the unit hypersphere. In variational autoencoders, the hyperspherical latent space has been shown to perform better than the Euclidean space (Xu & Durrett, 2018; Davidson et al., 2018). Directly matching uniformly sampled points on the unit hypersphere is known to provide good representations (Bojanowski & Joulin, 2017), agreeing with our intuition that uniformity is a desirable property. Mettes et al. (2019) optimizes prototype representations on the unit hypersphere for classification. Hyperspherical face embeddings greatly outperform the unnormalized counterparts (Parkhi et al., 2015; Liu et al., 2017; Wang et al., 2017; Schroff et al., 2015). Its empirical success suggests that the unit hypersphere is indeed a nice feature space. In this work, we formally investigate the interplay between the hypersphere geometry and the popular contrastive representation learning.

The problem of uniformly distributing points on the unit hypersphere is a well-studied one. It is often defined as minimizing the total pairwise potential w.r.t. a certain kernel function (Borodachov et al., 2019; Landkof, 1972), e.g., the Thomson problem of finding the minimal electrostatic potential energy configuration of electrons (Thomson, 1904), and minimization of the Riesz ss-potential (Götz & Saff, 2001; Hardin & Saff, 2005; Liu et al., 2018). The uniformity metric we propose is based on the Gaussian potential, which can be used to represent a very general class of kernels and is closely related to the universally optimal point configurations (Borodachov et al., 2019; Cohn & Kumar, 2007). Additionally, the best-packing problem on hyperspheres (often called the Tammes problem) is also well studied (Tammes, 1930).

Preliminaries on Unsupervised Contrastive Representation Learning

The popular unsupervised contrastive representation learning method (often referred to as contrastive learning in this paper) learns representations from unlabeled data. It assumes a way to sample positive pairs, representing similar samples that should have similar representations. Empirically, the positive pairs are often obtained by taking two independently randomly augmented versions of the same sample, e.g. two crops of the same image (Wu et al., 2018; Hjelm et al., 2018; Bachman et al., 2019; He et al., 2019; Chen et al., 2020a).

Distributions pdatap_{\mathsf{data}} and pposp_{\mathsf{pos}} should satisfy

Symmetry: ∀x,y, ppos(x,y)=ppos(y,x)\forall x,y,~{}p_{\mathsf{pos}}(x,y)=p_{\mathsf{pos}}(y,x).

The term contrastive loss has also been generally used to refer to various objectives based on positive and negative samples, e.g., in Siamese networks (Chopra et al., 2005; Hadsell et al., 2006). In this work, we focus on the specific form in Equation (1) that is widely used in modern unsupervised contrastive representation learning literature.

Without the norm constraint, the softmax\mathtt{softmax} distribution can be made arbitrarily sharp by simply scaling all the features. Wang et al. (2017) provided an analysis on this effect and argued for the necessity of normalization when using feature vector dot products in a cross entropy loss, as is in Eqn. (1). Experimentally, Chen et al. (2020a) also showed that normalizing outputs leads to superior representations.

Many empirical works are motivated by the InfoMax principle of maximizing I(f(x);f(y))I(f(x);f(y)) for (x,y)∼ppos(x,y)\sim p_{\mathsf{pos}} (Tian et al., 2019; Bachman et al., 2019; Wu et al., 2020). Usually they interpret Lcontrastive\mathcal{L}_{\mathsf{contrastive}} in Eqn. (1) as a lower bound of I(f(x);f(y))I(f(x);f(y)) (Oord et al., 2018; Hjelm et al., 2018; Bachman et al., 2019; Tian et al., 2019). However, this interpretation is known to have issues in practice, e.g., maximizing a tighter bound often leads to worse downstream task performance (Tschannen et al., 2019). Therefore, instead of viewing it as a bound, we investigate the exact behavior of directly optimizing Lcontrastive\mathcal{L}_{\mathsf{contrastive}} in the following sections.

Feature Distribution on the Hypersphere

The contrastive loss encourages learned feature representation for positive pairs to be similar, while pushing features from the randomly sampled negative pairs apart. Conventional wisdom says that representations should extract the most shared information between positive pairs and remain invariant to other noise factors (Linsker, 1988; Tian et al., 2019; Wu et al., 2020; Bachman et al., 2019). Therefore, the loss should prefer two following properties:

Alignment: two samples forming a positive pair should be mapped to nearby features, and thus be (mostly) invariant to unneeded noise factors.

Uniformity: feature vectors should be roughly uniformly distributed on the unit hypersphere Sm−1\mathcal{S}^{m-1}, preserving as much information of the data as possible.

To empirically verify this, we visualize CIFAR-10 (Torralba et al., 2008; Krizhevsky et al., 2009) representations on S1\mathcal{S}^{1} (m=2m=2) obtained via three different methods:

Supervised predictive learning: An encoder and a linear classifier are jointly trained from scratch with cross entropy loss on supervised labels.

Unsupervised contrastive learning: An encoder is trained w.r.t. Lcontrastive\mathcal{L}_{\mathsf{contrastive}} with τ=0.5\tau=0.5 and M=256M=256.

All three encoders share the same AlexNet based architecture (Krizhevsky et al., 2012), modified to map input images to 22-dimensional vectors in S1\mathcal{S}^{1}. Both predictive and contrastive learning use standard data augmentations to augment the dataset and sample positive pairs.

Figure 3 summarizes the resulting distributions of validation set features. Indeed, features from unsupervised contrastive learning (bottom in Figure 3) exhibit the most uniform distribution, and are closely clustered for positive pairs.

The form of the contrastive loss in Eqn. (1) also suggests this. We present informal arguments below, followed by more formal treatment in Section 4.2. From the symmetry of pp, we can derive

which is akin to maximizing pairwise distances with a LogSumExp\mathtt{LogSumExp} transformation. Intuitively, pushing all features away from each other should indeed cause them to be roughly uniformly distributed.

For further analysis, we need a way to measure alignment and uniformity. We propose the following two metrics (losses).

The alignment loss is straightforwardly defined with the expected distance between positive pairs:

1.2 Uniformity

and define the uniformity loss as the logarithm of the average pairwise Gaussian potential:

The average pairwise Gaussian potential is nicely tied with the uniform distribution on the unit hypersphere.

σd\sigma_{d} denotes the normalized surface area measure on Sd\mathcal{S}^{d}.

First, we show that the uniform distribution is the unique distribution that minimize the expected pairwise potential.

For M(Sd)\mathcal{M}(\mathcal{S}^{d}) the set of Borel probability measures on Sd\mathcal{S}^{d}, σd\sigma_{d} is the unique solution of

In addition, as number of points goes to infinity, distributions of points minimizing the average pairwise potential converge weak∗ to the uniform distribution. Recall the definition of the weak∗ convergence of measures.

For each N>0N>0, the NN point minimizer of the average pairwise potential is

The normalized counting measures associated with the {uN∗}N=1∞\{\mathbf{u}^{*}_{N}\}_{N=1}^{\infty} sequence converge weak∗ to σd\sigma_{d}.

Designing an objective minimized by the uniform distribution is in fact nontrivial. For instance, average pairwise dot products or Euclidean distances is simply optimized by any distribution that has zero mean. Among kernels that achieve uniformity at optima, the Gaussian kernel is special in that it is closely related to the universally optimal point configurations and can also be used to represent a general class of other kernels, including the Riesz ss-potentials. We refer readers to Borodachov et al. (2019) and Cohn & Kumar (2007) for in-depth discussions on these topics. Moreover, as we show below, Luniform\mathcal{L}_{\mathsf{uniform}}, defined with the Gaussian kernel, has close connections with Lcontrastive\mathcal{L}_{\mathsf{contrastive}}.

Empirically, we evaluate the average pairwise potential of various finite point collections on S1\mathcal{S}^{1} in Figure 4. The values nicely align with our intuitive understanding of uniformity.

We further discuss properties of Luniform\mathcal{L}_{\mathsf{uniform}} and characterize its optimal value and range in the appendix.

2 Limiting Behavior of Contrastive Learning

We first define the notion of optimal encoders for each of these two metrics.

We say an encoder ff is perfectly aligned if f(x)=f(y)f(x)=f(y) a.s. over (x,y)∼ppos(x,y)\sim p_{\mathsf{pos}}.

We say an encoder ff is perfectly uniform if the distribution of f(x)f(x) for x∼pdatax\sim p_{\mathsf{data}} is the uniform distribution σm−1\sigma_{m-1} on Sm−1\mathcal{S}^{m-1}.

We analyze the asymptotics with infinite negative samples. Existing empirical work has established that larger number of negative samples consistently leads to better downstream task performances (Wu et al., 2018; Tian et al., 2019; He et al., 2019; Chen et al., 2020a), and often uses very large values (e.g., M=65536M=65536 in He et al. (2019)). The following theorem nicely confirms that optimizing w.r.t. the limiting loss indeed requires both alignment and uniformity.

For fixed τ>0\tau>0, as the number of negative samples M→∞M\rightarrow\infty, the (normalized) contrastive loss converges to

The first term is minimized iff ff is perfectly aligned.

If perfectly uniform encoders exist, they form the exact minimizers of the second term.

For the convergence in Equation (2), the absolute deviation from the limit decays in O(M−1/2)\mathcal{O}(M^{-1/2}).

The proof of Theorem 1 in the appendix connects the asymptotic Lcontrastive\mathcal{L}_{\mathsf{contrastive}} form with minimizing average pairwise Gaussian potential, i.e., minimizing Luniform\mathcal{L}_{\mathsf{uniform}}. Compared with the second term of Equation (2), Luniform\mathcal{L}_{\mathsf{uniform}} essentially pushes the log⁡\log outside the outer expectation, without changing the minimizer (perfectly uniform encoders). However, due to its pairwise nature, Luniform\mathcal{L}_{\mathsf{uniform}} is much simpler in form and avoids the computationally expensive softmax\mathtt{softmax} operation in Lcontrastive\mathcal{L}_{\mathsf{contrastive}} (Goodman, 2001; Bengio et al., ; Gutmann & Hyvärinen, 2010; Grave et al., 2017; Chen et al., 2018).

When pdatap_{\mathsf{data}} is uniform over finite samples {x1,x2,…,xN}\{x_{1},x_{2},\dots,x_{N}\} (e.g., a collected dataset), the second term in Equation (2) can be alternatively viewed as a resubstitution entropy estimator of f(x)f(x) (Ahmad & Lin, 1976), where xx follows the underlying distribution pnaturep_{\mathsf{nature}} that generates {xi}i=1N\{x_{i}\}_{i=1}^{N}, via a von Mises-Fisher (vMF) kernel density estimation (KDE):

p^vMF-KDE\hat{p}_{\mathsf{vMF\text{-}KDE}} is the KDE based on samples {f(xj)}j=1N\{f(x_{j})\}_{j=1}^{N} using a vMF kernel with κ=τ−1\kappa=\tau^{-1},

ZvMFZ_{\mathsf{vMF}} is the normalization constant for vMF distribution with κ=τ−1\kappa=\tau^{-1},

H^\hat{H} denotes the resubstitution entropy estimator,

I^\hat{I} denotes the mutual information estimator based on H^\hat{H}, since ff is a deterministic function.

Many empirical works are motivated by the InfoMax principle, i.e., maximizing I(f(x);f(y))I(f(x);f(y)) for (x,y)∼ppos(x,y)\sim p_{\mathsf{pos}}. However, the interpretation of Lcontrastive\mathcal{L}_{\mathsf{contrastive}} as a lower bound of I(f(x);f(y))I(f(x);f(y)) is known to be inconsistent with its actual behavior in practice (Tschannen et al., 2019). Our results instead analyze the properties of Lcontrastive\mathcal{L}_{\mathsf{contrastive}} itself. Considering the identity I(f(x);f(y))=H(f(x))−H(f(x)∣f(y))I(f(x);f(y))=H(f(x))-H(f(x)\mathrel{|}f(y)), we can see that while uniformity indeed favors large H(f(x))H(f(x)), alignment is stronger than merely desiring small H(f(x)∣f(y))H(f(x)\mathrel{|}f(y)). In particular, both Theorem 1 and the above connection with maximizing an entropy estimator provide alternative interpretations and motivations that Lcontrastive\mathcal{L}_{\mathsf{contrastive}} optimizes for aligned and information-preserving encoders.

Finally, even for the case where only a single negative sample is used (i.e., M=1M=1), we can still prove a weaker result, which we describe in details in the appendix.

Experiments

In this section, we empirically verify the hypothesis that alignment and uniformity are desired properties for representations. Recall that our two metrics are

We conduct extensive experiments with convolutional neural network (CNN) and recurrent neural network (RNN) based encoders on four popular representation learning benchmarks with distinct types of downstream tasks:

STL-10 (Coates et al., 2011) classification on AlexNet-based encoder outputs or intermediate activations with a linear or kk-nearest neighbor (kk-NN) classifier.

NYU-Depth-V2 (Nathan Silberman & Fergus, 2012) depth prediction on CNN encoder intermediate activations after convolution layers.

ImageNet and ImageNet-100 (random 100100-class subset of ImageNet) classification on CNN encoder penultimate layer activations with a linear classifier.

BookCorpus (Zhu et al., 2015) RNN sentence encoder outputs used for Moview Review Sentence Polarity (MR) (Pang & Lee, 2005) and Customer Product Review Sentiment (CR) (Wang & Manning, 2012) binary classification tasks with logisitc classifiers.

For image datasets, we follow the standard practice and choose positive pairs as two independent augmentations of the same image. For BookCorpus, positive pairs are chosen as neighboring sentences, following Quick-Thought Vectors (Logeswaran & Lee, 2018).

We perform majority of our analysis on STL-10 and NYU-Depth-V2 encoders, where we calculate Lcontrastive\mathcal{L}_{\mathsf{contrastive}} with negatives being other samples within the minibatch following the standard practice (Hjelm et al., 2018; Bachman et al., 2019; Tian et al., 2019; Chen et al., 2020a), and Luniform\mathcal{L}_{\mathsf{uniform}} as the logarithm of average pairwise feature potentials also within the minibatch. Due to their simple forms, these two losses can be implemented in PyTorch (Paszke et al., 2019) with less than 1010 lines of code, as shown in Figure 6.

To investigate alignment and uniformity properties on recent contrastive learning methods and larger datasets, we also analyze ImageNet and ImageNet-100 encoders trained with Momentum Contrast (MoCo) (He et al., 2019; Chen et al., 2020b), and BookCorpus encoders trained with Quick-Thought Vectors (Logeswaran & Lee, 2018), with these methods modified to also allow Lalign\mathcal{L}_{\mathsf{align}} and Luniform\mathcal{L}_{\mathsf{uniform}}.

We optimize a total of 304304 STL-10 encoders, 6464 NYU-Depth-V2 encoders, 4545 ImageNet-100 encoders, and 108108 BookCorpus encoders without supervision. The encoders are optimized w.r.t. weighted combinations of Lcontrastive\mathcal{L}_{\mathsf{contrastive}}, Lalign\mathcal{L}_{\mathsf{align}}, and/or Luniform\mathcal{L}_{\mathsf{uniform}}, with varying

(possibly zero) weights on the three losses,

temperature τ\tau for Lcontrastive\mathcal{L}_{\mathsf{contrastive}},

α∈{1,2}\alpha\in\{1,2\} for Lalign\mathcal{L}_{\mathsf{align}},

t∈{1,2,…,8}t\in\{1,2,\dots,8\} for Luniform\mathcal{L}_{\mathsf{uniform}},

batch size (affecting the number of (negative) pairs for Lcontrastive\mathcal{L}_{\mathsf{contrastive}} and Luniform\mathcal{L}_{\mathsf{uniform}}),

number of training epochs and learning rate,

initialization (from scratch vs. a pretrained encoder).

See the appendix for more experiment details and the exact configurations used.

For each encoder, we measure the downstream task performance, and the Lalign\mathcal{L}_{\mathsf{align}}, Luniform\mathcal{L}_{\mathsf{uniform}} metrics on the validation set. Figure 5 visualizes the trends between both metrics and representation quality. We observe that the two metrics strongly agrees the representation quality overall. In particular, the best performing encoders are exactly the ones with low Lalign\mathcal{L}_{\mathsf{align}} and Luniform\mathcal{L}_{\mathsf{uniform}}, i.e., the lower left corners in Figure 5.

As shown in Tables 1 and 2, encoders trained with only Lalign\mathcal{L}_{\mathsf{align}} and Luniform\mathcal{L}_{\mathsf{uniform}} consistently outperform their Lcontrastive\mathcal{L}_{\mathsf{contrastive}}-trained counterparts, for both tasks. Theoretically, Theorem 1 showed that Lcontrastive\mathcal{L}_{\mathsf{contrastive}} optimizes alignment and uniformity asymptotically with infinite negative samples. This empirical performance gap suggests that directly optimizing these properties can be superior in practice, when we can only have finite negatives.

Figure 7 shows how the final encoder changes in response to optimizing differently weighted combinations of Lalign\mathcal{L}_{\mathsf{align}} and Luniform\mathcal{L}_{\mathsf{uniform}} on STL-10. The trade-off between the Lalign\mathcal{L}_{\mathsf{align}} and Luniform\mathcal{L}_{\mathsf{uniform}} indicates that perfect alignment and perfect uniformity are likely hard to simultaneously achieve in practice. However, the inverted-U-shaped accuracy curve confirms that both properties are indeed necessary for a good encoder. When Lalign\mathcal{L}_{\mathsf{align}} is weighted much higher than Luniform\mathcal{L}_{\mathsf{uniform}}, degenerate solution occurs and all inputs are mapped to the same feature vector (exp⁡Luniform=1)\exp\mathcal{L}_{\mathsf{uniform}}=1). However, as long as the ratio between two weights is not too large (e.g., <4<4), we observe that the representation quality remains relatively good and insensitive to the exact weight choices.

We take an encoder trained with Lcontrastive\mathcal{L}_{\mathsf{contrastive}} using a suboptimal temperature τ=2.5\tau=2.5, and finetune it according to Lalign\mathcal{L}_{\mathsf{align}} and/or Luniform\mathcal{L}_{\mathsf{uniform}}. Figure 8 visualizes the finetuning trajectories. When only one of alignment and uniformity is optimized, the corresponding metric improves, but both the other metric and performance degrade. However, when both properties are optimized, the representation quality steadily increases. These trends confirm the causal effect of alignment and uniformity on the representation quality, and suggest that directly optimizing them can be a reasonable choice.

MoCo (He et al., 2019) and Quick-Thought Vectors (Logeswaran & Lee, 2018) are contrastive representation learning variants that have nontrivial differences with directly optimizing Lcontrastive\mathcal{L}_{\mathsf{contrastive}} in Equation (1). MoCo introduces a memory queue and a momentum encoder. Quick-Thought Vectors uses two different encoders to encode each sentence in a positive pair, only normalizes encoder outputs during evaluation, and does not use random sampling to obtain minibatches. After modifying them to also allow Lalign\mathcal{L}_{\mathsf{align}} and Luniform\mathcal{L}_{\mathsf{uniform}}, we train these methods on ImageNet-100 and BookCorpus, respectively. Figure 9 shows that Lalign\mathcal{L}_{\mathsf{align}} and Luniform\mathcal{L}_{\mathsf{uniform}} metrics are still correlated with the downstream task performances. Tables 3 and 4 show that directly optimizing them also leads to comparable or better representation quality. Table 5 also shows improvements on full ImageNet when we use Lalign\mathcal{L}_{\mathsf{align}} and Luniform\mathcal{L}_{\mathsf{uniform}} to train MoCo v2 (Chen et al., 2020b) (an improved version of MoCo). These results suggest that alignment and uniformity are indeed desirable properties for representations, for both image and text modalities, and are likely connected with general contrastive representation learning methods.

Discussion

Alignment and uniformity are often alluded to as motivations for representation learning methods (see Figure 1). However, a thorough understanding of these properties is lacking in the literature.

Are they in fact related to the representation learning methods? Do they actually agree with the representation quality (measured by downstream task performance)?

In this work, we have presented a detailed investigation on the relation between these properties and the popular paradigm of contrastive representation learning. Through theoretical analysis and extensive experiments, we are able to relate the contrastive loss with the alignment and uniformity properties, and confirm their strong connection with downstream task performances. Remarkably, we have revealed that directly optimizing our proposed metrics often leads to representations of better quality.

Below we summarize several suggestions for future work.

Acknowledgements

We thank Philip Bachman, Ching-Yao Chuang, Justin Solomon, Yonglong Tian, and Zhenyang Zhang for many helpful comments and suggestions. Tongzhou Wang was supported by the MIT EECS Merrill Lynch Graduate Fellowship. We thank Yangjun Ruan for helping us realize a minor issue with STL-10 scatter plot (Figure 5, now fixed).

Major Changelog

Added results on full ImageNet and MoCo v2.

Added discussions on the range of Luniform\mathcal{L}_{\mathsf{uniform}}.

Corrected Theorem 1’s convergence rate to O(M−1/2)\mathcal{O}(M^{-1/2}).

Removed from Figures 5 and LABEL:supp:tbl:stl10-big two STL-10 encoders that should not be included due to their usage of other regularizers (not shown). This does not affect the observed relation among Lalign\mathcal{L}_{\mathsf{align}}, Luniform\mathcal{L}_{\mathsf{uniform}}, and downstream performance. All other text and discussions stay unchanged.

References

Appendix A Proofs and Additional Theoretical Analysis

In this section, we present proofs for propositions and theorems in main paper Sections 4.1.2 and 4.2.

In Section A.2, we present a proof for Theorem 1. Theorem 1 describes the asymptotic behavior of Lcontrastive\mathcal{L}_{\mathsf{contrastive}} as the number of negative samples MM approaches infinity. The theorem is strongly related to empirical contrastive learning, given an error term (deviation from the limit) decaying in O(M−1/2)\mathcal{O}(M^{-1/2}) and that empirical practices often use a large number of negatives (e.g., M=65536M=65536 in He et al. (2019)) based on the observation that using more negatives consistently leads to better representation quality (Wu et al., 2018; Tian et al., 2019; He et al., 2019). Our proof further reveals connections between Lcontrastive\mathcal{L}_{\mathsf{contrastive}} and Luniform\mathcal{L}_{\mathsf{uniform}} which is defined via the Gaussian kernel.

Finally, also in Section A.2, we present a weaker result on the setting where only a single negative is used in Lcontrastive\mathcal{L}_{\mathsf{contrastive}} (i.e., M=1M=1).

To prove Proposition 1 and 2, we utilize the strict positive definiteness (Bochner, 1992; Stewart, 1976) of the Gaussian kernel GtG_{t}:

From there, we apply a known result about such kernels, from which the two propositions directly follow.

A symmetric and lower semi-continuous kernel KK on A×AA\times A (where AA is infinite and compact) is called strictly positive definite if for every finite signed Borel measure μ\mu supported on AA whose energy

is well defined, we have IK[μ]≥0I_{K}[\mu]\geq 0, where equality holds only if μ≡0\mu\equiv 0 on the σ\sigma-algebra of Borel subsets of AA.

Let M(Sd)\mathcal{M}(\mathcal{S}^{d}) be the set of Borel probability measures on Sd\mathcal{S}^{d}.

We are now in the place to apply the following two well-known results, which we present by restating Proposition 4.4.1, Theorem 6.2.1 and Corollary 6.2.2 of Borodachov et al. (2019) in weaker forms. We refer readers to Borodachov et al. (2019) for their proofs.

For t>0t>0, the Gaussian kernel Gt(u,v)≜e−t∥u−v∥22=e2t⋅uTv−2tG_{t}(u,v)\triangleq e^{-t\left\lVert u-v\right\rVert_{2}^{2}}=e^{2t\cdot u^{\mathsf{T}}v-2t} is strictly positive definite on Sd×Sd\mathcal{S}^{d}\times\mathcal{S}^{d}.

Consider kernel Kf ⁣:Sd×Sd→(−∞,+∞]K_{f}\colon\mathcal{S}^{d}\times\mathcal{S}^{d}\rightarrow(-\infty,+\infty] of the form,

If KfK_{f} is strictly positive definite on Sd×Sd\mathcal{S}^{d}\times\mathcal{S}^{d} and IKf[σd]I_{K_{f}}[\sigma_{d}] is finite, then σd\sigma_{d} is the unique measure (on Borel subsets of Sd\mathcal{S}^{d}) in the solution of min⁡μ∈M(Sd)IKf[μ]\min_{\mu\in\mathcal{M}(\mathcal{S}^{d})}I_{K_{f}}[\mu], and the normalized counting measures associated with any KfK_{f}-energy minimizing sequence of NN-point configurations on Sd\mathcal{S}^{d} converges weak∗ to σd\sigma_{d}.

In particular, this conclusion holds whenever ff has the property that −f′(t)-f^{\prime}(t) is strictly completely monotone on (0,4](0,4] and IKf[σd]I_{K_{f}}[\sigma_{d}] is finite.

σd\sigma_{d} is the unique solution (on Borel subsets of Sd\mathcal{S}^{d}) of

This is a direct consequence of Lemmas 1 and 2. ∎

For each N>0N>0, the NN point minimizer of the average pairwise potential is

The normalized counting measures associated with the {uN∗}N=1∞\{\mathbf{u}^{*}_{N}\}_{N=1}^{\infty} sequence converge weak∗ to σd\sigma_{d}.

This is a direct consequence of Lemmas 1 and 2. ∎

It’s not obvious what the optimal value of Luniform\mathcal{L}_{\mathsf{uniform}} is. In the following proposition, we characterize the exact range of the expected Gaussian potential and how it evolves as dimensionality increases. The situation for Luniform\mathcal{L}_{\mathsf{uniform}} directly follows as a corollary.

For t>0t>0, the expected pairwise Gaussian potential w.r.t. Borel probability measure μ∈M(Sd)\mu\in\mathcal{M}(\mathcal{S}^{d})

has range [e−2t\prescript0F1(;d+12;t2),1][e^{-2t}\prescript{}{0}{F}_{1}(;\frac{d+1}{2};t^{2}),1], where \prescript0F1\prescript{}{0}{F}_{1} is the confluent hypergeometric limit function defined as

where we have used the Pochhammer symbol (a)n={1\mboxifn=0a(a+1)(n+2)…(a+n−1)\mboxifn≥1.(a)_{n}=\begin{cases}1&\mbox{if }n=0\\ a(a+1)(n+2)\dots(a+n-1)&\mbox{if }n\geq 1.\end{cases}

The minimum e−2t\prescript0F1(;d+12;t2)e^{-2t}\prescript{}{0}{F}_{1}(;\frac{d+1}{2};t^{2}) is achieved iff μ=σd\mu=\sigma_{d} (on Borel subsets of Sd\mathcal{S}^{d}). Furthermore, this value strictly decreases as dd increases, converging to e−2te^{-2t} in the limit of d→∞d\rightarrow\infty.

The maximum is achieved iff μ\mu is a Dirac delta distribution, i.e., μ=δu\mu=\delta_{u} (on Borel subsets of Sd\mathcal{S}^{d}), for some u∈Sdu\in\mathcal{S}^{d}.

We know from Proposition 1 that σd\sigma_{d} uniquely achieves the minimum, given by the following integral ratio

The denominator, with some trigonometric identities, can be more straightforwardly evaluated as

where we have used the following identity based on the Poisson formula for Bessel functions and the relationship between \prescript0F1\prescript{}{0}{F}_{1} and Bessel functions:

where we have used the definition of \prescript0F1\prescript{}{0}{F}_{1} in Equation 5 to expand the formula.

Notice that each summand strictly decreases as d→∞d\rightarrow\infty. So must the total sum.

For the asymptotic behavior at d→∞d\rightarrow\infty, it only remains to show that

For the purpose of applying the Dominated Convergence Theorem (DCT) (on the counting measure). We consider the following summable series

with each term bounding the corresponding one in Equation 6:

Hence, the asymptotic lower range is e−2te^{-2t}.

Obviously, Dirac delta distributions δu\delta_{u}, u∈Sdu\in\mathcal{S}^{d} would achieve a maximum of 11. We will now show that all Borel probability measures μ\mu s.t. IGt[μ]=1I_{G_{t}}[\mu]=1 are delta distributions.

Suppose that such a μ\mu is not a Dirac delta distribution. Then, we can take distinct x,y∈supp⁡(μ)⊆Sdx,y\in\operatorname{supp}(\mu)\subseteq\mathcal{S}^{d}, and open neighborhoods around xx and vv, Nx,Ny∈SdN_{x},N_{y}\in\mathcal{S}^{d} such that they are small enough and disjoint:

Hence, only Dirac delta distributions attain the maximum.

Furthermore, the lower bound strictly decreases as the output dimension mm increases, attaining the following asymptotic value

If we ignore the log⁡\prescript0F1(;m2;t2)\log\prescript{}{0}{F}_{1}(;\frac{m}{2};t^{2}) term, informally, the optimal value of −2t-2t roughly says that any pair of feature vectors on Sd\mathcal{S}^{d} has distance about 2\sqrt{2}, i.e., are nearly orthogonal to each other. Indeed, vectors of high dimensions are usually nearly orthogonal, which is also consistent with the asymptotic result in Equation 7.

Figures 11 and 11 visualize how \prescript0F1\prescript{}{0}{F}_{1} and the optimal Luniform\mathcal{L}_{\mathsf{uniform}} (given by perfectly uniform encoders) evolve.

In practice, when Luniform\mathcal{L}_{\mathsf{uniform}} calculated using expectation over (a batch of) empirical samples {xi}i=1B\{x_{i}\}_{i=1}^{B}, B>1B>1, the range in Corollary 1 is indeed valid, since it bounds over all distributions:

However, often Luniform\mathcal{L}_{\mathsf{uniform}} is empirically estimated without considering distances between a vector and itself (e.g., in Figure 6 and in our experiment settings as described in Appendix B):

While both quantities converge to the correct value in the limit, the lower bound is not always true for this one, because it is not the expected pairwise Gaussian kernel based on some distribution. Note the following relation:

We can derive a valid lower bound using Equation 8: for \prescript0F1(;m2;t2)>e2tB\prescript{}{0}{F}_{1}(;\frac{m}{2};t^{2})>\frac{e^{2t}}{B},

Since this approaches fails for cases that \prescript0F1(;m2;t2)≤e2tB\prescript{}{0}{F}_{1}(;\frac{m}{2};t^{2})\leq\frac{e^{2t}}{B}, we can combine it with the naive lower bound −4t-4t, and have

By definition, Luniform\mathcal{L}_{\mathsf{uniform}} always non-positive. As shown above, different Luniform\mathcal{L}_{\mathsf{uniform}} empirical estimates may admit different lower bounds. However, in our experience, for reasonably large batch sizes, adding an offset of 2t2t often ensures a non-negative loss that is near zero at optimum. When output dimensionality mm is low, it might be useful to add an additional offset of −log⁡\prescript0F1(;m2;t2)-\log\prescript{}{0}{F}_{1}(;\frac{m}{2};t^{2}), which can be computed with the help of the SciPy package function scipy.special.hyp0f1(m/2, t**2) (Virtanen et al., 2020).

A.2 Proofs and Additional Results for Section 4.2

The following lemma directly follows Theorem 3.3 and Remarks 3.4 (b)(i) of Serfozo (1982). We refer readers to Serfozo (1982) for its proof.

Let AA be a compact second countable Hausdorff space. Suppose

{μn}n=1∞\{\mu_{n}\}_{n=1}^{\infty} is a sequence of finite and positive Borel measures supported on AA that converges weak∗ to some finite and positive Borel measure μ\mu (which is same as vague convergence since AA is compact);

{fn}n=1∞\{f_{n}\}_{n=1}^{\infty} is a sequence of Borel measurable functions that converges continuously to a Borel measurable ff;

{fn}n\{f_{n}\}_{n} are uniformly bounded over AA.

For fixed τ>0\tau>0, as the number of negative samples M→∞M\rightarrow\infty, the (normalized) contrastive loss converges to

The first term is minimized iff ff is perfectly aligned.

If perfectly uniform encoders exist, they form the exact minimizers of the second term.

For the convergence in Equation 2, the absolute deviation from the limit (i.e., the error term) decays in O(M−1/2)\mathcal{O}(M^{-1/2}).

We first show the convergence stated in Equation 2 along with its speed (result 3), and then the relations between the two limiting terms and the alignment and uniformity properties (results 1 and 2).

Proof of the convergence in Equation 2 and the O(M−1/2)\mathcal{O}(M^{-1/2}) decay rate of its error term (result 3).

by the strong law of large numbers (SLLN) and the Continuous Mapping Theorem.

where we justify the switching of expectation and limit by the convergence stated in Equation 10, the boundedness of euTv/τe^{u^{\mathsf{T}}v/\tau} (where u,v∈Sd,τ>0u,v\in\mathcal{S}^{d},\tau>0), and the Dominated Convergence Theorem (DCT).

where the first inequality follows the Intermediate Value Theorem and the e1/τe^{1/\tau} upper bound on the absolute derivative of log⁡\log between the two points, and the last equality follows the Berry-Esseen Theorem given the bounded support of ef(xi−)Tf(x)/τe^{f(x^{-}_{i})^{\mathsf{T}}f(x)/\tau} as following: for i.i.d. random variables YiY_{i} with bounded support ⊂[−a,a]\subset[-a,a], zero mean and σY2≤a2\sigma^{2}_{Y}\leq a^{2} variance, we have

where the constant CaC_{a} only depends on aa (which controls both the second and the third moment).

Proof of result 1: The first term is minimized iff ff is perfectly aligned.

Then the result follows directly the definition of perfect alignment, and the existence of perfectly aligned encoders (e.g., an encoder that maps every input to the same output vector).

Proof of result 2: If perfectly uniform encoders exist, they form the exact minimizers of the second term.

For simplicity, we define the following notation:

∀μ∈M(Sd)\forall\mu\in\mathcal{M}(\mathcal{S}^{d}), u∈Sdu\in\mathcal{S}^{d}, we define the continuous and Borel measurable function

with its range bounded in [e−1/τ,e1/τ][e^{-1/\tau},e^{1/\tau}].

Then the second term can be equivalently written as

where pdata∘f−1∈M(Sd)p_{\mathsf{data}}\circ f^{-1}\in\mathcal{M}(\mathcal{S}^{d}) is the probability measure of features, i.e., the pushforward measure of pdatap_{\mathsf{data}} via ff.

We now consider the following relaxed problem, where the minimization is taken over M(Sd)\mathcal{M}(\mathcal{S}^{d}), all possible Borel probability measures on the hypersphere Sd\mathcal{S}^{d}:

Our strategy is to show that the unique minimizer of Equation 13 is σd\sigma_{d}, from which the result 2 directly follows. The rest of the proof is structured in three parts.

We show that minimizers of Equation 13 exist, i.e., the above infimum is attained for some μ∈M(Sd)\mu\in\mathcal{M}(\mathcal{S}^{d}).

Let {μm}m=1∞\{\mu_{m}\}_{m=1}^{\infty} be a sequence in M(Sd)\mathcal{M}(\mathcal{S}^{d}) such that the infimum of Equation 13 is reached in the limit:

We want to show that μ∗\mu^{*} attains the limit (and thus the infimum), i.e.,

In view of Lemma 3, since Sd\mathcal{S}^{d} is a compact second countable Hausdorff space and {log⁡Uμn}n\{\log U_{\mu_{n}}\}_{n} is uniformly bounded over Sd\mathcal{S}^{d}, it remains to prove that {log⁡Uμn}n\{\log U_{\mu_{n}}\}_{n} is continuously convergent to log⁡Uμ∗\log U_{\mu^{*}}.

Let δn=xn−x\delta_{n}=x_{n}-x. By simply expanding UμnU_{\mu_{n}} and μμ∗\mu_{\mu^{*}}, we have

Since both the upper and the lower bound converge to Uμ∗(x)U_{\mu^{*}}(x) (by the weak ∗ convergence of {μn}n\{\mu_{n}\}_{n} to μ∗\mu^{*}), Uμn(xn)U_{\mu_{n}}(x_{n}) must as well. We have proved the continuous convergence of {log⁡Uμn}n\{\log U_{\mu_{n}}\}_{n} to log⁡Uμ∗\log U_{\mu^{*}}.

Therefore, the limit in Equation 14 holds. The infimum is thus attained at μ∗\mu^{*}:

We show that Uμ∗U_{\mu^{*}} is constant μ∗\mu^{*}-almost surely for any minimizer μ∗\mu^{*} of Equation 13.

Let μ∗\mu^{*} be any solution of Equation 13:

Consider the Borel sets where μ∗\mu^{*} has positive measure: T≜{T∈B(Sd) ⁣:μ∗(T)>0}\mathcal{T}\triangleq\{T\in\mathcal{B}(\mathcal{S}^{d})\colon\mu^{*}(T)>0\}. For any T∈TT\in\mathcal{T}, let μT∗\mu^{*}_{T} denote the conditional distribution of μ∗\mu^{*} on TT, i.e., ∀A∈B(Sd)\forall A\in\mathcal{B}(\mathcal{S}^{d}),

Note that for any such T∈TT\in\mathcal{T}, the mixture (1−α)μ∗+αμT∗(1-\alpha)\mu^{*}+\alpha\mu^{*}_{T} is a valid probability distribution (i.e., in M(Sd)\mathcal{M}(\mathcal{S}^{d})) for α∈(−μ∗(T),1)\alpha\in(-\mu^{*}(T),1), an open interval containing .

where the Leibniz rule along with the boundedness of Uμ∗U_{\mu^{*}} and UμTn∗U_{\mu^{*}_{T_{n}}} together justify the exchanges of integration and differentiation.

Let {Tn}n=1∞\{T_{n}\}_{n=1}^{\infty} be a sequence of sets in T\mathcal{T} such that

where the supremum must exist since Uμ∗U_{\mu^{*}} is bounded above.

Because Uμ∗U_{\mu^{*}} is a continuous and Borel measurable function, we have {u ⁣:Uμ∗(u)>U∗}∈B(Sd)\{u\colon U_{\mu^{*}}(u)>U^{*}\}\in\mathcal{B}(\mathcal{S}^{d}) and thus

otherwise {u ⁣:Uμ∗(u)>U∗}∈T\{u\colon U_{\mu^{*}}(u)>U^{*}\}\in\mathcal{T}, contradicting the definition of U∗U^{*} as the supremum.

Asymptotically, Uμ∗U_{\mu^{*}} is constant μTn∗\mu^{*}_{T_{n}}-almost surely:

where the inequality follows the boundedness of Uμ∗U_{\mu^{*}} and that μTn∗({u ⁣:Uμ∗(u)>U∗})=0\mu^{*}_{T_{n}}(\{u\colon U_{\mu^{*}}(u)>U^{*}\})=0.

Therefore, given the continuity of log⁡\log and the boundedness of Uμ∗U_{\mu^{*}}, we have

Equation 15 gives that ∀n=1,2,…\forall n=1,2,\dots,

where the inequality follows the boundedness of UμTn∗Uμ∗\frac{U_{\mu_{T_{n}}^{*}}}{U_{\mu^{*}}} and that μ∗({u ⁣:Uμ∗(u)>U∗})=0\mu^{*}(\{u\colon U_{\mu^{*}}(u)>U^{*}\})=0.

Taking the limit of n→∞n\rightarrow\infty on both sides, we have

where the last inequality holds because the supremum taken over T⊃{Sd}\mathcal{T}\supset\{\mathcal{S}^{d}\}.

Since 1=11=1, all inequalities must be equalities. In particular,

That is, for any solution μ∗\mu^{*} of Equation 13, Uμ∗U_{\mu^{*}} must be constant μ∗\mu^{*}-almost surely.

We show that σd\sigma_{d} is the unique minimizer of the relaxed problem in Equation 13.

Let S⊂M(Sd)S\subset\mathcal{M}(\mathcal{S}^{d}) be the set of measures where the above property holds:

The problem in Equation 13 is thus equivalent to minimizing over SS:

By Proposition 1 and τ>0\tau>0, we know that the uniform distribution σd\sigma_{d} is the unique solution to

Since σd∈S\sigma_{d}\in S, it must also be the unique solution to Equation 13.

Finally, if perfectly uniform encoders exist, σd\sigma_{d} is realizable, and they are the exact encoders that realize it. Hence, in such cases, they are the exact minimizers of

The first term of Equation 2 is equivalent with Lalign\mathcal{L}_{\mathsf{align}} when α=2\alpha=2, up to a constant and a scaling. In the above proof, we showed that the second term favors uniformity, via the feature distribution that minimizes the pairwise Gaussian kernel (see Equation 16):

which can be alternatively viewed as the relaxed problem of optimizing for the uniformity loss Luniform\mathcal{L}_{\mathsf{uniform}}:

The relaxation comes from the observation that Equation 17 minimizes over all feature distributions on Sd\mathcal{S}^{d}, while Equation 18 only considers the realizable ones.

In view of the Proposition 1 and the proof of Theorem 1, we know that the uniform distribution σd\sigma_{d} is the unique minimizer of both of the following problems:

So pushing the log⁡\log inside the outer integral doesn’t change the solution. However, if we push the log⁡\log all the way inside the inner integral, the problem becomes equivalent with minimizing the norm of the mean, i.e.,

If perfectly aligned and uniform encoders exist, they form the exact minimizers of the contrastive loss Lcontrastive(f;τ,M)\mathcal{L}_{\mathsf{contrastive}}(f;\tau,M) for fixed τ>0\tau>0 and M=1M=1.

By the definition of perfect alignment, the equality in Equation 19 is satisfied iff ff is perfectly aligned.

−f′(t)=12τe−t2τ1+e−t2τ=12τ(1−(1+e−t2τ)−1)-f^{\prime}(t)=\frac{1}{2\tau}\frac{e^{-\frac{t}{2\tau}}}{1+e^{-\frac{t}{2\tau}}}=\frac{1}{2\tau}(1-(1+e^{-\frac{t}{2\tau}})^{-1}) is strictly completely monotone on (0,+∞)(0,+\infty):

In view of Lemma 2, we have that the equality in Equation 20 is satisfied iff the feature distribution induced by ff (i.e., the pushforward measure pdata∘f−1p_{\mathsf{data}}\circ f^{-1}) is σd\sigma_{d}, that is, in other words, ff is perfectly uniform.

where equality is satisfied iff ff is perfectly aligned and uniform. This concludes the proof. ∎

We remark that the statement in Theorem 2 is weaker than the previous Theorem 1. Theorem 2 is conditioned on the existence perfectly aligned and uniform encoders. It only shows that Lcontrastive(f;τ,M=1)\mathcal{L}_{\mathsf{contrastive}}(f;\tau,M=1) favors alignment under the condition that perfect uniformity is realizable, and vice versa. In Theorem 1, Lcontrastive\mathcal{L}_{\mathsf{contrastive}} decomposes into two terms, each favoring alignment and uniformity. Therefore, the decomposition in Theorem 1 is exempof t from this constraint.

Appendix B Experiment Details

All experiments are performed on 1-4 NVIDIA Titan Xp, Titan X PASCAL, Titan RTX, or 2080 Ti GPUs.

For CIFAR-10, STL-10 and NYU-Depth-V2 experiments, we use the following settings, unless otherwise stated in Tables LABEL:supp:tbl:stl10-big and LABEL:supp:tbl:nyudepth-big below:

Standard data augmentation procedures are used for generating positive pairs, including resizing, cropping, horizontal flipping, color jittering, and random grayscale conversion. This follows prior empirical work in contrastive representation learning (Wu et al., 2018; Tian et al., 2019; Hjelm et al., 2018; Bachman et al., 2019).

Neural network architectures follow the corresponding experiments on these datasets in Tian et al. (2019). For NYU-Depth-V2 evaluation, the architecture of the depth prediction CNN is described in Table 6.

We use minibatch stochastic gradient descent (SGD) with 0.90.9 momentum and 0.00010.0001 weight decay.

We use linearly scaled learning rate (0.120.12 per 256256 batch size) (Goyal et al., 2017).

CIFAR-10 and STL-10: Optimization is done over 200200 epochs, with learning rate decayed by a factor of 0.10.1 at epochs 155155, 170170, and 185185.

NYU-Depth-V2: Optimization is done over 400400 epochs, with learning rate decayed by a factor of 0.10.1 at epochs 310310, 340340, and 370370.

Encoders are optimized over the training split. For evaluation, we freeze the encoder, and train classifiers / depth predictors on the training set samples, and test on the validation split.

CIFAR-10 and STL-10: We use standard train-val split. Linear classifiers are trained with Adam (Kingma & Ba, 2014) over 100100 epochs, with β1=0.5,β2=0.999,ϵ=10−8\beta_{1}=0.5,\beta_{2}=0.999,\epsilon=10^{-8}, 128128 batch size, and an initial learning rate of 0.0010.001, decayed by a factor of 0.20.2 at epochs 6060 and 8080.

NYU-Depth-V2: We use the train-val split on the 14491449 labeled images from Nathan Silberman & Fergus (2012). Depth predictors are trained with Adam (Kingma & Ba, 2014) over 120120 epochs, with β1=0.5,β2=0.999,ϵ=10−8\beta_{1}=0.5,\beta_{2}=0.999,\epsilon=10^{-8}, 128128 batch size, and an initial learning rate of 0.0030.003, decayed by a factor of 0.20.2 at epochs 7070, 9090, 100100, and 110110.

At each SGD iteration, a minibatch of KK positive pairs is sampled {(xi,yi)}i=1K\{(x_{i},y_{i})\}_{i=1}^{K}, and the three losses for this minibatch are calculated as following:

Lcontrastive\mathcal{L}_{\mathsf{contrastive}}: For each xix_{i}, the sample contrastive loss is taken with the positive being yiy_{i}, and the negatives being {yj}j≠i\{y_{j}\}_{j\neq i}. For each yiy_{i}, the sample loss is computed similarly. The minibatch loss is calculated by aggregating these 2K2K terms:

This calculation follows empirical practices and is similar to Oord et al. (2018); Hénaff et al. (2019), and end-to-end in He et al. (2019).

Lalign\mathcal{L}_{\mathsf{align}}: The minibatch alignment loss is straightforwardly computed as

Luniform\mathcal{L}_{\mathsf{uniform}}: The minibatch uniform loss is calculated by considering each pair of {xi}i\{x_{i}\}_{i} and {yi}i\{y_{i}\}_{i}:

Tables LABEL:supp:tbl:stl10-big and LABEL:supp:tbl:nyudepth-big below describe the full specifications of all 304304 STL-10 and 6464 NYU-Depth-V2 encoders. These experiment results are visualized in main paper Figure 5, showing a clear connection between representation quality and Lalign\mathcal{L}_{\mathsf{align}} & Luniform\mathcal{L}_{\mathsf{uniform}} metrics.

B.2 ImageNet and ImageNet-100 with Momentum Contrast (MoCo) Variants

{f(xi)i}i=1K\{f(x_{i})_{i}\}_{i=1}^{K} be the batched query features encoded by the current up-to-date encoder ff (i.e., q\mathtt{q} in Algorithm 1 of He et al. (2019)),

{fEMA(yi)}i=1K\{f_{\textsf{EMA}}(y_{i})\}_{i=1}^{K} be the batched key features encoded by the exponential moving average encoder fEMAf_{\textsf{EMA}} (i.e., k\mathtt{k} in Algorithm 1 of He et al. (2019)),

{queuej}j=1N\{\mathtt{queue}_{j}\}_{j=1}^{N} be the feature queue, where NN is the queue size.

Lalign\mathcal{L}_{\mathsf{align}} and Luniform\mathcal{L}_{\mathsf{uniform}} for this minibatch are calculated as following:

Lalign\mathcal{L}_{\mathsf{align}}: The minibatch alignment loss is computed as disparity between features from the two encoders:

Luniform\mathcal{L}_{\mathsf{uniform}}: We experiment with two forms of Luniform\mathcal{L}_{\mathsf{uniform}}:

Only computing pairwise distance between {f(xi)}i\{f(x_{i})\}_{i} and {queuej}j\{\mathtt{queue}_{j}\}_{j}:

Also computing pairwise distance inside {f(xi)}i\{f(x_{i})\}_{i}:

B.2.1 ImageNet-100 with MoCo

We use the same ImageNet-100 sampled by Tian et al. (2019), containing the 100100 randomly selected classes listed in Table 7.

Our MoCo experiment settings below mostly follow He et al. (2019) and the unofficial implementation by Tian (2019), because the official implementation was not released at the time of performing these analyses:

Standard data augmentation procedures are used for generating positive pairs, including resizing, cropping, horizontal flipping, color jittering, and random grayscale conversion, following Tian (2019).

Encoder architecture is ResNet50 (He et al., 2016).

We use minibatch stochastic gradient descent (SGD) with 128128 batch size, 0.030.03 initial learning rate, 0.90.9 momentum and 0.00010.0001 weight decay.

Optimization is done over 240240 epochs, with learning rate decayed by a factor of 0.10.1 at epochs 120120, 160160, and 200200.

We use 0.9990.999 exponential moving average factor, following He et al. (2019).

For evaluation, we freeze the encoder, and train a linear classifier on the training set samples, and test on the validation split. Linear classifiers are trained with minibatch SGD over 6060 epochs, with 256256 batch size, and an initial learning rate of 1010, decayed by a factor of 0.20.2 at epochs 3030, 4040, and 5050.

LABEL:supp:tbl:imagenet100-big below describes the full specifications of all 4545 ImageNet-100 encoders. These experiment results are visualized in main paper Figure 9(a), showing a clear connection between representation quality and Lalign\mathcal{L}_{\mathsf{align}} & Luniform\mathcal{L}_{\mathsf{uniform}} metrics.

B.2.2 ImageNet with MoCo v2

Our MoCo v2 experiment settings directly follow Chen et al. (2020b) and the official implementation (Chen et al., 2020c):

Standard data augmentation procedures are used for generating positive pairs, including resizing, cropping, horizontal flipping, color jittering, random grayscale conversion, and random Gaussian blurring, following Chen et al. (2020c).

Encoder architecture is ResNet50 (He et al., 2016).

We use minibatch stochastic gradient descent (SGD) with 256256 batch size, 0.030.03 initial learning rate, 0.90.9 momentum and 0.00010.0001 weight decay.

Optimization is done over 200200 epochs, with learning rate decayed by a factor of 0.10.1 at epochs 120120 and 160160.

We use 0.9990.999 exponential moving average factor, 6553665536 queue size, 128128 feature dimensions.

For evaluation, we freeze the encoder, and train a linear classifier on the training set samples, and test on the validation split. Linear classifiers are trained with minibatch SGD over 100100 epochs, with 256256 batch size, and an initial learning rate of 3030, decayed by a factor of 0.10.1 at epochs 6060 and 8080.

Unlike the MoCo experiments on ImageNet-100, which were based on unofficial implementations for reasons stated in Sec. B.2.1, the MoCo v2 experiments on full ImageNet were based on the official implementation by Chen et al. (2020c). We provide a reference implementation that can fully reproduce the results in Table 5 at https://github.com/SsnL/moco_align_uniform, where we also provide a model checkpoint (trained using Lalign\mathcal{L}_{\mathsf{align}} and Luniform\mathcal{L}_{\mathsf{uniform}}) of 67.694%67.694\% validation top1 accuracy.

B.3 BookCorpus with Quick-Thought Vectors Variants

Since the original BookCorpus dataset (Zhu et al., 2015) is not distributed anymore, we use the unofficial code by Kobayashi (2019) to recreate our copy. Our copy ended up containing 52,799,51352{,}799{,}513 training sentences and 50,00050{,}000 validation sentences, compared to the original copy used by Quick-Thought Vectors (Logeswaran & Lee, 2018), which contains 45,786,40045{,}786{,}400 training sentences and 50,00050{,}000 validation sentences.

With Quick-Thought Vectors, the positive pairs are the neighboring sentences. At each optimization iteration, let

{xi}i=1K\{x_{i}\}_{i=1}^{K} be the KK consecutive sentences forming this minibatch, where KK be the minibatch size,

ff and gg be the two RNN sentence encoders.

The original Quick-Thought Vectors (Logeswaran & Lee, 2018) does not l2l2-normalize on encoder outputs during training the encoder. Here we describe the calculation of Lcontrastive\mathcal{L}_{\mathsf{contrastive}}, Lalign\mathcal{L}_{\mathsf{align}}, and Luniform\mathcal{L}_{\mathsf{uniform}} for l2l2-normalized encoders, in our modified Quick-Thought Vectors method. Note that this does not affect evaluation since features are l2l2-normalized before using in downstream tasks, following the original Quick-Thought Vectors (Logeswaran & Lee, 2018). For a minibatch, these losses are calculated as following:

Lcontrastive\mathcal{L}_{\mathsf{contrastive}} with temperature:

This is almost identical with the original contrastive loss used by Quick-Thought Vectors, except that this does not additionally manually masks out the entries f(xi)Tg(xi)f(x_{i})^{\mathsf{T}}g(x_{i}) with zeros, which is unnecessary with l2l2-normalization.

Lalign\mathcal{L}_{\mathsf{align}}: The minibatch alignment loss is computed as disparity between features from the two encoders encoding neighboring sentences (assuming K>=2K>=2):

Luniform\mathcal{L}_{\mathsf{uniform}}: We combine the uniformity losses for each of ff and gg by summing them (instead of averaging since ff and gg are two different encoders):

Our experiment settings below mostly follow the official implementation by Logeswaran & Lee (2018):

Sentence encoder architecture is bi-directional Gated Recurrent Unit (GRU) (Cho et al., 2014) with inputs from a 620620-dimensional word embedding trained jointly from scratch.

We use Adam (Kingma & Ba, 2014) with β1=0.9,β2=0.999,ϵ=10−8\beta_{1}=0.9,\beta_{2}=0.999,\epsilon=10^{-8}, 400400 batch size, 0.00050.0005 constant learning rate, and 0.50.5 gradient norm clipping.

Optimization is done during 11 epoch over the training data.

For evaluation on a binary classification task, we freeze the encoder, and fit a logistic classifier with l2l2 regularization on the encoder outputs. A 1010-fold cross validation is performed to determine the regularization strength among {1,2−1,…,2−8}\{1,2^{-1},\dots,2^{-8}\}, following Kiros et al. (2015) and Logeswaran & Lee (2018). The classifier is finally tested on the validation split.

LABEL:supp:tbl:bookcorpus-big below describes the full specifications of all 108108 BookCorpus encoders along with 66 settings that lead to training instability (i.e., NaN\mathtt{NaN} occurring). These experiment results are visualized in main paper Figure 9(b), showing a clear connection between representation quality and Lalign\mathcal{L}_{\mathsf{align}} & Luniform\mathcal{L}_{\mathsf{uniform}} metrics. For the unnormalized encoders, the features are normalized before calculated Lalign\mathcal{L}_{\mathsf{align}} and Luniform\mathcal{L}_{\mathsf{uniform}} metrics, since they are nonetheless still normalized before being used in downstream tasks (Logeswaran & Lee, 2018).