S4L: Self-Supervised Semi-Supervised Learning

Xiaohua Zhai, Avital Oliver, Alexander Kolesnikov, Lucas Beyer

Introduction

Modern computer vision systems demonstrate outstanding performance on a variety of challenging computer vision benchmarks, such as image recognition , object detection , semantic image segmentation , etc. Their success relies on the availability of a large amount of annotated data that is time-consuming and expensive to acquire. Moreover, applicability of such systems is typically limited in scope defined by the dataset they were trained on.

Many real-world computer vision applications are concerned with visual categories that are not present in standard benchmark datasets, or with applications of dynamic nature where visual categories or their appearance may change over time. Unfortunately, building large labeled datasets for all these scenarios is not practically feasible. Therefore, it is an important research challenge to design a learning approach that can successfully learn to recognize new concepts by leveraging only a small amount of labeled examples. The fact that humans quickly understand new concepts after seeing only a few (labeled) examples suggests that this goal is achievable in principle.

Notably, a large research effort is dedicated towards learning from unlabeled data that, in many realistic applications, is much less onerous to acquire than labeled data. Within this effort, the field of self-supervised visual representation learning has recently demonstrated the most promising results . Self-supervised learning techniques define pretext tasks which can be formulated using only unlabeled data, but do require higher-level semantic understanding in order to be solved. As a result, models trained for solving these pretext tasks learn representations that can be used for solving other downstream tasks of interest, such as image recognition.

Despite demonstrating encouraging results , purely self-supervised techniques learn visual representations that are significantly inferior to those delivered by fully-supervised techniques. Thus, their practical applicability is limited and as of yet, self-supervision alone is insufficient.

We further experimentally investigate whether our S4LS^{4}L methods could further benefit from regularizations proposed by the semi-supervised literature, and discover that they are complementary, i.e. combining them leads to improved results.

Our main contributions can be summarized as follows:

We propose a new family of techniques for semi-supervised learning with natural images that leverage recent advances in self-supervised representation learning.

We demonstrate that the proposed self-supervised semi-supervised (S4LS^{4}L) techniques outperform carefully tuned baselines that are trained with no unlabeled data, and achieve performance competitive with previously proposed semi-supervised learning techniques.

We further demonstrate that by combining our best S4LS^{4}L methods with existing semi-supervised techniques, we achieve new state-of-the-art performance on the semi-supervised ILSVRC-2012 benchmark.

Related Work

In this work we build on top of the current state-of-the-art in both fields of semi-supervised and self-supervised learning. Therefore, in this section we review the most relevant developments in these fields.

Semi-supervised learning describes a class of algorithms that seek to learn from both unlabeled and labeled samples, typically assumed to be sampled from the same or similar distributions. Approaches differ on what information to gain from the structure of the unlabeled data.

Given the wide variety of semi-supervised learning techniques proposed in the literature, we refer to for an extensive survey. For more context, we focus on recent developments based on deep neural networks.

Many of the initial results on semi-supervised learning with deep neural networks were based on generative models such as denoising autoencoders , variational autoencoders and generative adversarial networks . More recently, a line of research showed improved results on standard baselines by adding consistency regularization losses computed on unlabeled data. These consistency regularization losses measure discrepancy between predictions made on perturbed unlabeled data points. Additional improvements have been shown by smoothing predictions before measuring these perturbations. Approaches of these kind include Π\Pi-Model , Temporal Ensembling , Mean Teacher and Virtual Adversarial Training . Recently, fast-SWA showed improved results by training with cyclic learning rates and measuring discrepancy with an ensemble of predictions from multiple checkpoints. By minimizing consistency losses, these models implicitly push the decision boundary away from high-density parts of the unlabeled data. This may explain their success on typical image classification datasets, where points in each clusters typically share the same class.

Two additional important approaches for semi-supervised learning, which have shown success both in the context of deep neural networks and other types of models are Pseudo-Labeling , where one imputes approximate classes on unlabeled data by making predictions from a model trained only on labeled data, and conditional entropy minimization , where all unlabeled examples are encouraged to make confident predictions on some class.

2 Self-supervised Learning

Self-supervised learning is a general learning framework that relies on surrogate (pretext) tasks that can be formulated using only unsupervised data. A pretext task is designed in a way that solving it requires learning of a useful image representation. Self-supervised techniques have a variety of applications in a broad range of computer vision topics .

In this paper we employ self-supervised learning techniques that are designed to learn useful visual representations from image databases. These techniques achieve state-of-the-art performance among approaches that learn visual representations from unsupervised images only. Below we provided a non-comprehensive summary of the most important developments in this direction.

Doersch et al. propose to train a CNN model that predicts relative location of two randomly sampled non-overlapping image patches . Follow-up papers generalize this idea for predicting a permutation of multiple randomly sampled and permuted patches.

Beside the above patch-based methods, there are self-supervised techniques that employ image-level losses. Among those, in the authors propose to use grayscale image colorization as a pretext task. Another example is a pretext task that predicts an angle of the rotation transformation that was applied to an input image.

Some techniques go beyond solving surrogate classification tasks and enforce constraints on the representation space. A prominent example is the exemplar loss from that encourages the model to learn a representation that is invariant to heavy image augmentations. Another example is , that enforces additivity constraint on visual representation: the sum of representations of all image patches should be close to representation of the whole image. Finally, proposes a learning procedure that alternates between clustering images in the representation space and learning a model that assigns images to their clusters.

Methods

In this section we present our self-supervised semi-supervised learning (S4LS^{4}L) techniques. We first provide a general description of our approach. Afterwards, we introduce specific instantiations of our approach.

We focus on the semi-supervised image classification problem. Formally, we assume an (unknown) data generating joint distribution p(X,Y)p(X,Y) over images and labels. The learning algorithm has access to a labeled training set DlD_{l}, which is sampled i.i.d. from p(X,Y)p(X,Y) and an unlabeled training set DuD_{u}, which is sampled i.i.d. from the marginal distribution p(X)p(X).

The semi-supervised methods we consider in this paper have a learning objective of the following form:

where Ll\mathcal{L}_{l} is a standard cross-entropy classification loss of all labeled images in the dataset, Lu\mathcal{L}_{\textrm{u}} is a loss defined on unsupervised images (we discuss its particular instances later in this section), ww is a non-negative scalar weight and θ\theta is the parameters for model fθ(⋅)f_{\theta}(\cdot). Note that the learning objective can be extended to multiple unsupervised losses.

We now describe our self-supervised semi-supervised learning techniques. For simplicity, we present our approach in the context of multiclass image recognition, even though it can be easily generalized to other scenarios, such as dense image segmentation.

It is important to note that in practice the objective 1 is optimized using a stochastic gradient descent (or a variant) that uses mini-batches of data to update the parameters θ\theta. In this case the size of a supervised mini-batch xl,yl⊂Dlx_{l},y_{l}\subset D_{l} and an unsupervised mini-batch xu⊂Dux_{u}\subset D_{u} can be arbitrary chosen. In our experiments we always default to simplest possible option of using minibatches of equal sizes.

We also note that we can choose whether to include the minibatch xlx_{l} into the self-supervised loss, i.e. apply Lself\mathcal{L}_{\textrm{self}} to the union of xux_{u} and xlx_{l}. We experimentally study the effect of this choice in our experiments Section 4.4.

We demonstrate our framework on two prominent self-supervised techniques: predicting image rotation and exemplar . Note, that with our framework, more self-supervised losses can be explored in the future.

S4LS^{4}L-Rotation. The key idea of rotation self-supervision is to rotate an input image then predict which rotation degree was applied to these rotated images. The loss is defined as:

We also apply the self-supervised loss to the labeled images in each minibatch. Since we process rotated supervised images in this case, we suggest to also apply a classification loss to these images. This can be seen as an additional way to regularize a model in a regime when a small amount of labeled image are available. We measure the effect of this choice later in Section 4.4.

S4LS^{4}L-Exemplar. The idea of exemplar self-supervision is to learn a visual representation that is invariant to a wide range of image transformations. Specifically, we use “Inception” cropping , random horizontal mirroring, and HSV-space color randomization as in to produce 8 different instances of each image in a minibatch. Following , we implement Lu\mathcal{L}_{u} as the batch hard triplet loss with a soft margin. This encourages transformation of the same image to have similar representations and, conversely, encourages transformations of different images to have diverse representations.

Similarly to the rotation self-supervision case, Lu\mathcal{L}_{u} is applied to all eight instances of each image.

2 Semi-supervised Baselines

In the following section, we compare S4LS^{4}L to several leading semi-supervised learning algorithms that are not based on self-supervised objectives. We now describe the approaches that we compare to.

Our proposed objective 1 is applicable for semi supervised learning methods as well, where the loss Lu\mathcal{L}_{u} is the standard semi supervised loss as described below.

Virtual Adversarial Training (VAT) : The idea is making the predicted labels robust around input data point against local perturbation. It approximates the maximal change in predictions within an ϵvat\epsilon_{\text{vat}} vicinity of unlabeled data points, where ϵvat\epsilon_{\text{vat}} is a hyperparameter. Concretely, the VAT loss for a model fθf_{\theta} is:

While computing Δx\Delta x directly is not tractable, it can be efficiently approximated at the cost of an extra forward and backwards pass for every optimization step. .

Conditional Entropy Minimization (EntMin) : This works under the assumption that unlabeled data indeed has one of the classes that we are training on, even when the particular class is not known during training. It adds a loss for unlabeled data that, when minimized, encourages the model to make confident predictions on unlabeled data. Specifically, the conditional entropy minimization loss for a model fθf_{\theta} (treating fθf_{\theta} as a conditional distribution of labels over images) is:

Alone, the EntMin loss is not useful in the context of deep neural networks because the model can easily become extremely confident by increasing the weights of the last layer. One way to resolve this is to encourage the model predictions to be locally-Lipschitz, which VAT does. Therefore, we only consider VAT and EntMin combined, not just EntMin alone, i.e. Lu=wvatLvat+wentminLentmin\mathcal{L}_{u}=w_{vat}\mathcal{L}_{\textrm{vat}}+w_{\textrm{entmin}}\mathcal{L}_{\textrm{entmin}}.

Pseudo-Label is a simple approach: Train a model only on labeled data, then make predictions on unlabeled data. Then enlarge your training set with the predicted classes of the unlabeled data points whose predictions are confident past some threshold of confidence. Re-train your model with this enlarged labeled dataset. While shows that in a simple ”two moons” dataset, psuedo-label fails to learn a good model, in many real datasets this approach does show meaningful gains.

ILSVRC-2012 Experiments and Results

In this section, we present the results of our main experiments. We used the ILSVRC-2012 dataset due to its widespread use in self-supervised learning methods, and to see how well semi-supervised methods scale.

Since the test set of ILSVRC-2012 is not available, and numbers from the validation set are usually reported in the literature, we performed all hyperparameter selection for all models that we trained on a custom train/validation split of the public training set. This custom split contains 1 231 1211\,231\,121 training and 50 04650\,046 validation images. We then retrain the model using the best hyperparameters on the full training set (1 281 1671\,281\,167 images), possibly with fewer labels, and report final results obtained on the public validation set (50 00050\,000 images).

We always define epochs in terms of the available labeled data, i.e. one epoch corresponds to one full pass through the labeled data, regardless of how many unlabeled examples have been seen. We optimize our models using stochastic gradient descent with momentum on minibatches of size 256256 unless specified otherwise. While we do tune the learning rate, we keep the momentum fixed at 0.90.9 across all experiments. Table 1 summarizes our main results.

Whenever new methods are introduced, it is crucial to compare them against a solid baseline of existing methods. The simplest baseline to which any semi-supervised learning method should be compared to, is training a plain supervised model on the available labeled data.

Oliver et al. discovered that reported baselines trained on labeled examples alone are unfairly weak, perhaps given that there is not a strong community behind tuning those baselines. They provide strong supervised-only baselines for SVHN and CIFAR-10, and show that the gap shown by the use of unlabeled data is smaller than reported.

In total, we trained several thousand models on our custom training/validation split of the public training set of ILSVRC-2012. In summary, it is crucial to tune both weight decay and training duration while, perhaps surprisingly, model architecture, depth, and width only have a small influence on the final results. We thus use a standard, unmodified ResNet50v2 as model, trained with weight decay of 10−310^{-3} for 200200 epochs, using a standard learning rate of 0.10.1, ramped up from for the first five epochs, and decayed by a factor of 1010 at epochs 140140, 160160, and 180180. We train in total for 200200 epochs. The standard augmentation procedure of random cropping and horizontal flipping is used during training, and predictions are made using a single central crop keeping aspect ratio.

2 Semi-supervised Baselines

We train semi-supervised baseline models using (1) Pseudo-Label, (2) VAT, and (3) VAT+EntMin. To the best of our knowledge, we present the first evaluation of these techniques on ILSVRC-2012.

Pseudo-Label Using the plain supervised learning models from Section 4.1, we assign pseudo-labels to the full dataset. Then, in a second step, we train a ResNet50v2 from scratch following standard practice, i.e. with learning rate 0.10.1, weight decay 10−410^{-4}, and 100100 epochs on the full (pseudo-labeled) dataset.

We try both using all predictions as pseudo-labels, as well as using only those predictions with a confidence above 0.50.5. Both perform closely on our validation set, and we choose no filtering for the final model for simplicity.

Besides the previously mentioned hyperparameters common to all methods, VAT needs tuning ϵvat\epsilon_{\text{vat}}. Since it corresponds to a distance in pixel space, we use a simple heuristic for defining a range of values to try for ϵvat\epsilon_{\text{vat}}: values should be lower than half the distance between neighbouring images in the dataset. Based on this heuristic, we try values of ϵvat∈{50,50⋅2−1/3,50⋅2−2/3,25}\epsilon_{\text{vat}}\in\{50,50\cdot 2^{-1/3},50\cdot 2^{-2/3},25\} and found ϵvat≈40\epsilon_{\text{vat}}\approx 40 to work best.

VAT+EntMin VAT is intended to be used together with an additional entropy minimization (EntMin) loss. EntMin adds a single hyperparameter to our best VAT model: the weight of the EntMin loss, for which we try wentmin∈{0,0.03,0.1,0.3,1}w_{\text{entmin}}\in\{0,0.03,0.1,0.3,1\}, without re-tuning ϵvat\epsilon_{\text{vat}}.

3 Self-supervised Baselines

Previous work has evaluated features learned via self-supervision on the unlabeled data in a “semi-supervised” way by either freezing the features and learning a linear classifier on top, or by using the self-supervised model as an initialization and fine-tuning, using a subset of the labels in both cases. In order to compare our proposed way to do self-supervised semi-supervised learning to these common evaluations, we train a rotation and an exemplar model following the best practice from but with standard width (“4×4\times” in ).

Following our established protocol, we tune the weight decay and learning rate for the logistic regression, although interestingly the standard values from of 10−410^{-4} weight decay and 0.10.1 learning rate worked best.

For training our full self-supervised semi-supervised models (S4LS^{4}L), we follow the same protocol as for our semi-supervised baselines, i.e. we use the best settings of the plain supervised baseline and only tune the learning rate, weight decay, and weight of the newly introduced loss. We found that for both S4LS^{4}L-Rotation and S4LS^{4}L-Exemplar, the self-supervised loss weight w=1w=1 worked best (though not by much) and the optimal weight decay and learning rate were the same as for the supervised baseline.

The results shown in Table 1 show that our proposed way of doing self-supervised semi-supervised learning is indeed effective for the two self-supervision methods we tried. We hypothesize that such approaches can be designed for other self-supervision objectives.

We additionally verified that our proposed method is not sensitive to the random seed, nor the split of the dataset, see Appendix B for details.

Since we found that different types of models perform similarly well, the natural next question is whether they are complementary, in which case a combination would lead to an even better model, or whether they all reach a common “intrinsic” performance plateau.

In this section, we thus describe our Mix Of All Models (MOAM). In short: in a first step, we combine S4LS^{4}L-Rotation and VAT+EntMin to learn a 4×4\times wider model. We then use this model in order to generate pseudo labels for a second training step, followed by a final fine-tuning step. Results of the final model, as well as the models obtained in the two intermediate steps, are reported in Table 2 along with previous results reported in the literature.

Transfer of Learned Representations

Self-supervision methods are typically evaluated in terms of how generally useful their learned representation is. This is done by treating the learned model as a fixed feature extractor, and training a linear logistic regression model on top the features it extracts on a different dataset, usually Places205 . We perform such an evaluation on our S4LS^{4}L models in order to gain some insight into the generality of the learned features, and how they compare to those obtained by pure self-supervision.

We closely follow the protocol defined by . The representation is extracted from the pre-logits layer. We use stochastic gradient descent (SGD) with momentum for training these linear evaluation models with a minibatch size of 2048 and an initial learning rate of 0.1, warmed up in the first epoch.

Is a Tiny Validation Set Enough?

Current standard practice in semi-supervised learning is to use a subset of the labels for training on a large dataset, but still perform model selection using scores obtained on the full validation set.To make matters worse, in the case of ILSVRC-2012, this validation set is used both to select hyperparameters as well as to report final performance. Remember that we avoid this by creating a custom validation set from part of the training set for all hyperparameter selections. But having a large labeled validation set at hand is at odds with the promised practicality of semi-supervised learning, which is all about having only few labeled examples. This fact has been acknowledged by , but has been mostly ignored in the semi-supervised literature. Oliver et al. questions the viability of tuning with small validation sets by comparing the estimated model accuracy on small validation sets. They find that the variance of the estimated accuracy gap between two models can be larger than the actual gap between those models, hinting that model selection with small validation sets may not be viable. That said, they did not empirically evaluate whether it’s possible to find the best model with a small validation set, especially when choosing hyperparameters for a particular semi-supervised method.

Discussion and Future Work

In this paper, we have bridged the gap between self-supervision methods and semi-supervised learning by suggesting a framework (S4LS^{4}L) which can be used to turn any self-supervision method into a semi-supervised learning algorithm.

We instantiated two such methods: S4LS^{4}L-Rotation and S4LS^{4}L-Exemplar and have shown that they perform competitively to methods from the semi-supervised literature on the challenging ILSVRC-2012 dataset. We further showed that S4LS^{4}L methods are complementary to existing semi-supervision techniques, and MOAM, our proposed combination of those, leads to state-of-the-art performance.

Nevertheless, we hope that this work inspires other researchers in the field of self-supervision to consider extending their methods into semi-supervised methods using our S4LS^{4}L framework, as well as researchers in the field of semi-supervised learning to take inspiration from the vast amount of recently proposed self-supervision methods.

Acknowledgements. We thank the Google Brain Team in Zürich, and especially Sylvain Gelly for discussions.

References

Appendix A Detailed Results of the Supervised Baselines

We present the results in the form of what we call “hypersweep curves” in Figures 4 and 5.

Each plot shows a large collection of models – each point on each plot is a fully trained model. The curves are sorted by accuracy, allowing testing sensitivity to different hyperparameters, not only comparing the best model.

For each curve, we plot the accuracy of models where one of the hyperparameters is fixed.

Which value of a hyperparameter performs best by looking at which curve’s rightmost point is highest.

How sensitive the model is to a hyperparameter in the best case by looking at how far apart the curves are from eachother at their rightmost point.

How robust a hyperparameter is on average by looking at how similar the curves are overall.

How independent a specific hyperparameter value is from all others by looking at the curve’s shape, and whether curves cross-over (strong interplay) or not (strong independence).

While the results shown in Figure 4 use the full (custom) validation set, those in Figure 5 were computed using the validation set of size 10001000, i.e. with only one image per class. As we have shown in Section 7, this is sufficient to determine the best hyperparameters, and we encourage the community to follow this more realistic protocol.

As can be seen, weight decay and number of training epochs are the two things which matter most when training using only a fraction of ILSVRC-2012.

While we trained thousands of models in order to rigorously test multiple hypotheses (such as that of reducing model capacity), almost all boost in performance could have been achieved in just a few dozen trials with intuitively important hyperparameters (weight decay and epochs), which would take about a week on a modern four-GPU machine.

There are two factors of randomness of a semi supervised model: (1) labeled subset sampling, (2) run with different seeds. In order to estimate the randomness in the performance we train 9 models with random data subsets and random seeds for our proposed S4LS^{4}L method. Table 3 presents the detailed results. Overall, we observe that standard deviation is fairly small across both subsets and different runs and, therefore, our empirical evaluation provides robust comparison of various techniques.

Appendix C More Results in the Transfer Setup

In this section we present more results from the transfer evaluation task on Places205 . Table 4 shows the results for the models mentioned in our main paper. For each method, we select the best model and evaluate its transfer to Places205.

We follow the same setup as to train a linear models with SGD on top of frozen representations. The only difference is the training epochs, we train for 30 epochs in total with learning rate decayed at 10 and 20 epochs respectively. The learning rate is linearly ramped up for the first epoch. Kolesnikov et.al. train for 520 epochs with learning rate decays at 480 and 500 epochs. The schedule used in our paper is much shorter because of our finding that representation learned with labels are more separable and converges significantly faster. (See in Section 6 of the main paper for details.) To make fair comparison with the self-supervised models, results in Table 4 with 0%0\% labels are trained for 520 epochs to ensure their convergence.

From the plain supervised baselines, we observe that either more labels or wider networks lead to more transferable representations. Surprisingly, we found that pseudo labels outperforms the other two semi-supervised baselines in the transfer setup. On the 1%1\% labels evaluation setup, pseudo labels achieves the best result comparing to the other methods. With 10%10\% labels, S4LS^{4}L is comparable to the semi-supervised baselines, and our MOAM clearly outperforms all other models trained on 10%10\% of labels. More interestingly, the MOAM (full) model on 10%10\% is slightly better than the 100%100\% supervised baseline with the same 4×4\times wider network. This indicates that learning a model with multiple losses may lead to representations that generalize better to unseen tasks.