Temporal Ensembling for Semi-Supervised Learning

Samuli Laine, Timo Aila

Introduction

It has long been known that an ensemble of multiple neural networks generally yields better predictions than a single network in the ensemble. This effect has also been indirectly exploited when training a single network through dropout (srivastava2014), dropconnect (dropconnect), or stochastic depth (stochdepth) regularization methods, and in swapout networks (swapout), where training always focuses on a particular subset of the network, and thus the complete network can be seen as an implicit ensemble of such trained sub-networks. We extend this idea by forming ensemble predictions during training, using the outputs of a single network on different training epochs and under different regularization and input augmentation conditions. Our training still operates on a single network, but the predictions made on different epochs correspond to an ensemble prediction of a large number of individual sub-networks because of dropout regularization.

This ensemble prediction can be exploited for semi-supervised learning where only a small portion of training data is labeled. If we compare the ensemble prediction to the current output of the network being trained, the ensemble prediction is likely to be closer to the correct, unknown labels of the unlabeled inputs. Therefore the labels inferred this way can be used as training targets for the unlabeled inputs. Our method relies heavily on dropout regularization and versatile input augmentation. Indeed, without neither, there would be much less reason to place confidence in whatever labels are inferred for the unlabeled training data.

We describe two ways to implement self-ensembling, Π\Pi-model and temporal ensembling. Both approaches surpass prior state-of-the-art results in semi-supervised learning by a considerable margin. We furthermore observe that self-ensembling improves the classification accuracy in fully labeled cases as well, and provides tolerance against incorrect labels.

The recently introduced transform/stability loss of sajjadi16 is based on the same principle as our work, and the Π\Pi-model can be seen as a special case of it. The Π\Pi-model can also be seen as a simplification of the Γ\Gamma-model of the ladder network by ladder, a previously presented network architecture for semi-supervised learning. Our temporal ensembling method has connections to the bootstrapping method of reed14 targeted for training with noisy labels.

Self-ensembling during training

We present two implementations of self-ensembling during training. The first one, Π\Pi-model, encourages consistent network output between two realizations of the same input stimulus, under two different dropout conditions. The second method, temporal ensembling, simplifies and extends this by taking into account the network predictions over multiple previous training epochs.

We shall describe our methods in the context of traditional image classification networks. Let the training data consist of total of NN inputs, out of which MM are labeled. The input stimuli, available for all training data, are denoted xix_{i}, where i∈{1…N}i\in\{1\ldots N\}. Let set LL contain the indices of the labeled inputs, ∣L∣=M|L|=M. For every i∈Li\in L, we have a known correct label yi∈{1…C}y_{i}\in\{1\ldots C\}, where CC is the number of different classes.

In our implementation, the unsupervised loss weighting function w(t)w(t) ramps up, starting from zero, along a Gaussian curve during the first 80 training epochs. See Appendix LABEL:sec:Training_parameters for further details about this and other training parameters. In the beginning the total loss and the learning gradients are thus dominated by the supervised loss component, i.e., the labeled data only. We have found it to be very important that the ramp-up of the unsupervised loss component is slow enough—otherwise, the network gets easily stuck in a degenerate solution where no meaningful classification of the data is obtained.

Our approach is somewhat similar to the Γ\Gamma-model of the ladder network by ladder, but conceptually simpler. In the Π\Pi-model, the comparison is done directly on network outputs, i.e., after softmax activation, and there is no auxiliary mapping between the two branches such as the learned denoising functions in the ladder network architecture. Furthermore, instead of having one “clean” and one “corrupted” branch as in Γ\Gamma-model, we apply equal augmentation and noise to the inputs for both branches.

As shown in Section 3, the Π\Pi-model combined with a good convolutional network architecture provides a significant improvement over prior art in classification accuracy.

2 Temporal ensembling

Analyzing how the Π\Pi-model works, we could equally well split the evaluation of the two branches in two separate phases: first classifying the training set once without updating the weights θ\theta, and then training the network on the same inputs under different augmentations and dropout, using the just obtained predictions as targets for the unsupervised loss component. As the training targets obtained this way are based on a single evaluation of the network, they can be expected to be noisy. Temporal ensembling alleviates this by aggregating the predictions of multiple previous network evaluations into an ensemble prediction. It also lets us evaluate the network only once during training, gaining an approximate 2x speedup over the Π\Pi-model.

An intriguing additional possibility of temporal ensembling is collecting other statistics from the network predictions ziz_{i} besides the mean. For example, by tracking the second raw moment of the network outputs, we can estimate the variance of each output component zi,jz_{i,j}. This makes it possible to reason about the uncertainty of network outputs in a principled way (Gal16). Based on this information, we could, e.g., place more weight on more certain predictions vs. uncertain ones in the unsupervised loss term. However, we leave the exploration of these avenues as future work.

Results

Our network structure is given in Table LABEL:tbl:network, and the test setup and all training parameters are detailed in Appendix LABEL:sec:Training_parameters. We test the Π\Pi-model and temporal ensembling in two image classification tasks, CIFAR-10 and SVHN, and report the mean and standard deviation of 10 runs using different random seeds.

Although it is rarely stated explicitly, we believe that our comparison methods do not use input augmentation, i.e., are limited to dropout and other forms of permutation-invariant noise. Therefore we report the error rates without augmentation, unless explicitly stated otherwise. Given that the ability of an algorithm to extract benefit from augmentation is also an important property, we report the classification accuracy using a standard set of augmentations as well. In purely supervised training the de facto standard way of augmenting the CIFAR-10 dataset includes horizontal flips and random translations, while SVHN is limited to random translations. By using these same augmentations we can compare against the best fully supervised results as well. After all, the fully supervised results should indicate the upper bound of obtainable accuracy.