Swapout: Learning an ensemble of deep architectures

Saurabh Singh, Derek Hoiem, David Forsyth

Introduction

This paper describes swapout, a stochastic training method for general deep networks. Swapout is a generalization of dropout and stochastic depth methods. Dropout zeros the output of individual units at random during training, while stochastic depth skips entire layers at random during training. In comparison, the most general swapout network produces the value of each output unit independently by reporting the sum of a randomly selected subset of current and all previous layer outputs for that unit. As a result, while some units in a layer may act like normal feedforward units, others may produce skip connections and yet others may produce a sum of several earlier outputs. In effect, our method averages over a very large set of architectures that includes all architectures used by dropout and all used by stochastic depth.

Our experimental work focuses on a version of swapout which is a natural generalization of the residual network . We show that this results in improvements in accuracy over residual networks with the same number of layers.

Improvements in accuracy are often sought by increasing the depth, leading to serious practical difficulties. The number of parameters rises sharply, although recent works such as have addressed this by reducing the filter size . Another issue resulting from increased depth is the difficulty of training longer chains of dependent variables. Such difficulties have been addressed by architectural innovations that introduce shorter paths from input to loss either directly or with additional losses applied to intermediate layers . At the time of writing, the deepest networks that have been successfully trained are residual networks (1001 layers ). We show that increasing the depth of our swapout networks increases their accuracy.

There is compelling experimental evidence that these very large depths are helpful, though this may be because architectural innovations introduced to make networks trainable reduce the capacity of the layers. The theoretical evidence that a depth of 1000 is required for practical problems is thin. Bengio and Dellaleau argue that circuit efficiency constraints suggest increasing depth is important, because there are functions that require exponentially large shallow networks to compute . Less experimental interest has been displayed in the width of the networks (the number of filters in a convolutional layer). We show that increasing the width of our swapout networks leads to significant improvements in their accuracy; an appropriately wide swapout network is competitive with a deep residual network that is 1.5 orders of magnitude deeper and has more parameters.

Contributions: Swapout is a novel stochastic training scheme that can sample from a rich set of architectures including dropout, stochastic depth and residual architectures as special cases. Swapout improves the performance of the residual networks for a model of the same depth. Wider but much shallower swapout networks are competitive with very deep residual networks.

Related Work

Convolutional neural networks have a long history (see the introduction of ). They are now intensively studied as a result of recent successes (e.g. ). Increasing the number of layers in a network improves performance if the network can be trained. A variety of significant architectural innovations improve trainability, including: the ReLU ; batch normalization ; and allowing signals to skip layers.

Our method exploits this skipping process. Highway networks use gated skip connections to allow information and gradients to pass unimpeded across several layers . Residual networks use identity skip connections to further improve training ; extremely deep residual networks can be trained, and perform well . In contrast to these architectures, our method skips at the unit level (below), and does so randomly.

Our method employs randomness at training time. For a review of the history of random methods, see the introduction of , which shows that entirely randomly chosen features can produce an SVM that generalizes well. Randomly dropping out unit values (dropout ) discourages co- adaptation between units. Randomly skipping layers (stochastic depth) during training reliably leads to improvements at test time, likely because doing so regularizes the network. The precise details of the regularization remain uncertain, but it appears that stochastic depth represents a form of tying between layers; when a layer is dropped, other layers are encouraged to be able to replace it. Each method can be seen as training a network that averages over a family of architectures during inference. Dropout averages over architectures with “missing” units and stochastic depth averages over architectures with “missing” layers. Other successful recent randomized methods include dropconnect which generalizes dropout by dropping individual connections instead of units (so dropping several connections together), and stochastic pooling (which regularizes by replacing the deterministic pooling by randomized pooling). In contrast, our method skips layers randomly at a unit level enjoying the benefits of each method.

Recent results show that (a) stochastic gradient descent with sufficiently few steps is stable (in the sense that changes to training data do not unreasonably disrupt predictions) and (b) dropout enhances that property, by reducing the value of a Lipschitz constant (, Lemma 4.4). We show our method enjoys the same behavior as dropout in this framework.

Like dropout, the network trained with swapout depends on random variables. A reasonable strategy at test time with such a network is to evaluate multiple instances (with different samples used for the random variables) and average. Reliable improvements in accuracy are achievable by training distinct models (which have distinct sets of parameters), then averaging predictions , thereby forming an explicit ensemble. In contrast, each of the instances of our network in an average would draw from the same set of parameters (we call this an implicit ensemble). Srivastava et al. argue that, at test time, random values in a dropout network should be replaced with expectations, rather than taking an average over multiple instances (though they use explicit ensembles, increasing the computational cost). Considerations include runtime at test; the number of samples required; variance; and experimental accuracy results. For our model, accurate values of these expectations are not available. In Section 4, we show that (a) swapout networks that use estimates of these expectations outperform strong comparable baselines and (b) in turn, these are outperformed by swapout networks that use an implicit ensemble.

Swapout

We use capital letters to represent tensors and ⊙\odot to represent element- wise product (broadcasted for scalars). We use boldface 0\mathbf{0} and 1\mathbf{1} to represent tensors of 0 and 1 respectively. A network block is a set of simple layers in some specific configuration e.g. a convolution followed by a ReLU or a residual network block . Several such potentially different blocks can be connected in the form of a directed acyclic graph to form the full network model.

Dropout kills individual units randomly; stochastic depth skips entire blocks of units randomly. Swapout allows individual units to be dropped, or to skip blocks randomly. Implementing swapout is a straightforward generalization of dropout. Let XX be the input to some network block that computes F(X)F(X). The uu’th unit produces F(u)(X)F^{(u)}(X) as output. Let Θ\Theta be a tensor of i.i.d. Bernoulli random variables. Dropout computes the output YY of that block as

It is natural to think of dropout as randomly selecting an output from the set F(u)={0,F(u)(X)}\mathcal{F}^{(u)}=\{0,F^{(u)}(X)\} for the uu’th unit.

Swapout generalizes dropout by expanding the choice of F(u)\mathcal{F}^{(u)}. Now write {Θi}\{\Theta_{i}\} for NN distinct tensors of iid Bernoulli random variables indexed by ii and with corresponding parameters {θi}\{\theta_{i}\}. Let {Fi}\{F_{i}\} be corresponding tensors consisting of values already computed somewhere in the network. Note that one of these FiF_{i} can be XX itself (identity). However, FiF_{i} are not restricted to being a function of XX and we drop the XX to indicate this. Most natural choices for FiF_{i} are the outputs of earlier layers. Swapout computes the output of the layer in question by computing

and so, for unit uu, we have F(u)={F1(u),F2(u),…,F1(u)+F2(u),…,∑iFi(u)}\mathcal{F}^{(u)}=\{F^{(u)}_{1},F^{(u)}_{2},\ldots,F^{(u)}_{1}+F^{(u)}_{2},\ldots,\sum_{i}F^{(u)}_{i}\}. We study the simplest case where

so that, for unit uu, we have F(u)={0,X(u),F(u)(X),X(u)+F(u)(X)}\mathcal{F}^{(u)}=\{0,X^{(u)},F^{(u)}(X),X^{(u)}+F^{(u)}(X)\}. Thus, each unit in the layer could be:

a feedforward unit (choose F(u)(X)F^{(u)}(X));

or a residual network unit (choose X(u)+F(u)(X)X^{(u)}+F^{(u)}(X)).

Since a swapout network can clearly imitate a residual network, and since residual networks are currently the best-performing networks on various standard benchmarks, we perform exhaustive experimental comparisons with them.

If one accepts the view of dropout and stochastic depth as averaging over a set of architectures, then swapout extends the set of architectures used. Appropriate random choices of Θ1\Theta_{1} and Θ2\Theta_{2} yield: all architectures covered by dropout; all architectures covered by stochastic depth; and block level skip connections. But other choices yield unit level skip and residual connections.

Swapout retains important properties of dropout. Swapout discourages co-adaptation by dropping units, but also by on occasion presenting units with inputs that have come from earlier layers. Dropout has been shown to enhance the stability of stochastic gradient descent (, lemma 4.4). This applies to swapout in its most general form, too. We extend the notation of that paper, and write LL for a Lipschitz constant that applies to the network, ∇f(v)\nabla f(v) for the gradient of the network ff with parameters vv, and D∇f(v)D\nabla f(v) for the gradient of the dropped out version of the network.

1 Inference in Stochastic Networks

A model trained with swapout represents an entire family of networks with tied parameters, where members of the family were sampled randomly during training. There are two options for inference. We could either replace random variables with their expected values, as recommended by Srivastava et al. (deterministic inference). Alternatively, we could sample several members of the family at random, and average their predictions (stochastic inference).

Srivastava et al. argue that deterministic inference is significantly less expensive in computation. We believe that Srivastava et al. may have overestimated how many samples are required for an accurate average, because they use distinct dropout networks in the average (Figure 11 in ). Our experience of stochastic inference with swapout has been positive, with the number of samples needed for good behavior small (Figure 2). Furthermore, computational costs of inference are smaller when each instance of the network uses the same parameters

2 Baseline comparison methods

We compare with ResNet architectures as described in (referred to as v1) and in (referred to as v2).

Dropout:

We use standard dropout (replace equation 3 with equation 1).

Layer Dropout:

We replace equation 3 by Y=X+Θ(1×1)F(X)Y=X+\Theta^{(1\times 1)}F(X). Here Θ(1×1)\Theta^{(1\times 1)} is a single Bernoulli random variable shared across all units.

SkipForward:

Equation 3 introduces two stochastic parameters Θ1\Theta_{1} and Θ2\Theta_{2}. We also explore and compare with a simpler architecture, SkipForward, that introduces only one parameter but samples from a smaller set F(u)={X(u),F(u)(X)}\mathcal{F}^{(u)}=\{X^{(u)},F^{(u)}(X)\} as below.

Experiments

We experiment extensively on the CIFAR-10 dataset and demonstrate that a model trained with swapout outperforms a comparable ResNet model. Further, a 32 layer wider model matches the performance of a 1001 layer ResNet on both CIFAR-10 and CIFAR-100 datasets.

We experiment with ResNet architectures as described in (referred to as v1) and in (referred to as v2). However, our implementation (referred to as ResNet Ours) has the following modifications which improve the performance of the original model (Table 1). Between blocks of different feature sizes we subsample using average pooling instead of strided convolutions and use projection shortcuts with learned parameters. For final prediction we follow a scheme similar to Network in Network . We replace average pooling and fully connected layer by a 1x1 convolution layer followed by global average pooling to predict the logits that are fed into the softmax.

Layers in ResNets are arranged in three groups with all convolutional layers in a group containing equal number of filters. We represent the number of filters in each group as a tuple with the smallest size as (16, 32, 64) (as used in for CIFAR-10). We refer to this as width and experiment with various multiples of this base size represented as W×1W\times 1, W×2W\times 2 etc.

Training:

We train using SGD with a batch size of 128, momentum of 0.9 and weight decay of 0.0001. Unless otherwise specified, we train all the models for a total 256 epochs. Starting from an initial learning rate of 0.1, we drop it by a factor of 10 after 196 epochs and then again after 224 epochs. We do the standard augmentation of left-right flips and random translations of up to four pixels. For translation, we pad the images by 4 pixels on all the sides and sample a random 32x32 crop. All the images in a mini-batch use the same crop. Note that dropout slows convergence (, A.4), and swapout should do so too for similar reasons. Thus using the same training schedule for all the methods should disadvantage swapout.

Models trained with Swapout consistently outperform baselines:

Table 1 compares Swapout with various 20 layer baselines. Models trained with Swapout consistently outperform all other models of similar architecture.

The stochastic training schedule matters:

Different layers in a swapout network could be trained with different parameters of their Bernoulli distributions (the stochastic training schedule). Table 2 shows that different stochastic training schedules have a significant affect on the performance. We report the performance with deterministic as well as stochastic inference. These schedules differ in how the values of parameters θ1\theta_{1} and θ2\theta_{2} of the Bernoulli random variables in equation 3 are set for the different layers. Note that θ1=θ2=0.5\theta_{1}=\theta_{2}=0.5 corresponds to the maximum stochasticity. A schedule with less randomness in the early layers (bottom row) performs the best. This is expected because Swapout adds per unit noise and early layers have the largest number of units. Thus, low stochasticity in early layers significantly reduces the randomness in the system. We use this schedule for all the experiments unless otherwise stated.

Swapout improves over ResNet architecture:

From Table 3 it is evident that networks trained with Swapout consistently show better performance than corresponding ResNets, for most choices of width investigated, using just the deterministic inference. This difference indicates that the performance improvement is not just an ensemble effect.

Stochastic inference outperforms deterministic inference:

Table 3 shows that the stochastic inference scheme outperforms the deterministic scheme in all the experiments. Prediction for each image is done by averaging the results of 30 stochastic forward passes. This difference is not just due to the widely reported effect that an ensemble of networks is better as networks in our ensemble share parameters. Instead, stochastic inference produces more accurate expectations and interacts better with batch normalization.

Stochastic inference needs few samples for a good estimate:

Figure 2 shows the estimated accuracies as a function of the number of forward passes per image. It is evident that relatively few samples are enough for a good estimate of the mean. Compare Figure-11 of , which implies ∼50\sim{50} samples are required.

Increase in width leads to considerable performance improvements:

The number of filters in a convolutional layer is its width. Table 3 shows that the performance of a 20 layer model improves considerably as the width is increased both for the baseline ResNet v2 architecture as well as the models trained with Swapout. Swapout is better able to use the available capacity than the corresponding ResNet with similar architecture and number of parameters. Table 4 compares models trained with Swapout with other approaches on CIFAR-10 while Table 5 compares on CIFAR-100. On both datasets our shallower but wider model compares well with 1001 layer ResNet model.

Swapout uses parameters efficiently:

Persistently over tables 1, 3, and 4, Swapout models with fewer parameters outperform other comparable models. For example, Swapout v2(32) W×4W\times 4 gets 4.76% error with 7.43M parameters in comparison to the ResNet version at 4.91% with 10.2M parameters.

Experiments on CIFAR-100 confirm our results:

Table 5 shows that Swapout is very effective as it improves the performance of a 20 layer model (ResNet Ours) by more than 2%. Widening the network and reducing the stochasticity leads to further improvements. Further, a wider but relatively shallow model trained with Swapout (22.72%; 7.46M params) is competitive with the best performing, very deep (1001 layer) latest ResNet model (22.71%;10.2M params).

Discussion and future work

Swapout is a stochastic training method that shows reliable improvements in performance and leads to networks that use parameters efficiently. Relatively shallow swapout networks give comparable performance to extremely deep residual networks.

We have shown that different stochastic training schedules produce different behaviors, but have not searched for the best schedule in any systematic way. It may be possible to obtain improvements by doing so. We have described an extremely general swapout mechanism. It is straightforward using equation 2 to apply swapout to inception networks (by using several different functions of the input and a sufficiently general form of convolution); to recurrent convolutional networks (by choosing FiF_{i} to have the form F∘F∘F…F\circ F\circ F\ldots); and to gated networks. All our experiments focus on comparisons to residual networks because these are the current top performers on CIFAR-10 and CIFAR-100. It would be interesting to experiment with other versions of the method.

As with dropout and batch normalization, it is difficult to give a crisp explanation of why swapout works. We believe that our results support the idea that swapout causes some form of improvement in the optimization process. This is because relatively shallow networks with swapout reliably work as well as or better than quite deep alternatives; and because swapout is notably and reliably more efficient in its use of parameters than comparable deeper networks. Unlike dropout, swapout will often propagate gradients while still forcing units not to co-adapt. Furthermore, our swapout networks involve some form of tying between layers. When a unit sometimes sees layer ii and sometimes layer i−ji-j, the gradient signal will be exploited to encourage the two layers to behave similarly. The reason swapout is successful likely involves both of these points.

This work is supported in part by ONR MURI Awards N00014-10-1-0934 and N00014-16-1-2007. We would like to thank NVIDIA for donating some of the GPUs used in this work.

References