ShakeDrop Regularization for Deep Residual Learning

Yoshihiro Yamada, Masakazu Iwamura, Takuya Akiba, Koichi Kise

I Introduction

Recent advances in generic object recognition have been achieved using deep neural networks. Since ResNet created the opportunity to use very deep convolutional neural networks (CNNs) of over a hundred layers by introducing the building block, its improvements, such as Wide ResNet , PyramidNet , and ResNeXt have broken records for the lowest error rates.

The development of such base network architectures, however, is not sufficient to reduce the generalization error (i.e., difference between the training and test errors) due to over-fitting. In order to improve test errors, regularization methods which are processes to introduce additional information to CNNs have been proposed . Widely used regularization methods include data augmentation , stochastic gradient descent (SGD) , weight decay , batch normalization (BN) , label smoothing , adversarial training , mixup , and dropout . Because the generalization errors when regularization methods are used are still large, effective regularization methods have been studied.

Recently, an effective regularization method which achieved the lowest test error called Shake-Shake regularization was proposed. It is an interesting method, which, in training, disturbs the calculation of the forward pass using a random variable, and also that of the backward pass using a different random variable. Its effectiveness was proven by an experiment on ResNeXt, to which Shake-Shake was applied (hereafter, this type of combination is denoted by “ResNeXt + Shake-Shake”), which achieved the lowest error rate on CIFAR-10/100 datasets . Shake-Shake, however, has the following two drawbacks: (i) it can be applied to ResNeXt only, and (ii) the reason it is effective has not yet been identified.

The current paper addresses these problems. For problem (i), we propose a novel powerful regularization method called ShakeDrop regularization, which is more effective than Shake-Shake. Its main advantage is that it has the potential to be applied not only to ResNeXt (hereafter, three-branch architectures) but also ResNet, Wide ResNet, and PyramidNet (hereafter, two-branch architectures). The main difficulty to overcome is unstable training . We solve this problem by proposing a new stabilizing mechanism for difficult-to-train networks. For problem (ii), in the process of deriving ShakeDrop, we provide an intuitive interpretation of Shake-Shake. Additionally, we present the mechanism in which ShakeDrop works. Through experiments using various base network architectures and parameters, we demonstrate the conditions under which ShakeDrop successfully works.

This paper is an extended version of ICLR workshop paper .

II Regularization methods for the ResNet family

In this section, we present two regularization methods for the ResNet family, both of which are used to derive the proposed method.

Shake-Shake regularization is an effective regularization method for ResNeXt. It is illustrated in Fig. 1. The basic ResNeXt building block, which has a three-branch architecture, is given as

where xx and G(x)G(x) are the input and output of the building block, respectively, and F1(x)F_{1}(x) and F2(x)F_{2}(x) are the outputs of two residual branches.

Let α\alpha and β\beta be independent random coefficients uniformly drawn from the uniform distribution on the interval $$. Then Shake-Shake is given as

where train-fwd and train-bwd denote the forward and backward passes of training, respectively. Expected values E[α]=E[1−α]=0.5E[\alpha]=E[1-\alpha]=0.5. Equation (2) means that the calculation of the forward pass is multiplied by random coefficient α\alpha and that of the backward pass by another random coefficient β\beta. The values of α\alpha and β\beta are drawn for each image or batch. In this paper, we suggest training for longer than usual (more precisely, six times as long as usual).

In the training of neural networks, if the output of a residual branch is multiplied by coefficient α\alpha in the forward pass, then it is natural to multiply the gradient by the same coefficient (i.e., α\alpha) in the backward pass. Hence, compared with the standard approach, Shake-Shake makes the gradient β/α\beta/\alpha times as large as the correctly calculated gradient on one branch and (1−β)/(1−α)(1-\beta)/(1-\alpha) times on the other branch. It seems that the disturbance prevents the network parameters from being captured in local minima. However, the reason why such a disturbance is effective has not been sufficiently identified.

RandomDrop regularization (a.k.a., Stochastic Depth and ResDrop) is a regularization method originally proposed for ResNet, and also applied to PyramidNet . It is illustrated in Fig. 1. The basic ResNet building block, which has a two-branch architecture, is given as

where F(x)F(x) is the output of the residual branch. RandomDrop makes the network appear to be shallow in learning by dropping some stochastically selected building blocks. The lthl^{\textrm{th}} building block from the input layer is given as

where bl∈{0,1}b_{l}\in\{0,1\} is a Bernoulli random variable with the probability P(bl=1)=E[bl]=plP(b_{l}=1)=E[b_{l}]=p_{l}. In this paper, we recommend the linear decay rule to determine plp_{l}, which is given as

where LL is the total number of building blocks and pLp_{L} is the initial parameter. We suggest using pL=0.5p_{L}=0.5.

RandomDrop can be regarded as a simplified version of dropout . The main difference is that RandomDrop drops layers, whereas dropout drops elements.

III Proposed Method

The proposed ShakeDrop, illustrated in Fig. 1, is given as

where blb_{l} is a Bernoulli random variable with probability P(bl=1)=E[bl]=plP(b_{l}=1)=E[b_{l}]=p_{l} given by the linear decay rule (5) in each layer, and α\alpha and β\beta are independent uniform random variables in each element. The most effective ranges of α\alpha and β\beta were experimentally found to be different from those of Shake-Shake, and are α=0\alpha=0, β∈\beta\in and α∈\alpha\in, β∈\beta\in. Further details of the parameters are presented in Sections LABEL:sec:preliminary_experiments and V.

In the training phase, blb_{l} controls the behavior of ShakeDrop. If bl=1b_{l}=1, then (6) is deformed as

that is, ShakeDrop is equivalent to the original network (e.g., ResNet). If bl=0b_{l}=0, then (6) is deformed as

that is, the calculation of F(x)F(x) is perturbed by α\alpha and β\beta.

III-B Derivation of ShakeDrop

We provide an intuitive interpretation of Shake-Shake; to the best of our knowledge, it has not been provided yet. As shown in (2) (and in Fig. 1), in the forward pass, Shake-Shake interpolates the outputs of two residual branches (i.e., F1(x)F_{1}(x) and F2(x)F_{2}(x)) with random weight α\alpha. DeVries and Taylor demonstrated that the interpolation of two data in the feature space can synthesize reasonable augmented data; hence the interpolation in the forward pass of Shake-Shake can be interpreted as synthesizing reasonable augmented data. The use of random weight α\alpha enables us to generate many different augmented data. By contrast , in the backward pass, a different random weight β\beta is used to disturb the updating parameters, which is expected to help to prevent parameters from being caught in local minima by enhancing the effect of SGD .

III-B2 Single-branch Shake Regularization

The regularization mechanism of Shake-Shake relies on two or more residual branches; hence, it can only be applied to three-branch network architectures (i.e., ResNeXt). To achieve a similar regularization to Shake-Shake on two-branch architectures (i.e., ResNet, Wide ResNet, and PyramidNet), we need a different mechanism from interpolation in the forward pass that can synthesize augmented data in the feature space. In fact, DeVries and Taylor demonstrated not only interpolation but also noise addition in the feature space, which generates reasonable augmented data. Hence, following Shake-Shake, we apply random perturbation to the output of a residual branch (i.e., F(x)F(x) of (3)); that is, it is given as

We call this regularization method Single-branch Shake. It is illustrated in Fig. 1. Single-branch Shake is expected to be as effective as Shake-Shake. However, it does not work well in practice. For example, in our preliminary experiments, we applied it to 110-layer PyramidNet with α∈\alpha\in and β∈\beta\in following Shake-Shake. However, the result on the CIFAR-100 dataset was significantly bad (i.e., an error rate of 77.99%).

III-B3 Stabilization of training

In this section, we consider what caused the failure of Single-branch Shake. A natural guess is that Shake-Shake has a stabilizing mechanism that Single-branch Shake does not have. The mechanism is “two residual branches.” We present an argument to verify whether this is the case. As presented in Section II, in training, Shake-Shake makes the gradients of two branches β/α\beta/\alpha times and (1−β)/(1−α)(1-\beta)/(1-\alpha) times as large as the correctly calculated gradients. Thus, when α\alpha is close to zero or one, it cannot converge (ruin) training because it could make a gradient prohibitively largeThis idea is supported by an experiment that limited the ranges of α\alpha and β\beta in Shake-Shake . When α\alpha and β\beta were kept close (more precisely, on the number line, α\alpha and β\beta were on the same side of 0.5, such as α=0.1\alpha=0.1 and β=0.2\beta=0.2), Shake-Shake achieved relatively high accuracy. However, when α\alpha and β\beta were kept far apart (α\alpha and β\beta were on the opposite sides of 0.5, such as α=0.1\alpha=0.1 and β=0.7\beta=0.7), the accuracy was relatively low. This indicates that when β/α\beta/\alpha or (1−β)/(1−α)(1-\beta)/(1-\alpha) were large, training could become less stable.. However, two residual branches of Shake-Shake work as a fail-safe system; that is, even if the coefficient on one branch is large, the other is kept small. Hence, training on at least one branch is not ruined. Single-branch Shake, however, does not have such a fail-safe system.

From the discussion above, the failure of Single-branch Shake was caused by the perturbation being too strong and the lack of a stabilizing mechanism. Because weakening the perturbation would just weaken the effect of regularization, we need a method to stabilize unstable learning under strong perturbation.

We propose using the mechanism of RandomDrop to solve the issue. RandomDrop is designed to make a network apparently shallow to avoid the problems of vanishing gradients, diminishing feature reuse, and a long training time. In our scenario, the original use of RandomDrop does not have a positive effect because a shallower version of a strongly perturbed network (e.g., a shallow version of “PyramidNet + Single-branch Shake”) would also suffer from strong perturbation. Thus, we use the mechanism of RandomDrop as a probabilistic switch for the following two network architectures:

the original network (e.g., PyramidNet), which corresponds to (7), and

a network that suffers from strong perturbation (e.g., “PyramidNet + Single-branch Shake”), which corresponds to (8).

By mixing them up, as shown in Fig. 2, it is expected that (i) when the original network is selected, learning is correctly promoted, and (ii) when the network with strong perturbation is selected, learning is disturbed.

To achieve good performance, the two networks should be well balanced, which is controlled by parameter pLp_{L}. We discuss this issue in Section LABEL:sec:preliminary_experiments.

III-C Relationship with existing regularization methods

In this section, we discuss the relationship between ShakeDrop and existing regularization methods. Among them, SGD and weight decay are commonly used techniques in the training of deep neural networks. Although they were not designed for regularization, researchers have indicated that they have generalization effects . BN is a strong regularization technique that has been widely used in recent network architectures. ShakeDrop is appended to these regularization methods.

ShakeDrop differs from RandomDrop and dropout in the following two ways: they do not explicitly generate new data and they do not update network parameters based on noisy gradients. ShakeDrop coincides with RandomDrop when α=β=0\alpha=\beta=0 instead of the recommended parameters.

Some methods regularize by generating new data. They are summarized in Table I. Data augmentation and adversarial training synthesize data in the (input) data space. They differ in how they generate data. The former uses manually designed means, such as random crop and horizontal flip, whereas the latter automatically generates data that should be used for training to improve generalization performance. Label smoothing generates (or changes) labels for existing data. The methods mentioned above generate new data using a single sample. By contrast, some methods require multiple samples to generate new data. Mixup , BC learning , and RICAP generate new data and their corresponding class labels by interpolating two or more data. Although they generate new data in the data space, manifold mixup also does it in the feature space. Compared with ShakeDrop, which generates data in the feature space using a single sample, none of these regularization methods are in the same category, except for Shake-Shake.

Note that the selection of regularization methods is not always exclusive. We have successfully used ShakeDrop combined with mixup (see Section V-D). Although regularization methods in the same category may not be used together (e.g., “mixup and BC learning” and “ShakeDrop and Shake-Shake”), those of different categories may be used together. Thus, developing the best method in a category is meaningful.

α=1\alpha=1 indicates that the forward pass is normal. Hence, no regularization effect is expected. When β=0\beta=0, the network parameters of the layers selected for perturbat ion (i.e., the layers with bl=0b_{l}=0) are not updated. In layers other than the selected layers, the network parameters are updated as usual. One exception is that, as the network parameters of the selected layers are not updated, other layers compensate for the amount that should be updated on the selected layers. Cases IV and IV contain (1,0)(1,0). They were slightly worse than the best cases .

What does (α,β)=(−1,1)(\alpha,\beta)=(-1,1) do? (What the meaning of α=−1\alpha=-1?) When α=−1\alpha=-1, in the selected layers, the calculation of the forward pass is perturbed by α=−1\alpha=-1. Then, the effect of perturbation is propagated to the succeeding layers. Hence, not only the selected layers but also their succeeding layers are perturbed. In the backward pass, when α\alpha is negative, the network parameters of the selected layers are updated toward the opposite direction to usual. Because of this, the network parameters of the selected layers are strongly perturbed by negative α\alpha. This can be a destructive update. In layers other than the selected layers, it is less probable that the update of the network parameters is destructive because they follow the normal update rule (equivalent to α=1\alpha=1). Cases IV and IV contain (−1,1)(-1,1). The former was slightly worse than the best and the latter was significantly bad.

What does (α,β)=(−1,0)(\alpha,\beta)=(-1,0) do? As this is a combination of α=−1\alpha=-1 and β=0\beta=0, their combined effect occurs. Following the case of α=−1\alpha=-1 mentioned above, the calculation of the forward pass is perturbed, and its effect is propagated to the succeeding layers. In the backward pass, following the case of β=0\beta=0, the network parameters of only the selected layers are not updated. This can avoid destructive updates caused by negative α\alpha. Hence, (−1,0)(-1,0) is expected to be effective. Cases IV and IV contained (−1,0)(-1,0), and the former was the best.

By extending the discussion above, we can interpret the behavior of ShakeDrop using α=0\alpha=0 and β∈\beta\in, which was the most effective on ResNet. When α=0\alpha=0, in the forward pass, the outputs of the selected layers are identical to the inputs. In the backward pass, the amount of updating of the network parameters is perturbed by β\beta.

The proposed ShakeDrop was compared with RandomDrop and Shake-Shake in addition to the vanilla network (without regularization) on ResNet, Wide ResNet, ResNeXt, and PyramidNet. Implementation details are available in Appendix A.

Table III-C shows the conditions and experimental results on CIFAR datasets . In the table, method names are followed by the components of their building blocks. We used the parameters of ShakeDrop found in Section LABEL:sec:preliminary_experiments; that is, the original networks used α=0,β∈\alpha=0,\beta\in and the modified networks in which the residual branches end with BN (e.g., EraseReLU versions) used α∈,β∈\alpha\in,\beta\in. In ResNet and two-branch ResNeXt, in addition to the original form, EraseReLU versions were examined. In Wide ResNet, BN was added to the end of residual branches so that the residual branches ended with BN. In three-branch ResNeXt, we examined two approaches, referred to as “Type A” and “Type B,” to apply RandomDrop and ShakeDrop. “Type A” and “Type B” indicate that the regularization unit was inserted after and before the addition unit for residual branches, respectively; that is, on the forward pass of the training phase, Type A is given by

where D(⋅)D(\cdot) is a perturbation unit of RandomDrop or ShakeDrop, and Type B is given by

where D1(⋅)D_{1}(\cdot) and D2(⋅)D_{2}(\cdot) are individual perturbation units.

Table III-C shows that ShakeDrop can be applied not only to three-branch architectures (ResNeXt) but also two-branch architectures (ResNet, Wide ResNet, and PyramidNet), and ShakeDrop outperformed RandomDrop and Shake-Shake, except for some cases. In Wide ResNet with BN, although ShakeDrop improved the error rate compared with the vanilla network, it did not compared with RandomDrop. This is because the network only had 28 layers. As shown in the RandomDrop paper , RandomDrop is less effective on a shallow network and more effective on a deep network. We observed the same phenomenon in ShakeDrop, and ShakeDrop is more sensitive than RandomDrop. See Section V-E for more detail.

V-B Comparison on the ImageNet dataset

We also conducted experiments on the ImageNet classification dataset using ResNet, ResNeXt, and PyramidNet of 152 layers. The implementation details are presented in Appendix A. We used the best parameters found on the CIFAR datasets, except for pLp_{L}. We experimentally selected pL=0.9p_{L}=0.9.

Table III-C shows the experimental results. Contrary to the CIFAR cases, the EraseReLU versions were worse than the original networks, which does not support the claim of the EraseReLU paper . On ResNet and ResNeXt, in both the original and EraseReLU versions, ShakeDrop clearly outperformed RandomDrop and the vanilla network (ShakeDrop gained 0.84% and 0.15% compared with the vanilla network in the original networks, respectively). On PyramidNet, ShakeDrop outperformed the vanilla network (ShakeDrop gained 0.60% compared with the vanilla network) and also RandomDrop (ShakeDrop gained by 0.29% compared with RandomDrop). Therefore, on ResNet, ResNeXt, and PyramidNet, ShakeDrop clearly outperformed RandomDrop and the vanilla network.

V-C Comparison on the COCO dataset

From the results in Sections V-A and V-B, we considered that ShakeDrop promoted the generality of feature extraction and we evaluated the generality on the COCO dataset . We used Faster R-CNN and Mask R-CNN with the ImageNet pre-trained original version ResNet of 152 layers in Section V-B. The implementation details are presented in Appendix A.

Table LABEL:tab:COCO shows the experimental results. On Faster R-CNN and Mask R-CNN, ShakeDrop clearly outperformed RandomDrop and the vanilla network. Therefore, ShakeDrop promoted the generality of feature extraction not only for image classification but also detection and instance segmentation.

V-D Simultaneous use of ShakeDrop with mixup

As mentioned in Section III-C, we have successfully used ShakeDrop combined with mixup. Table. LABEL:tbl:mixup+ShakeDrop shows the results. In most cases, ShakeDrop further improved the error rates of the base neural networks to which mixup was applied. This indicates that ShakeDrop is not a rival to other regularization methods, such as mixup, but a “collaborator.”

As mentioned in Section II, it has been experimentally found that RandomDrop is more effective on deeper networks (see the figure on the right in Fig. 8 of ). We performed similar experiments on ShakeDrop and RandomDrop to compare their sensitivity to the depth of networks.

Table LABEL:tab:pl_depth shows that the error rates varied over both pLp_{L} and the network depth. ShakeDrop with a large pLp_{L} tended to be effective in shallower networks. The same observation was obtained in the experimental study on the relationship between pLp_{L} of RandomDrop and generalization performance . We recommend a large pLp_{L} for shallower network architectures.

We proposed a new stochastic regularization method called ShakeDrop which, in principle, can be applied to the ResNet family. Through experiments on the CIFAR and ImageNet datasets, we confirmed that, in most cases, ShakeDrop outperformed existing regularization methods of the same category, that is, Shake-Shake and RandomDrop.

All networks were trained using back-propagation by SGD with the Nesterov accelerated gradient and momentum method . Four GPUs (on CIFAR) and eight GPUs (on ImageNet) were used for learning acceleration: because of parallel processing, different observations of blb_{l}, α\alpha, and β\beta were obtained on each GPU. For example, the ll-th layer on a GPU could be perturbed, whereas the layer was not perturbed on other GPUs (ll is an arbitrary number). Additionally, even if the layer was perturbed on multiple GPUs, the different observations of α\alpha and β\beta could be used depending on each GPU.

All implementations used in the experiments were based on the publicly available code of ResNethttps://github.com/facebook/fb.resnet.torch, ResNeXthttps://github.com/facebookresearch/ResNeXt, PyramidNethttps://github.com/jhkim89/PyramidNet, Wide ResNethttps://github.com/szagoruyko/wide-residual-networks, Shake-Shakehttps://github.com/xgastaldi/shake-shake, and Faster/Mask R-CNNhttps://github.com/facebookresearch/maskrcnn-benchmark. We changed their various learning conditions to make them as common as possible on CIFAR (in Section V-A). Table X shows the main changes. The implementation is available at https://github.com/imenurok/ShakeDrop.

The experimental conditions for each type of dataset are described below.

CIFAR datasets The input images of CIFAR datasets were processed in the following manner. The original images of 32×3232\times 32 pixels were color-normalized and then horizontally flipped with a 50% probability. Then, they were zero-padded to be 40×4040\times 40 pixels and randomly cropped to be images of 32×3232\times 32 pixels. On PyramidNet, the initial learning rate was set to 0.1 on CIFAR-10 and 0.5 on CIFAR-100 following the PyramidNet paper . Other than PyramidNet, the initial learning rate was set to 0.1. The initial learning rate was decayed by a factor of 0.1 at 150 epochs and 225 epochs of the entire learning process (300 epochs), respectively. Additionally, a weight decay of 0.0001, momentum of 0.9, and batch size of 128 were used on four GPUs. “MSRA” was used as the filter parameter initializer. We evaluated the top-1 errors without any ensemble technique. Linear decay parameter pL=0.5p_{L}=0.5 was used following the RandomDrop paper . ShakeDrop used parameters of α=0,β=\alpha=0,\beta= (Original) and α=,β=\alpha=,\beta= (EraseReLU on ResNet and ResNeXt, Wide ResNet with BN, and PyramidNet) with the pixel-level update rule.

ImageNet dataset The input images of ImageNet were processed in the following manner. The original image was distorted using a random aspect ratio and randomly cropped to an image size of 224×224224\times 224 pixels. Then, the image was horizontally flipped with a 50% probability and standard color noise was added. On PyramidNet, the initial learning rate was set to 0.5. The initial learning rate was decayed by a factor of 0.1 at 6060, 9090, and 105105 epochs of the entire learning process (120 epochs) following . Additionally, a batch size of 128 was used on eight GPUs. Other than PyramidNet, the initial learning rate was set to 0.1. The initial learning rate was decayed by a factor of 0.1 at 3030, 6060, and 8080 epochs of the entire learning process (90 epochs) following . Additionally, a batch size of 256 was used on eight GPUs. A weight decay of 0.0001 and momentum of 0.9 were used. “MSRA” was used as the filter parameter initializer. We evaluated the top-1 errors without any ensemble technique on the single 224×224224\times 224 image that was cropped from the center of an image resized with the shorter side 256256. pL=0.9p_{L}=0.9 was used as the linear decay parameter. ShakeDrop used parameters of α=0,β=\alpha=0,\beta= (Original) and α=,β=\alpha=,\beta= (EraseReLU on ResNet and ResNeXt, and PyramidNet) with the pixel-level update rule.

COCO dataset Input images of COCO were processed in the following manner. We trained models on the union of the 80k training set and 35k val subset, and evaluated the models on the remaining 5k val subset. We used ResNet-152 for the backbone network and FPN for the predictor network. To use ResNet-152 as a feature extractor, we used the expected value E(bl+α−blα)E(b_{l}+\alpha-b_{l}\alpha) instead of ShakeDrop regularization. According to the experimental condition of the ImageNet dataset, the original image was color-normalized with the means and standard deviations of ImageNet dataset images. The initial learning rate was set to 0.2. The initial learning rate was decayed by a factor of 0.1 at 60,00060,000 and 80,00080,000 iterations of the entire learning process (90,000 iterations). Additionally, a batch size of 16 was used on eight GPUs. A weight decay of 0.0001 was used. The other experimental conditions were set according to maskrcnn-benchmark8.

Acknowledgment

We thank Maxine Garcia, PhD, from Edanz Group (www.edanzediting.com/ac) for editing a draft of this manuscript.

References