Normalization Propagation: A Parametric Technique for Removing Internal Covariate Shift in Deep Networks

Devansh Arpit, Yingbo Zhou, Bhargava U. Kota, Venu Govindaraju

Introduction and Motivation

Ioffe & Szegedy (2015) identified an important problem involved in training deep networks, viz., Internal Covariate Shift. It refers to the problem of shifting distribution of the input of every hidden layer in a deep neural network. This idea is borrowed from the concept of covariate shift (Shimodaira, 2000), where this problem is faced by a single input-output learning system. Consider the last layer of a deep network being used for classification; this layer essentially tries to learn P(Y∣X)P(Y|X), where YY is the class label random variable (r.v.) and XX is the layer input r.v. However, learning a fixed P(Y∣X)P(Y|X) becomes a problem if P(X)P(X) changes continuously. As a result, this slows down training convergence.

Batch Normalization (BN) addresses this problem by normalizing the distribution of every hidden layer’s input. In order to do so, it calculates the pre-activation mean and standard deviation using mini-batch statistics at each iteration of training and uses these estimates to normalize the input to the next layer. While this approach leads to a significant performance jump by addressing internal covariate shift, its estimates of mean and standard-deviation of hidden layer input for validation rely on mini-batch statistics, which are not representative of the entire data distribution (especially during initial training iterations). This is because the mini-batch statistics of input to hidden layers depends on the output from previous layers, which in turn depend on the previous layer parameters that keep shifting during training, and a moving average of these estimates are used for validation. Finally, due to involvement of batch statistics, BN is inapplicable with batch-size 11.

In this paper, we propose a simple parametric normalization technique for addressing internal covariate shift that does not depend on batch statistics for normalizing the input to hidden layers and is less severely affected by the problem of shifting parameters during validation. In fact, we show that it is unnecessary to explicitly calculate mean and standard-deviation from mini-batches for normalizing the input to hidden layers even for training. Instead, a data independent estimate of these normalization components are available in closed form for every hidden layer, assuming the pre-activation values follow Gaussian distribution and that the weight matrix of hidden layers are roughly incoherent. We show how to forward propagate the normalization property (of the data distribution) to all hidden layers by exploiting the knowledge of the distribution of the pre-activation values (Gaussian) and some algebraic manipulations. Hence we call our approach Normalization Propagation.

Background

It has long been known in Deep Learning community that input whitening and decorrelation helps in speeding up the training process. In fact, it is explicitly mentioned in (LeCun et al., 2012) that this whitening should be performed before every layer so that the input to the next layer has zero mean. From the perspective of Internal Covariate Shift, what is required for the network to learn a hypothesis P(Y∣X)P(Y|X) at any given layer at every point in time during training, is for the distribution P(X)P(X) of the input to that layer to be fixed. While whitening could be used for achieving this task at every layer, it would be a very expensive choice (cubic order of input size) since whitening dictates computing the Singular Value Decomposition (SVD) of the input data matrix. However, Desjardins et al. (2015) suggest to overcome this problem by approximating this SVD by: a) using sub-sampled training data to compute this SVD; b) computing it every few number of iteration and relying on the assumption that this SVD approximately holds for the iterations in between. In addition, each hidden layer’s input is then whitened by re-parametrizing a subset of network parameters that are involved in gradient descent. As mentioned in Ioffe & Szegedy (2015), this re-parametrizing may lead to effectively cancelling/attenuating the effect of the gradient update step since these two operations are done independently.

Batch Normalization (Ioffe & Szegedy, 2015) addresses both the above problems. First, they propose a strategy for normalizing the data at hidden layers such that the gradient update step accounts for this normalization. Secondly, this normalization is performed for units of each hidden layer independently (thus avoiding whitening) using mini-batch statistics. Specifically, this is achieved by normalizing the pre-activation u=WTx\mathbf{u}=\mathbf{W}^{T}\mathbf{x} of all hidden layers as,

where uiu_{i} denotes the ithi^{th} element of u\mathbf{u} and the expectation/variance is calculated over the training mini-batch B\mathcal{B}. Notice since W\mathbf{W} is a part of this normalization, it becomes a part of the gradient descent step as well. However, a problem common to both the above approaches is that of shifting network parameters upon which their approximation of input normalization for hidden layers depends.

Normalization Propagation (NormProp) Derivation

We will now describe the idea behind NormProp. At a glance the problem at hand seems cyclic because estimating the mean and standard deviation of the input distribution to any hidden layer requires the input distribution of its previous layer (and hence its parameters) to be fixed to the optimal value before hand. However, as we will now show, we can side-step this naive approach and get an approximation of the true unbiased estimate using the knowledge that the pre-activation to every hidden layer follows a Gaussian distribution and some algebraic manipulation over the properties of the weight matrix. For the derivation below, we will focus on networks with ReLU activation, and later discuss how to extend our algorithm to other activation functions.

2 Mean and Standard-deviation Normalization for First Hidden Layer

The above proposition tells us two things. First, the covariance matrix of the pre-activation Σ\mathbf{\Sigma} is approximately canonical (diagonal covariance matrix) if the above error bound can be made tight, and that this tightness can be controlled by certain properties of the weight matrix W\mathbf{W}. Second, if we want to normalize each element of the vector u\mathbf{u} to have unit standard-deviation, then our best bet is to divide each ui\mathbf{u}_{i} by the corresponding weight length ∥Wi∥2\lVert\mathbf{W}_{i}\rVert_{2} if we ensure tight canonical error bound. This is because the closest estimate of a diagonal variance for Σ\mathbf{\Sigma} is αi∗=∥Wi∥22\alpha_{i}^{*}=\lVert\mathbf{W}_{i}\rVert_{2}^{2} (σ=1\sigma=1 in our case).

At this point, we have normalized the pre-activation u\mathbf{u} to have zero mean and unit variance (divide each pre-activation element by corresponding ∥Wi∥2\lVert\mathbf{W}_{i}\rVert_{2}). As a result, the output of the first hidden layer (\mboxReLU(u)\mbox{ReLU}(\mathbf{u})) is Rectified Gaussian distribution. Notice that the above bound ensures the dimensions of u\mathbf{u} and hence (\mboxReLU(u)\mbox{ReLU}(\mathbf{u})) are roughly uncorrelated. Thus, if we subtract the distribution mean from \mboxReLU(u)\mbox{ReLU}(\mathbf{u}) and divide by its standard deviation, we will have reduced the dynamics of the second layer to be identical to that of the first layer. The mean and standard deviation of the aforementioned Rectified Gaussian is,

Hence in order to normalize the post-activation \mboxReLU(u)\mbox{ReLU}(\mathbf{u}) to have zero mean and unit standard, the above calculated values can be used. Finally, in the case of Pooling (in Conv-Nets), we essentially take a block of post-activated units and take average or maximum of these values. If we consider each such unit to be independent then the distribution after pooling, will have a different mean and standard deviation. However, in reality, each of these units are highly correlated since they involve computation over either overlapping or spatially close patches. Therefore, we found that the distribution statistics do not get affected significantly and hence we do not recompute mean and standard deviation post-pooling.

3 Propagation to Higher Layers

With the above two operations, the dynamics of the second hidden layer become identical to that of the first hidden layer. By induction, repeating these two operations for every layer, viz.–1) divide every hidden layer’s pre-ReLU-activation by its corresponding ∥Wi∥2\lVert\mathbf{W}_{i}\rVert_{2}, where W\mathbf{W} is the corresponding layer’s weight matrix, 2) subtract and divide 1/2π\sqrt{1/2\pi} and 12(1−1π)\sqrt{\frac{1}{{2}}\left(1-\frac{1}{\pi}\right)} (respectively) from every hidden layer’s post-ReLU-activation– we ensure that the input to every layer is a canonical distribution. While training, all these normalization operations are back-propagated.

4 Effect of NormProp on Jacobian of Hidden Layers

It has been discussed in Saxe et al. (2013); Ioffe & Szegedy (2015) that Jacobian of hidden layers with singular values close to one improves training convergence in deep networks. While BN has been shown to intuitively achieve this condition, we will now show more rigorously that NormProp (approximately) indeed achieves this condition.

Finally taking an expectation of JJT\mathbf{J}\mathbf{J}^{T} over the distribution of x\mathbf{x} which is Normal, we get,

NormProp: Implementation Details

We have all the ingredients required to filter out the steps for Normalization Propagation for training any deep neural network with ReLU activation though out its hidden layers. Like BN, NormProp can be used alongside any optimization algorithm (eg. Stochastic Gradient Descent with/without momentum) for training deep networks.

Since the core idea of NormProp is to propagate the data normalization through hidden layers, we offer two alternative choices either one of which can be used for normalizing the input to a NormProp network. As we will describe, both options are justified in their respective scenario.

1. Global Data Normalization: In cases when the entire dataset – approximately representing the true data distribution – is available at hand, we compute the global mean and standard deviation for each feature element. Then the first step for NormProp is to subtract element-wise mean calculated over the entire dataset from each sample. Similarly divide each feature element by the element-wise standard-deviation. Ideally it is required by NormProp that all input dimensions be statistically uncorrelated, a property achieved by whitening for instance, but we suggest element-wise normalization as an approximation since it is computationally cheaper. Notice this precludes the dilemma of what range the input should be scaled to before passing through the network.

2. Batch Data Normalization: In many real world scenario, streaming data is available and thus it is not possible to compute an unbiased estimate of global mean and standard deviation at any given point in time. In such cases, we propose to instead batch-normalize every mini-batch training data fed to the network. Again, we perform the normalization of each feature element independently for computational purposes. Notice this normalization is only performed at the data level, all hidden layers are still normalized by the NormProp strategy which is not affected by shifting model parameters as compared to BN. Moreover, Batch Data Normalization also serves as a regularization since each data sample gets a different representation each time depending on the mini-batch it comes with. Thus by using the Batch Data Normalization strategy we actually benefit from the regularization aspect of BN but also overcome its drawbacks by computing the hidden layer mean and standard-deviation without depending on batch statistics. Notice this strategy is most effective when the incoming data is well shuffled.

2 Initialize Network Parameters

We use Normalized Initialization (Glorot & Bengio, 2010) for setting the initial values of all the weight matrices, both fully connected and convolutional. Bias vectors are initialized to zeros and scaling vectors (described in the next subsection) can either be initialized to 11 or as described in the next subsection.

3 Propagate Normalization

Similar to BN, we also make use of gradient-based-learnable scaling and bias parameters γ\mathbf{\gamma} and β\mathbf{\beta} during implementation. We will now describe our normalization in detail for both fully connected and convolutional layers.

Now in the case of NormProp, the output oi{o}_{i} becomes,

Here we initialize each γi\gamma_{i} to 1/1.211/1.21 in order to make the Jacobian close to one as suggested by our analysis in section 3.4 for ReLU activation. Thus we call this number the Jacobian factor. We found this initializing using Jacobian factor helps training with larger learning rates without diverging. However, one can also choose to treat the initialization value as a hyper-parameter.

3.2 Convolutional Layers

where ∗\mathbf{*} denotes the convolution operation. Now in the case of NormProp, the output feature map oi\mathbf{o}_{i} becomes,

where each element of γi{\gamma}_{i} is again initialized to 1/1.211/1.21. Notice each γi{\gamma}_{i} is multiplied to all outputs from the same corresponding filter and similarly all the scalars as well as the bias vector are broadcasted to all the dimensions. Pooling is done after this normalization process the same way as done traditionally.

4 Training

The network is trained using Back Propagation. While doing so, all the normalizations also get back-propagated at every layer.

Optimization: We use Stochastic Gradient Descent with momentum (set to 0.90.9) for training. Data shuffling also leads to performance improvement (this however, is true in general while training deep networks).

Learning Rate: We found learning speeds up by reducing the learning rate by half whenever the training error starts saturating. Also, we found larger initial learning rate for larger batch size improves performance.

Regularizations: We use weight decay along with the loss function; we found a small coefficient value of 0.0005−0.0050.0005-0.005 is necessary during training. We found Dropout does not help during training; we believe this might be because Dropout changes the distribution of output of the layer it is applied, which affects NormProp.

5 Validation and Testing

Validation and test procedures are identical for NormProp. While validation/testing, each sample is first normalized using mean and standard deviation which are calculated depending on how the train data is normalized during training. In case we use Global Data Normalization during training, we simply use the same global estimate of mean and standard deviation to normalize each test/validation sample. On the other hand, if Batch Data Normalization is used during training, a running estimate of mean and standard deviation is maintained during training which is then used to normalize every test/validation sample. Finally, the input is forward propagated though the network with learned parameters using the same strategy described in section 4.3.

6 Extension to other Activation Functions

Even though our paper shows how to overcome the problem of Internal Covariate Shift specifically for networks with ReLU activation throughout, we have in essence proposed a general framework for propagating normalization done at data level to all hidden layers. All that is needed for extending NormProp to other activation functions is to compute the distribution mean (c2c_{2}) and standard deviation (c1c_{1}) of output after the activation function of choice, similar to what is shown in remark 1. Thus the general form of output for any given activation σ(.)\sigma(.) becomesUsing the appropriate Jacobian Factor allows the use of larger learning rate; however, NormProp works without it as well. (shown for convolution layer as an example),

This activation can be both parameter based or fixed. For instance, a parameter based activation is Parametric ReLU (PReLU, He et al. (2015)) (with parameter aa) given by,

Then the post PReLU distribution statistics is given by,

Notice the distribution mean and standard deviation depends on the parameter aa and thus will be involved in the normalization process. In case of non-parameter based activations (eg. Tanh, Sigmoid), one can either choose to analytically compute the statistics (like we did for ReLU) or compute these values empirically by simulation since the input distribution to the activation is a fixed Normal distribution. Thus NormProp is a general framework which can be extended to any activation function of choice.

Empirical Results and Observations

We want to verify the following: a) performance comparison of NormProp when using Global Data Normalization vs. Batch Data Normalization; b) NormProp alleviates the problem of Internal Covariate Shift more accurately compared to BN; c) thus, convergence stability of NormProp is better than BN; d) effect of batch-size on the behaviour of NormProp, especially batch-size 11 (BN not applicable). Finally we report classification result on various datasets using NormProp and BN.

Datasets: We use the following datasets, 1) CIFAR-1010 (Krizhevsky, 2009)– It consists of 60,00060,000 32×3232\times 32 real world color images in 10 classes split into 50,00050,000 train and 10,00010,000 test images. We use 50005000 images from train set for validation and remaining for training. 2) CIFAR-100100– It has the same number of train and test samples as CIFAR-1010 but it has 100100 classes. For training, we use hyperparameters same as those for CIFAR-1010. 3) SVHN (Netzer et al., 2011)– It consists of 32×3232\times 32 color images of house numbers collected by Google Street View. It has 73,25773,257 train images, 26,03226,032 test images and an additional 5,31,1315,31,131 train images. Similar to the protocol in (Goodfellow et al., 2013), we select 400 samples per class from the train set and 200 samples per class from the extra set as validation and use the remaining images of the train and extra sets for training.

Experimental Protocols (For experiments in sections 5.1 through 5.4): We use CIFAR-1010 with the following Network in Network (Lin et al., 2014) architectureWe use the following shorthand for a) conv layer: C(number of filters, filter size, stride size, padding); b) pooling: P(kernel size, stride, padding, pool mode) C(192,5,1,2)−C(160,1,1,0)−P(3,2,1,max)−C(96,1,1,0)−C(192,5,1,2)−C(192,1,1,0)−P(3,2,1,avg)−C(192,1,1,0)−C(192,5,1,0)−C(192,1,1,2)−C(10,1,1,0)−P(8,8,0,avg)C(192,5,1,2)-C(160,1,1,0)-P(3,2,1,\text{max})-C(96,1,1,0)-C(192,5,1,2)-C(192,1,1,0)-P(3,2,1,\text{avg})-C(192,1,1,0)-C(192,5,1,0)-C(192,1,1,2)-C(10,1,1,0)-P(8,8,0,\text{avg}). For any specified initial learning rate, we reduce it by half every 1010 epochs. We use Stochastic Gradient Descent with momentum 0.90.9. We use test set during validation for convergence analysis.

Since we offer two alternate ways to normalize data (section 4.1) fed to a NormProp network, we evaluate both strategies with different batch sizes to see the difference in performance. We use batch sizesNotice this batch size has nothing to do with the data normalization strategies in discussion. Different batch sizes are used only for adding more variation in experiments. 5050 and 100100 using initial learning rates 0.050.05 and 0.080.08 respectively. The results are shown in figure 1. The performance Even though the numbers are very close, the best accuracy of 90.35%90.35\% is achieved by Batch Data Normalization using batch size 5050. using both strategies is very similar for both batch sizes, converging in only 3030 epochs. This shows the robustness and applicability of NormProp in both streaming data as well as block data scenario. However, since Batch Data Normalization strategy is a more practical choice, we stick to this strategy for the rest of the experiments when using NormProp.

2 NormProp vs. BN– Internal Covariate Shift

The fundamental goal of our paper (as well as that of Batch Normalization Ioffe & Szegedy, 2015) is to alleviate the problem of Internal Covariate Shift. This implies preventing the distribution of hidden layer inputs from shifting while the network is being trained. In deep networks, the features generated by higher layers are completely dependent on the lower features since all the information in data is propagated from lower to higher layers. Thus the problem of Internal Covariate Shift in lower hidden layers is expected to affect the overall performance more severely as compared to the same problem in higher layers.

In order to study the effect of normalization by NormProp and BN on hidden layers, we train two separate networks using each strategy and an additional network without any normalization as our baseline. After every training epoch, we record the mean of the input distribution (over the validation set) to a single randomly chosen (but fixed) unit in each hidden layer. We use batch size 5050 and an initial learning rate of 0.050.05 for NormProp and BN, and 0.00010.0001 for training the network without any normalization (larger learning rates cause divergence). For the 99 layer convolutional networks we train, the input mean to the last 88 layers against training epoch are shown in figure 2. There are three important observations in these figures: a) NormProp achieves significantly more stable input distribution for lower hidden layers compared to BN, thus facilitating good lower level representation; b) the input distribution for all hidden layers converge after  32~{}32 epochs for NormProp. On the other hand, the input distribution to the second layer for BN remains un-converged even after 5050 epochs; c) on an average, input distribution to all layers converge closer to zero for NormProp (avg. 0.190.19) as compared to BN (avg. 0.330.33). Finally the performance of the network trained without any normalization is in-comparable to the normalized ones due to large variations in the hidden layer input distribution (especially the lower layers). This experiment also serves to show the Canonical Error Bound (proposition 1) is small since the input statistics to hidden layers are roughly preserved.

3 Convergence Stability of NormProp vs. BN

As a result of alleviating Internal Covariate Shift more accurately during validation as compared to BN, NormProp is expected to achieve a more stable convergence. We confirm this intuition by recording the validation accuracy while the network is being trained. We use batch size 5050 and initial learning rates 0.050.05. The plot is shown in figure 3We observed identical trends on SVHN and CIFAR-100100. Additionally, we also experimented optimizing with SGD without momentum and RMS prop (Tieleman & Hinton, 2012). We found in general (for most mini-batch sizes) the performance of RMS prop was worse than SGD with momentum while that of SGD without momentum was worse than both. On the other hand, RMSProp generally performed better than SGD but SGD with batch size 1 was very similar to SGD-Momentum.. Clearly NormProp achieves a more stable convergence in general, but especially during initial training. This is because NormProp achieves a more stable hidden layer input distribution computed for validation.

4 Effect of Batch-size on NormProp

We want to see the effect of batch-size used during training with NormProp. Since it is also possible to train with batch size 11 (using Global Data Normalization at data layer), we compare the validation performance of NormProp during training for various batch sizes including 11. The plots are shown in figure 4. The performance of NormProp is largely unaffected by batch size although lower batch sizes seem to yield better performance.

5 Results on various Datasets

We evaluate NormProp and BN on CIFAR-1010, CIFAR-100100 and SVHN datasets, but also report existing state-of-the-art (SOTA) results. For all the datasets and both methods, we use the same architecture as mentioned in the experimental protocol above except for CIFAR-100100, the last convolutional layer is C(100,1,1,0)C(100,1,1,0) instead of C(10,1,1,0)C(10,1,1,0). For CIFAR datasets we use batch size 5050 and an initial learning rate of 0.050.05 and reduce it by half after every 2525 epochs and train for 200200 epochs. Since SVHN is a much larger dataset, we only train for 2525 epochs with batch size 100100 and an initial learning rate of 0.080.08 and reduce it by half after every 55 epochs. We use Stochastic gradient descent with momentum (0.90.9). For CIFAR-1010 and CIFAR-100100, we train using both without data augmentation and with data augmentation (horizontal flipping only); and no data augmentation for SVHN. We did not pre-process any of the datasets. The results are shown in table 1. We find NormProp consistently achieves either better or competitive performance compared to BN, but also beats existing SOTA results.

6 Training Speed

Since there is no need for estimating the running average values of input mean and standard deviation for hidden layers for NormProp algorithm, it expected to be faster compared to Batch Normalization. So we record the time taken for NormProp and BN for 1 epoch of training on CIFAR-1010 dataset using the experimental protocol used for above experiments. On an NVIDIA GeForce GTX Titan X GPU with Intel i7-3930K CPU and 32GB Ram machine, NormProp takes ∼84\sim 84 sec while BN takes ∼96\sim 96 sec.

Conclusion

We have proposed a novel algorithm for addressing the problem of Internal Covariate Shift involved during training deep neural networks that overcomes certain drawbacks of Batch Normalization (BN). Specifically, we propose a parametric approach (NormProp) that avoids estimating the mean and standard deviation of hidden layers’ input distribution using input data mini-batch statistics (that involve shifting network parameters). Instead, NormProp relies on normalizing the statistics of the given dataset and conditioning the weight matrix which ensures normalization done for the dataset is propagated to all hidden layers. Thus NormProp does not need to maintain a moving average estimate of batch statistics of hidden layer inputs for validation/test phase, thus being more representative of the entire data distribution (especially during initial training period when parameters change drastically). This also enables the use of batch size 11 for training. Although we have shown how to apply NormProp in detail for networks with ReLU activation, we have discussed (section 4.6) how to extend it for other activations as well. We have empirically shown NormProp achieves more stable convergence and hidden layer input distribution over validation set during training, and better/competitive classification performance compared with BN while being faster by omitting the need to compute mini-batch estimate of mean/standard-deviation for hidden layers’ input. In conclusion, our approach is applicable alongside any activation function and cost objectives for improving training convergence.

References

Appendix A Proofs

On the other hand, the covariance of u\mathbf{u} is given by,

α2\mathbf{\alpha}^{2} in the above equation denotes element-wise square of elements of α\alpha. Finally minimizing w.r.t αi\alpha_{i} ∀i∈{1,…,m}\forall i\in\{1,\ldots,m\}, leads to αi∗=σ2∥Wi∥22\alpha_{i}^{*}=\sigma^{2}\lVert\mathbf{W}_{i}\rVert_{2}^{2}. Substituting this into equation 18, we get,

For the definition of XX and YY, we have,

Substituting var⁡(Z)=1−2π\operatorname{var}(Z)=1-\frac{2}{\pi} yields the claimed result. ∎

Substituting var⁡(Z)=1−2π\operatorname{var}(Z)=1-\frac{2}{\pi} yields the claimed result. ∎