Reconciling Modern Deep Learning with Traditional Optimization Analyses: The Intrinsic Learning Rate
Zhiyuan Li, Kaifeng Lyu, Sanjeev Arora
Introduction
The training of modern neural networks involves Stochastic Gradient Descent (SGD) with an appropriate learning rate schedule. The formula of SGD with weight decay can be written as:
where is the weight decay factor (or -regularization coefficient), and are the learning rate and batch gradient at the -th iteration.
Traditional analysis shows that SGD approaches a stationary point of the training loss if the learning rate is set to be sufficiently small depending on the smoothness constant and noise scale. In this viewpoint, if we reduce the learning rate by a factor , the end result is the same, and just takes times as many steps. SGD with very tiny step sizes can be thought of as Gradient Descent (GD) (i.e., gradient descent with full gradient), which in the limit of infinitesimal step size approaches Gradient Flow (GF).
However, it is well-known that using only small learning rates or large batch sizes (while fixing other hype-parameters) may lead to worse generalization [Bengio, 2012, Keskar et al., 2017]. From this one concludes that finite (not too small) learning rate —alternatively, noise in the gradient estimate, or small batch sizes— play an important role in generalization, and many authors have suggested that the noise helps avoid sharp minima [Hochreiter and Schmidhuber, 1997, Keskar et al., 2017, Li et al., 2018, Izmailov et al., 2018, He et al., 2019]. Formal understanding of the effect involves modeling SGD via a Stochastic Differential Equation (SDE) in the continuous time limit [Li and Tai, 2019]:
where is the covariance matrix of the noise at . Several works have adopted this SDE view and given some rigorous analysis of the effect of noise [Smith and Le, 2018, Chaudhari and Soatto, 2018, Shi et al., 2020].
While this SDE view is well-established, we will note in this paper that the past works (both theory and experiments) often draw intuition from shallow nets and do not help understand modern architectures, which can be very deep and crucially rely upon normalization schemes such as Batch Normalization (BN) [Ioffe and Szegedy, 2015], Group Normalization (GN) [Wu and He, 2018], Weight Normalization (WN) [Salimans and Kingma, 2016]. We will discuss in Section 4 that these normalization schemes are incompatible to the traditional view points in the following senses. First, normalization makes the loss provably non-smooth around origin, so GD could behave (and does behave, in our experiments) significantly differently from its continuous counterpart, GF, if weight decay is turned on. For example, GD may oscillate between zero loss and high loss and thus cannot persistently achieve perfect interpolation. Second, there is experimental evidence suggesting that the above SDE may be far from mixing for normalized networks in normal training budgets. Lastly, assumptions about the noise in the gradient being a fixed Gaussian turn out to be unrealistic.
In this work, we incorporate effects of normalization in the SDE view to study the complex interaction between BN, weight decay, and learning rate schedule. We particularly focus on Step Decay, one of the most commonly-used learning rate schedules. Here the training process is divided into several phases . In each phase , the learning rate is kept as a constant , and the constants are decreasing with the phase number . Our main experimental observation is the following one, which is formally stated as a conjecture in the context of SDE in Section 5.2.
If trained for sufficiently long time during some phase (), a neural network with BN and weight decay will eventually reach an equilibrium distribution in the function space. This equilibrium only depends on the product , and is independent of the history in the previous phases. Furthermore, the time that the neural net stays at this equilibrium will not affect its future performance.
In this work, we identify a new “intrinsic LR” parameter based on 1.1. The main contributions are the following:
We theoretically analyse how intrinsic LR controls the evolution of effective speed of learning and how it leads to the equilibrium. This is done through incorporating BN and weight decay into the classical framework of Langevin Dynamics (Section 5).
Based on our theory, we empirically observed that small learning rates can perform equally well, which challenges the popular belief that good generalization requires large initial LR (Section 6).
Finally, we make a conjecture, called Fast Equilibrium Conjecture, based on mathematical intuition (Section 5) and experimental evidence (Section 6): the number of steps for reaching equilibrium in 1.1 scales inversely to the intrinsic LR, as opposed to the mixing time upper bound for Langevin dynamics [Bovier, 2004, Shi et al., 2020]. This gives a new perspective in understanding why BN is effective in deep learning.
Related Works
The generalization issue of large batch size / small learning rate has been observed as early as [Bengio, 2012, LeCun et al., 2012]. [Keskar et al., 2017] argued that the cause is that large-batch training tends to converge to sharp minima, but [Dinh et al., 2017] noted that sharp minima can also generalize well due to invariance in ReLU networks. [Li et al., 2019] theoretically analysed the effect of large learning rate in a synthetic dataset to argue that the magnitude of learning rate changes the learning order of patterns in non-homogeneous dataset. To close the generalization gap between large-batch and small-batch training, several works proposed to use a large learning rate in the large-batch training to keep the scale of gradient noise [Hoffer et al., 2017, Smith and Le, 2018, Chaudhari and Soatto, 2018, Smith et al., 2018]. Shallue et al. demonstrated through a systematic empirical study that there is no simple rule for finding the optimal learning rate and batch size as the generalization error could largely depend on other training metaparameters. None of these works have found that training without a large initial learning rate can generalize equally well in presence of BN.
Batch Normalization.
Batch Normalization is proposed in [Ioffe and Szegedy, 2015]. While the original motivation is to reduce Internal Covariate Shift (ICS), [Santurkar et al., 2018] challenged this view and argued that the effectiveness of BN comes from a smoothening effect on the training objective. [Bjorck et al., 2018] empirically observed that the higher learning rate enabled by BN is responsible for the better generalization. [Kohler et al., 2019] studied the direction-length decoupling effect in BN and designed a new optimization algorithm with faster convergence for learning 1-layer or 2-layer neural nets. Another line of works focus on the effect of scale-invariance induced by BN as well as other normalization schemes. [Hoffer et al., 2018a] observed that the effective learning rate of the parameter direction is . [Arora et al., 2019b] identified an auto-tuning behavior for the effective learning rate and [Cai et al., 2019] gave a more quantitative analysis for linear models. In presence of BN and weight decay, [van Laarhoven, 2017] showed that the gradient noise causes the norm to grow and the weight decay causes to shrink, and the effective learning rate eventually reaches a constant value if the noise scale stays constant. [Zhang et al., 2019] validated this phenomenon in the experiments. [Li and Arora, 2020] rigorously proved that weight decay is equivalent to an exponential learning rate schedule.
Preliminaries
Normalization Schemes and Scale-invariance.
Scale-invariance implies the gradient and Hessian are inversely proportional to respectively, meaning that the smoothness is unbounded near . This can be seen by taking gradients with respect to on both sides of :
The gradient can also be proved to be perpendicular to , that is, holds for all . This property can also be seen as a corollary of Euler’s Homogeneous Function Theorem. In the deep learning literature, [Arora et al., 2019b] used this in the analysis of the auto-tuning behavior of normalization schemes.
Approximating SGD by SDE.
Folklore view of landscape exploration.
There is evidence that the training loss has many global minima (or near-minima), whose test loss values can differ radically. The basins around these global minima are separated by “hills” and only large noise can let SGD jump from one to another, while small noise will only make the network oscillate in the same basin around a minimum. The regularization effect of large LR/large noise happens because (1) sharp minima have worse generalization (2) noise prevents getting into narrow basins and thus biases exploration towards flatter basins. (But this view is known to be simplistic, as noted in many papers.)
Apparent Incompatibility between BN and Traditional View Points
In this section, we discuss how BN leads to issues with the traditional optimization view of gradient flow and SDE. This motivates our new view in Section 5.
It’s well known that if LR is smaller than the inverse of the smoothness, then trajectory of gradient descent will be close to that of gradient flow. But for normalized networks, the loss function is scale-invariant and thus provably non-smooth (i.e., smoothness becomes unbounded) around origin [Li et al., 2019]. (By contrast, without WD, the SGD moves away from origin [Arora et al., 2019b] since norm increases monotonically.) We will show that this nonsmoothness is very real and makes training unstable and even chaotic for full batch SGD with any nonzero learning rate. And yet convergence of gradient flow is unaffected.
Consider a toy scale-invariant loss, . Since loss only depends on , WD has no effect on it. Even with WD turned on, Gradient Flow (i.e., infinitesimal updates) will lead to monotone decrease in . But Figure 1(a) in the appendix shows that dynamics for GD with WD are chaotic: as similar trajectories approach the origin, tiny differences are amplified and they diverge.
Modern deep nets with BN + WD (the standard setup) also exhibit instability close to zero loss. See Figures 1(b) and 1(c), where deep nets being trained on small datasets exhibit oscillation between zero loss and high loss. In any significant period with low loss (i.e., almost full accuracy), gradient is small but WD continues to reduce the norm, and resulting non-smoothness leads to large increase in loss.
Problems with random walk/SDE view of SGD.
The standard story about the role of noise in deep learning is that it turns a deterministic process into a geometric random walk in the landscape, which can in principle explore the landscape more thoroughly, for instance by occasionally making loss-increasing steps. Rigorous analysis of this walk is difficult since the mathematics of real-life training losses is not understood. But assuming the noise in SDE is a fixed Gaussian, the stationary distribution of the random walk can be shown to be the familiar Gibbs distribution over the landscape. See [Shi et al., 2020] for a recent account, where SDE is shown to converge to equilibrium distribution in time for some term depending upon loss function. This convergence is extremely slow for small LR and thus way beyond normal training budget.
Recent experiments have also suggested the walk does not reach this initialization-independent equilibrium within normal training time. Stochastic Weight Averaging (SWA) [Izmailov et al., 2018] shows that the loss landscape is nearly convex along the trajectory of SGD with a fixed hyper-parameter choice, e.g., if the two network parameters from different epochs are averaged, the test loss is lower. This reduction can go on for 10 times more than the normal training budget as shown in Figure 2. However, the accuracy improvement is a very local phenomenon since it doesn’t happen for SWA between solutions obtained from different initialization, as shown in [Draxler et al., 2018, Garipov et al., 2018]. This suggests the networks found by SGD within normal training budget highly depends on the initialization, and thus SGD doesn’t mix in the parameter space.
For the case where WD is turned off, [Arora et al., 2019b] proves that the norm of weight is monotone increasing, thus the mixing in parameter space provably doesn’t exist for SGD with BN.
SDE-based framework for modeling SGD on Normalized Networks
For SGD with learning rate and weight decay , we define to be the effective weight decay. This is actually the original definition of weight decay [Hanson and Pratt, 1989] and is also proposed (based upon experiments) in [Loshchilov and Hutter, 2019] as a way to improve generalization for Adam and SGD. In Section 5.1, we will suggest calling the intrinsic learning rate because it controls trajectory in a manner similar to learning rate. Now we can rewrite update rule (2) and its corresponding SDE as
When the loss function is scale-invariant, the gradient noise is inherently anisotropic and position-dependent: Lemma B.1 in the appendix shows the noise lies in the subspace perpendicular to and blows up close to the origin. To get an SDE description closer to the canonical format, we reparametrize parameters to unit norm. Define , , where stands for the -norm of a vector . The following Lemma is proved in the appendix using Itô’s Lemma: {restatable}theoremmainlemma The evolution of the system can be described as:
This shows that the dynamics only depends on the ratio , which also suggests that initial LR is of limited importance, indistinguishable from scale of initialization. Now define . ( was called the effective learning rate in [Hoffer et al., 2018a, Zhang et al., 2019, Arora et al., 2019b].) This simplifies the equations:
(10) can be alternatively written as the following, which shows that squared effective LR is a running average of the norm squared of gradient noise.
ExperimentallyFigures 6 and 7 shows that after a certain length of time the relationship holds approximately, up to a small multiplicative constant. Since is the running average of , the magnitude of the noise, it suggests for different regions of the landscape explored by SGD with different intrinsic LR , the noise scales don’t differ a lot. we find that the trace of noise is approximately constant. This is the assumption of the next lemma (much weaker than assumption of fixed gaussian noise in past works).
If for all encountered in the trajectory, then
The lemma again suggests that the initial effective LR decided together by LR and norm only has a temporary effect on the dynamics: no matter how large is the initial effective LR, after time, the effective LR always converges to the stationary value .
2 A conjecture about mixing time in function space
As mentioned earlier, there is evidence that SGD may take a very long time to mix in the parameter space. However, we observed that test/train errors converge in expectation soon after the norm converges in expectation, which only takes time by our theoretical analysis. More specifically, we find experimentally that the number of steps of SGD (or the length of time for SDE) after the norm converges doesn’t significantly affect the expectation of any known statistics related to the training procedure, including train/test errors, and even the output distribution on every single datapoint. This suggests the neural net reaches an “equilibrium state” in function space.
The above findings motivate us to define the notion of “equilibrium state” rigorously and make a conjecture formally. For learning rate schedule and effective weight decay schedule , we define to be the marginal distribution of in the following SDE when :
It is worth to note that this conjecture obviously does not hold for some pathological initial distributions, e.g., all the neurons are initially dead. But we can verify that our conjecture holds for many initial distributions that can occur in training neural nets, e.g., random initialization, or the distribution after training with certain schedule for certain number of epochs. It remains a future work to theoretically identify the specific condition for this conjecture.
Interesting, we empirically find that the above conjecture still holds even if we are allowed to fine-tune the model before producing the output. This can be modeled by starting another SDE from :
In the appendix we provide the discrete version of the above conjecture by viewing each step of SGD as one step transition of a Markov Chain. By this means we can also extend the conjecture to SGD with momentum and even Adam.
3 What happens in real life training – An interpretation
Let’s first recap Step Decay – there are phases and LR in phase is . Below we will explain or give better interpretation for some phenomena in real life training related to Step Decay.
Usually here LR is dropped by something like a factor . As shown above, the instant effect is to reduce effective LR by a factor of , but it gradually equilibriates to the value , which is only reduced by a factor of . Hence there is a slow rise in error after every drop, as observed in previous works [Zagoruyko and Komodakis, 2016, Zhang et al., 2019, Li and Arora, 2020]. This rise could be beneficial since it coincides with equilibrium state in function space.
Intrisic LR and the final LR decay step:
However, the final LR decay needs to be treated differently. It is customary do early stopping, that finish very soon after the final LR decay, when accuracy is best. The above paragraph can help to explain this by decomposing the training after the final LR decay into two stage. In the first stage, the effective LR is very small, so the dynamics is closer to the classical gradient flow approximation, which can settle into a local basin. In the second stage, the effective LR increases to the stationary value and brings larger noise and worse performance. This decomposition also applies to earlier LR decay operations, but the phenomenon is more significant for the final LR decay because the convergence time is much longer.
Since each phase in Step decay except the last one is allowed to reach equilibrium, the above conjecture suggests the generalization error of Step Decay schedule only depends on the intrinsic LR for its last equilibrium, namely the second-to-last phase. Thus, Step Decay could be abstracted into the following general two-phase training paradigm, where the only hyper-parameter of SGD that affects generalization is the intrinsic LR, :
SDE Phase. Reach the equilibrium of (momentum) SGD with (as fast as possible).
Gradient Flow Phase. Decay the learning rate by a large constant, e.g., 10, and set . Train until loss is zero.
The above training paradigm says for good generalization, what we need is only reaching the equilibrium of small (intrinsic) LR and then decay LR and stop quickly. In other words, the initial large LR should not be necessary to achieve high test accuracy. Indeed our experiments show that networks trained directly with the small intrinsic LR, though necessarily for much longer due to slower mixing, also achieve the same performance. See Figure 3 for SGD with momentum and Figure 10 for vanilla SGD.
So what’s the benefit of early large learning rates?
Empirically we observed that initial large (intrinsic) LR leads to a faster convergence of the training process to the equilibrium. See the red lines in Figure 3. A natural guess for the reason is that directly reaching the equilibrium of small intrinsic LR from the initial distribution is slower than to first reaching the equilibrium of a larger intrinsic LR and then the equilibrium of the target small intrinsic LR. This has been made rigorous for the mixing time of SDE in parameter space [Shi et al., 2020]. In our setting, we show in Appendix D that this argument makes sense at least for norm convergence: the initial large LR reduces the gap between the current norm and the stationary value corresponding to the small LR, in a much shorter time. In Figure 8, we show that the early large learning rate is crucial for the learnability of normalized networks with initial distributions with extreme magnitude. Intriguingly, though without a theoretical analysis, early large learning rate experimentally (see Figure 10(c)) accelerates norm convergence and convergence to equilibrium even with momentum.
But is it worth waiting for the equilibrium of small (intrinsic) LR?
In Figure 3(d) we show that different equilibrium does lead to different performance after the final LR decay. Given this experimental result we speculate the basins of different scales in the optimization landscape seems to be nested, i.e., a larger basin can contain multiple smaller basins of different performances. And reaching the equilibrium of a smaller intrinsic LR seems to be a stronger regularization method, though it also costs much more time.
Batch size and linear scaling rule:
However, this analysis has the following weaknesses and thus we treat batch size as a fixed hyper-parameter in this paper: (1) It’s less general. For example, it doesn’t work for BN, especially when batch size goes to 0, as can be significantly different from due to the noise in batch statistics; (2) Unlike the equivalence between LR and WD, which holds exactly even for SGD [Li et al., 2019], the equivalence between batch size and intrinsic LR , only holds in the regime where and are small and is less meaningful for the more interesting direction, i.e., acceleration via large-batch training [Smith et al., 2020].
Experimental Evidence of Theory
In this subsection we aim to show that the equilibrium only depends on the intrinsic LR, , and is independent of the initial distribution of the weights and individual values of and .
We use a simple 4-layer CNN for MNIST. To highlight the effect of scale-invariance, we make the CNN scale-invariant by fixing the last linear layer as well as the affine parameters in every BN. Figure 4(a) shows the train/test errors for three different schedules, Const, LR-decay and WD-decay. Each error curve is averaged over 500 independent runs, where we call each run as a trial. Const initiates the training with and . LR-decay initiates the training with 4 times larger LR and decreases LR by a factor of every epochs. WD-decay initiates the training with 4 times larger WD and decreases WD by a factor of every epochs. All these three schedules share the same intrinsic LR from epoch to , and thus reach the same train/test errors in this phase as we have conjectured. Moreover, after we setting and at epoch for fine-tuning, all the schedules show the same curve of decreasing train/test errors, which verifies the strong form of our conjecture.
Figure 4(b) measures the total variation between predictions of neural nets trained for 120 and 200 epochs with different schedules. Given a pair of distributions of neural net parameters (e.g., the distributions of neural net parameters after training with LR-decay and WD-decay for epochs), we enumerate each input image from the test set and compute the total variation between the empirical distributions of the class prediction on for weights sampled from , where are estimated through the trials. Figure 4(b) shows that the average total variation over all test inputs decrease with the number of trials, which again suggests that the mixing happens in the function space.
CIFAR-10 Experiments.
We use PreResNet32 for CIFAR10 with data augmentation and the batch size is 128. We modify the downsampling part according to the Appendix C in [Li et al., 2019] and fix the last layer and , in every BN, to ensure the scale invariance. In Figure 5 we focus on the comparison between the performance of the networks within and after leaving the equilibrium, where the networks are initialized differently via different LR/WD schedules before switching to the same intrinsic LR. We repeat this experiment with VGG16 on CIFAR-10 (Figure 11) and PreResNet32 on CIFAR-100 (Figure 12) in appendix and get the same results. A direct comparison between the effect of LR and WD can be found in Figure 6.
2 Reaching Equilibrium only takes O(1/(λη))𝑂1𝜆𝜂O(1/(\lambda\eta)) steps
In Figure 6 we show that convergence of norm is a good measurement for reaching equilibrium, and it takes longer time for smaller intrinsic LR . The two networks are trained with the same sequence of intrinsic learning rates, where the first schedule (blue) decays LR by 2 at epoch , and the second schedule decays WD factor by 2 at the same epoch list. Note that the effective LR almost has the same trend as the training accuracy. Since in each phase, the effective LR , we conclude that the convergence of norm suggests SGD reaches the equilibrium.
In Figure 7 we provide experimental evidence that the mixing time to equilibrium in function space scales to . Note in Equation 12, the convergence of norm also depends on the initial value. Thus in order to reduce the effect of initialization on the time towards equilibrium, we use the setting of Figure 3 in [Li et al., 2019], where we first let the networks with the same architecture reach the equilibrium of different intrinsic LRs, and we decay the LR by and multiplying the WD factor by simultaneously. In this way the intrinsic LR is not changed and the equilibrium is still the same. However, the effective LR is perturbed far away from the equilibrium, i.e. multiplied by . And we measure how long does it takes SGD to recover the network back to the equilibrium and we find it to be almost linear in .
Conclusion and Open Questions
We pointed that use of normalization in today’s state-of-art architectures today leads to a mismatch with traditional mathematical views of optimization. To bridge this gap we develop the mathematics of SGD + BN + WD in scale-invariant nets, in the process identifying a new hyper-parameter “intrinsic learning rate”, , for which appears to determine trajectory evolution and network performance after reaching equilibrium. Experiments suggest time to equilibrium in function space is only , dramatically lower than the usual exponential upper bound for mixing time in parameter space. Our fast equilibrium conjecture about this may guide future theory. The conjecture suggests a more general two-phase training paradigm, which could be potentially interesting to practitioners and lead to better training.
Our theory shows that convergence of norm is a good sign for having reached equilibrium. However, we still lack a satisfying measure of the progress of training, since empirical risk is not good. Finally, it would be good to understand why reaching equilibrium helps regularization.
Acknowledgement
ZL and SA acknowledge support from NSF, ONR, Simons Foundation, Schmidt Foundation, Mozilla Research, Amazon Research, DARPA and SRC.
References
Appendix A Batch Normalization
Batch normalization (BN) [Ioffe and Szegedy, 2015] is one of the most commonly-used normalization schemes. Given a mini-batch of inputs from the last layer, a batch normalization layer first normalizes the inputs by subtracting the mean and dividing the variance , and then applies a linear transformation with trainable parameters and :
Typically BN is placed between the linear transformation and activation function. This makes the loss invariant to the re-scaling of weights in the linear transformation preceding the BN. If we fix the weights in the last linear layer as suggested by [Hoffer et al., 2018b] and put BN after every linear transformation, then the loss is invariant to all its parameters (see Appendix C of [Li and Arora, 2020] for more details).
Appendix B Missing derivation and proofs
For scale-invariant loss function , we have the following lemma on the covariance matrix of gradient noise:
If is scale-invariant with respect to , then
for any .
.
Note that the expectation of scale-invariant functions is scale-invariant, so is scale-invariant. The first bullet can be proved by combining the definition of and
To prove Section 5.1, we will need to use the Itô’s lemma, which is stated below:
Suppose is a vector of Itô’s processes s.t.
we have for any twice differentiable function ,
Recall the following original SDE in the space of . Below we will prove Section 5.1 by Itô’s Lemma.
We can prove (6) and (7) by Itô’s Lemma. For (7), note that , we have
By scale-invariance and Lemma B.1, , , . So we can simplify the formula to conclude that
By scale-invariance and Lemma B.1, , , which means the column span of is orthogonal to . Thus we can apply Lemma B.3 below and get
By scale-invariance, this can be simplified to the following formula:
which proves (6), since is arbitrary. ∎
where .
On the other hand, note , we have
from which we conclude . ∎
By (11), we can upper bound and lower bound by
Therefore, we have . ∎
yields the same trajectory in function space as Equation 4 for scale invariant loss . In fact, they also correspond to the same surrogate SDE Equations 9 and 10, where the exponent in the rate schedule is the intrinsic LR.
The following SDE with exponential LR is equivalent to Equations 9 and 10, where .
By Itô’s Lemma, let , we have
where the last step is by scale-invariance.
Note has the same direction as , i.e. , we can apply Section 5.1 to get Equations 6 and 7, and thus get Equations 9 and 10, with . ∎
Appendix C Extension to Other Optimization Algorithms
In this subsection we use momentum SGD as an example to show how does the discrete version of the fast equilibrium conjecture look like. Throughout this section we will assume all the momentum factors are constant, and we only care about the role of LR and WD factor in the discrete dynamics.
For fixed LR and WD , the formula of SGD with momentum can be written as follows:
We can decouple the effect of WD from SGD by replacing by :
By scale invariance of , letting , we have
which means the effect of in the new parametrization is no more than rescaling the initialization. This motivates as to define as the effective WD, or intrinsic LR.
Unlike vanilla SGD, the evolution of norm for momentum SGD is more complicated. However, a folklore intuition is that, if the gradient of loss changes slowly, one can approximate momentum SGD by vanilla SGD with LR . Therefore, we propose the following discrete version of fast equilibrium conjecture.
For LR schedule and WD schedule , we define to be the marginal distribution of in the following dynamical system when :
For SGD with momentum, modern neural nets converge to the equilibrium distribution in time in the following sense. Given two initial distributions for , constant LR and effective WD schedules , there exists a mixing time , where , such that for any input data from some input domain ,
for all , where , .
C.2 Adam
[Loshchilov and Hutter, 2019] found that using the parametrization of achieves better generalization and a more separable hyper-parameter search space for SGD and Adam, which are named SGDW and AdamW respectively. So far we have justified the role of intrinsic LR for SGD(W). The theorem below shows that the notion of the intrinsic LR also holds for AdamW, while the learning rate has no more power than initialization scale.
For fixed scale-invariant losses , constant schedule multiplier and , multiplying the initial weight and the learning rate by the same constant would not change the trajectory of AdamW (Algorithm 1) in function space.
The notation of LR is slightly different in [Loshchilov and Hutter, 2019] than in the main paper, where is LR and is the Schedule Multiplier. By using schedule multiplier, AdamWAlgorithm 1 can decay LR and WD factor simultaneously. This notation is only used for the statement of the above theorem and its proof.
The proof is based on induction. It suffices to prove that for two history and satisfying that , for all and some , and evolving with and respectively, the following holds:
Note that by scale invariance of , , therefore by definition , i.e. it’s independent of scaling of the history. We now can conclude that
Appendix D Discussion on the Benefit of Early Large Intrinsic LR
Fast equilibrium conjecture says that the equilibrium can be reached in steps for all reasonable initializations. Indeed, Equation 12 indicates that there is also a logarithmic dependency on , i.e., if the initial effective LR is far from the effective LR at equilibrium, then the mixing time can be larger by a multiplicative constant compared to good initial effective LRs. Below we show this constant improvement coming from a good initialization matters a lot for real life training (meaning the training budget is limited), and the usage of initial large intrinsic LR helps SGD to reach a better initialization for the final phase, and thus allow faster mixing to the final equilibrium.
In this section we give experimental evidence that how the fast equilibrium conjecture led by BatchNorm + WD makes the Step Decay training schedule robust to various different initialization methods. In detail, we compare the following 4 types of initialization: Neural Tangent Kernel (NTK) initialization [Jacot et al., 2018, Arora et al., 2019a], Kaiming initialization [He et al., 2015] , Kaiming initialization multiplied by 1000 and Kaiming initialization multiplied by 0.001. In Figure 8 we show that the initial large (intrinsic) learning rate in Step Decay is very necessary to ensure SGD reach the equilibrium of small (intrinsic) LR within the normal training budget, and thus achieving good test accuracy.
Comparison between the above four initialization.
Briefly speaking, these methods are quite similar as they all initialize each parameter by i.i.d. Gaussian, and the only difference is the variance of the gaussian distribution in each layer. For Kaiming initialization, the variance is roughly , and for NTK initialization, the variance is always but there is an additional multiplier of per layer, where is the number of the input channels/neurons that layer. Strictly speaking, NTK initialization is a re-parametrization of Kaiming initialization, rather than a different initialization method, as the additional multiplier indeed changes the architecture. Note that the NTK initialization and Kaiming initialization are always the same in function space. Due the scale invariance led by BatchNorm, all the scaled version of Kaiming initialization are the same as the original Kaiming initialization in function space.
A Theoretical Analysis on Norm Convergence.
Although the convergence of norm is not equivalent to the convergence in function space, analysing the convergence of norm can provide insights into how large LR helps training. Now we theoretically analyse the effect of early large LR on the convergence rate of norm. We compare the following two processes with the same initial norm squared :
Train the neural net with intrinsic LR , then decay it to after the norm converges.
For simplicity, we consider the case that , which means in the first process eventually converges to ; other cases can be transformed to this case by re-scaling the initialization.
For the first process, initially. By Lemma 5.1, converges to in
time. For the second process, initially, and first converges to in time. After LR decay, instantly becomes as is inversely proportional to LR squared. Then we only need another time to make the effective LR converges again. Overall, the second process takes
time. Comparing the second process with the first process, we can see that the large initial LR reduces the dependence of convergence time on the initial norm. It is worth to note that is typically larger than (which equals to ) in Figure 8. Therefore, a large initial LR also leads to faster convergence time without tuning the initialization scale.
Explanation for different convergence rates in Figure 8:
The 4 settings about LR schedules and WD can be interpreted using Equation 15 as the choices of . Let , means starting with , while means starting with the default LR, , which is times larger than . For the rest 3 initializations other than kaiming initialization, from Figure 8, we can see that the initial norm are all exponentially largeAs discussed earlier, NTK initialization has larger weight norm. For 0.001 kaiming initialization, the reason is more subtle: the initial norm are indeed super small, thus leading to huge initial gradient, and therefore the norm grows quickly in the first few iterations., making a large constant. Thus the total steps of training has to be for the effective learning rate to grow and the training to proceed. This could also be seen directly from the ratio of the slopes of the log norm square, which is {\color[rgb]{0,0,1}1}:{\color[rgb]{1,.5,0}10}:{\color[rgb]{0,1,0}5}:{\color[rgb]{1,0,0}50}.
Appendix E Supplementary Figures and Tables for Section 6
This subsection provides supplementary materials to justify that the equilibrium is independent of initialization.
Table 1 shows the LR and WD of each random schedule in Figure 5. Figure 11 and Figure 12 are experiments in similar settings as Figure 5 to show that the equilibrium is independent of initialization for VGG16 on CIFAR-100.
We also validate our claim in the case that we initialize the training with a single possible initial point in a similar setting as Figure 4. That is, we first randomly sample a parameter from the distribution for random initialization, and use it to initialize CNNs in all the independent runs for estimating the equilibrium. Figure 9 shows that CNNs still converge to the equilibrium even if the initial parameter is fixed to the same random sample.
E.2 Equilibrium Can be Reached in O(1/λη)𝑂1𝜆𝜂O(1/\lambda\eta) Steps
In this subsection we provide more experimental evidence that the mixing time to equilibrium in function space scales to . Note in (12), the convergence of norm also depends on the initial value. Thus in order to reduce the effect of initialization on the time towards equilibrium, we use the setting of Figure 3 in [Li et al., 2019], where we first let the networks with the same architecture reach the equilibrium of different intrinsic LRs, and we decay the LR by and multiplying the WD factor by simultaneously. In this way the intrinsic LR is not changed and the equilibrium is still the same. However, the effective LR is perturbed far away from the equilibrium, i.e. multiplied by . And we measure how long does it takes SGD to recover the network back to the equilibrium and we find it to be almost linear in .