AdaBelief Optimizer: Adapting Stepsizes by the Belief in Observed Gradients
Juntang Zhuang, Tommy Tang, Yifan Ding, Sekhar Tatikonda, Nicha Dvornek, Xenophon Papademetris, James S. Duncan
Introduction
Modern neural networks are typically trained with first-order gradient methods, which can be broadly categorized into two branches: the accelerated stochastic gradient descent (SGD) family , such as Nesterov accelerated gradient (NAG) , SGD with momentum and heavy-ball method (HB) ; and the adaptive learning rate methods, such as Adagrad , AdaDelta , RMSProp and Adam . SGD methods use a global learning rate for all parameters, while adaptive methods compute an individual learning rate for each parameter.
Compared to the SGD family, adaptive methods typically converge fast in the early training phases, but have poor generalization performance . Recent progress tries to combine the benefits of both, such as switching from Adam to SGD either with a hard schedule as in SWATS , or with a smooth transition as in AdaBound . Other modifications of Adam are also proposed: AMSGrad fixes the error in convergence analysis of Adam, Yogi considers the effect of minibatch size, MSVAG dissects Adam as sign update and magnitude scaling, RAdam rectifies the variance of learning rate, Fromage controls the distance in the function space, and AdamW decouples weight decay from gradient descent. Although these modifications achieve better accuracy compared to Adam, their generalization performance is typically worse than SGD on large-scale datasets such as ImageNet ; furthermore, compared with Adam, many optimizers are empirically unstable when training generative adversarial networks (GAN) .
To solve the problems above, we propose “AdaBelief”, which can be easily modified from Adam. Denote the observed gradient at step as and its exponential moving average (EMA) as . Denote the EMA of and as and , respectively. is divided by in Adam, while it is divided by in AdaBelief. Intuitively, is the “belief” in the observation: viewing as the prediction of the gradient, if deviates much from , we have weak belief in , and take a small step; if is close to the prediction , we have a strong belief in , and take a large step. We validate the performance of AdaBelief with extensive experiments. Our contributions can be summarized as:
We propose AdaBelief, which can be easily modified from Adam without extra parameters. AdaBelief has three properties: (1) fast convergence as in adaptive gradient methods, (2) good generalization as in the SGD family, and (3) training stability in complex settings such as GAN.
We theoretically analyze the convergence property of AdaBelief in both convex optimization and non-convex stochastic optimization.
We validate the performance of AdaBelief with extensive experiments: AdaBelief achieves fast convergence as Adam and good generalization as SGD in image classification tasks on CIFAR and ImageNet; AdaBelief outperforms other methods in language modeling; in the training of a W-GAN , compared to a well-tuned Adam optimizer, AdaBelief significantly improves the quality of generated images, while several recent adaptive optimizers fail the training.
Methods
By the convention in , we use the following notations:
: exponential moving average (EMA) of
: is the EMA of , is the EMA of
: is the learning rate, default is ; is a small number, typically set as
: smoothing parameters, typical values are
are the momentum for and respectively at step , and typically set as constant (e.g.
Adam and AdaBelief are summarized in Algo. 1 and Algo. 2, where all operations are element-wise, with differences marked in blue. Note that no extra parameters are introduced in AdaBelief. Specifically, in Adam, the update direction is , where is the EMA of ; in AdaBelief, the update direction is , where is the EMA of . Intuitively, viewing as the prediction of , AdaBelief takes a large step when observation is close to prediction , and a small step when the observation greatly deviates from the prediction. represents bias-corrected value. Note that an extra is added to during bias-correction, in order to better match the assumption that is bouded below (the lower bound is at leat ). For simplicity, we omit the bias correction step in theoretical analysis.
2 Intuitive explanation for benefits of AdaBelief
Update formulas for SGD, Adam and AdaBelief are:
Note that we name as the “learning rate” and as the “stepsize” for the th parameter. With a 1D example in Fig. 1, we demonstrate that AdaBelief uses the curvature of loss functions to improve training as summarized in Table 1, with a detailed description below:
(1) In region 1g1 in Fig. 1, the loss function is flat, hence the gradient is close to 0. In this case, an ideal optimizer should take a large stepsize. The stepsize of SGD is proportional to the EMA of the gradient, hence is small in this case; while both Adam and AdaBelief take a large stepsize, because the denominator ( and ) is a small value.
(2) In region 1g2, the algorithm oscillates in a “steep and narrow” valley, hence both and is large. An ideal optimizer should decrease its stepsize, while SGD takes a large step (proportional to ). Adam and AdaBelief take a small step because the denominator ( and ) is large.
(3) In region 1g3, we demonstrate AdaBelief’s advantage over Adam in the “large gradient, small curvature” case. In this case, and are large, but and are small; this could happen because of a small learning rate . In this case, an ideal optimizer should increase its stepsize. SGD uses a large stepsize (); in Adam, the denominator is large, hence the stepsize is small; in AdaBelief, denominator is small, hence the stepsize is large as in an ideal optimizer.
To sum up, AdaBelief scales the update direction by the change in gradient, which is related to the Hessian. Therefore, AdaBelief considers curvature information and performs better than Adam.
In practice, the bias correction step will further reduce the error between the EMA and its expectation if is a stationary process . Note that:
An example of the analysis above is summarized in Fig. 2. From Eq. 3 and Eq. 4, note that in Adam, ; this is because the update of only uses the amplitude of and ignores its sign, hence the stepsize for the and direction is the same . AdaBelief considers both the magnitude and sign of , and , hence takes a large step in the direction and a small step in the direction, which matches the behaviour of an ideal optimizer.
In this section, we demonstrate that when the gradient has low variance, the update direction in Adam is close to “sign descent”, hence deviates from the gradient. This is also mentioned in .
In this case, Adam behaves like a “sign descent”; in 2D cases the update is to the axis, hence deviates from the true gradient direction. The “sign update” effect might cause the generalization gap between adaptive methods and SGD (e.g. on ImageNet) . For AdaBelief, when the variance of is the same for all coordinates, the update direction matches the gradient direction; when the variance is not uniform, AdaBelief takes a small (large) step when the variance is large (small).
In this section, we validate intuitions in Sec. 2.2. Examples are shown in Fig. 3, and we refer readers to more video exampleshttps://www.youtube.com/playlist?list=PL7KkG3n9bER6YmMLrKJ5wocjlvP7aWoOu for better visualization. In all examples, compared with SGD with momentum and Adam, AdaBelief reaches the optimal point at the fastest speed. Learning rate is for all optimizers. For all examples except Fig. 3(d), we set the parameters of AdaBelief to be the same as the default in Adam , , and set momentum as 0.9 for SGD. For Fig. 3(d), to match the assumption in Sec. 2.2, we set for both Adam and AdaBelief, and set momentum as for SGD.
Consider the loss function and a starting point near the axis. This setting corresponds to Fig. 2. Under the same setting, AdaBelief takes a large step in the direction, and a small step in the direction, validating our analysis. More examples such as are in the supplementary videos.
For an inseparable loss, AdaBelief outperforms other methods under the same setting.
For an inseparable loss, AdaBelief outperforms other methods under the same setting.
Optimization trajectory under default setting for the Beale function in 2D and 3D.
Optimization trajectory under default setting for the Rosenbrock function.
3 Convergence analysis in convex and non-convex optimization
Similar to , for simplicity, we omit the de-biasing step (analysis applicable to de-biased version). Proof for convergence in convex and non-convex cases is in the appendix.
Suppose in Theorem (2.1), then we have:
(Convergence for non-convex stochastic optimization) Under the assumptions:
is differentiable; ; is also lower bounded.
At step , the algorithm can access a bounded noisy gradient, and the true gradient is also bounded. .
as in , are constants independent of and , and is a constant independent of .
If and assumptions for Theorem 2.2 are satisfied, we have:
Theorem 2.2 implies the convergence rate for AdaBelief in the non-convex case is , which is similar to Adam-type optimizers . Note that regret bounds are derived in the worst possible case, while empirically AdaBelief outperforms Adam mainly because the cases in Sec. 2.2 occur more frequently. It is possible that the above bounds are loose. Also note that we assume , in code this requires to use element wise maximum between and in the denominator.
Experiments
We performed extensive comparisons with other optimizers, including SGD , AdaBound , Yogi , Adam , MSVAG , RAdam , Fromage and AdamW . The experiments include: (a) image classification on Cifar dataset with VGG , ResNet and DenseNet , and image recognition with ResNet on ImageNet ; (b) language modeling with LSTM on Penn TreeBank dataset ; (c) wasserstein-GAN (WGAN) on Cifar10 dataset. We emphasize (c) because prior work focuses on convergence and accuracy, yet neglects training stability.
We performed a careful hyperparameter tuning in experiments. On image classification and language modeling we use the following:
AdaBelief: We use the default parameters of Adam: . SGD, Fromage: We set the momentum as , which is the default for many networks such as ResNet and DenseNet. We search learning rate among . Adam, Yogi, RAdam, MSVAG, AdaBound: We search for optimal among , search for as in SGD, and set other parameters as their own default values in the literature. AdamW: We use the same parameter searching scheme as Adam. For other optimizers, we set the weight decay as ; for AdamW, since the optimal weight decay is typically larger , we search weight decay among . For the training of a GAN, we set for AdaBelief in a small GAN with vanilla CNN generator, and use for a larger spectral normalization GAN (SN-GAN) with a ResNet generator; for other methods, we search for among , and search for among . We set learning rate as for all methods. Note that the recommended parameters for Adam and for RMSProp are within the search range.
We experiment with VGG11, ResNet34 and DenseNet121 on Cifar10 and Cifar100 dataset. We use the official implementation of AdaBound, hence achieved an exact replication of . For each optimizer, we search for the optimal hyperparameters, and report the mean and standard deviation of test-set accuracy (under optimal hyperparameters) for 3 runs with random initialization. As Fig. 4 shows, AdaBelief achieves fast convergence as in adaptive methods such as Adam while achieving better accuracy than SGD and other methods.
We then train a ResNet18 on ImageNet, and report the accuracy on the validation set in Table 2. Due to the heavy computational burden, we could not perform an extensive hyperparameter search; instead, we report the result of AdaBelief with the default parameters of Adam () and decoupled weight decay as in ; for other optimizers, we report the best result in the literature. AdaBelief outperforms other adaptive methods and achieves comparable accuracy to SGD (70.08 v.s. 70.23), which closes the generalization gap between adaptive methods and SGD. Experiments validate the fast convergence and good generalization performance of AdaBelief.
We experiment with LSTM on the Penn TreeBank dataset , and report the perplexity (lower is better) on the test set in Fig. 5. We report the mean and standard deviation across 3 runs. For both 2-layer and 3-layer LSTM models, AdaBelief achieves the lowest perplexity, validating its fast convergence as in adaptive methods and good accuracy. For the 1-layer model, the performance of AdaBelief is close to other optimizers.
Stability of optimizers is important in practice such as training of GANs, yet recently proposed optimizers often lack experimental validations. The training of a GAN alternates between generator and discriminator in a mini-max game, and is typically unstable ; SGD often generates mode collapse, and adaptive methods such as Adam and RMSProp are recommended in practice . Therefore, training of GANs is a good test for the stability.
We experiment with one of the most widely used models, the Wasserstein-GAN (WGAN) and the improved version with gradient penalty (WGAN-GP) using a small model with vanilla CNN generator. Using each optimizer, we train the model for 100 epochs, generate 64,000 fake images from noise, and compute the Frechet Inception Distance (FID) between the fake images and real dataset (60,000 real images). FID score captures both the quality and diversity of generated images and is widely used to assess generative models (lower FID is better). For each optimizer, under its optimal hyperparameter settings, we perform 5 runs of experiments, and report the results in Fig. 6 and Fig. 7. AdaBelief significantly outperforms other optimizers, and achieves the lowest FID score.
Besides the small model above, we also experiment with a large model using a ResNet generator and spectral normalization in the discriminator (SN-GAN). Results are summarized in Table. 3. Compared with a vanilla GAN, all FID scores are lower because the SN-GAN is more advanced. Compared with other optimizers, AdaBelief achieves the lowest FID with both large and small GANs.
Recent research on optimizers tries to combine the fast convergence of adaptive methods with high accuracy of SGD. AdaBound achieves this goal on Cifar, yet its performance on ImageNet is still inferior to SGD . Padam closes this generalization gap on ImageNet; writing the update as , SGD sets , Adam sets , and Padam searches between 0 and 0.5 (outside this region Padam diverges ). Intuitively, compared to Adam, by using a smaller , Padam sacrifices the adaptivity for better generalization as in SGD; however, without good adaptivity, Padam loses training stability. As in Table 4, compared with Padam, AdaBelief achieves a much lower FID score in the training of GAN, meanwhile achieving slightly higher accuracy on ImageNet classification. Furthermore, AdaBelief has the same number of parameters as Adam, while Padam has one more parameter hence is harder to tune.
1 Extra experiments
We conducted extra experiments according to public discussions after publication. Considering the heavy computation burden, we did not compare Adabelief with all other optimizers as in previous sections, instead we only compare with the result from the official implementations. For details and code to reproduce results, please refer to our github page.
We conducted experiments with a Transformer model . The code is modified from the official implementation of AdaHessian . As shown in Table. 5, AdaBelief outperforms Adam and RAdam. AdaBelief is slightly worse than AdaHessian, this could be caused by the "block-averaging" trick in AdaHessian, which could be used in AdaBelief but requires extra coding efforts.
We show the results on PASCAL VOC object deetection with a Faster-RCNN model . The results are reported in , and shown in Table. 6. AdaBelief outperforms other optimizers including Adam, RAdam and SGD, and achieves a higher mAP.
We conducted reinforcement learning with SAC (soft actor critic) in the PFRL package . We plot the reward vs. training step in Fig. 8. AdaBelief achieves higher reward value than Adam.
Related works
This work considers the update step in first-order methods. Other directions include Lookahead which updates “fast” and “slow” weights separately, and is a wrapper that can combine with other optimizers; variance reduction methods which reduce the variance in gradient; and LARS which uses a layer-wise learning rate scaling. AdaBelief can be combined with these methods. Other variants of Adam have also been proposed (e.g. NosAdam , Sadam and Adax ).
Besides first-order methods, second-order methods (e.g. Newton’s method , Quasi-Newton method and Gauss-Newton method , L-BFGS , Natural-Gradient , Conjugate-Gradient ) are widely used in conventional optimization. Hessian-free optimization (HFO) uses second-order methods to train neural networks. Second-order methods typically use curvature information and are invariant to scaling but have heavy computational burden, and hence are not widely used in deep learning.
Conclusion
We propose the AdaBelief optimizer, which adaptively scales the stepsize by the difference between predicted gradient and observed gradient. To our knowledge, AdaBelief is the first optimizer to achieve three goals simultaneously: fast convergence as in adaptive methods, good generalization as in SGD, and training stability in complex settings such as GANs. Furthermore, Adabelief has the same parameters as Adam, hence is easy to tune. We validate the benefits of AdaBelief with intuitive examples, theoretical convergence analysis in both convex and non-convex cases, and extensive experiments on real-world datasets.
Broader Impact
Optimization is at the core of modern machine learning, and numerous efforts have been put into it. To our knowledge, AdaBelief is the first optimizer to achieve fast speed, good generalization and training stability. Adabelief can be used for the training of all models that can numerically estimate parameter gradients, hence can boost the development and application of deep learning models. This work mainly focuses on the theory part, and the social impact is mainly determined by each application rather than by optimizer.
Acknowledgments and Disclosure of Funding
This research is supported by NIH grant R01NS035193.
References
Appendix
A. Detailed Algorithm of AdaBelief
By the convention in , we use the following notations:
: is the learning rate, default is ; is a small number, typically set as
: smoothing parameters, typical values are
: exponential moving average (EMA) of
: is the EMA of , is the EMA of
B. Convergence analysis in convex online learning case (Theorem 2.1 in main paper)
For the ease of notation, we absorb into . Equivalently, . For simplicity, we omit the debiasing step in theoretical analysis as in . Our analysis can be applied to the de-biased version as well.
Note that since . Use and to denote the th dimension of and respectively. From lemma (.1), using and , we have:
Note that and , rearranging inequality (B. Convergence analysis in convex online learning case (Theorem 2.1 in main paper)), we have:
Now bound in Formula (B. Convergence analysis in convex online learning case (Theorem 2.1 in main paper)), assuming .
Apply formula (B. Convergence analysis in convex online learning case (Theorem 2.1 in main paper)) to (B. Convergence analysis in convex online learning case (Theorem 2.1 in main paper)), we have:
Suppose in Theorem (.2), then we have:
Proof: By sum of arithmetico-geometric series, we have:
Plugging (7) into (B. Convergence analysis in convex online learning case (Theorem 2.1 in main paper)), we can derive the results above. ∎
C. Convergence analysis for non-convex stochastic optimization (Theorem 2.2 in main paper)
A1, is differentiable and has gradient, . is also lower bounded.
A2, at time , the algorithm can access a bounded noisy gradient, the true gradient is also bounded. .
Suppose assumptions A1-A3 are satisfied, is chosen such that . For some constant , \Big{|}\Big{|}\alpha_{t}\frac{m_{t}}{\sqrt{s_{t}}}\Big{|}\Big{|}\leq G,\forall t. Then Adam-type algorithms yield
where are constants independent of and , is a constant independent of , the expectation is taken all randomness corresponding to . Furthermore, let denote the minimum possible value of effective stepsize at time over all possible coordinate and past gradients . The convergence rate of Adam-type algorithm is given by
where is defined through the upper bound of RHS of (.3), and
Proof: We provide the proof from in next section for completeness. ∎
Proof: We first derive an upper bound of the RHS of formula (.3), then derive a lower bound of the LHS of (.3).
Next we derive the lower bound of LHS of (.3).
Combining (Assumptions), (Assumptions), (Assumptions) and (13) to (.3), we have:
If and assumptions for Theorem .3 are satisfied, we have:
By assumption, , then we have
D. Proof of Theorem .3
Let in the Algorithm, consider the sequence
Suppose that the conditions in Theorem (.3) hold, then
Suppose that the condition in Theorem .3 hold, in (26) can be bounded as:
Suppose the conditions in Theorem .3 are satisfied, then in (28) can be bounded as
Suppose assumptions in Theorem .3 are satisfied, then in (29) can be bounded as:
Suppose the assumptions in Theorem .3 are satisfied, then in (30) can be bounded as:
Suppose the assumptions in Theorem .3 are satisfied, then in (27) are bounded as:
We provide the proof from for completeness. We combine Lemma .5, .6, .7, .8, .9, .10 and .11 to bound the objective.
E. Bayesian interpretation of AdaBelief
We analyze AdaBelief from a Bayesian perspective.
We skip the proof, which is a direct application of the Bayes rule in the Gaussian distribution case as in . If is averaged across a batch of size , we can replace with .
According to Theorem .12, the gradient descent direction with maximum expected gain is:
From a practical perspective, can be interpreted as a numerical term to avoid division by 0; from the Bayesian perspective, represents our prior on , with a larger indicating a larger . Note that as the network evolves with training, the distribution of the gradient is distorted (an example with Adam is shown in Fig. 2 of ), hence the Gaussian prior might not match the true distribution. To solve the mismatch between prior and the true distribution, it might be reasonable to use a weak prior during late stages of training (e.g., let grow at late training phases, and when reduces to a uniform prior). We only provide a Bayesian perspective here, and leave the detailed discussion to future works.
F. Experimental Details
We performed experiments based on the official implementationhttps://github.com/Luolc/AdaBound of AdaBound , and exactly replicated the results of AdaBound as reported in . We then experimented with different optimizers under the same setting: for all experiments, the model is trained for 200 epochs with a batch size of 128, and the learning rate is multiplied by 0.1 at epoch 150. We performed extensive hyperparameter search as described in the main paper. In the main paper we only report test accuracy; here we report both training and test accuracy in Fig. 1 and Fig. 2. AdaBelief not only achieves the highest test accuracy, but also a smaller gap between training and test accuracy compared with other optimizers such as Yogi.
Image Classification on ImageNet
We experimented with a ResNet18 on ImageNet classication task. For SGD, we use the same learning rate schedule as , with an initial learning rate of 0.1, and multiplied by 0.1 at epoch 30 and 60; for AdaBelief, we use an initial learning rate of 0.001, and decayed it at epoch 70 and 80. Weight decay is set as for both cases. To match the settings in [loshchilov2018fixing] and , we use decoupled weight decay. As shown in Fig. 3, AdaBelief achieves an accuracy very close to SGD, closing the generalization gap between adaptive methods and SGD. Meanwhile, when trained with a large learning rate (0.1 for SGD, 0.001 for AdaBelief), AdaBelief achieves faster convergence than SGD in the initial phase.
Robustness to hyperparameters
We test the performances of AdaBelief and Adam with different values of varying from to in a log-scale grid. We perform experiments with a ResNet34 on Cifar10 dataset, and summarize the results in Fig. 4. Compared with Adam, AdaBelief is slightly more sensitive to the choice of , and achieves the highest accuracy at the default valiue ; AdaBelief achieves accuracy higher than for all values, consistently outperforming Adam which achieves an accuracy around .
We test the performance of AdaBelief with different learning rates. We experiment with a VGG11 network on Cifar10, and display the results in Fig. 5. For a large range of learning rates from to , compared with Adam, AdaBelief generates higher test accuracy curve, and is more robust to the change of learning rate.
Experiments with LSTM on language modeling
We experiment with LSTM models on Penn-TreeBank dataset, and report the results in Fig. 6. Our experiments are based on this implementation https://github.com/salesforce/awd-lstm-lm. Results are measured across 3 runs with independent initialization. For completeness, we plot both the training and test curves.
We use the default parameters for 2-layer and 3-layer models; for 1-layer model we set and set other parameters as default. For simple models (1-layer LSTM), AdaBelief’s perplexity is very close to other optimizers; on complicated models, AdaBelief achieves a significantly lower perplexity on the test set.
Experiments with GAN
We experimented with a WGAN and WGAN-GP . The code is based on several public github repositories https://github.com/pytorch/examples,https://github.com/eriklindernoren/PyTorch-GAN. We summarize network structure in Table 1. For WGAN, the weight of discriminator is clipped within ; for WGAN-GP, the weight for gradient-penalty is set as 10.0, as recommended by the original implementation. For each optimizer, we perform 5 independent runs. We train the model for 100 epochs, generate 64,000 fake samples (60,000 real images in Cifar10), and measure the Frechet Inception Distance (FID) between generated samples and real samples. Our implementation on FID heavily relies on an open-source implementationhttps://github.com/mseitzer/pytorch-fid. We report the FID scores in the main paper, and demonstrate fake samples in Fig. 7 and Fig. 8 for WGAN and WGAN-GP respectively.
We also experimented with Spectral Normalization GAN based on a public repository https://github.com/POSTECH-CVLab/PyTorch-StudioGAN. For this experiment, we set and use the rectification technique as in RAdam. Other hyperparamters and training schemes are the same as in the repository.