An Alternative View: When Does SGD Escape Local Minima?

Robert Kleinberg, Yuanzhi Li, Yang Yuan

Introduction

Nowadays, stochastic gradient descent (SGD), as well as its variants (Adam , Momentum , Adagrad , etc.) have become the de facto algorithms for training neural networks. SGD runs iterative updates for the weights xtx_{t}: xt+1=xt−ηvtx_{t+1}=x_{t}-\eta v_{t}, where η\eta is the step sizeIn this paper, we use step size and learning rate interchangeably.. vtv_{t} is the stochastic gradient that satisfies E[vt]=∇f(xt)E[v_{t}]=\nabla f(x_{t}), and is usually computed using a mini-batch of the dataset.

In the regime of convex optimization, SGD is proved to be a nice tradeoff between accuracy and efficiency: it requires more iterations to converge, but fewer gradient evaluations per iteration. Therefore, for the standard empirical risk minimizing problems with nn points and smoothness LL, to get to ϵ\epsilon-close to x∗x^{*}, GD needs O(Ln/ϵ)O(Ln/\epsilon) gradient evaluations , but SGD with reduced variance only needs O(nlog⁡1ϵ+Lϵ)O(n\log\frac{1}{\epsilon}+\frac{L}{\epsilon}) gradient evaluations . In these scenarios, noise is a by-product of cheap gradient computation, and does not help training.

By contrast, for non-convex optimization problems like training neural networks, noise seems crucial. It is observed that with the help of noisy gradients, SGD does not only converge faster, but also converge to a better solution compared with GD . To formally understand this phenomenon, people have analyzed the role of noise in various settings. For example, it is proved that noise helps to escape saddle points , gives better generalization , and also guarantees polynomial hitting time of good local minima under some assumptions .

However, it is still unclear why SGD could converge to better local minima than GD. Empirically, in additional to the gradient noise, the step size is observed to be a key factor in optimization. More specifically, small step size helps refine the network and converge to a local minimum, while large step size helps escape the current local minimum and go towards a better one . Thus, standard training schedule for modern networks uses large step size first, and shrinks it later . While using large step sizes to escape local minima matches with intuition, the existing analysis on SGD for non-convex objectives always considers the small-step-size settings .

To formalize this intuition, instead of analyzing the sequence xt→xt+1x_{t}\rightarrow x_{t+1}, let us look at the sequence yt→yt+1y_{t}\rightarrow y_{t+1}, where yty_{t} is defined to be xt−η∇f(xt)x_{t}-\eta\nabla f(x_{t}), as in the preceding paragraph. The SGD algorithm never computes these vectors yty_{t}, but we are only using them as an analysis tool. From the equation xt+1=yt−ηωtx_{t+1}=y_{t}-\eta\omega_{t} we obtain the following update rule relating yt+1y_{t+1} to yty_{t}.

This alternative view helps to explain why SGD converges to a good local minimum, even when ff has many other sharp local minima. Intuitively, sharp local minima are eliminated by the convolution operator that transforms ff to gtg_{t}, since convolution has the effect of smoothing out short-range fluctuations. This reasoning ensures that SGD converges to a good local minimum under much weaker conditions, because instead of imposing convexity or one-point convexity requirements on ff itself, we only require those properties to hold for the smoothed functions obtained from ff by convolution. We can formalize the foregoing argument using the following assumption.

For point yy, since the direction x∗−yx^{*}-y points to x∗x^{*}, by having positive inner product with x∗−yx^{*}-y, we know the direction −η∇f(yt−ηωt)-\eta\nabla f(y_{t}-\eta\omega_{t}) in (1) approximately points to x∗x^{*} in expectation (See more discussion on one point convexity in Appendix). Therefore, yty_{t} will converge to x∗x^{*} with decent probability:

Notice that our main theorem not only says SGD will get close to x∗x^{*}, but also says with constant probability, SGD will stay close to x∗x^{*} for the future T2T_{2} steps. As we will see in Section 5, we observe that Assumption 2 holds along the SGD trajectory for the modern neural networks when the noise comes from real data mini-batches. Moreover, the SGD trajectory matches with our theory prediction in practice.

Our main theorem can also help explain why SGD could escape “sharp” local minima and converge to “flat” local minima in practice . Indeed, the sharp local minima have small loss value and small diameter, so after convolved with the noise kernel, they easily disappear, which means Assumption 2 holds. However, flat local minima have large diameter, so they still exists after convolution. In that case, our main theorem says, it is more likely that SGD will converge to flat local minima, instead of sharp local minima.

Previously, people already realized that the noise in the gradient could help SGD to escape saddle points or achieve better generalization . With the help of noise, SGD can also be viewed as doing approximate Bayesian inference or variational inference . Besides, it is proved that SGD with extra noise could “hit” a local minimum with small loss value in polynomial time under some assumptions . However, the extra noise is too big to guarantee convergence, and that model cannot deal with escaping sharp local minima.

Escaping sharp local minima for neural network is important, because it is conjectured (although controversial ) that flat local minima may lead to better generalization . It is also observed that the correct learning rate schedule (small or large) is crucial for escaping bad local minima . Furthermore, solutions that are farther away from the initialization may lead to wider local minima and better generalization . Under a Bayesian perspective, it is shown that the noise in stochastic gradient could drive SGD away from sharp minima, which decides the optimal batch size . There are also explanations for why small batch methods prefers flat minima while large batch methods are not, by investigating the canonical quadratic sums problem .

To visualize the loss surface of neural network, a common practice is projecting it onto a one dimensional line , which was observed to be convex. For the simple two layer neural network, a local one point strongly convexity property provably holds under Gaussian input assumption .

Motivating Example

Main Theorem

Assume that we are running SGD on the sequence {xt}\{x_{t}\}. Recall the update rule (1) for yty_{t}. Our main theorem says that {yt}\{y_{t}\} is converging to x∗x^{*} and will stay around x∗x^{*} afterwards.

This corollary can be easily generalized to shrink the learning rate multiple times.

Our main theorem is based on the important assumption that the step size is bounded. If the step size is too big, even if the whole function ff is one point convex (a stronger assumption than Assumption 2), and we run full gradient descent, we may not keep getting closer to x∗x^{*}, as we show below.

For function ff, if ∀x,⟨−∇f(x),x∗−x⟩≤c′∥x∗−x∥22\forall x,\langle-\nabla f(x),x^{*}-x\rangle\leq c^{\prime}\|x^{*}-x\|_{2}^{2}, and we are at the point xtx_{t}. If we run full gradient descent with step size η>2c′∥xt−x∗∥22∥∇f(xt)∥22\eta>\frac{2c^{\prime}\|x_{t}-x^{*}\|_{2}^{2}}{\|\nabla f(x_{t})\|_{2}^{2}}, we have ∥xt+1−x∗∥22≥∥xt−x∗∥22\|x_{t+1}-x^{*}\|_{2}^{2}\geq\|x_{t}-x^{*}\|_{2}^{2}.

The proof is straightforward and we defer it to Appendix C. ∎

This theorem can be best illustrated with Figure 4. If η\eta is too big, although the gradient (the arrow) is pointing to the approximately correct direction, xt+1x_{t+1} will be farther away from x∗x^{*} (going outside of the x∗x^{*}-centered ball).

Although this theorem analyzes the simple full gradient case, SGD is similar. In the high dimensional case, it is natural to assume that most of the noise will be orthogonal to the direction of xt−x∗x_{t}-x^{*}, therefore with additional noise inside the stochastic gradient, a large step size will drive xt+1x_{t+1} away from x∗x^{*} more easily.

Therefore, our paper provides a theoretical explanation for why picking step size is so important (too big or too small will not work). We hope it could lead to more practical guidelines in the future.

Proof for Theorem 4

In the proof, we will use the following lemma.

Step 1. Since Assumption 2 holds, we show that SGD always makes progress towards x∗x^{*} in expectation, plus some noise.

Step 2. Since SGD makes progress in every step, after many steps, SGD gets very close to x∗x^{*} in expectation. By Markov inequality, this event holds with large probability.

Notice that since η<cL2\eta<\frac{c}{L^{2}}, we have λ=2ηc−η2L2>ηc>0\lambda=2\eta c-\eta^{2}L^{2}>\eta c>0. Recall b≜η2r2(1+ηL)2b\triangleq\eta^{2}r^{2}(1+\eta L)^{2}, we get:

Let Gt=(1−λ)−t(∥yt−x∗∥22−bλ)G_{t}=(1-\lambda)^{-t}(\|y_{t}-x^{*}\|_{2}^{2}-\frac{b}{\lambda}), we get:

That means, GtG_{t} is a supermartingale. We have

Since T1≥log⁡(λ∥y0−x∗∥22b)λT_{1}\geq\frac{\log\left(\frac{\lambda\|y_{0}-x^{*}\|_{2}^{2}}{b}\right)}{\lambda}, we get:

By Markov inequality, we know with probability at least 0.90.9,

For notational simplicity, for the analysis below we relabel the point yT1y_{T_{1}} as y0y_{0}. Therefore, at time we already have ∥y0−x∗∥22≤20bλ\|y_{0}-x^{*}\|_{2}^{2}\leq\frac{20b}{\lambda}.

Step 3. Conditioned on the event that we are close to x∗x^{*}, below we show that if for t0>t≥0t_{0}>t\geq 0, yty_{t} is close to x∗x^{*}, then yt0y_{t_{0}} is also close to x∗x^{*} with high probability.

Let ζ=9T24\zeta=\frac{9T_{2}}{4}. Let event Et={∀τ≤t,∥yτ−x∗∥≤μbλ=δ}\mathfrak{E}_{t}=\{\forall\tau\leq t,\|y_{\tau}-x^{*}\|\leq\mu\sqrt{\frac{b}{\lambda}}=\delta\}, where μ\mu is a parameter satisfies μ≥max⁡{8,42log⁡12(ζ)}\mu\geq\max\{8,42\log^{\frac{1}{2}}(\zeta)\}. If with probability 59\frac{5}{9}, Et\mathfrak{E}_{t} holds for every t≤T2t\leq T_{2}, we are done.

Apply Azuma inequality (Theorem 4), for any ζ>0\zeta>0, we know

Therefore, with probability 1−1ζ1-\frac{1}{\zeta},

Step 4. The inequality above says, if Et−1\mathfrak{E}_{t-1} holds, i.e., for all τ≤t−1,∥yτ−x∗∥≤δ\tau\leq t-1,\|y_{\tau}-x^{*}\|\leq\delta, then with probability 1−1ζ1-\frac{1}{\zeta}, GtG_{t} is bounded. If we can show from the upper bound of GtG_{t} that ∥yt−x∗∥≤δ\|y_{t}-x^{*}\|\leq\delta is also true, we automatically get Et\mathfrak{E}_{t} holds. In other words, that means if Et−1\mathfrak{E}_{t-1} holds, then Et\mathfrak{E}_{t} holds with probability 1−1ζ1-\frac{1}{\zeta}. Therefore, by applying this claim T2T_{2} times, we get ET2\mathfrak{E}_{T_{2}} holds with probability 1−T2ζ=591-\frac{T_{2}}{\zeta}=\frac{5}{9}. Combining with inequality (3), we know with probability at least 1/21/2, the theorem statement holds. Thus, it remains to show that ∥yt−x∗∥≤δ\|y_{t}-x^{*}\|\leq\delta.

The second last inequality holds because we know 11−(1−λ)2=12λ−λ2≤1λ≤1ηc\frac{1}{1-(1-\lambda)^{2}}=\frac{1}{2\lambda-\lambda^{2}}\leq\frac{1}{\lambda}\leq\frac{1}{\eta c}, since λ=2ηc−η2L2≤2ηc<1\lambda=2\eta c-\eta^{2}L^{2}\leq 2\eta c<1, and λ>ηc\lambda>\eta c.

It remains to prove the following lemma, which we defer to Appendix B.

Therefore, ∥yt−x∗∥≤δ\|y_{t}-x^{*}\|\leq\delta. Combining the 4 steps together, we have proved the theorem.

Empirical Observations

In this section, we explore the loss surfaces of modern neural networks, and show that they enjoy many nice one point convex properties. Therefore, our main theorem could be used for explaining why SGD works so well in practice.

It is well known that the loss surface of neural network is highly non-convex, with numerous local minima. However, we observe that the loss surface is consisted of many one point convex basin region, while each time SGD traverses one of such regions.

See Figure 5(a) for details. We ran experiments on Resnet (3434 layers, ≈1.2\approx 1.2M parameters), Densenet (100100 layers, ≈0.8\approx 0.8M parameters) on cifar10 and cifar100, each for 5 trials with 300300 epochs and different initializations. For the start of every epoch xtx_{t} in each trial, we compute the inner product between the negative gradient −∇f(xt)-\nabla f(x_{t}) and the direction x300−xtx_{300}-x_{t}. In Figure 5(a), we plot the minimum value for every epoch among 55 trials for each setting. Notice that except for the starting period of densenet on Cifar-10, all the other networks in all trials have positive inner products, which shows that the trajectory of SGD (except the starting period) is one point convex with respect to the final solutionSimilar observations were implicitly observed previously .. In these experiments, we have used the standard step size schedule (0.10.1 initially, 0.010.01 after epoch 150150, and 0.0010.001 after epoch 225225). However, we got the same observation when using smoothly decreasing step sizes (shrink by 0.990.99 per epoch).

2 The neighborhood of the trajectory is one point convex

Having a one point convex trajectory for 55 trials does not suffice to show SGD always has a simple and easy trajectory, due to the randomness of the stochastic gradient. Indeed, by a slight random perturbation, SGD might be in a completely different trajectory that is far from being one point convex to the final solution. However, in this subsection, we show that it is not the case, as the SGD trajectory is one point convex after convolving with uniform ball with radius 0.50.5. That means, the whole neighborhood of the SGD trajectory is one point convex with respect to the final solution.

In this experiment, we tried Resnet (3434 layers, ≈1.2\approx 1.2M parameters), Densenet (100100 layers, ≈0.8\approx 0.8M parameters) on cifar10 and cifar100We also tried VGG with ≈1\approx 1M parameters, but does not have similar observations. This might be why Resnet and Densenet are slightly easier to optimize.. For every epoch in each setting, we take one point and look at its neighborhood with radius 0.50.5 (upper bound of the length of one SGD step, as we will show below). We take 100100 random points inside each neighborhood to verify Assumption 2We also tried to sample points that are one SGD step away to represent the neighborhood, and got similar observations.. More specifically, for every random point ww in the neighborhood of xtx_{t}, we computer ⟨−∇f(w),x300−xt⟩\langle-\nabla f(w),x_{300}-x_{t}\rangle. Figure 5(b) shows the mean value (solid line), as well as upper and lower bound of the inner product (shaded area). As we can see, the inner products for all epochs in every setting have small variances, and are always positive. Although we could not verify Assumption 2 by computing the exact expectation due to limited computational resources, from Figure 5(b) and Hoeffding bound (Lemma 6), we conclude that Assumption 2 should hold with high probability.

Figure 5(c) shows the norm of the stochastic gradients, including both the mean value (solid lines), as well as upper and lower bounds (shaped area). For all settings, the stochastic gradients are always less than 55 before epoch 150150 with learning rate 0.10.1, and less than 1515 afterwards with learning rate 0.010.01. Therefore, multiplying step size with gradient norm, we know SGD step length is always bounded by 0.50.5.

Notice that the gradient norm gets bigger when we get closer to the final solution (after epoch 150150). This further explains why shrinking step size is important.

3 Loss surface is locally a “slope”

Even with the observation that the whole neighborhood along the SGD trajectory is one point convex with respect to the final solution, there exists a chicken-and-egg concern, as the final target is generated using the SGD trajectory.

In this subsection, we show that the one point convexity is a pretty “global” property. We were running Resnet and Densenet on Cifar10, but with smaller networks (each with about 10K10K parameters). For each network, if we fix the first 1010 epochs, and generate 5050 SGD trajectories with different random seeds for 140140 epochs and 0.10.1 learning rate, we get 5050 different final solutions (they are pretty far away from each other, with minimum pairwise distance 4040). For each network, if we look at the inner product between the negative gradient of any epoch of any trajectories, and the vector pointing to any final solutions, we find that the inner products are almost always positive. (only 0.1%0.1\% of the inner products are not positive for Densenet, and only 22 out of 343,000343,000 inner products are not positive for Resnet).

This indicates that the loss surface is “skewed” to the similar direction, and our observation that the whole SGD trajectory is one point convex w.r. to the last point is not a coincidence. Based on our Theorem 4, such loss surface is very friendly to SGD optimization, even with a few exceptional points that are not one point convex with respect to the final solution.

Notice that in general, it is not possible that all the negative gradients of all points are one point convex with respect to multiple target points. For example, if we take 1D1D interpolation between any two target points, we could easily find points that have negative gradients only pointing to one target point. However, based on our simulation, empirically SGD almost never traverse those regions.

4 Spectrum of the local minima

From the previous subsections, we know that the loss surface of neural network has great one point convex properties. It seems that by our Theorem 4, SGD will almost always converge to a few target points (or regions). However, empirically SGD converges to very different target points. In this subsection, we argue that this is because of the learning rate is too big for SGD to converge (Theorem 3). On the other hand, whenever we shrink the learning rate to 0.010.01, Theorem 4 immediately applies and SGD converges to a local minimum.

In this experiment, we were running smaller version of Resnet and Densenet (each with about 10K10K parameters) on Cifar10 and Cifar100. For each setting, we first train the network with step size 0.10.1 for 300300 epochs, then we pick different epochs as the new starting points for finding nearby local minima using smaller learning rates with additional 150150 epochs.

See Figure 6(a) and Figure 6(b). Starting from different epochs, we got local minima with decreasing validation loss and training loss.

To show that these local minima are not from the same region, we also plot the distance of the local minima to the (unique) initialized point. As we can see, as we pick later epochs as the starting points, we get local minima that are farther away from the initialization with better quality (also observed in ).

Furthermore, we observe that for every local minimum, the whole trajectory is always one point convex to that local minimum. Therefore, the time for shrinking learning rate decides the quality of the final local minimum. That is, using large step size initially avoids being trapped into a bad local minimum, and whenever we are distant enough from the initialization, we can shrink the step size and converge to a good local minimum (due to one point convexity by Theorem 4).

Conclusion

In this paper, we take an alternative view of SGD that it is working on the convolved version of the loss function. Under this view, we could show that when the convolved function is one point convex with respect to the final solution x∗x^{*}, SGD could escape all the other local minima and stay around x∗x^{*} with constant probability.

To show our assumption is reasonable, we look at the loss surface of modern neural networks, and find that SGD trajectory has nice local one point convex properties, therefore the loss surface is very friend to SGD optimization. It remains an interesting open question to prove local one point convex property for deep neural networks.

Acknowledgement

The authors want to thank Zhishen Huang for pointing out a mistake in an early version of this paper, and want to thank Gao Huang, Kilian Weinberger, Jorge Nocedal, Ruoyu Sun, Dylan Foster and Aleksander Madry for helpful discussions. This project is supported by a Microsoft Azure research award and Amazon AWS research award.

References

Appendix A Discussions on one point convexity

If ff is δ\delta-one point strongly convex around x∗x^{*} in a convex domain D\mathcal{D}, then x∗x^{*} is the only local minimum point in D\mathcal{D} (i.e., global minimum).

To see this, for any fixed x∈Dx\in\mathcal{D}, look at the function g(t)=f(tx∗+(1−t)x)g(t)=f(tx^{*}+(1-t)x) for t∈t\in, then g′(t)=⟨∇f(tx∗+(1−t)x),x∗−x⟩g^{\prime}(t)=\langle\nabla f(tx^{*}+(1-t)x),x^{*}-x\rangle. The definition of δ\delta-one point strongly convex implies that the right side is negative for t∈(0,1]t\in(0,1]. Therefore, g(t)>g(1)g(t)>g(1) for t>0t>0. This implies that for every point yy on the line segment joining xx to x∗x^{*}, we have f(y)>f(x∗)f(y)>f(x^{*}), so x∗x^{*} is the only local minimum point.

Appendix B Proof for Lemma 5

On the left hand side there are three summands. Below we show that each of them is bounded by μ2b3λ\frac{\mu^{2}b}{3\lambda}We made no effort to optimize the constants here..

Since μ≥max⁡{8,42log⁡12(ζ)}\mu\geq\max\{8,42\log^{\frac{1}{2}}(\zeta)\}, we know 21bλ≤63b3λ<82b3λ≤μ2b3λ\frac{21b}{\lambda}\leq\frac{63b}{3\lambda}<\frac{8^{2}b}{3\lambda}\leq\frac{\mu^{2}b}{3\lambda}. Next, we have

Adding the three summands together, we get the claim. ∎

Appendix C Proof for Theorem 3

Recall that we have xt+1=xt−η∇f(xt)x_{t+1}=x_{t}-\eta\nabla f(x_{t}). Since we have ⟨−∇f(xt),x∗−xt⟩≤c′∥x∗−xt∥22\langle-\nabla f(x_{t}),x^{*}-x_{t}\rangle\leq c^{\prime}\|x^{*}-x_{t}\|_{2}^{2}, then

Where the last inequality holds since we know η>2c′∥xt−x∗∥22∥∇f(xt)∥22\eta>\frac{2c^{\prime}\|x_{t}-x^{*}\|_{2}^{2}}{\|\nabla f(x_{t})\|_{2}^{2}}. ∎