The Goldilocks zone: Towards better understanding of neural network loss landscapes

Stanislav Fort, Adam Scherlis

Introduction

A neural networks is fully specified by its architecture – connections between neurons – and a particular choice of weights {W}\{W\} and biases {b}\{b\} – free parameters of the model. Once a particular architecture is chosen, the set of all possible value assignments to these parameters forms the objective landscape – the configuration space of the problem. Given a specific dataset and a task, a loss function LL characterizes how unhappy we are with the solution provided by the neural network whose weights are populated by the parameter assignment PP. Training a neural network corresponds to optimization over the objective landscape, searching for a point – a configuration of weights and biases – producing a loss as low as possible.

The dimensionality, DD, of the objective landscape is typically very high, reaching hundreds of thousands even for the most simple of tasks. Due to the complicated mapping between the individual weight elements and the resulting loss, its analytic study proves challenging. Instead, the objective landscape has been explored numerically. The high dimensionality of the objective landscape brings about considerable geometrical simplifications that we utilize in this paper.

2 Related Work

Neural network training is a large-scale non-convex optimization task, and as such provides space for potentially very complex optimization behavior. (?), however, demonstrated that the structure of the objective landscape might not be as complex as expected for a variety of models, including fully-connected neural networks (?) and convolutional neural networks (?). They showed that the loss along the direct path from the initial to the final configuration typically decreases monotonically, encountering no significant obstacles along the way. The general structure of the objective landscape has been a subject of a large number of studies (?; ?).

A vital part of successful neural network training is a suitable choice of initialization. Several approaches have been developed based on various assumptions, most notably the so-called Xavier initialization (?), and He initialization (?), which are designed to prevent catastrophic shrinking or growth of signals in the network. We address these procedures, noticing their geometric similarity, and relate them to our theoretical model and empirical findings.

3 Our Contributions

Using our observations, we demonstrate a close connection between the Goldilocks zone, measures of local convexity/prevalence of positive curvature, the suitability of a network initialization, and the ability to optimize while constrained to low-dimensional hypersurfaces. Extending the experiments in (?) to different radii and generalizing to hyperspheres, we are able to demonstrate that the main predictor of the success of optimization on a (d≪D)(d\ll D)-dimensional sub-manifold is the amount of its overlap with the Goldilocks zone. We therefore demonstrate that the concept of intrinsic dimension from (?) is radius- and therefore initialization-dependent. Using the realization that common initialization techniques (?; ?), due to properties of high-dimensional Gaussian distributions, initialize neural networks in the same particular region, we conclude that the Goldilocks zone contains an exceptional amount of very suitable initialization configurations.

This paper is structured as follows: We begin by introducing the notion of random hyperplanes and continue with building up the theoretical basis to explain our observations in Section 2. We report the results of our experiments and discuss their implications in Section 3. We conclude with a summary and future outlook in Section 4.

Measurements on Random Hyperplanes and Their Theory

As demonstrated in Section 2.2, common initialization schemes choose points at an approximately fixed radius r=∣P⃗∣r=|\vec{P}|. In the low-dimensional hyperplane limit d≪Dd\ll D, it is exceedingly unlikely that the hyperplane has a significant overlap with the radial direction r^\hat{r}, and we can therefore visualize the hyperplane as a tangent plane at radius ∣P∣|P|. This is illustrated in Figure 1.

For implementation reasons, we decided to generate sparse, nearly-orthogonal projection matrices MM by choosing dd vectors, each having a random number nn of randomly placed non-zero entries, each being equally likely ±1/n\pm 1/\sqrt{n}. Due to the low-dimensional regime d≪Dd\ll D, such matrix is sufficiently near-orthogonal for our purposes, as validated numerically.

2 Gaussian Initializations on a Thin Spherical Shell

Common initialization procedures (?; ?) populate the network weight matrices {W}\{W\} with elements drawn independently from a Gaussian distribution with mean μ=0\mu=0 and a standard deviation σ(W)\sigma(W) dependent on the dimensionality of the particular matrix. For a matrix WW with elements {Wij}\left\{W_{ij}\right\}, the probability of each element having value ww is P(Wij=w)∝exp⁡(−w22σ2)P(W_{ij}=w)\propto\exp\left(-\frac{w^{2}}{2\sigma^{2}}\right). The joint probability distribution is therefore

where we define r2≡∑ijWij2r^{2}\equiv\sum_{ij}W^{2}_{ij}, i.e. the Euclidean norm where we treat each element of the matrix as a coordinate. The probability density at a radius rr therefore corresponds to

where NN is the number of elements of the matrix WW, i.e. the number of coordinates. The probability density for high NN peaks sharply at radius r∗=N−1 σr_{*}=\sqrt{N-1}\,\sigma. That means that the bulk of the random initializations of WW will lie around this radius, and we can therefore visualize such random initialization as assigning a point to a thin spherical shell in the configuration space, as illustrated in Figure 1.

Each matrix WW is initialized according to its own dimensions and therefore ends up on a shell of a different radius in its respective coordinates. We compensate for this by introducing the metric η\eta discussed in Section 2.1. Biases are initialized at zeros. There is a factor O(D)\mathcal{O}(\sqrt{D}) fewer biases than weights in a typical fully-connected network, therefore we ignore biases in our theoretical treatment.

3 Hessian

We would like to explain the observed relationship between several sets of quantities involving second derivatives of the loss function on random, low-dimensional hyperplanes and hyperspheres. The loss at a point x⃗=x0⃗+ε⃗\vec{x}=\vec{x_{0}}+\vec{\varepsilon} can be approximated as

The first-order term involves the gradient, defined as gi=∂L/∂xig_{i}=\partial L/\partial x_{i}. The second-order term uses the second derivatives encapsulated in the Hessian matrix defined as Hij=∂2L/∂xi∂xjH_{ij}=\partial^{2}L/\partial x_{i}\partial x_{j}. The Hessian characterizes the local curvature of the loss function. For a direction v⃗\vec{v}, v⃗THv⃗>0\vec{v}^{T}H\vec{v}>0 implies that the loss is convex along that direction, and conversely v⃗THv⃗<0\vec{v}^{T}H\vec{v}<0 implies that it is concave. We can diagonalize the Hessian matrix to its eigenbasis, in which its only non-zero components {hi}\{h_{i}\} lie on its diagonal. We refer to its eigenvectors as the principal directions, and its eigenvalues {hi}\{h_{i}\} as the principal curvatures.

After restricting the loss function to a dd-dimensional hyperplane, we can compute a new Hessian, HdH_{d}, which comes with its own eigenbasis and eigenvalues. Each hyperplane, therefore, has its own set of principal directions and principal curvatures.

4 Curvature Statistics and Connections to Experiment

We can make sense of these observations using two tools: the fact that high-dimensional multivariate Gaussian distributions have correlation functions that approximate those of the uniform distribution on a hypersphere (?) and the fact (via Wick’s Theorem) that for multivariate Gaussian random variables (X1,X2,X3,X4)(X_{1},X_{2},X_{3},X_{4}),

The variance of the curvature in the vv-direction is

Therefore, when DD is large, the principal directions have the same average curvature that randomly-chosen directions do, but with much more variation. This explains why a slight excess of positive eigenvalues hih_{i} can correspond to an overwhelming majority of positive curvatures in random directions.

Because ∣∣Hd∣∣∝d||H_{d}||\propto d, the principal curvatures on dd-hyperplanes have standard deviation ∣∣Hd∣∣/d∝d||H_{d}||/\sqrt{d}\propto\sqrt{d}, while the average principal curvature stays constant. This explains why smaller dd hyperplanes have a larger excess of positive eigenvalues in the Goldilocks zone than larger dd hyperplanes (Figure 3(a)): their principal curvatures have the same (positive) average value, but vary less, on hyperplanes of smaller dd.

5 Radial Dependence of Loss

Therefore, above some critical rr, the loss will grow as a power law with exponent nLn_{L}. Below this critical value, it will be nearly flat. We confirm this empirically, as shown in Figure 5(a). This scaling of the loss does not apply to tanh⁡\tanh activations, which have a much more bounded linear regime and do not produce arbitrarily large outputs, which might explain the weaker presence of the Goldilocks zone for them, as seen in Figure 4.

6 Radial Features and the Laplacian

Results and Discussion

An unusually high fraction (>1/2>1/2) of positive eigenvalues of the Hessian at randomly initialized points on randomly oriented, low-dimensional hyperplanes intersecting them. The fraction increased as we optimized within the respective hyperplanes, and decreased with an increasing dimension of the hyperplane, as predicted in Section 2. The effect appeared at a well defined range of coordinate space radii we refer to as the Goldilocks zone, as shown in Figure 3(a) and illustrated in Figure 1.

Common initialization schemes (such as Xavier (?) and He (?)), initialize neural networks well within the Goldilocks zone, precisely at the radius at which measures of local convexity/prevalence of positive curvature peak (see Figures 3 and 4).

Hints that selecting initialization points for high measures of local convexity leads to statistically significantly faster convergence (see Figure 6), and correlates well with low initial loss.

An illustration of the Goldilocks zone and its relationship to the common initialization radius is shown in Figure 1.

Using our observations, we draw the following conclusions:

There exists a thick, hollow, spherical shell of unusually high local convexity/prevalence of positive curvature we refer to as the Goldilocks zone (see Figure 1 for an illustration, and Figures 3 and 4 for experimental data).

When optimizing on a random, low-dimensional hypersurface of dimensionality dd, the overlap between the Goldilocks zone and the hypersurface is the main predictor of final accuracy reached, as demonstrated in Figure 2.

The small variance between final accuracy reached by optimization constrained to random, low-dimensional hyperplanes is related to the hyperplanes being a) normal to r^\hat{r}, b) well within the Goldilocks zone, and c) the Goldilocks zone being angularly very isotropic.

Common initialization schemes (such as Xavier (?) and He (?)), initialize neural networks well within the Goldilocks zone, consistently matching the radius at which measures of convexity peak.

Conclusion

We would like to thank Yihui Quek and Geoff Penington from Stanford University for useful discussions.

References