The Goldilocks zone: Towards better understanding of neural network loss landscapes
Stanislav Fort, Adam Scherlis
Introduction
A neural networks is fully specified by its architecture – connections between neurons – and a particular choice of weights and biases – free parameters of the model. Once a particular architecture is chosen, the set of all possible value assignments to these parameters forms the objective landscape – the configuration space of the problem. Given a specific dataset and a task, a loss function characterizes how unhappy we are with the solution provided by the neural network whose weights are populated by the parameter assignment . Training a neural network corresponds to optimization over the objective landscape, searching for a point – a configuration of weights and biases – producing a loss as low as possible.
The dimensionality, , of the objective landscape is typically very high, reaching hundreds of thousands even for the most simple of tasks. Due to the complicated mapping between the individual weight elements and the resulting loss, its analytic study proves challenging. Instead, the objective landscape has been explored numerically. The high dimensionality of the objective landscape brings about considerable geometrical simplifications that we utilize in this paper.
2 Related Work
Neural network training is a large-scale non-convex optimization task, and as such provides space for potentially very complex optimization behavior. (?), however, demonstrated that the structure of the objective landscape might not be as complex as expected for a variety of models, including fully-connected neural networks (?) and convolutional neural networks (?). They showed that the loss along the direct path from the initial to the final configuration typically decreases monotonically, encountering no significant obstacles along the way. The general structure of the objective landscape has been a subject of a large number of studies (?; ?).
A vital part of successful neural network training is a suitable choice of initialization. Several approaches have been developed based on various assumptions, most notably the so-called Xavier initialization (?), and He initialization (?), which are designed to prevent catastrophic shrinking or growth of signals in the network. We address these procedures, noticing their geometric similarity, and relate them to our theoretical model and empirical findings.
3 Our Contributions
Using our observations, we demonstrate a close connection between the Goldilocks zone, measures of local convexity/prevalence of positive curvature, the suitability of a network initialization, and the ability to optimize while constrained to low-dimensional hypersurfaces. Extending the experiments in (?) to different radii and generalizing to hyperspheres, we are able to demonstrate that the main predictor of the success of optimization on a -dimensional sub-manifold is the amount of its overlap with the Goldilocks zone. We therefore demonstrate that the concept of intrinsic dimension from (?) is radius- and therefore initialization-dependent. Using the realization that common initialization techniques (?; ?), due to properties of high-dimensional Gaussian distributions, initialize neural networks in the same particular region, we conclude that the Goldilocks zone contains an exceptional amount of very suitable initialization configurations.
This paper is structured as follows: We begin by introducing the notion of random hyperplanes and continue with building up the theoretical basis to explain our observations in Section 2. We report the results of our experiments and discuss their implications in Section 3. We conclude with a summary and future outlook in Section 4.
Measurements on Random Hyperplanes and Their Theory
As demonstrated in Section 2.2, common initialization schemes choose points at an approximately fixed radius . In the low-dimensional hyperplane limit , it is exceedingly unlikely that the hyperplane has a significant overlap with the radial direction , and we can therefore visualize the hyperplane as a tangent plane at radius . This is illustrated in Figure 1.
For implementation reasons, we decided to generate sparse, nearly-orthogonal projection matrices by choosing vectors, each having a random number of randomly placed non-zero entries, each being equally likely . Due to the low-dimensional regime , such matrix is sufficiently near-orthogonal for our purposes, as validated numerically.
2 Gaussian Initializations on a Thin Spherical Shell
Common initialization procedures (?; ?) populate the network weight matrices with elements drawn independently from a Gaussian distribution with mean and a standard deviation dependent on the dimensionality of the particular matrix. For a matrix with elements , the probability of each element having value is . The joint probability distribution is therefore
where we define , i.e. the Euclidean norm where we treat each element of the matrix as a coordinate. The probability density at a radius therefore corresponds to
where is the number of elements of the matrix , i.e. the number of coordinates. The probability density for high peaks sharply at radius . That means that the bulk of the random initializations of will lie around this radius, and we can therefore visualize such random initialization as assigning a point to a thin spherical shell in the configuration space, as illustrated in Figure 1.
Each matrix is initialized according to its own dimensions and therefore ends up on a shell of a different radius in its respective coordinates. We compensate for this by introducing the metric discussed in Section 2.1. Biases are initialized at zeros. There is a factor fewer biases than weights in a typical fully-connected network, therefore we ignore biases in our theoretical treatment.
3 Hessian
We would like to explain the observed relationship between several sets of quantities involving second derivatives of the loss function on random, low-dimensional hyperplanes and hyperspheres. The loss at a point can be approximated as
The first-order term involves the gradient, defined as . The second-order term uses the second derivatives encapsulated in the Hessian matrix defined as . The Hessian characterizes the local curvature of the loss function. For a direction , implies that the loss is convex along that direction, and conversely implies that it is concave. We can diagonalize the Hessian matrix to its eigenbasis, in which its only non-zero components lie on its diagonal. We refer to its eigenvectors as the principal directions, and its eigenvalues as the principal curvatures.
After restricting the loss function to a -dimensional hyperplane, we can compute a new Hessian, , which comes with its own eigenbasis and eigenvalues. Each hyperplane, therefore, has its own set of principal directions and principal curvatures.
4 Curvature Statistics and Connections to Experiment
We can make sense of these observations using two tools: the fact that high-dimensional multivariate Gaussian distributions have correlation functions that approximate those of the uniform distribution on a hypersphere (?) and the fact (via Wick’s Theorem) that for multivariate Gaussian random variables ,
The variance of the curvature in the -direction is
Therefore, when is large, the principal directions have the same average curvature that randomly-chosen directions do, but with much more variation. This explains why a slight excess of positive eigenvalues can correspond to an overwhelming majority of positive curvatures in random directions.
Because , the principal curvatures on -hyperplanes have standard deviation , while the average principal curvature stays constant. This explains why smaller hyperplanes have a larger excess of positive eigenvalues in the Goldilocks zone than larger hyperplanes (Figure 3(a)): their principal curvatures have the same (positive) average value, but vary less, on hyperplanes of smaller .
5 Radial Dependence of Loss
Therefore, above some critical , the loss will grow as a power law with exponent . Below this critical value, it will be nearly flat. We confirm this empirically, as shown in Figure 5(a). This scaling of the loss does not apply to activations, which have a much more bounded linear regime and do not produce arbitrarily large outputs, which might explain the weaker presence of the Goldilocks zone for them, as seen in Figure 4.
6 Radial Features and the Laplacian
Results and Discussion
An unusually high fraction () of positive eigenvalues of the Hessian at randomly initialized points on randomly oriented, low-dimensional hyperplanes intersecting them. The fraction increased as we optimized within the respective hyperplanes, and decreased with an increasing dimension of the hyperplane, as predicted in Section 2. The effect appeared at a well defined range of coordinate space radii we refer to as the Goldilocks zone, as shown in Figure 3(a) and illustrated in Figure 1.
Common initialization schemes (such as Xavier (?) and He (?)), initialize neural networks well within the Goldilocks zone, precisely at the radius at which measures of local convexity/prevalence of positive curvature peak (see Figures 3 and 4).
Hints that selecting initialization points for high measures of local convexity leads to statistically significantly faster convergence (see Figure 6), and correlates well with low initial loss.
An illustration of the Goldilocks zone and its relationship to the common initialization radius is shown in Figure 1.
Using our observations, we draw the following conclusions:
There exists a thick, hollow, spherical shell of unusually high local convexity/prevalence of positive curvature we refer to as the Goldilocks zone (see Figure 1 for an illustration, and Figures 3 and 4 for experimental data).
When optimizing on a random, low-dimensional hypersurface of dimensionality , the overlap between the Goldilocks zone and the hypersurface is the main predictor of final accuracy reached, as demonstrated in Figure 2.
The small variance between final accuracy reached by optimization constrained to random, low-dimensional hyperplanes is related to the hyperplanes being a) normal to , b) well within the Goldilocks zone, and c) the Goldilocks zone being angularly very isotropic.
Common initialization schemes (such as Xavier (?) and He (?)), initialize neural networks well within the Goldilocks zone, consistently matching the radius at which measures of convexity peak.
Conclusion
We would like to thank Yihui Quek and Geoff Penington from Stanford University for useful discussions.