Comparing Dynamics: Deep Neural Networks versus Glassy Systems

M. Baity-Jesi, L. Sagun, M. Geiger, S. Spigler, G. Ben Arous, C. Cammarota, Y. LeCun, M. Wyart, G. Biroli

Introduction

The training process of a deep neural network (DNN) shares very strong similarities with the physical dynamics of disordered systems: the loss function plays the role of the energy, the weights are the degrees of freedom, and the dataset corresponds to the parameters defining the energy function. The randomness in the data is akin to what is called “quenched disorder” in the physics literature.In statistical physics, the term “quenched” refers to coefficients randomly picked at the preparation of the system and kept constant during its evolution. Training is routinely performed by the Stochastic Gradient Descent (SGD), which consists in starting from random initial conditions and then letting the weights evolve dynamically towards configurations corresponding to low loss values. This process is, in fact, similar to what is called “a quench” in physics. The quenching protocol corresponds to a sudden decrease of the thermal noise, usually done by lowering the temperature of the thermal bath, for a system which is initially prepared in equilibrium at very high temperature. The study of the dynamics induced by quenches has been one of the most important topics of out-of-equilibrium physics of the last decades (Biroli, 2016). The main model considered in the literature is based on stochastic Langevin equations, reminiscent of SGD and corresponding to an evolution governed by gradient descent plus random noise. Since the initial temperature is very high, the initial conditions for the dynamics are random, featureless and uncorrelated with the quenched disorder if present, again in strong analogy with DNNs. Disordered systems are known to display glassy dynamics after a quench, which means that the system gets stuck for long times in local minima (Biroli, 2016; Bouchaud et al., 1998; Berthier & Biroli, 2011; Cugliandolo, 2003). Given the similarity between the training of DNNs and quenching of disordered systems, it may seem surprising that meaningful local minima with perfect accuracy on the training set are found (Zhang et al., 2016).

In the current literature, several explanations are proposed to explain this paradox. Two quite different points of view emerge from it. One is that even though the loss function displays a very large number of local minima with different loss values, the dynamics during the training process allows the system to decrease the loss without barrier crossing and to converge towards quite low local minima that allow good generalization. In other words, the loss landscape is very rough, however, this doesn’t damage the performance of the system. In this direction, (Choromanska et al., 2015) proposed an analogy with mean-field glassy systems. In such systems, it was shown by theoretical physics methods (Cugliandolo & Kurchan, 1993), backed up by rigorous results (Ben Arous et al., 2006), that dynamics corresponding to gradient descent or stochastic versions of it tend without barrier crossing to the widest and the highest minima, despite the existence of deeper local and global minima. A complementary point of view, proposed in (Baldassi et al., 2016), is that there exist rare and wide minima which have large basins of attraction and are reached without any substantial barrier crossing by the training dynamics.

Another quite different point of view is that deep neural networks work in a regime in which there are actually no spurious local minima that can trap the system during the training process. Several rigorous and numerical works, including but not limited to (Freeman & Bruna, 2016; Hoffer et al., 2017; Soudry & Carmon, 2016), suggest that the loss function, despite being non-convex, is characterized by a connected level set as long as one considers loss values above the global minimum. From this perspective, the dynamical evolution induced by the stochastic gradient descent corresponds to falling down in the loss landscape without barrier crossing. In this case, it is the absence of bad local minima, and consequently, the absence of roughness and glassy dynamics, that solves the previous paradox.

Beyond the above two seemingly contradictory pictures on the structure of the loss landscape, there is also a rich literature discussing the path the dynamical process takes during the training process. For instance, Dauphin et al. (2014) claims that it is the existence of numerous saddle points that lie on the dynamical paths that present itself as a form of an obstacle to find deeper local minimum. Several other works, including Lee et al. (2016), claim that gradient-based training avoids such obstacles even if they do exist. And finally, Lipton (2016) demonstrates how the weights travel large distances through the flat basins by looking at the principle components of the evolution of the weights.

Establishing conclusively these scenarios in realistic cases is a challenge. Exact calculations of the statistical properties of critical points are hampered by the increased computational complexity of over-parametrized models and the possible degeneracy of critical points. Some guidance is provided by empirical results. In fact, simulations in Sagun et al. (2014) demonstrate that different dynamical processes on the loss landscape can indeed perform similarly regardless of the effect of the noise of SGD, thus suggesting that barrier crossing indeed does not take place. The works Keskar et al. (2016) and Jastrzebski et al. (2017) claim that by tuning the hyper-parameters of the system one can locate local minima with different qualities, thus providing indications of the roughness of the loss landscape. The results of Chaudhari et al. (2016) demonstrate that wider and possibly rarer basins can be found by averaging out the values of several parallel optimizers.

At the moment, it is still not clear what approach provides a good answer. It could be actually that the correct one contains ingredients from all the perspectives cited above. In this work, we address this problem by taking advantage of knowledge gained in the field of glassy out-of-equilibrium systems in the last decades (Bray, 2002; Biroli, 2016; Bouchaud et al., 1998). Our approach is twofold: (1) probing the training dynamics through the measurement of one and two-point correlation functions, as done in physics, we infer properties of the loss landscape in which the system is evolving, (2) comparing the results obtained for mean-field glasses to measurements performed for realistic DNNs we test the analogy between these systems.

Our Contribution: The analysis is performed for several different architectures, see Sec. 3, varying from specific toy models to ResNets (He et al., 2016) which are evaluated on popular datasets such as MNIST and CIFAR. We decided to focus both on a simple architecture and on more competitive ones. The former is close to a model where, for a large-enough hidden layer, there is a proof of the non-existence of bad local minima (Freeman & Bruna, 2016), and the latter are a relatively more realistic one with relevant performances on the given task. The dynamical behavior we found is similar for all cases: After an initial exploration of high-loss configurations, the system starts its descent in the “loss landscape”, and displays a particular kind of glassy dynamics, called aging, see Sec. 2. Our results suggest that the slowness of the dynamics in this stage is not related to the crossing of large barriers but instead to the emergence of an increasingly large number of flat directions (Sagun et al., 2017). At long times, a stationary regime where aging is interrupted and the system becomes almost stationary sets in. We present evidences that this dynamical regime corresponds to diffusion, not necessarily isotropic (as suggested by (Jastrzebski et al., 2017)), at or close to the bottom of the loss landscape. We compare these behaviors to the ones of the pp-spin spherical model, which is one of the most studied mean-field glass models. We find that although the first regimes share similarities with the dynamics of mean-field glasses after a quench, the final regime does not. This suggests a qualitative different geometrical characterization of the bottom of the loss landscape and, accordingly, of the dynamics within it.

Basic facts on glassy dynamics

Two main observables have been identified as central to characterize the slow dynamics of physical systems. The first one is the energy as a function of time. When a system is quenched from high to low temperature the energy decreases and slowly approaches an asymptotic value. The functional dependence can be a power law of time, as in the Ising model (Bray, 2002), or even a power of the logarithm of time as in several disordered systems, in particular glasses (Berthier & Biroli, 2011). This dependence is called “slow” by comparison with an exponential relaxation which is typical of high-temperature phasesThe existence of conserved quantities can produce a power-law dependence even in high-temperature phases.. In Figure 1(a) we show the characteristic behavior of the energy as a function of time for a quench from high to low temperatures in the pp-spin spherical model, which was highlighted in the context of DNNs through an analogy in (Choromanska et al., 2015) and through phenomenological comparison in (Sagun et al., 2014). The degrees of freedom of the pp-spin model are σi\sigma_{i}, the NN components of a vector belonging to the NN-dimensional sphere of radius N\sqrt{N}. Its energy reads for p=3p=3:

where the sum runs over all the possible 33-tuples and Ji1,i2,i3J_{i_{1},i_{2},i_{3}} are i.i.d. Gaussian random variables with zero mean and variance 3/N23/N^{2}. The dynamical evolution is governed by the stochastic Langevin equation. This model has a dynamical transition at a temperature Td≃0.612T_{d}\simeq 0.612, see (Castellani & Cavagna, 2005) for a review. The plot in Figure 1(a) corresponds to a quench from Ti=∞T_{i}=\infty to Tf=0.5T_{f}=0.5; it is obtained by integrating numerically the Cugliandolo-Kurchan equations (Cugliandolo & Kurchan, 1993; Ben Arous et al., 2006).

Slow dynamics and aging are distinctive features of any glassy system. Particularly, in the pp-spin spherical model, and in other models of glasses, the slow dynamics observed after a quenchThis dynamical regime corresponds to large time-scales that do not diverge with NN. There is a second regime of time-scales, that diverge exponentially with the number of degrees of freedom (Montanari & Semerjian, 2006; Ben Arous & Jagannath, 2017), in which barrier crossing does take place. In practice, except for small systems (Baity-Jesi et al., 2018), this second regime cannot be accessed numerically since the corresponding time-scales are too big. is not due to barrier crossing but to the emergence of almost flat directions (Castellani & Cavagna, 2005). As explained in (Kurchan & Laloux, 1996), this phenomenon is due to the peculiarity of gradient descent in very high-dimensions; in this case the system is always confined at the border of the basins of attraction, and the Hessian at long times contains a decreasing number of negative eigenvalues, thus leading to an increasingly slow dynamics.

Models and Results

We present our core results in two parts: time dependence of the loss function (Sec. 3.1), and identifying different regimes through the two-point correlation function (Sec. 3.2). We start by describing the models used for evaluation: We did not remark any significant difference in the presence of explicit regularization, so we present the results where no regularization is used.

A - Toy Model: The network contains only 1 hidden layer with 10410^{4} hidden nodes. The non-linear function on the hidden layer is ReLU. The output layer is filtered through a sigmoid. The loss function is a mean square error. The total number of weights is around 3×1083\times 10^{8}.

B - Fully Connected: A simple network with three fully connected layers, of sizes 100, 100 and 10, respectively. The non-linear functions are ReLUs, and the loss function is the negative log-likelihood of soft-max outputs. The total number of weights is about 9×1049\times 10^{4}.

C - Small Net: A simple convolutional network with two conv-layers that has 10 and 20 filters in the first and second layer, respectively. It is followed by two fully-connected layers of sizes 100 and 10. The non-linear functions in the hidden layers are ReLUs, and the loss function is the negative log-likelihood of soft-max outputs. The total number of weights is around 6×1046\times 10^{4}.

D - ResNet18: The final model is a ResNet with 18 hidden layers. The total number of weights is around 2×1072\times 10^{7}.

We have chosen networks with various levels of complexity. All networks are initialized in the standard procedures of the PyTorch library (version 0.3.0). The toy model is inspired by the one introduced in (Freeman & Bruna, 2016) which is shown not to have any barriers if the hidden layer is large enough. The training is carried out by SGD that takes a single learning rate that remains unchanged until the end of the computation. The training process runs for a fixed given number of iterations which is deemed to be ‘long enough’ for all practical purposes. For most cases, this means that training kept running long after the perfect accuracy was reached on the training set. All the networks have been trained on multiple datasets: MNIST, CIFAR-10, CIFAR-100, and multiple sets of parameters.

We first focus on the time-dependence of the loss function over the training, and we compare it to the one of the energy in glassy systems. For the sake of completeness, we also show the accuracy. We plot the loss values as a function of the logarithm of time, measured in units of iterations so that the unit time step corresponds to a single update of the weights. This choice is different from the wall time or number of epochs which is often used. Although less common in machine learning, the logarithmic scale highlights the slow dynamics and the time dependenceA positive side effect of a logarithmic representation is that the measurements can be exponentially spaced. As a consequence, the numerical overhead of the measurements goes to zero as the simulation time increases. Since the relevant time scales are logarithmic, this implies no loss of information.. The results obtained for the four networks described above are shown in Figures 2(a), 2(b), 2(c), 2(d). There are several features worth noticing. We can remark three regimes. The first one goes from the beginning of the training up to a time t1t_{1}, where the loss and accuracy stay roughly constant. At t=t1t=t_{1} the loss starts decreasing roughly linearly in log⁡(t)\log(t), and concomitantly the accuracy increases in a similar way. This second regime persists until a time t2t_{2}, at which the train loss approaches zero. In the final regime beyond t2t_{2} the speed of decay sharply decreases. The cross-over times t1t_{1} and t2t_{2} are indicated in Figures 2(a), 2(b), 2(c), 2(d). In Sec. 3.2 we show that t1t_{1} and t2t_{2} can also be identified through the evolution of the mean-square displacement.

This behavior is similar to the ones found in disordered systems, see e.g. Figure 1(a). There are however two main differences. First, in several cases the decrease in the second part is actually slower for the DNNs compared to the power-law of the pp-spin modelThe power law decrease of the energy was established in (Cugliandolo & Kurchan, 1993) and is well verified numerically.. Second, and more importantly, the loss reaches asymptotically (i.e. after t2t_{2}) its lowest possible value. This is not the case in the pp-spin model in which instead the energy converges asymptotically to one of the highest and widest minima (Cugliandolo & Kurchan, 1993; Castellani & Cavagna, 2005). Actually, a pp-spin model with a number of degrees of freedom comparable to the number MM of weights that are used in deep learning (in our examples M=104−107M=10^{4}-10^{7}) would take an exponentially long time to go beyond the highest and widest minima and reach the bottom of the landscape (Castellani & Cavagna, 2005; Berthier & Biroli, 2011). This is a first indication that the dynamics involved in the training of deep neural networks, although slow, does not correspond to the crossing of large barriers, which would instead lead to much longer time-scales.

In summary, the reason for the slowing down of the dynamics during training is apparently not due to barrier crossing but instead likely related to an increasingly large amount of flat directions that become available to the system during its descent in the loss landscape, as found numerically in (LeCun et al., 1998; Sagun et al., 2017). This is actually similar in the pp-spin spherical model to the first dynamical regime of aging dynamics that follows a quench. However, in this case the system does not reach the lowest possible values of the loss, as it happens to loss functions during training, but remains trapped in higher and wide local minima.

2 Further evidence: Two-time correlation functions

where the sum runs over all the weights wiw_{i} of the network, and MM is their total number.

To a large extent, the training dynamics at large times can be explained in terms of diffusion in the weight space. A hallmark of a diffusing system is a motion purely driven by the noise DD (Crank, 1979). We estimate the noise in SGD with the variance of the loss function’s gradientFor reasons of numerical efficiency, for some models DD is calculated on a (sufficiently large) subset of the training set., which reads (details on the definition of the noise can be found in several resources, see, for example, Li et al. (2015)):

Both the aging and the diffusive regimes are present and qualitatively similar in all the analyzed networks. The fact that a slow aging dynamics is also present in model A (Toy Model), that supposedly has no barriers (see Sec. 3), strengthens the conclusion that the dynamics slows down because of the emergence of flat directions that ultimately lead to diffusion at or close to the bottom of the landscape. A deeper analysis of the finer properties of the diffusive regime will be studied in a forthcoming publication.

Discussion

In this work we have analyzed the training dynamics of DNNs by methods developed in physics for out-of-equilibrium disordered systems. We have studied the time dependence of the loss value and the mean-square displacements of the weights and compared them to their counterparts in physical systems, in particular the 3-spin spherical spin-glass. The analysis of the time-dependence of the loss function and the mean square displacement indicates that there are at least three time regimes in the training process: one corresponding to an initial exploration of the energy/loss landscape, followed by a decrease of the loss, in which the system displays aging dynamics, and a final regime in which the dynamics appears to be almost stationary and diffusive. Barrier crossing does not seem to play any role. The slowing down can be instead traced back to an increasingly large amount of flat directions that become available to the system during its descent in the loss landscape.

The non-existence of such barrier crossings has been already proposed in the machine learning literature and some indirect evidences where obtained in numerical works. In (Freeman & Bruna, 2016), it is shown that in certain networks one can connect two different solutions by a path in the weight space in such a way that the loss doesn’t increase by much, and the amount of increase diminishes as the size of the network grows. In a related perspective on the loss surface, (Sagun et al., 2016) and (Sagun et al., 2017) demonstrate separate cases where the straight line between two weight configurations at the bottom of the loss landscape evaluates to the same loss value, in other words there are no barriers between these two points.

On the basis of these results, we conjecture the existence of a phase transition between two regimes: (i) an easy phase corresponding to over-parametrized networks, in which bad local minima do not play any role, dynamics is governed by a massive amount of flat directions, and learning is achieved; (ii) a hard phase corresponding to under-parametrized networks, in which the landscape is rough, dynamics is glassy and the network does not learn well. Whether learning is possible in this case but it would take a huge amount of time to find the good minima is an interesting question.

This scenario has tantalizing similarities with the one found in several combinatorial optimization problems in which easy, hard and impossible algorithmic phases have been found, see e.g. (Monasson et al., 1999; Mézard et al., 2002; Krz̧akała et al., 2007; Zdeborová & Krzakala, 2016; Achlioptas & Coja-Oghlan, 2008). When degrees of freedom are continuous, the transition between these phases can be associated with the emergence of many flat directions in the energy landscape, a well-known example is the jamming transition of disordered solids (Wyart, 2005; Liu et al., 2010). A detailed investigation of this scenario for DNNs is ongoing and will be presented in a future publication.

Acknowledgements

We thank Valentina Ros for useful conversations. We thank Utku Evci and Uğur Güney for providing the initial version of the code that we used in our numerical simulations. This work was partially supported by the grant from the Simons Foundation (♯\sharp454935 Giulio Biroli, ♯\sharp454953 Matthieu Wyart, ♯\sharp454951 David Reichman). M.W. thanks the Swiss National Science Foundation for support under Grant No. 200021-165509. M.B.-J. was partially supported through Grant No. FIS2015-65078-C2-1-P, jointly funded by MINECO (Spain) and FEDER (European Union). CC acknowledges support from the King’s Worldwide Partnership Fund.

References