The Future is Log-Gaussian: ResNets and Their Infinite-Depth-and-Width Limit at Initialization
Mufan Bill Li, Mihai Nica, Daniel M. Roy
Introduction
The characterization of infinite-width dynamics of gradient descent (GD) in terms of the so-called Neural Tangent Kernel (NTK) represented a major breakthrough in our understanding of deep learning in the large-width regime. Before the identification of infinite-width limits, the theoretical study of deep learning had long been hindered by the apparent analytical intractability of gradient descent and variants acting on the nonconvex objectives used to train neural networks. Despite this progress, evidence suggests that deep neural networks can outperform their infinite-width limits in practice , particularly when the depth of the network is large. These observations motivate the study of other approximations that may close the gap.
Several alternative limits have been proposed. Around the time of the discovery of the NTK limit, mean-field limits were also characterized , and more recently have been linked with the NTK limit . describe a family of infinite-width limits indexed by the scaling limits of initial weight variance, weight rescaling, and learning rates. This family includes both the NTK and mean field limits. One motivation for studying these alternative limits is that they yield a notion of feature learning, which provably does not occur in the NTK limit .
Despite the variety of these limits, one common feature is that the depth of the network (i.e., the number of its layers) is treated as a constant as the width of the network is allowed to grow. Indeed, for fixed width, , the Gaussian approximations at initialization worsen as the depth, , increases. While real-world networks are fairly wide, their relative depth is not trivial. were the first to compute an infinite-depth-and-width limit for fully connected networks. While the training dynamics of this limit are still not completely understood, we now know that, in the infinite-depth-and-width limit, the neural tangent kernel is random, and the derivative of the kernel is nonzero at initialization. This means training does not correspond to that of a linear model, like it does in the NTK limit .
In addition to these theoretical corrections to the Gaussian limit, practitioners have begun to notice that even basic properties of standard neural networks do not line up with those predicted by the infinite-width limit. Perhaps the most basic is that the gradients of the network early in training are not Gaussian, but instead are approximately log-Gaussian . In fact, one should see log-Gaussian behaviour agrees with the theoretical predictions of , although there has yet to be a careful empirical comparison made between the precise predictions coming from infinite-depth-and-width models and real world networks.
In practice, however, fully connected networks are not often used without architectural modifications. In particular, residual connections, ushered into widespread use after the description of the ResNet architecture , produced very deep architectures that were practically useful with optimization techniques available at the time. Initialization schemes for ResNets have been studied in the infinite-width limit or with modifications .
In this work, we consider the infinite-depth-and-width limit of fully connected architectures with residual connections (“Vanilla ResNets”) and standard initializations. Here, the analysis is complicated by the effect of skip connections, which introduce interlayer correlations that have a non-negligible effect in this limit. Surprisingly, we observe a counter-intuitive but fundamental phenomenon, whereby these skip connection cause the network to be hypoactivated, meaning that less than half of the neurons are activated on initialization. This fact undermines key assumptions that underlie other infinite-depth-and-width limit studies of architectures without residual connections or with nonstandard modifications, such as post-activation residual connections.
Hypoactivation is a roadblock that all theoretical research into standard ResNet architectures must contend with—it is an unavoidable property of the standard architecture and the root of many technical difficulties. In order to sidestep this roadblock, we introduce a conjecture that bounds the size of hypoactivation and effect on interlayer correlation. The conjecture is inspired by empirical evidence from Monte Carlo simulations, as well as several simplified analyses that ignore certain technical difficulties. The conjecture introduces what we believe to be the minimal assumption necessary to allow rigorous theoretical work to proceed. It also defines an important open problem in the study of limits of residual architectures.
To demonstrate the utility of the conjecture, we show that it leads to precise predictions. In particular, we prove a limit theorem characterizing the exact marginal distribution of the output at initialization in the infinite-depth-and-width limit, up to order . Our limit result shows that ResNets have log-Gaussian behaviour on initialization, and like fully connected networks , the behaviour is determined by the depth-to-width aspect ratio. This corroborates recent empirical observations about deep ResNets . Since real world networks are finite, the question of how well this approximates finite behaviour is of paramount importance. Based on Monte Carlo simulations, we find excellent agreement between our predictions and finite networks (see Figure 1). Moreover, for very deep networks (e.g. ) the infinite-depth-and-width prediction is extremely different than the infinite-width prediction. More surprisingly, however, is that even at comparatively small depth-to-width ratio (e.g. ) the two limits are already significantly different. Furthermore, we also observe that the effects due to hypoactivation and interlayer correlation are non-negligible; these effects are precisely the difference between Vanilla ResNets and so-called “Balanced Resnets” in Figure 1. Perhaps most importantly, we observe that real network outputs exhibit exponentially larger variance than predicted by infinite-width limits. This type of variance at initialization is known to cause exploding and vanishing gradients and other types of training failures . Our result also implies that the output neurons of the network are not independent as predicted by the Gaussian infinite-width limit. See Figure 3.
In order to maintain the same skip connections from layer to layer, but render the activation patterns completely independent from layer to layer on initialization, we introduce the Balanced ResNet architecture, where the sign of each neuron’s activation function is randomized. We demonstrate that this exponentially decreases the variance of the network output on initialization. Moreover, this independence between neurons also makes the model more amenable to theoretical analysis and opens the door to future understanding of network behaviour. Although it is beyond the scope of this paper, a small preliminary empirical study (Appendix C) suggests that standard training regimes are not negatively affected by replacing standard ResNet architectures with Balanced ones.
We summarize our main contributions as follows:
We identify and characterize a fundamental property of ResNets that we call hypoactivation: less than half of the ReLU neurons are activated. Based on empirical evidence, we formulate a precise minimal conjecture bounding the effect of hypoactivation that permits us to make precise, rigorous estimates for other properties of ResNets.
We prove a limit theorem which shows that the output of ResNets on initialization exhibits log-Gaussian behaviour with parameterization depending on the depth-to-width ratio .
We provide empirical evidence from Monte Carlo simulations showing our theory provides more accurate predictions for simple properties of finite networks compared to the predictions made by the Gaussian infinite-width limit.
We introduce the Balanced ResNet architecture, which corrects the hypoactivation and variance due to layerwise correlation from Vanilla ResNets. We also prove that the output for this architecture is log-Gaussian on initialization with exponentially lower variance. This simple modification can be applied to any neural network that uses ReLU activations.
Main Results
In terms of the notation in Table 1, a Vanilla ResNet with fully connected first/last layers and hidden layers of width is defined by
Note that factors of in the hidden layer are equivalent to intializing according to the so-called He initialization . Other intializations correspond to changing the coefficient . This setup is similar to that of “Stable ResNets” , where the infinite-width limit is studied.
At the same time, we also find the covariance between the activations of various layers does not vanish in the infinite-depth-and-width limit. This motivates the definition of the total interlayer covariance correction
The conjecture is well supported by Monte-Carlo simulations (see Figure 4). We provide a more detailed discussion and a precise statement of 5 in Section 4. Assuming the conjecture holds, we prove a limit theorem about the distribution of . Informally, this says that is approximately a log-Gaussian scalar times an independent Gaussian vector
where has iid entries, and and are defined by
The precise statement, including asymptotic error bounds, is as follows.
For any choice of hyperparameters , and every input , the output at initialization has a marginal distribution which can be written in the form
where and are as in (5), and moreover converges in distribution to a Gaussian random variable with mean and variance given by (7) in this limit.
The main ideas of the proof of 1 is given in Section 5 and the detailed proof is given in Appendix B. We also provide more explicit formulas for below.
Assume 5 is true. Then in the same infinite-depth-and-width limit as 1, the total hypoactivation and total interlayer covariance obey
Here is a constant depending on and
where first appeared in , and is such that .
The Gaussian infinite-width limit predicts that the marginals of should have the form of (6) with being identically zero. As depends exponentially on , the infinite-depth-and-width limit predicts the variance is exponentially larger than the infinite-width limit. See Section 3.1 for a detailed discussion and Figure 3 for verification against finite networks.
Using Monte Carlo simulations, we estimate the constant . The result of 1 is then compared against finite networks in Figure 2. Full proof can be found Appendix B.
The idea of using a scaling parameter in the skip connections has been noted in empirical papers and also studied under simplified assumptions on the number of activation in each layer . Our result shows that by choosing , the prefactor in (6) does not grow with depth thereby enabling deeper networks to be trained.
For a balanced ResNet, the same result as 1 given in (6) still holds, but with the mean and variance of given simply by
Consequences of Theorems 1 & 4 and Comparison to Infinite-Width Limit
By the basic fact it follows from Theorems 1 & 4 that, when the inputs has , the mean size scale of any neuron is approximately
(Note that the terms with cancel out!) When , this is constant for Balanced ResNets. In contrast, Vanilla ResNets have a complicated dependence on the network depth and width due to the hypoactivation and correlations terms. This means the behaviour is , which is somewhat surprising. A more serious issue is the variance which Theorems 1 & 4 predict to be
Balanced ResNets suffer from this variance problem less because the interlayer correlation term is zero. Since this variance reduction happens at the exponential scale in, the difference can be significant; for networks with , the contribution is a factor of for Vanilla ResNets vs. for Balanced ResNets. See Figure 3 for a comparison of these theoretically predicted properties vs experiments with finite networks.
2 Correlated Output Neurons
Since the same random variable multiplies the entire vector , the individual neurons in the output layer are not independent. For example, Theorem 1 and 4 predict that the squared entries have strictly positive correlation given by for any two neurons where . This tends to as grows. The effect of correlated output neurons persists for Balanced ResNet but is reduced again due to the lower variance. This is very different from the infinite-width limit, which predicts that individual neurons should be independent Gaussians. This prediction of the theorem matches finite networks closely; see Figure 3.
Conjecture 5: Hypoactivation and Layerwise Correlations
Proof Ideas for Theorems 1 & 4
A key element of the proof is the following property of Gaussian random matrices. If which has iid entries, then for any vector , we have
where is a vector whose entries are iid random variables. Because of the fully connected first and last layer of the network, (14) implies that
Hence only depends on . (Equivalently, has the distribution of when .) With this definition, (15) also shows is proportional to , establishing the first part of 1.
Taking the of (17) exhibits as a sum of these weakly correlated random variables. Here we note that various tail estimates for the same or related quantities have been developed , however these estimates are not precise enough to pinpoint the exact limiting distribution. In contrast, we are able to derive the exact limiting distribution via a Central Limit Theorem (CLT) for weakly correlated sums . The proof of 1 is completed by computing the mean, variance and covariance of terms using 5. For 4, the final calculation is simplified by (10) which shows the terms are uncorrelated. The detailed proof is given in Appendix B.
Acknowledgement
We would like to thank Blair Bilodeau, Gintare Karolina Dziugaite, Mahdi Haghifam, Yani A. Ioannou, James Lucas, Jeffrey Negrea, Mengye Ren, and Ekansh Sharma for helpful discussions and draft feedback. We would also like to thank the anonymous NeurIPS reviewers for insightful feedback. In particular, one identified numerous relations to existing work, and another helped us identify the uniformity requirement in 5. ML is supported by Ontario Graduate Scholarship and the Vector Institute. MN is supported by an NSERC Discovery Grant. DMR is supported in part by an NSERC Discovery Grant, Ontario Early Researcher Award, and a stipend provided by the Charles Simonyi Endowment.
References
Appendix A Appendix
We stated our main result with fixed , but the result easily extends to the case where vary from layer to layer. This allows comparison between our result and the infinite width limits , where they have and allow to varying layer by layer. The statement in that setting is modified as follows.
Suppose are sequences such that is uniformly bounded away from . Define the network by
Consider the limit where both the network depth and hidden layer width in such a way that the ratio converges to a constant. In this limit, the distribution of the network output for a given input is given by
The corresponding limit theorem for balanced ResNets also holds; the hypoactivation and layer-wise covariance term in (20) vanish.
The behaviour is a complicated function of the sequences . It would be interesting to use these theoretical results to guide the choice of parameters and investigate how this effects training behaviour.
Appendix B Proof of main results
To simplify the exposition of the proofs, we will assume without loss of generality that . The general case can be reduced to the case by dividing by in each layer and rescaling the parameters to and .
Taking expectation of both sides of (22), we have
where we have used in the last line. ∎
B.2 Variance Calculation
We will use the decomposition and compute the the two terms individually.
In the second term, since and , we have .
To compute , notice that is a sum of two terms. The two terms are uncorrelated since . Hence, the expectation of the cross term in is zero and we can compute
We have used the fact about random variables that . Finally then:
This can be calculated directly using properties of the unit sphere, but the proof is complicated by the fact that the entries are not independent. Instead, there is an elementary proof using the following equality in distribution:
Taking norm and expectation of this equality in distribution gives . Since and are independent, we can factor and rearrange to obtain
Since the entries of are independent of each other, it is easier to compute using and this identity instead of using . Moreover, because Gaussian distribution are symmetrically distributed, we have that is independent of, Hence:
and so . Similarly, using we now compute by looking at diagonal and off-diagonal terms as follows
Using , we finally obtain from which the claimed variance formula follows. ∎
B.3 Uniform distribution on spheres and Gaussian random variables
This is a direct corollary to Theorem 2 of , which uses Stirling’s formula to show that the total variation distance between the random variables and is at most . ∎
The result then follows by using the independence of and , and the formula for the -th moment of random variable, namely . ∎
The proof is immediate writing the difference in expectation as an integral and then comparing to by the results of the previous two lemmas. ∎
B.4 Pairwise covariances
By expanding the norms into sums, , we can compute the covariance by summing over all pairs of coordinates . There are two types of terms to consider. (Note: we use the notation .)
Diagonal Terms\contourwhiteDiagonal Terms: for We first use the positive homogeneity of to extract a factor of from both terms
We now use the approximation of by as in Section B.3 to obtain
where we have used the definition of from (23).
Off diagonal terms\contourwhiteOff diagonal terms for ,,
As in the calculation for the diagonal term, we now again use the approximation that is approximately marginally distributed like and the positive homogoneity of to obtain
Renaming to to match the notation of (23), we finally obtain
Summing the diagonal and off-diagonal terms, we find the total covariance is:
which gives the claimed formula for the covariance. ∎
This follows immediately from the approximation for the covariance in 5 and 16. ∎
By the series expansion for as , it follows that as . This exponential decay in explains why the total covariance correction remains even as as .
B.5 Central Limit Theorem
where the constant in the big notation is uniform in . Moreover, is asymptotically Gaussian in the infinite width and depth limit.
from which the desired mean and variance formula for follows.
The proof is analogous to the proof of 1; the calculation for the activation of the layers is simplified by the fact that which neurons activated in each layer are uncorrelated by (10). ∎
B.6 Input-Output Gradient for Balanced ResNets
For any input , the output of a Balanced ResNet can be written as
The derivative of with respect to any input has the distribution
This follows immediately since is constant on the neighbourhood . ∎
has the distribution as the output at any input with .
By the previous results, both are equal in distribution to for any unit vector , where is the distribution of the random matrix in (33). ∎
Appendix C Experiments
Throughout the paper, the Monte Carlo simulations were computed on a single NVIDIA Titan-XP GPU. The main tools used in the neural network simulations are the JAX library (Apache 2.0 License) and the PyTorch library (BSD 3-Clause License). Furthermore, the Python libraries numpy (BSD 3-Clause License), plotnine (based on ggplot2 , GNU GPLv2 License), and pandas (BSD 3-Clause License) tremendously helpful. We used Python version 3.6.8 from Anaconda 3 (3-clause BSD License) and Jupyter notebook (3-Clause BSD License).
These experiments, displayed in Figure 6, investigate the difference between training performance of the Vanilla ResNet from (1) and the Balanced ResNet defined in Section 2.1. This is beyond the theory proven in our work which concerns statistical properties of the network on initialization. Both of these architectures are fully-connected networks with skip connections between layers. We observe that both architectures perform similarly in standard training regimes where the networks are much wider than they are deep.
C.2 Convolutional ResNet CIFAR-10 Experiments
These experiments investigate how the Balanced ResNet architecture modification, namely randomly flipping between and , effects training for deep convolutional ResNets (We call these C-ResNets here to distinguish from the fully connected ones studied in detail in the paper). Although the theory in this paper only are proven for Vanilla ResNets (which are fully connected with skip connections), we expect that the Balanced ResNet tweak does not negatively effect performance in standard training regimes and may allow better initialization for very deep models.
For this experiment, we used the ResNet implementations from the GitHub repository by (MIT License). To create a Balanced Convolutional ResNet (Balanced C-ResNet), we modified the class PreActBlock to add flipped ReLUs into the network channel-wise, thinking of channels as the natural generalization of neurons for convolution neural networks. More specifically, we are interested in an input tensor of dimension , where is for batch size, is for channel, is for height, and is for width. Before feeding into a ReLU non-linearity, we will multiply by a vector of iid random signs , so that the ReLU output is
The experiment results are reported in Table 1 and Figure 7. Here, we demonstrate that using the Balanced ResNet architecture does not reduce the performance of existing training regimes. For the deeper ResNet101, the Balanced ResNet seems to perform slightly better early in training but finishes worse at the end of training.
Note that even the deeper ResNet101 tested here is still relatively shallow compared to its “width”; most of the layers are either or channels by neurons per channel which represent many more neurons in each hidden layer than the depth of the network, . The possible advantage of the Balanced ResNet idea is that it will enable the training architectures which are even deeper compared to their width, which are currently not trainable or difficult to train due to initialization issues. A detailed empirical study exploring this idea is needed to study how Balanced ResNets perform beyond the theory proven in this paper.
C.3 Density Plot Calculations
In this section, we describe the calculations required for plotting Figure 1. Firstly, we need to estimate the hypoactivation constant from 2 using Monte Carlo simulations. For the choice of , we find the constant , which we use for estimating the mean. See Figure 8 for further simulations with varying values, and Figure 9 for simulations demonstrating these constants provide accurate prediction for mean and variance.
Next, using the choice of and , we can write
Here we observe that , and therefore we can compute the density of with a coordinate change
To compute the density of the infinite width prediction, we first observe that in this limit . Therefore it’s sufficient to simply compute the variance.
To this goal, we follow the calculations of with the variance recursion formula (slightly modified to include )
C.4 Additional Monte Carlo Simulations
These additional Monte Carlo simulations provide further comparisons between the infinite depth-and-width limit predictions and finite networks. See Figure 9.