The Future is Log-Gaussian: ResNets and Their Infinite-Depth-and-Width Limit at Initialization

Mufan Bill Li, Mihai Nica, Daniel M. Roy

Introduction

The characterization of infinite-width dynamics of gradient descent (GD) in terms of the so-called Neural Tangent Kernel (NTK) represented a major breakthrough in our understanding of deep learning in the large-width regime. Before the identification of infinite-width limits, the theoretical study of deep learning had long been hindered by the apparent analytical intractability of gradient descent and variants acting on the nonconvex objectives used to train neural networks. Despite this progress, evidence suggests that deep neural networks can outperform their infinite-width limits in practice , particularly when the depth of the network is large. These observations motivate the study of other approximations that may close the gap.

Several alternative limits have been proposed. Around the time of the discovery of the NTK limit, mean-field limits were also characterized , and more recently have been linked with the NTK limit . describe a family of infinite-width limits indexed by the scaling limits of initial weight variance, weight rescaling, and learning rates. This family includes both the NTK and mean field limits. One motivation for studying these alternative limits is that they yield a notion of feature learning, which provably does not occur in the NTK limit .

Despite the variety of these limits, one common feature is that the depth of the network (i.e., the number of its layers) is treated as a constant as the width of the network is allowed to grow. Indeed, for fixed width, nn, the Gaussian approximations at initialization worsen as the depth, dd, increases. While real-world networks are fairly wide, their relative depth is not trivial. were the first to compute an infinite-depth-and-width limit for fully connected networks. While the training dynamics of this limit are still not completely understood, we now know that, in the infinite-depth-and-width limit, the neural tangent kernel is random, and the derivative of the kernel is nonzero at initialization. This means training does not correspond to that of a linear model, like it does in the NTK limit .

In addition to these theoretical corrections to the Gaussian limit, practitioners have begun to notice that even basic properties of standard neural networks do not line up with those predicted by the infinite-width limit. Perhaps the most basic is that the gradients of the network early in training are not Gaussian, but instead are approximately log-Gaussian . In fact, one should see log-Gaussian behaviour agrees with the theoretical predictions of , although there has yet to be a careful empirical comparison made between the precise predictions coming from infinite-depth-and-width models and real world networks.

In practice, however, fully connected networks are not often used without architectural modifications. In particular, residual connections, ushered into widespread use after the description of the ResNet architecture , produced very deep architectures that were practically useful with optimization techniques available at the time. Initialization schemes for ResNets have been studied in the infinite-width limit or with modifications .

In this work, we consider the infinite-depth-and-width limit of fully connected architectures with residual connections (“Vanilla ResNets”) and standard initializations. Here, the analysis is complicated by the effect of skip connections, which introduce interlayer correlations that have a non-negligible effect in this limit. Surprisingly, we observe a counter-intuitive but fundamental phenomenon, whereby these skip connection cause the network to be hypoactivated, meaning that less than half of the neurons are activated on initialization. This fact undermines key assumptions that underlie other infinite-depth-and-width limit studies of architectures without residual connections or with nonstandard modifications, such as post-activation residual connections.

Hypoactivation is a roadblock that all theoretical research into standard ResNet architectures must contend with—it is an unavoidable property of the standard architecture and the root of many technical difficulties. In order to sidestep this roadblock, we introduce a conjecture that bounds the size of hypoactivation and effect on interlayer correlation. The conjecture is inspired by empirical evidence from Monte Carlo simulations, as well as several simplified analyses that ignore certain technical difficulties. The conjecture introduces what we believe to be the minimal assumption necessary to allow rigorous theoretical work to proceed. It also defines an important open problem in the study of limits of residual architectures.

To demonstrate the utility of the conjecture, we show that it leads to precise predictions. In particular, we prove a limit theorem characterizing the exact marginal distribution of the output at initialization in the infinite-depth-and-width limit, up to order O(dn−2)O(dn^{-2}). Our limit result shows that ResNets have log-Gaussian behaviour on initialization, and like fully connected networks , the behaviour is determined by the depth-to-width aspect ratio. This corroborates recent empirical observations about deep ResNets . Since real world networks are finite, the question of how well this approximates finite behaviour is of paramount importance. Based on Monte Carlo simulations, we find excellent agreement between our predictions and finite networks (see Figure 1). Moreover, for very deep networks (e.g. d/n=1d/n=1) the infinite-depth-and-width prediction is extremely different than the infinite-width prediction. More surprisingly, however, is that even at comparatively small depth-to-width ratio (e.g. d/n=0.1d/n=0.1) the two limits are already significantly different. Furthermore, we also observe that the effects due to hypoactivation and interlayer correlation are non-negligible; these effects are precisely the difference between Vanilla ResNets and so-called “Balanced Resnets” in Figure 1. Perhaps most importantly, we observe that real network outputs exhibit exponentially larger variance than predicted by infinite-width limits. This type of variance at initialization is known to cause exploding and vanishing gradients and other types of training failures . Our result also implies that the output neurons of the network are not independent as predicted by the Gaussian infinite-width limit. See Figure 3.

In order to maintain the same skip connections from layer to layer, but render the activation patterns completely independent from layer to layer on initialization, we introduce the Balanced ResNet architecture, where the sign of each neuron’s activation function is randomized. We demonstrate that this exponentially decreases the variance of the network output on initialization. Moreover, this independence between neurons also makes the model more amenable to theoretical analysis and opens the door to future understanding of network behaviour. Although it is beyond the scope of this paper, a small preliminary empirical study (Appendix C) suggests that standard training regimes are not negatively affected by replacing standard ResNet architectures with Balanced ones.

We summarize our main contributions as follows:

We identify and characterize a fundamental property of ResNets that we call hypoactivation: less than half of the ReLU neurons are activated. Based on empirical evidence, we formulate a precise minimal conjecture bounding the effect of hypoactivation that permits us to make precise, rigorous estimates for other properties of ResNets.

We prove a limit theorem which shows that the output of ResNets on initialization exhibits log-Gaussian behaviour with parameterization depending on the depth-to-width ratio d/nd/n.

We provide empirical evidence from Monte Carlo simulations showing our theory provides more accurate predictions for simple properties of finite networks compared to the predictions made by the Gaussian infinite-width limit.

We introduce the Balanced ResNet architecture, which corrects the hypoactivation and variance due to layerwise correlation from Vanilla ResNets. We also prove that the output for this architecture is log-Gaussian on initialization with exponentially lower variance. This simple modification can be applied to any neural network that uses ReLU activations.

Main Results

In terms of the notation in Table 1, a Vanilla ResNet with fully connected first/last layers and dd hidden layers of width nn is defined by

Note that factors of 2n−1\sqrt{2n^{-1}} in the hidden layer are equivalent to intializing according to the so-called He initialization . Other intializations correspond to changing the coefficient λ\lambda. This setup is similar to that of “Stable ResNets” , where the infinite-width limit is studied.

At the same time, we also find the covariance between the activations of various layers does not vanish in the infinite-depth-and-width limit. This motivates the definition of the total interlayer covariance correction

The conjecture is well supported by Monte-Carlo simulations (see Figure 4). We provide a more detailed discussion and a precise statement of 5 in Section 4. Assuming the conjecture holds, we prove a limit theorem about the distribution of zoutz^{\text{out}}. Informally, this says that zoutz^{\text{out}} is approximately a log-Gaussian scalar times an independent Gaussian vector

where Z⃗\vec{Z} has iid N(0,1)\mathcal{N}(0,1) entries, and β\beta and cc are defined by

The precise statement, including asymptotic error bounds, is as follows.

For any choice of hyperparameters nin,nout,n,d,α,λ{n_{\text{in}}},{n_{\text{out}}},n,d,\alpha,\lambda, and every input xx, the output zoutz^{\text{out}} at initialization has a marginal distribution which can be written in the form

where β\beta and cc are as in (5), and moreover GG converges in distribution to a Gaussian random variable with mean and variance given by (7) in this limit.

The main ideas of the proof of 1 is given in Section 5 and the detailed proof is given in Appendix B. We also provide more explicit formulas for htotal,Itotalh_{\text{total}},I_{\text{total}} below.

Assume 5 is true. Then in the same infinite-depth-and-width limit as 1, the total hypoactivation htotalh_{\text{total}} and total interlayer covariance ItotalI_{\text{total}} obey

Here Cα,λC_{\alpha,\lambda} is a constant depending on α,λ\alpha,\lambda and

where J2(θ)J_{2}(\theta) first appeared in , and θk\theta_{k} is such that cos⁡(θk)=αk/(α2+λ2)k/2\cos(\theta_{k})=\alpha^{k}/(\alpha^{2}+\lambda^{2})^{k/2} .

The Gaussian infinite-width limit predicts that the marginals of zoutz^{\text{out}} should have the form of (6) with GG being identically zero. As zoutz^{\text{out}} depends exponentially on GG, the infinite-depth-and-width limit predicts the variance is exponentially larger than the infinite-width limit. See Section 3.1 for a detailed discussion and Figure 3 for verification against finite networks.

Using Monte Carlo simulations, we estimate the constant Cα,λC_{\alpha,\lambda}. The result of 1 is then compared against finite networks in Figure 2. Full proof can be found Appendix B.

The idea of using a scaling parameter α=1/2\alpha=1/\sqrt{2} in the skip connections has been noted in empirical papers and also studied under simplified assumptions on the number of activation in each layer . Our result shows that by choosing α2+λ2=1\alpha^{2}+\lambda^{2}=1, the prefactor in (6) does not grow with depth thereby enabling deeper networks to be trained.

For a balanced ResNet, the same result as 1 given in (6) still holds, but with the mean and variance of GG given simply by

Consequences of Theorems 1 & 4 and Comparison to Infinite-Width Limit

By the basic fact E[exp⁡(N(μ,σ2))]=exp⁡(μ+12σ2)\mathbf{E}[\exp(\mathcal{N}(\mu,\sigma^{2}))]=\exp(\mu+\frac{1}{2}\sigma^{2}) it follows from Theorems 1 & 4 that, when the inputs xx has ∥x∥=nin\left\|x\right\|=\sqrt{{n_{\text{in}}}}, the mean size scale of any neuron zioutz^{\text{out}}_{i} is approximately

(Note that the terms with β\beta cancel out!) When α2+λ2=1\alpha^{2}+\lambda^{2}=1, this is constant for Balanced ResNets. In contrast, Vanilla ResNets have a complicated dependence on the network depth dd and width nn due to the hypoactivation and correlations terms. This means the behaviour is exp⁡(Cd/n)\exp(Cd/n), which is somewhat surprising. A more serious issue is the variance which Theorems 1 & 4 predict to be

Balanced ResNets suffer from this variance problem less because the interlayer correlation term is zero. Since this variance reduction happens at the exponential scale in, the difference can be significant; for networks with d/n=1d/n=1, the contribution is a factor of ≈e5.5≈250\approx e^{5.5}\approx 250 for Vanilla ResNets vs. ≈e2.5≈10\approx e^{2.5}\approx 10 for Balanced ResNets. See Figure 3 for a comparison of these theoretically predicted properties vs experiments with finite networks.

2 Correlated Output Neurons

Since the same random variable GG multiplies the entire vector zoutz^{\text{out}}, the individual neurons in the output layer are not independent. For example, Theorem 1 and 4 predict that the squared entries have strictly positive correlation given by Corr((ziout)2,(zjout)2)=(exp⁡(σ2)−1)/(3exp⁡(σ2)−1)\mathbf{Corr}\left(\left(z^{\text{out}}_{i}\right)^{2},\left(z^{\text{out}}_{j}\right)^{2}\right)=(\exp(\sigma^{2})-1)/(3\exp(\sigma^{2})-1) for any two neurons i≠ji\neq j where σ2=Var(G)\sigma^{2}=\mathbf{Var}(G). This tends to 1/31/3 as d/nd/n grows. The effect of correlated output neurons persists for Balanced ResNet but is reduced again due to the lower variance. This is very different from the infinite-width limit, which predicts that individual neurons should be independent Gaussians. This prediction of the theorem matches finite networks closely; see Figure 3.

Conjecture 5: Hypoactivation and Layerwise Correlations

Proof Ideas for Theorems 1 & 4

A key element of the proof is the following property of Gaussian random matrices. If WW which has iid N(0,1)\mathcal{N}(0,1) entries, then for any vector xx, we have

where gg is a vector whose entries are iid N(0,1)\mathcal{N}(0,1) random variables. Because of the fully connected first and last layer of the network, (14) implies that

Hence G≔ln⁡((∥zd∥2/n)⋅(∥x∥2/nin)−1⋅(α2+λ2)−d)G\coloneqq\ln\left((\left\|z^{d}\right\|^{2}/n)\cdot(\left\|x\right\|^{2}/{n_{\text{in}}})^{-1}\cdot\left(\alpha^{2}+\lambda^{2}\right)^{-d}\right) only depends on n,d,α,λn,d,\alpha,\lambda. (Equivalently, GG has the distribution of ln⁡(∥zd∥2/n⋅(α2+λ2)−d)\ln(\left\|z^{d}\right\|^{2}/n\cdot\left(\alpha^{2}+\lambda^{2}\right)^{-d}) when z0=gz^{0}=g.) With this definition, (15) also shows zoutz^{\text{out}} is proportional to exp⁡(G/2)\exp(G/2), establishing the first part of 1.

Taking the ln⁡\ln of (17) exhibits ln⁡(∥zd∥2/n)\ln(\left\|z^{d}\right\|^{2}/n) as a sum of these weakly correlated random variables. Here we note that various tail estimates for the same or related quantities have been developed , however these estimates are not precise enough to pinpoint the exact limiting distribution. In contrast, we are able to derive the exact limiting distribution via a Central Limit Theorem (CLT) for weakly correlated sums . The proof of 1 is completed by computing the mean, variance and covariance of terms using 5. For 4, the final calculation is simplified by (10) which shows the terms are uncorrelated. The detailed proof is given in Appendix B.

Acknowledgement

We would like to thank Blair Bilodeau, Gintare Karolina Dziugaite, Mahdi Haghifam, Yani A. Ioannou, James Lucas, Jeffrey Negrea, Mengye Ren, and Ekansh Sharma for helpful discussions and draft feedback. We would also like to thank the anonymous NeurIPS reviewers for insightful feedback. In particular, one identified numerous relations to existing work, and another helped us identify the uniformity requirement in 5. ML is supported by Ontario Graduate Scholarship and the Vector Institute. MN is supported by an NSERC Discovery Grant. DMR is supported in part by an NSERC Discovery Grant, Ontario Early Researcher Award, and a stipend provided by the Charles Simonyi Endowment.

References

Appendix A Appendix

We stated our main result with fixed α,λ\alpha,\lambda, but the result easily extends to the case where α,λ\alpha,\lambda vary from layer to layer. This allows comparison between our result and the infinite width limits , where they have α=1\alpha=1 and allow λi\lambda_{i} to varying layer by layer. The statement in that setting is modified as follows.

Suppose αi,λi\alpha_{i},\lambda_{i} are sequences such that λi\lambda_{i} is uniformly bounded away from . Define the network by

Consider the limit where both the network depth d→∞d\to\infty and hidden layer width n→∞n\to\infty in such a way that the ratio dn\frac{d}{n} converges to a constant. In this limit, the distribution of the network output zoutz^{\text{out}} for a given input xx is given by

The corresponding limit theorem for balanced ResNets also holds; the hypoactivation and layer-wise covariance term in (20) vanish.

The behaviour is a complicated function of the sequences αi,λi\alpha_{i},\lambda_{i}. It would be interesting to use these theoretical results to guide the choice of parameters αi,λi\alpha_{i},\lambda_{i} and investigate how this effects training behaviour.

Appendix B Proof of main results

To simplify the exposition of the proofs, we will assume without loss of generality that α2+λ2=1\alpha^{2}+\lambda^{2}=1. The general case can be reduced to the case α2+λ2=1\alpha^{2}+\lambda^{2}=1 by dividing by α2+λ2\sqrt{\alpha^{2}+\lambda^{2}} in each layer and rescaling the parameters α,λ\alpha,\lambda to αα2+λ2\frac{\alpha}{\sqrt{\alpha^{2}+\lambda^{2}}} and λα2+λ2\frac{\lambda}{\sqrt{\alpha^{2}+\lambda^{2}}}.

Taking expectation of both sides of (22), we have

where we have used α2+λ2=1\alpha^{2}+\lambda^{2}=1 in the last line. ∎

B.2 Variance Calculation

We will use the decomposition Var(X)=E[(X−α2)2]+(E[X]−α2)2\mathbf{Var}\left(X\right)=\mathbf{E}\left[(X-\alpha^{2})^{2}\right]+\left(\mathbf{E}[X]-\alpha^{2}\right)^{2} and compute the the two terms individually.

In the second term, since E[g1]=0\mathbf{E}[g_{1}]=0 and E[∥g∥2]=n\mathbf{E}[\left\|g\right\|^{2}]=n, we have E[X]−α2=λ2E[2∥φ+(z)∥2]\mathbf{E}[X]-\alpha^{2}=\lambda^{2}\mathbf{E}\left[2\left\|\varphi_{+}(z)\right\|^{2}\right].

To compute E[(X−α2)2]\mathbf{E}[(X-\alpha^{2})^{2}], notice that X−α2=λ22∥φ+(z)∥2⋅1n∥g∥2+2αλ1n∥2φ+(z)∥g1X-\alpha^{2}=\lambda^{2}2\left\|\varphi_{+}(z)\right\|^{2}\cdot\frac{1}{n}\left\|g\right\|^{2}+2\alpha\lambda\frac{1}{\sqrt{n}}\left\|\sqrt{2}\varphi_{+}(z)\right\|g_{1} is a sum of two terms. The two terms are uncorrelated since Cov(∥g∥2,g1)=0\mathbf{Cov}(\left\|g\right\|^{2},g_{1})=0. Hence, the expectation of the cross term in (X−α2)2\left(X-\alpha^{2}\right)^{2} is zero and we can compute

We have used the fact about χn2\chi^{2}_{n} random variables that E[∥g∥4]=n(n+2)\mathbf{E}\left[\left\|g\right\|^{4}\right]=n(n+2). Finally then:

This can be calculated directly using properties of the unit sphere, but the proof is complicated by the fact that the entries ui,uju_{i},u_{j} are not independent. Instead, there is an elementary proof using the following equality in distribution:

Taking norm and expectation of this equality in distribution gives E[2∥φ+(u)∥2∥g∥2]=E[2∥φ+(g)∥2]\mathbf{E}\left[2\left\|\varphi_{+}(u)\right\|^{2}\left\|g\right\|^{2}\right]=\mathbf{E}\left[2\left\|\varphi_{+}(g)\right\|^{2}\right]. Since ∥g∥\left\|g\right\| and uu are independent, we can factor and rearrange to obtain

Since the entries of gg are independent of each other, it is easier to compute using gg and this identity instead of using uu. Moreover, because Gaussian distribution are symmetrically distributed, we have that {g12,…,gn2}\left\{g_{1}^{2},\ldots,g_{n}^{2}\right\} is independent of,1{g1>0},…,1{gn>0}\mathtt{1}\left\{g_{1}>0\right\},\ldots,\mathtt{1}\left\{g_{n}>0\right\} Hence:

and so E[2∥φ+(u)∥2]=E[2∥φ+(g)∥2]E[∥g∥2]=nn=1\mathbf{E}\left[2\left\|\varphi_{+}(u)\right\|^{2}\right]=\frac{\mathbf{E}\left[2\left\|\varphi_{+}(g)\right\|^{2}\right]}{\mathbf{E}\left[\left\|g\right\|^{2}\right]}=\frac{n}{n}=1. Similarly, using E[4∥φ+(u)∥4]=E[4∥φ+(g)∥4]E[∥g∥4]\mathbf{E}\left[4\left\|\varphi_{+}(u)\right\|^{4}\right]=\frac{\mathbf{E}\left[4\left\|\varphi_{+}(g)\right\|^{4}\right]}{\mathbf{E}\left[\left\|g\right\|^{4}\right]} we now compute by looking at diagonal and off-diagonal terms as follows

Using E[∥g∥4]=n(n+2)\mathbf{E}\left[\left\|g\right\|^{4}\right]=n(n+2), we finally obtain E[4∥φ+(u)∥4]=n(n+5)n(n+2)\mathbf{E}\left[4\left\|\varphi_{+}(u)\right\|^{4}\right]=\frac{n(n+5)}{n(n+2)} from which the claimed variance formula follows. ∎

B.3 Uniform distribution on spheres and Gaussian random variables

This is a direct corollary to Theorem 2 of , which uses Stirling’s formula to show that the total variation distance between the random variables nui\sqrt{n}u_{i} and ZZ is at most 2(1+3n−3−1)=O(1/n)2\left(\sqrt{1+\frac{3}{n-3}}-1\right)=O(1/n). ∎

The result then follows by using the independence of ∥g∥\left\|g\right\| and uu, and the formula for the pp-th moment of χn2\chi^{2}_{n} random variable, namely ∥g∥2p=n(n+2)⋯(n+2p−2)\left\|g\right\|^{2p}=n(n+2)\cdots(n+2p-2). ∎

The proof is immediate writing the difference in expectation as an integral and then comparing ∫f(x)(ρnui(x)−ρZ(x))dx\intop f(x)\left(\rho_{\sqrt{n}u_{i}}(x)-\rho_{Z}(x)\right)dx to ∫(Ax2p+B)(ρnui(x)−ρZ(x))dx\intop\left(Ax^{2p}+B\right)\left(\rho_{\sqrt{n}u_{i}}(x)-\rho_{Z}(x)\right)dx by the results of the previous two lemmas. ∎

B.4 Pairwise covariances

By expanding the norms into sums, ∥x∥2=∑i=1nxi2\left\|x\right\|^{2}=\sum_{i=1}^{n}x_{i}^{2}, we can compute the covariance by summing over all pairs of coordinates ui,cos⁡(θ)uj+sin⁡(θ)n−12gju_{i},\cos(\theta)u_{j}+\sin(\theta)n^{-\frac{1}{2}}g_{j}. There are two types of terms to consider. (Note: we use the notation φ+2(x)=(φ+(x))2\varphi_{+}^{2}(x)=(\varphi_{+}(x))^{2}.)

Diagonal Terms\contourwhiteDiagonal Terms: Cov(φ+2(ui),φ+2(cos⁡(θ)ui+sin⁡(θ)n−12gi))\mathbf{Cov}\left(\varphi_{+}^{2}(u_{i}),\varphi_{+}^{2}(\cos(\theta)u_{i}+\sin(\theta)n^{-\frac{1}{2}}g_{i})\right) for 1≤i≤n1\leq i\leq n We first use the positive homogeneity of φ+\varphi_{+} to extract a factor of n\sqrt{n} from both terms

We now use the approximation of nui\sqrt{n}u_{i} by Z∼N(0,1)Z\sim\mathcal{N}(0,1) as in Section B.3 to obtain

where we have used the definition of Jˉ2(θ)\bar{J}_{2}(\theta) from (23).

Off diagonal terms\contourwhiteOff diagonal terms Cov(φ+2(ui),φ+2(cos⁡(θ)uj+sin⁡(θ)n−12gj))\mathbf{Cov}\left(\varphi_{+}^{2}(u_{i}),\varphi_{+}^{2}(\cos(\theta)u_{j}+\sin(\theta)n^{-\frac{1}{2}}g_{j})\right) for 1≤i≤n1\leq i\leq n,1≤j≤n1\leq j\leq n,i≠ji\neq j

As in the calculation for the diagonal term, we now again use the approximation that nuj\sqrt{n}u_{j} is approximately marginally distributed like Z∼N(0,1)Z\sim\mathcal{N}(0,1) and the positive homogoneity of φ+\varphi_{+} to obtain

Renaming gjg_{j} to WW to match the notation of (23), we finally obtain

Summing the diagonal and off-diagonal terms, we find the total covariance is:

which gives the claimed formula for the covariance. ∎

This follows immediately from the approximation for the covariance in 5 and 16. ∎

By the series expansion for arccos⁡(x)=π2−x+O(x3)\arccos(x)=\frac{\pi}{2}-x+O(x^{3}) as x→0x\to 0, it follows that Jˉ2(θk)−Jˉ2(π−θk)=8αk/π+O(α2k)\bar{J}_{2}(\theta_{k})-\bar{J}_{2}(\pi-\theta_{k})=8\alpha^{k}/\pi+O\left(\alpha^{2k}\right) as k→∞k\to\infty. This exponential decay in kk explains why the total covariance correction ItotalI_{\text{total}} remains O(d/n)O(d/n) even as as d→∞d\to\infty.

B.5 Central Limit Theorem

where the constant in the big O(⋅)O(\cdot) notation is uniform in dd. Moreover, SS is asymptotically Gaussian in the infinite width and depth limit.

from which the desired mean and variance formula for SS follows.

The proof is analogous to the proof of 1; the calculation for the activation of the layers is simplified by the fact that which neurons activated in each layer are uncorrelated by (10). ∎

B.6 Input-Output Gradient for Balanced ResNets

For any input xx, the output of a Balanced ResNet can be written as

The derivative of zoutz^{\text{out}} with respect to any input xix_{i} has the distribution

This follows immediately since M(x)M(x) is constant on the neighbourhood A(x)\mathcal{A}(x). ∎

∂∂xizout\frac{\partial}{\partial x_{i}}z^{\text{out}} has the distribution as the output zoutz^{\text{out}} at any input xx with ∥x∥=1\left\|x\right\|=1.

By the previous results, both are equal in distribution to MuMu for any unit vector uu, where MM is the distribution of the random matrix in (33). ∎

Appendix C Experiments

Throughout the paper, the Monte Carlo simulations were computed on a single NVIDIA Titan-XP GPU. The main tools used in the neural network simulations are the JAX library (Apache 2.0 License) and the PyTorch library (BSD 3-Clause License). Furthermore, the Python libraries numpy (BSD 3-Clause License), plotnine (based on ggplot2 , GNU GPLv2 License), and pandas (BSD 3-Clause License) tremendously helpful. We used Python version 3.6.8 from Anaconda 3 (3-clause BSD License) and Jupyter notebook (3-Clause BSD License).

These experiments, displayed in Figure 6, investigate the difference between training performance of the Vanilla ResNet from (1) and the Balanced ResNet defined in Section 2.1. This is beyond the theory proven in our work which concerns statistical properties of the network on initialization. Both of these architectures are fully-connected networks with skip connections between layers. We observe that both architectures perform similarly in standard training regimes where the networks are much wider than they are deep.

C.2 Convolutional ResNet CIFAR-10 Experiments

These experiments investigate how the Balanced ResNet architecture modification, namely randomly flipping between φ+\varphi_{+} and φ−\varphi_{-}, effects training for deep convolutional ResNets (We call these C-ResNets here to distinguish from the fully connected ones studied in detail in the paper). Although the theory in this paper only are proven for Vanilla ResNets (which are fully connected with skip connections), we expect that the Balanced ResNet tweak does not negatively effect performance in standard training regimes and may allow better initialization for very deep models.

For this experiment, we used the ResNet implementations from the GitHub repository by (MIT License). To create a Balanced Convolutional ResNet (Balanced C-ResNet), we modified the class PreActBlock to add flipped ReLUs into the network channel-wise, thinking of channels as the natural generalization of neurons for convolution neural networks. More specifically, we are interested in an input tensor YY of dimension (b,c,h,w)(b,c,h,w), where bb is for batch size, cc is for channel, hh is for height, and ww is for width. Before feeding into a ReLU non-linearity, we will multiply YY by a vector of iid random signs {sj}j∈[c]\{s_{j}\}_{j\in[c]}, so that the ReLU output is

The experiment results are reported in Table 1 and Figure 7. Here, we demonstrate that using the Balanced ResNet architecture does not reduce the performance of existing training regimes. For the deeper ResNet101, the Balanced ResNet seems to perform slightly better early in training but finishes 0.2%0.2\% worse at the end of training.

Note that even the deeper ResNet101 tested here is still relatively shallow compared to its “width”; most of the layers are either 256256 or 10241024 channels by 8×88\times 8 neurons per channel which represent many more neurons in each hidden layer than the depth of the network, 101101. The possible advantage of the Balanced ResNet idea is that it will enable the training architectures which are even deeper compared to their width, which are currently not trainable or difficult to train due to initialization issues. A detailed empirical study exploring this idea is needed to study how Balanced ResNets perform beyond the theory proven in this paper.

C.3 Density Plot Calculations

In this section, we describe the calculations required for plotting Figure 1. Firstly, we need to estimate the hypoactivation constant Cα,λC_{\alpha,\lambda} from 2 using Monte Carlo simulations. For the choice of α=λ=1/2\alpha=\lambda=1/\sqrt{2}, we find the constant Cα,λ≈−0.876C_{\alpha,\lambda}\approx-0.876, which we use for estimating the mean. See Figure 8 for further simulations with varying α,λ\alpha,\lambda values, and Figure 9 for simulations demonstrating these constants provide accurate prediction for mean and variance.

Next, using the choice of ∥xin∥=1\left\|x_{in}\right\|=1 and α2+λ2=1\alpha^{2}+\lambda^{2}=1, we can write

Here we observe that ∥Z⃗∥2∼χ2(nout)\left\|\vec{Z}\right\|^{2}\sim\chi^{2}(n_{out}), and therefore we can compute the density of ln⁡∥Z⃗∥2\ln\left\|\vec{Z}\right\|^{2} with a coordinate change

To compute the density of the infinite width prediction, we first observe that in this limit zout∼N(0,σ2Inout)z^{\text{out}}\sim\mathcal{N}(0,\sigma^{2}I_{n_{out}}). Therefore it’s sufficient to simply compute the variance.

To this goal, we follow the calculations of with the variance recursion formula (slightly modified to include α\alpha)

C.4 Additional Monte Carlo Simulations

These additional Monte Carlo simulations provide further comparisons between the infinite depth-and-width limit predictions and finite networks. See Figure 9.