Random Neural Networks in the Infinite Width Limit as Gaussian Processes

Boris Hanin

Introduction

In the last decade or so neural networks, originally introduced in the 1940’s and 50’s , have become indispensable tools for machine learning tasks ranging from computer vision to natural language processing and reinforcement learning . Their empirical success has raised many new mathematical questions in approximation theory , probability (see §1.2.2 for some references), optimization/learning theory and so on. The present article concerns a fundamental probabilistic question about arguably the simplest networks, the so-called fully connected neural networks, defined as follows:

This article considers the mapping xα↦zα(L+1)x_{\alpha}\mapsto z_{\alpha}^{(L+1)} when the network’s weights and biases are chosen independently at random and the hidden layer widths n1,…,nLn_{1},\ldots,n_{L} are sent to infinity while the input dimension n0,n_{0}, output dimension nL+1n_{L+1}, and network depth LL are fixed. In this infinite width limit, akin to the large matrix limit in random matrix theory (see §1.2), neural networks with random weights and biases converge to Gaussian processes (see §1.4 for a review of prior work). Unlike prior work Theorem 1.2, our main result, is that this holds for general non-linearities σ\sigma and distributions of network weights (cf §1.3).

Moreover, in addition to establishing convergence of wide neural networks to a Gaussian process under weak hypotheses, the present article gives a mathematical take aimed at probabilists of some of the ideas developed in the recent monograph . This book, written in the language and style of theoretical physics by Roberts and Yaida, is based on research done jointly with the author. It represents a far-reaching development of the breakthrough work of Yaida , which was the first to systematically explain how to compute finite width corrections to infinite width Gaussian process limit of random neural networks for arbitrary depth, width, and non-linearity. Previously, such finite width (and large depth) corrections were only possible for some special observables in linear and ReLU networks . The present article deals only with the asymptotic analysis of random neural networks as the width tends to infinity, leaving to future work a probabilistic elaboration of the some aspects of the approach to finite width corrections from .

The rest of this article is organized as follows. First, in §1.2 we briefly motivate the study of neural networks with random weights. Then, in §1.3 we formulate our main result, Theorem 1.2. Before giving its proof in §2, we first indicate in §1.4 the general idea of the proof and its relation to prior work.

2 Why Random Neural Networks?

Beyond illuminating the properties of networks at the start of training, the analysis of random neural networks can reveal a great deal about networks after training as well. Indeed, on a heuristic level, just as the behavior of the level spacings of the eigenvalues of large random matrices is a surprisingly good match for emission spectra of heavy atoms , it is not unreasonable to believe that certain coarse properties of the incredibly complex networks used in practice will be similar to those of networks with random weights and biases. More rigorously, neural networks used in practice often have many more tunable parameters (weights and biases) than the number of datapoints from the training dataset. Thus, at least in certain regimes, neural network training provably proceeds by an approximate linearization around initialization, since no one parameter needs to move much to fit the data. This so-called NTK analysis shows, with several important caveats related to network size and initialization scheme, that in some cases the statistical properties of neural networks at the start of training are the key determinants of their behavior throughout training.

2.2 Motivation from Random Matrix Theory

Finally, beyond studying linear networks, random matrix theory questions naturally appear in neural network theory via non-linear analogs of the Marchenko-Pastur distribution for empirical covariance matrices of zα(L+1)z_{\alpha}^{(L+1)} when α∈A\alpha\in A ranges over a random dataset of inputs as well as through the spectrum of the input-output Jacobian and the NTK .

3 Main Result

We further assume the network biases are iid GaussianAs explained in §1.4 the universality results in this article are simply not true if the biases are drawn iid from a fixed non-Gaussian distribution. and independent of the weights:

for this limiting process satisfies the layerwise recursion

where the distribution of (z1;α(1),z1;β(1))(z_{1;\alpha}^{(1)},z_{1;\beta}^{(1)}) is determined via (1.3) by the distribution of weights and biases in the first layer and hence is not universal.

We prove Theorem 1.2 in §2. First, we explain the main idea and review prior work.

4 Theorem 1.2: Discussion, Main Idea, and Relation to Prior Work

At a high level, the proof of Theorem 1.2 (specifically the convergence of finite-dimensional distributions) proceeds as follows:

Conditional on the mapping xα↦zα(L)x_{\alpha}\mapsto z_{\alpha}^{(L)}, the components of the neural network output xα↦zα(L+1)x_{\alpha}\mapsto z_{\alpha}^{(L+1)} are independent sums of nLn_{L} independent random fields (see (1.3)), and hence, when nLn_{L} is large, are each approximately Gaussian by the CLT.

The conditional covariance in the CLT from step 11 is random at finite widths (it depends on zα(L)z_{\alpha}^{(L)}). However, it has the special form of an average over j=1,…,nLj=1,\ldots,n_{L} of the same function applied to each component zj;α(L)z_{j;\alpha}^{(L)} of the vector zα(L)z_{\alpha}^{(L)} of pre-activations at the last hidden layer. We call such objects collective observables (see §2.1.2 and (1.10)).

The LLN from step 3 allows us to replace the random conditional covariance matrix from steps 1 and 2 by its expectation, asymptotically as n1,…,nLn_{1},\ldots,n_{L} tend to infinity.

We turn to giving a few more details on steps 1-4 and reviewing along the way the relation of the present article to prior work. The study of the infinite width limit for random neural networks dates back at least to Neal , who considered networks with one hidden layer:

where i=1,…,n2.i=1,\ldots,n_{2}. In the shallow L=1L=1 setting of Neal if in addition n2=1n_{2}=1, then neglecting the bias b1(2)b_{1}^{(2)} for the moment, the scalar field z1;α(2)z_{1;\alpha}^{(2)} is a sum of iid random fields with finite moments, and hence the asymptotic normality of its finite-dimensional distributions follows immediately from the multidimensional CLT. Modulo tightness, this explains why z1;α(2)z_{1;\alpha}^{(2)} ought to converge to a Gaussian field. Even this simple case, however, holds several useful lessons:

If the distribution of the bias b1(2)b_{1}^{(2)} is fixed independent of n1n_{1} is and non-Gaussian, then the distribution of z1;α(2)z_{1;\alpha}^{(2)} will not be Gaussian, even in the limit when n1→∞n_{1}\rightarrow\infty.

If the first layer biases bj(1)b_{j}^{(1)} are drawn iid from a fixed distribution μb\mu_{b} and σ\sigma is non-linear, then higher moments of μb\mu_{b} will contribute to the variance of each neuron post-activation σ(zj;α(1))\sigma(z_{j;\alpha}^{(1)}), causing the covariance of the Gaussian field at infinite width to be non-universal.

Unlike in deeper layers, as long as n0n_{0} is fixed, the distribution of each neuron pre-activation zj;α(1)z_{j;\alpha}^{(1)} in the first layer will not be Gaussian, unless the weights and biases in layer 11 are themselves Gaussian. This explains why, in the initial condition (1.8) the distribution is non-Gaussian in the first layer.

In light of the first two points, what should one assume about the bias distribution? There are, it seems, two options. The first is to assume that the variance of the biases tends to zero as n1→∞n_{1}\rightarrow\infty, putting them on par with the weights. The second, which we adopt in this article, is to declare all biases to be Gaussian.

The first trick in proving Theorem 1.2 for general depth and width appears already when L=1L=1 but the output dimension n2n_{2} is at least two.Neal states erroneously on page 38 of his thesis that zi;α(2)z_{i;\alpha}^{(2)} and zj;α(2)z_{j;\alpha}^{(2)} will be independent because the weights going into them are independent. This is not true at finite width but becomes true in the infinite width limit. In this case, even for a single network input xαx_{\alpha}, at finite values of the network width n1n_{1} different components of the random n2n_{2}-dimensional vector zα(2)z_{\alpha}^{(2)} are not independent, due to their shared dependence on the vector zα(1)z_{\alpha}^{(1)}. The key observation, which to the author’s knowledge was first presented in , is to note that the components of zα(2)z_{\alpha}^{(2)} are independent conditional on the first layer (i.e. on zα(1)z_{\alpha}^{(1)}) and are approximately Gaussian when n1n_{1} is large by the CLT. The conditional variance, which captures the main dependence on zα(1)z_{\alpha}^{(1)}, has the following form:

This is an example of what we’ll call a collective observable, an average over all neurons in a layer of the same function applied to the pre-activations at each neuron (see §2.1.2 for the precise definition). In the shallow L=1L=1 setting, Σαα(2)\Sigma_{\alpha\alpha}^{(2)} is a sum of n1n_{1} iid random variables with finite moments. Hence, by the LLN, it converges almost surely to its mean as n1→∞n_{1}\rightarrow\infty. This causes the components of zα(2)z_{\alpha}^{(2)} to become independent in the infinite width limit, since the source of their shared randomness, Σαα(L+1)\Sigma_{\alpha\alpha}^{(L+1)}, can be replaced asymptotically by its expectation.

We conclude this section by pointing the reader to several other related strands of work. The first are articles such as , which quantify the magnitude of the difference

The second is the series of articles starting with the work of Yang , which develops the study not only of initialization but also certain aspects of inference with infinitely wide networks using what Yang terms tensor programs. As part of that series, the article establishes that in the infinite width limit many different architectures become Gaussian processes. However, the arguments in those articles are significantly more technical than the ones presented here since they are focused on building the foundation for the tensor program framework. At any rate, to the best of the author’s knowledge, no prior article addresses universality of the Gaussian process limit with respect to the weight distribution in deep networks (for shallow networks with L=1L=1 this was considered by Neal in ). Finally, that random neural networks converge to Gaussian processes in the infinite width limit under various restrictions but for architectures other than fully connected is taken up in .

Proof of Theorem 1.2

Let us recall the notation. Namely, we fix a network depth L≥1L\geq 1, an input dimension n0≥1,n_{0}\geq 1, an output dimension nL+1≥1n_{L+1}\geq 1, hidden layer widths n1,…,nL≥1n_{1},\ldots,n_{L}\geq 1 and a non-linearity σ\sigma satisfying (2.5). We further assume that the networks weights and biases are independent and random as in (1.4) and (1.6). To prove Theorem 1.2 we must show that the random fields xα↦zα(L+1)x_{\alpha}\mapsto z_{\alpha}^{(L+1)} converge weakly in distribution to a Gaussian process in the limit where n1,…,nLn_{1},\dots,n_{L} tend to infinity. We start with the convergence of finite-dimensional distributions. Let us therefore fix a collection

between the entries in each row satisfies the recursion (1.7) with initial condition (1.8).

𝐿1z_{\alpha}^{(L+1)}). For every L≥1, ϵ>0L\geq 1,\,\epsilon>0 there exists C=C(ϵ,σ,T,L,Cb,CW)>0C=C(\epsilon,\sigma,T,L,C_{b},C_{W})>0 so that

We continue to assume (as in the statement of Theorem 1.2) that all biases are Gaussian:

is the matrix defined by the recursion (1.7) with initial condition (1.8). Writing

where for any α,β∈A\alpha,\beta\in A the conditional covariance is

Using (2.4) and the explicit form of the characteristic function of a Gaussian reveals

The crucial observation is that each entry of the conditional covariance matrix ΣA(L+1)\Sigma_{A}^{(L+1)} is an average over j=1,…,nLj=1,\ldots,n_{L} of the same fixed function applied to the vector zj;A(L)z_{j;A}^{(L)}. While zj;A(L)z_{j;A}^{(L)} are not independent at finite values of n1,…,nL−1n_{1},\ldots,n_{L-1} for L>1L>1, they are sufficiently weakly correlated that a weak law of large numbers still holds:

Fix n0,nL+1n_{0},n_{L+1}. There exists a ∣A∣×∣A∣\left|A\right|\times\left|A\right| PSD matrix

Lemma 2.3 is a special case of Lemma 2.4 (see §2.1.2). ∎

Lemma 2.3 implies that ΣA(L+1)\Sigma_{A}^{(L+1)} converges in distribution to KA(L+1)K_{A}^{(L+1)}. In view of (2.6) and the definition of weak convergence this immediately implies (2.2). It therefore remains to check that KA(L+1)K_{A}^{(L+1)} satisfies the desired recursion. For this, note that at any values of n1,…,nLn_{1},\ldots,n_{L} we find

where the law of (zi;α(1),zi;β(1))(z_{i;\alpha}^{(1)},z_{i;\beta}^{(1)}) is determined by the distribution μW\mu_{W} of weights in layer 11 and does not depend on n1n_{1}. This confirms the initial condition (1.8). Otherwise, if L>1L>1, the convergence of finite-dimensional distributions that we’ve already established yields

Since σ\sigma is continuous we may invoke the continuous mapping theorem to conclude that

1.2 Collective Observables with Gaussian Weights: Generalizing Lemma 2.3

Hence, we have the following convergence in probability

Hence, zi;α(1)z_{i;\alpha}^{(1)} have finite moments since bi(1)b_{i}^{(1)} are iid Gaussian and W^ij(1)\widehat{W}_{ij}^{(1)} are mean with finite higher moments. In particular, since ff is polynomially bounded, we find for every n1n_{1} that

which is finite and independent of n1n_{1}, confirming (2.7). Further, On1,f;A(1)\mathcal{O}_{n_{1},f;A}^{(1)} is the average of n1n_{1} iid random variables with all moments finite. Hence, (2.8) follows by the weak law of large numbers, completing the proof of the base case.

Since the weights and biases in layer L+1L+1 are Gaussian and independent of FL\mathcal{F}_{L}, we find

where ΣA(L+1)\Sigma_{A}^{(L+1)} is the conditional covariance defined in (2.5) and GG is an ∣A∣\left|A\right|-dimensional standard Gaussian. The key point is that ΣA(L+1)\Sigma_{A}^{(L+1)} is a collective observable at layer LL. Hence, by the inductive hypothesis, there exists a PSD matrix Σ‾A(L+1)\overline{\Sigma}_{A}^{(L+1)} such that ΣA(L+1)\Sigma_{A}^{(L+1)} converges in probability to Σ‾A(L+1)\overline{\Sigma}_{A}^{(L+1)} as n1,…,nL→∞n_{1},\ldots,n_{L}\rightarrow\infty. To establish (2.7) it therefore suffices in view of (2.9) to check that

where the right hand side is finite since ff is polynomially bounded and all polynomial moments of GG are finite. To establish (2.11), let us invoke the Skorohod representation theorem to find a common probability space on which there are versions of ΣA(L+1)\Sigma_{A}^{(L+1)} – which by an abuse of notation we will still denote by ΣA(L+1)\Sigma_{A}^{(L+1)} – that converge to Σ‾A(L+1)\overline{\Sigma}_{A}^{(L+1)} almost surely. Next, note that since ff is polynomially bounded we may repeatedly apply ab≤12(a2+b2)ab\leq\frac{1}{2}(a^{2}+b^{2}) to find

where pp is a polynomial in the entries of (ΣA(L+1))1/2(\Sigma_{A}^{(L+1)})^{1/2}, qq a polynomial in the entries of GG, and the polynomials p,qp,q don’t depend on n1,…,nLn_{1},\ldots,n_{L}. The continuous mapping theorem shows that

Thus, since all moments of Gaussian are finite, (2.11) follows from the generalized dominated convergence theorem. It remains to argue that (2.8) holds at layer L+1L+1. To do this, we write

since we already showed that (2.7) holds at layer L+1L+1. Next, recall that, conditional on FL\mathcal{F}_{L}, neurons in layer L+1L+1 are independent. The law of total variance and Cauchy-Schwartz yield

Using (2.10) and the polynomial estimates (2.12) on ff, we conclude that the conditional expectation on the previous line is some polynomially bounded function of the components of (ΣA(L+1))1/2(\Sigma_{A}^{(L+1)})^{1/2}. Hence, we may apply dominated convergence as above to find

1.3 Proof of Proposition 2.1 for General Weights

To check (2.15), let us define an intermediate object:

where the entries Wij(L+1),∙W_{ij}^{(L+1),{\tiny{\bullet}}} of W(L+1),∙W^{(L+1),{\tiny{\bullet}}} are iid Gaussian with mean and variance CW/nLC_{W}/n_{L}. That is, we take the vector σ(zα(L))\sigma(z_{\alpha}^{(L)}) of post-activations from layer LL obtained by using general weights in layers 1,…,L1,\ldots,L and use Gaussian weights only in layer L+1L+1. Our first step in checking (2.15) is to show that it this relation holds when zA(L+1)z_{A}^{(L+1)} is replaced by zA(L+1),∙z_{A}^{(L+1),{\tiny{\bullet}}}.

This is a standard Lindeberg swapping argument. Namely, for each α∈A\alpha\in A and k=0,…,nLk=0,\ldots,n_{L} define

where the first kk entries of each row of W(L+1),kW^{(L+1),k} are iid Gaussian with mean and variance CW/nLC_{W}/n_{L}, while the remaining entries are (CW/nL)−1/2(C_{W}/n_{L})^{-1/2} times iid draws W^ij(L+1)\widehat{W}_{ij}^{(L+1)} from the general distribution μ\mu of network weights, as in (1.4) and (1.5). With this notation, we have

consider the third order Taylor expansion of gg around ZZ.

where W~ik(L+1)∼N(0,1)\widetilde{W}_{ik}^{(L+1)}\sim\mathcal{N}(0,1). Then, Taylor expanding gg to third order around Zk=zi;α(L+1),kZ_{k}=z_{i;\alpha}^{(L+1),k} and, using that the first two moments of (CWnL−1)1/2W^ij(L)(C_{W}n_{L}^{-1})^{1/2}\widehat{W}_{ij}^{(L)} match those of N(0,CWnL−1)\mathcal{N}(0,C_{W}n_{L}^{-1}), we find that

To make use of Lemma 2.5 let us consider any collective observable OnL,f;A(L)\mathcal{O}_{n_{L},f;A}^{{}_{(L)}} at layer LL. Recall that by (2.9) and (2.13) both the mean and variance of OnL,f;A(L)\mathcal{O}_{n_{L},f;A}^{{}_{(L)}} depend only on the distributions of finitely many components of the vector zA(L)z_{A}^{(L)}. By the inductive hypothesis we therefore find

where the right hand side means that we consider the same collective observable but for z~A(L)\widetilde{z}_{A}^{(L)} instead of zA(L)z_{A}^{(L)}, which exists by Lemma 2.4. Similarly, again using Lemma 2.4, we have

This follows from (2.17) and the inductive hypothesis. Indeed, by construction, conditional on the filtration FL\mathcal{F}_{L} defined by weights and biases in layers up to LL (see (2.3)), the ∣A∣\left|A\right|-dimensional vectors zi;A(L+1),∙z_{i;A}^{(L+1),{\tiny{\bullet}}} are iid Gaussians:

where ΣA(L+1)\Sigma_{A}^{(L+1)} is the conditional covariance matrix from (2.5). The key point, as in the proof with all Gaussian weights, is that each entry of the matrix ΣA(L+1)\Sigma_{A}^{(L+1)} is a collective observable at layer LL. Moreover, since the weights and biases in the final layer are Gaussian for zA(L+1),∙z_{A}^{(L+1),{\tiny{\bullet}}} the conditional distribution of g(zA(L+1),∙)g(z_{A}^{(L+1),{\tiny{\bullet}}}) given FL\mathcal{F}_{L} is completely determined by ΣA(L+1)\Sigma_{A}^{(L+1)}. In particular, since gg is bounded and continuous, we find that

2 Tightness: Proof of Proposition 2.2

In this section, we provide a proof of Proposition 2.2. In the course of showing tightness, we will need several elementary Lemmas, which we record in the §2.2.1. We then use them in §2.2.2 to complete the proof of Proposition 2.2.

In particular, for some constant CC depending on T0,λT_{0},\lambda, we have

The second Lemma we need is an elementary inequality.

Let a,b,c≥0a,b,c\geq 0 be real numbers and k≥1k\geq 1 be an integer. We have

Further, breaking into cases depending on whether 0≤a≤10\leq a\leq 1 or 1≤a1\leq a we find that

Combining (2.21) with (2.22) we see as desired that any a,b,c≥0a,b,c\geq 0

The next Lemma is also an elementary estimate.

Fix an integer k≥1k\geq 1, and suppose X1,…,XkX_{1},\ldots,X_{k} are non-negative random variables. There exists a positive integer q=q(k)q=q(k) such that

The proof is by induction on kk. For the base cases when k=1k=1, we may take q=1q=1 and when k=2k=2 we may take q=2q=2 by Cauchy-Schwartz. Now suppose we have proved the claim for all k=1,2,…,Kk=1,2,\ldots,K for some K≥3K\geq 3. Note that 1≤v⌈(K+1)/2⌉≤K1\leq v\lceil{(K+1)/2}\rceil\leq K. So we may use Cauchy-Schwartz and the inductive hypothesis to obtain

where q=max⁡{q(⌈12(K+1)⌉),q(K−⌈12(K+1)⌉−1)}q=\max\left\{q\left(\lceil\frac{1}{2}(K+1)\rceil\right),q\left(K-\lceil\frac{1}{2}(K+1)\rceil-1\right)\right\}. ∎

The next Lemma is an elementary result about the moments of marginals of iid random vectors.

We will use the following result of Łatała [33, Thm. 2, Cor. 2, Rmk. 2]. Suppose XiX_{i} are independent random variables and pp is a positive even integer. Then

where ≃\simeq means bounded above and below up to universal multiplicative constants. Let us fix a unit vector u=(u1,…,un)∈Sn−1u=\left(u_{1},\ldots,u_{n}\right)\in S^{n-1} and apply this to Xi=uiwiX_{i}=u_{i}w_{i}. Since wiw_{i} have mean and pp is even we find

Note that for each k=2,…,pk=2,\ldots,p we have

Hence, using that log⁡(1+x)≤x\log(1+x)\leq x we find

Note that for 2<t2<t, there is a universal constant C>0C>0 so that

Thus, there exists a constant C′>0C^{\prime}>0 so that

Combining this with (2.23) completes the proof. ∎

The final Lemma we need is an integrability statement for the supremum of certain non-Gaussian fields over low-dimensional sets.

The proof is a standard chaining argument. For each y∈T1y\in T_{1} write Πk(y)\Pi_{k}(y) for the closest point to yy in a 2−k2^{-k} net in T1T_{1} and assume without loss of generality that the diameter of T1T_{1} is bounded above by 11 and that Π0(y)=y0\Pi_{0}(y)=y_{0} for all y∈T1y\in T_{1}. We have using the usual chaining trick that

By Lemma 2.8, there exists qq depending only on kk so that for any q1,…,qkq_{1},\ldots,q_{k} we have

We seek to bound each expectation on the right hand side in (2.26). To do this, write

Note that the supremum is only over a finite set of cardinality at most

for some c>0c>0 depending only T0,λT_{0},\lambda. This is because, by assumption T1T_{1} is the image of T0T_{0} under a λ\lambda-Lipschtiz map and Lipschitz maps preserve covering numbers. Thus, by a union bound,

But for any y∈T1y\in T_{1} and any s>0,p≥1s>0,p\geq 1 we have

Putting this all together we find for any p≥2max⁡{q+2,cn0}p\geq 2\max\left\{q+2,cn_{0}\right\} that

Thus, substituting this into (2.25) yields

Appealing to Lemma 2.9 completes the proof of Lemma 2.10. ∎

2.2 Proof of Proposition 2.2 Using Lemmas from §2.2.1

Fix xα≠xβ∈T1x_{\alpha}\neq x_{\beta}\in T_{1} and define

Write WiW_{i} for the ii-th row of WW and bib_{i} for the ii-th component of bb. Since ab≤12(a2+b2)ab\leq\frac{1}{2}(a^{2}+b^{2}) and σ\sigma is absolutely continuous, we have

Since σ′\sigma^{\prime} is polynomially bounded by assumption (1.2), we find by Markov’s inequality that there exists an even integer k≥2k\geq 2 so that for any C>1C>1

Our goal is now to show that the numerator in (2.30) is bounded above by a constant that depends only on T0,n0,λT_{0},n_{0},\lambda. For this, let us fix any y0∈T^y_{0}\in\widehat{T} and apply Lemma 2.7 as follows:

Substituting this and the analogous estimate for ∣W⋅ξ∣4\left|W\cdot\xi\right|^{4} into (2.30), we see that since all moments of the entries of the weights and biases exist, there exists a constant C′>0C^{\prime}>0 depending on λ,T0,k\lambda,T_{0},k so that

Substituting this into (2.31) and taking CC sufficiently large completes the proof of Lemma 2.6. ∎

with probability at least 1−ϵ/(L+1)1-\epsilon/(L+1). Thus, the image

Proceeding in this way, with probability at least 1−ϵ1-\epsilon we that

Since nL+1n_{L+1} is fixed and finite, this confirms (2.27). It remains to check the uniform boundedness condition in (2.1). For this note that for any fixed xβ∈Kx_{\beta}\in K by Lemma 2.4, we have

Thus, by Markov’s inequality, ∣∣zβ(L+1)∣∣\left|\left|z_{\beta}^{(L+1)}\right|\right| is bounded above with high probability. Combined with the equi-Lipschitz condition \eqrefE:equi−lip\eqref{E:equi-lip}, which we just saw holds with high probability on KK, we conclude that for each ϵ>0\epsilon>0 there exists C>0C>0 so that

References