The edge of chaos: quantum field theory and deep neural networks
Kevin T. Grosvenor, Ro Jefferson
Introduction
The remarkable empirical success of deep neural networks in domains ranging from image classification to natural language processing has far outpaced our theoretical understanding. In practice, these networks are generally treated as black boxes, with state-of-the-art performance relying primarily on heuristics and trial-and-error rather than on any fundamental theory BahriRev ; Roberts:2021fes . Indeed, this state of affairs has even led some experts in the field to compare modern machine learning to alchemy Alchemy . As has been pointed out in, e.g., Roberts:2021lll , top-down insights from fields such as cognitive science may be fruitfully complemented by bottom-up approaches from theoretical physics towards uncovering the fundamental principles governing learning and intelligence in minds and machines.
Recently, physicists have begun exploring the parallels between deep neural networks and quantum field theory, leading to a number of interesting works on what is sometimes called the NN-QFT correspondence Roberts:2021fes ; Yaida:2019sjo ; Dyer:2019uzd ; Halverson:2020trp ; Erbin:2021kqf ; Maiti:2021fpy .We use the phrase “quantum field theory” to make manifest the precise analogies with well-studied techniques in QFT that we explore below, but we emphasize that there is nothing intrinsically “quantum” about the networks under consideration. It would perhaps be more accurate to say “statistical field theory,” or one may simply think of this as QFT in Euclidean signature, cf. sec. 5. We also thank Dan Roberts and Sho Yaida for pointing out that previous works that we have included under the “NN-QFT” umbrella did not involve a bona fide field theory in the sense of a well-defined continuum limit, nor a direct correspondence between elements of a neural network and elements of a QFT (see appendix A), but we have here used the term broadly as in Halverson:2020trp to refer to the general philosophy of establishing parallels between neural networks and QFT, with the aim of using the latter to understand the former. See also bondesan2021hintons for an orthogonal approach from quantum mechanics. One key aspect of this approach is the fact that many modern architectures admit a Gaussian process limit: as the number of neurons (per layer, i.e., the width) , the network is described by a Gaussian distribution, and hence may be modelled as a free (non-interacting) quantum field theory. This is essentially a consequence of the central limit theorem (CLT); see for example Halverson:2020trp for a brief summary and plentiful references of the Gaussian process limit in this context, including the original thesis neal1995 .
While the infinite-width limit provides an analytically tractable approximation that has led to important progress (see, e.g., jacot2020neural ; lee2018deep ; sohldickstein2020infinite ; yang2021feature ), it fails to capture crucial aspects of real-world networks which must of necessity be of finite width. For example, the lack of interactions – i.e., intralayer correlations – in the Gaussian limit implies that representations in these idealized networks do not evolve during gradient-based learning Roberts:2021fes ; LeeXiao . (Note that this does not contradict the well-known universal approximation theorem for infinite-width networks: the latter merely states that there exists some network (that is, a particular set of weights and biases) that approximates a given function, but says nothing about the learning dynamics. The statement that the representation (specifically, the neural tangent kernel) does not evolve at infinite width means that if the network does not happen to start in the correct state, it will never evolve to it; see sections 6.3.3, 10.1.2, or chapter 11 of Roberts:2021fes for details.) Additionally, for networks exhibiting an order-to-chaos phase transition (see below), the location of the critical point predicted by the central limit theorem differs from empirical observations schoenholz2017deep ; Erdmenger:2021sot . Understanding networks away from the infinite-width limit is therefore of central importance for any theory of deep learning with practical applications.
Conveniently, the art of carefully backing away from the limit has been perfected by physicists in the form of quantum field theory. The machinery of Feynman diagrams, which we shall later employ, provides a compact and efficient means of computing perturbative contributions from interactions. In the most basic terms, the idea is to solve a complicated (non-Gaussian) theory by perturbing about the free (Gaussian) theory in terms of some small parameter. This allows one to effectively turn on interactions – i.e., finite-width effects – in a controlled manner, which in statistical language corresponds to the inclusion of higher cumulants. In finite-width networks, such cumulants appear in subsequent layers in a manner precisely analogous to what one observes along the renormalization group (RG) flow, due to the marginalization over neurons in previous layers (i.e., tracing-out hidden degrees of freedom) Roberts:2021fes ; roDLRG .
Another important aspect that many neural networks share with models well-studied in physics is the phenomenon of criticality, by which we mean a continuous phase transition separating an ordered and disordered or chaotic phase. The idea that the edge of chaos holds key advantages for computation and learning dates back to at least the classic paper by Langton LANGTON199012 , and was further developed in the influential work by Bertschinger, Legenstein, and Natschläger Bert1 ; Bert2 . It has since appeared in many fields ranging from theoretical neuroscience and neurophysiology RCHIALVO2004756 ; BeingCrit ; relevant , to biological and complex systems Chialvo_2018 ; chialvo2019controlling , and was also explored in an early model of bioplausible neural networks in mistakes . More recently, poole2016exponential ; schoenholz2017deep ; xiao2018dynamical ; chen2018dynamical ; gilboa2019dynamical ; Erdmenger:2021sot demonstrated that networks initialized at criticality are trainable to far greater depths than those lying further into either phase. To see this, one identifies a correlation length that sets the scale at which correlations between local degrees of freedom – e.g., the activations of two different neurons, or the spins of two different magnetic dipoles – decay with separation or depth through the network. Intuitively, both the ordered and chaotic phases are bad for learning, since correlations about the input data – i.e., the structure one is attempting to learn – are energetically damped or washed-out by noise, respectively. However, a critical point is characterized by a divergent correlation length, which allows the persistence of structured information on all scales. We refer the interested reader to BahriRev ; TongSFT ; roCrit for topical reviews of this phenomenon, or to Ashton for a beautiful visual demonstration in the case of the 2d Ising model.
In this work, using techniques from statistical field theory helias2019statistical ; Sompo , we explicitly construct the quantum field theory corresponding to a general class of deep neural networks encompassing both recurrent (RNN) and feedforward (multilayer perceptron, MLP) architectures, in order to leverage well-known tools in perturbative QFT to compute the finite-width corrections to the correlation length. An original motivation was to determine whether these effects would shift the location of the critical point, as one expects on both theoretical and empirical grounds. Perhaps surprisingly, while the effective correlation length does appear to receive corrections from all orders in the perturbative expansion, the location of the critical point itself appears unchanged. This conclusion should however be regarded with skepticism, since the theory breaks down at the critical point itself. As is generally the case in QFT, the perturbative analysis is only valid at weak coupling, which we shall see corresponds to a constraint on the variance of the weight initializations, . Consequently, the resulting theory is only strictly valid in a subregion of phase space that shrinks as we approach the edge of chaos from the ordered phase. Nonetheless, the theory exhibits a remarkable similarity with the well-known vector model tHooft:1973alw ; Brezin:1972se ; Coleman:1985rnk ; brezin1993large ; tHooft:2002ufq ; Zinn-Justin:2002ecy ; Moshe:2003xn that we elaborate on in more detail below, and offers a new perspective on the nascent NN-QFT correspondence that may be fruitfully developed further.
As this work was nearing completion, the excellent book Roberts:2021fes appeared, which shares a similar philosophy of applying ideas from physics to understand the structure and properties of deep neural networks, and in particular also considers perturbative corrections in depth/width; see also the previous work Yaida:2019sjo . Relative to our approach, the main difference is that they did not develop a correspondence with any QFT; rather, the finite-width corrections in their case are seen to arise by starting with the Gaussian distribution of preactivations in the first layer, and fixing the couplings in subsequent layers to track the appearance of effective interactions, i.e., higher cumulants that appear as a result of marginalizing over hidden degrees of freedom, in the analogy with RG mentioned above. In contrast, we are concerned with the bulkThat is, we work away from the input/output layers. At a technical level, we shall assume time- (i.e., layer-) translation invariance in order to facilitate obtaining closed-form expressions for the correlation functions. behavior of the network after the layer-to-layer Hamiltonian has reached a steady state (which must occur some finite depth after the input layer), and the exact form of the interactions are included in the full action of the QFT by construction. A central feature shared by our analyses is that the ratio of depth to width, , is the perturbative expansion parameter; accordingly, we likewise work in the limit with fixed (though in our case this quantity, specifically , is dimensionful, and hence one must include an order-1 constant, as will be discussed in detail below). As observed in Roberts:2021fes , this is the regime of most deep neural networks used in practice. We view our approaches as complementary, and warmly recommend Roberts:2021fes for readers interested in a pedagogical treatment of many ideas at this interface of theoretical physics and machine learning. We note that finite-width corrections to the feature kernel were also considered recently in zavatoneveth2021asymptotics .
Another RG-inspired approach appeared in the earlier work Halverson:2020trp , which was further developed in two more papers Erbin:2021kqf ; Maiti:2021fpy that appeared while this project was underway (see also antognini2019finite ). As alluded above, Halverson:2020trp explores the relation between QFTs and the Gaussian process (i.e., ) limit of neural networks in detail. The basic idea from Wilsonian effective field theory is to write down the most general action consistent with the assumed symmetries and properties of the model (e.g., translation symmetry, locality), and fit the associated couplings (that is, coefficients of the interaction terms) by experiment. In particular, they posit an initially Gaussian action, and show that the deviations from the infinite-width limit can be accurately modelled by the addition of a quartic interaction term, whose contribution is suppressed by . The irrelevance of higher (e.g., six-point) interactions can be understood from an RG perspective, which was developed more thoroughly in Erbin:2021kqf . These authors also introduced the phrase “neural network phenomenology” to emphasize the fact that the form of the action in all of the above approaches is posited on general grounds, with couplings determined either by experiment or by tracking effective interactions. In contrast, the key difference in the present work is again that we explicitly construct the action of the field theory from the (stochastic) differential equation governing the network dynamics, so that the precise form of the interactions follows directly. It would be very interesting to understand the relation between these approaches in more detail, and in particular whether applying the RG techniques in Halverson:2020trp ; Erbin:2021kqf to the type of bottom-up model we develop here would lead to further insights into criticality in deep neural networks.
Turning to the titular edge of chaos, we are inspired by the aforementioned works poole2016exponential ; schoenholz2017deep ; xiao2018dynamical ; chen2018dynamical ; gilboa2019dynamical ; Erdmenger:2021sot examining criticality in various deep network architectures. However, while many of these papers used the phrase “mean-field theory”, they did not actually rely on any MFT analysis: as mentioned above, Gaussianity arises simply as a consequence of the central limit theorem (CLT). In fact, one important lesson of our analysis is that the correlation functions obtained from the CLT do not agree with the MFT for random networks, even in the infinite-width limit. For the aforementioned works, the distinction is inconsequential, since both approximations appear to yield the same condition for criticality,For this reason, it is unclear whether the CLT result for the critical condition in schoenholz2017deep and related works effectively includes the corrections we compute in sec. 4, since we expect that the latter do not alter the Gaussianity of the theory; see sec. 5 for further discussion. but it is essential to the development of a rigorous NN-QFT correspondence. Prior to considering finite-width corrections in sec. 4, we shall first show in sec. 3 that we recover the criticality condition in the previous works from a true MFT treatment, which we will subsequently reaffirm as the tree-level contribution in perturbative quantum field theory in sec. 4. For readers unfamiliar with statistical field theory, we emphasize that the use of MFT in this context is itself not a novel development: the basic approach we employ was pioneered in the classic paper Sompo , and we have benefited greatly from the pedagogical review helias2019statistical .For other applications of MFT to artificial intelligence, see for example Gabri_2020 and references therein. Nonetheless, while the basic starting point of our work is not fundamentally new, we have considerably developed the perturbative analysis, and applied a variety of old ideas in novel ways.
The remainder of this paper is organized as follows: in sec. 2, we construct the QFT corresponding to the general class of networks under consideration using standard techniques in statistical field theory helias2019statistical . Specifically, we will obtain the partition functionFor our machine learning readers, this is essentially the moment-generating function, with which we can – in principle – compute any observable we wish to know; see roCum for a pedagogical exploration of these relationships. describing self-averaging random networks, and derive equations for the correlation functions in the MFT approximation. In sec. 3, we will construct the double-copy MFT system in order to examine the growth of correlations between different samples of networks from the ensemble. The edge of chaos is then determined by the point at which the largest Lyapunov exponent becomes positive, which yields precisely the criticality condition discussed above. In sec. 4, we go beyond the MFT regime, and consider the perturbative expansion of the full (single-copy) QFT. We first explicitly compute the tree-level propagator – that is, the two-point correlation function in the absence of quantum correctionsAgain, to preempt any sensationalist headlines that “quantum physics explains deep learning”, we remind the reader that there is nothing fundamentally quantum at work: the language is meant to be evocative of the loop corrections with which physicists are intimately familiar, but these are purely classical corrections due to the cumulants associated with turning on interactions in a controlled manner. – identify the correlation length, and show that the condition for criticality precisely agrees with the previous MFT result, as expected. We then compute the first two correction terms to the two-point correlation function, which contribute at and , respectively, for both linear (subsec. 4.3) and non-linear (subsec. 4.4) models.Our analysis is phrased at the more general level of recursive neural networks (RNNs), which includes feedforward networks as a simple case. In the former, the total time serves as an infra-red (IR) cutoff, while for feedforward networks this role is played by the total length , as in Roberts:2021fes . Here the use of Feynman diagrams vastly simplifies the computations: we will develop an elegant diagrammatic recursion relation, which enables us to sum the infinite series of diagrams at each order to yield a finite contribution. With the loop-corrected propagator in hand, we are then able to identify the finite-width corrections to the effective correlation length in the weak-coupling regime. We close in sec. 5 with some discussion and comments on future directions. For convenience, we have summarized the elements of the nascent NN-QFT dictionary in appendix A. Appendix B contains some identities used in evaluating the Feynman diagrams encountered in sec. 4.4, while appendix C contains explicit expressions for the loop-corrected correlation length for nonlinear models which were deemed too lengthy to include in the main text.
The statistical field theory approach
Rather than writing down a general action and fixing the couplings to match empirical results, we will instead employ a bottom-up approach, in which we start with the differential equation describing an RNN – which includes vanilla feedforward neural networks (i.e., multilayer perceptrons, MLPs) as a special case – and explicitly construct the corresponding quantum (statistical) field theory. As mentioned in the introduction, the approach we follow has a long and diverse history, and we shall draw heavily from the pedagogical review helias2019statistical . Indeed, the main technical differences are the inclusion of the data-dependent term, and the explicit constant in (2). That said, the presentation in helias2019statistical focuses on many other topics that are not relevant to our goal, and we have distilled the full derivation in order to make the paper self-contained, as well as to establish our notation.
The basic strategy is as follows: starting with the stochastic differential equation describing the update rule for the network, the probability for a particular sequence of network states is obtained by marginalizing over the stochasticity. We then introduce an auxiliary field to impose the constraint that the network continue to satisfy the update rule while allowing the fields to fluctuate (i.e., to go off-shell). Upon adding source terms, we obtain the partition function in a form familiar to physicists, from which we can in principle compute any quantity of interest, e.g., correlation functions. In subsec. 2.2, we make the theory explicit by introducing an ansatz for the trainable parameters, which are similarly elevated to fluctuating field variables. The MFT in subsec. 2.3 is then simply the leading-order saddlepoint approximation of the resulting QFT. Various technical assumptions introduced along the way are summarized in appendix A.1. Throughout this work, we have endeavored to make each step in the construction as transparently explicit as possible, in the hopes that the reader will find the length more pedagogical than daunting.
Our starting point is the continuous-time formulation of a recurrent neural network (RNN) as a stochastic differential equation (SDE) lim2021noisy
The SDE (1) is interpreted under the Itô discretization convention as
The field-theoretic approach begins by constructing a moment generating functional for the state of the system, i.e., the partition function. Proceeding with (1), the first step is to observe that if the noise is drawn independently for each timestep, then the probability of a particular path (by which we mean the sequence of states ) may be written asStrictly speaking, we should view this as a conditional distribution on a particular sequence of inputs , since this will be treated as external data below.
where we have included the possibility of specifying a non-zero initial condition in the form of the Kronecker delta term. The Dirac delta in (4) thus imposes that the state satisfies the given SDE (3) at time , given the noise and the solution at the previous timestep (i.e., the process is Markovian).
Now, given the integral expression for the Dirac delta function,
From the first line, one immediately sees that differentiating with respect to yields the moments of , e.g., \partial_{j_{t}\eta}Z[j]\big{|}_{j_{t}=0}=\langle h_{t}\rangle_{h} and so on.
where the operator insertions , need not be time-ordered.
2 Self-averaging random networks
Recall that the function appearing in the partition function (12) depends on the trainable parameters , cf. (2). Henceforth we shall consider a class of so-called random neural networks, in which these parameters are initialized as independent Gaussian variables,
and similarly for . Note that as usual, the variances of are scaled by the width, to ensure that the activations at subsequent layers remain at large . The partition function then describes an ensemble of networks, and we expect that in the large- limit, the observables of interest are captured by the ensemble average . The idea is that the network is self-averaging, in that while the particular instantiation varies from one realization of , to the next, the physical properties of the ensemble should approach that described by the average over the (random) couplings, provided the fluctuations around the mean values are sufficiently small. This statement is closely related to the central limit theorem, and becomes exact as helias2019statistical .As we will discuss in more detail below, the CLT does not appear to yield equivalent results to MFT in the present case, even in the infinite-width limit. The ensemble average is then
where for each are the normalized Gaussian probability density functions (16), and similarly for , which we have absorbed into the measures , for compactness. Substituting the rule (2) into the partition function (12) and integrating over the disorder and the bias , we obtain
and the integration over and is contained in the interaction term
where in going to the third line, the exponential has factorized into a product of moment generating functions for and , with an implicit sum over repeated neuron indices . Due to the i.i.d. condition (15), each of these is a simple product of Gaussian integrals, e.g.,
In the large- (i.e., infinite width) limit, the value of these contributions will approach their respective means. The mean-field theory approximation to the ensemble average is then obtained as the leading-order saddlepoint contribution from , treated as independent fields that we allow to fluctuate.
To proceed, we enforce the constraints (24) via delta functions as in (4),
3 Mean-field theory approximation
We now perform the aforementioned saddlepoint approximation, from which the MFT is obtained at leading order. Let us express the partition function schematically as . Provided the exponential decays sufficiently rapidly, the integral will be dominated by the minimum value . The saddlepoint approximation simply consists of Taylor expanding about the minimum:
where the prime denotes (functional) differentiation with respect to , and the linear term vanishes by definition, i.e., . Note that while this may be a very poor approximation to when , it will still give an accurate approximation to due to the exponential suppression of higher-order terms. If we take only the leading-order contribution, then the partition function is given entirely by the prefactor, . It then remains simply to determine .
Substituting the Gaussian cumulant generating function (33) into the MFT action (32), we have
The propagators are then obtained as the matrix elements of the Green function for the matrix operator , defined as the right-inverse
which has three non-trivial components; denoting , we have
Note that : what this tells us is that the expression for the transition amplitude from to must be determined self-consistently from these equations, i.e., that the correlation depends on the correlations in the other variables, via the interactions induced by the disorder we integrated out above.
Let us now assume that the system exhibits time translation symmetry, so that . This allows us to swap derivatives à la . Performing this little slight of hand on the first of the equations above yields
If we then act on the third equation with , we obtain
Given our assumption of (time) translation invariance, it is convenient to define . Then the off-diagonal terms are given by
while for the diagonal term, (41) becomes
This is the analogue of eq. (147) in helias2019statistical ; the treatment above simply extends this framework to the stochastic recursive networks of interest. Upon setting , , and setting all the variances to 0, our result reduces to the non-stochastic, purely feed-forward case considered in the seminal work Sompo1988 , based on the early MFT analysis in Sompo1982 . While we believe this to be the historical origin of the (use of the phrase) “mean-field approximation” that has recently appeared in the machine learning literature, we emphasize again that this does not necessarily correspond to the limit actually used in poole2016exponential ; schoenholz2017deep and related works. We shall show this explicitly in the next section, where correlation functions in the infinite-width limit retain an contribution, which for some parameter values represents a substantial correction to the tree-level (MFT) result.
Now, our primary interest is in the propagator . In particular, in section 3, this solution is taken as a background around which chaotic fluctuations are treated perturbatively (not to be confused with the perturbative quantum field theory treatment in sec. 4). However, we do not actually require an explicit solution to (43); it suffices to obtain expressions for , , and in order to express the equation for in a more tractable form.
To that end, observe that , are Gaussian random variables with covariance matrix . That is, dropping the individual neuron indices and adopting the shorthand , these are drawn from the bivariate normal distribution with
where in the second equality we have introduced the compact notation , (for ). Furthermore, is simply the expectation value of a particular function of these variables, namely , with respect to this distribution:
where is the bivariate normal measure obtained by diagonalizing (44),
where the Pearson correlation coefficient is defined as . If we then define the new integration variables , such that
where is the (factorized) standard Gaussian measure,
By Price’s theorem Price for Gaussian processes,Perhaps the easiest way to see this is to work with the original bivariate measure (46), which satisfies . Since the fall-off of the Gaussian measure ensures vanishing boundary terms, integration by parts can then be used to move the derivatives to . Since (50) cannot depend on the choice of integration variables, this proves the theorem for as well. See also Papoulis theorem 6.7. we may express this as
The motivation for this is that it allows us to define the potential
in terms of which we can express the differential equation (43) for the autocorrelation as
where , and the prime denotes differentiation with respect to . This is now a self-consistentInsofar as the value determines the potential via (51). kinematic expression for , where the left-hand side is a time-derivative of the kinetic term, hence the identification of as the potential energy. In this language, the delta function enforces a discontinuity in the velocity at . As alluded above, we do not require an explicit solution to this expression, but the form will prove convenient below.
The edge of chaos
In the previous section, we obtained a second-order differential equation for the two-point correlator , (43). With this in hand, we now wish to determine where the edge of chaos lies in phase space. There are at least two ways of proceeding: one is to explicitly solve for by making some assumptions about the various correlators on the right-hand side. We will do this in the perturbative expansion in sec. 4, which then allows us to identify the correlation length in the presence of finite-width corrections. However, if one is content with MFT at , then there is an alternative way forwards that does not require a solution for at all: the basic idea, as pioneered in, e.g., Sompo , is to examine the response of the system to fluctuations, whose stability is governed by the largest Lyapunov exponent. Following helias2019statistical , our strategy is to construct a double-copy of the mean-field theory obtained above, and consider the correlator between neurons in different copies as an infinitesimal fluctuation about the MFT correlator between neurons in the same copy. At the most basic level, the idea is to fix to be the same in both copies – which have initially identical parameters – and examine how (or rather, ) differs as a function of time; a similar strategy was used in poole2016exponential ; schoenholz2017deep . A formal parallel with the time-independent Schrödinger equation emerges, which relates the largest Lyapunov exponent to the ground-state energy of the system; as we will see, the onset of chaos corresponds to the point at which the ground-state energy becomes negative.
Denote the MFT correlator , and extend this notation to two identical copies of the system in the obvious way, namely , where label the copies (so that reduces to the single-copy case considered above). Note that at large ,
as a consequence of the self-averaging behavior discussed above. We may then consider the mean-squared distance between two trajectories (i.e., between two identically-prepared copies of the system):
Note that in general ; only at equal times do we have , so that reduces to the usual mean-squared distance. Our goal in this section is to understand the dynamics of . In particular, we expect that if the system is chaotic, should have a characteristic divergence governed by the largest Lyapunov exponent.
Since the copies are initially uncoupled, we may immediately write down the partition function analogous to (9)
We now proceed as above, adding labels to the auxiliary variables in (24) in order to accommodate the additional mixed fields with , i.e.,
We now perform the saddlepoint approximation as before, and keep the leading-order contribution to obtain the mean-field result. The minimality condition yields the following equations of motion for the auxilliary fields:
Substituting the minima (61) into the action, we obtain the MFT for two identical copies of the network:
where, in an extension of the single-copy notation (35), we have defined
We now seek the propagators of the double-copy MFT, defined as the matrix elements of the right-inverse of :
This yields the following three equations for the non-vanishing components of :
Formally, this is the same result we obtained in the single-copy case, cf. (41); the difference lies in the sum over all in the action (62), which picks up cross-correlations between the copies. As before, it is convenient to express these equations for the components of in terms of the temporal difference (where possible), so that the off-diagonal elements are given by
while the non-zero diagonal element becomesWe note that this corresponds to eq. (167) of helias2019statistical ; we have merely derived it via the standard approach.
Observe that the components of this equation are precisely (43), as expected, and hence the solution for the autocorrelation is the same as for the single-copy case above. We may say that in this case, the system resides at a fixed point at which the trajectories for all time, and hence (since ).
Our interest is then in the stability of this fixed point. In particular, instability to fluctuations will cause to grow, indicating chaotic dynamics. Conversely, a stable fixed point implies that the system eventually relaxes back to , so that perfect correlation is restored. Following helias2019statistical , our strategy will be to take the autocorrelation – given implicitly by (69) with , which is equivalent to the single-copy solution (43) – as a fixed background solution, and compute the cross-correlator to linear order in the expansionWe have changed to rather than to avoid confusion with the copy labels , so that henceforth .
where is some small expansion parameter.
To proceed, it is convenient to introduce the following notation for the two-point correlators of :
where , are the components of the covariance matrix of the single-copy in (44), and , are the corresponding components for the double-copy:Note that the lower off-diagonal component is .
We can then Taylor expand (69) and identify terms order by order to determine the linear contribution . On the left-hand side, we simply substitute (70) for . On the right-hand side, we take the input data to be the same for both copies, i.e., , so that and ; then we only need to expand :
where the second line follows via Price’s theorem (50), with . Upon substituting these expansions into (69), we recover (43) at , which leaves
2 The largest Lyapunov exponent
We have found that at linear order in fluctuations about the fixed point, the mean-squared distance between copies (54) is
with given by (74). We now wish to solve (74) in order to determine how the distance behaves as a function of time.
Since (74) is precisely of the same form as that in helias2019statistical , we may simply rephrase their derivation here. We begin by expressing (74) in terms of the “lightcone coordinates”
To solve this equation, we make the separation ansatz
One can think of this as a Euclidean analogue of the usual plane-wave ansatz with phase factor ; the significance of the constant will be discussed below. We then have
and is the second derivative of the potential defined in (51):Note that is the potential energy for the autocorrelation , while is the potential energy for the fluctuation amplitude .
Note , and hence the potential term , depend on via the autocorrelation .
Formally, (80) is the time-independent Schrödinger equation with Hamiltonian
Since we seek normalizable (i.e., bound) states, and the energies of bound states are quantized, the solution will be characterized by a discrete set of eigenvalues,
As per our ansatz (79), this in turn implies a discretized set of characteristic or Lyapunov exponents which control the growth of . In particular, the growth rate will be governed by the largest Lyapunov exponent, given by the ground state energy :
The unstable regime is characterized by , so that grows exponentially with time. Note that if , this will true for all possible values of the ground state energy . Therefore, in order to study networks at the edge of stability, we will henceforth consider the case with (cf. helias2019statistical , which considered ). We must then determine the point at which the ground state energy becomes negative, . Following helias2019statistical , our strategy will be to construct a normalizable solution with , and derive a necessary condition on the existence of lower-energy ground states.
To proceed, observe that if we differentiate (52) with respect to for , and denote , we obtain
where the second step follows from applying the chain rule to . Comparing this with the Schrödinger equation (80), we see that is an eigensolution of everywhere except at , with . To construct a valid solution for all , we must address the discontinuity at the origin caused by the delta function in (52). To do so, define
and impose smoothness at (i.e., that belong to differentiability class ). Denoting the approach to zero from above and below respectively by , the latter condition amounts to the constraint that :
Therefore, the existence of such a solution requires
This is a necessary but not sufficient condition for the solution to be valid: since has zero total energy by construction, we must also impose that the potential be less than or equal to zero at the extremum given by (89); see helias2019statistical for an in-depth discussion on this point. Recall from (83) or footnote 29 that the potential for this solution is . Hence,
When this inequality is saturated, must be the ground state with . Away from saturation, the energy of the ground state is . Intuitively, as the potential becomes more negative, there is more room for lower-energy states. While this does not technically suffice to show the existence of a strictly negative energy ground state, it does provide a necessary condition on the edge of stability for the system, since the corresponding Lyapunov exponent is greater than or equal to zero, . In fact, this is consistent with the condition identified in previous literature: recalling the definition of the two-point correlator (cf. (71)) in terms of the Gaussian integral (48) with (i.e., ) we may equivalently express (90) in the form
where the second integral over has evaluated to unity. Note that the quantity on the left-hand side is exactly that denoted in poole2016exponential ; schoenholz2017deep and serves as a probe of stability for the fixed-point in their analysis. Comparing this expression with eq. (7) in poole2016exponential (which is eq. (5) in schoenholz2017deep ), we see that our result precisely recovers the condition for chaos in Gaussian random networks, which corresponds to and , cf. (2). However, as alluded above, and will be discussed in more detail below, mean-field theory is not exact in the infinite-width limit; indeed, it is perhaps surprising that the MFT result (91) agrees with the result from the CLT obtained in the aforementioned works. In general, we must consider the full QFT, in which the MFT correlator represents only the tree-level contribution.
Before turning to perturbative QFT in the next section, we note that in chen2018dynamical it was reported that in the case of RNNs, the injection of time-series dataThe authors of chen2018dynamical refer to this as “noise” from the inputs, but we have reserved that term for the true noise encoded in , cf. (1). destroys the ordered phase, and consequently there is no order-to-chaos phase transition. This arises due to an extra factor that appears in their analogue of (91) containing possible correlations in . In our case, however, these are contained in the correlator , which does not affect the condition for criticality in the MFT approximation (though it does of course affect the explicit form of ). Further studies are therefore needed to elucidate the potential role of the data , or the introduction of other forms of noise more generally, in modifying the edge of chaos in different network models.
Perturbative corrections
As discussed in the introduction, the network becomes a Gaussian process in the infinite-width () limit. At finite , deviations from Gaussianity require the addition of interaction terms, corresponding to the fact that higher cumulants no longer precisely vanish. In the language of quantum field theory, these correspond to loop corrections to the leading-order or tree-level result above, which have the potential to shift the “classical” edge of stability, which is the chief object of interest here. Indeed, empirical evidence schoenholz2017deep ; Erdmenger:2021sot suggests that the critical point in real-world networks is noticeably displaced from the CLT prediction, which has practical relevance for initialization. Even away from criticality, the correlation length – which sets the depth scale beyond which trainability sharply falls off – deviates substantially from the large- result—see for example fig. 5 in schoenholz2017deep , fig. 2 in xiao2018dynamical , or fig. 6 in Erdmenger:2021sot . It is therefore of significant practical as well as theoretical interest to quantify the deviations from Gaussianity in networks of finite width.
To that end, in this section we will compute both the leading and subleading corrections to the two-point correlator , which will then allow us to identify the loop-corrected correlation length for small . As discussed in the introduction and in more detail below, the result holds only at weak ’t Hooft coupling,In fact, we shall require both and to be sufficiently small relative to , as will be explained in more detail below. and the perturbative expansion assumes ; the latter is the regime of practical relevance for modern deep neural networks Roberts:2021fes .Note that the authors of Roberts:2021fes referred to as the “chaotic limit”, which differs from the use of that terminology here, but the important point is that the perturbative expansion cannot be performed if is large. Conversely, simply recovers the Gaussian theory, including the correction to be computed in subsec. 4.3.1. In this sense, the contributions are not bona fide interactions, but finite-temperature fluctuations in the statistical ensemble, as we explain in sec. 5. It would be interesting to explore these connections to theory in more detail, e.g., to see whether the analysis can be extended beyond the perturbative (weak-coupling) regime we consider here.
Since we will compute the two-point function explicitly, there is no need to employ the double-copy system used in the MFT analysis in the previous section; it is sufficient to examine the correlation length in a single copy of the full QFT for small times. The reason for this can be traced back to (70), which requires that the fluctuation be small relative to the background solution . In other words, we do not demand that the two systems will diverge for all time, or in the single-copy theory, that the correlator will grow exponentially without bound. In the language of RG, we are considering infinitesimal perturbations about the critical point by some relevant operator(s), which will trigger a flow to a new, non-trivial fixed point in parameter space. This behavior was observed in poole2016exponential , see in particular fig. 2 therein, in which networks in the chaotic phase converge to some parameter-dependent fixed point near, but not necessarily at, .
To keep the theory as simple as possible while still capturing the essential aspects, we shall consider the partition function for the (single-copy) theory with . Relative to (24) however, it is extremely convenient to modify the auxiliary field to
and swap the attachment of the prefactor within the delta function-imposition of the constraint, i.e.,
A consistent solution to these equations is
Note that is not constrained by its eom (which reduces to ), and thus we are free to select an arbitrary vev for the inputs; for simplicity (since we shall have in mind a nonlinear activation function with , see below), we shall take as well. We may then expand the fields around these background values by making the following shifts in the action:
and similarly for . Our theory is then
where the quadratic part of the action is – again with implicit sums over repeated neuron indices –
In the next subsection, we will obtain the propagators for the theory given by (105). Unlike the MFT analysis in the previous section however, we will solve for the propagators explicitly, in order to proceed to obtain the finite-width corrections in subsections (4.3.1) and (4.3.2).
To find the propagators of the theory (105), we first introduce the two-component fields
where the operators with are defined as
As in the previous MFT analysis, the propagators are obtained as the elements of the Green functions (i.e., the inverse operators) for the operators appearing in the Gaussian part of the action, cf. (65). For , we have
Substituting in the definition of the delta function (118) and imposing that the result hold for all momenta , we find the momentum space propagator
and for the non-zero off-diagonal element,
where we have performed the same trick as before, cf. (67), again assuming translation invariance of the correlators. As a quick sanity check on the consistency of our expansion, observe that (after adjusting for the difference in normalization, cf. (92)) we have precisely recovered the MFT results (68), (69), as expected, since the former correspond to the tree-level contribution in perturbation theory.
As usual, it is more convenient to work in momentum (i.e., frequency) space, so we shall proceed by Fourier transforming the above expressions to obtain the momentum space propagators. We shall use the standard high-energy theory convention in which factors of are attached to the momentum measure, and denote the momentum space functions with hats, i.e.,
where in the last step, the only non-vanishing contribution comes from the pole at ; recalling from our discussion below (85) that in the regime of interest, this lies in the lower half-plane, and hence requires in order to close the contour; the negative sign then arises from encircling the pole clockwise. Similarly,
Turning now to the off-diagonal term (116), we must contend with the fact that, from the eom (99), will be a function of when we perform the perturbative expansion. For consistency, let us work to fourth order in as above; then
The first term on the right-hand side is simply , where . For the remaining term, since we are expanding around a free (Gaussian) theory,That is, the full expectation values are computed with respect to the interacting theory (105), but when evaluating them, we expand in terms of the coupling (i.e., ), so that the leading contribution is (123), in which the expectation values are taken with respect to the Gaussian theory. we can apply Wick’s theorem to write this as a sum of products of two-point functions. Let us temporarily drop the neuron indices, and instead employ the shorthand ; then the four-point function simplifies to
where may be solved for self-consistently, cf. (135). Therefore, to this order,
where for compactness we have defined the effective variance
for reasons which will become clear in subsec. 4.3.1. Note that if we had worked to next (sixth) order in , we would obtain an additional term of the form
where the combined combinatoric factor for the -point correlator is (to see this, choose any point at random; there are ways to contract it, which leaves choices for the second pair, and so on). This would lead to a cubic differential equation for , in place of the much simpler linear equation (128). To avoid this – and the associated explosion of possible Feynman diagrams – we have limited ourselves to only the first two orders in the Taylor expansion; we comment on this truncation in more detail below. Hence, substituting (125) into (116), we have
where for simplicity we will henceforth assume that is constant (which we can absorb into , and hence set ), and have written as though it were time-translation invariant.This amounts to assuming a sufficient degree of temporal uniformity in the injected data. In fact, we shall shortly assume that is time-independent, in order to perform the inverse Fourier transform. In principle, this is not required for the theory, but one must otherwise provide the form of the correlator as an input to the following analysis to obtain a closed-form result. Fourier transforming with respect to , this then becomes
We can then Fourier transform this back to real space, where the sign of determines the contour (i.e., which pole we encircle):
The second term is trivial on account of the delta function:
Lastly, by the convolution theorem, the third and final term can be expressed as
This is as far as we can evaluate this term without further information about the data . To obtain a closed-form result, we shall take to be time-independent; cf. footnote 36. Summing these contributions, the result for the propagator is
The form of the two-point correlator allows us to immediately identify the correlation length at tree-level:
As discussed in the introduction, the critical point is defined by , which occurs when
As an aside, we note that we may express the first term in (135) as
2 Feynman rules
In order to compute the perturbative corrections to the tree-level propagator (135), we must deduce the Feynman rules for the theory. To that end, the interaction part of the action (107) may be written in terms of the two-component fields (108) as
At a glance, it appears that the only -dependence is in the coupling constants, such that a diagram with vertices will contribute at . However, in analogy with the vector model, we must sum over any internal neuron indices, which implies that internal loops have the potential to give rise to additional factors of that disrupts this simple counting. In fact, it turns out that the leading correction to the tree-level (MFT) result is due to an infinite number of cactus diagrams that all contribute at ; we shall begin with these in the next subsection.
Therefore, working to second-order in the expansion of (but keeping terms only up to quartic order in ; see above), we have the following Feynman rules for the theory:
propagating to : \quad G_{hh}(t-s)\;=\raisebox{-0.35pt}{\includegraphics[scale={1.2}]{Ghh.png}}
Consider all possible arrangements at the vertices; e.g., include a copy of the diagram with only if this produces a distinct diagram.
Combinatorics: consider all possible ways to contract internal legs; e.g., in the 5-pt vertex , we can form loops by connecting pairs of incident legs in 3 different ways, cf. (124).
For each internal loop with a free neuron index, we have an additional factor of .
Add the contributions of all diagrams that contribute at the desired order in .
3 Linear models
To organize the presentation, we shall first consider perturbative corrections for linear models, i.e., . That is, we consider only the leading term in the Taylor series (123). In this case we have only the 3-pt vertices in (141), which greatly simplifies the diagrammatic expansion. In the next two subsections, we will compute both the leading and subleading contributions to the two-point function (138). We will then generalize the analysis to include the 5-pt vertices corresponding to the subleading term in (123) in subsec. 4.4.
As in the model tHooft:1973alw ; Brezin:1972se ; Coleman:1985rnk ; brezin1993large ; tHooft:2002ufq ; Zinn-Justin:2002ecy ; Moshe:2003xn , the leading correction to the tree-level propagator comes from an infinite series of so-called cactus diagrams, each of which contributes at . Intuitively, these quantify fluctuations from typicality in the ensemble of networks, and vanish in the limit where the ’t Hooft coupling , as we discuss in more detail in sec. 5. Crucially, the contribution from these diagrams remains even as , and thus represents an important effect in networks of arbitrary size.This is the reason that MFT is generically not what one obtains from the CLT, contrary to some indications in the literature. In the Ising model for example, MFT amounts to replacing each degree of freedom, together with its nearest-neighbor interactions, with an effective degree of freedom – the so-called mean field – in which these interactions have been averaged over. This ignores long range interactions which become important near criticality, and is known to give incorrect results in low dimensions, while the CLT is agnostic about both. The first three diagrams in the infinite series, to linear order in , are illustrated in fig. 1. Note that in each case, we have choices for the neuron that runs in each loop, which precisely cancels the factor of from each pair of vertices.
where in the penultimate step, we have made the change of variables
where the exponent denotes that the base is convolved times. However, since convolutions correspond to products in momentum space, it is more convenient to instead work with the Fourier transform
where whether , , are position or momentum space propagators will be indicated by the argument ( or , respectively).
We can now sum the infinite series to obtain the correction from all cactus diagrams in fig. 1. Including the bare propagator as , we have
where we have substituted in the momentum space propagators (120), and summed the geometric series subject to the convergence condition
3.2 Infinite 𝒪(T/N)𝒪𝑇𝑁\mathcal{O}(T/N) mushrooms
The subleading correction to the Gaussian propagator is the order at which we observe manifest finite-width effects. Here we will also find that the more relevant expansion parameter is rather than itself, in agreement with Roberts:2021fes . We shall first detail the computation of a generic, one-particle irreducible (1PI) diagram, and then use this to compute all contributions (including non-1PI diagrams) to the propagator at below.
whence the previous expression simplifies to
While it is possible to sum the infinite series to obtain the contribution from all 1PI mushroom diagrams, this does not yet include all contributions at , since we must also consider non-1PI diagrams—in particular, those involving an arbitrary -loop cactus, followed by an mushroom (or the reverse). To accomplish this task, let us introduce the following elegant recursion relation, as a warm-up for the nonlinear analysis in subsec. 4.4: let represent the propagator including all and corrections, and denote its momentum space expression by . The dots at either side of the circle are to indicate that the latter subsumes half of the external legs, that is, that the external legs can be either , or , . The diagrammatic recursion relation for is then
where and respectively denote the and contributions to in the expansion
where we have pulled out the factor of simply to make dimensionless, cf. (169). Note that includes the bare propagator, and that we can generate any or diagram by recursively substituting for , on the right-hand side. For example, to generate an cactus, we place the first diagram on the right-hand side (the bare propagator) in the space for in the second diagram; we then take this diagram and substitute it in again to generate an cactus, and so on. On the second line, the circles containing indicate that only diagrams may be recursively inserted there, since the presence of the explicit mushroom means that any further mushroom insertions would make the resulting diagram .
Now, given the analysis of the 1PI diagrams above, the computation of this expression is quite straightforward. Consider the second diagram on the right-hand side: this takes the form of (147) with and . Similarly, the sum of the two diagrams on the second line is given by (152) with and . Thus, (153) reads
which we may solve order by order in . Expanding as in (154) and solving for the contribution, we find
which is precisely the correction obtained in (148), as expected. We then use this to solve for the contribution:
We emphasize that this expression includes all perturbative contributions to linear models (that is, to leading order in the Taylor expansion (123)), including diagrams involving an infinite number of loops.
3.3 The loop-corrected correlation length
Adding the results (156) (equivalently (148)) and (157), we have thus found that the loop-corrected two-point function in momentum space, to leading order in , is
With the loop-corrected propagator in hand, we would now like to identify the correlation length in the presence of both fluctuations and finite-width effects. At a glance, this is obstructed by the mixed exponential and polynomial dependence on . However, as explained at the beginning of this section, we do not require an expression for the correlation length valid at all times; rather, it suffices to identify an effective correlation length that governs the evolution of the correlator for small (i.e., small perturbations about the critical point, as in schoenholz2017deep and related works, as well as our MFT treatment in sec. 3).A sufficient but not necessary condition for this is , which certainly holds near the critical point. Thus, we can identify a meaningful loop-corrected correlation length by first expanding the exponential in (159), collecting terms order by order in , and re-exponentiating to obtain a purely exponential term of the form . To linear order in , we find
where we have identified the loop-corrected correlation length
4 Nonlinear models
We thus have a recursive sequence of cactus diagrams, which we can evaluate using a similar recursive approach to that in the previous subsection. Let represent the loop-corrected propagator including all possible cacti, allowing both 3-pt and 5-pt interactions, and denote its momentum space expression by . That is, we expand
cf. (154), where here we have used in place of to denote the inclusion of the subleading term in the expansion of the nonlinearity (123). Then the recursion relation for all diagrams is simply
Implicitly, we also include a copy of the last diagram with the branch attached to the other side of the stem.
Now, the second diagram on the right-hand side is of course simply (147) with and . For the third, let us first consider the case in which we replace the in the branch with the bare propagator, resulting in a closed propagator or petal. Since the temporal endpoints are the same, this simply contributes a factor of , where the negative sign stems from the coupling , cf. (142). The factor of in the latter is precisely cancelled by the 3 ways to contract the two pairs of -legs. Importantly, note that the neuron index in the petal is the same as that in the adjacent loop; this can be seen from the indices in (107). Therefore, while petalous vertices gain an extra loop, they contribute at the same order as their apetalous counterparts in the perturbative expansion. Therefore, this diagram – plus the copy with the petal attached to the opposite side of the stem – is given by (147) with and .
This lesson generalizes to the case in which the branch is a recursive cactus of arbitrary complexity: the delta functions within it will collapse all times to those at the vertex at which the branch joins the trunk, so that the entire branch is evaluated at . Hence, the expression for the third diagram on the right-hand side of (163) – plus the copy with the branch attached to the opposite side of the stem – is formally identical to the petalous cactus just described, but with , where includes the bare propagator . The recursion relation (163) then reads
where we have defined the loop-corrected variance
Solving this for and recalling from (151) that , we obtain the -corrected propagator, including both 3-pt () and 5-pt () interactions:
Comparing this with (148), we see that the effect of the 5-pt vertex is to modify the effective variance we defined in (126) to (165).
4.2 Branching mushrooms
As before, we must regulate this divergence by imposing a finite cutoff , so that the above diagram contains a contribution that scales like . At a glance, this appears pathological, but instead highlights the fact that the weak coupling regime is governed by both and ; in particular, in addition to the convergence condition encountered above, we also require . To see this, observe that by dimensional analysis, the various components of the action (106) have the following mass (inverse time) dimensions:
Unfortunately, these last four diagrams are considerably more complicated than any we have considered above, since they involve a mixture of products and convolutions in both position and momentum space. Let us begin with the first, in comparison to the momentum space expression (150): we have the same factor of from the top-most together with the remaining (regulated) integral over in the cap, and a net factor of ( from the vertices, and from the internal neuron index). The two vertices will contribute a collective factor of , while the two vertices will contribute , cf. (141). The propagators themselves carry a net factor of , and the freedom to twist the in the stem contributes an additional factor of 2. Note that we do not have the freedom to twist the horizontal propagator, but can instead change the orientation of the cap, so that diagrams with either or running along the underside top-most bulb are allowed. We also have a combinatoric factor of from the possible ways of connecting the three internal propagators, and – as mentioned above – can insert a self-similar to any of them. Finally, the external legs contribute as before, so that
For the next diagram in fig. 4, we again have a collective factor of from the vertices, a factor of from the two propagators plus the freedom to twist the one in the stem, an from the external legs, and (in momentum space) from the possible orientations of the cap, and in place of each of the internal propagators. However, while the top-most propagator again evaluates to , we are not quite free to integrate over the remaining vertex position (time), since this will be encountered by the intervertex propagator, which will thus be evaluated at (similar to how the delta functions collapse the temporal indices on the branches to , in the diagrams above). Since will generically be some sum of exponentials plus a constant piece, this last is the only divergent contribution that requires us to regulate the integral, and hence carries the sought-after factor of . The other, exponential terms will carry no such factors,More precisely, the regulators in all convergent integrals are exponentially suppressed; see sec. 5. and hence contribute at order in the expansion; we therefore drop them, and keep only the constant part of this convolution, which we denote (so-named because the constant part of the propagator is always the bias-dependent term). Additionally, the other difference relative to the previous diagram is that in place of from the intervertex combinatorics, we have : 3 for which of the legs from one vertex to connect with (3 additional choices for) a leg from the other vertex, and 2 ways to connect the remaining two pairs. Thus, including the overall factor of from the internal neuron index, we have
The third diagram in fig. 4 offers a moment’s respite: it is precisely the same as (172). We thus move to the fourth and final diagram, which has from the cap and a net factor of from the vertices and neuron index as in the first diagram, and the combinatoric factor of as in the second/third. As for the propagators themselves, observe that by labelling the vertices on either side of the stem , and the vertex on the underside of the horizontal propagator , the two possible orientations of the diagram readRecall from (112) that the temporal indices on either end of the propagators are identified by virtue of the delta functions.
which one can see by traversing the diagram from left to right, and
which one can see by traversing the diagram from right to left, with and . Therefore,
where in the last step we have Fourier transformed to momentum space. Prior to this, since in the first line is evaluated in position space, (121) and (122) imply
With the form of these 1PI diagrams in hand, we can now proceed to consider all possible contributions to from both 1PI and non-1PI diagrams. To do so, let us extend the diagrammatic notation introduced in subsubsec. 4.3.2 to include 5-pt interactions, so that represents the loop-corrected propagator up to and subleading order in the expansion of the nonlinearity (123), and denote its momentum space expression by , cf. (162). We then have the recursion relation
Let us consider the various contributions to this expression in turn. From our analysis in the previous subsubsection, we know that the first two cactus-like diagrams are collectively given by (147) with and . Similarly, for the third cactus diagram, we have (since allowing would make this diagram ) and ; note that we are including both possible choices for the side at which to attach the branches. The next two diagrams on the second row are collectively given by (152) with , , and . Similarly, the first diagram on the third row is given by (167) with ; note that there is no possibility to attach branches, since the only explicit 5-pt vertex is already exhausted. Again, we are including both possible orientations for these three diagrams. Lastly, we have the four diagrams in fig. 4 with and no further branching possibilities, which we have already given in (171), (172), and (175). Collecting results, we thus obtain the following recursive expression for :
which is readily solved, at least formally, to yield
where for compactness we have suppressed the dependence, it being understood that these expressions are written in momentum space. The expression (179) is the analogue of (155) in the presence of 5-pt () interactions arising from the subleading term in the Taylor expansion of the nonlinearity, cf. (123), and can likewise be solved order by order in . At , (179) reduces to
which is precisely the -corrected propagator obtained in (166), as expected.
with replaced by (180), and where for compactness we have absorbed the pseudo-initial conditions into
where was defined in (165). It then remains to compute , , and , which we do in appendix B; the results are given in (201), (202), and (204), respectively. The explicit expression for that we obtain upon substituting these into (181) is however exceedingly lengthy, so we refrain from displaying it here. Nonetheless, we now have all the ingredients to obtain the loop-corrected propagator to linear order in , cf. (162), which enables us to define an effective correlation length in the next subsubsection.
4.3 The loop-corrected correlation length
Having computed the loop-corrected propagator (162) with (180) and (181), including both the leading and subleading terms in the expansion (123) of the nonlinearity , we now wish to repeat the analysis in subsubsec. 4.3.3, in which we identified a loop-corrected correlation length by expanding in small . That is, upon inverse Fourier transforming to
we again obtain an expression involving multiple different exponentials. Following the logic below (159), we therefore approximate the loop-corrected propagator for small as
where the constants , are given by
and the loop-corrected correlation length at small is
where in the present case, one finds that . Note that taking the limit in simply isolates the constant part of the corresponding expression, since all exponentials are decaying. As above, the explicit expressions for the components of on the right-hand sides of (185) and (186) are rather lengthy, but are given in appendix C.
Discussion
In this work, we have explicitly constructed the quantum (statistical) field theory corresponding to the general class of networks described by (1), which includes both RNNs and deep MLPs. Relative to previous works on the nascent NN-QFT correspondence that take a phenomenological perspective, we have pursued a more fundamental or bottom-up approach, using well-known methods in statistical field theory to derive the partition function for an ensemble of random networks. This has led us to several intriguing parallels with well-studied vector models and the perturbative expansion in Yang-Mills/QCD, and it would be interesting to study these formal similarities in more detail. For convenience, we have summarized the elements of the NN-QFT dictionary in appendix A.
At a physical level, the interpretation of the corrections is fairly clear. As expected on general grounds, and explored quantitatively in Halverson:2020trp ; Yaida:2019sjo ; Roberts:2021fes , the Gaussian () limit does not suffice to describe real-world networks at finite width, whose distributions (i.e., actions) include effective interaction terms corresponding to higher cumulants. These finite-width effects are of both theoretical and practical relevance, and the depth-to-width ratio was identified as an important emergent scale in Roberts:2021fes . Note however that, as mentioned in sec. 4.4, we can in fact have corrections at any order in with . While those with – such as the term that we have computed here – will dominate in the limit, terms with can become comparably important in networks of finite size; for example, if and , then . Physically, the effect is the same: non-Gaussianities accumulate with increasing depth, and are suppressed with increasing width. However, this suggests that the precise interplay between the two may be more subtle.Note that since is dimensionful, the depth of the network is measured in units of , cf. (169); e.g., an MLP has layers.
Additionally, we have identified a dominant correction that persists even at infinite width. It is unclear whether this is implicitly included in the results from the central limit theorem (CLT) in poole2016exponential ; schoenholz2017deep and related work, but regardless represents an important and novel contribution in the field-theoretic approach. While we have not proven this,An in principle straightforward albeit computationally prohibitive way to do this would be to compute all (or a convincing number of) higher-point correlators and show that they reduce to sums of products of two-point correlators as per Wick’s theorem. we expect – based on consistency with the CLT, i.e., the vanishing of interactions in the limit – that the contribution does not alter the Gaussian nature of the distribution, and instead merely changes the variance from (126) to (165). Therefore, at a physical level, the contribution does not represent bona fide interactions per se, but rather quantifies fluctuations from typicality in the ensemble of networks.
To see this, note that since we are working in Euclidean signature, our action is proportional to , where is the inverse temperature.Throughout this paper, we have followed the theoretical physics convention of working in civilized units, in which fundamental constants such as , , , etc. are unity. As in the model, the ’t Hooft coupling in order that the action be dimensionless, and thus higher-order terms in the perturbative expansion carry more factors of . In QFT in Lorentzian signature, this reflects the fact that loops represent quantum effects, with the order of the loop diagram in some sense quantifying the distance from classicality at which that effect arises. For QFT at finite temperature, i.e., statistical field theory, , and the loops can be understood as thermal (statistical) fluctuations around the zero-temperature () background. The ’t Hooft coupling thus controls the size of the quantum/statistical fluctuations, so that the classical/MFT result is recovered in the limit when the weight initialization becomes deterministic. We emphasize again that this effect is present even when , as can be seen from (161) or (186). That is, MFT is not obtained by simply taking , but rather by ignoring all fluctuations, including both genuine interactions (which disappear in the Gaussian limit, and are hence quantified by powers of ) and thermal fluctuations (which are extrinsic, i.e., independent of system size, and hence ).
As mentioned in the introduction, the location of the critical point predicted by the CLT (i.e., at ) differs from empirical observations; see for example fig. 6 of Erdmenger:2021sot , in which the theoretical result – derived via the CLT, following poole2016exponential ; schoenholz2017deep – lies at , whereas the empirical result for networks with appears to lie at . However, as we remarked beneath (161) and (186), we are unable to observe a shift in the location of the critical point at either or . In the linear case, (161) does not exhibit any new divergences, while in the nonlinear case, the new divergences in (186) lie to the right of the tree-level divergence, i.e., further into the strong coupling regime; see the discussion at the end of sec. 4.4. In principle, we see no reason that these finite-width corrections could not have shifted the edge of chaos to the left, into the weak coupling regime. In practice however, this may represent a limitation of the present perturbative approach.
There are many ways in which the first-principles approach pursued here could be further explored and improved. For example, we have only computed the (quantum-corrected) two-point function , whereas in principle there is no obstruction to computing higher-point functions or more complicated observables of interest. This leads to the question of whether such a theory is renormalizable: we have shown that the infinite series of corrections to the two-point function converge at weak coupling only to linear order in , and while we see no obvious obstruction to convergence at higher orders, we have not proven that this continues to hold. Should it fail, then we have an effective theory valid only for sufficiently low-point correlators, to finite order in perturbation theory. By analogy, general relativity provides an extremely useful effective theory of gravity in the weak coupling regime, but is non-renormalizable, and hence breaks down at sufficiently high energies; the relevant question is then whether the neural networks of interest happen to lie at the metaphorical black hole singularity.
A related question is whether nonperturbative effects might extend this analysis beyond the weak coupling regime, and in particular, whether these shift the edge of chaos or otherwise contribute to the correlation length. The potential presence of nonperturbative effects was examined recently in zavatoneveth2021exact , which demonstrated that the exact output distributions for linear and ReLU networks with zero bias exhibit heavy tails. Of course, for networks with sufficiently small (such at those with illustrated in fig. 1 therein), we expect that perturbation theory will cease to apply, since the network is no longer Gaussian in the first place. However, while Gaussianity is recovered in the limit, one can see the appearance of heavy tails for larger values of (e.g., with in fig. 1 of zavatoneveth2021exact ) that cannot be accounted for by an effect of the type we compute above, i.e., a simple shift in the variance (note that in the aforementioned figure, the finite- curves intersect the Gaussian in two locations). It would be interesting to study this in more detail, e.g., to quantify the minimum network size at which the perturbative approach breaks down, or to investigate whether other nonperturbative effects persist at large .
In closing, we offer this work as a first-principles contribution to the rapidly emerging NN-QFT correspondence Roberts:2021fes ; Yaida:2019sjo ; Dyer:2019uzd ; Halverson:2020trp ; Erbin:2021kqf ; Maiti:2021fpy . While there is still much to learn, we believe this and similar approaches from theoretical physics can shed light on the underlying mathematical principles of deep neural networks, and have the potential to provide practical guidance for the development of more sophisticated machine learning techniques.
Acknowledgements
It is a pleasure to thank James Giammona, Jim Halverson, Anindita Maiti, Dan Roberts, Keegan Stoner, and Sho Yaida for comments on a draft of this manuscript, as well as Johanna Erdmenger, Boris Hanin, and Soon Hoe Lim for discussions. K.T.G. acknowledges financial support from the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under Germany’s Excellence Strategy through the Würzburg-Dresden Cluster of Excellence on Complexity and Topology in Quantum Matter ct.qmat (EXC 2147, project id 390858490), as well as the Hallwachs-Röntgen Postdoc Program of ct.qmat. K.T.G. has also received funding from the European Union’s Horizon 2020 research and innovation programme under the Marie Skłodowska-Curie grant agreement No 101024967.
Appendix A The NN-QFT dictionary
As explained in the introduction, we have generally used the phrase “NN-QFT correspondence” broadly to refer to the general philosophy of using techniques and ideas from QFT to understand deep neural networks. However, we have gone beyond previous works in directly constructing a bona fide field theory, allowing us to precisely identify various elements on either side of the correspondence:Here, we use the term “layer” to refer to either a single layer of an MLP, or a single timestep of an RNN. “Width” then refers to the number of neurons per layer; see the comment below (3). Let us also reiterate that the basic tools used in our construction are not new, and can be traced at least back to the work by Sompolinsky helias2019statistical ; Sompo ; nonetheless, the correspondence has not been explicitly stated in these terms.
Let us add a few comments about this NN-QFT dictionary. First, as remarked in the main text, the perturbative analysis applies in the regime with fixed. This agrees with previous results Roberts:2021fes that the behaviour of the network is controlled by the ratio of depth to width, rather than either of these parameters independently: if with fixed, the network becomes a Gaussian process (i.e., we turn off interactions), but is effectively shallow; conversely, if with fixed, we lose all control of the perturbative expansion (i.e., effective interactions become too strong). As emphasized in Roberts:2021fes , networks of practical interest typically have small depth to width ratios, and therefore we expect real-world networks to fall within this perturbative regime. (Note that also serves as an IR regulator on otherwise divergent integrals, cf. footnote 39).
However, our analysis has uncovered an ambiguity in the “infinite-width limit” as presented in the literature, namely that previous works have often conflated MFT and the CLT. As mentioned in the introduction, and in more detail in footnote 37, these are generally inequivalent. In the present case for example, the MFT (i.e., tree-level) result is not recovered in the infinite-width limit, since the correction is independent of (i.e., intensive, as opposed to extensive). See sec. 5 for further discussion of the physical interpretation of the and contributions.
As an additional constraint on the validity of our analysis, we found that (and, in the case of non-linear models, also ) must be sufficiently small in order for the infinite series of Feynman diagrams appearing at each order in the perturbative expansion to converge, cf. (149) and (170). This is consistent with our observation that plays the role of the ’t Hooft coupling, since the latter controls the weak/strong regime in Yang-Mills theory: perturbative quantum field theory of the kind we explore here is only tractable in the weak coupling regime, i.e., . At a practical level, this implies that our perturbative analysis can only be applied to the left of the critical point in , and breaks down precisely when . For non-linear models, there is an additional constraint that be sufficiently small to avoid an IR divergence, but this is less restrictive, since networks of practical interest typically have .
Finally, we emphasize that the continuum limit is taken in the layer index, not the neural index, as explained in footnote 12. There is no meaningful notion of distance within a given layer. However, there is a meaningful notion of distance between different layers, since the networks under study do not admit skip connections; i.e., layer 1 is “closer” to layer 2 than it is to layer 3, and any effects of layer 1 on layer 3 are mediated via layer 2. At a technical level, this is the reason that has units of length in our analysis, while remains dimensionless, hence the appearance of in the dimensionless expansion parameter (in contrast, both and are dimensionless in Roberts:2021fes , no continuum limit was taken, and there is no field theory). Physically, this gives rise to the correspondence between a single layer of neurons in the network, and an -component field in the QFT. Hence we have -dimensional field theory, consistent with the dimensional analysis in (169). The interaction between components as the field evolves in time (i.e., in subsequent layers) is ultimately responsible for the free neurons running in the internal loops of the Feynman diagrams discussed in sec. 4.
Throughout this work, we have endeavored to be as explicit and as general as possible; however, we have made a number of technical assumptions along the way, most of which can be grouped into two broad categories:
Conditions which are required for our approach to hold on mathematical grounds (e.g., for convergence), and
Assumptions which are not strictly necessary, but which facilitate obtaining closed form expressions.
In (128), we have restricted to the case where the stochastic function is constant, and absorbed this into . While this is not strictly necessary for the single-copy theory, it would be unusual for the stochasticity to depend on the current state of the system. In any case however, it is necessary for the double-copy analysis in sec. 3 that the stochasticity be common to both copies, which generically cannot hold if it depends on the states, cf. (55). If this were not the case, then the systems would deviate due to the different stochasticities even in the ostensibly ordered phase.
Turning now to the second category above, let us briefly summarize the various simplifying assumptions here, in roughly chronological order:
In (15), we have chosen to work with (Gaussian) random networks largely for analytical tractability, since the Gaussian form of the initializations allows many integrals to be performed analytically. We note however that this class of networks is a standard ansatz in the literature, e.g., poole2016exponential ; schoenholz2017deep , and in any case the distributions will be Gaussian at large due to the CLT.
To simplify (22) and related integrals, we have further taken the means of the initializations in (15) to be zero, again consistent with the cited literature.
In (33), we chose the stochastic increments to also be i.i.d. Gaussian with zero mean. In principle however, one could consider other forms of stochasticity; see lim2021noisy and references therein.
In solving for the Green functions, we assumed that the system exhibits time translation symmetry, cf. (40). While this seems like a strong assumption, we expect it to hold to good approximation in the bulk of the network (i.e., away from the boundaries), where our analysis is concerned; see sec. 1, in particular footnote 2.
In the perturbative analysis in sec. 4, we have set in order to focus on the core elements of the theory. As seen previously, these can in principle be treated analogously to and .
We have set the vevs in (102) to zero, as this is the simplest solution to the equations of motion. However, other, more complicated vacua may exist; we hope to explore this in future work. Similarly, we have chosen the data to also have zero vev, though this is an external parameter that one can fix based on the dataset in question.
Another important restriction on the generality of our perturbative treatment is that we have restricted to as our activation function. As discussed in Roberts:2021fes (see in particular sec. 2.2), in order to achieve criticality, we desire an activation function which is smooth, and which passes through the origin, i.e., . While there are other activation functions which posses this property (e.g., , SWISH, GELU), none are as widely used as ; insofar as the latter captures the key properties we wish to explore, we have limited ourselves to for concreteness. This however is not strictly necessary, and it would be interesting to consider other activation functions in this context. An alternative route would be to simply work with a generic Taylor expansionWe thank Harold Erbin for this suggestion.
where we have taken for reasons just explained, and rescaled so that the coefficient of the linear term is unity. The case considered herein then corresponds to taking and , and dropping all higher powers of . As discussed below (139), these are negligible for activation functions of this type.
Appendix B Convolution identities
In this appendix, we evaluate the convolutions that appear in the perturbative (finite-width) corrections to the correlation function in subsec. 4.4.2. Specifically, to evaluate (181), we require , , and .
Let us introduce the convenient shorthand notation
where is the inverse Fourier transform; for compactness, we will often suppress the argument of , it being understood that this function acts on momentum space. We now begin by collecting a few useful identities that will enable us to straightforwardly evaluate the rather complicated expressions for the various convolutions in the main text above. First, consider the convolution
which follows from the convolution theorem and (188). Next, consider the convolution of with the product
Lastly, observe that convolving with is the identity operation; that is, if is an arbitrary smooth distribution on momentum space, then
where the in the measure is due to our Fourier conventions in (117). Note in particular that this formula applies for .
In the notation (188), the bare (tree-level/MFT) propagator given in (130) may be written
Observing that and applying (190), we may express this as
On the second line of (196), we have defined the coefficients for later convenience.
With the above expressions in hand, consider the convolution
with the coefficients defined in (196). Applying (189) and (192), this becomes
We can now use this expression to evaluate the three desired convolutions mentioned at the beginning of this appendix.
In the notation (188), , cf. (151). Hence, to compute in momentum space, we simply convolve each term in (199) with . By (189), this will result in convolutions of the form
Recall from the discussion above (172) that refers to the non-exponential part of , evaluated at . That is, we keep only the term in that requires regulating by imposing a cutoff as in footnote 39. From the form of (199) and the inverse Fourier transform (188), we see that the only term that thus contributes is proportional to , hence:
where we have substituted in the coefficient defined in (196).
Finally, to evaluate , we must convolve (199) with (196):
Applying (189) and (192) as before and collecting terms, this becomes
The convolutions (201), (202), and (204) can then be substituted into (181) to obtain an explicit expression for .
Appendix C Explicit expressions for the loop-corrected propagator
In this appendix, we collect the explicit expressions for the various components of the loop-corrected propagator and associated correlation length in the nonlinear case, in the small- approximation (184):
where the constants were defined as
and the loop-corrected correlation length is
where , are the inverse Fourier transforms of (180) and (181), respectively. The former is straightforward and relatively compact:
where the coefficients , , and defined in (196) are
and , were defined in (137) and (195), respectively,
Unfortunately, is substantially more complicated, but can nonetheless be obtained by using the shorthand (188) and associated identities introduced in appendix B to write (181) as an expression linear in , each instance of which can then be trivially inverse Fourier transformed to . After a great deal of algebra, we find that the constant portion appearing in is
As for the value at appearing in the coefficient , let us express the five terms in the inverse Fourier transform of (181) separately. Noting that