Nonperturbative renormalization for the neural network-QFT correspondence

Harold Erbin, Vincent Lahoche, Dine Ousmane Samary

Introduction and outline

Deep learning and neural networks (NNs) have experienced a rapid development in the last decade, with an ever-increasing number of remarkable applications. In many cases, these systems outperform humans and ordinary algorithms. However, there are still many challenges to be solved: in particular, most neural networks work as a black box and require a huge number of examples during the learning phases. More generally, there is no complete theoretical understanding of why deep learning works so well and how to improve it further. For example, it is not clear how training can be made more efficient and fast, how knowledge can be transferred to other tasks or how to choose hyperparameters systematically. This lack of reliability poses, in certain cases, important ethical problems. Indeed, with the growing use of AI for making decisions (for example, in banking, employment, medicine, military, etc.), it is crucial to be able to explain the choices of the AI in a transparent way . Moreover, having a black box is also a drawback for scientific discovery since the goal of science is to interpret and explain, and knowledge can grow only from understanding . Our paper is part of the lively field of explainable AI , where physics do have a role to play .

A natural path for studying neural networks is provided by theoretical physics: it offers an array of tools useful to describe a wide range of complex systems . In the recent years, evidence has accumulated in favor of a scenario involving a particular form of "coarse-graining", making contact with a familiar tool for physicists: the Wilsonian renormalization group (RG).

A macroscopic ideal gas is completely described by the ideal gas law, and a macroscopic fluid is well described by the Navier-Stokes equation. Both equations ignore the microscopic atomic and molecular interactions and provide a coarse-grained description. This idea was fully developed by Wilson’s renormalization group, formalizing the general feature that long range behavior of physical systems does not require an understanding of the nature and interactions of its microscopic building blocks. Through a very impressive argumentation, Wilson showed that it is possible to explain the apparent universality of physical systems near critical points from the observation that, up to the accuracy of physical predictions, the specific microscopic details can be absorbed through a few effective couplings, defining an effective large scale theory. Despite the fact that the RG was born in the era of critical phenomena, it turned out to be a very general framework, largely responsible for the success of field theory descriptions of long range distance physics, both in condensed matter physics and in high energy physics .

The ability of the RG to explain long range universality can be traced from information geometry . The RG coarse-graining is performed on the eigenvalues of the (free) Fisher information metric, which is a local version of the Kullback-Leibler (KL) divergence DKL(p∣∣q)D_{\text{KL}}(p||q) (or relative entropy) and which provides a reasonable measure of distinguishability between two probability distributions pp and qq . From coarse graining, and in absence of singular structures, KL divergence decreases, as well as distinguishability between distributions, as to become smaller than any experimental precision. Beyond this limit, we cannot distinguish the two distributions, as different as they may have been originally. The ability of the RG to extract the relevant features from a large set of interacting microscopic degrees of freedom is a compelling argument for a link with deep learning. In fact, it is natural to expect a relation with any procedure able to extract relevant features from a massive data set, as it is the case, for instance, in principal component analysis (PCA), where some recent works stressed such a connection between signal detection and RG .

In this paper, we aim at developing further the correspondence between quantum field theory (QFT) and neural networks, called the NN-QFT correspondence . The main objective is to provide a description of the NN behavior using the non-perturbative RG and the corresponding effective field theory.In this paper, we will mostly use both “quantum field theory” (QFT) and “effective field theory” interchangeably, the context makes it clear if we speak of the microscopic (ultraviolet, UV) or effective theory (low-energy). We are working in Euclidean signature, in which case the term “statistical field theory” is sometimes used, to make clear that it describes thermal and not quantum fluctuations. However, we will use QFT for uniformity. This positions our paper in a growing tradition of papers describing how the behavior of NNs can be understood through a more and less sophisticated coarse-graining, which can itself be related to a RG. Strong evidence in favor of a correspondence between RG and deep learning has been stressed for Restricted Boltzmann machines (RBM), whose architecture exhibits similarities with the Ising model (a theoretical model for ferromagnets) . Historically, the Ising model was precisely the conceptual cradle of the RG through Kadanoff’s "block-spin" method, which can be viewed as an elementary version of the general Wilson coarse-graining . The use of a field theoretical formalism is not a novelty as well . In fact, this is expected since physics has shown that field theories are a general feature for systems involving emergent collective dynamics. For example, they appeared to provide a good understanding of the qualitative behavior of NNs through the spin-glass formalism .

We follow the correspondence between NN and QFT pioneered by Halverson, Maiti and Stoner . Its originality with respect to other approaches lies in the observation that, under very general conditions, NN with infinitely wide layers are described by a Gaussian process (GP) due to central limit theorem . Realistic architectures never involve an infinite number of hyperparameters NN, and their behavior fails to be well described by a GP. However, this useful but purely theoretical limit allows approaching the non-Gaussian process (NGP) as a perturbation from the large NN limit, which is assumed to receive 1/N1/N corrections which, for NN large enough, can be computed perturbatively. In , the correspondence has been developed in the case of a fully connected network with a single hidden layer of width NN. They developed the field theoretical machinery necessary to describe neural networks (see Appendix A for a summary). This includes computing correlation functions of outputs, obtained in QFT by constructing Green functions from Feynman rules in perturbation theory. Effective interactions (also called couplings) can then be extracted by comparing the NN correlation functions and the QFT Green functions. Finally, they introduced a RG flow from a cut-off on the volume of the input data, from the assumption that the effective field theory must be insensitive to the choice of the volume, up to a global rescaling of couplings entering in its definition. In the QFT language, this corresponds to an infrared (IR, large volume) cut-off: a major difference in our paper is that we will use an ultraviolet (UV, data resolution) cut-off (see Section 2.1.3).

The relation between different effective models can be translated locally through a set of β\beta-functions which describe the evolution of couplings when the cut-off changes, with universal features of the theory emerging from the flow. In this paper, we are aiming at proposing a non-perturbative formalism, based on the Morris-Wetterich equation , to investigate the NN-QFT correspondence beyond the perturbative regime, i.e. beyond the large NN regime. This means that our analysis does not requite the coupling constants to be small and that our equations are given in a 1/N1/N expansion. As mentioned above, our framework differs from the one used in in that we introduce a true partial integration of the degrees of freedom procedure, without any assumption on the expected large volume behavior of the corresponding effective field theory. Among the major differences with respect to the situation with ordinary QFT, the full (effective or IR) 22-point function is known theoretically, including non-Gaussian effects, whereas the free propagator (microscopic or UV) is not known. This unconventional setting allows going beyond standard limitations of the non-perturbative framework, in particular to close the infinite hierarchical system of equation describing the RG flow and keeping the full momentum dependence of the correlation functions following the Blaizot-Mendez-Wschebor method .

Another unconventional aspect concerns the notions of power-counting and locality, which are traditionally inherited from the background space-time (which we will call “data-space” in the case of neural networks). In the case of QFT for neural networks, such a relation appears as an additional hypothesis that no experience motivates "a priori". Other properties such as rotation and permutation invariances of the point components may not make sense for neural network data. However, recent works in the context of background independent quantum gravity have shown that the notions of scales and power-counting are more primitive than that of space-time, and such that locality can be derived from power-counting itself, ensuring moreover that the RG exists and is well-defined. In this article, we will discuss how these ideas can be relevant for neural networks.

A RG can be then constructed by following the standard method, partially integrating on the degrees of freedom and starting with those associated with the highest scales (UV). However, the situation for the NN-QFT correspondence is quite different compared to usual studies of the non-perturbative RG: indeed, we are able to solve exactly the 44-point vertex function while keeping the full momentum-dependence without approximation on the 22-point function (since it is already known exactly). We can also solve almost exactly for the other momentum-dependent nn-point vertex functions (when two momenta are equal and the others vanish, we can also find an exact solution without approximation). In this paper, we consider two versions of the RG. In a first approach, called passive, the notion of scale is fixed by the resolution chosen to describe the data. In that approach, the standard deviation of the hidden weights, σW\sigma_{W}, is viewed as a reference mass scale. The resulting evolution equation provides an explicit realization of equivalence classes of networks having the same output (up to the machine precision), as the data is coarse-grained. In a second approach, called active, the RG flow is constructed by viewing σW\sigma_{W} as a running scale. In such a way, the equivalence class is between networks having the same output, keeping the data resolution fixed. This implies that, for fixed NN, neural networks with different σW\sigma_{W} can be viewed as belonging to the same RG trajectory. In particular, this implies that the renormalization flow can be used to make predictions for any σW\sigma_{W} given the results for one of them. We illustrate this by describing the behavior of the quartic coupling constant of the effective field theory and check numerically the flow equations. In this paper, we focus on the analytic result and we plan to extend the numerical aspects in future works.

In Section 2, we discuss some general concepts about the field theories which may be used to describe neural networks. In particular, we comment on the definition of the data-space, IR and UV regimes, (non-)locality and its consequences on scaling and power-counting. At the end, we describe the passive and active points of views for the renormalization group. In Section 3 and 4, we derive the passive and active RG flow equations respectively. In Appendix A, we review the numerical simulations from and provide some additional details. Finally, Appendix B contains the details of technical computations.

NN-QFT, locality, scaling and RG

In this section, we present the framework of the NN-QFT correspondence proposed in . As explained in the introduction, we focus on the Gaussian networkThe term “Gaussian” refers to the fact that the kernel is a Gaussian kernel, not that we have a Gaussian process in the infinite-width limit. (or Gauss-net) which have a translation invariant kernel. In this section, we first recall the main ideas of the correspondence (some numerical results from are reproduced in Appendix A).

Then, we discuss the role played by non-local interactions.They were not considered in the original version of but additional discussion has been added in a subsequent version during the preparation of this manuscript. In particular, we describe the different ways to relax locality and how this naturally leads to break the rotation invariance of the data. The most general QFT in the latter case are called random tensor field theories (or group field theories), which are generalization of random matrix field theories.

We are also revising the concept of power-counting, preferring a notion intrinsic to the renormalization group compared to the one used in , which is inherited from a background “data-space”. In the latter case, a classical scale dimension is attributed to the data and dimensional analysis is performed by requiring that the action is dimensionless (such that its exponential can serve as a weight in the path integral). However, it is not clear how to extend this notion in the presence of non-local interactions. We introduce two notions of scales which emerge from the analysis: the first is attached to the data and called “working precision”, and the second is attached to the network and called the “observation scale”. We consider two versions of the RG, flowing in these two parameter scales. We conclude the section with a short presentation of the Wetterich-Morris formalism for non-perturbative RG and a discussion about the RG version considered in the reference paper . Note that we voluntary use the same notations and conventions to make the comparison with their results easier.

In , the authors proposed a general QFT framework to describe the statistical behavior of neural networks (NN), working in the function-space rather than parameter-space (which can be viewed as a duality ). The original motivation stems from the observation that neural networks in the infinite-width limit are described by a random Gaussian process : the latter can also be described by a free (or Gaussian) quantum field theory.We refer to for a gentle introduction to QFT with neural networks in mind. When the width is finite, the random process is not Gaussian and one can expect the NN to be mapped to an interacting field theory, which has been checked in .

where the weights WiW_{i} and biases bib_{i} characterize the affine transformation of each layer and σ\sigma acts element-wise. The weights W0W_{0} and W1W_{1} follow centered Gaussian distributions N(0,σW2/din)\mathcal{N}(0,\sigma_{W}^{2}/d_{\text{in}}) and N(0,σW2/N)\mathcal{N}(0,\sigma_{W}^{2}/N) respectively, and both biases b0b_{0} and b1b_{1} are drawn from centered Gaussian distributions N(0,σb)\mathcal{N}(0,\sigma_{b}). The input data xx is a dind_{\text{in}}-dimensional vector, while we take dout=1d_{\text{out}}=1 for the output data for simplicity. As a consequence, W0W_{0} is a (din,N)(d_{\text{in}},N)-matrix, W1W_{1} a (N,1)(N,1)-matrix, b0b_{0} a NN-vector, and b1b_{1} a scalar. The Gauss-net activation is slightly peculiar because it acts as an exponential of the layer output normalized by the data of the previous layer:

Finally, we stress that the NN is randomly initialized and that we will not consider the effect of training.

Information on the neural network can be extracted by considering correlations of the outputs: they are encoded by the “experimental” correlation (or Green) functions Gexp(n)G^{(n)}_{\text{exp}} :

where the statistical averageThe notation ⟨⋅⟩\langle\cdot\rangle should not be confused with the expectation value in QFT: we will always use it to denote the statistical average over a set of networks. is taken over a large number of neural networks with identical NN and parameter distributions. The numerical evaluation of these quantities is explained in Appendix A.

1.2 Large N𝑁N: free field theory

where the factor ZZ ensures that the expression is normalized when integrating over the full functional space

The origin of this Gaussian behavior in the limit N→∞N\to\infty can be traced from central limit theorem: since fθ(x)f_{\theta}(x) is formally a sum of NN identically distributed random terms which self-average. The function f(x)f(x) splits into two contributions:

where fb(x)≡b1f_{b}(x)\equiv b_{1} being essentially NN independent variables and following the Gaussian law N(0,σb)\mathcal{N}(0,\sigma_{b}), whereas fW(x)f_{W}(x) goes toward a Gaussian distribution only for large NN. Formally, it reads as:

where x1x_{1} is given by (2.2) (such that fWf_{W} depends on W0,W1W_{0},W_{1} and b0b_{0}). As stated above, for large NN, one expects that such a quantity self-averages around its mean, and thus that fluctuations are small:

where the last equality follows from the assumptions that initial distributions for θ\theta are centered and non-correlated. Hence, the statistical properties of fWf_{W} are essentially given by a centered Gaussian distribution, up to 1/N1/N corrections. Obviously, the random nature of fWf_{W} is inherited from the initial parameter distribution, however the asymptotic Gaussian behavior arises from the law of large numbers.

The 22-point correlation (or Green) function

is the inverse of the Gaussian kernel Ξ(x,y)\Xi(x,y) which appears in the free action (2.4):

However, according to (2.6), it is also possible to decompose KK as:

where KWK_{W} is the 22-point function associated to fWf_{W}. It corresponds to the Fisher information metric in the information geometry language, and is fixed from the choice of the activation function.

In this paper, we essentially focus on translation invariant kernels KW(x,y)≡KW(∣x−y∣)K_{W}(x,y)\equiv K_{W}(|x-y|), which is achieved by the Gauss-net architecture (2.2), corresponding to the kernel:

where ∣x−y∣:=∑i(x−y)i2|x-y|:=\sqrt{\sum_{i}(x-y)_{i}^{2}} denotes the ordinary Euclidean distance between xx and yy.

the corresponding probability distribution being given by the exponential law P[f]∝e−Skin[f]P[f]\propto e^{-S_{\text{kin}}[f]} in (2.4). The nn-point correlation (or Green) functions are defined as:

In the free theory, G0(n)G_{0}^{(n)} is completely determined in terms of G2(x,y)=K(x,y)G_{2}(x,y)=K(x,y) through Wick’s theorem, and vanishes for nn odd . Hence, this implies that:

1.3 Data-space and momentum space

In this subsection, we discuss some definitions related to the data-spaceAs we will see, its properties may be sufficiently different from the usual space of positions – spacetime – appearing in usual QFT to find another name. corresponding to the neural network input xx, and how they differ from .

It is generally more convenient to work in Fourier (or momentum) space. The allowed momenta p=(p1,…,pdin)p=(p_{1},\ldots,p_{d_{\text{in}}}) lie in the first Brillouin region:

Note that we assume periodic boundary conditions. This is not a problem for LL large enough; for small LL,We will precise “small with respect to what” in a moment. we can simply repeat the data set a large number of times to obtain a large enough effective volume to make the boundary conditions irrelevant.

In the rest of this paper, we use the following definitions.

We call (2N0)din(2N_{0})^{d_{\text{in}}} the discrete volume and a0a_{0} the working precision.

the basis functions eipxe^{ipx} being normalized such that ∑xei(p1−p2)x=N0δp1p2\sum_{x}e^{i(p_{1}-p_{2})x}=N_{0}\delta_{p_{1}p_{2}}. Note that (2.18) holds for any discrete function on the lattice. In the continuum limit, for small a0a_{0} and N0N_{0} large such that LL remains fixed, discrete sums can be replaced by integrals. Moreover, for volume large enough, integrals becomes standard Fourier transform. We call this limit the thermodynamic limit following the standard terminology in physics, and we focus on this regime in our investigations. Taking the Fourier transform of the 22-point function:

where px:=∑i=1din pixipx:=\sum_{i=1}^{d_{\text{in}}}\,p_{i}x_{i}, we get for (2.12):This type of kinetic term is reminiscent of pp-adic string theory .

Note that in this continuum approximation, the Dirac delta δ(p)\delta(p) has to be understood as a shorthand notation for a Kronecker delta (2π)−dinVδp0(2\pi)^{-d_{\text{in}}}\mathcal{V}\delta_{p0}. Translation invariance is crucial to obtain a kernel (2.20) which depends on a single momentum, and reflection invariance implies that it must be a function of p2p^{2}. It would be interesting to understand how to generalize our computations for kernels which are not translation invariant .

Up to O(p4)\mathcal{O}(p^{4}) corrections, the propagator looks like the canonical propagator of a free scalar field theory:

In the QFT terminology, Z0Z_{0} and m02m^{2}_{0} are respectively the wave function renormalization and bare mass. One can rescale the field to to set Z0=1Z_{0}=1, in which case the mass becomes:

The mass mˉ02\bar{m}_{0}^{2} defines the typical mass scale and its inverse defines the (IR) correlation length ξ\xi, or the typical observation scale

Note that the large volume limit is defined with respect to this correlation length, i.e. L≫ξL\gg\xi. Beside the existence of an intrinsic length scale, the system the propagator at large distance behaves like (mˉ02+p2)−1(\bar{m}_{0}^{2}+p^{2})^{-1}.

Assigning the label xx (position space) to the original data and pp (momentum space) to the Fourier conjugate may seem arbitrary. Indeed, while the machine precision provides a natural UV cut-off and an associated identification of the data-space as position space (since UV corresponds to small distances in that space), signals in Fourier space are also represented in the computer up to the machine precision. However, in that case, using the machine precision as a UV cut-off would not match the usual intuition in QFT. Given a translation-invariant kernel, another possibility is to identify the momentum space as the space where the propagator is diagonal, such that the propagator in position space depends on the distance ∣x−x′∣|x-x^{\prime}|.

1.4 Finite-N𝑁N corrections and interactions

For a Gaussian process, as we have seen earlier, correlation functions G(2n)G^{(2n)} for n>1n>1 can be decomposed as a sum of product of 22-point functions thanks to the Wick theorem . For NN large but finite, the distribution is not exactly Gaussian, and the correlation functions do not match with Gaussian predictions.

The deviation of the QFT and experimental correlation functions from the Gaussian case are denoted as:

Note that G0(n)G_{0}^{(n)} are still the large NN Green functions defined in (2.14). Importantly, we identify the exact 22-point function G(2)G^{(2)} with the kernel KK which equals G0(2)G_{0}^{(2)}. Since it contains (quantum) corrections due to the interactions, the free 22-point Green function computed from the kinetic term only is not known (in standard QFT, the converse is true, see Section 2.3.2 for a discussion). We will see in Section 2.2.3 that connected functions ΔGc(2n)\Delta G^{(2n)}_{c} behaves as:

which has also been investigated analytically and numerically in (see also Appendix A).The original version did not contain this discussion which has been added while the present manuscript was in preparation. This scaling is consistent with the fact that the exact 22-point function is independent of NN.

Qualitatively, this is reminiscent of what happens for the Ising model in large dimension. The local magnetization self-averages because the number of closest neighbors is large, and the statistical properties remain (quasi)-Gaussian. For space dimension dd large but finite, thermodynamical quantities can be computed as power series in 1/d1/d, which do not affect universal quantities, as soon as d>4d>4. For d<4d<4 however, the decoupling of physical scales breaks down and the Gaussian approximation is not suitable .

The same scenario is expected to be true for neural networks. For finite NN, the distribution does not obey Wick theorem, and correlations functions receive contributions which do not reduce to products of 22-point functions, and the classical action must include non-Gaussian contributions, i.e. products of ff of degrees higher than 22. However, as soon as NN remains large enough, deviations from the Gaussian behavior are expected to remain small. In the classical action, these corrections materialize as product of mm fields (m>2m>2), that we call interactions:

2 Locality, scaling(s), and power-counting

In this subsection, we precise the class of interactions which are assumed to suitably reproduce the non-Gaussian properties of correlations. Moreover, we discuss the scaling behaviors, especially relevant for the renormalization group investigations in the next section.

However, this makes various assumptions which may not be valid for a general neural network QFT. For this reason, we will make them explicit and explain how to gradually lift them to consider the most general QFT. Deciding which assumptions to use should be dictated by numerical evidences: in particular, it was found in that (2.29) is sufficient for the activation functions and range of input parameters considered there (see Appendix A for more details). This approach can be considered as neural network phenomenology, in the sense that we are writing a model to match observations, but we can also use this model to check theoretical facts such as dualities .

The first assumption is locality of the interaction: the fields appearing in the monomial f(x)nf(x)^{n} can be taken at different points (for simplicity, we consider a single coupling in SintS_{\text{int}}):

This breaks locality because fields at different points in space(time) can interact together. In fact, since gg is a constant, this happens for arbitrarily large distances. Note that this preserves translation invariance xi→xi+ax_{i}\to x_{i}+a.

The next natural step is to replace gg by a coupling function, i.e. a function of space but independent of the field. Going back to (2.29) where all fields are at the same points, we can write a local action with a coupling function:

It was argued in from technical naturalness that g(x)g(x) must be approximately constant since a coupling function g(x)g(x) breaks the translation invariance of the action. However, this is correct only when assuming locality of the action: replacing gg by a coupling function in (2.30) gives the non-local action :

However, translation invariance can be preserved if gg depends only on the distances between the points:

In order for this to make sense, the Fourier transform of the kernel κ(x,y)\kappa(x,y) must be an entire analytic function (with rapid decay if one wants to ensure UV finiteness). This corresponds to a coupling function

Smeared fields naturally appear in string theory and are responsible for its well-behaved UV behavior . In fact, for the Gauss-net, rescaling the field ff to remove the exponential from the kinetic term (2.20) is equivalent to smearing the field (as pointed out earlier by comparing with the pp-adic string ).

However, it is not clear a priori that the data-space possesses this symmetry: it may not be possible to exchange two data components or even consider linear combinations if the components are not homogeneous. Despite the fact that the free theory supports such a symmetry, the rotational invariance has no meaning for a neural network in general. Moreover, symmetries of the free theory can be broken by interactions, which are necessary to fully characterize the system. Conditions under which input and output symmetries can be present has been analyzed in . In general, one can start by assuming no symmetry in order to describe the most general model, and then adapt to what the numerical experiments are indicating.

where xix_{i}, yiy_{i} and ziz_{i} are the components of the 33-dimensional points xx, yy and zz. Note that nothing prevents to use only two points, for example setting y=zy=z and integrating only over xx and yy, or more generally to repeat the same component in any number of fields (for an early example, see ).

Such general theories are too wild and it is hard to make sense of them. A controllable subclass is provided by random tensor field theories . In this case, the fields are tensors, each component of the positions being seen as a (continuous) index and indices can be contracted pairwise only (which is achieved by integrating over the component, since the index is continuous), such that a given component can appear at most twice. An intuitive way to represent it is to assign a color to each component, and Feynman diagrams can be written in terms of strand graphs (generalizing ribbon graphs from matrix models). For instance, a possible quartic interaction for din=3d_{\text{in}}=3 is:

We will see in the next section that tensor field theories are particularly interesting in the RG approach because, under some additional conditions, they possess a natural background-independent power-counting (see next subsection).

We conclude this section by clarifying a subtlety concerning QFT in curved spaces. In this case, the Euclidean (or Poincaré) group is not a global symmetry (symmetries of the action are given by the isometry group of the background space) and one may ask what is the difference with tensor field theories. The point is that this group is still a local symmetry (general relativity can be seen as gauging the Poincaré group) such that the properties discussed above continue to hold. Indeed, one can always consider the tangent space associated to a point: since it is isomorphic to flat space, it means that the coordinates are still homogeneous.

2.2 ΛΛ\Lambda-scaling and power-counting

We are aiming to construct a field theory which admits a well-defined RG flow. In standard QFT, the rigorous construction of such a flow requires essentially three basics ingredients: 1) a scale decomposition, 2) a locality principle, 3) a power-counting.

Let us illustrate heuristically on a simple example how contractibility and power-counting allows fixing the scaling dimension of couplings. Consider the following classical action:

for some cut-off Λ\Lambda for large momenta. In the same way, the first correction for gg, say δ(2)g\delta^{(2)}g involves the following integral:

the upper index referring to the number of vertices involved in the Feynman diagram. Now, to obtain a well-defined power-counting, the correction for gg has to scale with Λ\Lambda in the same way as gg itself. This is solved by g∼Λ4−dg\sim\Lambda^{4-d}, and we say that the Λ\Lambda-scaling of gg is [g]Λ=4−d[g]_{\Lambda}=4-d. This moreover implies δm2∼Λ2\delta m^{2}\sim\Lambda^{2}, and thus [m2]Λ=2[m^{2}]_{\Lambda}=2. Now, we have to check that it is coherent to all orders of the perturbative expansion. To this end, let us consider a Feynman graph GV\mathcal{G}_{V}, of order VV contributing to the perturbative expansion through the amplitude AGV∼gVΛω(GV)\mathcal{A}_{\mathcal{G}_{V}}\sim g^{V}\Lambda^{\omega(\mathcal{G}_{V})}. Contracting along a spanning tree TV⊂GV\mathcal{T}_{V}\subset\mathcal{G}_{V}, we reduce the original number of propagator edges LL to L−V+1L-V+1, and the resulting graph looks like an effective (local) vertex, having L−V+1L-V+1 loops of length one (tadpoles). Each tadpole behaves like ∫dp/(p2+m2)\int dp/(p^{2}+m^{2}), and thus scales as Λd−2\Lambda^{d-2}. The divergent degree for the contracted graph GV\TV\mathcal{G}_{V}\backslash\mathcal{T}_{V} is therefore:

Because the contraction procedure removes V−1V-1 propagator edges, it increases ω(GV)\omega(\mathcal{G}_{V}) by 2(V−1)2(V-1): ω(GV\TV)=ω(GV)+2(V−1)\omega(\mathcal{G}_{V}\backslash\mathcal{T}_{V})=\omega(\mathcal{G}_{V})+2(V-1). Finally, because the interaction is quartic, we have the relation 2L=4V−N2L=4V-N, NN being the number of external edges. Finally, we get:

Each vertex contributes a factor Λd−4\Lambda^{d-4}, and the scaling g∼Λ4−dg\sim\Lambda^{4-d} ensures that all the quantum corrections have the same scaling. Moreover, setting N=2N=2, we get ω=2\omega=2, in agreement with the one-loop scaling dimension for mass.

Obviously, because this theory is local in the usual sense, the derived scaling dimensions are exactly the same as the one derived from the standard dimensional analysis of the classical action. The two methods however do not coincide for non-local interactions such that (2.38). We argue that this more abstract way to think about locality, scaling and power-counting is more appropriate in a context where the construction of the theory space is not guided by experimental evidences, such that it seems more appropriate to work from the outset within a sufficiently broad framework to accommodate future developments in formalism. However, the exploration of these aspects for NNs is beyond the scope of this paper, since standard locality seems to hold for the Gauss-net kernel .

2.3 N𝑁N-scaling

There exists another scaling dimension, called NN-scaling, associated to the behavior of correlation functions with respect to the width NN of the hidden layer. The Gaussian universality for large NN ensures that the couplings gng_{n} behave as gn∼N−α(n)g_{n}\sim N^{-\alpha(n)} for some positive function α(n)\alpha(n).

The computation can be done by returning to the definition (2.7) and using the fact that W1W_{1} follows a centered Gaussian distribution with variance σW2/N\sigma_{W}^{2}/N. For instance, we find:

which is of order 11. The computation of higher correlation functions can be done using a similar strategy, from the assumptions that the x1(i)x_{1}^{(i)} with different index ii are statistically independent variables. This in particular ensures that:

from this observation, a tedious calculation which is given in shows that the connected 44-point function Gc(4)(x1,x2,x3,x4)G^{(4)}_{c}(x_{1},x_{2},x_{3},x_{4}) has to scale as 1/N1/N, and more generally that α(n)=n/2−1\alpha(n)=n/2-1.

This analytic result also shows the limitations of the approach. Indeed, we expect that a more fundamental method will be able to predict the weights of the interactions. Moreover, the derivation assumes the relation (2.46), and thus the independence of the x1(i)x_{1}^{(i)} having different indices ii, but such an assumption seems to be in conflict with an interaction such that (2.29), which morally must introduce couplings mixing different outputs from the definition (2.7) of fWf_{W}. One may expect that these difficulties could be solved by working with a random vector of size NN, with components φi(x)\varphi_{i}(x) rather than with the function fW(x)f_{W}(x), defining it as an observable fW(x):=⟨φi⟩f_{W}(x):=\langle\varphi_{i}\rangle i.e. the vacuum of the corresponding theory. However, the construction of such a theory is going beyond the scope of this paper, and we plan to investigate it in a forthcoming work.

3 Renormalization group

The RG is probably one of the most important concepts discovered in physics during the last century and forms together with field theory the reference framework of modern physics, from condensed matter to high energies. Pioneered in the works of Wilson and Kadanoff , RG is based on the idea of organizing the theory according to length scales, integrating out short distance degrees of freedom following a recursive procedure called coarse-graining and providing an effective description for the long distance degrees of freedom, through an effective action where microscopic interactions are hidden in effective interactions. Note that RG is in fact a semi-group, which is non-invertible. Thus at each step, information is lost, and RG can be viewed as a systematic procedure to extract large scale relevant features.

To illustrate the physics underlying the Wilson procedure, and before making contact with NN field theory, let us consider a physical system made of a single real scalar field ϕ\phi whose configuration probability follows the exponential form p[ϕ]=e−S[ϕ]p[\phi]=e^{-S[\phi]}, for some classical action S[ϕ]S[\phi]. To have a concrete example in mind, we can take for ϕ\phi the real field described by the classical action (2.40). All the statistical properties of the distribution can be derived from the generating functional (partition function):

This integral being over all configurations for ϕ(x)\phi(x), all the degrees of freedom are integrated out in one step. Equation (2.47) provides a canonical definition of what is microscopic and what is macroscopic, two limits that we conventionally call ultraviolet (UV) and infrared (IR):

In the UV limit, no fluctuations are integrated out. The field configurations are therefore fixed from the extrema of the classical action SS.

In the IR limit, all fluctuations are integrated out. The configurations are fixed by a new action Γ\Gamma, called effective action.

The effective action Γ\Gamma is in turn defined as the Legendre transform of the free energy W[j]:=ln⁡Z[j]\mathcal{W}[j]:=\ln Z[j],

the classical field Ψ\Psi being defined as Ψ(x):=δW/δj(x)\Psi(x):=\delta\mathcal{W}/\delta j(x).

bounded by UV and IR effective physics (Figure 1).

Let us illustrate how that works on the concrete example of the scalar field ϕ\phi described by action (2.40). In that case, μ(E)\mu(E) corresponds to the spectrum of the Laplacian Δ\Delta, whose eigenmodes are Fourier modes, and E≡pE\equiv p. Assuming continuity of the spectrum, we can consider infinitesimal coarse-graining, integrating out slices of infinitesimal thickness. This leads to a differential equation describing how the couplings change as the reference scale changes. Formally, this can be done as follows. We assume the existence of an upper bound for pp, say Λ\Lambda, and we call μΛ(p)\mu_{\Lambda}(p) the spectrum of the free 22-point function with cut-off Λ\Lambda, KΛ(p)K_{\Lambda}(p). As the cut-off Λ\Lambda moves, degrees of freedom are added or removed from the spectrum. Thus, let us consider the bare action “at scale Λ\Lambda”:

where V[ϕ]\mathcal{V}[\phi] includes interactions following our definition of section 2.1. Now let us consider the running cut-off Λ(s)=sΛ\Lambda(s)=s\Lambda, for s∈s\in, which interpolate between the UV scale s=1s=1, and IR scale s=0s=0. If KΛ(s)(E)K_{\Lambda(s)}(E) is at least C(1)\mathcal{C}^{(1)} in ss, we can consider the variation at first order from ss to s′=s+δs^{\prime}=s+\delta:

The following decomposition can be translated as a partial integration from the original partition function using the functional identity:

and χ\chi denotes the degrees of freedom integrated out. Indeed, defining:

we show that the identity (2.52) can be rewritten as:

The classical action at scale Λ(s′)\Lambda(s^{\prime}) formally looks like the action at the scale Λ(s)\Lambda(s). What is different between them is the interaction, which comes at scale Λ(s′)\Lambda(s^{\prime}) from a partial integration over the field χ\chi. The transformation (2.54) can be translated in a differential equation for δ\delta small enough. Indeed, in this limit, the modes χ\chi have a large mass, and can be treated perturbatively. Thus, expanding VΛ(s)[ϕ+χ]\mathcal{V}_{\Lambda(s)}[\phi+\chi] in powers of χ\chi and keeping only terms of order 22, we get Polchinski’s equation :

This equation is formally “exact”. However, this has the reputation to be very hard to solve for many reasons. The first one is that it takes place in a functional space of infinite dimension. If we decide to work in a reduced phase space, taking into account only the most relevant interactions, difficulties appear, instabilities with respect to the considered truncation appear as soon as we try to get beyond the perturbative sector, which is precisely what we are aiming at in this paper. For this reason, and as it is the case for the largest part of non-perturbative investigations in the literature , we will prefer to use the Wetterich formalism, better for dealing with non-perturbative approximations. We will discuss this method in the next section.

3.2 Renormalization group(s) for the NN-QFT

The analogy between NNs and RG is evident: both are aiming at extracting relevant features from a massive number of degrees of freedom. RG shows that microscopic details can be ignored to describe long distance physics, and that microscopic theories can be indistinguishable from their common large distance properties. Extracting regularities from large sets of data is exactly what machine learning does; and, as we recalled in the introduction, the question of the relevance of the renormalization group in artificial intelligence is growing in the literature . However, the effective field theory that we presented in the first part offers a new framework to discuss aspects related to the renormalization group in the study of the behavior of NNs . As the previous section stressed out, the field theory that we consider exhibits strong similarities with theories usually considered by physicists: the long distance (i.e. large volume, small momenta) limit (2.22) of the free propagator being the same as for the usual scalar field ϕ\phi described by the action (2.40). This formal similarity will serve as a guide in the construction of the renormalization group, and it is very tempting to carry out a coarse-graining in momenta, exactly as for the scalar field ϕ\phi in the previous section. We will discuss two different coarse-graining strategies which we call respectively passive and active RGs. But before we go into them in detail, let’s make a few general remarks about what distinguishes this NN field theory from ordinary theories.

In the standard scenario, what is known is the UV theory i.e. the classical action. This action itself is viewed as an effective description, valid at some fundamental scale and ignoring the details about nature and physics of microscopic degrees of freedom underlying the physical world. The choice of the classical action is constrained by predictivity (which promotes just-renormalizable theories), consistency with quantum effects (compensation of anomalies in gauge theories, for instance), and the effective structures at the scale at which the theory is defined, which generally implies some symmetries (rotation, reflection, gauge invariance…). In this respect, the RG aims to provide an approximation of the exact quantum theory, and to compare it with experiments. The case of the field theory that we consider differs from this general picture in its relations between UV and IR scales. The propagator (2.12) is exact and defined in the deep IR. From a RG point of view, the knowledge of this propagator takes into account all the fluctuations at all scales. But, for finite NN, the knowledge of the 22-point functions is not sufficient to reproduce higher correlations functions, and non-Gaussian interactions are required in the classical action to reproduce experimental correlation functions. But, due to these interactions, the flow of the different ingredients entering in the definition of the classical action becomes non-trivial, with the consequence that both SkinS_{\text{kin}} and SintS_{\text{int}} in the UV are unknown. Thus, in some sense, the situation is the inverse of what we do in ordinary field theory: we have to infer the form of the UV theory (or more likely a class of UV theory) from the knowledge of only a part of the IR theory. By construction, such an inference cannot lead to a single solution, but a class of solutions which have to satisfy the following requirements:

reproduce the exact 22-point functions up to the experimental precision;

reproduce the deviations from Wick’s theorem, due to interactions, and which are less and less perturbative as NN becomes small, once again up to irrelevant corrections with respect to the experimental precision.

Any measurement in physics comes with a finite precision: hence, two effective descriptions are considered to be equivalent and sufficient to describe something if the predictions agree up to the experiment precision. The precision is also finite in numerical simulations, and this explains why that we are able to infer only an equivalent class of models rather than a point in the theory space. In the first section, we showed that the relative relevance of the interactions is not the same such that irrelevant interactions contribute below the machine precision threshold, meaning that we have no way to distinguish between several initial conditions whose trajectories are sufficiently close in the IR (see Figure 2). This argument allows working, in a first approximation, within a finite subspace of the full theory space, focusing on interactions having the largest canonical dimension.

Because of the existence of an intrinsic length scale ξ\xi defined in (2.25), we can think to partially integrate microscopic degrees of freedom with respect to this length scale to construct a proper RG flow following standard field theory. In this picture, what is playing the role of a microscopic scale is the working precision (see Section 2.1.3), which introduces a cut-off in momentum integration Λ=1/a0\Lambda=1/a_{0}. We can then construct a coarse graining procedure from grid size dilatation (see Figure 3).

Note that from such a procedure, one needs to have ξ≫a0\xi\gg a_{0}. Because the maximal value for pp is p∞=2π/a0p_{\infty}=2\pi/a_{0}, we can show that it implies p∞ξ≫1p_{\infty}\xi\gg 1, which invalidates the expansion (2.21). However, it may happen that such an expansion holds in a sufficiently large domain. A necessary condition is that the expansion (2.21) holds for the smallest (nonzero) momentum p0=2π/a0N0p_{0}=2\pi/a_{0}N_{0}, implying:

which is the condition defining the large volume limit.

A dilatation procedure as described in Figure 3 induces a renormalization group by partial integration of momenta into the windows ∼]1/a0′,1/a0]\sim]1/a_{0}^{\prime},1/a_{0}]. The existence of two complementary limits ξ≪L\xi\ll L and ξ≫a0\xi\gg a_{0} is reminiscent of a crossover scale behavior, between a deep ultraviolet (UV) limit p∼1/a0p\sim 1/a_{0} and a deep infrared limit (p∼1/Lp\sim 1/L), which we will study separately in the next section. Note that such a crossover scale appears generally in situations where two very different mass scales appear, ensuring decouplingSuch a decoupling is at the origin of the large mass expansion in field theory. of some effects associated to the larger one when experiments focus on the first one . Here, what plays the role of a large mass is the inverse of the typical observation scale ξ\xi; in the very large mass limit, p∞≪(ξ)−1p_{\infty}\ll(\xi)^{-1}, and the IR sector recovers all the physics. In the opposite limit, p0≫(ξ)−1p_{0}\gg(\xi)^{-1}, everything is UV, and an expansion such that (2.21) does not hold. In other words, for p≪(ξ)−1p\ll(\xi)^{-1}, one expects that quantum effects are suppressed with powers of (ξ)−1(\xi)^{-1}. This observation can be a source of improvement for approximations used to solve the RG flow equation (3.3) in the next section. In particular, we understand that contributions coming from higher couplings will tend to stay small if they are at the transition scale (ξ)−1(\xi)^{-1}. Section 3 is devoted to this RG strategy.

In the process described above, the observation scale ξ\xi is kept fixed and the working precision is changed. Conversely, we can keep the data (i.e. the working precision) fixed and change the observation scale. If the first version is essentially passive with respect to the NNs (i.e. the latter is not changed), this strategy is, in contrast, active (see Figure 4). Indeed, remembering the expression (2.25), ξ\xi is completely determined in terms of σW\sigma_{W}, the standard deviation of the weight distribution. Hence, flowing in the observation scale is equivalent to changing the weight standard deviation, and thus the neural network.

Physically, if we think of a thermodynamic system like a ferromagnet, such a strategy is equivalent to turning the thermostat’s knob to lower the temperature towards the critical regime. This alternative point of view is the subject of the section 4.

The active RG is closer to the RG version considered in the reference than the passive scheme, as the flow equations derived in section 4 show explicitly. However, despite this formal contact, our approach differs by its very construction. While, from their point of view, the RG is the mathematical explanation of a principle of invariance with respect to a certain volume “cut-off”, our RG is the result of a procedure of partial integration of the degrees of freedom of the field.

Indeed, the RG flow is usually performed with respect to a UV cut-off (spacetime /data-space resolution) and not an IR cut-off (volume). In , the large volume cut-off was introduced because the 22-point function diverges at large distance (at least, for the ReLU-net), which reminds the short-distance divergence of the canonical propagator in particle QFT. Moreover, one can ask whether the data-space should be identified with the position or momentum space in usual spacetime QFT, and in principle this could depend on the problem. From our arguments in Section 2.1.3, it seems more natural to identify the data-space with the position space (except in the case where the data is already the Fourier transform of a space(time) process). IR divergences are also present in particle QFT, and they are cured following different methods according to their origin. The first case are massless particles, for which a refined definition of amplitudes is needed . An infrared cut-off (such as a mass) can be introduced at intermediate stages to regulate the integral, but it is not a renormalization parameter. In practice, the divergences of the ReLU-net arise from a similar origin (singularity in the propagator for large distance / zero-momentum). Second, IR divergences appear for internal on-shell propagators: they translate the fact that quantum effects shift the vacuum and masses of the fields. Resummation of quantum effects through renormalization leads to finite results . Third, some quantities can diverge for infinite volume for example when studying phase transition: in that case, the usual method is to study the theory with different values of a volume cut-off and to extrapolate to infinite volume (thermodynamic limit) . However, this is not a renormalization flow. For these reasons, we take a more conservative approach and identify small resolution in data space with the UV limit, and perform the RG flow for the associated cut-off.

It is also noted that in [1, sec. 4.3] that the Gauss-net does not require renormalization because the 22-point function is exponentially decaying with the distance, such that all integrals are convergent. In fact, the previous paragraph shows that renormalization is still needed in this case because its role is not only to handle properly (spurious) UV divergences, but also to take into account quantum effects (some of which lead to IR divergences). Said another way, renormalization provides a mapping between the bare and physical parameters (at a given energy scale) there is always a renormalization flow in the space of couplings. Indeed, the bare parameters describe the properties of the fields without interactions: they are not physical because fields do not live in isolation and any measurement implies an interaction. A famous example of a perfectly finite theory but which has an infinite number of finite (such that predictivity is not lost) counter-terms and non-trivial RG flow (with the so-called stub length) is string field theory .

Flowing through NN-QFT theory space: the passive RG

In this section, we show how the passive RG within Wetterich formalism allows predicting the behavior of correlation functions for a fully connected NN with a single hidden layer. We start with a short presentation of the Wetterich formalism, before turning on to applications. We will consider separately two different regimes, the deep IR regime k≪(ξ)−1k\ll(\xi)^{-1} where the effective propagator can be suitably approximated with an ordinary Laplacian ∼(−Δ+m2)−1\sim(-\Delta+m^{2})^{-1}, and the UV regime k∼(ξ)−1k\sim(\xi)^{-1}, where the propagator follows the exponential law ∼e−Δ/m2/m2\sim e^{-\Delta/m^{2}}/m^{2}.

In section 2.3.1, we provided a formal introduction to Wilson’s ideas for the RG. In this section, we present another incarnation, the so-called Wetterich formalism , which focuses on the effective action for integrated degrees of freedom rather than on the effective classical action for the remaining degrees of freedom, as it is the case in (2.57). We focus on the passive RG as presented in the section 2.3.2. Let Λ=1/a0\Lambda=1/a_{0} be some reference working precision and k∈[0,Λ]k\in[0,\Lambda]. Assuming that we performed partial integration up to the scale kk, we denote as Γk\Gamma_{k} the effective action for those averaged degrees of freedom. Obviously, it must satisfy the boundary conditions:

Γk=Λ=S\Gamma_{k=\Lambda}=S, no fluctuations are integrated out, and the effective action reduces to the classical action.

Γk=0=Γ\Gamma_{k=0}=\Gamma, all fluctuations are integrated out and we recover the full effective action Γ\Gamma defined in (2.48).

The Wetterich formalism aims to construct a smooth interpolation between these two limits. To this end, it is convenient to modify the classical action with a scale dependent mass term ΔSk\Delta S_{k}, which reads in momentum space:

The substitution S→S+ΔSkS\to S+\Delta S_{k} defines a kk-dependent partition function ZkZ_{k} through the definition (2.47). The shape of the scale-dependent mass rk(p2)r_{k}(p^{2}) is designed to freeze low momenta modes p2<k2p^{2}<k^{2}, decoupling them from long distance physics whereas high energy modes p2>k2p^{2}>k^{2} remain essentially unaffected. Moreover, in order to recover the full classical action Γ\Gamma for k=0k=0, rk(p2)r_{k}(p^{2}) has to vanish in that limit. In the same way, it has to become very large in the opposite limit, for k→Λk\to\Lambda, in order to satisfy the UV boundary condition Γk→Λ→S\Gamma_{k\to\Lambda}\to S (all the fluctuations are frozen). The interpolating functional Γk\Gamma_{k} is defined as:

As kk varies from kk to k−δkk-\delta k, effective couplings involved in the effective action change. To obtain the differential equation governing the behavior of Γk\Gamma_{k}, as kk varies, we can differentiate the definition (3.2) with respect to kk. After a tedious calculation whose details can be found in , we get the following functional equation:

where Γk(n)\Gamma^{(n)}_{k} denotes the nn-th functional derivative with respect to the classical field Ψ(x):=δln⁡Zk/δj(x)\Psi(x):=\delta\ln Z_{k}/\delta j(x). This equation, up to the formal character of its derivation, is as exact as equation (2.57) is. It defines a trajectory through a functional space and is as hard to solve as the equation (2.57). Approximations are required to make the underlying physics tractable. The standard strategy, called truncation, is to identify a relevant finite-dimensional subspace of the full theory space, and to project the flow equation (3.3) onto it. Working with equation (3.3) has the great advantage that this projection procedure does not require to assume that couplings are small, and thus allows investigating approximate but non-perturbative solutions of the RG flow.

2 Local potential approximation in the deep IR

The local potential approximation (LPA) is one of the most popular approximation procedures to solve the exact RG flow equation (3.3). This approximation focuses on the region of the full theory space spanned by local interactions in the sense of (2.29). For the investigations in this section, we assume p2≪2σW2/dinp^{2}\ll 2\sigma_{W}^{2}/d_{\text{in}}, which is our reference mass scale. This implies:

which defines the IR regime (see Section 2.3.2). Note that due to the scaling behavior of derivative contributions, one expects that the validity of this description survives in the weak UV regime: 1/a0≫k≫(ξ)−11/a_{0}\gg k\gg(\xi)^{-1} due to the large river effect which states, that in a suitable vicinity of the Gaussian fixed point and in the absence of singularities along the flow, the latter projects itself into the subspace spanned by the most relevant couplings .

To begin, we focus on the simplest truncation around sixtic interactions, discarding from our analysis contributions arising from higher couplings. This is equivalent to setting:

where, to avoid confusion with the example given in Section 2.3, we denote as Ψ\Psi the classical field. Note that such an expansion around Ψ=0\Psi=0 is named symmetric phase expansion, and we call symmetric phase the domain of the full phase space where it remains valid. It may happen that such an expansion breaks down, in cases where Ψ=0\Psi=0 becomes an unstable vacuum. This is the case when phase transitions are encountered. In this section, we focus on the symmetric phase, and discuss more elaborate formalisms in the next section. Approximation (3.5) ensures that we keep effects up to order 1/N21/N^{2}.

To be more concrete, we assume that Γk[Ψ]\Gamma_{k}[\Psi] can be decomposed as a sum of two contributions:

The kinetic contribution Γk,kin[Ψ]\Gamma_{k,\text{kin}}[\Psi] keeps all the quadratic terms in Γk[Ψ]\Gamma_{k}[\Psi].

The effective potential Uk[Ψ]U_{k}[\Psi] gathers the non-Gaussian contributions in the expansion of Γk[Ψ]\Gamma_{k}[\Psi].

Without lost of generality, the kinetic contribution can be written as:

The kernel Kk(p2)\mathcal{K}_{k}(p^{2}) is a priori difficult to track. Fortunately, because we are aiming to deal with IR effects, the momentum pp is expected to be small, justifying to expand K(p)\mathcal{K}(p) in power of p2p^{2}:

The first term of this expansion define the running mass, and we denote it as m2(k)m^{2}(k). In the same way the second term of the expansion is called running wave function renormalization, and we denote it as Z(k)Z(k). Nerveless, it is easy to check that in the symmetric phase Z(k)Z(k) does not depend on the running scale kk (see below). Thus we must have Z(k)=Z0=1Z(k)=Z_{0}=1. This scheme defines the derivative expansion , and for this section we focus on the two first terms:

To keep only effects up to order 1/N21/N^{2}, we consider the following truncation for the effective potential:

This form follows from the expression of a local interaction of order nn in position space: the Fourier transformation to momentum space introduces one momentum pjp_{j} for each field together with a delta function for momentum conservation, and a sum (since the pjp_{j} take discrete values) over each value of the momentum. The final piece is the regulator rkr_{k}. From the choice of the kinetic truncation (3.9), it is suitable to use the modified version of the standard optimized Litim’s regulator :

for which analytic computations are possible. In this equation, θ\theta is the step function such that θ(x>0)=1\theta(x>0)=1 and θ(x<0)=0\theta(x<0)=0. The flow equations can be deduced from the exact RG equation (3.3) by taking successive derivatives with respect to the classical field Ψ\Psi. Taking the second derivative gives the flow equation for Γk(2)(p⃗1,p⃗2)\Gamma_{k}^{(2)}(\vec{p}_{1},\vec{p}_{2}).

where, on the RHS, functions are computed for Ψ=0\Psi=0. From the truncation (3.9), we must have:

The fourth derivative Γk(4)(p1,p2,p3,p4)\Gamma_{k}^{(4)}(p_{1},p_{2},p_{3},p_{4}) can be easily computed from the truncation (3.10), leading to:

replacing Ψ=0\Psi=0 at the end of the computation. Thus, setting p1=0p_{1}=0 on both sides of equation (3.13), we get after some calculationsDetails are given on Appendix B.:

In equation (3.13), the only dependence on the external momenta p1p_{1} and p2p_{2} on left-hand side is through the conservation delta δp1,−p2\delta_{p_{1},-p_{2}} arising from the structure of the four-point function vertex Γk(4)\Gamma_{k}^{(4)}. Thus, the field strength Z(k)Z(k), whose flow equation could be deduced by taking derivatives on both sides of equation (3.13) with respect to p12p_{1}^{2}, vanish identically.

In the same way, taking the fourth and sixth derivatives with respect to MM of the flow equation (3.3), and from the condition (3.15), we get schematically:

One expects that such an approximation remains valid for k2≫4π2/L2k^{2}\gg 4\pi^{2}/L^{2}, with LL being large (see Figure 5). Thus, defining,

we get the autonomous system (β2n:=kduˉ2n/dk\beta_{2n}:=kd\bar{u}_{2n}/dk):

From these equations, it is obvious that the behavior of the flow depends on the dimension dind_{\text{in}}. For instance, for din>4d_{\text{in}}>4, all the couplings are irrelevant and trajectories return toward the Gaussian region, the uˉ2\bar{u}_{2} axis being the only direction of instability. In contrast, for din<4d_{\text{in}}<4, some couplings become relevant, and trajectories are repelled from the Gaussian region. u4u_{4} is the first one to become relevant, for din>3d_{\text{in}}>3; for din<3d_{\text{in}}<3, u6u_{6} becomes relevant as well. Figure 6 illustrates the behavior of the RG flow for several dimensions. We have integrated numerically the flow equations for u2=1u_{2}=1, u4=−0.5u_{4}=-0.5, u6=0.01u_{6}=0.01 in Figure 7.

2.2 Beyond the symmetric phase

In this section, we consider another approximation scheme for the effective potential UkU_{k}. Focusing on the IR regime, we assume that Ψ(p)\Psi(p) essentially reduces to its zero component (the macroscopic field):

and, defining χ:=Ψ02/2\chi:=\Psi_{0}^{2}/2, we expand the effective potential per unit volume in power series around χ=κ(k)\chi=\kappa(k):

Within this parametrization, we identify directly κ\kappa with the (non-zero) vacuum, which runs with the scale kk. The two-point function Γk(2)\Gamma^{(2)}_{k} is moreover defined as:

Note that we introduced the field strength renormalization Z(k)Z(k) because its own flow is nonzero as soon as κ≠0\kappa\neq 0, i.e. broken phase effect introduces an anomalous dimension. As a technical device, we move the mass contribution in the effective potential. For a uniform field configuration, we must have:

Therefore, taking the derivative with respect to t:=ln⁡(k/Λ)t:=\ln(k/\Lambda) (X˙:=kdX/dk\dot{X}:=kdX/dk) and writing Uk′(χ)=∂Uk/∂Ψ0\mathcal{U}_{k}^{\prime}(\chi)=\partial\mathcal{U}_{k}/\partial\Psi_{0}, we get from (3.3):

As in the previous section, we use the Litim regulatorWhich is optimal in the following sense: the functional renormalization group (FRG) equations are defined only if the effective propagator GG is well-defined; that is, if G−1G^{-1} has no zero modes (and therefore do not develop IR divergences). This can be achieved by demanding that G−1G^{-1} has a sufficiently large gap, i.e. a sufficiently large minimum. See for more details., but modify it to deal with the running field strength Z(k)Z(k):

For kk large enough, we may use the same integral approximation as for (3.22), but taking into account that KdinK_{d_{\text{in}}} must now depend on the anomalous dimension ηk\eta_{k} because of the factor Z(k)Z(k) in (3.33):

As in the previous section, we introduce dimensionless quantities (labeled with overlines) as:

Note that all these changes of variable make sense from the requirement that all the terms in the potential must have the same dimension (the dimensions of gg and hh have been fixed). The derivative on the RHS in equation (3.36) is taken at χ\chi fixed. Therefore, we have:

where in the RHS the derivative is taken with χˉ\bar{\chi} fixed. We obtain:

The flow equations can be deduced from the normalization conditions at scale kk:

Hence, because Uˉ˙k[χˉ=κˉ]=−gˉ κˉ˙\dot{\bar{\mathcal{U}}}_{k}[\bar{\chi}=\bar{\kappa}]=-\bar{g}\,\dot{\bar{\kappa}}, we obtain for κˉ\bar{\kappa},

and after a tedious calculation, we obtain for u4u_{4} and u6u_{6}:

The computation of the anomalous dimension is long and provided in Appendix B. The result is:

For the Litim’s regulator (3.33), the anomalous dimension ηk\eta_{k} in the local potential approximation with kinetic truncation up to order p2p^{2} is given by:

In this section, we are aiming to discuss the deep UV regime 1/a0>k≳(ξ)−11/a_{0}>k\gtrsim(\xi)^{-1}, where the expansion (2.21) is not valid. For this regime, the derivative expansion breaks down as well, and a local approximation for interactions is no longer justified. We present a method, inspired from the Blaizot-Mendez-Wschebor (BMW) formalism , which considerably improves the accuracy of truncations in regimes where the momentum-dependence of vertex functions is as relevant as the purely local ones, i.e. relevant enough to invalidate the derivative expansion. In this section, we summarize the essential results through three compact statements, in order to focus on the results, and leave the technical details in Appendix B.

The procedure that we propose is based on the following three approximations (see also ):

We parametrize the 22-point function Γk(2)(p,p′)\Gamma_{k}^{(2)}(p,p^{\prime}) with a single parameter, the running mass m2(k)m^{2}(k), such that:

such that Γk(2)(p,p′)\Gamma_{k}^{(2)}(p,p^{\prime}) reduces to the exact 22-point function in the deep IR for m2(k=0)=σW2/2dinm^{2}(k=0)=\sigma_{W}^{2}/2d_{\text{in}}.

We assume that vertices are slowly varying with respect to the momenta qq running through the effective loops in the flow equation. The allowed windows of momenta being such that q2≲k2q^{2}\lesssim k^{2}, for kk small enough with respect to other momenta, we require:

The third approximation is about the propagator entering in the flow equation. For qq in the windows of momenta allowed by ∂krk(q2)\partial_{k}r_{k}(q^{2}), we must have:

where θ\theta is the Heaviside step function and α\alpha a positive number, expected to be of order 11.

To complete these approximations we need to choose a suitable regulator. In principle, we could always use the Litim regulator (or any other regulator used in the literature). However, due to the parameterization of the phase space that we have chosen, and in particular the expression of the 22-point function, this regulator loses its crucial advantage which consists in freezing all the fluctuations below the scale kk. The Litim optimal condition is moreover expected to be a relevant constraint to define a regulator, especially in the symmetric phase, and we have the following statement:

satisfies all the requirements for a regulator as soon as m2(k)>0m^{2}(k)>0, freezes out all fluctuations with momentum q2<k2q^{2}<k^{2}, and is optimal in Litim’s sense.

The physical discussion motivating this choice being a little technical, we provide it in Appendix B. Note that Litim’s condition, which relies on the existence of an optimized gap for the inverse 22-point function Γk(2)+rk(p2)\Gamma^{(2)}_{k}+r_{k}(p^{2}), is not an absolute criterion in regard to the reliability of the results. Indeed, some choices are expected to provide an optimal bound for the gap, which may have an influence on the computation of physical quantities like critical exponents . Working within the set of “optimized regulators” in Litim’s sense, we may complete the optimization argument with a principle of minimal sensitivity , requiring that physical quantity has to be stationary with respect to some parameters spanning a family of regulators. This can be done for instance by replacing rk→βrkr_{k}\to\beta r_{k}, and to vary the physical quantities with respect to β\beta. This will be the only optimization scheme that we will discuss in this paper.

Within these approximations, the equation for the 22-point function reads:

which, from the observation that Γk(n+1)(p1,⋯ ,pn,0)≡∂Γk(n)(p1,⋯ ,pn)/∂Ψ0\Gamma_{k}^{(n+1)}(p_{1},\cdots,p_{n},0)\equiv\partial\Gamma_{k}^{(n)}(p_{1},\cdots,p_{n})/\partial\Psi_{0}, leads to a closed equation for Γk(2)\Gamma_{k}^{(2)}:

This is the standard BMW strategy. Our approach will however be a little different. First, we work in the symmetric phase Ψ0=0\Psi_{0}=0. Second, we exploit the fact that the 22-point function in our parametrization depends only on a single parameter (the mass) to close the hierarchy around the 66-point function, thus removing the need for the usual assumption of a proportionality relation between the 66 and 44 points contributions in the flow equation of Γk(4)\Gamma_{k}{(4)} (see ). On the contrary, we will be able to deduce an expression for the 66-point function from the knowledge of the 44-point function, itself deduced from the flow equation of the 22-point function. The only relevant parameters at sufficiently large times being the local parameters, u2nu_{2n}, whose flow equations are deduced from the derivative expansion. The derivation of these equations being technical, we provide it in Appendix B, summarizing them in the following statement:

Truncating around sixtic interactions (i.e. up to O(1/N2)\mathcal{O}(1/N^{2}) effects) in the deep UV, neglecting the momentum dependence of effective vertices in the computation of effective loops and for external momenta large enough, the flow equations for the local couplings uˉ2\bar{u}_{2}, uˉ4\bar{u}_{4} and uˉ6\bar{u}_{6} are:

and Fn(uˉ2):=din∫01rdin+2n−1er2−1uˉ2drF_{n}(\bar{u}_{2}):=d_{\text{in}}\int_{0}^{1}r^{d_{\text{in}}+2n-1}e^{\frac{r^{2}-1}{\bar{u}_{2}}}dr. Furthermore, defining the minimal dimensionless vertex functions as (x:=p/kx:=p/k):

It is moreover interesting to note that for external momenta large enough, the knowledge of fˉk(x,0)\bar{f}_{k}(x,0) allows reconstructing the 44-point function. Once again, we put the proof in Appendix B, and summarize the result in a compact statement:

For momenta large enough (pi2∼(ξ)−2(k)p_{i}^{2}\sim(\xi)^{-2}(k)), the 44-point vertex Γk(4)(p1,p2,p3,p4)\Gamma_{k}^{(4)}(p_{1},p_{2},p_{3},p_{4}) can be suitably approximated as:

Flowing though the neural network space: the active RG

In this section, following the discussion of section 2.3.2, we consider the active RG, viewing the networks parameters 2σW2/din2\sigma_{W}^{2}/d_{\text{in}} as an UV cut-off rather than a running mass. First, we derive the corresponding flow equation. We show that nn-point functions exhibit a purely scaling behavior, and that the corresponding β\beta-functions reduce to the linear dimensional contributions. As discussed in section 2.3.2, this RG is formally the same as the one used in , up to the lack of explicit scaling for mass, explaining why our scaling dimensions are different. Second, we investigate the content of the flow equations that we obtained. As pointed out in the section 2.3.2, the major advantage of this approach is to avoid introducing a working precision, or any special structure regarding the nature of the data.

Let us consider a network (σW,σb)(\sigma_{W},\sigma_{b}), and define Λ2:=2σW2/din\Lambda^{2}:=2\sigma_{W}^{2}/d_{\text{in}}. Within this suggestive notation, the exact propagator looks like a UV regularized free propagator:

Λ\Lambda playing the role of a UV cut-off, which suppresses large momenta. The question is therefore: what happen if we smoothly change the parameter σW\sigma_{W}?In fact, this situation is familiar in string field theory, where Λ\Lambda is called the stub parameter . Formally, this is equivalent to moving the UV cut-off, which can be translated as a chain of equivalence relations between classical actions through the differential equation (2.57), all of them having the same long distance physics. One can think for the evolution equation of the classical action to something like equation (2.57), i.e.

Such an equation however assumes that KΛK_{\Lambda} is the free propagator. Rather, in our construction, it has to be understood as the effective propagator, taking into account fluctuations. Therefore, we have to construct a coarse graining with a fixed shape of the effective propagator, the corresponding free propagator remaining unknown. There is a pragmatic way to do this. We formally introduce a regulator ΔSk\Delta S_{k} in the classical action. This leads to the Wetterich equation (3.3), but with the additional constraint that:

This equation simply means that we relate the running scale kk to the standard deviation of the neural network weights as

and that we keep the shape of the 22-point function fixed along the RG trajectory (if it exists), fixing Γk(2)(p,−p)\Gamma_{k}^{(2)}(p,-p) as soon as rk(p2)r_{k}(p^{2}) is given. Let us show how this condition allows closing the hierarchical flow equations. Let us consider the flow equation (4.5) for Γk(2)(p1,p2)\Gamma_{k}^{(2)}(p_{1},p_{2}). Neglecting the momentum dependence of the effective vertex Γk(4)(p1,p2,q,−q)\Gamma_{k}^{(4)}(p_{1},p_{2},q,-q) with respect to the momenta "qq" running through the effective loop following the discussion of Section 3.3, we get:

Note that this approach implicitly assumes kk is small enough to justify the replacement: Γk(4)(p,−p,q,−q)→Γk(4)(p,−p,0,0)\Gamma_{k}^{(4)}(p,-p,q,-q)\to\Gamma_{k}^{(4)}(p,-p,0,0) in (4.5). Hence, the resulting flow equations are expected to be exact for reference scales in the IR.

Because the LHS can be explicitly computed from (4.3), we therefore obtain:

In the same way, the flow equation for Γk(4)(p,−p,0,0)\Gamma_{k}^{(4)}(p,-p,0,0) allows in principle to compute Γk(6)(p,−p,0,0,0,0)\Gamma_{k}^{(6)}(p,-p,0,0,0,0) within the same approximation. Let us illustrate how this works. Let us consider a given network defining the “fundamental scale” k≡Λ0k\equiv\Lambda_{0}. We can measure the 44-point function at zero momentum Γk=Λ0(4)(0,0,0,0)≡u4(Λ0)δ(0)\Gamma_{k=\Lambda_{0}}^{(4)}(0,0,0,0)\equiv u_{4}(\Lambda_{0})\delta(0). This condition in turn fixes the value of r˙k(0)\dot{r}_{k}(0). For instance, let us consider the following explicit example, working with the slightly modified Litim’s regulator:

Straightforwardly, we have rk(0)=αk2r_{k}(0)=\alpha k^{2} and r˙k(0)=2αk2\dot{r}_{k}(0)=2\alpha k^{2}, and the previous equality reads as:

where we introduced the dimensionless variable x:=q/kx:=q/k, and

Introducing the dimensionless coupling uˉ4:=Λ0din−4u4\bar{u}_{4}:=\Lambda_{0}^{d_{\text{in}}-4}u_{4}, and solving on α\alpha, we thus obtain:

Because u4=O(1/N)u_{4}=\mathcal{O}(1/N), and I2I_{2} is a pure number, α\alpha is close to 11 for large NN. However, α\alpha increases as NN decreases, and for uˉ4∼2/I2\bar{u}_{4}\sim 2/I_{2} the approximation breaks down. Note that we could expect this not to be a limitation of the approach in itself, but a limitation of the Litim regulator, however, a moment of reflection shows that such a singular behavior is in fact very general, and independent on the choice of the regulator. Under the condition (4.10), the problem (4.6) is well posed but trivial: it reduces to a pure scaling behavior. Indeed, given (4.6), we have:

The flow is entirely fixed by dimensional analysis, and the flow equation for u4u_{4} reduces to its linear contribution:

In turn, this equation determines Γk(6)∼(Γk(4))2∼O(1/N2)\Gamma_{k}^{(6)}\sim(\Gamma_{k}^{(4)})^{2}\sim\mathcal{O}(1/N^{2}). The effective loop behaves like k6−2dink^{6-2d_{\text{in}}}, times a factor which is kk independent. Hence, we deduce:

meaning that u6u_{6} follows a purely scaling behavior as well.

We can use (4.4) to write the flow equations in terms of the standard deviation σW\sigma_{W}:

where now u4u_{4} and u6u_{6} are seen as functions of σW\sigma_{W}. As displayed in Figures 8 and 9, the numerical simulations match to a good precision with the solution to this equation (see Appendix A for the computations of u4u_{4}).

Finally, let us make a remark in regard to the results obtained in the reference . Indeed, the authors arrived to the equations (4.14) with σW\sigma_{W} replaced by an IR cut-off from perturbation theory, whose validity assumes NN to be large enough. What is puzzling with this calculation is that the RG predictions work even for small NN, where we expect that perturbation theory breaks down. Our derivation solves this paradox: using a non-perturbative framework, we are able to show that coupling constants follow scaling laws (4.11), without assumption on the sizes of the coupling constants (however, note that both derivations have been performed with different activation functions, such that it would be interesting to check how general (4.14) is).

Conclusion and outlooks

In this paper, we have pushed further the use of the renormalization group for the NN-QFT correspondence , which states that a neural network can be represented by a quantum field theory. In the infinite limit of the hidden layer width NN, the neural network is described by a Gaussian process and mapped to a free field theory, and interactions translate finite-NN corrections. The main difference with usual QFTs used in physics stems from the choice of the kernel (or propagator), itself inherited from the choice of activation function in the neural network. Since it encodes important properties on the theory (IR and UV divergences, scaling, etc.), it is important to ask how the data-space of the neural network inputs differ from usual spacetime. As a consequence, the usual assumptions on interaction locality may not be appropriate. In the first part of this paper, we have discussed several of these aspects, providing an interpretation slightly different from the one in .However, let us stress that this divergence in interpretation does not change in any way the computations and numerical results in both sides.

We have then described how to build a non-perturbative renormalization group flow following Wetterich-Morris formalism. We introduce two different points of view: in the passive case, the UV cut-off is related to the data resolution, while in the active case, it is given in terms of the standard deviation of the NN weights. The main difference with is that they postulate a global scale invariance with respect to a large volume cut-off (IR, in the language of our paper) on the data-space. Intriguingly, their results agree strongly with numerical simulations, for the small widths, whereas perturbation theory is expected to fail, and even if a scale invariance with respect to the volume is not expected. In this paper, we solve this paradox by developing a renormalization group based on an explicit coarse-graining, and derive flow equations from a process of partial integration of the field degrees of freedom. We find that the active point of view is formally identified with the flow in , thus justifying, in an explicitly non-perturbative framework, the agreement between theory and experience found in the paper. A natural extension of this work is to include non-local interactions using tensor models. Another possible direction is to generalize the derivation to other networks such as ReLU-net .

On the numerical side, the main result of our paper is the flow equation (4.14) which shows that the weight standard deviation σW\sigma_{W} can be interpreted as a running cut-off in terms of which the couplings of the NN-QFT change. This means that given the couplings for a specific value of σW\sigma_{W}, it is possible to compute analytically the couplings for any other value of σW\sigma_{W} without doing any numerical simulation. We have verified this statement using numerical simulations (Figures 8 and 9). In this paper, we have focused the analysis on the analytical computations in the QFT side: we plan to analyze the equations numerically in future works.

From a function-space perspective, it is natural to understand the learning process as a RG flow induced in a suitable theory space. It would very interesting to investigate how the notions presented in this paper could generalize to describe this process and how the couplings change under learning.

Acknowledgments

We are very grateful for useful discussions with James Halverson, Anindita Maiti, and Keegan Stoner. We would like to thank particularly Anindita Maiti for help in reproducing the numerical results from . This project has received funding from the European Union’s Horizon 2020 research and innovation program under the Marie Skłodowska-Curie grant agreement No 891169. This work is also supported by the National Science Foundation under Cooperative Agreement PHY-2019786 (The NSF AI Institute for Artificial Intelligence and Fundamental Interactions, http://iaifi.org/).

Appendix A Numerical simulations

In this appendix, we explain how to compute numerically correlation functions for neural networks defined in Section 2.1.1 and how to extract relevant information. In particular, we reproduce the numerical results from and provide additional details. The code is written in Python and is available at https://github.com/melsophos/nnqft. Throughout this appendix, we take:

In order to evaluate correlation functions, we consider nnetsn_{\text{nets}} neural networks fαf_{\alpha}. For each of them, the weights and biases are drawn independently from the distributions N(0,σW2/N)\mathcal{N}(0,\sigma_{W}^{2}/N) and N(0,σb)\mathcal{N}(0,\sigma_{b}), where NN is the width of the hidden layer. We will take :

Then, the experimental nn-point correlation functions are computed as (2.3):

We define the difference with the large NN Green functions G0(n)G^{(n)}_{0} as (see Section 2.1.4):

and the normalized nn-point functions as:

Note that no absolute value has been taken until now, and the result can be positive or negative. The large NN Green functions are computed with Wick theorem from the Gauss-net kernel (2.11). For example, the 44-point function is given by:

In order to reduce variance of the results, we will compute the Green functions by averaging over nbagsn_{\text{bags}}, each made of nnetsn_{\text{nets}} networks:

where Gexp(n)(x1,…,xn)∣AG^{(n)}_{\text{exp}}(x_{1},\ldots,x_{n})|_{A} means that the correlation functions is computed with the bag AA. This also allows extracting standard deviations if needed.

We will compute the correlation functions for the following points :

Given a nn-point correlation function, we compute it for all the possible combinations of nn points x(i)x^{(i)} with i=1,…,6i=1,\ldots,6, including identical entries. Since the experimental Green functions are symmetric by construction (as are the QFT Green functions), we consider only combinations which are inequivalent up to permutations. For example, we will compute the following 22-point functions:

For n=2,4,6n=2,4,6, there are respectively ncomb=21,126,462n_{\text{comb}}=21,126,462 inequivalent combinations. We denote by ⟨⋅⟩\langle\cdot\rangle the average of a quantity over all possible combinations of points, and by ⟨∣⋅∣⟩\langle|\cdot|\rangle the average of the absolute value.We use the same notation as the average over the networks. However, there is no ambiguity since the latter is used in this section only to compute Green functions and never appears after.

The numerical Green functions are exact Green functions because they contain already all quantum corrections from loop diagrams. Hence, it is more natural to write a 1PI effective field theory and determine the coefficients by matching the Green functions computed from 1PI Feynman diagrams. Moreover, the Wetterich formalism from Sections 3 and 4 gives relations for the 1PI couplings. We consider the following 1PI interactions to describe the neural network:

where SkinS_{\text{kin}} is the large NN free action (2.13). We consider a local Lagrangian because it turns out that it reproduces well the experimental Green functions for the points considered previously . In the notations of , we have u4=4! λu_{4}=4!\,\lambda and u6=6! κu_{6}=6!\,\kappa. However, the interpretation is slightly different compared to which writes a microscopic action. The interactions are associated with the part of the kinetic operator ΞW\Xi_{W} corresponding to the weight only, since the bias part is always Gaussian and independent of NN . Hence, propagators attached to vertices are KWK_{W} instead of KK: the latter appear only in the disconnected 22-point propagators.

We now turn our attention to the computation of the experimental Green functions. We take:

Since we know the exact 22-point function G2=KG_{2}=K, we must have:

Similarly, we know from (2.27) that higher-order Green functions must decrease as NN increases:

We check that it is indeed the case by plotting the values of m2m_{2}, m4m_{4} and m6m_{6} for the different combinations of points (A.8). The figures 10 (not present in ) show that these values go toward as NN increases for n=4,6n=4,6. On the other hand, the values for n=2n=2 don’t have any specific pattern, which is expected, since Gexp(2)G^{(2)}_{\text{exp}} should be independent of NN.

We can simplify further this information and extract a single number. To do this, we take the absolute value of the normalized deviations (A.5) and average over the different combinations of points (A.8) to get ⟨∣mn∣⟩\langle|m_{n}|\rangle. Moreover, to get an idea of how small are the normalized deviations, we define a background as follows: we compute the standard deviation of mnm_{n} over all bags of neural networks for all combinations of points (A.8), and then average over the latter. The idea is to compare the normalized error encoded by mnm_{n} with its numerical fluctuations over different bags, represented by the standard variation. On the figures 11, we reproduce the results from : ⟨∣mn∣⟩\langle|m_{n}|\rangle for n=4,6n=4,6 is below the background only for small NN and for N=1000N=1000, and it is always below the background for n=2n=2. In principle, ⟨∣mn∣⟩\langle|m_{n}|\rangle should always be below the background for higher NN (which was not studied in the original paper and which we could not reach for computational reasons) so the current test is not very sharp; Figures 10 give a cleaner assessment.

Next, we can compute u4(x1,x2,x3,x4)u_{4}(x_{1},x_{2},x_{3},x_{4}). Using Feynman rules, it can be obtained by subtracting the disconnected contributions (equal to G0(4)G^{(4)}_{0} and built from the 1PI 22-point function) from the full 44-point function to extract the contact interaction

where KWK_{W} was defined in (2.11) (see for more details). Importantly, this equation is really an equality and not an approximation as in : since we are working with 1PI diagrams, there are no quantum corrections and any nn-point Green function is built from vertices of order n′≤nn^{\prime}\leq n. Higher-order vertices m>nm>n appear only in loop diagrams, which are not present. Hence, this allows determining all 1PI couplings exactly in a recursive way. Our results agree quantitatively for u4u_{4} with those of because the loop corrections are subleading in the large NN expansion. However, this may give different results for u6u_{6} since the latter receive loop corrections from the microscopic quartic vertex.

We find that u4u_{4} is constant to a very good precision when evaluated over all combinations of points (A.8). In Figure 12, we display the values of u4u_{4} averaged over all combinations and the corresponding standard deviation and find that its absolute value decreases as NN increases, reproducing the results . Importantly, we find that u4u_{4} is negative, which was not indicated in (their figure 4 has an implicit absolute value needed to use the log-scale). As a consequence, the effective action (A.10) must include a sixtic contribution, however small, for the path integral to be stable: truncating to quartic interactions as in leads to an exponential growth of the weight. A preliminary analysis of the passive flow equations (Section 3) indicate that they can be integrated over a large range of kk only if the initial conditions satisfy u4<0u_{4}<0 and u6>0u_{6}>0, otherwise the flow diverges.

The final numerical test we perform in this paper is to compute u4u_{4} as a function of NN and σW\sigma_{W} (Figures 8, 9 and 13). We consider the following values of σW\sigma_{W}:

We see that u4u_{4} decreases as σW\sigma_{W} and NN increase and that the values are well predicted by the active RG flow equations (4.14). As such, knowing u4u_{4} for a single σW\sigma_{W} at fixed NN allows computing it for any other σW\sigma_{W}.

Appendix B Proofs and technical discussions

The term involving the field strength ZZ in the truncation takes the form

where ZZ is assumed to depend on MM. Because we furthermore assume it to be independent of p2p^{2}, we can define ZZ operationally as:

The flow equation for Γk(2)\Gamma_{k}^{(2)} can be deduced from the Wetterich equation,

In the local potential approximation, the vertices are momentum-independent. Therefore, the contribution involving Γk(4)\Gamma^{(4)}_{k} can be discarded, leading to:

where, according to local potential approximation, we evaluate the RHS over uniform configurations. The derivative is then easy to compute, leading to:

The expression of Γk(3)(0,0,0)\Gamma^{(3)}_{k}(0,0,0) can be easily obtained by taking the third derivative of the effective potential with respect to MM:

Note that the renormalized vertex Γ(3)ˉk(0,0,0)\bar{\Gamma^{(3)}}_{k}(0,0,0) has to be defined as (the factor ZZ will be explained below):

We focus on small and positive pp along axis 11. The integral decomposes as In(k,p)=In(+)(k,p)+In(−)(k,p)I_{n}(k,p)=I_{n}^{(+)}(k,p)+I_{n}^{(-)}(k,p), where:

Because p>0p>0, in the negative branch, (q1+p)2<k2−∑i=2qi2(q_{1}+p)^{2}<k^{2}-\sum_{i=2}q_{i}^{2}, we have:

which is independent of pp. In the positive branch, in contrast, we get:

where the bounds in the integrals refer to the integral over coordinate 11, and M2(q⊥):=M2+Zq⊥2\textbf{M}^{2}(q_{\bot}):=M^{2}+Zq_{\bot}^{2}, for q⊥:=(q2,⋯ ,qdin)q_{\bot}:=(q_{2},\cdots,q_{d_{\text{in}}}). Note that we omitted the Heaviside functions. Taking the first derivative with respect to pp, we get:

Next, we take the second derivative and we set p=0p=0. The contribution coming from derivative of the interior of the integral vanishes, because the remaining integration over q⊥q_{\bot} is empty. Thus, only the variation of the bound contributes; assuming θ(0)=1\theta(0)=1, we get after a tedious calculation:

As for the strict LPA, we introduced dimensionless quantities (and then, explain the origin of the factor ZZ in front of (B.7)). Now, we have to take into account the wave function renormalization. Recovering 1/21/2 in front of the kinetic action requires:

where in this expression uˉ2=−2u4κ\bar{u}_{2}=-2u_{4}\kappa refers to the effective mass. This relation implies κ=κˉkdin−2Z−1\kappa=\bar{\kappa}k^{d_{\text{in}}-2}Z^{-1}. After some simplifications, we get:

The explicit expression for Mˉ2\bar{M}^{2} can be easily derived within the local potential approximation:

B.2 Discussion about Claim 1

The regulator is obviously positive definite as soon as m2>0m^{2}>0. For p∈[0,k]p\in[0,k], because m2(ep2/m2−1)≥p2m^{2}(e^{p^{2}/m^{2}}-1)\geq p^{2}, we must have rk(p2)≤k2(1−p2/k2)r_{k}(p^{2})\leq k^{2}(1-p^{2}/k^{2}), therefore:

Thus, low momenta fluctuations are frozen, decoupling from long distance physics, whereas high momenta modes are unchanged and integrated out. Note that the last condition makes rkr_{k} an infrared regulator, which prevents infrared divergences along the flow. Finally, rk→0→0+r_{k\to 0}\to 0^{+}, meaning that the original model is formally recovered in the deep infrared limit. All these properties ensure that the boundaries interpolation conditions Γk→∞→S\Gamma_{k\to\infty}\to S and Γk→0→Γ\Gamma_{k\to 0}\to\Gamma holds using rkr_{k}. Now, let us show that rkr_{k} is optimal in the Litim’s sense . The Wetterich equation (3.3) can be singular if the effective propagator diverges, or equivalently, if its inverse Γk(2)+rk\Gamma^{(2)}_{k}+r_{k} vanishes. To avoid this difficulty, rkr_{k} has to prevent the existence of zero-modes. In other words, the RG requires the existence of a “gap”, and Litim’s criterion for optimization is to maximize the gap. This can be done from the observation that only the field-independent part of the inverse 22-point function is relevant to discuss optimization, i.e. F(p2):=p2+Zk−1rk(p2)F(p^{2}):=p^{2}+Z^{-1}_{k}r_{k}(p^{2}), in view to establish a (weakly) model-independent criterion. It is suitable to introduce y=p2/k2y=p^{2}/k^{2}. The optimal value for the gap, C0C_{0}, is:

By construction, the regulator is expected to be efficient for p2∼k2p^{2}\sim k^{2}, and we fix the normalization such that F(y=α)=1F(y=\alpha)=1 for some α∈]0,1[\alpha\in]0,1[. For a large enough family of regulators, this condition imposes that C0≤1C_{0}\leq 1. If F(y)F(y) reaches its absolute minimum for y=0y=0, the regulator cannot be optimal from definition. Thus we may have F(0)≥C0F(0)\geq C_{0}. Without loss of generality we may choose F(0)≥1F(0)\geq 1. For a regulator which attributes the same size to the IR fluctuations p2<k2p^{2}<k^{2}, this reduces to an equality; this is the case for the regulator (3.49). Indeed, in that case F(y)=m2k2ek2/m2F(y)=\frac{m^{2}}{k^{2}}e^{k^{2}/m^{2}} for y≤1y\leq 1, F(y)=m2k2eyk2/m2F(y)=\frac{m^{2}}{k^{2}}e^{yk^{2}/m^{2}} for y≥1y\geq 1: the previous argument holds, up to the normalization factor m2k2ek2/m2\frac{m^{2}}{k^{2}}e^{k^{2}/m^{2}}.

B.3 Proof of proposition 2

Because we focus on the symmetric phase, odd effective vertices have to vanish identically Γk(n)=0\Gamma_{k}^{(n)}=0 as n=2p+1n=2p+1. Moreover, the effective propagator Gk(p,p′)G_{k}(p,p^{\prime}) has to be diagonal: Gk−1(p,p′)=(gk(p2)+rk(p2))δ(p+p′)G_{k}^{-1}(p,p^{\prime})=(g_{k}(p^{2})+r_{k}(p^{2}))\delta(p+p^{\prime}). From our ansatz, gk(p)g_{k}(p) is given by:

Taking the second derivative of the exact RG equation (3.3), we get for g˙k\dot{g}_{k}:

Because the windows of momenta allowed by the function r˙k(q2)\dot{r}_{k}(q^{2}) is limited to the region q2<k2q^{2}<k^{2} by construction, the symmetric function Γk(4)(p,−p,q,−q)=:fk(p,q)\Gamma_{k}^{(4)}(p,-p,q,-q)=:f_{k}(p,q) can be expanded in powers of q/kq/k. At leading order, setting q=0q=0 and taking into account the definition 1, the equation simplifies as:

The derivative of the regulator can be computed from the definition 1 as well. We get, for p2<k2p^{2}<k^{2}:From its definition, the regulator vanishes for p2>k2p^{2}>k^{2}, and it is also the case for its derivative.

In the RG transformation, a rescaling of the lattice is required after partial integration of degrees of freedom to ensure preservation of the IR physics. This is equivalent to assuming the existence of a proper rescaling of the coupling, turning the flow equation in an autonomous system. From the equation above, we see in particular that the mass has to be rescaled as:See also the discussion at the beginning of the Section 3.3.

and the rescaling for the couplings f(p,0)f(p,0) and gk(p2)g_{k}(p^{2}) follows:

respectively defining the local 44-point coupling and effective mass. Within these dimensionless couplings, equation (B.27) becomes:

where β2n:=uˉ˙2n\beta_{2n}:=\dot{\bar{u}}_{2n}. Within this approximation, and introducing x:=p/kx:=p/k, the flow equation for g˙k(p2)\dot{g}_{k}(p^{2}) takes the form:

the integral being restricted in the interior of the sphere y2<1y^{2}<1. Thus, defining:

From this equation, we easily deduce that fˉk(x,0)\bar{f}_{k}(x,0) must be a function of x2x^{2}. Setting x=0x=0 on both sides, we get an algebraic closed equation for β2\beta_{2}:

The solution exhibits a singularity, which has to be taken into account when solving the flow equation. Through the definition (B.24), the solution to this equation provides gk(p2)g_{k}(p^{2}). Taking into account that, we find from the chain rule:

where gˉk′:=∂gˉk(x2)/∂x2\bar{g}_{k}^{\prime}:=\partial\bar{g}_{k}(x^{2})/\partial x^{2}. We thus obtain for fˉk(x,0)\bar{f}_{k}(x,0):

These equations depend on local couplings uˉ2\bar{u}_{2} and uˉ4\bar{u}_{4}. The flow of uˉ2\bar{u}_{2} is fixed by the flow equation (B.34), but requires the knowledge of uˉ4\bar{u}_{4}. It can be obtained using standard LPA, equations (3.18) and (3.19) for a sixtic truncation, which discard contributions or order 1/N31/N^{3} from the NN-scaling. We get, for uˉ4\bar{u}_{4} and uˉ6\bar{u}_{6}:

Finally, from the flow equation for Γk(4)(p1,p2,p3,p4)\Gamma_{k}^{(4)}(p_{1},p_{2},p_{3},p_{4}) (equation (3.18)), setting p1=−p2=pp_{1}=-p_{2}=p and p3=p4=0p_{3}=p_{4}=0, we get:

For reader familiar with QFT, the origin of the function Rk(x)R_{k}(x) can be traced from the ss-, tt- and uu-channels. In fact, the (Γk(4))2(\Gamma_{k}^{(4)})^{2} contributions have the following structure (the intermediate fat dotted edge materializing the effective loop between effective vertices, see equation (3.18)):

The first term corresponds to uˉ4fˉk(x,0)\bar{u}_{4}\bar{f}_{k}(x,0), the second defines Rk(x)R_{k}(x). A direct inspection shows that Rk∼∫dq r˙k(q2)Gk(−q−p)Gk(q)R_{k}\sim\int dq\,\dot{r}_{k}(q^{2})G_{k}(-q-p)G_{k}(q), which becomes small for pp large enough. We thus obtain the approximation for hˉk(x)\bar{h}_{k}(x) for xx large enough:

B.4 Discussion about Claim 2

The 44-point function Γk(4)(p1,p2,p3,p4)\Gamma_{k}^{(4)}(p_{1},p_{2},p_{3},p_{4}) has to be symmetric under any permutation of the four external momenta p1p_{1}, p2p_{2}, p3p_{3} and p4p_{4}. Moreover, because we assume local interactions as building blocks, external momenta have to be conserved: p1+p2+p3+p4=0p_{1}+p_{2}+p_{3}+p_{4}=0. Let us assume that Γk(4)\Gamma_{k}^{(4)} is the analytic continuation with respect to some couplings from a perturbative solution Γk,pert(4)\Gamma_{k,\text{pert}}^{(4)}, defined as the formal sum of an asymptotic perturbative series:

where the first sum runs over one-particle irreducible (1PI) Feynman diagrams G4\mathcal{G}_{4} having four external points. The product runs over vertices v∈Gv\in G, 2n(v)2n(v) denoting the number of fields involved in the interaction having coupling constant g2n(v)g_{2n(v)}. Finally, AG\mathcal{A}_{G} is the Feynman amplitude associated with the graph GG. Note that all Feynman amplitudes arise with a global Dirac delta δ(p1+p2+p3+p4)\delta(p_{1}+p_{2}+p_{3}+p_{4}) ensuring momentum conservation. We recall that Feynman diagrams provide a graphical representations of the Wick contractions involved in the perturbative expansion around the Gaussian theory. A typical Feynman graph is a set of vertices and edges, vertices corresponding to interactions and edges to the Wick contractions between pairs of fields. The momentum dependence of the 44-point function can be investigated from the structure of Feynman graphs labeling its perturbative expansion. First, we assume that the the theory involves only 44-points vertices. At one loop, Γk,pert(4)(p1,p2,p3,p4)\Gamma_{k,\text{pert}}^{(4)}(p_{1},p_{2},p_{3},p_{4}) has the following structure:

each term corresponding to the allowed permutations of the external momentaThe so called ss-, tt- and uu-channels.. Explicitly, the relevant one-loop diagrams have the following structure:

the loop being proportional to ∫dqK(q2)K((q+p1+p2)2)\int dq\mathcal{K}(q^{2})\mathcal{K}((q+p_{1}+p_{2})^{2}). It can be easily checked that the decomposition (B.45) keeps the same form after including sixtic interactions. Our aim is to prove that such a decomposition remains a suitable approximation for Γk(4)\Gamma_{k}^{(4)} beyond one-loop. To this end, we will make use of a renormalization group argument, using the explicit expression (B.37). Assuming that (B.45) holds beyond one loop, and there exists a function γk(p)\gamma_{k}(p) such that:

Combining it with the relation (B.37), we get:

and fk(0)≡u4=3γk(0)f_{k}(0)\equiv u_{4}=3\gamma_{k}(0), thus:

The flow equation for Γk(4)(p1,p2,p3,p4)\Gamma_{k}^{(4)}(p_{1},p_{2},p_{3},p_{4}) reads graphically as:

the cyclic permutation covering the three pairings (p1,p2)(p_{1},p_{2}), (p1,p3)(p_{1},p_{3}) and (p1,p4)(p_{1},p_{4}), the solid black edge materializing the effective propagator r˙k(p2) Gk(p2)\dot{r}_{k}(p^{2})\,G_{k}(p^{2}), whereas the dotted edge corresponds to Gk(p2)G_{k}(p^{2}). Let us investigate the structure of the (Γk(4))2(\Gamma_{k}^{(4)})^{2} contribution. From our assumption, and neglecting the dependence of the effective vertex on the momentum qq running through the effective loop, we get:

where Lk(2)(p1+p2):=∫dqr˙k(q2)Gk((q+p1+p2)2)Gk(q2)L_{k}^{(2)}(p_{1}+p_{2}):=\int dq\dot{r}_{k}(q^{2})G_{k}((q+p_{1}+p_{2})^{2})G_{k}(q^{2}). For pi2p_{i}^{2}’s close to the running horizon ξ−2(k):=k2u2(k)\xi^{-2}(k):=k^{2}u_{2}(k), fk(p)f_{k}(p) becomes small as the explicit expression (B.37) shows. Thus γk(pi)∼−u4/6\gamma_{k}(p_{i})\sim-u_{4}/6,

and (Γk(4))2(\Gamma_{k}^{(4)})^{2} contributions ensure stability of the assumption about Γk(4)\Gamma_{k}^{(4)}. To check consistency with Γk(6)\Gamma_{k}^{(6)}, we have to consider the corresponding flow equation:

up to permutation of the external momenta. Assuming we are close to the quartic sector, we focus on the last contribution. Setting p5=−p6=qp_{5}=-p_{6}=q, we have two relevant configurations to investigate. The first one is for p5p_{5} and p6p_{6} hooked to the same vertex. In that case, the loop depends only on two momenta, hooked to another vertex, say p1+p2p_{1}+p_{2} in the following:

where we discarded the dependence of the effective vertices on the momentum running through the effective loop, and we assumed γk(p)=γk(−p)\gamma_{k}(p)=\gamma_{k}(-p). Such a contribution in the first term on the RHS of the equation (B.50) does not break the ansatz for Γk(4)\Gamma_{k}^{(4)}, the remaining momentum qq could be set to zero outside of the tadpole. The second configuration is for p5p_{5} and p6p_{6} hooked to different effective vertices. It is however easy to check that for large external momenta with respect to the IR cut-off kk, these contributions are suppressed. For instance, we have:

and for p42p_{4}^{2} large enough with respect to kk, this contribution is less relevant than the first in the flow equation for Γk(4)\Gamma_{k}^{(4)}. From the same argument, it is easy to check that the second kind of contributions in the flow of Γk(6)\Gamma_{k}^{(6)}, involving Γk(6)\Gamma_{k}^{(6)} and Γk(4)\Gamma_{k}^{(4)} does not break the ansatz for Γk(4)\Gamma_{k}^{(4)} in the range of momenta that we consider.

References