Nonperturbative renormalization for the neural network-QFT correspondence
Harold Erbin, Vincent Lahoche, Dine Ousmane Samary
Introduction and outline
Deep learning and neural networks (NNs) have experienced a rapid development in the last decade, with an ever-increasing number of remarkable applications. In many cases, these systems outperform humans and ordinary algorithms. However, there are still many challenges to be solved: in particular, most neural networks work as a black box and require a huge number of examples during the learning phases. More generally, there is no complete theoretical understanding of why deep learning works so well and how to improve it further. For example, it is not clear how training can be made more efficient and fast, how knowledge can be transferred to other tasks or how to choose hyperparameters systematically. This lack of reliability poses, in certain cases, important ethical problems. Indeed, with the growing use of AI for making decisions (for example, in banking, employment, medicine, military, etc.), it is crucial to be able to explain the choices of the AI in a transparent way . Moreover, having a black box is also a drawback for scientific discovery since the goal of science is to interpret and explain, and knowledge can grow only from understanding . Our paper is part of the lively field of explainable AI , where physics do have a role to play .
A natural path for studying neural networks is provided by theoretical physics: it offers an array of tools useful to describe a wide range of complex systems . In the recent years, evidence has accumulated in favor of a scenario involving a particular form of "coarse-graining", making contact with a familiar tool for physicists: the Wilsonian renormalization group (RG).
A macroscopic ideal gas is completely described by the ideal gas law, and a macroscopic fluid is well described by the Navier-Stokes equation. Both equations ignore the microscopic atomic and molecular interactions and provide a coarse-grained description. This idea was fully developed by Wilson’s renormalization group, formalizing the general feature that long range behavior of physical systems does not require an understanding of the nature and interactions of its microscopic building blocks. Through a very impressive argumentation, Wilson showed that it is possible to explain the apparent universality of physical systems near critical points from the observation that, up to the accuracy of physical predictions, the specific microscopic details can be absorbed through a few effective couplings, defining an effective large scale theory. Despite the fact that the RG was born in the era of critical phenomena, it turned out to be a very general framework, largely responsible for the success of field theory descriptions of long range distance physics, both in condensed matter physics and in high energy physics .
The ability of the RG to explain long range universality can be traced from information geometry . The RG coarse-graining is performed on the eigenvalues of the (free) Fisher information metric, which is a local version of the Kullback-Leibler (KL) divergence (or relative entropy) and which provides a reasonable measure of distinguishability between two probability distributions and . From coarse graining, and in absence of singular structures, KL divergence decreases, as well as distinguishability between distributions, as to become smaller than any experimental precision. Beyond this limit, we cannot distinguish the two distributions, as different as they may have been originally. The ability of the RG to extract the relevant features from a large set of interacting microscopic degrees of freedom is a compelling argument for a link with deep learning. In fact, it is natural to expect a relation with any procedure able to extract relevant features from a massive data set, as it is the case, for instance, in principal component analysis (PCA), where some recent works stressed such a connection between signal detection and RG .
In this paper, we aim at developing further the correspondence between quantum field theory (QFT) and neural networks, called the NN-QFT correspondence . The main objective is to provide a description of the NN behavior using the non-perturbative RG and the corresponding effective field theory.In this paper, we will mostly use both “quantum field theory” (QFT) and “effective field theory” interchangeably, the context makes it clear if we speak of the microscopic (ultraviolet, UV) or effective theory (low-energy). We are working in Euclidean signature, in which case the term “statistical field theory” is sometimes used, to make clear that it describes thermal and not quantum fluctuations. However, we will use QFT for uniformity. This positions our paper in a growing tradition of papers describing how the behavior of NNs can be understood through a more and less sophisticated coarse-graining, which can itself be related to a RG. Strong evidence in favor of a correspondence between RG and deep learning has been stressed for Restricted Boltzmann machines (RBM), whose architecture exhibits similarities with the Ising model (a theoretical model for ferromagnets) . Historically, the Ising model was precisely the conceptual cradle of the RG through Kadanoff’s "block-spin" method, which can be viewed as an elementary version of the general Wilson coarse-graining . The use of a field theoretical formalism is not a novelty as well . In fact, this is expected since physics has shown that field theories are a general feature for systems involving emergent collective dynamics. For example, they appeared to provide a good understanding of the qualitative behavior of NNs through the spin-glass formalism .
We follow the correspondence between NN and QFT pioneered by Halverson, Maiti and Stoner . Its originality with respect to other approaches lies in the observation that, under very general conditions, NN with infinitely wide layers are described by a Gaussian process (GP) due to central limit theorem . Realistic architectures never involve an infinite number of hyperparameters , and their behavior fails to be well described by a GP. However, this useful but purely theoretical limit allows approaching the non-Gaussian process (NGP) as a perturbation from the large limit, which is assumed to receive corrections which, for large enough, can be computed perturbatively. In , the correspondence has been developed in the case of a fully connected network with a single hidden layer of width . They developed the field theoretical machinery necessary to describe neural networks (see Appendix A for a summary). This includes computing correlation functions of outputs, obtained in QFT by constructing Green functions from Feynman rules in perturbation theory. Effective interactions (also called couplings) can then be extracted by comparing the NN correlation functions and the QFT Green functions. Finally, they introduced a RG flow from a cut-off on the volume of the input data, from the assumption that the effective field theory must be insensitive to the choice of the volume, up to a global rescaling of couplings entering in its definition. In the QFT language, this corresponds to an infrared (IR, large volume) cut-off: a major difference in our paper is that we will use an ultraviolet (UV, data resolution) cut-off (see Section 2.1.3).
The relation between different effective models can be translated locally through a set of -functions which describe the evolution of couplings when the cut-off changes, with universal features of the theory emerging from the flow. In this paper, we are aiming at proposing a non-perturbative formalism, based on the Morris-Wetterich equation , to investigate the NN-QFT correspondence beyond the perturbative regime, i.e. beyond the large regime. This means that our analysis does not requite the coupling constants to be small and that our equations are given in a expansion. As mentioned above, our framework differs from the one used in in that we introduce a true partial integration of the degrees of freedom procedure, without any assumption on the expected large volume behavior of the corresponding effective field theory. Among the major differences with respect to the situation with ordinary QFT, the full (effective or IR) -point function is known theoretically, including non-Gaussian effects, whereas the free propagator (microscopic or UV) is not known. This unconventional setting allows going beyond standard limitations of the non-perturbative framework, in particular to close the infinite hierarchical system of equation describing the RG flow and keeping the full momentum dependence of the correlation functions following the Blaizot-Mendez-Wschebor method .
Another unconventional aspect concerns the notions of power-counting and locality, which are traditionally inherited from the background space-time (which we will call “data-space” in the case of neural networks). In the case of QFT for neural networks, such a relation appears as an additional hypothesis that no experience motivates "a priori". Other properties such as rotation and permutation invariances of the point components may not make sense for neural network data. However, recent works in the context of background independent quantum gravity have shown that the notions of scales and power-counting are more primitive than that of space-time, and such that locality can be derived from power-counting itself, ensuring moreover that the RG exists and is well-defined. In this article, we will discuss how these ideas can be relevant for neural networks.
A RG can be then constructed by following the standard method, partially integrating on the degrees of freedom and starting with those associated with the highest scales (UV). However, the situation for the NN-QFT correspondence is quite different compared to usual studies of the non-perturbative RG: indeed, we are able to solve exactly the -point vertex function while keeping the full momentum-dependence without approximation on the -point function (since it is already known exactly). We can also solve almost exactly for the other momentum-dependent -point vertex functions (when two momenta are equal and the others vanish, we can also find an exact solution without approximation). In this paper, we consider two versions of the RG. In a first approach, called passive, the notion of scale is fixed by the resolution chosen to describe the data. In that approach, the standard deviation of the hidden weights, , is viewed as a reference mass scale. The resulting evolution equation provides an explicit realization of equivalence classes of networks having the same output (up to the machine precision), as the data is coarse-grained. In a second approach, called active, the RG flow is constructed by viewing as a running scale. In such a way, the equivalence class is between networks having the same output, keeping the data resolution fixed. This implies that, for fixed , neural networks with different can be viewed as belonging to the same RG trajectory. In particular, this implies that the renormalization flow can be used to make predictions for any given the results for one of them. We illustrate this by describing the behavior of the quartic coupling constant of the effective field theory and check numerically the flow equations. In this paper, we focus on the analytic result and we plan to extend the numerical aspects in future works.
In Section 2, we discuss some general concepts about the field theories which may be used to describe neural networks. In particular, we comment on the definition of the data-space, IR and UV regimes, (non-)locality and its consequences on scaling and power-counting. At the end, we describe the passive and active points of views for the renormalization group. In Section 3 and 4, we derive the passive and active RG flow equations respectively. In Appendix A, we review the numerical simulations from and provide some additional details. Finally, Appendix B contains the details of technical computations.
NN-QFT, locality, scaling and RG
In this section, we present the framework of the NN-QFT correspondence proposed in . As explained in the introduction, we focus on the Gaussian networkThe term “Gaussian” refers to the fact that the kernel is a Gaussian kernel, not that we have a Gaussian process in the infinite-width limit. (or Gauss-net) which have a translation invariant kernel. In this section, we first recall the main ideas of the correspondence (some numerical results from are reproduced in Appendix A).
Then, we discuss the role played by non-local interactions.They were not considered in the original version of but additional discussion has been added in a subsequent version during the preparation of this manuscript. In particular, we describe the different ways to relax locality and how this naturally leads to break the rotation invariance of the data. The most general QFT in the latter case are called random tensor field theories (or group field theories), which are generalization of random matrix field theories.
We are also revising the concept of power-counting, preferring a notion intrinsic to the renormalization group compared to the one used in , which is inherited from a background “data-space”. In the latter case, a classical scale dimension is attributed to the data and dimensional analysis is performed by requiring that the action is dimensionless (such that its exponential can serve as a weight in the path integral). However, it is not clear how to extend this notion in the presence of non-local interactions. We introduce two notions of scales which emerge from the analysis: the first is attached to the data and called “working precision”, and the second is attached to the network and called the “observation scale”. We consider two versions of the RG, flowing in these two parameter scales. We conclude the section with a short presentation of the Wetterich-Morris formalism for non-perturbative RG and a discussion about the RG version considered in the reference paper . Note that we voluntary use the same notations and conventions to make the comparison with their results easier.
In , the authors proposed a general QFT framework to describe the statistical behavior of neural networks (NN), working in the function-space rather than parameter-space (which can be viewed as a duality ). The original motivation stems from the observation that neural networks in the infinite-width limit are described by a random Gaussian process : the latter can also be described by a free (or Gaussian) quantum field theory.We refer to for a gentle introduction to QFT with neural networks in mind. When the width is finite, the random process is not Gaussian and one can expect the NN to be mapped to an interacting field theory, which has been checked in .
where the weights and biases characterize the affine transformation of each layer and acts element-wise. The weights and follow centered Gaussian distributions and respectively, and both biases and are drawn from centered Gaussian distributions . The input data is a -dimensional vector, while we take for the output data for simplicity. As a consequence, is a -matrix, a -matrix, a -vector, and a scalar. The Gauss-net activation is slightly peculiar because it acts as an exponential of the layer output normalized by the data of the previous layer:
Finally, we stress that the NN is randomly initialized and that we will not consider the effect of training.
Information on the neural network can be extracted by considering correlations of the outputs: they are encoded by the “experimental” correlation (or Green) functions :
where the statistical averageThe notation should not be confused with the expectation value in QFT: we will always use it to denote the statistical average over a set of networks. is taken over a large number of neural networks with identical and parameter distributions. The numerical evaluation of these quantities is explained in Appendix A.
1.2 Large N𝑁N: free field theory
where the factor ensures that the expression is normalized when integrating over the full functional space
The origin of this Gaussian behavior in the limit can be traced from central limit theorem: since is formally a sum of identically distributed random terms which self-average. The function splits into two contributions:
where being essentially independent variables and following the Gaussian law , whereas goes toward a Gaussian distribution only for large . Formally, it reads as:
where is given by (2.2) (such that depends on and ). As stated above, for large , one expects that such a quantity self-averages around its mean, and thus that fluctuations are small:
where the last equality follows from the assumptions that initial distributions for are centered and non-correlated. Hence, the statistical properties of are essentially given by a centered Gaussian distribution, up to corrections. Obviously, the random nature of is inherited from the initial parameter distribution, however the asymptotic Gaussian behavior arises from the law of large numbers.
The -point correlation (or Green) function
is the inverse of the Gaussian kernel which appears in the free action (2.4):
However, according to (2.6), it is also possible to decompose as:
where is the -point function associated to . It corresponds to the Fisher information metric in the information geometry language, and is fixed from the choice of the activation function.
In this paper, we essentially focus on translation invariant kernels , which is achieved by the Gauss-net architecture (2.2), corresponding to the kernel:
where denotes the ordinary Euclidean distance between and .
the corresponding probability distribution being given by the exponential law in (2.4). The -point correlation (or Green) functions are defined as:
In the free theory, is completely determined in terms of through Wick’s theorem, and vanishes for odd . Hence, this implies that:
1.3 Data-space and momentum space
In this subsection, we discuss some definitions related to the data-spaceAs we will see, its properties may be sufficiently different from the usual space of positions – spacetime – appearing in usual QFT to find another name. corresponding to the neural network input , and how they differ from .
It is generally more convenient to work in Fourier (or momentum) space. The allowed momenta lie in the first Brillouin region:
Note that we assume periodic boundary conditions. This is not a problem for large enough; for small ,We will precise “small with respect to what” in a moment. we can simply repeat the data set a large number of times to obtain a large enough effective volume to make the boundary conditions irrelevant.
In the rest of this paper, we use the following definitions.
We call the discrete volume and the working precision.
the basis functions being normalized such that . Note that (2.18) holds for any discrete function on the lattice. In the continuum limit, for small and large such that remains fixed, discrete sums can be replaced by integrals. Moreover, for volume large enough, integrals becomes standard Fourier transform. We call this limit the thermodynamic limit following the standard terminology in physics, and we focus on this regime in our investigations. Taking the Fourier transform of the -point function:
where , we get for (2.12):This type of kinetic term is reminiscent of -adic string theory .
Note that in this continuum approximation, the Dirac delta has to be understood as a shorthand notation for a Kronecker delta . Translation invariance is crucial to obtain a kernel (2.20) which depends on a single momentum, and reflection invariance implies that it must be a function of . It would be interesting to understand how to generalize our computations for kernels which are not translation invariant .
Up to corrections, the propagator looks like the canonical propagator of a free scalar field theory:
In the QFT terminology, and are respectively the wave function renormalization and bare mass. One can rescale the field to to set , in which case the mass becomes:
The mass defines the typical mass scale and its inverse defines the (IR) correlation length , or the typical observation scale
Note that the large volume limit is defined with respect to this correlation length, i.e. . Beside the existence of an intrinsic length scale, the system the propagator at large distance behaves like .
Assigning the label (position space) to the original data and (momentum space) to the Fourier conjugate may seem arbitrary. Indeed, while the machine precision provides a natural UV cut-off and an associated identification of the data-space as position space (since UV corresponds to small distances in that space), signals in Fourier space are also represented in the computer up to the machine precision. However, in that case, using the machine precision as a UV cut-off would not match the usual intuition in QFT. Given a translation-invariant kernel, another possibility is to identify the momentum space as the space where the propagator is diagonal, such that the propagator in position space depends on the distance .
1.4 Finite-N𝑁N corrections and interactions
For a Gaussian process, as we have seen earlier, correlation functions for can be decomposed as a sum of product of -point functions thanks to the Wick theorem . For large but finite, the distribution is not exactly Gaussian, and the correlation functions do not match with Gaussian predictions.
The deviation of the QFT and experimental correlation functions from the Gaussian case are denoted as:
Note that are still the large Green functions defined in (2.14). Importantly, we identify the exact -point function with the kernel which equals . Since it contains (quantum) corrections due to the interactions, the free -point Green function computed from the kinetic term only is not known (in standard QFT, the converse is true, see Section 2.3.2 for a discussion). We will see in Section 2.2.3 that connected functions behaves as:
which has also been investigated analytically and numerically in (see also Appendix A).The original version did not contain this discussion which has been added while the present manuscript was in preparation. This scaling is consistent with the fact that the exact -point function is independent of .
Qualitatively, this is reminiscent of what happens for the Ising model in large dimension. The local magnetization self-averages because the number of closest neighbors is large, and the statistical properties remain (quasi)-Gaussian. For space dimension large but finite, thermodynamical quantities can be computed as power series in , which do not affect universal quantities, as soon as . For however, the decoupling of physical scales breaks down and the Gaussian approximation is not suitable .
The same scenario is expected to be true for neural networks. For finite , the distribution does not obey Wick theorem, and correlations functions receive contributions which do not reduce to products of -point functions, and the classical action must include non-Gaussian contributions, i.e. products of of degrees higher than . However, as soon as remains large enough, deviations from the Gaussian behavior are expected to remain small. In the classical action, these corrections materialize as product of fields (), that we call interactions:
2 Locality, scaling(s), and power-counting
In this subsection, we precise the class of interactions which are assumed to suitably reproduce the non-Gaussian properties of correlations. Moreover, we discuss the scaling behaviors, especially relevant for the renormalization group investigations in the next section.
However, this makes various assumptions which may not be valid for a general neural network QFT. For this reason, we will make them explicit and explain how to gradually lift them to consider the most general QFT. Deciding which assumptions to use should be dictated by numerical evidences: in particular, it was found in that (2.29) is sufficient for the activation functions and range of input parameters considered there (see Appendix A for more details). This approach can be considered as neural network phenomenology, in the sense that we are writing a model to match observations, but we can also use this model to check theoretical facts such as dualities .
The first assumption is locality of the interaction: the fields appearing in the monomial can be taken at different points (for simplicity, we consider a single coupling in ):
This breaks locality because fields at different points in space(time) can interact together. In fact, since is a constant, this happens for arbitrarily large distances. Note that this preserves translation invariance .
The next natural step is to replace by a coupling function, i.e. a function of space but independent of the field. Going back to (2.29) where all fields are at the same points, we can write a local action with a coupling function:
It was argued in from technical naturalness that must be approximately constant since a coupling function breaks the translation invariance of the action. However, this is correct only when assuming locality of the action: replacing by a coupling function in (2.30) gives the non-local action :
However, translation invariance can be preserved if depends only on the distances between the points:
In order for this to make sense, the Fourier transform of the kernel must be an entire analytic function (with rapid decay if one wants to ensure UV finiteness). This corresponds to a coupling function
Smeared fields naturally appear in string theory and are responsible for its well-behaved UV behavior . In fact, for the Gauss-net, rescaling the field to remove the exponential from the kinetic term (2.20) is equivalent to smearing the field (as pointed out earlier by comparing with the -adic string ).
However, it is not clear a priori that the data-space possesses this symmetry: it may not be possible to exchange two data components or even consider linear combinations if the components are not homogeneous. Despite the fact that the free theory supports such a symmetry, the rotational invariance has no meaning for a neural network in general. Moreover, symmetries of the free theory can be broken by interactions, which are necessary to fully characterize the system. Conditions under which input and output symmetries can be present has been analyzed in . In general, one can start by assuming no symmetry in order to describe the most general model, and then adapt to what the numerical experiments are indicating.
where , and are the components of the -dimensional points , and . Note that nothing prevents to use only two points, for example setting and integrating only over and , or more generally to repeat the same component in any number of fields (for an early example, see ).
Such general theories are too wild and it is hard to make sense of them. A controllable subclass is provided by random tensor field theories . In this case, the fields are tensors, each component of the positions being seen as a (continuous) index and indices can be contracted pairwise only (which is achieved by integrating over the component, since the index is continuous), such that a given component can appear at most twice. An intuitive way to represent it is to assign a color to each component, and Feynman diagrams can be written in terms of strand graphs (generalizing ribbon graphs from matrix models). For instance, a possible quartic interaction for is:
We will see in the next section that tensor field theories are particularly interesting in the RG approach because, under some additional conditions, they possess a natural background-independent power-counting (see next subsection).
We conclude this section by clarifying a subtlety concerning QFT in curved spaces. In this case, the Euclidean (or Poincaré) group is not a global symmetry (symmetries of the action are given by the isometry group of the background space) and one may ask what is the difference with tensor field theories. The point is that this group is still a local symmetry (general relativity can be seen as gauging the Poincaré group) such that the properties discussed above continue to hold. Indeed, one can always consider the tangent space associated to a point: since it is isomorphic to flat space, it means that the coordinates are still homogeneous.
2.2 ΛΛ\Lambda-scaling and power-counting
We are aiming to construct a field theory which admits a well-defined RG flow. In standard QFT, the rigorous construction of such a flow requires essentially three basics ingredients: 1) a scale decomposition, 2) a locality principle, 3) a power-counting.
Let us illustrate heuristically on a simple example how contractibility and power-counting allows fixing the scaling dimension of couplings. Consider the following classical action:
for some cut-off for large momenta. In the same way, the first correction for , say involves the following integral:
the upper index referring to the number of vertices involved in the Feynman diagram. Now, to obtain a well-defined power-counting, the correction for has to scale with in the same way as itself. This is solved by , and we say that the -scaling of is . This moreover implies , and thus . Now, we have to check that it is coherent to all orders of the perturbative expansion. To this end, let us consider a Feynman graph , of order contributing to the perturbative expansion through the amplitude . Contracting along a spanning tree , we reduce the original number of propagator edges to , and the resulting graph looks like an effective (local) vertex, having loops of length one (tadpoles). Each tadpole behaves like , and thus scales as . The divergent degree for the contracted graph is therefore:
Because the contraction procedure removes propagator edges, it increases by : . Finally, because the interaction is quartic, we have the relation , being the number of external edges. Finally, we get:
Each vertex contributes a factor , and the scaling ensures that all the quantum corrections have the same scaling. Moreover, setting , we get , in agreement with the one-loop scaling dimension for mass.
Obviously, because this theory is local in the usual sense, the derived scaling dimensions are exactly the same as the one derived from the standard dimensional analysis of the classical action. The two methods however do not coincide for non-local interactions such that (2.38). We argue that this more abstract way to think about locality, scaling and power-counting is more appropriate in a context where the construction of the theory space is not guided by experimental evidences, such that it seems more appropriate to work from the outset within a sufficiently broad framework to accommodate future developments in formalism. However, the exploration of these aspects for NNs is beyond the scope of this paper, since standard locality seems to hold for the Gauss-net kernel .
2.3 N𝑁N-scaling
There exists another scaling dimension, called -scaling, associated to the behavior of correlation functions with respect to the width of the hidden layer. The Gaussian universality for large ensures that the couplings behave as for some positive function .
The computation can be done by returning to the definition (2.7) and using the fact that follows a centered Gaussian distribution with variance . For instance, we find:
which is of order . The computation of higher correlation functions can be done using a similar strategy, from the assumptions that the with different index are statistically independent variables. This in particular ensures that:
from this observation, a tedious calculation which is given in shows that the connected -point function has to scale as , and more generally that .
This analytic result also shows the limitations of the approach. Indeed, we expect that a more fundamental method will be able to predict the weights of the interactions. Moreover, the derivation assumes the relation (2.46), and thus the independence of the having different indices , but such an assumption seems to be in conflict with an interaction such that (2.29), which morally must introduce couplings mixing different outputs from the definition (2.7) of . One may expect that these difficulties could be solved by working with a random vector of size , with components rather than with the function , defining it as an observable i.e. the vacuum of the corresponding theory. However, the construction of such a theory is going beyond the scope of this paper, and we plan to investigate it in a forthcoming work.
3 Renormalization group
The RG is probably one of the most important concepts discovered in physics during the last century and forms together with field theory the reference framework of modern physics, from condensed matter to high energies. Pioneered in the works of Wilson and Kadanoff , RG is based on the idea of organizing the theory according to length scales, integrating out short distance degrees of freedom following a recursive procedure called coarse-graining and providing an effective description for the long distance degrees of freedom, through an effective action where microscopic interactions are hidden in effective interactions. Note that RG is in fact a semi-group, which is non-invertible. Thus at each step, information is lost, and RG can be viewed as a systematic procedure to extract large scale relevant features.
To illustrate the physics underlying the Wilson procedure, and before making contact with NN field theory, let us consider a physical system made of a single real scalar field whose configuration probability follows the exponential form , for some classical action . To have a concrete example in mind, we can take for the real field described by the classical action (2.40). All the statistical properties of the distribution can be derived from the generating functional (partition function):
This integral being over all configurations for , all the degrees of freedom are integrated out in one step. Equation (2.47) provides a canonical definition of what is microscopic and what is macroscopic, two limits that we conventionally call ultraviolet (UV) and infrared (IR):
In the UV limit, no fluctuations are integrated out. The field configurations are therefore fixed from the extrema of the classical action .
In the IR limit, all fluctuations are integrated out. The configurations are fixed by a new action , called effective action.
The effective action is in turn defined as the Legendre transform of the free energy ,
the classical field being defined as .
bounded by UV and IR effective physics (Figure 1).
Let us illustrate how that works on the concrete example of the scalar field described by action (2.40). In that case, corresponds to the spectrum of the Laplacian , whose eigenmodes are Fourier modes, and . Assuming continuity of the spectrum, we can consider infinitesimal coarse-graining, integrating out slices of infinitesimal thickness. This leads to a differential equation describing how the couplings change as the reference scale changes. Formally, this can be done as follows. We assume the existence of an upper bound for , say , and we call the spectrum of the free -point function with cut-off , . As the cut-off moves, degrees of freedom are added or removed from the spectrum. Thus, let us consider the bare action “at scale ”:
where includes interactions following our definition of section 2.1. Now let us consider the running cut-off , for , which interpolate between the UV scale , and IR scale . If is at least in , we can consider the variation at first order from to :
The following decomposition can be translated as a partial integration from the original partition function using the functional identity:
and denotes the degrees of freedom integrated out. Indeed, defining:
we show that the identity (2.52) can be rewritten as:
The classical action at scale formally looks like the action at the scale . What is different between them is the interaction, which comes at scale from a partial integration over the field . The transformation (2.54) can be translated in a differential equation for small enough. Indeed, in this limit, the modes have a large mass, and can be treated perturbatively. Thus, expanding in powers of and keeping only terms of order , we get Polchinski’s equation :
This equation is formally “exact”. However, this has the reputation to be very hard to solve for many reasons. The first one is that it takes place in a functional space of infinite dimension. If we decide to work in a reduced phase space, taking into account only the most relevant interactions, difficulties appear, instabilities with respect to the considered truncation appear as soon as we try to get beyond the perturbative sector, which is precisely what we are aiming at in this paper. For this reason, and as it is the case for the largest part of non-perturbative investigations in the literature , we will prefer to use the Wetterich formalism, better for dealing with non-perturbative approximations. We will discuss this method in the next section.
3.2 Renormalization group(s) for the NN-QFT
The analogy between NNs and RG is evident: both are aiming at extracting relevant features from a massive number of degrees of freedom. RG shows that microscopic details can be ignored to describe long distance physics, and that microscopic theories can be indistinguishable from their common large distance properties. Extracting regularities from large sets of data is exactly what machine learning does; and, as we recalled in the introduction, the question of the relevance of the renormalization group in artificial intelligence is growing in the literature . However, the effective field theory that we presented in the first part offers a new framework to discuss aspects related to the renormalization group in the study of the behavior of NNs . As the previous section stressed out, the field theory that we consider exhibits strong similarities with theories usually considered by physicists: the long distance (i.e. large volume, small momenta) limit (2.22) of the free propagator being the same as for the usual scalar field described by the action (2.40). This formal similarity will serve as a guide in the construction of the renormalization group, and it is very tempting to carry out a coarse-graining in momenta, exactly as for the scalar field in the previous section. We will discuss two different coarse-graining strategies which we call respectively passive and active RGs. But before we go into them in detail, let’s make a few general remarks about what distinguishes this NN field theory from ordinary theories.
In the standard scenario, what is known is the UV theory i.e. the classical action. This action itself is viewed as an effective description, valid at some fundamental scale and ignoring the details about nature and physics of microscopic degrees of freedom underlying the physical world. The choice of the classical action is constrained by predictivity (which promotes just-renormalizable theories), consistency with quantum effects (compensation of anomalies in gauge theories, for instance), and the effective structures at the scale at which the theory is defined, which generally implies some symmetries (rotation, reflection, gauge invariance…). In this respect, the RG aims to provide an approximation of the exact quantum theory, and to compare it with experiments. The case of the field theory that we consider differs from this general picture in its relations between UV and IR scales. The propagator (2.12) is exact and defined in the deep IR. From a RG point of view, the knowledge of this propagator takes into account all the fluctuations at all scales. But, for finite , the knowledge of the -point functions is not sufficient to reproduce higher correlations functions, and non-Gaussian interactions are required in the classical action to reproduce experimental correlation functions. But, due to these interactions, the flow of the different ingredients entering in the definition of the classical action becomes non-trivial, with the consequence that both and in the UV are unknown. Thus, in some sense, the situation is the inverse of what we do in ordinary field theory: we have to infer the form of the UV theory (or more likely a class of UV theory) from the knowledge of only a part of the IR theory. By construction, such an inference cannot lead to a single solution, but a class of solutions which have to satisfy the following requirements:
reproduce the exact -point functions up to the experimental precision;
reproduce the deviations from Wick’s theorem, due to interactions, and which are less and less perturbative as becomes small, once again up to irrelevant corrections with respect to the experimental precision.
Any measurement in physics comes with a finite precision: hence, two effective descriptions are considered to be equivalent and sufficient to describe something if the predictions agree up to the experiment precision. The precision is also finite in numerical simulations, and this explains why that we are able to infer only an equivalent class of models rather than a point in the theory space. In the first section, we showed that the relative relevance of the interactions is not the same such that irrelevant interactions contribute below the machine precision threshold, meaning that we have no way to distinguish between several initial conditions whose trajectories are sufficiently close in the IR (see Figure 2). This argument allows working, in a first approximation, within a finite subspace of the full theory space, focusing on interactions having the largest canonical dimension.
Because of the existence of an intrinsic length scale defined in (2.25), we can think to partially integrate microscopic degrees of freedom with respect to this length scale to construct a proper RG flow following standard field theory. In this picture, what is playing the role of a microscopic scale is the working precision (see Section 2.1.3), which introduces a cut-off in momentum integration . We can then construct a coarse graining procedure from grid size dilatation (see Figure 3).
Note that from such a procedure, one needs to have . Because the maximal value for is , we can show that it implies , which invalidates the expansion (2.21). However, it may happen that such an expansion holds in a sufficiently large domain. A necessary condition is that the expansion (2.21) holds for the smallest (nonzero) momentum , implying:
which is the condition defining the large volume limit.
A dilatation procedure as described in Figure 3 induces a renormalization group by partial integration of momenta into the windows . The existence of two complementary limits and is reminiscent of a crossover scale behavior, between a deep ultraviolet (UV) limit and a deep infrared limit (), which we will study separately in the next section. Note that such a crossover scale appears generally in situations where two very different mass scales appear, ensuring decouplingSuch a decoupling is at the origin of the large mass expansion in field theory. of some effects associated to the larger one when experiments focus on the first one . Here, what plays the role of a large mass is the inverse of the typical observation scale ; in the very large mass limit, , and the IR sector recovers all the physics. In the opposite limit, , everything is UV, and an expansion such that (2.21) does not hold. In other words, for , one expects that quantum effects are suppressed with powers of . This observation can be a source of improvement for approximations used to solve the RG flow equation (3.3) in the next section. In particular, we understand that contributions coming from higher couplings will tend to stay small if they are at the transition scale . Section 3 is devoted to this RG strategy.
In the process described above, the observation scale is kept fixed and the working precision is changed. Conversely, we can keep the data (i.e. the working precision) fixed and change the observation scale. If the first version is essentially passive with respect to the NNs (i.e. the latter is not changed), this strategy is, in contrast, active (see Figure 4). Indeed, remembering the expression (2.25), is completely determined in terms of , the standard deviation of the weight distribution. Hence, flowing in the observation scale is equivalent to changing the weight standard deviation, and thus the neural network.
Physically, if we think of a thermodynamic system like a ferromagnet, such a strategy is equivalent to turning the thermostat’s knob to lower the temperature towards the critical regime. This alternative point of view is the subject of the section 4.
The active RG is closer to the RG version considered in the reference than the passive scheme, as the flow equations derived in section 4 show explicitly. However, despite this formal contact, our approach differs by its very construction. While, from their point of view, the RG is the mathematical explanation of a principle of invariance with respect to a certain volume “cut-off”, our RG is the result of a procedure of partial integration of the degrees of freedom of the field.
Indeed, the RG flow is usually performed with respect to a UV cut-off (spacetime /data-space resolution) and not an IR cut-off (volume). In , the large volume cut-off was introduced because the -point function diverges at large distance (at least, for the ReLU-net), which reminds the short-distance divergence of the canonical propagator in particle QFT. Moreover, one can ask whether the data-space should be identified with the position or momentum space in usual spacetime QFT, and in principle this could depend on the problem. From our arguments in Section 2.1.3, it seems more natural to identify the data-space with the position space (except in the case where the data is already the Fourier transform of a space(time) process). IR divergences are also present in particle QFT, and they are cured following different methods according to their origin. The first case are massless particles, for which a refined definition of amplitudes is needed . An infrared cut-off (such as a mass) can be introduced at intermediate stages to regulate the integral, but it is not a renormalization parameter. In practice, the divergences of the ReLU-net arise from a similar origin (singularity in the propagator for large distance / zero-momentum). Second, IR divergences appear for internal on-shell propagators: they translate the fact that quantum effects shift the vacuum and masses of the fields. Resummation of quantum effects through renormalization leads to finite results . Third, some quantities can diverge for infinite volume for example when studying phase transition: in that case, the usual method is to study the theory with different values of a volume cut-off and to extrapolate to infinite volume (thermodynamic limit) . However, this is not a renormalization flow. For these reasons, we take a more conservative approach and identify small resolution in data space with the UV limit, and perform the RG flow for the associated cut-off.
It is also noted that in [1, sec. 4.3] that the Gauss-net does not require renormalization because the -point function is exponentially decaying with the distance, such that all integrals are convergent. In fact, the previous paragraph shows that renormalization is still needed in this case because its role is not only to handle properly (spurious) UV divergences, but also to take into account quantum effects (some of which lead to IR divergences). Said another way, renormalization provides a mapping between the bare and physical parameters (at a given energy scale) there is always a renormalization flow in the space of couplings. Indeed, the bare parameters describe the properties of the fields without interactions: they are not physical because fields do not live in isolation and any measurement implies an interaction. A famous example of a perfectly finite theory but which has an infinite number of finite (such that predictivity is not lost) counter-terms and non-trivial RG flow (with the so-called stub length) is string field theory .
Flowing through NN-QFT theory space: the passive RG
In this section, we show how the passive RG within Wetterich formalism allows predicting the behavior of correlation functions for a fully connected NN with a single hidden layer. We start with a short presentation of the Wetterich formalism, before turning on to applications. We will consider separately two different regimes, the deep IR regime where the effective propagator can be suitably approximated with an ordinary Laplacian , and the UV regime , where the propagator follows the exponential law .
In section 2.3.1, we provided a formal introduction to Wilson’s ideas for the RG. In this section, we present another incarnation, the so-called Wetterich formalism , which focuses on the effective action for integrated degrees of freedom rather than on the effective classical action for the remaining degrees of freedom, as it is the case in (2.57). We focus on the passive RG as presented in the section 2.3.2. Let be some reference working precision and . Assuming that we performed partial integration up to the scale , we denote as the effective action for those averaged degrees of freedom. Obviously, it must satisfy the boundary conditions:
, no fluctuations are integrated out, and the effective action reduces to the classical action.
, all fluctuations are integrated out and we recover the full effective action defined in (2.48).
The Wetterich formalism aims to construct a smooth interpolation between these two limits. To this end, it is convenient to modify the classical action with a scale dependent mass term , which reads in momentum space:
The substitution defines a -dependent partition function through the definition (2.47). The shape of the scale-dependent mass is designed to freeze low momenta modes , decoupling them from long distance physics whereas high energy modes remain essentially unaffected. Moreover, in order to recover the full classical action for , has to vanish in that limit. In the same way, it has to become very large in the opposite limit, for , in order to satisfy the UV boundary condition (all the fluctuations are frozen). The interpolating functional is defined as:
As varies from to , effective couplings involved in the effective action change. To obtain the differential equation governing the behavior of , as varies, we can differentiate the definition (3.2) with respect to . After a tedious calculation whose details can be found in , we get the following functional equation:
where denotes the -th functional derivative with respect to the classical field . This equation, up to the formal character of its derivation, is as exact as equation (2.57) is. It defines a trajectory through a functional space and is as hard to solve as the equation (2.57). Approximations are required to make the underlying physics tractable. The standard strategy, called truncation, is to identify a relevant finite-dimensional subspace of the full theory space, and to project the flow equation (3.3) onto it. Working with equation (3.3) has the great advantage that this projection procedure does not require to assume that couplings are small, and thus allows investigating approximate but non-perturbative solutions of the RG flow.
2 Local potential approximation in the deep IR
The local potential approximation (LPA) is one of the most popular approximation procedures to solve the exact RG flow equation (3.3). This approximation focuses on the region of the full theory space spanned by local interactions in the sense of (2.29). For the investigations in this section, we assume , which is our reference mass scale. This implies:
which defines the IR regime (see Section 2.3.2). Note that due to the scaling behavior of derivative contributions, one expects that the validity of this description survives in the weak UV regime: due to the large river effect which states, that in a suitable vicinity of the Gaussian fixed point and in the absence of singularities along the flow, the latter projects itself into the subspace spanned by the most relevant couplings .
To begin, we focus on the simplest truncation around sixtic interactions, discarding from our analysis contributions arising from higher couplings. This is equivalent to setting:
where, to avoid confusion with the example given in Section 2.3, we denote as the classical field. Note that such an expansion around is named symmetric phase expansion, and we call symmetric phase the domain of the full phase space where it remains valid. It may happen that such an expansion breaks down, in cases where becomes an unstable vacuum. This is the case when phase transitions are encountered. In this section, we focus on the symmetric phase, and discuss more elaborate formalisms in the next section. Approximation (3.5) ensures that we keep effects up to order .
To be more concrete, we assume that can be decomposed as a sum of two contributions:
The kinetic contribution keeps all the quadratic terms in .
The effective potential gathers the non-Gaussian contributions in the expansion of .
Without lost of generality, the kinetic contribution can be written as:
The kernel is a priori difficult to track. Fortunately, because we are aiming to deal with IR effects, the momentum is expected to be small, justifying to expand in power of :
The first term of this expansion define the running mass, and we denote it as . In the same way the second term of the expansion is called running wave function renormalization, and we denote it as . Nerveless, it is easy to check that in the symmetric phase does not depend on the running scale (see below). Thus we must have . This scheme defines the derivative expansion , and for this section we focus on the two first terms:
To keep only effects up to order , we consider the following truncation for the effective potential:
This form follows from the expression of a local interaction of order in position space: the Fourier transformation to momentum space introduces one momentum for each field together with a delta function for momentum conservation, and a sum (since the take discrete values) over each value of the momentum. The final piece is the regulator . From the choice of the kinetic truncation (3.9), it is suitable to use the modified version of the standard optimized Litim’s regulator :
for which analytic computations are possible. In this equation, is the step function such that and . The flow equations can be deduced from the exact RG equation (3.3) by taking successive derivatives with respect to the classical field . Taking the second derivative gives the flow equation for .
where, on the RHS, functions are computed for . From the truncation (3.9), we must have:
The fourth derivative can be easily computed from the truncation (3.10), leading to:
replacing at the end of the computation. Thus, setting on both sides of equation (3.13), we get after some calculationsDetails are given on Appendix B.:
In equation (3.13), the only dependence on the external momenta and on left-hand side is through the conservation delta arising from the structure of the four-point function vertex . Thus, the field strength , whose flow equation could be deduced by taking derivatives on both sides of equation (3.13) with respect to , vanish identically.
In the same way, taking the fourth and sixth derivatives with respect to of the flow equation (3.3), and from the condition (3.15), we get schematically:
One expects that such an approximation remains valid for , with being large (see Figure 5). Thus, defining,
we get the autonomous system ():
From these equations, it is obvious that the behavior of the flow depends on the dimension . For instance, for , all the couplings are irrelevant and trajectories return toward the Gaussian region, the axis being the only direction of instability. In contrast, for , some couplings become relevant, and trajectories are repelled from the Gaussian region. is the first one to become relevant, for ; for , becomes relevant as well. Figure 6 illustrates the behavior of the RG flow for several dimensions. We have integrated numerically the flow equations for , , in Figure 7.
2.2 Beyond the symmetric phase
In this section, we consider another approximation scheme for the effective potential . Focusing on the IR regime, we assume that essentially reduces to its zero component (the macroscopic field):
and, defining , we expand the effective potential per unit volume in power series around :
Within this parametrization, we identify directly with the (non-zero) vacuum, which runs with the scale . The two-point function is moreover defined as:
Note that we introduced the field strength renormalization because its own flow is nonzero as soon as , i.e. broken phase effect introduces an anomalous dimension. As a technical device, we move the mass contribution in the effective potential. For a uniform field configuration, we must have:
Therefore, taking the derivative with respect to () and writing , we get from (3.3):
As in the previous section, we use the Litim regulatorWhich is optimal in the following sense: the functional renormalization group (FRG) equations are defined only if the effective propagator is well-defined; that is, if has no zero modes (and therefore do not develop IR divergences). This can be achieved by demanding that has a sufficiently large gap, i.e. a sufficiently large minimum. See for more details., but modify it to deal with the running field strength :
For large enough, we may use the same integral approximation as for (3.22), but taking into account that must now depend on the anomalous dimension because of the factor in (3.33):
As in the previous section, we introduce dimensionless quantities (labeled with overlines) as:
Note that all these changes of variable make sense from the requirement that all the terms in the potential must have the same dimension (the dimensions of and have been fixed). The derivative on the RHS in equation (3.36) is taken at fixed. Therefore, we have:
where in the RHS the derivative is taken with fixed. We obtain:
The flow equations can be deduced from the normalization conditions at scale :
Hence, because , we obtain for ,
and after a tedious calculation, we obtain for and :
The computation of the anomalous dimension is long and provided in Appendix B. The result is:
For the Litim’s regulator (3.33), the anomalous dimension in the local potential approximation with kinetic truncation up to order is given by:
In this section, we are aiming to discuss the deep UV regime , where the expansion (2.21) is not valid. For this regime, the derivative expansion breaks down as well, and a local approximation for interactions is no longer justified. We present a method, inspired from the Blaizot-Mendez-Wschebor (BMW) formalism , which considerably improves the accuracy of truncations in regimes where the momentum-dependence of vertex functions is as relevant as the purely local ones, i.e. relevant enough to invalidate the derivative expansion. In this section, we summarize the essential results through three compact statements, in order to focus on the results, and leave the technical details in Appendix B.
The procedure that we propose is based on the following three approximations (see also ):
We parametrize the -point function with a single parameter, the running mass , such that:
such that reduces to the exact -point function in the deep IR for .
We assume that vertices are slowly varying with respect to the momenta running through the effective loops in the flow equation. The allowed windows of momenta being such that , for small enough with respect to other momenta, we require:
The third approximation is about the propagator entering in the flow equation. For in the windows of momenta allowed by , we must have:
where is the Heaviside step function and a positive number, expected to be of order .
To complete these approximations we need to choose a suitable regulator. In principle, we could always use the Litim regulator (or any other regulator used in the literature). However, due to the parameterization of the phase space that we have chosen, and in particular the expression of the -point function, this regulator loses its crucial advantage which consists in freezing all the fluctuations below the scale . The Litim optimal condition is moreover expected to be a relevant constraint to define a regulator, especially in the symmetric phase, and we have the following statement:
satisfies all the requirements for a regulator as soon as , freezes out all fluctuations with momentum , and is optimal in Litim’s sense.
The physical discussion motivating this choice being a little technical, we provide it in Appendix B. Note that Litim’s condition, which relies on the existence of an optimized gap for the inverse -point function , is not an absolute criterion in regard to the reliability of the results. Indeed, some choices are expected to provide an optimal bound for the gap, which may have an influence on the computation of physical quantities like critical exponents . Working within the set of “optimized regulators” in Litim’s sense, we may complete the optimization argument with a principle of minimal sensitivity , requiring that physical quantity has to be stationary with respect to some parameters spanning a family of regulators. This can be done for instance by replacing , and to vary the physical quantities with respect to . This will be the only optimization scheme that we will discuss in this paper.
Within these approximations, the equation for the -point function reads:
which, from the observation that , leads to a closed equation for :
This is the standard BMW strategy. Our approach will however be a little different. First, we work in the symmetric phase . Second, we exploit the fact that the -point function in our parametrization depends only on a single parameter (the mass) to close the hierarchy around the -point function, thus removing the need for the usual assumption of a proportionality relation between the and points contributions in the flow equation of (see ). On the contrary, we will be able to deduce an expression for the -point function from the knowledge of the -point function, itself deduced from the flow equation of the -point function. The only relevant parameters at sufficiently large times being the local parameters, , whose flow equations are deduced from the derivative expansion. The derivation of these equations being technical, we provide it in Appendix B, summarizing them in the following statement:
Truncating around sixtic interactions (i.e. up to effects) in the deep UV, neglecting the momentum dependence of effective vertices in the computation of effective loops and for external momenta large enough, the flow equations for the local couplings , and are:
and . Furthermore, defining the minimal dimensionless vertex functions as ():
It is moreover interesting to note that for external momenta large enough, the knowledge of allows reconstructing the -point function. Once again, we put the proof in Appendix B, and summarize the result in a compact statement:
For momenta large enough (), the -point vertex can be suitably approximated as:
Flowing though the neural network space: the active RG
In this section, following the discussion of section 2.3.2, we consider the active RG, viewing the networks parameters as an UV cut-off rather than a running mass. First, we derive the corresponding flow equation. We show that -point functions exhibit a purely scaling behavior, and that the corresponding -functions reduce to the linear dimensional contributions. As discussed in section 2.3.2, this RG is formally the same as the one used in , up to the lack of explicit scaling for mass, explaining why our scaling dimensions are different. Second, we investigate the content of the flow equations that we obtained. As pointed out in the section 2.3.2, the major advantage of this approach is to avoid introducing a working precision, or any special structure regarding the nature of the data.
Let us consider a network , and define . Within this suggestive notation, the exact propagator looks like a UV regularized free propagator:
playing the role of a UV cut-off, which suppresses large momenta. The question is therefore: what happen if we smoothly change the parameter ?In fact, this situation is familiar in string field theory, where is called the stub parameter . Formally, this is equivalent to moving the UV cut-off, which can be translated as a chain of equivalence relations between classical actions through the differential equation (2.57), all of them having the same long distance physics. One can think for the evolution equation of the classical action to something like equation (2.57), i.e.
Such an equation however assumes that is the free propagator. Rather, in our construction, it has to be understood as the effective propagator, taking into account fluctuations. Therefore, we have to construct a coarse graining with a fixed shape of the effective propagator, the corresponding free propagator remaining unknown. There is a pragmatic way to do this. We formally introduce a regulator in the classical action. This leads to the Wetterich equation (3.3), but with the additional constraint that:
This equation simply means that we relate the running scale to the standard deviation of the neural network weights as
and that we keep the shape of the -point function fixed along the RG trajectory (if it exists), fixing as soon as is given. Let us show how this condition allows closing the hierarchical flow equations. Let us consider the flow equation (4.5) for . Neglecting the momentum dependence of the effective vertex with respect to the momenta "" running through the effective loop following the discussion of Section 3.3, we get:
Note that this approach implicitly assumes is small enough to justify the replacement: in (4.5). Hence, the resulting flow equations are expected to be exact for reference scales in the IR.
Because the LHS can be explicitly computed from (4.3), we therefore obtain:
In the same way, the flow equation for allows in principle to compute within the same approximation. Let us illustrate how this works. Let us consider a given network defining the “fundamental scale” . We can measure the -point function at zero momentum . This condition in turn fixes the value of . For instance, let us consider the following explicit example, working with the slightly modified Litim’s regulator:
Straightforwardly, we have and , and the previous equality reads as:
where we introduced the dimensionless variable , and
Introducing the dimensionless coupling , and solving on , we thus obtain:
Because , and is a pure number, is close to for large . However, increases as decreases, and for the approximation breaks down. Note that we could expect this not to be a limitation of the approach in itself, but a limitation of the Litim regulator, however, a moment of reflection shows that such a singular behavior is in fact very general, and independent on the choice of the regulator. Under the condition (4.10), the problem (4.6) is well posed but trivial: it reduces to a pure scaling behavior. Indeed, given (4.6), we have:
The flow is entirely fixed by dimensional analysis, and the flow equation for reduces to its linear contribution:
In turn, this equation determines . The effective loop behaves like , times a factor which is independent. Hence, we deduce:
meaning that follows a purely scaling behavior as well.
We can use (4.4) to write the flow equations in terms of the standard deviation :
where now and are seen as functions of . As displayed in Figures 8 and 9, the numerical simulations match to a good precision with the solution to this equation (see Appendix A for the computations of ).
Finally, let us make a remark in regard to the results obtained in the reference . Indeed, the authors arrived to the equations (4.14) with replaced by an IR cut-off from perturbation theory, whose validity assumes to be large enough. What is puzzling with this calculation is that the RG predictions work even for small , where we expect that perturbation theory breaks down. Our derivation solves this paradox: using a non-perturbative framework, we are able to show that coupling constants follow scaling laws (4.11), without assumption on the sizes of the coupling constants (however, note that both derivations have been performed with different activation functions, such that it would be interesting to check how general (4.14) is).
Conclusion and outlooks
In this paper, we have pushed further the use of the renormalization group for the NN-QFT correspondence , which states that a neural network can be represented by a quantum field theory. In the infinite limit of the hidden layer width , the neural network is described by a Gaussian process and mapped to a free field theory, and interactions translate finite- corrections. The main difference with usual QFTs used in physics stems from the choice of the kernel (or propagator), itself inherited from the choice of activation function in the neural network. Since it encodes important properties on the theory (IR and UV divergences, scaling, etc.), it is important to ask how the data-space of the neural network inputs differ from usual spacetime. As a consequence, the usual assumptions on interaction locality may not be appropriate. In the first part of this paper, we have discussed several of these aspects, providing an interpretation slightly different from the one in .However, let us stress that this divergence in interpretation does not change in any way the computations and numerical results in both sides.
We have then described how to build a non-perturbative renormalization group flow following Wetterich-Morris formalism. We introduce two different points of view: in the passive case, the UV cut-off is related to the data resolution, while in the active case, it is given in terms of the standard deviation of the NN weights. The main difference with is that they postulate a global scale invariance with respect to a large volume cut-off (IR, in the language of our paper) on the data-space. Intriguingly, their results agree strongly with numerical simulations, for the small widths, whereas perturbation theory is expected to fail, and even if a scale invariance with respect to the volume is not expected. In this paper, we solve this paradox by developing a renormalization group based on an explicit coarse-graining, and derive flow equations from a process of partial integration of the field degrees of freedom. We find that the active point of view is formally identified with the flow in , thus justifying, in an explicitly non-perturbative framework, the agreement between theory and experience found in the paper. A natural extension of this work is to include non-local interactions using tensor models. Another possible direction is to generalize the derivation to other networks such as ReLU-net .
On the numerical side, the main result of our paper is the flow equation (4.14) which shows that the weight standard deviation can be interpreted as a running cut-off in terms of which the couplings of the NN-QFT change. This means that given the couplings for a specific value of , it is possible to compute analytically the couplings for any other value of without doing any numerical simulation. We have verified this statement using numerical simulations (Figures 8 and 9). In this paper, we have focused the analysis on the analytical computations in the QFT side: we plan to analyze the equations numerically in future works.
From a function-space perspective, it is natural to understand the learning process as a RG flow induced in a suitable theory space. It would very interesting to investigate how the notions presented in this paper could generalize to describe this process and how the couplings change under learning.
Acknowledgments
We are very grateful for useful discussions with James Halverson, Anindita Maiti, and Keegan Stoner. We would like to thank particularly Anindita Maiti for help in reproducing the numerical results from . This project has received funding from the European Union’s Horizon 2020 research and innovation program under the Marie Skłodowska-Curie grant agreement No 891169. This work is also supported by the National Science Foundation under Cooperative Agreement PHY-2019786 (The NSF AI Institute for Artificial Intelligence and Fundamental Interactions, http://iaifi.org/).
Appendix A Numerical simulations
In this appendix, we explain how to compute numerically correlation functions for neural networks defined in Section 2.1.1 and how to extract relevant information. In particular, we reproduce the numerical results from and provide additional details. The code is written in Python and is available at https://github.com/melsophos/nnqft. Throughout this appendix, we take:
In order to evaluate correlation functions, we consider neural networks . For each of them, the weights and biases are drawn independently from the distributions and , where is the width of the hidden layer. We will take :
Then, the experimental -point correlation functions are computed as (2.3):
We define the difference with the large Green functions as (see Section 2.1.4):
and the normalized -point functions as:
Note that no absolute value has been taken until now, and the result can be positive or negative. The large Green functions are computed with Wick theorem from the Gauss-net kernel (2.11). For example, the -point function is given by:
In order to reduce variance of the results, we will compute the Green functions by averaging over , each made of networks:
where means that the correlation functions is computed with the bag . This also allows extracting standard deviations if needed.
We will compute the correlation functions for the following points :
Given a -point correlation function, we compute it for all the possible combinations of points with , including identical entries. Since the experimental Green functions are symmetric by construction (as are the QFT Green functions), we consider only combinations which are inequivalent up to permutations. For example, we will compute the following -point functions:
For , there are respectively inequivalent combinations. We denote by the average of a quantity over all possible combinations of points, and by the average of the absolute value.We use the same notation as the average over the networks. However, there is no ambiguity since the latter is used in this section only to compute Green functions and never appears after.
The numerical Green functions are exact Green functions because they contain already all quantum corrections from loop diagrams. Hence, it is more natural to write a 1PI effective field theory and determine the coefficients by matching the Green functions computed from 1PI Feynman diagrams. Moreover, the Wetterich formalism from Sections 3 and 4 gives relations for the 1PI couplings. We consider the following 1PI interactions to describe the neural network:
where is the large free action (2.13). We consider a local Lagrangian because it turns out that it reproduces well the experimental Green functions for the points considered previously . In the notations of , we have and . However, the interpretation is slightly different compared to which writes a microscopic action. The interactions are associated with the part of the kinetic operator corresponding to the weight only, since the bias part is always Gaussian and independent of . Hence, propagators attached to vertices are instead of : the latter appear only in the disconnected -point propagators.
We now turn our attention to the computation of the experimental Green functions. We take:
Since we know the exact -point function , we must have:
Similarly, we know from (2.27) that higher-order Green functions must decrease as increases:
We check that it is indeed the case by plotting the values of , and for the different combinations of points (A.8). The figures 10 (not present in ) show that these values go toward as increases for . On the other hand, the values for don’t have any specific pattern, which is expected, since should be independent of .
We can simplify further this information and extract a single number. To do this, we take the absolute value of the normalized deviations (A.5) and average over the different combinations of points (A.8) to get . Moreover, to get an idea of how small are the normalized deviations, we define a background as follows: we compute the standard deviation of over all bags of neural networks for all combinations of points (A.8), and then average over the latter. The idea is to compare the normalized error encoded by with its numerical fluctuations over different bags, represented by the standard variation. On the figures 11, we reproduce the results from : for is below the background only for small and for , and it is always below the background for . In principle, should always be below the background for higher (which was not studied in the original paper and which we could not reach for computational reasons) so the current test is not very sharp; Figures 10 give a cleaner assessment.
Next, we can compute . Using Feynman rules, it can be obtained by subtracting the disconnected contributions (equal to and built from the 1PI -point function) from the full -point function to extract the contact interaction
where was defined in (2.11) (see for more details). Importantly, this equation is really an equality and not an approximation as in : since we are working with 1PI diagrams, there are no quantum corrections and any -point Green function is built from vertices of order . Higher-order vertices appear only in loop diagrams, which are not present. Hence, this allows determining all 1PI couplings exactly in a recursive way. Our results agree quantitatively for with those of because the loop corrections are subleading in the large expansion. However, this may give different results for since the latter receive loop corrections from the microscopic quartic vertex.
We find that is constant to a very good precision when evaluated over all combinations of points (A.8). In Figure 12, we display the values of averaged over all combinations and the corresponding standard deviation and find that its absolute value decreases as increases, reproducing the results . Importantly, we find that is negative, which was not indicated in (their figure 4 has an implicit absolute value needed to use the log-scale). As a consequence, the effective action (A.10) must include a sixtic contribution, however small, for the path integral to be stable: truncating to quartic interactions as in leads to an exponential growth of the weight. A preliminary analysis of the passive flow equations (Section 3) indicate that they can be integrated over a large range of only if the initial conditions satisfy and , otherwise the flow diverges.
The final numerical test we perform in this paper is to compute as a function of and (Figures 8, 9 and 13). We consider the following values of :
We see that decreases as and increase and that the values are well predicted by the active RG flow equations (4.14). As such, knowing for a single at fixed allows computing it for any other .
Appendix B Proofs and technical discussions
The term involving the field strength in the truncation takes the form
where is assumed to depend on . Because we furthermore assume it to be independent of , we can define operationally as:
The flow equation for can be deduced from the Wetterich equation,
In the local potential approximation, the vertices are momentum-independent. Therefore, the contribution involving can be discarded, leading to:
where, according to local potential approximation, we evaluate the RHS over uniform configurations. The derivative is then easy to compute, leading to:
The expression of can be easily obtained by taking the third derivative of the effective potential with respect to :
Note that the renormalized vertex has to be defined as (the factor will be explained below):
We focus on small and positive along axis . The integral decomposes as , where:
Because , in the negative branch, , we have:
which is independent of . In the positive branch, in contrast, we get:
where the bounds in the integrals refer to the integral over coordinate , and , for . Note that we omitted the Heaviside functions. Taking the first derivative with respect to , we get:
Next, we take the second derivative and we set . The contribution coming from derivative of the interior of the integral vanishes, because the remaining integration over is empty. Thus, only the variation of the bound contributes; assuming , we get after a tedious calculation:
As for the strict LPA, we introduced dimensionless quantities (and then, explain the origin of the factor in front of (B.7)). Now, we have to take into account the wave function renormalization. Recovering in front of the kinetic action requires:
where in this expression refers to the effective mass. This relation implies . After some simplifications, we get:
The explicit expression for can be easily derived within the local potential approximation:
B.2 Discussion about Claim 1
The regulator is obviously positive definite as soon as . For , because , we must have , therefore:
Thus, low momenta fluctuations are frozen, decoupling from long distance physics, whereas high momenta modes are unchanged and integrated out. Note that the last condition makes an infrared regulator, which prevents infrared divergences along the flow. Finally, , meaning that the original model is formally recovered in the deep infrared limit. All these properties ensure that the boundaries interpolation conditions and holds using . Now, let us show that is optimal in the Litim’s sense . The Wetterich equation (3.3) can be singular if the effective propagator diverges, or equivalently, if its inverse vanishes. To avoid this difficulty, has to prevent the existence of zero-modes. In other words, the RG requires the existence of a “gap”, and Litim’s criterion for optimization is to maximize the gap. This can be done from the observation that only the field-independent part of the inverse -point function is relevant to discuss optimization, i.e. , in view to establish a (weakly) model-independent criterion. It is suitable to introduce . The optimal value for the gap, , is:
By construction, the regulator is expected to be efficient for , and we fix the normalization such that for some . For a large enough family of regulators, this condition imposes that . If reaches its absolute minimum for , the regulator cannot be optimal from definition. Thus we may have . Without loss of generality we may choose . For a regulator which attributes the same size to the IR fluctuations , this reduces to an equality; this is the case for the regulator (3.49). Indeed, in that case for , for : the previous argument holds, up to the normalization factor .
B.3 Proof of proposition 2
Because we focus on the symmetric phase, odd effective vertices have to vanish identically as . Moreover, the effective propagator has to be diagonal: . From our ansatz, is given by:
Taking the second derivative of the exact RG equation (3.3), we get for :
Because the windows of momenta allowed by the function is limited to the region by construction, the symmetric function can be expanded in powers of . At leading order, setting and taking into account the definition 1, the equation simplifies as:
The derivative of the regulator can be computed from the definition 1 as well. We get, for :From its definition, the regulator vanishes for , and it is also the case for its derivative.
In the RG transformation, a rescaling of the lattice is required after partial integration of degrees of freedom to ensure preservation of the IR physics. This is equivalent to assuming the existence of a proper rescaling of the coupling, turning the flow equation in an autonomous system. From the equation above, we see in particular that the mass has to be rescaled as:See also the discussion at the beginning of the Section 3.3.
and the rescaling for the couplings and follows:
respectively defining the local -point coupling and effective mass. Within these dimensionless couplings, equation (B.27) becomes:
where . Within this approximation, and introducing , the flow equation for takes the form:
the integral being restricted in the interior of the sphere . Thus, defining:
From this equation, we easily deduce that must be a function of . Setting on both sides, we get an algebraic closed equation for :
The solution exhibits a singularity, which has to be taken into account when solving the flow equation. Through the definition (B.24), the solution to this equation provides . Taking into account that, we find from the chain rule:
where . We thus obtain for :
These equations depend on local couplings and . The flow of is fixed by the flow equation (B.34), but requires the knowledge of . It can be obtained using standard LPA, equations (3.18) and (3.19) for a sixtic truncation, which discard contributions or order from the -scaling. We get, for and :
Finally, from the flow equation for (equation (3.18)), setting and , we get:
For reader familiar with QFT, the origin of the function can be traced from the -, - and -channels. In fact, the contributions have the following structure (the intermediate fat dotted edge materializing the effective loop between effective vertices, see equation (3.18)):
The first term corresponds to , the second defines . A direct inspection shows that , which becomes small for large enough. We thus obtain the approximation for for large enough:
B.4 Discussion about Claim 2
The -point function has to be symmetric under any permutation of the four external momenta , , and . Moreover, because we assume local interactions as building blocks, external momenta have to be conserved: . Let us assume that is the analytic continuation with respect to some couplings from a perturbative solution , defined as the formal sum of an asymptotic perturbative series:
where the first sum runs over one-particle irreducible (1PI) Feynman diagrams having four external points. The product runs over vertices , denoting the number of fields involved in the interaction having coupling constant . Finally, is the Feynman amplitude associated with the graph . Note that all Feynman amplitudes arise with a global Dirac delta ensuring momentum conservation. We recall that Feynman diagrams provide a graphical representations of the Wick contractions involved in the perturbative expansion around the Gaussian theory. A typical Feynman graph is a set of vertices and edges, vertices corresponding to interactions and edges to the Wick contractions between pairs of fields. The momentum dependence of the -point function can be investigated from the structure of Feynman graphs labeling its perturbative expansion. First, we assume that the the theory involves only -points vertices. At one loop, has the following structure:
each term corresponding to the allowed permutations of the external momentaThe so called -, - and -channels.. Explicitly, the relevant one-loop diagrams have the following structure:
the loop being proportional to . It can be easily checked that the decomposition (B.45) keeps the same form after including sixtic interactions. Our aim is to prove that such a decomposition remains a suitable approximation for beyond one-loop. To this end, we will make use of a renormalization group argument, using the explicit expression (B.37). Assuming that (B.45) holds beyond one loop, and there exists a function such that:
Combining it with the relation (B.37), we get:
and , thus:
The flow equation for reads graphically as:
the cyclic permutation covering the three pairings , and , the solid black edge materializing the effective propagator , whereas the dotted edge corresponds to . Let us investigate the structure of the contribution. From our assumption, and neglecting the dependence of the effective vertex on the momentum running through the effective loop, we get:
where . For ’s close to the running horizon , becomes small as the explicit expression (B.37) shows. Thus ,
and contributions ensure stability of the assumption about . To check consistency with , we have to consider the corresponding flow equation:
up to permutation of the external momenta. Assuming we are close to the quartic sector, we focus on the last contribution. Setting , we have two relevant configurations to investigate. The first one is for and hooked to the same vertex. In that case, the loop depends only on two momenta, hooked to another vertex, say in the following:
where we discarded the dependence of the effective vertices on the momentum running through the effective loop, and we assumed . Such a contribution in the first term on the RHS of the equation (B.50) does not break the ansatz for , the remaining momentum could be set to zero outside of the tadpole. The second configuration is for and hooked to different effective vertices. It is however easy to check that for large external momenta with respect to the IR cut-off , these contributions are suppressed. For instance, we have:
and for large enough with respect to , this contribution is less relevant than the first in the flow equation for . From the same argument, it is easy to check that the second kind of contributions in the flow of , involving and does not break the ansatz for in the range of momenta that we consider.