Symmetry-via-Duality: Invariant Neural Network Densities from Parameter-Space Correlators
Anindita Maiti, Keegan Stoner, James Halverson
Introduction
Many systems in Nature, mathematics, and deep learning are described by densities over functions. In physics, it is central in quantum field theory (QFT) via the Feynman path integral, whereas in deep learning it explicitly arises via a correspondence between infinite networks and Gaussian processes.
More broadly, the density associated to a network architecture is itself of foundational importance. Though only a small collection of networks is trained in practice, due to computational limitations, a priori there is no reason to prefer one randomly initialized network over another (of the same architecture). In that case, ideally one would control the flow of the initialization density to the trained density, compute the trained mean , and use it to make predictions. Remarkably, may be analytically computed for infinite networks trained via gradient flow or Bayesian inference .
In systems governed by densities over functions, observables are strongly constrained by symmetry, which is usually determined via experiments. Examples include the Standard Model of Particle Physics, which has gauge symmetry (possibly with a discrete quotient), as well as certain multi-dimensional Gaussian Processes. In the absence of good experimental data or an explicit form for the density, it seems difficult to deduce much about its symmetries.
We introduce a mechanism for determining the symmetries of a neural network density via duality, even for an unknown density. A physical system is said to exhibit a duality when it admits two different, but equally fundamental, descriptions, called duality frames. Hallmarks of duality include the utility of one frame in understanding a feature of the system that is difficult to understand in the other, as well as limits of the system where one description is more tractable than the other. In neural networks, one sharp duality is Parameter-Space / Function-Space duality: networks may be thought of as instantiations of a network architecture with fixed parameter densities, or alternatively as draws from a function space density. In GP limits where a discrete hyperparameter (e.g. the width), the number of parameters is infinite and the parameter description unwieldy, but the function space density is Gaussian and therefore tractable. Conversely, when , the function space density is generally non-perturbative due to large non-Gaussianities, yet the network has few parameters.
We demonstrate that symmetries of network densities may be determined via the invariance of correlation functions computed in parameter space. We call this mechanism symmetry-via-duality, and it is utilized to demonstrate numerous cases in which transformations of neural networks (or layers) at input or output leave the correlation functions invariant, implying the invariance of the functional density. It also has implications for learning, which we test experimentally. For a summary of our contributions and results, see Section (5).
Symmetries of neural networks are a major topic of study in recent years. Generalizing beyond mere invariance of networks, equivariant networks have aided learning in a variety of contexts, including gauge-equivariant networks and their their utilization in generative models , for instance in applications to Lattice QCD . See also for symmetries and duality in ML and physics.
Of closest relation to our work is the appearance of symmetries in generative models, where invariance of a generative model density is often desired. It may be achieved via draws from a simple symmetric input density on and an equivariant network , which ensures that the induced output density is invariant. In Lattice QCD applications, this is used to ensure that gauge fields are sampled from the correct -invariant density , due to the -equivariance of a trained network .
In contrast, in our work it is the network itself that is sampled from an invariant density over functions. That is, if one were to cast our work into a lattice field theory context, it is the networks themselves that are the fields, and symmetry arises from symmetries in the density over networks. Notably, nowhere in our paper do we utilize equivariance.
Modeling Densities for Non-Gaussian Processes.
One motivation for understanding symmetries of network densities is that it constrains the modeling of neural network non-Gaussian process densities using techniques from QFT (see also ), as well as exact non-Gaussian network priors on individual inputs. Such finite- densities arise for network architectures admitting a GP limit as , and they admit a perturbative description when is large-but-finite. These functional symmetry considerations should also place constraints on NTK scaling laws and preactivation distribution flows studied in parameter space for large-but-finite networks.
Symmetry Invariant Densities via Duality
The -point correlation functions (or correlators) of neural network outputs are then
and (if the corresponding densities are known) they may be computed in either parameter- or function-space. These functions are the moments of the density over functions. When the output dimension , we may write output indices explicitly, e.g. , in which case the correlators are written .
Neural network symmetries are a focus of this work. Consider a continuous transformation
i.e. the transformed network at is a function of the old network at . We say there is a classical symmetry if is invariant under the transformation, which in physics is usually phrased in terms of the action . If the functional measure Df is also invariant, it is said that there is a quantum symmetry and the correlation functions are constrained by
See appendix for the elementary proof. In physics, if but is non-trivial the symmetry is called internal, and if is trivial but it is called a spacetime symmetry. Instead, we will call them output and input symmetries to identify the part of the neural network that is transformed; examples include rotations of outputs and translations of inputs. Of course, if is composed with other functions to form a larger neural network, then input and output refer to those of the layer .
Our goal in this paper is to determine symmetries of network densities via the constraint 3. For a discussion of functional densities, see Appendix E.
Further study in this direction is motivated by first deriving a result for one of the simplest function-space densities: a Gaussian process.
If the NNGP has zero mean, then for all and the even-point functions may be computed in terms of the kernel via Wick’s theorem,
where the Wick contractions are defined by the set
We write as . A network transformation induces an action on each index of each Kronecker delta in (4). For instance, as where the last equality holds for . By this phenomenon, the even-point correlation functions (4) are invariant. Conversely, if the NNGP has a mean , it transforms with a single and is not invariant. From the NNGP correlation functions, the GP density has symmetry iff it has zero mean. This is not surprising, and could be shown directly by inverting the kernel to get the GP density and then checking its symmetry.
However, we see correlation functions contain symmetry information, which becomes particularly powerful when the correlation functions are known, but the network density is not.
Parameter-Space / Function-Space Duality.
To determine the symmetries of an unknown network density via correlation functions, we need a way to compute them. For this, we utilize duality.
A physical system is said to exhibit a duality if there are two different descriptions of the system, often with different degrees of freedom, that exhibit the same predictions either exactly (an exact duality) or in a limit, e.g., at long distances (an infrared duality); see for a review. Duality is useful precisely when one perspective, a.k.a. a duality frame, allows you to determine something about the system that would be difficult from the other perspective. Examples in physics include electric-magnetic duality, which in some cases allows a strongly interacting theory of electrons to be reinterpreted in terms of a weakly coupled theory of monopoles , and gauge-gravity duality , which relates gravitational and non-gravitational quantum theories via the holographic principle.
In the context of neural networks, the relevant duality frames are provided by parameter-space and function-space, yielding a Parameter-Space / Function-Space duality. In the parameter frame, a neural network is considered to be compositions of functions which themselves have parameters drawn from , whereas in the function frame, the neural network is considered as an entire function drawn from a function-space density . Of course, the choice of network architecture and densities determine , but they do not appear explicitly in it, giving two different descriptions of the system.
Symmetry-via-Duality.
Our central point is that symmetries of function-space densities may be determined from correlation functions computed in the parameter-space description, even if the function space density is not known. That is, it is possible to check (3) via correlators computed in parameter space; if so, then the product is invariant. Barring an appearance of the Green-Schwarz mechanism in neural networks, by which is invariant but Df and are not, this implies that is invariant. This leads to our main result.
with associated function space measure Df and density , as well as a transformation satisfying
Then is invariant, and is itself invariant if a Green-Schwarz mechanism is not effective.
The proof of the theorem follows from the proof of (3) in the Appendix and the fact that correlators may also be computed in parameter space. Additionally, there may be multiple such transformations that generate a group of invariant transformations, in which case is -invariant.
The schematic for each calculation is to transform the correlators by transforming some part of the network, such as the input or output, absorb the transformation into a transformation of parameters (which could be all ), and then show invariance of the correlation functions via invariance of . Thus,
Symmetries of derived via duality rely on symmetry properties of .
In what follows we will show that (7) holds in numerous well-studied neural networks for a variety of transformations, without requiring equivariance of the neural network. Throughout, we use to denote the parameter space partition function of all parameters of the network.
Example: SO(D)𝑆𝑂𝐷SO(D) Output Symmetry.
We now demonstrate in detail that a linear output layer leads to invariant network densities provided that its weight and bias distributions are invariant. The network is defined by where and is an -dimension postactivation with parameters . Consider an invertible matrix transformation acting as as ; we use Einstein summation convention here and throughout. The transformed correlation functions are
The result holds more generally, for any invariant and , which as we will discuss in Section 3 includes the case of correlated parameters, as is relevant for learning.
Example: SO(d)𝑆𝑂𝑑SO(d) Input Symmetry.
We now demonstrate an example of neural networks with density invariant under input rotations, provided that the input layer parameters are drawn from an invariant distribution.
We will take a linear input layer and turn off the bias for simplicity, since it may be trivially included as in the output symmetry above. The network function is , , and the input rotation acts as . The output of the input layer is the preactivation for the rest of the network , which has parameters . The transformed correlators are
SU(D)𝑆𝑈𝐷SU(D) Output Symmetry
We also demonstrate that a linear complex-valued output layer, given in details in Appendix (A.2), leads to invariant networks densities provided that last linear layer weight and bias distributions are invariant. For clarity we leave off the bias term; it may be added trivially similar to Eqn. (2). This network is defined by , and transforms as under an invertible matrix transformation by group element . A necessary condition for symmetry is that the only non-zero correlation functions have an equal number of and ’s, as in (32), which transform as
Example: Translation Input Symmetry and T-layers.
which follows by absorbing the shift via , using , and renaming variables. See appendix (A.4) for details. Thus, the density is translation invariant.
The -layer is compatible with training since it is differentiable everywhere except when the layer input is , i.e. an integer. When doing gradient descent, the mod operation is treated as the identity, which gives the correct gradient at all non-integer inputs to the layer, and thus training performs well on typical real-world datasets.
Symmetry-via-Duality and Deep Learning
We now various aspects of relating symmetry, deduced via duality, and learning.
It may sometimes be useful to preserve a symmetry (deduced via duality) during training that is present in the network density at initialization. Corollary 1.1 allows symmetry-via-duality to be utilized at any time, relying on the invariance properties of (and therefore ). The initialization symmetry is preserved if the invariance properties of that ensured symmetry at persist at all times.
While this interesting matter deserves a systematic study of its own, here we study it in the simple case of continuous time gradient descent. The parameter update , where is the loss function, induces a flow in governed by
the update equation for . If is invariant at initialization (), then the update is invariant provided that is invariant and the second term is invariant.
When these conditions are satisfied, the symmetry of the network initialization density is preserved throughout training. However, they must be checked on a case-by-case basis. As a simple example, consider again the output symmetry from Section 2. Absorbing the action on output into parameters as before, the is itself invariant, and therefore the first term in (12) is invariant when is invariant. If additionally
for -dependent invariants and , then the second term is invariant as well, yielding an invariant symmetry-preserving update. See Appendix (A.6) for a detailed example realizing these conditions.
Supervised Learning, Symmetry Breaking, and the One-point Function.
To develop these ideas and prepare for experiments, we frame the discussion in terms of an architecture with symmetry properties that are easily determined via duality: a network with a linear no-bias output layer from an -dimensional hidden layer to -dimensional network output. The network function is , , where is the post-activation of the last hidden layer, with parameters . Output weights in the final () layer are initialized as
where is a hyperparameter that will be varied in the experiments. By a simple extension of our output symmetry result, this network density has symmetry. The symmetry breaking is measured by the one-point function
and zero otherwise. Here, means as a function; for some , may evaluate to . This non-zero mean breaks symmetry in the first network components, and since symmetry is restored in the limit, and together control the amount of symmetry breaking.
Symmetry and Correlated Parameters.
Learning-induced flows in neural network densities lead to correlations between parameters, breaking any independence that might exist in the parameter priors. Since symmetry-via-duality relies only on an invariant , it can apply in the case of correlations, when does not factorize. In fact, this is the generic case: a density which is symmetric due to being constructed from group invariants is not, in general, factorizable. For instance, the multivariate Gaussian with is constructed from the invariant and leads to independent parameters due to factorization. However, provided a density normalization condition is satisfied, additional -invariant terms (or any other Casimir invariant) may be added to the exponent which preserve symmetry, but break independence for .
By construction these distributions are invariant under , and therefore the function-space density is as well. However, independence is broken for . Training could also mix in parameters from other layers, yielding a non-trivial joint distribution which is nevertheless invariant provided that the final layer parameter-dependence arises only through invariants.
Such independence-breaking networks provide another perspective on neural networks and GPs. Since the NNGP correspondence relies crucially on the central limit theorem, and therefore independence of an infinite number of parameters as , we may break the GP to a non-Gaussian process not only by taking finite-, but also by breaking independence, as with above. In this example, symmetry-via-duality requires neither the asymptotic limit nor the independence limit.
Independence breaking introduces potentially interesting non-Gaussianities into function-space densities, which likely admit an effective field theory (EFT) akin to the finite- EFT treatment developed in . We leave this treatment for future work.
Symmetry and the Neural Tangent Kernel.
Gradient descent training of a neural network is governed by the Neural Tangent Kernel (NTK) , . Since depends on concrete parameters associated to a fixed neural network draw, it is not invariant. However, the NTK converges in appropriate large- limits to a kernel that is deterministic, due to the appearance of ensemble averages, allowing for the study of symmetries of via duality.
This depends on the concrete draw and is not invariant. However, as , at initialization becomes the deterministic NTK
Such results are more general, arising similarly in other architectures according to Corollary 1.1. Though the deterministic NTK at is crucial in the linearized regime, invariance of may also hold during training, as may be studied in examples via invariance of . See for input symmetry of the deterministic NTK .
Experiments
We carry out two classes of experimentsWe provide an implementation of our code at https://github.com/keeganstoner/nn-symmetry. testing symmetry-via-duality. In the first, we demonstrate that the amount of symmetry in the network density at initialization affects training accuracy. In the second, presented in Appendix C, we demonstrate how symmetry may be tested via numerically computed correlators.
We now wish to test the ideas from Section 3 on how the amount of symmetry at initialization affects test accuracy after training, as controlled by the hyperparameters and ; see for another analysis of symmetry breaking and learning. We further specify the networks discussed there by choosing a single-layer network (, ) with ReLU non-linearities (i.e., is the post-activation of a linear layer), , weights of the first linear layer , and weights of the output initialized as in (14). All networks were trained on the Fashion-MNIST dataset for epochs with MSE loss, with one-hot (or one-cold) encoded class labels, leading to network outputs with dimension , and therefore symmetry at initialization.
In the first experiment, we study how the amount of rotational symmetry breaking at initialization affects test accuracy on Fashion-MNIST with one-hot encoded class labels. We vary the amount of symmetry breaking by taking and with increment. Each experiment is repeated times with learning rate . For each pair, the mean of the maximum test accuracy across all experiments is plotted in Figure 1 (LHS). We see that performance is highest for networks initialized with or , i.e. with an symmetric initialization density, and decreases significantly with increasing amounts of symmetry breaking (increasing and ), contrary to the intuition discussed in Section 3. See Appendix (D) for more details about the experiments.
If supervised learning breaks symmetry via developing a non-trivial one-point function (mean), but we see that symmetry at initialization helps training in this experiment, then what concept is missing?
It is that symmetry breaking at initialization could be in the wrong direction, i.e. the initialization mean is quite different from the desired trained mean, which (if well-trained) approximate ground truth labels. In our experiment, the initialization mean is
which will in general be non-zero along all output components. It is "in the wrong direction" since class labels are one-hot encoded and therefore have precisely one non-zero entry. Furthermore, even for means in the right direction, the magnitude could be significantly off.
We see from Figure 1 (right) that performance improves until , but then monotonically decreases for larger . By construction, the symmetry breaking is much closer to the correct direction than in the first experiment, but the magnitude of the initialization means affects performance: the closer they are to the ones in the one-cold encoding, the better the performance. The latter occurs for according to our calculation, which matches the experimental result.
Conclusion
We introduce symmetry-via-duality, a mechanism that allows for the determination of symmetries of neural network functional densities , even when the density is unknown. The mechanism relies crucially two facts: i) that symmetries of a statistical system may also be determined via their correlation functions; and ii) that the correlators may be computed in parameter space. The utility of parameter space in determining symmetries of the network density is a hallmark of duality in physical systems, in this case, Parameter-Space / Function-Space duality.
We demonstrated that invariance of correlation functions ensures the invariance of , which yields the invariance of the density itself in the absence of a Green-Schwarz mechanism. Symmetries were categorized into input and output symmetries, the analogs of spatial and internal symmetries in physics, and a number of examples of symmetries were presented, including and symmetries at both input and output. In all calculations, the symmetry transformation induces a transformation on the input or output that may be absorbed into a transformation of network parameters , and invariance of the correlation functions follows from invariance of . The invariance of also follows, since are by definition the only parameters that transform.
The mechanism may also be applied at any point during training, since it relies on the invariance of . If duality is used to ensure the symmetry of the network density at initialization, then the persistence of this symmetry during training requires that remains symmetric at all times. Under continuous time gradient descent, the flow equation for yields conditions preserving the symmetry of . We also demonstrated that symmetry could be partially broken in the initialization density, that symmetry-via-duality may also apply in the case of non-independent parameters, and that the Neural Tangent Kernel may be invariant under symmetry transformations.
Our analysis allows for different amounts of symmetry in the network density at initialization, leading to increasing constraints on the density with increasing symmetry. Accordingly, it is natural to ask whether this affects training. To this end, we performed Fashion-MNIST experiments with different amounts of network density symmetry at initialization. The experiments demonstrate that symmetry breaking helps training when the associated mean is in the direction of the class labels, and entries are of the same order of magnitude. However, if symmetry is broken in the wrong direction or with too large a magnitude, performance is worse than for networks with symmetric initialization density.
Acknowledgements
We thank Sergei Gukov, Joonho Kim, Neil Lawrence, Magnus Rattray, and Matt Schwartz for discussions. We are especially indebted to Sébastian Racanière, Danilo Rezende, and Fabian Ruehle for comments on the manuscript. J.H. is supported by NSF CAREER grant PHY-1848089. This work is supported by the National Science Foundation under Cooperative Agreement PHY-2019786 (The NSF AI Institute for Artificial Intelligence and Fundamental Interactions).
References
Appendix A Proofs and derivations
We begin by demonstrating the invariance of network correlation functions under transformations that leave the functional measure and density invariant. Consider a transformation
that leaves the functional density invariant, i.e.
where the second to last equality holds due to the invariance of the functional density. This completes the proof of (3); See, e.g., for a QFT analogy.
For completeness we wish to derive the same result for infinitesimal output transformations, where the parameters of the transformation depend on the neural network input; in physics language, these are called infinitesimal gauge transformations.
The NN output, transformed by an infinitesimal parameter , is , where ; is the generator of the transformation group. Corresponding output space log-likelihood transforms as , for a current that may be computed. The transformed -pt function at is given by,
where we obtain second and last equalities under the assumption that functional density is invariant, following (20), and invariance of function-space measure, , respectively.
Following the -independence of L.H.S. of (A.1), terms on R.H.S. must cancel each other, i.e.
for any infinitesimal function . Thus, the coefficient of in above integrand vanishes at all , and we have the following by divergence theorem
a statement of invariance of correlation functions under infinitesimal input-dependent transformations.
Thus, we obtain the following invariance under finite / infinitesimal, input-dependent/independent transformations, whenever ,
(29) is same as (3), completing the proof.
A.2 SU(D)𝑆𝑈𝐷SU(D) Output Symmetry
We show the detailed construction of invariant network densities, for networks with a complex linear output layer, when weight and bias distributions are invariant.
The network is defined by , for a final affine transformation on last postactivation ; and are the inputs and parameters until the final linear layer, respectively. As is the rotation group over complex numbers, invariant NN densities require complex-valued outputs, and this requires complex weights and biases in layer . Denoting the real and imaginary parts of complex weight and bias in layer as respectively, we obtain distributions as and . The simplest invariant structure is
and similarly for bias. To obtain an -invariant structure in as a sum of -invariant structures from products of , all three PDFs need to be exponential functions, with equal coefficients in . Therefore, starting with -invariant real and imaginary parts and , one can obtain the simplest invariant complex weight and bias distributions, given by , respectively.
We want to express the network density and its correlation functions entirely in terms of complex-valued outputs, weights and biases, therefore, we need to transform the measures of into measures over . As , for Jacobian of , we obtain , and . With this, the -pt function for any number of ’s and ’s becomes the following,
where is the normalization factor. We emphasize that the transformation of and only transforms two indices inside the trace in ; it is invariant in this case, and also when all four indices transform.
From the structure of and , only those terms in the integrand of (A.2), that are functions of and alone, and not in product with any number of individually, result in a non-zero integral. Thus, we have the only non-vanishing correlation functions from an equal number of ’s and ’s. We hereby redefine the correlation functions of this complex-valued network as
where can be any permutation of set .
A.3 SU(d)𝑆𝑈𝑑SU(d) Input Symmetry
We now show an example of neural networks densities invariant under input transformations, provided that input layer parameters are drawn from an invariant distribution.
We will take a linear input layer and turn off bias for simplicity, as it may be trivially included as in output symmetry. group acts on complex numbers, therefore network inputs and input layer parameters need to be complex, such a network function is . The distribution of is obtained from products of distributions of its real and imaginary parts . Following output symmetry demonstration, the simplest invariant is obtained when are both invariant, we get . The measure of is obtained from the measures over as , with . Following a similar analysis as (A.2), the only non-trivial correlation functions are
Under input rotations by , the correlation functions transform into
A.4 Translation Input Symmetry
We demonstrate an example of network densities that remain invariant under continuous translations on input space, when the input layer weight is deterministic and input layer bias is sampled from a uniform distribution on the circle, . We will map the weight term to the circle by taking it mod 1, (i.e. % 1).
The network output transforms into under translations of inputs , where . With a deterministic , the network parameters are given by , and . The transformed -pt function is
A.5 Sp(D)𝑆𝑝𝐷Sp(D) Output Symmetry
We also demonstrate an example of network densities that remain invariant under the compact symplectic group transformations on output space.
The compact symplectic is the rotation group of quaternions, just as is the rotation group of complex numbers. Thus, a network with linear output layer would remain invariant under compact symplectic group, if last linear layer weights and biases are quaternionic numbers, drawn from invariant distributions. We define the network output as before, with parameters and , such that Hermitian norms and are compact symplectic invariant by definition, where the conjugate of a quarternion is . The distributions of are obtained as products of distributions of the components and respectively. Following the symmetry construction, we can obtain the simplest invariant and when these are functions of the Hermitian norm, and PDF of each component is an exponential function of invariant term with equal coefficient, similarly with bias. Starting with , and , we get invariant quaternionic parameter distributions and . We also obtain the measures over from measures over and , e.g. . Following an analysis similar to (A.2), it can be shown that the only non-trivial correlation functions of this quaternionic-valued network are
for any permutation over . Under transformation of outputs , by in quaternionic basis, the correlation functions transform as
A.6 Preserving Symmetry During Training: Examples
We study further the example of an output symmetry from Section 2. Turning off the bias for simplicity, the network function is
with parameters ; transformations of may be absorbed into , i.e. .
The network density remains symmetric during training when the updates to preserve symmetry; for this example, we showed in Section 3 that it occurs when is invariant and
where is invariant because is and does not transform. Furthermore,
which satisfies the second condition in (39), since the first index is the one that transforms when absorbs the transformation of .
Appendix B More General SO𝑆𝑂SO Invariant Network Distributions
We will now give an example of an invariant non-Gaussian network distribution at infinite width, as parameters of the initialized invariant GP become correlated through training. Such a network distribution can be obtained up to perturbative corrections to the initialized network distribution, if the extent of parameter correlation is small.
Training may correlate last layer weights of a linear network output initialized with , such that at a particular training step, we get independent from , with small . Correlation functions of this network distribution can be obtained by perturbative corrections to correlation functions of the network distribution at initialization. For example, the -pt function of the correlated network distribution is given by
Appendix C SO(D)𝑆𝑂𝐷SO(D) Invariance in Experiments
The correlator constraint (3) gives testable necessary conditions for a symmetric density. Consider a single-layer fully-connected network, called Gauss-net due to having a Gaussian GP kernel defined by , where , , and , with activation .
To test for invariance via (3), we measure the average elementwise change in -pt functions before and after an transformation. To do this we generate -pt and -pt correlators at various for a number of experiments and act on them with random group elements of a given group. Each group element is generated by exponentiating a random linear combination of generators of the corresponding algebra, namely
for , , and are generators of Lie algebra; i.e. skew-symmetric matrices written in a simple basis The generators are obtained by choosing each of the independent planes of rotations to have a canonical ordering with index , determined by a -plane. Each plane of rotation has a generator matrix with , , for , rest . For instance, at , there are independent planes of rotation formed by direction pairs , and . For each plane defined by directions , general elements have for and variable . Expanding each in Taylor series; the coefficients of terms are taken to define the generators of as in (44). . For example, for we take the standard basis for ,
We define the elementwise deviation to capture the change in correlators due to transformations. Here is the -transformed -pt function; both and have the same rank.
Error bounds for deviation are determined as using the standard error propagation formulae, equals the average (elementwise) standard deviation of -pt functions across 10 experiments, and is calculated as the following
We take the deviation tensors over experiments for and transformations of -pt and -pt functions at respectively, both correlators of each experiment are calculated using and network outputs respectively. An element-wise average and standard deviation across deviation tensors are taken and then averaged over, to produce the mean of the -transformation deviation , and its error , respectively. We plot (the blue shaded area) in Figure (2), this signal lies well within predicted error bounds of (in orange), although deviates significantly from at low widths, in contradiction to width independence of (3). This is due to the smaller sample size of parameters in low-width networks, and therefore small fluctuations in the weight and bias draws lead to more significant deviations from the "true" distribution of these parameters. A nonzero mean of the parameters caused by fluctuations leads to a nonzero mean of the function distribution , thus breaking symmetry. We believe this is a computational artefact and does not contradict -invariance in (3).
Appendix D Experiment Details
The experiments in Section (4) were done using Fashion-MNIST under the MIT LicenseThe MIT License (MIT) Copyright © 2017 Zalando SE, https://tech.zalando.com, using data points for each epoch split with a batch size of , and test batch size of . Each experiment was run on a 240GB computing node through the Discovery Cluster at Northeastern University and took between 1 and 1.5 hrs to train 20 epochs. The experiments were repeated 20 times for each configuration, and were run with configurations for the left plot in Fig. (1), and for the right plot. The error in the left plot of Fig. (1) is shown in Fig. (3).
The experiments in Appendix (C) were done on the same cluster, with the same memory nodes. These took around 24 hours on 10 compute nodes to generate models for each of the values. The -pt functions then took another 6 hours on a single node each.
Appendix E Comments on Functional Densities
Functional integrals are often treated loosely by physicists: they use them to great effect and experimental agreement in practice, but they are not rigorously defined in general; see, e.g., .
We follow in this tradition in this work, but would like to make some further comments regarding cases that are well-defined, casting the discussion first into the language of Euclidean QFT, and then bringing it back to machine learning.
First, the standard Feynman functional path integral for a scalar field is
but in many cases the action is split into free and interacting pieces
where the free action is Gaussian and the interacting action is non-Gaussian. The free theory, which has , is a Gaussian process and is therefore well-defined. When interactions are turned on, i.e. the non-Gaussianities in are small relative to some scale, physicists compute correlation functions (moments of the functional density) in perturbation theory, truncating the expansion at some order and writing approximate moments of the interacting theory density in terms of a sum of well-defined Gaussian moments, including higher moments.
In this work we study functional densities associated to neural networks, and have both the perturbative and especially the lattice understanding in mind when we consider them. In particular, for readers uncomfortable with the lack of precision in defining a functional density, we emphasize that our results can also be understood on a lattice, though input symmetries may be discrete subgroups of those existing in the continuum limit. Furthermore, any concrete ML application involves a finite set of inputs, and for any fixed application in physics or ML one can simply choose the spacing between lattice points to be smaller than the experimental resolution.