Symmetry-via-Duality: Invariant Neural Network Densities from Parameter-Space Correlators

Anindita Maiti, Keegan Stoner, James Halverson

Introduction

Many systems in Nature, mathematics, and deep learning are described by densities over functions. In physics, it is central in quantum field theory (QFT) via the Feynman path integral, whereas in deep learning it explicitly arises via a correspondence between infinite networks and Gaussian processes.

More broadly, the density associated to a network architecture is itself of foundational importance. Though only a small collection of networks is trained in practice, due to computational limitations, a priori there is no reason to prefer one randomly initialized network over another (of the same architecture). In that case, ideally one would control the flow of the initialization density to the trained density, compute the trained mean μ(x)\mu(x), and use it to make predictions. Remarkably, μ(x)\mu(x) may be analytically computed for infinite networks trained via gradient flow or Bayesian inference .

In systems governed by densities over functions, observables are strongly constrained by symmetry, which is usually determined via experiments. Examples include the Standard Model of Particle Physics, which has gauge symmetry SU(3)×SU(2)×U(1)SU(3)\times SU(2)\times U(1) (possibly with a discrete quotient), as well as certain multi-dimensional Gaussian Processes. In the absence of good experimental data or an explicit form for the density, it seems difficult to deduce much about its symmetries.

We introduce a mechanism for determining the symmetries of a neural network density via duality, even for an unknown density. A physical system is said to exhibit a duality when it admits two different, but equally fundamental, descriptions, called duality frames. Hallmarks of duality include the utility of one frame in understanding a feature of the system that is difficult to understand in the other, as well as limits of the system where one description is more tractable than the other. In neural networks, one sharp duality is Parameter-Space / Function-Space duality: networks may be thought of as instantiations of a network architecture with fixed parameter densities, or alternatively as draws from a function space density. In GP limits where a discrete hyperparameter N→∞N\to\infty (e.g. the width), the number of parameters is infinite and the parameter description unwieldy, but the function space density is Gaussian and therefore tractable. Conversely, when N=1N=1, the function space density is generally non-perturbative due to large non-Gaussianities, yet the network has few parameters.

We demonstrate that symmetries of network densities may be determined via the invariance of correlation functions computed in parameter space. We call this mechanism symmetry-via-duality, and it is utilized to demonstrate numerous cases in which transformations of neural networks (or layers) at input or output leave the correlation functions invariant, implying the invariance of the functional density. It also has implications for learning, which we test experimentally. For a summary of our contributions and results, see Section (5).

Symmetries of neural networks are a major topic of study in recent years. Generalizing beyond mere invariance of networks, equivariant networks have aided learning in a variety of contexts, including gauge-equivariant networks and their their utilization in generative models , for instance in applications to Lattice QCD . See also for symmetries and duality in ML and physics.

Of closest relation to our work is the appearance of symmetries in generative models, where invariance of a generative model density is often desired. It may be achieved via draws from a simple symmetric input density ρ\rho on VV and an equivariant network fθ:V→Vf_{\theta}:V\to V, which ensures that the induced output density ρfθ\rho_{f_{\theta}} is invariant. In Lattice QCD applications, this is used to ensure that gauge fields are sampled from the correct GG-invariant density ρfθ\rho_{f_{\theta}}, due to the GG-equivariance of a trained network fθf_{\theta}.

In contrast, in our work it is the network fθf_{\theta} itself that is sampled from an invariant density over functions. That is, if one were to cast our work into a lattice field theory context, it is the networks themselves that are the fields, and symmetry arises from symmetries in the density over networks. Notably, nowhere in our paper do we utilize equivariance.

Modeling Densities for Non-Gaussian Processes.

One motivation for understanding symmetries of network densities is that it constrains the modeling of neural network non-Gaussian process densities using techniques from QFT (see also ), as well as exact non-Gaussian network priors on individual inputs. Such finite-NN densities arise for network architectures admitting a GP limit as N→∞N\to\infty, and they admit a perturbative description when NN is large-but-finite. These functional symmetry considerations should also place constraints on NTK scaling laws and preactivation distribution flows studied in parameter space for large-but-finite NN networks.

Symmetry Invariant Densities via Duality

The nn-point correlation functions (or correlators) of neural network outputs are then

and (if the corresponding densities are known) they may be computed in either parameter- or function-space. These functions are the moments of the density over functions. When the output dimension D>1D>1, we may write output indices explicitly, e.g. fi(x)f_{i}(x), in which case the correlators are written Gi1,…,in(n)(x1,…,xn)G^{(n)}_{i_{1},\dots,i_{n}}(x_{1},\dots,x_{n}).

Neural network symmetries are a focus of this work. Consider a continuous transformation

i.e. the transformed network f′f^{\prime} at xx is a function Φ\Phi of the old network ff at x′x^{\prime}. We say there is a classical symmetry if PfP_{f} is invariant under the transformation, which in physics is usually phrased in terms of the action Sf=−log PfS_{f}=-{\rm log}\,P_{f}. If the functional measure Df is also invariant, it is said that there is a quantum symmetry and the correlation functions are constrained by

See appendix for the elementary proof. In physics, if x=x′x=x^{\prime} but Φ\Phi is non-trivial the symmetry is called internal, and if Φ\Phi is trivial but x≠x′x\neq x^{\prime} it is called a spacetime symmetry. Instead, we will call them output and input symmetries to identify the part of the neural network that is transformed; examples include rotations of outputs and translations of inputs. Of course, if fθf_{\theta} is composed with other functions to form a larger neural network, then input and output refer to those of the layer ff.

Our goal in this paper is to determine symmetries of network densities via the constraint 3. For a discussion of functional densities, see Appendix E.

Further study in this direction is motivated by first deriving a result for one of the simplest function-space densities: a Gaussian process.

If the NNGP has zero mean, then G(2n+1)(x1,…,xn)=0G^{(2n+1)}(x_{1},\dots,x_{n})=0 for all nn and the even-point functions may be computed in terms of the kernel via Wick’s theorem,

where the Wick contractions are defined by the set

We write P∈Wick(2n)P\in\text{Wick}(2n) as P={(a1,b1),…,(an,bn)}P=\{(a_{1},b_{1}),\dots,(a_{n},b_{n})\}. A network transformation fi↦Rijfjf_{i}\mapsto R_{ij}f_{j} induces an action on each index of each Kronecker delta in (4). For instance, as δik↦RijRklδjl=(R RT)ik=δik\delta_{ik}\mapsto R_{ij}R_{kl}\delta_{jl}=(R\,R^{T})_{ik}=\delta_{ik} where the last equality holds for R∈SO(D)R\in SO(D). By this phenomenon, the even-point correlation functions (4) are SO(D)SO(D) invariant. Conversely, if the NNGP has a mean μ(x)=Gi11(x1)≠0\mu(x)=G_{i_{1}}^{1}(x_{1})\neq 0, it transforms with a single RR and is not invariant. From the NNGP correlation functions, the GP density has SO(D)SO(D) symmetry iff it has zero mean. This is not surprising, and could be shown directly by inverting the kernel to get the GP density and then checking its symmetry.

However, we see correlation functions contain symmetry information, which becomes particularly powerful when the correlation functions are known, but the network density is not.

Parameter-Space / Function-Space Duality.

To determine the symmetries of an unknown network density via correlation functions, we need a way to compute them. For this, we utilize duality.

A physical system is said to exhibit a duality if there are two different descriptions of the system, often with different degrees of freedom, that exhibit the same predictions either exactly (an exact duality) or in a limit, e.g., at long distances (an infrared duality); see for a review. Duality is useful precisely when one perspective, a.k.a. a duality frame, allows you to determine something about the system that would be difficult from the other perspective. Examples in physics include electric-magnetic duality, which in some cases allows a strongly interacting theory of electrons to be reinterpreted in terms of a weakly coupled theory of monopoles , and gauge-gravity duality , which relates gravitational and non-gravitational quantum theories via the holographic principle.

In the context of neural networks, the relevant duality frames are provided by parameter-space and function-space, yielding a Parameter-Space / Function-Space duality. In the parameter frame, a neural network is considered to be compositions of functions which themselves have parameters drawn from PθP_{\theta}, whereas in the function frame, the neural network is considered as an entire function drawn from a function-space density PfP_{f}. Of course, the choice of network architecture and densities PθP_{\theta} determine PfP_{f}, but they do not appear explicitly in it, giving two different descriptions of the system.

Symmetry-via-Duality.

Our central point is that symmetries of function-space densities may be determined from correlation functions computed in the parameter-space description, even if the function space density is not known. That is, it is possible to check (3) via correlators computed in parameter space; if so, then the product Df Pf\textit{Df}\,P_{f} is invariant. Barring an appearance of the Green-Schwarz mechanism in neural networks, by which Df Pf\textit{Df}\,P_{f} is invariant but Df and PfP_{f} are not, this implies that PfP_{f} is invariant. This leads to our main result.

with associated function space measure Df and density PfP_{f}, as well as a transformation f′(x)=Φ(f(x′))f^{\prime}(x)=\Phi(f(x^{\prime})) satisfying

Then Df Pf\textit{Df}\,P_{f} is invariant, and PfP_{f} is itself invariant if a Green-Schwarz mechanism is not effective.

The proof of the theorem follows from the proof of (3) in the Appendix and the fact that correlators may also be computed in parameter space. Additionally, there may be multiple such transformations that generate a group GG of invariant transformations, in which case PfP_{f} is GG-invariant.

The schematic for each calculation is to transform the correlators by transforming some part of the network, such as the input or output, absorb the transformation into a transformation of parameters θT⊂θ\theta_{T}\subset\theta (which could be all θ\theta), and then show invariance of the correlation functions via invariance of PθTP_{\theta_{T}}. Thus,

Symmetries of PfP_{f} derived via duality rely on symmetry properties of PθTP_{\theta_{T}}.

In what follows we will show that (7) holds in numerous well-studied neural networks for a variety of transformations, without requiring equivariance of the neural network. Throughout, we use ZθZ_{\theta} to denote the parameter space partition function of all parameters θ\theta of the network.

Example: S​O​(D)𝑆𝑂𝐷SO(D) Output Symmetry.

We now demonstrate in detail that a linear output layer leads to SO(D)SO(D) invariant network densities provided that its weight and bias distributions are invariant. The network is defined by fi(x)=Wijgj(x)+bif_{i}(x)=W_{ij}g_{j}(x)+b_{i} where i=1,…,Di=1,\dots,D and gjg_{j} is an NN-dimension postactivation with parameters θg\theta_{g}. Consider an invertible matrix transformation RR acting as as fi↦Rijfjf_{i}\mapsto R_{ij}f_{j}; we use Einstein summation convention here and throughout. The transformed correlation functions are

The result holds more generally, for any invariant PWP_{W} and PbP_{b}, which as we will discuss in Section 3 includes the case of correlated parameters, as is relevant for learning.

Example: S​O​(d)𝑆𝑂𝑑SO(d) Input Symmetry.

We now demonstrate an example of neural networks with density invariant under SO(d)SO(d) input rotations, provided that the input layer parameters are drawn from an invariant distribution.

We will take a linear input layer and turn off the bias for simplicity, since it may be trivially included as in the SO(D)SO(D) output symmetry above. The network function is fi(x)=gij(Wjkxk)f_{i}(x)=g_{ij}(W_{jk}x_{k}), W∼PWW\sim P_{W}, and the input rotation R∈SO(d)R\in SO(d) acts as xi↦xi′=Rijxjx_{i}\mapsto x_{i}^{\prime}=R_{ij}x_{j}. The output of the input layer is the preactivation for the rest of the network gg, which has parameters θg\theta_{g}. The transformed correlators are

S​U​(D)𝑆𝑈𝐷SU(D) Output Symmetry

We also demonstrate that a linear complex-valued output layer, given in details in Appendix (A.2), leads to SU(D)SU(D) invariant networks densities provided that last linear layer weight and bias distributions are invariant. For clarity we leave off the bias term; it may be added trivially similar to Eqn. (2). This network is defined by fi=Wijgj(x,θg)\mathbf{f}_{i}=\mathbf{W}_{ij}g_{j}(x,\theta_{g}), and transforms as fi↦Sijfj, fk†↦fl†Slk†\mathbf{f}_{i}\mapsto S_{ij}\mathbf{f}_{j},~{}\mathbf{f}^{\dagger}_{k}\mapsto\mathbf{f}^{\dagger}_{l}S^{\dagger}_{lk} under an invertible matrix transformation by SU(D)SU(D) group element SS. A necessary condition for symmetry is that the only non-zero correlation functions have an equal number of ff and f†f^{\dagger}’s, as in (32), which transform as

Example: Translation Input Symmetry and T-layers.

which follows by absorbing the shift via b′=(Wc %  1)+bb^{\prime}=(Wc\,\%\,\,1)+b, using db′=dbdb^{\prime}=db, and renaming variables. See appendix (A.4) for details. Thus, the density PfP_{f} is translation invariant.

The TT-layer is compatible with training since it is differentiable everywhere except when the layer input is ≡0 mod 1\equiv 0\text{ mod }1, i.e. an integer. When doing gradient descent, the mod operation is treated as the identity, which gives the correct gradient at all non-integer inputs to the layer, and thus training performs well on typical real-world datasets.

Symmetry-via-Duality and Deep Learning

We now various aspects of relating symmetry, deduced via duality, and learning.

It may sometimes be useful to preserve a symmetry (deduced via duality) during training that is present in the network density at initialization. Corollary 1.1 allows symmetry-via-duality to be utilized at any time, relying on the invariance properties of PθTP_{\theta_{T}} (and therefore PθP_{\theta}). The initialization symmetry is preserved if the invariance properties of PθP_{\theta} that ensured symmetry at t=0t=0 persist at all times.

While this interesting matter deserves a systematic study of its own, here we study it in the simple case of continuous time gradient descent. The parameter update dθi/dt=−∂L/∂θid\theta_{i}/dt=-\partial\mathcal{L}/\partial{\theta_{i}}, where L\mathcal{L} is the loss function, induces a flow in PθP_{\theta} governed by

the update equation for Pθ(t)P_{\theta}(t). If Pθ(t)P_{\theta}(t) is invariant at initialization (t=0t=0), then the update is invariant provided that ∂2L/(∂θi)2\partial^{2}\mathcal{L}/(\partial\theta_{i})^{2} is invariant and the second term is invariant.

When these conditions are satisfied, the symmetry of the network initialization density is preserved throughout training. However, they must be checked on a case-by-case basis. As a simple example, consider again the SO(D)SO(D) output symmetry from Section 2. Absorbing the action on output into parameters as before, the (∂/∂θi)(∂/∂θi)(\partial/\partial{\theta_{i}})(\partial/\partial{\theta_{i}}) is itself invariant, and therefore the first term in (12) is invariant when L\mathcal{L} is invariant. If additionally

for θ\theta-dependent invariants IPI_{P} and ILI_{\mathcal{L}}, then the second term is invariant as well, yielding an invariant symmetry-preserving update. See Appendix (A.6) for a detailed example realizing these conditions.

Supervised Learning, Symmetry Breaking, and the One-point Function.

To develop these ideas and prepare for experiments, we frame the discussion in terms of an architecture with symmetry properties that are easily determined via duality: a network with a linear no-bias output layer from an NN-dimensional hidden layer to DD-dimensional network output. The network function is fi(x)=Wijlgj(x)\smash{f_{i}(x)=W^{l}_{ij}g_{j}(x)}, i=1,…,Di=1,\dots,D, where gj(x)g_{j}(x) is the post-activation of the last hidden layer, with parameters θg\theta_{g}. Output weights in the final (lthl^{th}) layer are initialized as

where kk is a hyperparameter that will be varied in the experiments. By a simple extension of our SO(D)SO(D) output symmetry result, this network density has SO(D−k)SO(D-k) symmetry. The symmetry breaking is measured by the one-point function

and zero otherwise. Here, ≠0\neq 0 means as a function; for some xx, Gi(1)(x)G_{i}^{(1)}(x) may evaluate to . This non-zero mean breaks symmetry in the first kk network components, and since SO(D)SO(D) symmetry is restored in the μWl→0\mu_{W^{l}}\to 0 limit, μWl\mu_{W^{l}} and kk together control the amount of symmetry breaking.

Symmetry and Correlated Parameters.

Learning-induced flows in neural network densities lead to correlations between parameters, breaking any independence that might exist in the parameter priors. Since symmetry-via-duality relies only on an invariant PθP_{\theta}, it can apply in the case of correlations, when PθP_{\theta} does not factorize. In fact, this is the generic case: a density PθP_{\theta} which is symmetric due to being constructed from group invariants is not, in general, factorizable. For instance, the multivariate Gaussian Pθ=exp(−xixi/2σ2)\smash{P_{\theta}=\text{exp}(-x_{i}x_{i}/2\sigma^{2})} with i=1,…,mi=1,\dots,m is constructed from the SO(m)SO(m) invariant xixix_{i}x_{i} and leads to independent parameters due to factorization. However, provided a density normalization condition is satisfied, additional SO(l)SO(l)-invariant terms cn (xixi)nc_{n}\,(x_{i}x_{i})^{n} (or any other Casimir invariant) may be added to the exponent which preserve symmetry, but break independence for n>1n>1.

By construction these distributions are invariant under SO(D)SO(D), and therefore the function-space density is as well. However, independence is broken for λ≠0\lambda\neq 0. Training could also mix in parameters from other layers, yielding a non-trivial joint distribution which is nevertheless invariant provided that the final layer parameter-dependence arises only through SO(D)SO(D) invariants.

Such independence-breaking networks provide another perspective on neural networks and GPs. Since the NNGP correspondence relies crucially on the central limit theorem, and therefore independence of an infinite number of parameters as N→∞N\to\infty, we may break the GP to a non-Gaussian process not only by taking finite-NN, but also by breaking independence, as with λ≠0\lambda\neq 0 above. In this example, symmetry-via-duality requires neither the asymptotic N→∞N\to\infty limit nor the independence limit.

Independence breaking introduces potentially interesting non-Gaussianities into function-space densities, which likely admit an effective field theory (EFT) akin to the finite-NN EFT treatment developed in . We leave this treatment for future work.

Symmetry and the Neural Tangent Kernel.

Gradient descent training of a neural network fθf_{\theta} is governed by the Neural Tangent Kernel (NTK) , Θ^(x,x′)=∂θifθ(x) ∂θifθ(x′)\smash{\hat{\Theta}(x,x^{\prime})=\partial_{\theta_{i}}f_{\theta}(x)\,\partial_{\theta_{i}}f_{\theta}(x^{\prime})}. Since Θ^(x,x′)\hat{\Theta}(x,x^{\prime}) depends on concrete parameters associated to a fixed neural network draw, it is not invariant. However, the NTK converges in appropriate large-NN limits to a kernel Θ\Theta that is deterministic, due to the appearance of ensemble averages, allowing for the study of symmetries of Θ(x,x′)\Theta(x,x^{\prime}) via duality.

This depends on the concrete draw ff and is not invariant. However, as N→∞N\to\infty, Θ^\hat{\Theta} at initialization becomes the deterministic NTK

Such results are more general, arising similarly in other architectures according to Corollary 1.1. Though the deterministic NTK at t=0t=0 is crucial in the linearized regime, invariance of Θ\Theta may also hold during training, as may be studied in examples via invariance of PθP_{\theta}. See for SO(D)SO(D) input symmetry of the deterministic NTK Θ\Theta.

Experiments

We carry out two classes of experimentsWe provide an implementation of our code at https://github.com/keeganstoner/nn-symmetry. testing symmetry-via-duality. In the first, we demonstrate that the amount of symmetry in the network density at initialization affects training accuracy. In the second, presented in Appendix C, we demonstrate how symmetry may be tested via numerically computed correlators.

We now wish to test the ideas from Section 3 on how the amount of symmetry at initialization affects test accuracy after training, as controlled by the hyperparameters μW1\mu_{W^{1}} and kk; see for another analysis of symmetry breaking and learning. We further specify the networks discussed there by choosing a single-layer network (l=1l=1, d=1d=1) with ReLU non-linearities (i.e., gjg_{j} is the post-activation of a linear layer), N=50N=50, weights of the first linear layer W0∼N(0,1/d)W^{0}\sim\mathcal{N}(0,1/\sqrt{d}), and weights of the output initialized as in (14). All networks were trained on the Fashion-MNIST dataset for 2020 epochs with MSE loss, with one-hot (or one-cold) encoded class labels, leading to network outputs with dimension D=10D=10, and therefore symmetry SO(10−k)SO(10-k) at initialization.

In the first experiment, we study how the amount of rotational symmetry breaking at initialization affects test accuracy on Fashion-MNIST with one-hot encoded class labels. We vary the amount of symmetry breaking by taking k∈{0,2,4,6,10}k\in\{0,2,4,6,10\} and μW1∈{0.0,…,0.2}\mu_{W^{1}}\in\{0.0,\dots,0.2\} with .01.01 increment. Each experiment is repeated 2020 times with learning rate η=.001\eta=.001. For each (k,μW1)(k,\mu_{W^{1}}) pair, the mean of the maximum test accuracy across all 2020 experiments is plotted in Figure 1 (LHS). We see that performance is highest for networks initialized with μW1=0\mu_{W^{1}}=0 or k=0k=0, i.e. with an SO(D)SO(D) symmetric initialization density, and decreases significantly with increasing amounts of symmetry breaking (increasing μW1\mu_{W^{1}} and kk), contrary to the intuition discussed in Section 3. See Appendix (D) for more details about the experiments.

If supervised learning breaks symmetry via developing a non-trivial one-point function (mean), but we see that symmetry at initialization helps training in this experiment, then what concept is missing?

It is that symmetry breaking at initialization could be in the wrong direction, i.e. the initialization mean is quite different from the desired trained mean, which (if well-trained) approximate ground truth labels. In our experiment, the initialization mean is

which will in general be non-zero along all output components. It is "in the wrong direction" since class labels are one-hot encoded and therefore have precisely one non-zero entry. Furthermore, even for means in the right direction, the magnitude could be significantly off.

We see from Figure 1 (right) that performance improves until μW1∼.02−.04\mu_{W^{1}}\sim.02-.04, but then monotonically decreases for larger μW1\mu_{W^{1}}. By construction, the symmetry breaking is much closer to the correct direction than in the first experiment, but the magnitude of the initialization means affects performance: the closer they are to the D−1D-1 ones in the one-cold encoding, the better the performance. The latter occurs for μW1∼.02−.04\mu_{W^{1}}\sim.02-.04 according to our calculation, which matches the experimental result.

Conclusion

We introduce symmetry-via-duality, a mechanism that allows for the determination of symmetries of neural network functional densities PfP_{f}, even when the density is unknown. The mechanism relies crucially two facts: i) that symmetries of a statistical system may also be determined via their correlation functions; and ii) that the correlators may be computed in parameter space. The utility of parameter space in determining symmetries of the network density is a hallmark of duality in physical systems, in this case, Parameter-Space / Function-Space duality.

We demonstrated that invariance of correlation functions ensures the invariance of Df Pf\textit{Df}\,P_{f}, which yields the invariance of the density PfP_{f} itself in the absence of a Green-Schwarz mechanism. Symmetries were categorized into input and output symmetries, the analogs of spatial and internal symmetries in physics, and a number of examples of symmetries were presented, including SO(D)SO(D) and SU(D)SU(D) symmetries at both input and output. In all calculations, the symmetry transformation induces a transformation on the input or output that may be absorbed into a transformation of network parameters θT\theta_{T}, and invariance of the correlation functions follows from invariance of DθTPθTD\theta_{T}P_{\theta_{T}}. The invariance of DθPθD\theta P_{\theta} also follows, since θT\theta_{T} are by definition the only parameters that transform.

The mechanism may also be applied at any point during training, since it relies on the invariance of PθP_{\theta}. If duality is used to ensure the symmetry of the network density at initialization, then the persistence of this symmetry during training requires that PθP_{\theta} remains symmetric at all times. Under continuous time gradient descent, the flow equation for PθP_{\theta} yields conditions preserving the symmetry of PθP_{\theta}. We also demonstrated that symmetry could be partially broken in the initialization density, that symmetry-via-duality may also apply in the case of non-independent parameters, and that the Neural Tangent Kernel may be invariant under symmetry transformations.

Our analysis allows for different amounts of symmetry in the network density at initialization, leading to increasing constraints on the density with increasing symmetry. Accordingly, it is natural to ask whether this affects training. To this end, we performed Fashion-MNIST experiments with different amounts of network density symmetry at initialization. The experiments demonstrate that symmetry breaking helps training when the associated mean is in the direction of the class labels, and entries are of the same order of magnitude. However, if symmetry is broken in the wrong direction or with too large a magnitude, performance is worse than for networks with symmetric initialization density.

Acknowledgements

We thank Sergei Gukov, Joonho Kim, Neil Lawrence, Magnus Rattray, and Matt Schwartz for discussions. We are especially indebted to Sébastian Racanière, Danilo Rezende, and Fabian Ruehle for comments on the manuscript. J.H. is supported by NSF CAREER grant PHY-1848089. This work is supported by the National Science Foundation under Cooperative Agreement PHY-2019786 (The NSF AI Institute for Artificial Intelligence and Fundamental Interactions).

References

Appendix A Proofs and derivations

We begin by demonstrating the invariance of network correlation functions under transformations that leave the functional measure and density invariant. Consider a transformation

that leaves the functional density invariant, i.e.

where the second to last equality holds due to the invariance of the functional density. This completes the proof of (3); See, e.g., for a QFT analogy.

For completeness we wish to derive the same result for infinitesimal output transformations, where the parameters of the transformation depend on the neural network input; in physics language, these are called infinitesimal gauge transformations.

The NN output, transformed by an infinitesimal parameter ωa(x)\omega_{a}(x), is f′(x)=Φ(f(x′))=f(x′)+δωf(x′)f^{\prime}(x)=\Phi(f(x^{\prime}))=f(x^{\prime})+\delta_{\omega}f(x^{\prime}), where δωf(x′)=−iωa(x′)Taf(x′)\delta_{\omega}f(x^{\prime})=-i\omega_{a}(x^{\prime})T_{a}f(x^{\prime}); TaT_{a} is the generator of the transformation group. Corresponding output space log-likelihood transforms as S↦S−∫ dx′ ∂μ jaμ(x′) ωa(x′)S\mapsto S-\int\,dx^{\prime}\,\partial_{\mu}\,j^{\mu}_{a}(x^{\prime})\,\omega_{a}(x^{\prime}), for a current jaμ(x′)j^{\mu}_{a}(x^{\prime}) that may be computed. The transformed nn-pt function at O(ω)O(\omega) is given by,

where we obtain second and last equalities under the assumption that functional density is invariant, following (20), and invariance of function-space measure, Df ′=Df \textit{Df}\,^{\prime}=\textit{Df}\,, respectively.

Following the ω\omega-independence of L.H.S. of (A.1), O(ω)O(\omega) terms on R.H.S. must cancel each other, i.e.

for any infinitesimal function ω(x′)\omega(x^{\prime}). Thus, the coefficient of ω(x′)\omega(x^{\prime}) in above integrand vanishes at all x′x^{\prime}, and we have the following by divergence theorem

a statement of invariance of correlation functions under infinitesimal input-dependent transformations.

Thus, we obtain the following invariance under finite / infinitesimal, input-dependent/independent transformations, whenever Df =Df ′\textit{Df}\,=\textit{Df}\,^{\prime},

(29) is same as (3), completing the proof.

A.2 S​U​(D)𝑆𝑈𝐷SU(D) Output Symmetry

We show the detailed construction of SU(D)SU(D) invariant network densities, for networks with a complex linear output layer, when weight and bias distributions are SU(D)SU(D) invariant.

The network is defined by f(x)=L(g(x,θg))f(x)=L(g(x,\theta_{g})), for a final affine transformation LL on last postactivation g(x,θ)g(x,\theta); xx and θg\theta_{g} are the inputs and parameters until the final linear layer, respectively. As SU(D)SU(D) is the rotation group over complex numbers, SU(D)SU(D) invariant NN densities require complex-valued outputs, and this requires complex weights and biases in layer LL. Denoting the real and imaginary parts of complex weight W\mathbf{W} and bias b\mathbf{b} in layer LL as W1,W2,b1,b2W^{1},W^{2},b^{1},b^{2} respectively, we obtain W,b\mathbf{W},\mathbf{b} distributions as PW,W†=PW1PW2P_{\mathbf{W},\mathbf{W}^{\dagger}}=P_{W^{1}}P_{W^{2}} and Pb,b†=Pb1Pb2P_{\mathbf{b},\mathbf{b}^{\dagger}}=P_{b^{1}}P_{b^{2}}. The simplest SU(D)SU(D) invariant structure is

and similarly for bias. To obtain an SUSU-invariant structure in PW,W†P_{\mathbf{W},\mathbf{W}^{\dagger}} as a sum of SOSO-invariant structures from products of PW1,PW2P_{W^{1}},P_{W^{2}}, all three PDFs need to be exponential functions, with equal coefficients in PW1,PW2P_{W^{1}},P_{W^{2}}. Therefore, starting with SO(D)SO(D)-invariant real and imaginary parts W1,W2∼N(0,σW2)W^{1},W^{2}\sim\mathcal{N}(0,\sigma^{2}_{W}) and b1,b2∼N(0,σb2)b^{1},b^{2}\sim\mathcal{N}(0,\sigma^{2}_{b}), one can obtain the simplest SU(D)SU(D) invariant complex weight and bias distributions, given by PW,W†=exp⁡(−Tr[W†W]/2σW2)P_{\mathbf{W},\mathbf{W}^{\dagger}}=\exp(-\text{Tr}[\mathbf{W}^{\dagger}\mathbf{W}]/2\sigma^{2}_{W}), Pb,b†=exp⁡(−Tr[b†b]/2σb2)P_{\mathbf{b},\mathbf{b}^{\dagger}}=\exp(-\text{Tr}[\mathbf{b}^{\dagger}\mathbf{b}]/2\sigma^{2}_{b}) respectively.

We want to express the network density and its correlation functions entirely in terms of complex-valued outputs, weights and biases, therefore, we need to transform the measures of W1,W2,b1,b2W^{1},W^{2},b^{1},b^{2} into measures over W,b\mathbf{W},\mathbf{b}. As DW1DW2=∣J∣DWDW†DW^{1}DW^{2}=|J|D\mathbf{W}D\mathbf{W}^{\dagger}, for Jacobian of [W1W2]=[1212i2−i2][WW†]\begin{bmatrix}W^{1}\\ W^{2}\end{bmatrix}=\begin{bmatrix}\frac{1}{2}&\frac{1}{2}\\ \frac{i}{2}&-\frac{i}{2}\end{bmatrix}\begin{bmatrix}\mathbf{W}\\ \mathbf{W}^{\dagger}\end{bmatrix}, we obtain DW1DW2Db1Db2=∣J∣2DWDW†DbDb†DW^{1}DW^{2}Db^{1}Db^{2}=|J|^{2}D\mathbf{W}D\mathbf{W}^{\dagger}D\mathbf{b}D\mathbf{b}^{\dagger}, and ∣J∣2=1/4|J|^{2}=1/4. With this, the nn-pt function for any number of ff’s and f†f^{\dagger}’s becomes the following,

where ZθZ_{\theta} is the normalization factor. We emphasize that the transformation of fif_{i} and fj†f^{\dagger}_{j} only transforms two indices inside the trace in Tr(W†W)\text{Tr}(\mathbf{W}^{\dagger}\mathbf{W}); it is invariant in this case, and also when all four indices transform.

From the structure of PW,W†P_{\mathbf{W},\mathbf{W}^{\dagger}} and Pb,b†P_{\mathbf{b},\mathbf{b}^{\dagger}}, only those terms in the integrand of (A.2), that are functions of Wis†WisW^{\dagger}_{i_{s}}W_{i_{s}} and bit†bitb^{\dagger}_{i_{t}}b_{i_{t}} alone, and not in product with any number of Wiu,Wiu†,biu,biu†W_{i_{u}},W^{\dagger}_{i_{u}},b_{i_{u}},b^{\dagger}_{i_{u}} individually, result in a non-zero integral. Thus, we have the only non-vanishing correlation functions from an equal number of ff’s and f†f^{\dagger}’s. We hereby redefine the correlation functions of this complex-valued network as

where {p1,⋯ ,p2n}\{p_{1},\cdots,p_{2n}\} can be any permutation of set {1,⋯ ,2n}\{1,\cdots,2n\}.

A.3 S​U​(d)𝑆𝑈𝑑SU(d) Input Symmetry

We now show an example of neural networks densities invariant under SU(d)SU(d) input transformations, provided that input layer parameters are drawn from an SU(d)SU(d) invariant distribution.

We will take a linear input layer and turn off bias for simplicity, as it may be trivially included as in SU(D)SU(D) output symmetry. SU(d)SU(d) group acts on complex numbers, therefore network inputs and input layer parameters need to be complex, such a network function is fi=gij(Wjkxk)f_{i}=g_{ij}(\mathbf{W}_{jk}x_{k}). The distribution of W\mathbf{W} is obtained from products of distributions of its real and imaginary parts W1,W2W^{1},W^{2}. Following SU(D)SU(D) output symmetry demonstration, the simplest SU(d)SU(d) invariant PW,W†P_{\mathbf{W},\mathbf{W}^{\dagger}} is obtained when W1,W2∼N(0,σW2)W^{1},W^{2}\sim\mathcal{N}(0,\sigma^{2}_{W}) are both SO(d)SO(d) invariant, we get PW,W†=exp⁡(−Tr[W†W]/2σW2)P_{\mathbf{W},\mathbf{W}^{\dagger}}=\exp(-\text{Tr}[\mathbf{W}^{\dagger}\mathbf{W}]/2\sigma^{2}_{W}). The measure of W\mathbf{W} is obtained from the measures over W1,W2W^{1},W^{2} as DW1DW2=∣J∣DWDW†DW^{1}DW^{2}=|J|D\mathbf{W}D\mathbf{W}^{\dagger}, with ∣J∣=1/2|J|=1/2. Following a similar analysis as (A.2), the only non-trivial correlation functions are

Under input rotations xi↦Sijxj, xk†↦xl†Slk†x_{i}\mapsto S_{ij}x_{j},~{}x^{\dagger}_{k}\mapsto x^{\dagger}_{l}S^{\dagger}_{lk} by S∈SU(d)S\in SU(d), the correlation functions transform into

A.4 Translation Input Symmetry

We demonstrate an example of network densities that remain invariant under continuous translations on input space, when the input layer weight is deterministic and input layer bias is sampled from a uniform distribution on the circle, b∼U(S1)b\sim\mathcal{U}(S^{1}). We will map the weight term to the circle by taking it mod 1, (i.e. % 1).

The network output fi(x)=gij((Wjkxk) %  1+bj)f_{i}(x)=g_{ij}((W_{jk}x_{k})\,\%\,\,1+b_{j}) transforms into f′(x′)=gij((Wjkxk) %  1+bj′)f^{\prime}(x^{\prime})=g_{ij}((W_{jk}x_{k})\,\%\,\,1+b^{\prime}_{j}) under translations of inputs xk↦xk+ckx_{k}\mapsto x_{k}+c_{k}, where bj′=(Wjkck) %  1+bjb^{\prime}_{j}=(W_{jk}c_{k})\,\%\,\,1+b_{j}. With a deterministic WW, the network parameters are given by θ={ϕ,b}\theta=\{\phi,b\}, and Db′=DbDb^{\prime}=Db. The transformed nn-pt function is

A.5 S​p​(D)𝑆𝑝𝐷Sp(D) Output Symmetry

We also demonstrate an example of network densities that remain invariant under the compact symplectic group Sp(D)Sp(D) transformations on output space.

The compact symplectic Sp(D)Sp(D) is the rotation group of quaternions, just as SUSU is the rotation group of complex numbers. Thus, a network with linear output layer would remain invariant under compact symplectic group, if last linear layer weights and biases are quaternionic numbers, drawn from Sp(D)Sp(D) invariant distributions. We define the network output fi(x)=Wijgj(x,θg)+bif_{i}(x)=W_{ij}g_{j}(x,\theta_{g})+b_{i} as before, with parameters Wab=Wab,0+iWab,1+jWab,2+kWab,3W_{ab}=W_{ab,0}+iW_{ab,1}+jW_{ab,2}+kW_{ab,3} and ba=ba,0+iba,1+jba,2+kba,3b_{a}=b_{a,0}+ib_{a,1}+jb_{a,2}+kb_{a,3}, such that Hermitian norms Tr(W†W)=Wab†Wab=∑i=03Wab,i2\text{Tr}(W^{\dagger}W)=W^{\dagger}_{ab}W_{ab}=\sum_{i=0}^{3}W^{2}_{ab,i} and Tr(b†b)=ba†ba=∑i=03ba,i2\text{Tr}(b^{\dagger}b)=b^{\dagger}_{a}b_{a}=\sum_{i=0}^{3}b^{2}_{a,i} are compact symplectic Sp(D)Sp(D) invariant by definition, where the conjugate of a quarternion q=a+ib+jc+kdq=a+ib+jc+kd is q∗=a−ib−jc−kdq^{*}=a-ib-jc-kd. The distributions of W,bW,b are obtained as products of distributions of the components W0,W1,W2,W3W_{0},W_{1},W_{2},W_{3} and b0,b1,b2,b3b_{0},b_{1},b_{2},b_{3} respectively. Following the SU(D)SU(D) symmetry construction, we can obtain the simplest Sp(D)Sp(D) invariant PW,W†P_{W,W^{\dagger}} and Pb,b†P_{b,b^{\dagger}} when these are functions of the Hermitian norm, and PDF of each component PWiP_{W_{i}} is an exponential function of SO(D)SO(D) invariant term Tr(WiTWi)\text{Tr}(W^{T}_{i}W_{i}) with equal coefficient, similarly with bias. Starting with W0,W1,W2,W3∼N(0,σW2)W_{0},W_{1},W_{2},W_{3}\sim\mathcal{N}(0,\sigma^{2}_{W}), and b0,b1,b2,b3∼N(0,σb2)b_{0},b_{1},b_{2},b_{3}\sim\mathcal{N}(0,\sigma^{2}_{b}), we get Sp(D)Sp(D) invariant quaternionic parameter distributions PW,W†=exp⁡(−Tr(W†W)/2σW2)P_{W,W^{\dagger}}=\exp(-\text{Tr}(W^{\dagger}W)/2\sigma^{2}_{W}) and Pb,b†=exp⁡(−Tr(b†b)/2σb2)P_{b,b^{\dagger}}=\exp(-\text{Tr}(b^{\dagger}b)/2\sigma^{2}_{b}). We also obtain the measures over W,bW,b from measures over WiW_{i} and bib_{i}, e.g. DWDW†=∣J∣DW0DW1DW2DW3DWDW^{\dagger}=|J|DW_{0}DW_{1}DW_{2}DW_{3}. Following an analysis similar to (A.2), it can be shown that the only non-trivial correlation functions of this quaternionic-valued network are

for {p1,⋯ ,p2n}\{p_{1},\cdots,p_{2n}\} any permutation over {1,⋯ ,2n}\{1,\cdots,2n\}. Under Sp(D)Sp(D) transformation of outputs fi↦Sijfj, fk†↦fl†Slk†f_{i}\mapsto S_{ij}f_{j},~{}f^{\dagger}_{k}\mapsto f^{\dagger}_{l}S^{\dagger}_{lk}, by S∈Sp(D)S\in Sp(D) in quaternionic basis, the correlation functions transform as

A.6 Preserving Symmetry During Training: Examples

We study further the example of an SO(D)SO(D) output symmetry from Section 2. Turning off the bias for simplicity, the network function is

with parameters θ={W,θg}\theta=\{W,\theta_{g}\}; transformations of fif_{i} may be absorbed into WW, i.e. W=θTW=\theta_{T}.

The network density remains symmetric during training when the updates to PθP_{\theta} preserve symmetry; for this example, we showed in Section 3 that it occurs when L\mathcal{L} is invariant and

where ∂L/∂θg\partial\mathcal{L}/\partial\theta_{g} is invariant because L\mathcal{L} is and θg\theta_{g} does not transform. Furthermore,

which satisfies the second condition in (39), since the first index is the one that transforms when WW absorbs the transformation of fif_{i}.

Appendix B More General S​O𝑆𝑂SO Invariant Network Distributions

We will now give an example of an SO(D)SO(D) invariant non-Gaussian network distribution at infinite width, as parameters of the initialized SO(D)SO(D) invariant GP become correlated through training. Such a network distribution can be obtained up to perturbative corrections to the initialized network distribution, if the extent of parameter correlation is small.

Training may correlate last layer weights θij\theta_{ij} of a linear network output fi=θijgj(x,θg)f_{i}=\theta_{ij}g_{j}(x,\theta_{g}) initialized with θij∼N(0,σθ2)\theta_{ij}\sim\mathcal{N}(0,\sigma^{2}_{\theta}), such that at a particular training step, we get Pθ=e−12σθ2θαβ2−λθ θabθabθcdθcd\mathcal{P}_{\theta}=e^{-\frac{1}{2\sigma^{2}_{\theta}}\theta^{2}_{\alpha\beta}-\lambda_{\theta}\,\theta_{ab}\theta_{ab}\theta_{cd}\theta_{cd}} independent from PθgP_{\theta_{g}}, with small λθ\lambda_{\theta}. Correlation functions of this network distribution can be obtained by perturbative corrections to correlation functions of the network distribution at initialization. For example, the 22-pt function of the correlated network distribution is given by

Appendix C S​O​(D)𝑆𝑂𝐷SO(D) Invariance in Experiments

The correlator constraint (3) gives testable necessary conditions for a symmetric density. Consider a single-layer fully-connected network, called Gauss-net due to having a Gaussian GP kernel defined by f(x)=W1(σ(W0x+b0))+b1f(x)=W_{1}(\sigma(W_{0}x+b_{0}))+b_{1}, where W0∼N(0,σW2/d)W_{0}\sim\mathcal{N}(0,\sigma_{W}^{2}/\sqrt{d}), W1∼N(0,σW2/N)W_{1}\sim\mathcal{N}(0,\sigma_{W}^{2}/\sqrt{N}), and b0,b1∼N(0,σb2)b_{0},b_{1}\sim\mathcal{N}(0,\sigma_{b}^{2}), with activation σ(x)=exp(W0x+b0)/exp(2(σb2+σW2/d))\sigma(x)={\rm exp}(W_{0}x+b_{0})/\sqrt{{\rm exp}(2(\sigma_{b}^{2}+\sigma_{W}^{2}/d))}.

To test for SO(D)SO(D) invariance via (3), we measure the average elementwise change in nn-pt functions before and after an SO(D)SO(D) transformation. To do this we generate 22-pt and 44-pt correlators at various DD for a number of experiments and act on them with 10001000 random group elements of a given SO(D)SO(D) group. Each group element is generated by exponentiating a random linear combination of generators of the corresponding algebra, namely

for b=1,⋯ ,1000b=1,\cdots,1000, p=dim(SO(D))=D(D−1)/2p=\text{dim}(SO(D))=D(D-1)/2, αi∼U(0,1)\alpha_{i}\sim\mathcal{U}(0,1) and TiT^{i} are generators of so(D)\mathfrak{so}(D) Lie algebra; i.e. D×DD\times D skew-symmetric matrices written in a simple basis The generators are obtained by choosing each of the D(D−1)2\frac{D(D-1)}{2} independent planes of rotations to have a canonical ordering with index ii, determined by a (p,q)(p,q)-plane. Each ithi^{th} plane of rotation has a generator matrix [Ti]D×D[T^{i}]_{D\times D} with Tpqi=−1T^{i}_{pq}=-1, Tqpi=1T^{i}_{qp}=1, for p<qp<q, rest . For instance, at D=3D=3, there are 33 independent planes of rotation formed by direction pairs {2,3}\{2,3\}, {1,3}\{1,3\} and {1,2}\{1,2\}. For each ithi^{\text{th}} plane defined by directions {p,q}\{p,q\}, general SO(3)SO(3) elements [Ri]3×3[R^{i}]_{3\times 3} have Rpqi=−sin⁡θ,Rqpi=sin⁡θ,Rqqi=cos⁡θ,Rqqi=cos⁡θ,Rrri=1R^{i}_{pq}=-\sin\theta,R^{i}_{qp}=\sin\theta,R^{i}_{qq}=\cos\theta,R^{i}_{qq}=\cos\theta,R^{i}_{rr}=1 for p<q , r≠p,qp<q\,,\,r\neq p,q and variable θ\theta. Expanding each RiR^{i} in Taylor series; the coefficients of O(θ)\mathcal{O}(\theta) terms are taken to define the generators of so(3)\mathfrak{so}(3) as in (44). . For example, for D=3D=3 we take the standard basis for so(3)\mathfrak{so}(3),

We define the elementwise deviation Mn=abs(G′(n)−G(n))\mathcal{M}_{n}=\text{abs}({G}^{\prime(n)}-G^{(n)}) to capture the change in correlators due to SO(D)SO(D) transformations. Here Gi1,⋯ ,in′(n)(x1,…,xn):=Ri1p1⋯RinpnGp1,⋯ ,pn(n)(x1,…,xn){G}^{\prime(n)}_{i_{1},\cdots,i_{n}}(x_{1},\dots,x_{n}):=R_{i_{1}p_{1}}\cdots R_{i_{n}p_{n}}G^{(n)}_{p_{1},\cdots,p_{n}}(x_{1},\dots,x_{n}) is the SOSO-transformed nn-pt function; both Mn\mathcal{M}_{n} and G(n)G^{(n)} have the same rank.

Error bounds for deviation Mn\mathcal{M}_{n} are determined as δMn=(δG′(n))2+(δG(n))2\delta\mathcal{M}_{n}=\sqrt{(\delta{G}^{\prime(n)})^{2}+(\delta G^{(n)})^{2}} using the standard error propagation formulae, δG(n)\delta G^{(n)} equals the average (elementwise) standard deviation of nn-pt functions across 10 experiments, and δG′(n)\delta{G}^{\prime(n)} is calculated as the following

We take the deviation tensors Mn\mathcal{M}_{n} over 1010 experiments for SO(3)SO(3) and SO(5)SO(5) transformations of 22-pt and 44-pt functions at D=3,5D=3,5 respectively, both correlators of each experiment are calculated using 4⋅1064\cdot 10^{6} and 10610^{6} network outputs respectively. An element-wise average and standard deviation across 1010 deviation tensors Mn\mathcal{M}_{n} are taken and then averaged over, to produce the mean of the SOSO-transformation deviation μMn\mu_{\mathcal{M}_{n}}, and its error σMn\sigma_{\mathcal{M}_{n}}, respectively. We plot μMn±σMn\mu_{\mathcal{M}_{n}}\pm\sigma_{\mathcal{M}_{n}} (the blue shaded area) in Figure (2), this signal lies well within predicted error bounds of ±δMn\pm\delta_{\mathcal{M}_{n}} (in orange), although μMn\mu_{\mathcal{M}_{n}} deviates significantly from at low widths, in contradiction to width independence of (3). This is due to the smaller sample size of parameters in low-width networks, and therefore small fluctuations in the weight and bias draws lead to more significant deviations from the "true" distribution of these parameters. A nonzero mean of the parameters caused by fluctuations leads to a nonzero mean of the function distribution ⟨f⟩≠0\langle f\rangle\neq 0, thus breaking SO(D)SO(D) symmetry. We believe this is a computational artefact and does not contradict SOSO-invariance in (3).

Appendix D Experiment Details

The experiments in Section (4) were done using Fashion-MNIST under the MIT LicenseThe MIT License (MIT) Copyright © 2017 Zalando SE, https://tech.zalando.com, using 6000060000 data points for each epoch split with a batch size of 6464, and test batch size of 10001000. Each experiment was run on a 240GB computing node through the Discovery Cluster at Northeastern University and took between 1 and 1.5 hrs to train 20 epochs. The experiments were repeated 20 times for each configuration, and were run with 21×621\times 6 configurations for the left plot in Fig. (1), and 1111 for the right plot. The error in the left plot of Fig. (1) is shown in Fig. (3).

The experiments in Appendix (C) were done on the same cluster, with the same memory nodes. These took around 24 hours on 10 compute nodes to generate models for each of the DD values. The nn-pt functions then took another 6 hours on a single node each.

Appendix E Comments on Functional Densities

Functional integrals are often treated loosely by physicists: they use them to great effect and experimental agreement in practice, but they are not rigorously defined in general; see, e.g., .

We follow in this tradition in this work, but would like to make some further comments regarding cases that are well-defined, casting the discussion first into the language of Euclidean QFT, and then bringing it back to machine learning.

First, the standard Feynman functional path integral for a scalar field ϕ(x)\phi(x) is

but in many cases the action S[ϕ]S[\phi] is split into free and interacting pieces

where the free action SF[ϕ]S_{F}[\phi] is Gaussian and the interacting action Sint[ϕ]S_{\rm int}[\phi] is non-Gaussian. The free theory, which has Sint[ϕ]=0S_{\rm int}[\phi]=0, is a Gaussian process and is therefore well-defined. When interactions are turned on, i.e. the non-Gaussianities in Sint[ϕ]S_{\rm int}[\phi] are small relative to some scale, physicists compute correlation functions (moments of the functional density) in perturbation theory, truncating the expansion at some order and writing approximate moments of the interacting theory density in terms of a sum of well-defined Gaussian moments, including higher moments.

In this work we study functional densities associated to neural networks, and have both the perturbative and especially the lattice understanding in mind when we consider them. In particular, for readers uncomfortable with the lack of precision in defining a functional density, we emphasize that our results can also be understood on a lattice, though input symmetries may be discrete subgroups of those existing in the continuum limit. Furthermore, any concrete ML application involves a finite set of inputs, and for any fixed application in physics or ML one can simply choose the spacing between lattice points to be smaller than the experimental resolution.