Separation of Scales and a Thermodynamic Description of Feature Learning in Some CNNs
Inbar Seroussi, Gadi Naveh, Zohar Ringel
Introduction
Identifying slow or relevant variables is an essential step in analyzing large-scale non-linear systems. In the context of deep neural networks (DNNs), these should be some combinations of the individual weights that are weakly fluctuating and obey a closed set of equations. One potential set of such variables is the DNNs’ outputs themselves. Indeed, in the limit of infinitely over-parameterized DNNs these provide an elegant picture of deep learning based on a mapping to Gaussian Processes (GPs). However, these GP limits miss out on several qualitative aspects, such as feature learning and the fact that real-world DNNs are not nearly as over-parameterized as required for the GP description to hold . Obtaining a useful set of slow variables for describing deep learning at finite over-parameterization is thus an important open problem in the field.
Several works provide guidelines for this search. Noting that GP limits can have surprisingly good performance and that over-parameterization is natural to deep learning we are inclined to keep some elements of the GP picture. One such element is to work in function space and study pre-activation and outputs instead of weights whose posterior distribution becomes complicated even in the GP limit . Another element is the layer-wise composition of hidden layer kernels , using which one generates the output kernel of the GP . Such a layer-wise picture is also harmonious with the idea that DNN layers should not correlate strongly, to prevent co-adaptation . Recently, it was shown that in some limited settings, making the GP kernel "dynamical" or flexible, so that it adapts to the dataset, can account for differences between infinite and finite DNNs . Still, the task of finding an explicit and general set of equations describing this flexibility in DNNs remains unsolved.
In this work, we identify such slow variables and use these to derive an effective theory for deep learning capable of capturing various finite channel/width () effects in convolutional neural networks (CNNs) and fully connected neural networks (FCNs) such as feature learning. We argue that:
For, the erratic behavior of specific channels/neurons averages out and hidden layers coupled to each other only through two “slow" variables per layer: The second moment of the pre-activations (pre-kernel), , and the second moment of activations (post-kernel), , of the th layer. Furthermore, for mean square error (MSE) loss, FCNs in the so-called mean-field (MF) scaling (where the last layer weights are scaled down) or CNNs with a large read-out layer fan-in behave effectively as a GP with a data-aware kernel determined by the second moment of pre-activations in the penultimate layer.
In settings where the kernels have a large density of dominant eigenvalues, the posterior (or trained) pre-activations fluctuate in a nearly Gaussian manner. Following this we use a multivariate Gaussian variational approximation for the posterior pre-activations (pre-kernels) and derive explicit equations (equations of state), for the covariance matrices governing these pre-activations.
We identify an emergent feature learning scale (FLS) denoted by , proportional to the train MSE times over (or ). This scale controls the difference between the finite output kernel () and its limit and in this sense reflects feature learning. Due to the factor, can be or larger even for , e.g. for CNN architectures (see Fig. 1 panel c). The same holds, with replaced by , for FCNs in the MF scaling . Unlike perturbation theory , our theory tracks all orders of and treats only perturbatively. The separation of scales between and is thus central to our analysis. Its manifestation is the fact that feature learning shifts and stretches the dynamical variables in the theory (the pre-activations) in a considerable manner yet barely spoils their Gaussianity.
The predictions of our approach are tested on several toy and real-world examples using direct analytical approaches and numerical solutions to the equations of state. Our analysis takes a physics viewpoint on this complex non-linear problem. Rigorous mathematical proofs are left as an open problem for future research.
We note that there are several works showing evidence that the spectrum of the empirical weight correlation matrix show various tail effects and spikes . While in deeper layers we focus on pre-activations, the spectrum of input layer weights we obtained, is Gaussian but not independent as in Ref. . Hence it can produce a variety of spectral distributions for the covariance matrix, similar to the aforementioned ones. We note a recent interesting work arguing that the test-loss depends only on the mean and variance of hidden activations. There, however, the setting is of a fixed trained DNN and the statistics are over the input measure rather than over the DNN parameters as in our case. While quantitatively different, our approach is similar in spirit to the phenomenological layer-wise Gaussian Processes put forward in a recent work . Additional approaches for finite-width include perturbative correction around the infinite width limit to leading or higher orders . There is however mounting evidence from bounds on GP limits , numerical experiments , as well as the current work, that such perturbative expansions have slow convergence in practical regimes. In contrast our EoS are useful both numerically (see Sec. 4 and 2.2) and analytically (see Sec. 2.3) and in addition allow us to model pre-activation distributions in the wild via our pre-kernels (see Sec. 2.4).
Our theory can be applied to any finite number of convolutional, dense, or pooling layers. To illustrate its main aspects, let us focus on an -layer fully connected model with width (),
The main object we analyze is the equilibrium distribution of the GD+Noise algorithm in function space. Adopting physics notation, this distribution can be written as , where is the partition function, and is the action or negative log-posterior (see also Methods). Taking a Bayesian perspective, this probability distribution can also be viewed as a posterior distribution given measurements of having Gaussian noise with variance and a prior given by a finite-width random DNN. As shown in Supp. Mat. (1) our partition function is governed by the following action
Here, however, our focus is at large but finite width (). In this more complex regime, several corrections may appear: (i) The pre-activations’ average and covariance may deviate from those of a random DNN. (ii) , the covariance of activations in the layer, would not solely determine the covariance of pre-activations on the downstream layer , as upstream effects between and come into play. (iii) Inter-channel (or inter-neuron in the fully-connected case) and inter-layer correlations may appear. (iv) The fluctuations of pre-activations may deviate from that of a Gaussian. A priori, all these corrections may play similarly dominant roles, thereby making analysis cumbersome.
Results
This leaves us with corrections of types (i) and (ii). Interestingly, following these corrections to all orders leads to a tractable mean-field picture of learning. The latter is an augmentation of the standard correspondence between GPs and DNNs at infinite width (NNGP): Pre-activations in different layers or channels/neurons remain uncorrelated and Gaussian. Correlations only appear between different data-points (and latent pixels for CNNs) within the same layer and channel/neuron. We henceforth denote the covariance of pre-activations and activations at layer (up to normalization by the variance of the weights) by and and refer to these as pre-kernel and post-kernel, respectively. However in the NNGP viewpoint, is simply proportional to and fully determined by the upstream kernel () whereas here and differ and moreover depend both on the upstream and downstream kernels.
As their lack of dependence on width suggests, the first equation together with the definitions of the post-kernels are already present in the strict GP limit (). They are, respectively, the GP inference formula and standard kernel recursive equations of random DNNs with activation. The remaining equations are, to the best of our knowledge, novel and follow the changes to the pre-kernels and post-kernels at finite . These could be solved analytically in some simple cases (see subsection 2.3 for the case of two-layer CNN). We note that for non-anti-symmetric activation, one will also need to track the mean of each layer’s pre-activation (see Supp. Mat. (5)).
To get a qualitative impression of their role, one can consider the case where the penultimate layer () is linear, in which case . Consequently, , (where , with double index refers here to the Kronecker delta) and thus the second equation simplifies to
where . We note in passing that even for a non-linear penultimate layer, a similar term will arise from the expansion of in to linear order. From the above form, several insights can be drawn.
First, we argue that the above equation implies that the trained DNN is more susceptible to changes along than the DNN at . Noting how enters the action (Eq. 2), it controls the stiffness associated with fluctuations in . Hence makes fluctuations in the direction of more likely than they are according to . Since measures the discrepancy in train predictions, this effect reduces the discrepancy by making the DNN more responsive in these directions than it is at . The second term, proportional to amounts to a negligible reduction in fluctuations along eigenvectors of corresponding to eigenvalues which are larger than .
Using Eq. (5) one can also identify the aforementioned emergent feature learning scale (or FLS) namely, . This scale represents the magnitude of the leading term when one Taylor expands in . When or larger there is a significant change in the eigenvalues of compared to which indicates feature learning. On the other hand, when this quantity is small, we are closer to the GP regime (see Supp. Mat. (1.6)). To asses the scaling , one can consider the common situation where has some non-negligible overlap with dominant eigenvectors of whose eigenvalues are on the scale . Here we find , where MSE denotes the mean train MSE which enters here via . Due to its explicit dependency, and for at large — maybe even at very large and/or when the average MSE is rather small.
Figure 1 panel (c) shows the value of (or ) at which (i.e. or at which feature learning becomes a dominant effect) as a function of for several DNNs we study. The scale separation, demonstrated there by the fact that can be in regions where ’s is negligible, is central to our analytical approach.
This scale is also the reason that naive perturbation theory in fails at large , as it treats and on the same footing, since they both have a single negative power of . In contrast, our EoS treat the FLS non-perturbatively.
Last we stress that the EoS provide us with a concrete effective GP description for the entire DNN as well as its hidden layers. A priori one would expect that the normality of pre-activations, a large trait, will be lost at finite . Yet we find that pre-activations remain Gaussian and accommodate strong feature learning effects while maintaining accurate predictions. This unexpectedly simple behavior opens various reverse engineering possibilities wherein one infers the effective kernels from experiments and uses their spectrum and eigenvectors to rationalize about the DNN (see also Fig. 5).
2 Numerical Demonstration: 3 Layer FCN
Next, we test the agreement between the above results and statements and actual trained DNNs, starting from the 3-layer FCN defined in Eq. (1) with . We focus here on a student-teacher setting with or training data points drawn from iid Gaussian distributions with unit variance along each input dimension. The target was generated by a randomly drawn teacher FCN of the same type only with . The student was trained using an analog scaling to the MF scaling , wherein the output layer weights are scaled down by a factor of . Whereas for the CNNs discussed below, this choice of scaling was not required, for FCNs we found it necessary for getting any appreciable feature learning at (see Fig. 1 panel (c)).
As described in the methodology section 4, we trained FCNs using the GD+Noise algorithm until they reached equilibrium. We use these trained FCNs to calculate various average quantities under our partition function (Eq. (2)). Specifically, we focused on: (i) The normalized train-loss on the scale of , namely (ii) The eigenvalues () of the average , where we average over neurons, training seeds, and training time (the latter within the equilibrium region). (iii) The normalized overlap () between the discrepancy in prediction on the training set times the target, namely . We then used a JAX-based numerical solver for the EoS and compared it with the experiment.
As the results of Fig. 1 panel (b) show, the predictions of our EoS for all these three quantities converged well as we increased . Furthermore, they do so in a region where they differ considerably from their associated GP limit. Indeed, as shown in the Supp. Mat. (3) the top eigenvalue came out 2-3 times larger than it is in the GP limit. The associated eigenvector corresponded to the first layer weights of the teacher (). The rest of the eigenvalues remained at their GP limit values. Put together, this is a clear sign of strong feature learning.
Notably, however, this notion of feature learning does not involve compression. Indeed, since has the same variance as in the GP limit for directions perpendicular to , it does not compress the input by projecting it solely on the label relevant direction (). Instead, it exaggerates the fluctuation of student weights along, thereby making it statistically more likely that and with opposite sign of and will be further apart in the space of pre-activations.
Next, we study how behaves as a function of and (or ) for different architectures. Fig. 1 shows the value of (or ) at which . As contains a single inverse power of at ten times this value, would be and thus indicate only minor feature learning effects in our EoS. As diminish from this latter value, our EoS yield increasingly stronger feature learning effects. We find that both for CNNs in the standard scaling and for FCNs with MF scaling, the crossover to feature learning happens well within the validity region of our mean-field decoupling (i.e. large or ). In contrast, FCN with standard scaling shows this crossover when , which is outside the scope of our theory. In this aspect, we comment that there is evidence that FCNs with standard scale are inferior to those with mean-field scaling and perform similarly to GPs .
3 Analytical Solution of the EoS - Two Layer CNN
Having tested our EoS numerically, we turn to show they lend themselves, in simple settings, to a fully analytical calculation. Amongst other things, this will flesh out the non-perturbative nature of our results. To this end, we consider a simple non-linear CNN with 2 layers. Though bounds have been derived , we are not aware of any analytical predictions for the performance of finite non-linear 2-layer DNNs, let alone CNNs. It is therefore a natural first application of our approach. Specifically, we consider
The equations of state are given by (See Supp. Mat. (3))
Here we denote the discrepancy from the target by . The above equations for and could be solved numerically and compared with DNN training experiments. The results are shown in Fig. (2) in solid lines and match empirical values well.
To obtain fully analytical results, we proceed with several approximations for large . First, we approximate the spectrum of the matrix based on its continuum kernel version . This is closely related to the equivalent kernel approximation, which we adopt here along with its leading order correction . Similarly, we use large to replace the double summation by two integrals over the measure from which are drawn (). See Supp. Mat. (3.1) for further details and a discussion of the fully connected case ().
The latter approximation also fleshes out the importance of the FLS (). Technically, the two summations provide for the scaling and is related to the MSE via (see also Supp. Mat. (3)). The FLS controls the deviation between the pre-kernel and post-kernel of the penultimate layer with our EoS. Hence, in particular, it implies deviations from GP where these are not equal.
Following our approximations, the equations acquire the full rotation symmetry of the data-set measure, which amount to an independent orthogonal transformation of each . Furthermore, as shown in Supp. Mat. (3), at large , (the continuum function representing ) is linear given a linear target, , regardless of and hence so is . The above symmetry then implies that is only a function of and furthermore takes the simple form . The quantity thus measures the overlap between the discrepancy in predictions () and the target. Following this, the EoS are reduced to a non-linear equation in a single variable
Solving the above equation for , one obtains and hence and also (via Eq. (7)). Using the obtained one can calculate the DNN’s predictions on the test-set. The effect of the FLS is evident in the second equation where it controls the deviations from the GP limit. Here we also recall that is the train MSE over , thus as defined above contains the MSE factor mentioned in the introduction.
To test the theoretical predictions, we trained two such CNNs, with and varying channel number. Fig. 2, left panel, shows the empirical test-set values for (dots) compared with a numerical solution of the equations of state (solid lines) and their analytical solution (dashed lines). For the latter, we obtained analytically and performed the resulting GP inference with numerically. The inset tracks the train-set results which, in this case, are fully analytical and involve no numerical GP inference. Both predictions match empirical values quite well, even in the regime where test root MSE is roughly half that of a Gaussian Process (). The right panel shows the input layer weights, dotted with and with a normalized random vector. These remain Gaussian up to minor statistical noise. Further details can be found in the methods section.
To emphasize the non-perturbative nature of our Eq. (111), let us assume for the sake of negation that they agree with first order perturbation theory in (as in Refs. ). If so, we may replace in the above expression for by its GP value, as it already contains one negative power of and hence receives no further corrections at that order. Numerics show this value for . Plugging this in, one obtains . Clearly, this logic leads to a contradiction unless . In contrast, our theory provides highly accurate predictions for and well away from where admits a perturbation theory in . In Supp. Mat. (6.2) we report additional results on over its GP value.
4 Extensions to Deeper CNNs and Subsets of Real-World Data-sets
For truly deep CNNs and real-world datasets, obtaining fully analytical predictions for DNN performance is a challenging task even in the limit. Still, the EoS could be solved numerically and compared with experimental values. Furthermore, the quantities which underlie them could be examined and reasoned upon. We do so here in two richer settings, a 3-layer CNN trained with a teacher CNN and the Myrtle-5 CNN trained on a subset of CIFAR-10.
Our first setting extends that of the previous subsection by having an extra activated layer and a non-linear target function. Specifically, we consider a student CNN defined by
As the first test of our theory, we examine the fluctuations of pre-activations in the input and middle layers of the trained student CNN and check their normality. Specifically, for the input weights we obtain the histogram (over channels, equilibrium samples, and seeds) of , where is the teacher input weight and the histogram of where is a random vector. Teacher overlap has a variance of here, whereas random overlap variance, averaged over choices of ’s, was with an std of . For the hidden layer, we obtain the histogram of where are the teacher’s pre-activations as well as where is the pre-activation of a different randomly chosen teacher. Teacher overlap variance here was, whereas average student variance was, with an std of . Fig. 4 shows the associated histograms along with their fit to a Gaussian. The large and consistent differences in the variance of the fluctuations between teacher directions and random directions show that we are deep in the feature learning regime. Remarkably, the fluctuations remain almost perfectly Gaussian. The larger variance along teacher directions implies that by drawing DNNs from the trained DNN ensemble and diagonalizing their empirical covariance matrices, one is more likely to find dominant eigenvalues along these teacher directions.
We turn to verify the EoS and rationalize on the behavior pre-kernels. To this end, we average the empirical pre-activations, over channels and training seeds, to obtain an estimator for the pre-kernel and post-kernel of (i.e. and ) and that of the input weights (). We then obtain analytically using the 3rd equation from Eqs. 4, plugging in the empirical . Finally, we compare the empirical with that obtained from the last equation from Eqs. (4). Fig. 5 left panel plots the eigenvalues of , our predictions (), and the post-kernel of the input layer which is simply , showing a good match between the first two. Figure (5) right panel plots the eigenvalues of as predicted from , compared with its empirical value.
Next, we trained the myrtle-5 CNN, capable of good performance and containing both pooling layers and ReLU activations, with on a subset of CIFAR-10 (). As shown in Fig. 6, pre-activations show a strong deviation of trained DNNs from non-trained DNNs or DNNs at infinite channel/width, and at the same time show quite a good fit to Gaussian in most cases. This opens the possibility of reverse engineering the pre-kernels governing this trained network and using them to rationalize about the DNN, for instance by identifying their dominant eigenvectors.
The 2nd layer (as well as the input layer, (see Supp. Mat. (6.3))) show deviations from Gaussianity in the leading eigenvalue. This is expected since the kernels of these layers show quite a dilute dominant spectrum whereas VGA requires a contribution from many adjacent modes (see Supp. Mat. (1.3)). Interestingly, despite this non-Gaussianity in the leading eigenvalue of layers 1 and 2, Gaussainity is restored in the downstream layers 3 and 4. Correlations across layers and across channels within the same layer are very weak (largely on the order of ) and fully consistent with the mean-field decoupling underlying this work. Further technical details are found in Supp. Mat. (6.3).
Discussion
In this work, we presented what is, to the best of our knowledge, a novel mean-field framework for analyzing finite deep non-linear neural networks in the feature learning regime. Central to our analysis was a series of mean-field approximations, revealing that pre-activations are weakly correlated between layers and follow a Gaussian distribution within each layer with a pre-kernel . Using the latter together with the post-kernel induced by the upstream layer, explicit equations-of-state (EoS) governing the statistics of the hidden layers were given. These enabled us to derive, for the first time, analytical predictions for the performance of non-linear CNNs and deep non-linear FCNs in the feature learning regime. We further note that our EoS generalize straightforwardly to combined CNN-FCN architectures, pooling layers, and models with multiple outputs. Our theory can also be viewed from a Bayesian perspective, the GP process represented by the equation we find is a good approximation to the true posterior distribution generated by a large but finite width Bayesian neural network.
Various aspects of this work invite further study. Empirically, it would be interesting to better characterize the scope of models for which Langevin dynamics (or potentially ensemble-averaged NTK dynamics) leads to Gaussian pre-activations and overall GP-like behavior. Probing the “feature-learning-load“ of each layer, by experimentally measuring the differences between the kernels and , may also provide insights on generalization, transfer learning, and pruning, thus complementing other diagnostic tools suggested recently . For instance, transferring a layer with a small feature learning load may provide little benefit, and pruning a channel having a large overlap with a leading eigenvalue of may be harmful.
From the theory side, it is desirable to develop analytical techniques for solving the EoS as well as guarantees regarding the existence and uniqueness of solutions. In particular, exploring the possibility of spontaneous symmetry breaking of internal symmetries such as weight inversion. Providing a mathematical underpinning for the approximations involved here may lend itself to developing performance bounds on the Langevin algorithm and Bayesian neural networks. Similarly, one can consider using the empirical effective kernel () as a starting point to develop GP-based bounds on performance. Last, it is interesting to explore the approach to equilibrium of the training dynamics and adapt the approximations carried here to the NTK setting .
Methods
Here we present the main ingredients of our theory, leading to the EoS we find. Further details can be found in the Supp. Mat.
First, we provide decoupling of Eq. (2) into layer-wise neuron-wise terms, wherein each of the terms depends on the upstream and downstream layers only through channel-averaged second moments of activations and pre-activations (pre-kernel and post-kernel). Further details are found in Supp. Mat. (1).
Consider the non-linear terms in the action 2 which couple the different layers. This coupling is mediated through the channel/width-averaged quantities: indeed depends on through the channel/width averaged square term in , depends on through the average of , and depends on through the average of and so forth. For we expect these to be weakly fluctuating and well approximated by their mean-field values. This behavior propagates till the output layer, and in particular implies that the outputs fluctuate in a Gaussian manner, as previously conjectured . As for the dependency of on the variables, it is not through a channel/width averaged quantity. However, we find that in various scenarios, such as FCNs with MF scaling or CNNs with large , the fluctuations of are suppressed enabling us to replace by its average (see Supp. Mat. (1.7) and (2.2)). Following this, we obtain our mean-field action,
Notably, any coupling between the different layers is only through static mean-field quantities, namely the pre-kernels and-post kernels. In addition, all neuron-neuron couplings (and similarly, channel-channel couplings for CNNs) have been removed.
1.2 Intra-Layer Decoupling
Despite the simplified inter-layer coupling and intra-layer neuron coupling, the mean-field actions are still non-quadratic for all layers but the output layer. This non-linearity couples all the variables for the same neuron (channel in the CNN case) in a way that is roughly all-to-all in the data-point index. In atomic and nuclear physics, similar circumstances are well described by self-consistent Hartree-Fock approximations . In our setting, this approximation is directly analogous to a variational Gaussian approximation (VGA). In Supp. mat. (4) we argue that in the typical case where the diagonal of is much larger than the off-diagonal elements, the VGA is well controlled. Technically, we do so by showing, order by order in perturbation theory, that the diagrams accounted for by the VGA approximation dominate all other perturbation theory diagrams. In Supp. Mat. (3) we also establish this using different means for for the specific case of two-layer CNN with a single activated layer. We further comment that the VGA is exact for deep linear DNNs.
Accordingly, we now look for the Gaussian distribution, governed by a kernel which is the closest to the above non-quadratic action. In models with many hidden layers, this leads to the following "inverse kernel shift" behavior for
For non-anti-symmetric ones, see Supp. Mat. (5).
2 Experimental Details
Hyperparameters. For the 2-layer CNN experiments, we used , , and varying channel number. The training parameters (noise and weight-decay) were tuned such that and weight variance of over fan-in, for both layers at . The target was drawn once for all experiments using i.i.d. Gaussian centered random and with variances and respectively.
For the 3-layer CNN experiments, we took . The training parameters (noise and weight-decay) were scaled such that and weight variance of over fan-in for the inputs and hidden layer with no training data (at initialization). The weight variance of the read-out layer was over the fan-in. The target was drawn again once for all experiments from a teacher CNN with .
For all the myrtle-5 experiments, we used and ReLU activation. The training parameters (noise and weight-decay) were scaled such that and weight variance of over fan-in for all layers with no training data (at initialization).
For all the FCN experiments, we used equal width () and weight decay corresponding to variance, (with no training data) in the regular scaling. For the MF scaling, we took . The target was drawn again once for all experiments from a teacher CNN with . Specifically when calculating the emergent scale, we used independent of .
Equilibrium sampling. To obtain weakly correlated samples from the equilibrium distribution of the trained CNNs we used the following procedure. For the 2 and 3-layer CNNs, we used an adaptive learning rate scheduler: For the first epochs we used a learning rate , then we crank up the learning rate to . As of epoch , every epoch we estimate the fluctuations of the train-loss and check for spikes - events in which the train-loss was 5-times larger than the standard deviation in the past epochs. If a spike is observed, the learning rate is reduced by a factor of . This continues until epochs pass without any events. Then the learning rate is reduced again by a factor of two and remains fixed. Samples from these final stages were treated as equilibrium samples. We further checked that (i) different initialization seeds trained with this protocol reached the same train-loss statistics. (ii) No further reduction in train-loss occurred after the final learning rate reduction. For several runs, we also verified that increasing the last reduction of learning rate by an additional factor of did not have any appreciable effect on the loss. The initial was (w.r.t. a standard mean reduction MSE loss) and the final learning rate was typically . The runs terminated at epoch .
For the myrtle-5 CNN trained on CIFAR-10, we first ran several runs for epochs using the above procedure and examined those that reached the lowest train-loss. We then generated a fixed scheduler based on those more successful instances, running up to epochs. We again verified that further lowering the final learning rate has no appreciable effect on the training loss and that different seeds reach similar final train-loss. This ensures that we are indeed sampling from a valid equilibrium distribution.
For the 3-layer CNN and Myrtle-5, we found that auto-correlation times of pre-activations change considerably between the layers. While the read-out layer typically had an auto-correlation time of the order of epochs (at the lowest learning rates) the auto-correlation times for the input layers could reach or larger values. To overcome this issue, when analyzing pre-activations of these deeper DNNs we took an ensemble containing and different initialization seeds for the 3-layer CNN and Myrtle-5 respectively.
For the 3-layer FCN we used a fixed scheduler which starts at the maximal stable learning rate and reduces the learning rate by factors of at epochs and by a factor of at epochs (factor of in total). Equilibrium sampling was done between epochs.
Numerical solution of the equations of state. For the 2-layer CNN, the equations of state were solved using Newton-Krylov method, which does not require explicit gradients. To facilitate convergence, we adopted an annealing procedure: For , we obtain the solution using a GP initial value () for . The optimization outcome was then used as for the next lower value of . Using 12 CPU cores, this optimization took several hours. After obtaining as a function of , the resulting kernel was used in standard GP inference to obtain on the test-set. For the 3-layer FCN we used a more efficient JAX-based code to generate the kernels and kernel derivatives involved in the EoS, but otherwise followed the same procedure. Optimization took between several minutes to a few hours on one Titan-X GPU, depending on parameters.
Supplementary Information - Separation of Scales and a Thermodynamic Description of Feature Learning in Some CNNs
Derivation of the Mean-field Equations for a Fully Connected Network
In this section, we derive the equations of state for deep fully connected NNs with a finite number of layers, . In Sec. 6 we provide a sketch of the derivation for CNN architecture. The analysis can also be extended to other architectures, including pooling layers and skip connections.
The model is composed of a layer NN having activated hidden layers, and one linear readout layer. Specifically, we consider
Our main object of interest is the following equilibrium distribution of the Langevin dynamics algorithm with noise strength and weight decay written in function space (i.e. in terms of the DNNs outputs)
In practice, we sample from this distribution using Gradient descent (GD), at small learning rates, together with weight decay and noise on each weight derivative. The weight decay parameters are for layer . The first term on the r.h.s. is given by
where is viewed now as a random variable following the NN outputs on all different training points. The average is over the weights of the network at equilibrium. The weights’ distribution at equilibrium can be obtained explicitly. At , for these distributions (Eq. (21) and Eq. (20)) tend to a GP, however our interest here is at finite .
Eq. (20) and Eq. (21) can also be understood from a Bayesian perspective. Eq. (20) can be viewed as the posterior distribution assuming each sample is generated by a neural network model as in Eq. (19) and is corrupted by an additive i.i.d. Gaussian noise with variance . The prior distribution over the weights of the network is taken to be Gaussian with variance .
To obtain a more explicitly "layer-wise" representation of Eq. 21, we next condition over the pre-activations () of each layer using Bayes’ formula, we obtain the following Markov representation of Eq. (21)
The hidden layers probabilities above are defined as follows:
where we write for short for , the latter being the random variables describing argument of the activation function (pre-activation) of the th layer at neuron , on the data-point. Later we will use this Markov structure of the distribution to decouple the different layers.
We continue our analysis by using the Fourier identity, which replaces the above delta functions by auxiliary fields, . To this end, the probability distribution over the network output given the input data can be written as follows.
where we adopt here physics notation, and define, , as the action associated with this distribution. Collecting all terms, the action is defined as follows
To obtain the above expression, we performed Gaussian integration over the weights of all layers. In addition, we define the following matrices:
In the following section, we derive the inter-layer mean-field decoupling. In this context, the form of the action in the presence of the auxiliary fields (Eq. (27)) turns out to be useful.
2 Mean-Field Decoupling
Our first step is to understand the dependence of on for all . Following the mean-field idea introduced in Eq. (32) we first subtract average quantities and rewrite the action of the th neuron of the th layer such that :
Following Eq. (34) for and , we gather all the dependent terms and obtain the mean-field action of the th layer:
2.2 Mean-Field Decoupling - Output Layer
We next consider the coupling between the final/output layer and the penultimate layer. As done previously in the analysis of the hidden layers, we rewrite the action of the top layer as
Our mean-field decoupling is therefore valid at but finite. We stress that even in this regime, feature learning effects can still be of order, as these are controlled by an emergent scale () containing positive powers of . See Fig. 1(c) in the main text.
For the layer, we obtain from Eq. (35) the following action:
where the mean field average of is then,
We can also now identify the last layer of the mean-field action:
Combining all the layers Eq. (35), Eq. (39), and Eq. (42), we can write the mean-field action:
3 Variational Gaussian approximation
For the first layer, one is free to choose either pre-activations or the weights themselves as the variables, as these two are linear functions of one another. Here we will use the weights, , (the th row of the matrix ) and denote the covariance matrix of these weights as . The pre-kernel matrix of the input layer pre-activations is then given by .
The KL divergence between the above distribution and a Gaussian distribution for all with covariance is then
Note that, this is indeed the optimal distribution since taking the second derivative yields and since is positive definite as a covariance matrix is negative definite, therefore it is a global minimum.
3.2 l=1𝑙1l=1
We turn to find the pre-kernel of the input layer. Here, we find it more convenient to work with the covariance matrix of the weights rather than the pre-activation, as it is a more compact object having fewer indices. We follow the same procedure and minimize the KL divergence for all to find the closest Gaussian distribution with covariance ,
taking the derivative with respect to
Plugging in the above expression for and combining with Eq. (46) we have that:
where and .
4 Equations of State (EoS)
Collecting all the results above, we obtain the following closed set of equations determining all pre-kernel and post-kernel as well as the average output of the Langevin algorithm, : namely
where , , and . While not explicitly apparent, the above expression does converge to the GP limit as for all . Indeed, for very large , . Consequently, the term is . The term on the r.h.s. thus vanishes as .
We note that the second term in Eq. (53) and Eq. (54) has a more profound meaning in terms of the information transfer between the pre-kernel and post-kernel. Looking at the KL divergence between two centered multidimensional Gaussian with kernel and of the same dimension
Taking the derivative with respect to for and with respect to for , we then obtain that:
where in the second transition we rearranged terms and used the cyclical property of the trace. Substituting this relation leads to Eq. (4) provided in the main text.
4.2 Auxiliary field correlation
The second moment correlation then can be found by taking twice the derivative of the following free entropy and taking the source field to zero:
Plugging this expression back in Eq. (60), and using the fact that all matrices are symmetric, we have that
5 An Emergent Scale
Here we identify a quantity () whose scale characterizes the amount of feature learning in the DNN and its deviations from the GP limit. More specifically, when becomes comparable to , strong feature learning effects appear and perturbation theory in becomes impractical. Since would consist of a non-trivial combination of factors, we refer to it as an emergent scale. To define this scale, we work within our EoS. and ask when changes in a noticeable manner as we lower from the at fixed (the GP limit). More technically, we next solve the EoS using perturbation theory in and estimate the magnitude of the leading terms we obtain.
For simplicity, we focus on a setting where the penultimate layer is linear, due to this linearity we obtain and so:
Following the aforementioned perturbation theory in , we perform a first-order Taylor expansion of yielding
To evaluate the magnitude of the first term we multiply by from both sides to obtain
We define , and find:
As we argue below, the last term on the r.h.s is small compared to the first two. Putting it aside and recalling that the first term on the right-hand side is the zeroth order term in , we find that controls the ratio between the zeroth order contribute and the first order perturbative correction. Hence, once becomes order , first-order perturbation in becomes inaccurate. However, it can be further checked that a second-order perturbation will contain a contribution, and hence this is not just a problem in first-order perturbation theory. Rather, it is that low order perturbation theory becomes inaccurate.
Next, we argue that having non-negligible also implies feature learning, in the sense that changes from its value. Indeed, the quantity we are examining () involves both the discrepancy () and the kernel . Thus, a-priory may change just because changes. However, is a function of via . Thus, it cannot undergo any change if remains inert. Thus, we conclude that a change to must come from a change in and hence, by our definition, from feature learning. This combined with the previous paragraph also shows that strong feature learning effects are beyond the practical reach of straightforward perturbation theory.
We turn to estimate the magnitude of the term we neglected namely
Noting and that, following our EoS., , we can rewrite the above term up to corrections as
which is indeed negligible compared to at large .
6 Estimating Corrections to Mean-field Results
In the mean-field derivation, when focusing on the two last layers, we neglected the term
In this section, we study the effect of this term in perturbation theory and when it can be neglected.
Specifically, we treat the above term as a perturbation over the mean-field limit and calculate its leading order effect on the mean-field average of, which coincides with the mean-field average of which is directly related to the MSE. To expose this matter in its simplest form, we shall assume that the penultimate layer is linear. Consequently and in addition, VGA becomes exact (see Eq. 33 and definition of ).
Turning to an action formulation, we focus on the following mean-field action of the and variables together with the perturbation namely,
Perturbation in the term yields the following zeroth and first-order contributions
A key point is that making smaller makes bigger. Hence, our mean-field treatment works within the feature learning regime. Indeed, the emergent scale is defined as and hence does not involve directly. Decreasing can only make larger since . Thus, decreasing increases feature learning while making mean-field corrections more negligible. On this note, we comment that in our FCN experiments we found, similarly to Ref. , that mean-field scaled FCNs had better test performance compared to those which used standard scaling similar to
Mean-Field Equations for CNNs
In this section, we provide a sketch of the derivation of the equations of state for deep CNNs highlighting the differences between CNN and FCN. For simplicity, we follow here the derivation of a three-layer CNN. The generalization to any number of layers is straightforward. The model we consider is a three-layer CNN having two activated convolutional layers and one linear readout layer. Specifically, we consider
where the matrix represents all the samples, , and the first term on the r.h.s. is given by
where is viewed now as a random variable following the CNN outputs on all different training points, and denote average over the weights . The above probability can be viewed as the prior induced on by a finite random DNN with i.i.d. Gaussian weights and variance and respectively for each layer. At such priors tend to a GP, however our interest here is at finite .
By conditioning on the pre-activation output (), Eq. (76) can be re-written as
where the hidden layers probabilities used above are
the tensor consists of all the latter being the random variables describing outputs of the second layer at hidden pixel , channel on the data-point.
Similar to the FCN, we continue our analysis by using the Fourier identity, which replaces the delta function with auxiliary fields . The resulting action () is then,
Where the “channel” post-kernels for CNN contain also summation overstrides and are defined as:
We comment that by averaging over the auxiliary fields, , one obtains the following equivalent form, containing only the pre-activations and the outputs
We continue with the auxiliary variables and derive the inter-layer mean-field decoupling. Where for CNN the number of channels plays the role of width in FCN. As for FCN, we note that in Eq. (79), the hidden layer and the output layer depend on their respective upstream layers only through the "channel" post-kernels. Performing our mean-field decoupling as in Sec. 5.2 we obtain the resulting action
where the post-kernel can now be defined as the mean-field average of the “channel” post-kernel via the above mean-field action distribution:
The correlation functions similar to the FCN are then,
The last layer of auxiliary field correlation is
this is easily derive from Eq. (82). The average of using the mean-field action is then,
Despite reducing the full system into decoupled systems per channel and layer, the resulting action for all but the top layer is still non-Gaussian. Following the justifications discussed in the main text and for the FCN, we approximate the latter using the variational Gaussian approximation (VGA). Assuming for simplicity an antisymmetric activation function such as , the CNN has an internal symmetry, making each pre-activation positive output as likely as its negative. Also, at large enough we do not expect spontaneous symmetry breaking, thus our VGA will involve only a centered Gaussian. Specifically, we denote the optimal covariance of , as the pre-kernel of the second layer and the optimal covariance of , , is a connected to the first layer pre-kernel, .
Following the analysis in subsection. 5.3 with the above modification for CNN, we obtain the following closed set of equations determining all pre-kernels as well as the outputs namely
Following the FCN case, we again examine the EoS for the penultimate layer assuming that it is linear, solve them using perturbation theory in , and use the magnitude of the correction to estimate the scale at which feature learning becomes important. For our CNNs we have that which yields the following equation of state for
Next, we perform a leading order perturbation theory in (or equivalently in the second term on the r.h.s) yielding
As justified in the FCN case, we focus on the second term on the r.h.s. and look again at yielding,
We define the ratio of the second to the first term as the emergent scale. Notably for it coincides with the definition for FCNs. In addition to considering the 2-layer CNN studied in the main text and estimating it using the same approximations, it results in the same scale.
2 Estimating Mean-field corrections - CNNs
In the FCN case, we found that an additional ingredient, on top of large , is needed for our mean-field decoupling to hold - either mean-field scaling or a target with support only on weak eigenvalues. For the CNNs we have studied, and quite possibly for a much larger family of CNNs, this additional ingredient comes naturally from the read-out layer averages over latent pixels in the penultimate layers. Since these are expected to be somewhat independent, one can hope that summing over these terms is similar to increasing the number of channels by a factor of (the number of pixels in the penultimate layer). Here we establish this more concretely. Similar to the FCN we calculate the average of the discrepancy up to the second order:
For CNN second order term can be further simplified by using Wick theorem for pre-activations of the last layer when again we consider the case of linear activation function
A similar derivation to 5.6 yields the following correction to ()
To simplify this expression we next note that for data-sets in which for each there exists a "symmetry-partner" point wherein all coordinates in the fan-in of the ’th latent pixels are flipped - the second, trace-like, term must vanish for . This is due to the fact that is invariant under the action of the associated symmetry, whereas receives a minor sign whenever . As we expect this symmetry to be approximately realized as it is a symmetry of the underlying measure from which are drawn. Following this, we remove terms from the summation.
Next we notice that due to the approximate translation symmetry of the dataset, at large
becomes independent of . We thus replace it by its average and perform the remaining summation over, which now involves only the first term to obtain
Recall that . This resulting expression is very similar to its FCN version (with standard scaling) with one crucial difference, which is the appearance of the aforementioned factor. More specifically, the first summation is smaller than the zeroth term (). The scale controlling the mean-field decoupling is therefore . Thus, at large , we can have a reliable mean-field decoupling even when .
Toy Example - One Hidden Layer
requiring enough data-points to resolve the target sets (number of target parameters) while staying within the over-parameterized regime implies . We further consider the large-scale “thermodynamic” limit, where .
Similarly to the 3 layer case, the equations of states here are given by,
These can be viewed as non-linear equations in the variables making up the symmetric matrix . Having these variables determines the variables directly.
We begin with approximating the GP inference appearing in the last equation. This can be represented as where, following the standard GP prediction formula (for the training set). At large , can be approximated using the equivalence kernel (EK) approximation together with its perturbative corrections (non-perturbative approaches in could also be considered ). Within this approximation scheme, one considers as the continuum operator , diagonalized on the data-set measure (), leading to the following formula for in the strict EK limit
where are the eigenvalues and eigenfunctions of and . To obtain an explicit formula for, , we proceed by solving the eigenvalue problem. To this end, we first consider the kernel action on a general linear function (), where are vectors of size for all , and is drawn from the dataset measure,
where is a centered Gaussian with covariance matrix . In the second transition, we undo the kernel integral, where is a Gaussian measure. We then exchange the order of integration and sum over the weights and the data. We now do the integration over the data,
where on the r.h.s. we noted that as , is weakly fluctuating and close to its mean. Next, we perform the integral which, following this mean-field replacement, is now of the same type as the previous one. Overall this yields
Using again we replace by its mean under . Following this, we obtain that the action of preserves the space of linear function. Furthermore, we see that diagonalizing in this subspace reduces to diagonalizing . We thus reach the conclusion that since is a linear function, so must be and hence .
We turn to the quantity appearing in the equation of state for . For the moment, we omit fluctuation piece of and replace it by . Below, we will show that the contribution of this fluctuation term is negligible at large, which is our focus here. Following this, we approximate the two summations with the following expression
Next, we argue that at large , where ’s are some real numbers. Indeed, at large when replacing summations by integrals and all matrices by their continuum kernels, the full symmetry of the measure from which ’s are drawn becomes manifest in the equations. The latter amounts to an independent orthogonal rotation () of each, which leaves invariant. Recalling the previous result, that is linear, together with this symmetry, implies that where ’s are some real numbers. Consequently, is of the same form.
Using the above ansatz for, we can solve for the r.h.s. of Eq. (102). To this end, we again rewrite the r.h.s. by undoing the kernel integral and exchanging the order of the integration and the sum,
Next we use again the fact that is weakly fluctuating at large to perform the remaining integration over and obtain
Here one can also see a different justification for the VGA underlying our equations of state. Indeed, Eq. (102) is exactly the non-linear term in the mean-field-decoupled-action for the input layer. At large, it leads to Eq. 103 where replacing by its expectation value makes the term quadratic. As the first term in that action is quadratic in, the overall action becomes Gaussian, as the VGA assumes.
Following the above simplification, the equation for becomes
revealing that only the eigenvector along is affected by training. At large, we may thus replace by rendering the above an explicit formula for .
Next, we return to the eigenvalue equation for (Eq. 101) and use the fact that is an eigenvalue () of
Taking next the EK predictions along with its leading correction yields
where is the posterior covariance in the EK limit given by , where are ’s eigenvalues in the limit. In practice, we estimated numerically by diagonalizing large kernels. For found and .
Notably since came out proportional to we obtained a simplified form for containing only a single free parameter ()
In addition we may now replace all the factor appearing above with .
Altogether, this yields the following non-linear equation for the scalar quantity
Solving the above equation for , one obtains and using the equations of state. Using, we can calculate the DNNs predictions on the test set using standard GP inference. We note by passing that one can also estimate the result of this GP inference using EK. However, we found that just keeping a leading perturbative correction to the EK result, as done for the training set, resulted in discrepancies when compared to exact the GP inference formula. As estimating GP inference was not a main focus of the current work, we instead opted to perform this last GP inference on the test set numerically. In principle, other analytical methods for estimating GP inference could be used here . Finally, we note that when estimating and on a specific dataset they are defined as
where and refer to samples taken from the train and test datasets, respectively.
Last we turn to discuss the fluctuation term we omitted given by
We wish to compare its contribution to that of . To this end, we write in terms of its eigenvectors () and eigenvalues ()
aiming for an order of magnitude estimation, we perform the following two approximations: First, we approximate the eigenvalue by the leading eigenvalues of the continuum kernel time . Second, we take this continuum kernel to be GP kernel. The latter is justified by the fact that the feature-learning effects we found are large, but still do not correspond to an order of magnitude change. Following this, we obtain degenerate eigenvalues equal to and the corresponding ’s span all linear functions on input space (sampled on the training set).
Next, we estimate how adding such terms to the previous computation affects the equation for . First, we note that , in our two experiments. Next, we imagine repeating the computation of the previous section, with these extra terms corresponding to the various different . Notably each such term would enter the computation in the same exact manner to (see Eq. 102) namely
the only two differences are that (i) whereas (hence a factor on the r.h.s. was lost compared to Eq. 102) and (ii) can have any dependence on and not only through . For concreteness, let us span the continuum version of these by . Summed together, all these eigenvalues will end up augmenting the r.h.s. Eq. 106 into
comparing the first and last term on the r.h.s we find it is negligible for . Notably, even for our experiment at , we find this last term is negligible.
Here, we argue that the replacement involved in Eq. (102) is valid for . We further comment on some implications this has for the fully-connected case (). To show this, we consider a summation of the form
undoing the kernel integral as we have done before (see for example Eq. (99)) yields
For simplicity, we present the analysis for , and report the results for general by symmetry. Re-focusing on the relevant summation,
we consider the average (denoted below by a bar) and variance of over the dataset, where each sample is drawn from the measure , conditioning on the values of and . This yield,
To estimate the scale of these quantities, we focus simplicity on the regime where the error function is linear. This can be generalized by taking into account perturbative corrections, but these do not change the scale. Following this approximation and taking into account centered Gaussian i.i.d. measure with variance for, one finds,
Since and are high- dimensional vectors where and can be taken to be fixed with norm in our setting. Therefore, the norm of concentrates on its average value, , which is of order one, with fluctuation of order . Hence, the variance fluctuation are of order and mean fluctuation are of order . Thus, is required for replacing the summation by an integral for (the fully connected case). Similar analysis can be done for general leading to , and which then requires , where we assumed, without loss of generality, that and , as in the previous section. For our CNN model and hence this approximation is reasonable.
Let us consider the implications this has on the fully connected case. In the regime where , the behavior of the MSE will change drastically compared to . Indeed, as our previous results show, for our CNN experiments, the train MSE (over ) reaches of the corresponding at , and for the GP this quantity is order . Thus, while being small, this MSE is far from negligible. Specifically, in our CNN experiments, the emergent scale () is order .
In contrast, at , much like in experiments , the self-consistent equation predicts a negligible GP-DNN performance gap down to or order . Moreover, for the GP (which has a uniform prior over all the possible linear functions) actually performs very well. Specifically, the EK approximation yields a train MSE (over ) of order . The emergent scale () is of the same order at .
Several conclusions could be drawn here: (i) Taking in our experiments, there is no separation of scales between the emergent scale and . (ii) In a related manner, no appreciable label/target-aware feature learning will take place down to the scale where our inter-layer mean-field breaks down ().
Validity of the Variational Gaussian Approximation
Here, we provide some analytical support to the validity of the Gaussian variational approximation, used to obtain the equation of state. We present the variational treatment from a perturbation theory approach, as a partial summation of a subset of all perturbative corrections. We identify a qualitative difference between this subset and other perturbative corrections that we neglect. We apply our analysis to one of the typical hidden layers in the mean-field limit. This is easily generalized to all layers due to the recursive structure of the problem.
In Eq. (43), we introduce the mean-field probability distribution:
where is averaging with respect to the Gaussian measures induced by the kernel . For the layer below, This yields the following self-consistent equation for the post-kernel, and the pre-kernel, :
We now show in what sense this approximation is valid, i.e. when can we approximate the mean-field distribution by a Gaussian distribution. We start by calculating the interacting Green function of the process (second moment of the process).
where is the connected expectation with respect to and is the connected expectation with respect to Here, connected mean cumulant moments w.r.t the variables and . We perform our perturbation analysis, for simplicity, for the monomial activation function , with finite, as a characteristic example. This can be generalized to other smooth activation functions. The perturbative correction term, to the free moment, is as follows:
where is due to the choice of from each monomial activation function, the is due to the arrangement in pairs of the remaining fields. We denote by . Plugging back in Eq. (126) leads to the following self-consistent equation:
Plugging the definition of the matrix V, we have that,
The second transition is using Gaussian integration by parts. The resulting equation is very similar to the mean-field equation we find. The difference is that here instead of derivative by , the pre-kernel, we have a derivative of the post-kernel. In addition, the expectation is also with respect to the free theory. Indeed, the full variational treatment is self-consistent or, equivalently stated, it takes into account a larger set of diagrams (terms in perturbation theory) which amount to renormalizing the 2-point function from to . We argue however that doing so only improves the overall accuracy. Indeed, the expansion of Eq. (128), is the same as one would get from a Gaussian action consisting of plus a quadratic term of the form . The variational Gaussian approximation essentially finds the closest Gaussian distribution. Therefore, it can only improve upon this simpler approximation we took here.
Variational Gaussian Approximation for ReLU Activation
Here, we extend the previous VGA treatment, which assumed centered distributions, to non-centered ones. Indeed, for antisymmetric activation functions, the pre-activations appear schematically as , thus . Since appears only in quadratic order in the action, we find that is as likely as and its ensemble average is strictly zero. However, for , it is not the case. This requires us to extend the previous treatment by including extra variational parameters for the mean.
Concretely, let us focus on the VGA for the input layer of 3 layers, CNN. Our VGA for the probability is now defined by the variance of the Gaussian in weight space, as well as the mean . Repeating the previous analysis one obtains
where are the weights of channel, and is now defined by
which one can reduce to a one-dimensional integral following Ref. . Obtaining an explicit expression, or potentially a perturbation expansion in , is left for future work.
Taking the derivative of the KL-divergence with respect to and equating it to zero, one obtains
Similarly as one obtains the additional equation
where is the average of (for any ).
Further details on the numerical experiments
Here we report on several additional numerical results for FCNs. In particular, we provide details about the full spectrum of , the equilibration process, and the Gaussianity measures. We also report on some numerical experiments with FCN in the standard scaling.
We conducted further experiments with the above FCNs (, teacher-student with for the teacher) however with standard scaling (, regardless of ) rather than "MF" scaling (). Here, we expect feature learning to diminish . Considering the numerical solution of our EoS, we found that they essentially remain close to the GP limit for . Specifically, at leading eigenvalue came out , whereas in the GP limit we obtain . This is consistent with the fact that the emergent scale here (Figure 1. panel (b) main text) is of the order of . Figure 7 and Figure 8 presents the results for different width .
2 2-layer CNN experiment
Figure 9 shows the top eigenvalues normalized by , this is a complementary figure to Figure 1(b) introduced in the main text. Clearly, as the width increase, the eigenvalues of are getting closer to the GP kennel eigenvalues.
3 Myrtle-5 CNN on subsets of CIFAR-10 experiment
Here, we report the statistics of the pre-activations in Fourier space similar to Figure 6 in the main text, but now for all layers and projected on the 1st, 3rd and 10th eigenvectors of the Fourier space covariance matrix. Let us begin by giving some more details on the procedure for deriving these quantities. The pre-activations at some layer is of the form where is a data point index, is a channel index, and are pixel coordinates. We choose some wavenumber and transform these to Fourier space to yield which summarizes contributions from all pixels. We then compute the covariance matrix of these (an matrix), averaging across channels and seeds. Finally, we project on some eigenvector of the covariance matrix, and these are the quantities whose statistics we report in figures 10, 11.
There are several empirical observations to be made here:
Gaussianity generally increases as we go deeper into the network from the input to the output.
Gaussianity generally increases as we project on higher index eigenvectors (going from left to right across the columns).
Deviations from Gaussianity can appear in several ways: e.g. as multi-modality (e.g. top left panel), or as excessive kurtosis (e.g. 2nd-row left column).