Contrasting random and learned features in deep Bayesian linear regression
Jacob A. Zavatone-Veth, William L. Tong, Cengiz Pehlevan
I Introduction
Deep neural networks (NNs) display a rich and often-perplexing spectrum of generalization behaviors. Highly overparameterized NNs may possess the expressivity to fit random noise, yet in practice can still generalize well to unseen data . The ability of NNs to flexibly learn features from data is widely believed to be a critical contributor to their practical success , but the precise contributions of feature learning to their generalization behavior remain incompletely understood .
In recent years, intensive theoretical work has begun to elucidate the properties of deep networks in the limit of infinite hidden layer width. In this limit, a dramatic simplification occurs, and inference in deep networks is equivalent to kernel regression or classification . This correspondence has enabled detailed characterizations of inference at infinite width in both maximum-likelihood and fully Bayesian settings, providing new insights into the inductive biases that allow deep networks to overfit benignly . Yet, understanding inference in the kernel limit is not sufficient, because kernel descriptions cannot capture feature learning .
As a result, a growing number of recent works have aimed to study the behavior of networks near the kernel limit, with the hope that leading-order corrections to the large-width behavior might elucidate how width and depth affect inference . Some of these works focus on the properties of the function-space prior distribution , some consider maximum-likelihood inference with gradient descent , and some consider properties of the full Bayes posterior . This body of research has resulted in several conjectural conditions under which when narrower and deeper networks might perform better than their infinitely-wide cousins in the Bayesian setting, as measured by generalization for fixed data or by some alternative criterion based on entropic considerations .
However, previous studies of Bayesian neural network generalization near the kernel limit have not clearly differentiated the effect of width on feature learning from its other potential effects on inference. Concretely, it is not clear whether potential improvements in generalization afforded by the leading finite-width correction reflect the benefits of feature learning, or if a similar gain would be observed in random feature models, where only the readout layer is trained. Here, we explore how random and learned features affect generalization in the simplest class of Bayesian NNs—deep linear models—when trained on unstructured, noisy data. By developing a detailed understanding of this simple setting, one might hope to gain intuition that may prove useful in studying more complex networks .
In this work, we study the asymptotic generalization performance of deep linear Bayesian regression for data generated with an isotropic Gaussian covariate model. Using the replica trick , we compute learning curves for simple linear regression, deep linear Gaussian random feature (RF) models, and deep linear NNs. Our results are obtained using an isotropic Gaussian likelihood in the limit of small likelihood variance, which renders this analysis analytically tractable . Using alternative replica-free methods and numerical simulation, we show that the predictions obtained under a replica-symmetric (RS) Ansatz are accurate for all three model classes. In particular, the RS result for learning curves of NNs with hidden layers of equal widths is consistent with results obtained by Li and Sompolinsky using a different approximation method.
In the presence of label noise, both RF and NN models display sample-wise non-monotonicity in their learning curves. As we work in a high-dimensional limit, this non-monontonicity is of a particularly extreme form: the generalization error diverges at a particular data density. In keeping with modern deep learning parlance, we refer to this behavior as “double-descent,” though this monotonicity can arise from distinct effects in different settings . If one introduces a bottleneck layer that is narrower than the input dimension, an RF model will display model-wise double-descent behavior at fixed data density—or equivalently sample-wise double-descent at fixed width—even in the absence of label noise, while an NN model will not show this divergence. This distinct small-width behavior shows one advantage afforded by the flexibility to learn features. For both models, we analyze how optimal network architecture depends on data density and prior mismatch. We show that, at a given data density, RF models have a particular optimal width for fixed depth and optimal depth for fixed width that minimizes the generalization error. In contrast, it is always optimal to take an NN to be as wide or as narrow as possible, depending on the regime.
We further analyze models of arbitrary depth perturbatively in the limit in which the network depth and dataset size are small relative to the hidden layer widths, connecting these results to those of previous work on fixed-dataset perturbation theory . We find that the leading order correction to the large-width behavior of RF and NN models is identical, hence first-order perturbation theory for the generalization error cannot distinguish between random and learned features. To distinguish between training only the readout layer and training all layers, one must go to second order in perturbation theory. Therefore, at large widths, the ability to perform representation learning provides only a small advantage in generalization performance in these simple models relative to random features, which is invisible in first-order perturbation theory. In total, our results provide new insight into how the generalization behavior of deep Bayesian linear regression in high dimensions depends on architectural details. Moreover, they shed light onto which qualitative features of generalization behavior can or cannot be captured by low-order perturbative corrections .
II Problem setting
In this section, we introduce the three classes of regression models we consider in this work, as well as our generative data model. Our notation throughout is standard; we use to denote the Euclidean norm, to denote the identity matrix, and to denote the vector with all elements equal to one.
In this work, we consider three classes of scalar Bayesian linear regression models for a scalar-valued function of -dimensional inputs. All three of these model classes are of the form
Below, we list the three classes of models we consider, and introduce a two-letter abbreviation for each:
Simple Bayesian linear regression. For this model, the end-to-end weight vector is directly parameterized as
Previous works have extensively studied this model in both maximum-likelihood and fully Bayesian settings , hence we include it as a baseline against which we will compare our results for more complicated models.
Deep Bayesian random feature models. For these models, the weight vector is parameterized as
while the hidden layer weights are drawn from a fixed isotropic Gaussian distribution
Deep Bayesian linear neural networks. For these models, the weight vector is parameterized as
Though NNs are parameterized identically to the RF models above, they differ in that all of the weights are trainable, not only the readout. We again choose isotropic Gaussian prior distributions
From a physical perspective, the hidden layer weights in the RF model are ‘quenched’ disorder, whereas they are ‘annealed’ disorder in NNs .
II.2 Data model and the Bayes posterior
We train all models on a dataset of examples, generated according to a standard isotropic Gaussian covariate model . In this model, the example inputs are independent and identically distributed samples from a standard Gaussian distribution:
while the labels are generated by a ground truth linear model, possibly corrupted by additive Gaussian noise:
where sets the noise variance. The noise variables are independent and identically distributed as
For a dataset thusly generated, we introduce an isotropic Gaussian likelihood of variance :
where denotes the set of trainable parameters for a given model, and the normalization constant is implied. We will refer to as the ‘inverse temperature’ by standard analogy with statistical mechanics . Then, the partition function of the resulting Bayes posterior is given as
We denote expectations with respect to this Bayes posterior by .
II.3 Generalization error in the thermodynamic limit
Moreover, we focus on the zero-temperature limit , in which the likelihood tends to a constraint that the network interpolates the training set with probability one. In the noise-free case, this limiting likelihood is matched to the true generative model of the data, but it is clearly mismatched in the presence of label noise. This limit has been considered in several recent studies of deep linear Bayesian neural networks .
Our goal is to study the average-case generalization error of the resulting model, as measured by the deviation of its end-to-end weight vector from the true teacher weight vector :
We remark that (17) is the average-case error of the Gibbs estimator (i.e., a single sample from the posterior); one could instead consider the error of the mean estimator . For the LR and RF models, this corresponds to studying Bayesian minimum mean-squared error (MMSE) inference . As one has the thermal bias-variance decomposition
our results include an additional contribution to the generalization error from the posterior covariance of the end-to-end weight vector, which is not identically zero. If one considered an alternative low-temperature limit in which the prior variance is proportional to , then this additional contribution would vanish in the low-temperature limit. Our choice of scaling is motivated by the considerations described in our previous work , and is the one classically used in studies of the statistical mechanics of Bayesian inference . This choice is important as it affects the relationship of our results to those in the setting of ridge regression. As discussed in Appendices C and D, in our previous work , and in previous works of Advani and Ganguli and Barbier et al. , the zero-temperature limit of the MMSE estimator would in this case coincide with the ridge regression estimator.
We compute the limiting average generalization error using the replica method, a non-rigorous but powerful heuristic that has seen broad use in statistical mechanical studies of inference . As our main results can be understood independently of calculation through which they were obtained, we relegate the details to Appendices A and B. We note the important caveat that our main results are obtained under a replica-symmetric Ansatz. We expect this assumption to hold exactly for the LR and RF models by virtue of the concavity of their log-posteriors, but replica symmetry may be broken in deep linear NNs . We will not address this possibility analytically by considering Ansätze with broken replica symmetry , but will instead simply compare the RS predictions against results obtained through a combination of alternative analytical methods and numerics.
III Learning curves for the LR model
We begin by briefly describing the learning curve of the simple LR model. Our result extends the classic result of Krogh and Hertz for ridge regression in the ridgeless limit to the Bayesian setting:
For this simple model, the learning curve can also be computed directly by first evaluating the posterior average defining for a fixed realization of the disorder, and then averaging the result over the disorder in the zero-temperature limit (see Appendix C for details). The result of can be recovered from (19) by setting . We provide further discussion of the relationship between the Bayesian LR model in the zero-temperature limit and ridge regression in the ridgeless limit in Appendix C.
Therefore, as found in the ridge regression setting, the LR model exhibits sample-wise double-descent behavior—i.e., non-monotonicity in as a function of —in the presence of label noise. In the thermodynamic limit, the double-descent behavior is particularly striking: diverges as . In the absence of noise, decreases monotonically from to as , and then remains at zero for all . We remark that, for this and subsequent models, we will not conduct a detailed analysis of what happens precisely at exceptional points, e.g., . In the ridge regression setting, the phase transition at has recently been analyzed in detail by Canatar et al. . We also direct the interested reader to an expository note by Nakkiran for further intuitions on double-descent in ridge regression, and to work by Hastie et al. for a detailed rigorous analysis. We will take the model-wise double-descent behavior of the LR model as a benchmark for our subsequent analyses of the more complex RF and NN models.
IV Learning curves for the RF model
We validate the accuracy of this RS result by comparing it against the result of an alternative semi-analytical approach. As shown in Appendix C, the zero-temperature posterior average in (17) can be computed for a fixed realization of the disorder. Even without explicitly evaluating the disorder average, this shows that the RF model should display the three phases indicated by the RS result (20), and confirms the prediction for in which of the phases the learning curve should depend on the prior variance (see Appendix C). Importantly, the RF model learning curve (17) does not depend on the ordering of the hidden layer widths, which follows from the fact that the random Gaussian hidden layer weight matrices weakly commute . For conceptual clarity, we therefore refer to the cases in which different hidden layers are the narrowest as a single phase. To quantitatively test the accuracy of the RS result, the disorder average can be evaluated numerically using sampling (see Appendix G). As shown in Figures 1 and 2, we observe excellent agreement over a broad range of parameter values. These results are consistent with our expectation that the RS Ansatz should yield accurate results for the RF models .
While the LR model only exhibits double-descent behavior in the presence of label noise (19), the RF model can also exhibit double-descent behavior in the absence of label noise if any one of the hidden layers is narrower than the input dimension, i.e., . This phenomenon occurs in a model-wise fashion at fixed data density: if one considers a decreasing sequence of widths at fixed , will diverge as (Figure 1a,c,e). Equivalently, this divergence can be observed in a sample-wise fashion at fixed width, with as . Moreover, as illustrated in Figure 2, it is determined by the width of the narrowest hidden layer. If one adds more bottleneck layers, then the expression for the generalization error in the regime will formally include more poles (20), but these poles will not be visible as one varies the size of the training set or the width of the narrowest bottleneck.
IV.2 Large-width behavior
where denotes terms that include two or more factors of any combination of the layer widths.
where we have defined the re-scaled prior variance
Then, for , we can read off the full series expansion using the binomial theorem and the geometric series:
IV.3 Optimal width and depth
which is consistent with the result for optimal width at fixed depth given in (25). Therefore, much like we found in our analysis of optimal width, the optimal depth of an RF model is related to the match between the scale of the prior and of the target. This behavior is illustrated in Figure 3.
V Learning curves for the NN model
For the NN model, we do not obtain a simple closed-form solution for the RS learning curve at general depth. As shown in Appendix B.3, we find that the solution is of the form
We defer more detailed discussion of which root should be selected to Appendix B.3, where we show that one required condition on the solution is that
The special case of (28) for networks with hidden layers of equal widths follows from results obtained through a rather different approach in a recent study by Li and Sompolinsky . Concretely, they use an iterative saddle-point argument to approximate the posterior expectation in (17) for fixed data, and then apply that result to a random Gaussian covariate model under what amounts to the assumption that the quantity concentrates rapidly. In Appendix D, we provide a detailed discussion of the mapping between the polynomial condition in terms of which their result is expressed and the RS condition (29). In Appendix D, we also use a finite-size fixed-data approach derived from our previous work to show that the learning curve should be of the form (28). Concretely, this approach gives an expression for as the thermodynamic limit of a dataset average of a ratio of prior averages, with the remaining components of the learning curve exactly matching the RS prediction. Taken together, these result suggests that the RS prediction for the learning curve correctly captures at least the coarse behavior of generalization in NNs.
To further probe whether the RS prediction is quantitatively accurate, we evaluate the finite-size data average numerically. As shown in Figures 4 and 5, and in supplemental figures provided in Appendix G, we observe good agreement for two-layer networks. To probe the accuracy of the RS prediction for deeper networks, we solve the polynomial (29) numerically. As shown in Figures 4 and 5, we again observe good agreement. Therefore, both alternative heuristic analytical approaches and numerical results are consistent with the RS learning curve, suggesting that it provides a reasonably accurate picture of generalization in deep NNs.
Like the previously-studied models, we see that label noise can induce sample-wise double-descent, with as (Figure 4). However, unlike for the RF model, having relatively narrow hidden layers does not introduce the possibility of divergences other than at , as should remain bounded. This is illustrated in Figure 5, where we repeat the analysis of Figure 2, but do not observe similar model-wise divergences. Moreover, this means that the NN model does not display sample-wise divergences in the absence of label noise. Therefore, training the hidden layers affords the advantage of avoiding the possible model- and sample-wise divergences that can arise in RF models with narrow bottlenecks. This sharp contrast makes sense, since in the RF model the presence of layers width introduces a true bottleneck, while in the NN model one could in principle find a solution where, in all layers except the first, exactly one weight is nonzero, and the model essentially reduces to shallow linear regression. The existence of this solution reflects the fact that, from the standpoint of expressivity, NN models should be able to perform as well as LR models, and differences in performance reflect the behavior of the inference algorithm . Indeed, if and , we have the solution , and . Therefore, in this special case, the RS result predicts that depth has no effect on generalization performance. This behavior is clearly illustrated by Figure 5, where the generalization error of a three-layer NN remains constant as the widths of the two hidden layers are varied. Even at non-zero noise levels, Figure 4 illustrates that width has a relatively minimal effect of generalization performance when .
V.2 Large-width behavior
This limiting result has several interesting features. First, paralleling our analysis of the RF model at large widths, the closeness of the NN model’s learning curve to that of simple linear regression is determined by a combination of depth, dataset size and width. Second, not only do the RS learning curves for NN and RF models agree at infinite width, but the leading order corrections agree (i.e., the term that is linear in ; see (21)). Thus, if one tracked only the generalization error, one could not differentiate between training only the readout layer and training all of the layers simply by considering the leading order perturbative correction. One could of course distinguish between these two models by considering leading-order corrections to observables that explicitly measure task-relevant feature learning in early hidden layers, such as the kernels considered in our previous work .
V.3 Generalization gap between RF and NN models
Therefore, the next-to-leading order correction can distinguish between RF and NN models. Moreover, the gap in the generalization performance of the two models is, to the given order,
The coefficient of the leading term is always positive, hence at very large widths training both layers should produce a small benefit relative to simply training the readout. In the two-layer case, one can use the closed-form solution for the RS generalization error to show that the generalization gap is strictly positive, except at vanishing load or in the limit (see Appendix F.4). These results suggest that training all layers of a deep linear network can yield improved generalization relative to training only the last layer, even if the widths are large enough such that the RF model does not display double-descent in the absence of noise. See Figure 6 for an illustration of this behavior.
V.4 Optimal width and depth
VI Discussion and conclusions
In this work, we studied the statistical mechanics of inference in deep Bayesian linear models. We characterized the learning curves of deep linear random feature models and deep linear neural networks for isotropic Gaussian covariates, using a combination of the replica trick and replica-free methods. Our primary results for how deep Bayesian linear models with random and learned features differ or resemble may be summarized as follows:
In the presence of label noise, both RF and NN models display sample-wise double-descent (Figures 1 and 4). For RF models, the presence of a bottleneck layer with width less than the input dimension induces model-wise double-descent at fixed dataset size and sample-wise double descent at fixed width (Figures 1 and 2), while bottlenecks do not affect the double-descent behavior of NN models (Figures 4 and 5). In particular, NN models do not display model-wise double-descent, and do not display sample-wise double-descent in the absence of label noise.
For both RF and NN models, the effect of width on generalization depends on the match between the prior variance and the true scale of the targets, with wider networks yielding better generalization when the prior variance is less than the average target scale (Figures 3 and 7). For NN models, taking the network to be as wide or as narrow as possible is always optimal. In contrast, when the prior variance is greater than the average target scale, there is a particular width that yields optimal generalization in RF models.
Similarly, the optimal depth for both models depends on prior-target mismatch. Paralleling the case of optimal width, deeper models always perform worse when the prior variance is less than the average target scale (Figures 3 and 7). When prior variance is greater than the average target scale, shallower models perform better. In this regime, as in the case of optimal width, there is a particular depth that yields optimal RF model generalization for fixed width, prior variance, and data density.
VI.2 Prior work
Double-descent phenomena have recently garnered significant interest in deep learning . In high-dimensional random feature models like those considered in §IV, divergences in the generalization error can arise through interactions between randomness in the features and randomness in the training data . Moreover, as noted in our discussion of simple linear regression models in §III, divergences can arise in models without additional feature randomness—including kernel regressors with deterministic nonlinear features—due to overfitting of noisy labels . Disentangling the causes of non-monotonicitic generalization performance observed in experimental settings for realistic data models remains an interesting subject for further study .
The statistical mechanics of inference in shallow linear models with more general priors and likelihoods was investigated in detail by Advani and Ganguli , who showed a correspondence between the performance of Bayesian MMSE inference and a class of algorithms known as M-estimators. The effect of prior mismatch on the performance of the shallow MMSE estimator has also been considered in recent rigorous work by Barbier et al. . However, neither of these works considered the effect of depth on inference.
This regime has thus far proven challenging to access perturbatively, as large deviations from the kernel limit may emerge . Existing fixed-data approaches to regimes in which either the dataset size or the output dimension is not negligible relative to hidden layer width rely on saddle-point approximations that may break down when both of these parameters are large. New approaches will therefore be required to study networks in this limit non-perturbatively. With such results in hand, it will be interesting to test whether existing perturbative predictions do in fact capture qualitative features of how generalization depends on network architecture and other hyperparameters. For the simple models considered here, we found that small-sample-size perturbation theory does in fact yield largely correct predictions for when wider networks generalize better, even at large sample size.
VI.3 Outlook
We conclude by noting that our work has several important limitations, which will be interesting to address in future work. First, our approach is highly specialized to deep linear networks, and would not extend easily to nonlinear models. Though the utility of linear networks as a model system for studying the effect of depth on inference has been clearly established , rigorous characterization of the effect of nonlinearity on inference in deep Bayesian neural networks remains a largely open problem . Second, we have assumed that the covariates are drawn from an isotropic Gaussian distribution. Though this is a standard generative model in theoretical studies of inference , it is undoubtedly not reflective of real-world data. Extending results of this form to more realistic generative models will be an interesting objective for future work . We remark that some of the fixed-data results of Appendices C and D would extend immediately to anisotropic and non-Gaussian data provided that the requisite invertibility conditions hold. While BNNs are finding practical applications in physics and elsewhere , another important direction for future work will be to develop a rigorous theoretical understanding of how results on the generalization performance of BNNs, like those obtained here, relate to the generalization performance of networks trained with stochastic gradient-based algorithms, a link that remains incompletely understood . Finally, our replica theory approach is of course non-rigorous. For the RF model, we do not expect replica symmetry to be broken, and conjecture that our results might be rigorously justifiable . Moreover, our replica-free analytical approaches and numerical experiments suggest that our RS results for NNs are at the very least a reasonable approximation for their true generalization performance. With that in mind, careful exploration of the possibility of replica symmetry breaking will be an interesting topic for further investigation.
Note added. Following the appearance of our work in preprint form, results on the behavior of the ridge regression estimator—which in this case would coincide with the limiting MMSE estimator—for a model with a single layer of Gaussian linear random features were announced by Rocks and Mehta .
Appendix A Replica theory framework
In this appendix, we introduce the replica theory framework we use to compute learning curves. We direct the interested reader to for more details on replica theory. Our starting point is the partition function of the Bayes posterior:
In the limit of interest, we expect the quenched free energy
We first integrate out the data. Introducing replicas indexed by , the object of interest is the disorder-averaged replicated partition function:
where denotes the end-to-end weight vector with appropriate replica indices for a given model. Using the fact that the training examples are independent and identically distributed, we have
where we have defined the overlap matrix
Enforcing the definition of the order parameter matrix by introducing corresponding Lagrange multipliers , we therefore have
Here, the integrals over and are taken over the spaces of real and imaginary symmetric matrices, respectively. Our remaining task is to integrate out the weights, which we will do for each of the three models of interest in the following sections.
A.2 Integrating out the weights for the LR model
Using the assumption that and defining
we can write the averaged replicated partition function of a single-layer network as
A.3 Integrating out the weights for the RF model
via Fourier representations of the Dirac distribution with corresponding Lagrange multipliers , we can integrate out , yielding
It is easy to see that we can iterate this procedure forward through the network, introducing order parameters
along with corresponding Lagrange multipliers, yielding
Then, using the assumption that and defining
we can write the averaged replicated partition function as
We note that we have intentionally split the entropic contribution to the replica free energy into two pieces. The first, , reduces to the entropic contribution for simple linear regression upon fixing . The second, , captures the effect of depth.
A.4 Integrating out the weights for the NN model
As we did for the RF model, we start by integrating out . We introduce analogous order parameters
However, importantly, as the weights are annealed rather than quenched, the covariance of is replica-diagonal. For clarity of notation, we define the diagonal matrix
Then, introducing a corresponding diagonal matrix of Lagrange multipliers , we have
upon integrating out . We can see that we can follow much the same procedure to integrate out the remaining weights as we did for the random feature model, except for the fact that we only consider the replica-diagonal component of the overlaps, which are re-defined to include the replica indices of the hidden layer weights, i.e.,
where is the same as for the random feature model and the matrices and are constrained to be replica-diagonal. This difference reflects the fact that the hidden layer weights of the NN model are annealed, rather than being quenched as in the RF model.
A.5 The replica-symmetric Ansatz
In the thermodynamic limit, we evaluate the integral over the appropriate order parameters and Lagrange multipliers for each model via the method of steepest descent. Importantly, we note that the diagonal components of the order parameters give the posterior-averaged generalization errors of the replicas, as in the thermodynamic limit the mean value of these parameters is given by the saddle-point equations. Our eventual objective is therefore simply to evaluate the saddle-point values of in the zero-temperature limit, and we will not consider the resulting values of the free energy.
As usual in the replica method, we seek extrema in the limit of a constrained form, known as the replica-symmetric (RS) Ansatz . For all three models, the RS Ansatz for the variables and is simply
For a deep random feature model, the RS Ansatz for the remaining order parameters is
while, for a deep network, the RS Ansatz for the remaining order parameters is
as we consider only the replica-diagonal components of the overlaps . With this Ansatz, one can simplify the expressions for the free energy and the saddle-point equations in the limit . These manipulations are standard exercises using the properties of RS matrices , hence we will only report the results (in Appendix B).
We remark briefly on the conditions under which the RS order parameters make physical sense given their definitions. We must have and for all , as these quantities are the squares of norms of vectors. If for any , the norm of the end-to-end weight vector tends to zero, and we must have a trivial solution with . We must also have and for all . Moreover, if (respectively for some ), then we must have (respectively ), to obtain a nontrivial physical solution.
Appendix B Solution of the replica-symmetric saddle point equations
In this appendix, we analyze the replica-symmetric saddle point equations in the zero-temperature limit.
For simple linear regression, the RS saddle point is given by a 4-dimensional system of equations, which decouples into a two-dimensional nonlinear system for the replica-nonuniform components and ,
and a linear system for the replica-uniform components and :
Using the expression for as a function of , we obtain the quadratic condition
The two solutions to this quadratic equation have zero-temperature limits
For , is negative, and is therefore unphysical. If ,
These are the two low-temperature scalings we would expect to be self-consistent given the saddle point equations; we could alternatively derive the above solutions by assuming these scalings for .
Considering the replica-uniform components, we use the expressions for and as functions of and to write
For the solution with , we then have
hence, recalling that this scaling yields and is valid for ,
For the solution with with , we have
hence, recalling that this scaling yields and is valid for ,
Combining these results, we obtain a zero-temperature solution which gives the result for reported in the main text.
B.2 RF model
We can then solve the remaining equations for the replica-uniform components,
We first consider the replica nonuniform components. We start by noting that the equations for and yield
hence the equation for yields
With this observation in mind, we will eliminate the Lagrange multipliers . Formally defining for convenience, we have
which yields a self-consistent equation for
As in the single-layer case, it can be seen that the self-consistent scalings for in the zero-temperature limit are and . If we take with , we simply have , which gives
for all . As physical solutions have and all , this solution is sensible in the regime .
If we take for , we have
For the zeroth solution with , we have , and thus
This solution is therefore physical for and all . For the -th such solution, we have , hence
Thus, we have for all . For , , and we must have for all such that . Moreover, we must have , such that . We will obtain further conditions on the validity of these solutions from solving for the replica-uniform components.
B.2.2 Solving for the replica-uniform components
We now consider the linear system of equations (96) that determines the replica-uniform components in terms of the non-uniform components. We start by noting that we can express a function of alone, eliminating the dependence on using the equation for :
where we have used the fact that .
Similarly, we can simplify the initial difference condition to
hence, substituting in , we have
To simplify our remaining task, we define new variables such that
Given a solution to the recurrence for the variables , we then have a closed equation for :
With this solution in hand, we can then obtain via the relation .
We now consider the zero-temperature limits of interest. With , we have . The limiting equation for is then
Considering the recurrence for , we have
which can easily be iterated backward, yielding
hence, iterating one step backwards, we find that
It is now easy to see that we can iterate this process backwards, yielding
in particular. Using the initial difference condition to express in terms of , we then obtain a closed equation for :
As this equation is linear, it is easy to solve, yielding
under the assumption that for all . Then, we have
in terms of the solution for . Then, in terms of these solutions for , we have
We now consider the solutions with for . We first consider the solution with
For this solution, , and the limiting self-consistent equation for reduces to
for any . This is non-negative throughout the expected region of physical validity (), and therefore the overall solution makes sense given that in this regime. The recurrence for simplifies to
This yields a self-consistent equation for , which gives
for all , where the empty product is interpreted as unity. These results are positive throughout the region of physical validity we expect from our analysis of the replica-nonuniform components; recalling that for these solutions, no further conditions are imposed.
Iterating one step backward, we find that
It is then easy to see that we can iterate further back to obtain, for ,
hence the initial difference equation yields
For , this yields the solution
By the same reasoning as in our analysis of the case , we have the limiting closed set of equations
Using the fact that , we have
Then, for , we have
using the solution for obtained above. As for , we must have for in order for these solutions to be physical, hence we conclude that we must have for all . As for , we will not obtain further conditions on the physical validity of these solutions by solving for for , hence we will not attempt to do so.
B.3 NN model
where, as before, we have defined and for brevity. Unlike for the RF model, in this case the replica-uniform and replica-nonuniform components do not decouple nicely. However, we have fewer equations to solve. Moreover, we can exclude solutions with , as they will be trivial.
where we have noted that .
We now seek to eliminate the Lagrange multipliers and all of the order parameters except for . To do so, we will follow our earlier analysis of the RF model. We define
which will allow us to close the equations.
Using the abovementioned fact that we can write
we can see that this set of equations is analogous to what we obtained for the deviations from uniformity and in the RF case (with, in that case, ). Thus, using the results of our previous calculation, we conclude that
which gives us closed set of equations for , , , , and .
With the scaling , we have , and the condition on becomes
To determine the limiting condition on , we note that
Therefore, we have the polynomial condition
in order for it to be a nontrivial physical solution. This implies that we must have
For any , if , and if . However, noting that
the solution yields a non-negative value for , and is therefore physical for , while the solution yields a non-positive value, and is therefore unphysical.
order-by-order in with the Ansatz
It is easy to see that the zeroth-order condition yields
In this simplified setting, it is relatively straightforward to work out by hand or with the aid of Mathematica that
For solutions with and , we have the limiting equation
As , we must have , hence this solution makes sense for all . To solve for for these solutions, it is most convenient to express in terms of . Noting that
Given a candidate positive solution for , we can then determine for all via
for all (including ) in order for the candidate solution to be physical and non-trivial. As we are interested only in learning curves, we will not analyze this equation further.
Appendix C Direct computation of posterior expectations for LR and RF models
For the LR and RF models, we can evaluate the zero-temperature posterior expectation in the definition of analytically. In particular, writing
for brevity, we have the posterior moment generating function for :
where the implied constants of proportionality are independent of the source . Then, as . the posterior mean and covariance of the end-to-end weight vector are given as
respectively. We note that is simply the RF ridge regression estimator with ridge parameter , as
The thermal bias-variance decomposition of the zero-temperature generalization error is then given as
Under the stated assumptions, the matrices and are invertible with probability one, as is their product. Then, we have
We observe that, under the re-scaling of the feature map for any , is always constant, while is either identically zero or degree-two homogeneous in . This suggests that we should be able to read off the ridgeless results from our Bayesian replica results. We also note that we have recovered the three-region phase diagram indicated by our replica calculation.
For completeness, we also remark that we can use these results to directly compute without the use of the replica trick. For the LR model, we have
in this regime. This recovers the result of our replica computation.
We remark that a similar, albeit more complex, procedure would likely allow one to derive the learning curve for a deep RF model rigorously using properties of products of large Gaussian random matrices . However, from a physical perspective, the non-rigorous replica theory approach used here has the advantages of being more transparent and of allowing a relatively unified treatment of NN models.
Appendix D Direct computation of posterior expectations for NN models
In this appendix, we show that the zero-temperature posterior expectation in the definition of can be evaluated semi-analytically for NNs. Our approach mirrors that of our previous work in : we will integrate out the weights of the first hidden layer () exactly, yielding expressions for the posterior mean and variance of the end-to-end weight vector in terms of expectations over the remaining weights. These results follow by applying the results of to a test dataset of examples with trivial data matrix and then passing to the zero-temperature limit, but we will provide a detailed derivation for completeness.
for brevity, such that , we can write the posterior moment generating function of as
where we discard irrelevant constants of proportionality. This matrix Gaussian integral can be conveniently evaluated through vectorization . Using standard properties of the Kronecker product, we find that
By varying this result with respect to the source, we thus obtain
for brevity. This matches the result of applying ’s expressions to a trivial dataset with .
As for the RF model, we introduce a thermal bias-variance decomposition
For any set of hidden layer widths, is almost surely positive. Therefore, as the only matrix inverses present in these expressions are of the form , the NN model should have two phases: and .
If , then the matrix is invertible with probability one. Then, up to (divergent) multiplicative constants which will cancel in the ratios of expectations, we have the almost-sure pointwise limit
Similarly, we have the almost-sure limits
Therefore, noting that is almost surely a constant function of , we have
If , then the matrix is invertible with probability zero, but the matrix is invertible with probability one. By the Weinstein-Aronzjan identity,
hence the determinant factors in will yield a factor of in the zero-temperature limit. We must be more careful in considering the exponential term in . Letting the orthonormal eigendecomposition of be
we have the low-temperature Neumann series
hence the divergent null-space projector term does not depend on , and will therefore cancel in the ratio of expectations. Thus, we have
By a simple application of the push-through identity, we have the almost-sure pointwise limits
Therefore, noting that is once again almost surely a constant function of , we have
Comparing this result to the discussion of the LR model in Appendix C, we can see that the bias terms in each phase are identical to those for the LR model, hence we can apply the results given there for their dataset averages. This shows that the learning curve for the NN model is of the form (28), with
We remark that evaluation of the outer dataset average without resorting to the replica trick seems likely to be challenging.
For a network with a single hidden layer, we can evaluate the average over analytically. In this case, we have
for brevity. Using the fact that under the prior, we have
where is a modified Bessel function of the second kind . Thus, for a NN with a single hidden layer, we have
rapidly concentrates about its limiting mean value, which is
Then, by equations (26) and (27) of the main text of , or by equations (24), (26), and (30) of their Supplemental Material, the result of their approximation is that
with . We would like to show that this is consistent with the result of our RS calculation, which implies that should be a non-negative root of
which recovers Li and Sompolinsky’s result.
Appendix E Comparison to large-width perturbative calculations with fixed data
We remark that the results of show that it should be safe to interchange the limit with the high-dimensional limit and the expectation over data. Using the fact that is invertible with probability one in this regime, the disorder average of the thermal bias term yields
where we have again used the formula for the expectation of an inverse Wishart matrix with identity scale matrix . Similarly, the disorder average of the thermal variance term yields
Therefore, the perturbative result of implies a large-width disorder-averaged generalization error of
which agrees with the leading-order large-width solution of the RS result reported here. This makes sense, as we intuitively expect possible RSB effects to emerge at smaller width. Moreover, we remark that we have an exact correspondence between and and the averages of the thermal variance and bias terms, respectively. Finally, we note that the coefficient of the correction, which is the dataset average of , gives the dataset average of the condition for when increasing width helps generalization noted by .
Appendix F Detailed analysis of optimal network architecture
In this appendix, we provide a detailed analysis of how width and depth affect generalization in RF and NN models.
We first consider optimizing the width of a deep random feature model. In the regime , we have
For , , while for , . This point is therefore a minimum of .
To check whether this is indeed a local minimum, we compute the Hessian of at the stationary point, which is given by
Diagonalizing this matrix is trivial, yielding eigenvalue
with multiplicity one. Both of these eigenvalues are positive throughout the parameter region of interest, confirming that the Hessian is positive-definite at the stationary point. Moreover, substituting into the generalization error , we have
For RF models in the regime for , it is easy to see that is a monotonically decreasing function of if , while if , it is a monotonically increasing function of . However, in this regime, it is important to keep track of crossings in the ordering of different layer widths.
F.2 Optimal depth for RF models
which is bounded as in the regime of interest. This gives
Using the lower bound , which is strict for all , we have
F.3 Optimal width for NN models
For a two-layer NN in the regime , we have
F.4 Difference in generalization in two-layer NN and RF models
For two-layer networks, we have the RF-NN generalization gap
Appendix G Numerical methods
Below, we elaborate on the numerical methods used to validate the replica-symmetric learning curves. In all of the numerical simulations, we set the input dimensionality , and sample with a resolution of around 50 - 100 estimates per dimension. To produce the standard error bars, we sample 10 values per estimate. Numerical procedures were written with NumPy and SciPy . Plots were generated with Matplotlib .
Theoretical predictions for the generalization error in Bayesian random feature models can be computed directly from equation (20). Additionally, after sampling a set of initial weights, we use the results from Appendix C to directly compute the posterior expectations for the RF model. By then averaging the resulting error across multiple samples of weights, we numerically verify our theoretical curves.
G.2 Neural network model
Theoretical predictions for the generalization error in a Bayesian Neural Network can be computed directly from equation (28). We then use the results from Appendix D to directly compute the error for particular instantiations of weights, numerically verifying our theoretical results.