Neural Networks as Kernel Learners: The Silent Alignment Effect
Alexander Atanasov, Blake Bordelon, Cengiz Pehlevan
Introduction
Despite the numerous empirical successes of deep learning, much of the underlying theory remains poorly understood. One promising direction forward to an interpretable account of deep learning is in the study of the relationship between deep neural networks and kernel machines. Several studies in recent years have shown that gradient flow on infinitely wide neural networks with a certain parameterization gives rise to linearized dynamics in parameter space (Lee et al., 2019; Liu et al., 2020) and consequently a kernel regression solution with a kernel known as the neural tangent kernel (NTK) in function space (Jacot et al., 2018; Arora et al., 2019). Kernel machines enjoy firmer theoretical footing than deep neural networks, which allows one to accurately study their training and generalization (Rasmussen & Williams, 2006; Schölkopf & Smola, 2002). Moreover, they share many of the phenomena that overparameterized neural networks exhibit, such as interpolating the training data (Zhang et al., 2017; Liang & Rakhlin, 2018; Belkin et al., 2018). However, the exact equivalence between neural networks and kernel machines breaks for finite width networks. Further, the regime with approximately static kernel, also referred to as the lazy training regime (Chizat et al., 2019), cannot account for the ability of deep networks to adapt their internal representations to the structure of the data, a phenomenon widely believed to be crucial to their success.
In this present study, we pursue an alternative perspective on the NTK, and ask whether a neural network with an NTK that changes significantly during training can ever be a kernel machine for a data-dependent kernel: i.e. does there exist a kernel function for which the final neural network function is with coefficients that depend only on the training data? We answer in the affirmative: that a large class of neural networks at small initialization trained on approximately whitened data are accurately approximated as kernel regression solutions with their final, data-dependent NTKs up to an error dependent on initialization scale. Hence, our results provide a further concrete link between kernel machines and deep learning which, unlike the infinite width limit, allows for the kernel to be shaped by the data.
The phenomenon we study consists of two training phases. In the first phase, the kernel starts off small in overall scale and quickly aligns its eigenvectors toward task-relevant directions. In the second phase, the kernel increases in overall scale, causing the network to learn a kernel regression solution with the final NTK. We call this phenomenon the silent alignment effect because the feature learning happens before the loss appreciably decreases. Our contributions are the following
In Section 2, we demonstrate the silent alignment effect by considering a simplified model where the kernel evolves while small and then subsequently increases only in scale. We theoretically show that if these conditions are met, the final neural network is a kernel machine that uses the final, data-dependent NTK. A proof is provided in Appendix B.
In Section 3, we provide an analysis of the NTK evolution of two layer linear MLPs with scalar target function with small initialization. If the input training data is whitened, the kernel aligns its eigenvectors towards the direction of the optimal linear function early on during training while the loss does not decrease appreciably. After this, the kernel changes in scale only, showing this setup satisfies the requirements for silent alignment discussed in Section 2.
In Section 4, we extend our analysis to deep MLPs by showing that the time required for alignment scales with initialization the same way as the time for the loss to decrease appreciably. Still, these time scales can be sufficiently separated to lead to the silent alignment effect for which we provide empirical evidence. We further present an explicit formula for the final kernel in linear networks of any depth and width when trained from small initialization, showing that the final NTK aligns to task-relevant directions.
In Section 5, we show empirically that the silent alignment phenomenon carries over to nonlinear networks trained with ReLU and Tanh activations on isotropic data, as well as linear and nonlinear networks with multiple output classes. For anisotropic data, we show that the NTK must necessarily change its eigenvectors when the loss is significantly decreasing, destroying the silent alignment phenomenon. In these cases, the final neural network output deviates from a kernel machine that uses the final NTK.
Jacot et al. (2018) demonstrated that infinitely wide neural networks initialized with an appropriate parameterization and trained on mean square error loss evolve their predictions as a linear dynamical system with the NTK at initalization. A limitation of this kernel regime is that the neural network internal representations and the kernel function do not evolve during training. Conditions under which such lazy training can happen is studied further in (Chizat et al., 2019; Liu et al., 2020). Domingos (2020) recently showed that every model, including neural networks, trained with gradient descent leads to a kernel model with a path kernel and coefficients that depend on the test point . This dependence on makes the construction not a kernel method in the traditional sense that we pursue here (see Remark 1 in (Domingos, 2020)).
Phenomenological studies and models of kernel evolution have been recently invoked to gain insight into the difference between lazy and feature learning regimes of neural networks. These include analysis of NTK dynamics which revealed that the NTK in the feature learning regime aligns its eigenvectors to the labels throughout training, causing non-linear prediction dynamics (Fort et al., 2020; Baratin et al., 2021; Shan & Bordelon, 2021; Woodworth et al., 2020; Chen et al., 2020; Geiger et al., 2021; Bai et al., 2020). Experiments have shown that lazy learning can be faster but less robust than feature learning (Flesch et al., 2021) and that the generalization advantage that feature learning provides to the final predictor is heavily task and architecture dependent (Lee et al., 2020). Fort et al. (2020) found that networks can undergo a rapid change of kernel early on in training after which the network’s output function is well-approximated by a kernel method with a data-dependent NTK. Our findings are consistent with these results.
Stöger & Soltanolkotabi (2021) recently obtained a similar multiple-phase training dynamics involving an early alignment phase followed by spectral learning and refinement phases in the setting of low-rank matrix recovery. Their results share qualitative similarities with our analysis of deep linear networks. The second phase after alignment, where the kernel’s eigenspectrum grows, was studied in linear networks in (Jacot et al., 2021), where it is referred to as the saddle-to-saddle regime.
Unlike prior works (Dyer & Gur-Ari, 2020; Aitken & Gur-Ari, 2020; Andreassen & Dyer, 2020), our results do not rely on perturbative expansions in network width. Also unlike the work of Saxe et al. (2014), our solutions for the evolution of the kernel do not depend on choosing a specific set of initial conditions, but rather follow only from assumptions of small initialization and whitened data.
The Silent Alignment Effect and Approximate Kernel Solution
where is the learning rate. If one had access to the dynamics of throughout all , one could solve for the final learned function with integrating factors under conditions discussed in Appendix A
Here, , , and is the initial error on point . We see that the final function has contributions throughout the full training interval . The seminal work by Jacot et al. (2018) considers an infinite-width limit of neural networks, where the kernel function stays constant throughout training time. In this setting where the kernel is constant and , then we obtain a true kernel regression solution for a kernel which does not depend on the training data.
Much less is known about what happens in the rich, feature learning regime of neural networks, where the kernel evolves significantly during time in a data-dependent manner. In this paper, we consider a setting where the initial kernel is small in scale, aligns its eigenfunctions early on during gradient descent, and then increases only in scale monotonically. As a concrete phenomenological model, consider depth networks with homogenous activation functions with weights initialized with variance . At initialization (see Appendix B). We further assume that after time , the kernel only evolves in scale in a constant direction
The assumption that the kernel evolves early on in gradient descent before increasing only in scale may seem overly strict as a model of kernel evolution. However, we analytically show in Sections 3 and 4 that this can happen in deep linear networks initialized with small weights, and consequently that the final learned function is a kernel regression with the final NTK. Moreover, we show that for a linear network with small weight initialization, the final NTK depends on the training data in a universal and predictable way.
Kernel Evolution in 2 Layer Linear Networks
We will first study shallow linear networks trained with small initialization before providing analysis for deeper networks in Section 4. We will focus our discussion in this section on the scalar output case but we will provide similar analysis in the multiple output channel case in a subsequent section. We stress, though we have systematically tracked the error incurred at each step, we have focused on transparency over rigor in the following derivation. We demonstrate that our analytic solutions match empirical simulations in Appendix C.5.
Under gradient flow with learning rate , the weight matrices in each layer evolve as
The NTK takes the following form throughout training.
Note that while the second term, a simple isotropic linear kernel, does not reflect the nature of the learning task, the first term can evolve to yield an anisotropic kernel that has learned a representation from the data.
We next show that there are essentially two phases of training when training a two-layer linear network from small initialization on whitened-input data.
Phase I: An alignment phase which occurs for . In this phase the weights align to their low rank structure and the kernel picks up a rank-one term of the form . In this setting, since the network is initialized near , which is a saddle point of the loss function, the gradient of the loss is small. Consequently, the magnitudes of the weights and kernel evolve slowly.
Phase II: A data fitting phase which begins around . In this phase, the system escapes the initial saddle point and loss decreases to zero. In this setting both the kernel’s overall scale and the scale of the function increase substantially.
If Phase I and Phase II are well separated in time, which can be guaranteed by making small, then the final function solves a kernel interpolation problem for the NTK which is only sensitive to the geometry of gradients in the final basin of attraction. In fact, in the linear case, the kernel interpolation at every point along the gradient descent trajectory would give the final solution as we show in Appendix G. A visual summary of these phases is provided in Figure 2.
In this section we show how the kernel aligns to the correct eigenspace early in training. We focus on the whitened setting, where the data matrix has all of its nonzero singular values equal. We let represent the normalized component of in the span of the training data . We will discuss general in section 3.2. We approximate the dynamics early in training by recognizing that the network output is small due to the small initialization. Early on, the dynamics are given by:
Truncating terms order and higher, we can solve for the kernel’s dynamics early on in training
where is an initialization dependent quantity, see Appendix C.1. The bound on the error is obtained in Appendix C.2. We see that the kernel picks up a rank one-correction which points in the direction of the task vector , indicating that the kernel evolves in a direction sensitive to the target function . This term grows exponentially during the early stages of training, and overwhelms the original kernel with timescale . Though the neural network has not yet achieved low loss in this phase, the alignment of the kernel and learned representation has consequences for the transfer ability of the network on correlated tasks as we show in Appendix F.
1.2 Phase II: Spectral Learning
We now assume that the weights have approached their low rank structure, as predicted from the previous analysis of Phase I dynamics, and study the subsequent NTK evolution. We will show that, under the assumption of whitening, the kernel only evolves in overall scale.
First, following (Fukumizu, 1998; Arora et al., 2018; Du et al., 2018), we note the following conservation law which holds for all time. If we assume small initial weight variance , at initialization, and stays that way during the training due to the conservation law. This condition is surprisingly informative, since it indicates that is rank-one up to corrections. From the analysis of the alignment phase, we also have that . These two observations uniquely determine the rank one structure of to be . Thus, from equation 5 it follows that in Phase II, the kernel evolution takes the form
where . This demonstrates that the kernel only changes in overall scale during Phase II.
Once the weights are aligned with this scheme, we can get an expression for the evolution of analytically, , using the results of (Fukumizu, 1998; Saxe et al., 2014) as we discuss in C.4. This is a sigmoidal curve which starts at and approaches . The transition time where active learning begins occurs when . This analysis demonstrates that the kernel only evolves in scale during this second phase in training from the small initial value to its asymptote.
Hence, kernel evolution in this scenario is equivalent to the assumptions discussed in Section 2, with , showing that the final solution is well approximated by kernel regression with the final NTK. We stress that the timescale for the first phase , where eigenvectors evolve, is independent of the scale of the initialization , whereas the second phase occurs around . This separation of timescales for small guarantees the silent alignment effect. We illustrate these learning curves and for varying in Figure C.2.
2 Unwhitened data
Extension to Deep Linear Networks
We derive this formula in the Appendix E. We observe that the NTK consists of a rank- correction to the isotropic linear kernel with the rank-one spike pointing along the direction. This is true dynamically throughout training under the assumption of small . At convergence , which is the unique fixed point reachable through gradient descent. We discuss evolution of below. The alignment of the NTK with the direction increases with depth .
We now argue that in the case where the input data is whitened, the trained network function is again a kernel machine that uses the final NTK. The unit vector quickly aligns to since the first layer weight matrix evolves in the rank-one direction throughout training for a time dependent vector function . As a consequence, early in training the top eigenvector of the NTK aligns to . Due to gradient descent dynamics, grows only in the direction. Since the quickly aligns to due to growing only along the direction, then the global scalar function satisfies the dynamics in the whitened data case, which is consistent with the dynamics obtained when starting from the orthogonal initialization scheme of Saxe et al. (2014). We show in the Appendix E.1 that spectral learning occurs over a timescale on the order of , where is the time required to reach half the value of the initial loss. We discuss this scaling in detail in Figure 3, showing that although the timescale of alignment shares the same scaling with for , empirically alignment in deep networks occurs faster than spectral learning. Hence, the silent alignment conditions of Section 2 are satisfied. In the case where the data is unwhitened, the vector aligns with early in training. This happens since, early on, the dynamics for the first layer are for time dependent vector . However, for the same reasons we discussed in Section 3.2 the kernel must realign at late times, violating the conditions for silent alignment.
2 Multiple Output Channels
Silent Alignment on Real Data and ReLU nets
In this section, we empirically demonstrate that many of the phenomena described in the previous sections carry over to the nonlinear homogenous networks with small initialization provided that the data is not highly anisotropic. A similar separation in timescales is expected in the nonlinear -homogenous case since, early in training, the kernel evolves more quickly than the network predictions. This argument is based on a phenomenon discussed by Chizat et al. (2019). Consider an initial scaling of the parameters by . We find that the relative change in the loss compared to the relative change in the features has the form which becomes very large for small initialization as we show in Appendix I. This indicates, that from small initialization, the parameter gradients and NTK evolve much more quickly than the loss. This is a necessary, but not sufficient condition for the silent alignment effect. To guarantee the silent alignment, the gradients must be finished evolving except for overall scale by the time the loss appreciably decreases. However, we showed that for whitened data that nonlinear ReLU networks do in fact enjoy the separation of timescales necessary for the silent alignment effect in Figure 1. In even more realistic settings, like ResNet in Figure 1 (d), we also see signatures of the silent alignment effect since the kernel does not grow in magnitude until the alignment has stabilized.
We now explore how anisotropic data can interfere with silent alignment. We consider the partial whitening transformation: let the singular value decomposition of the data matrix be and construct a new partially whitened dataset , where . As the dataset becomes closer to perfectly whitened. We compute loss and kernel aligment for depth 2 ReLU MLPs on a subset of CIFAR-10 and show results in Figure 4. As the agreement between the final NTK and the learned neural network function becomes much closer, since the kernel alignment curve is stable after a smaller number of training steps. As the data becomes more anisotropic, the kernel’s dynamics become less trivial at later time: rather than evolving only in scale, the alignment with the target function varies in a non-trivial way while the loss is decreasing. As a consequence, the NN function deviates from a kernel machine with the final NTK.
Conclusion
We provided an example of a case where neural networks can learn a kernel regression solution while in the rich regime. Our silent alignment phenomenon requires a separation of timescales between the evolution of the NTK’s eigenfunctions and relative eigenvalues and a separate phase where the NTK grows only in scale. We demonstrate that, if these conditions are satisfied, then the final neural network function satisfies a representer theorem for the final NTK. We show analytically that these assumptions are realized in linear neural networks with small initialization trained on approximately whitened data and observe that the results hold for nonlinear networks and networks with multiple outputs. We demonstrate that silent alignment is highly sensitive to anisotropy in the input data.
Our results demonstrate that representation learning is not necessarily at odds with the learned neural network function being a kernel regression solution; i.e. a superposition of a kernel function on the training data. While we provide one mechanism for a richly trained neural network to learn a kernel regression solution through the silent alignment effect, perhaps other temporal dynamics of the NTK could also give rise to the neural network learning a kernel machine for a data-dependent kernel. Further, by asking whether neural networks behave as kernel machine for some data-dependent kernel function, one can hopefully shed light on their generalization properties and transfer learning capabilities (Bordelon et al., 2020; Canatar et al., 2021; Loureiro et al., 2021; Geiger et al., 2021) and see Appendix F.
CP acknowledges support from the Harvard Data Science Initiative. AA acknowledges support from an NDSEG Fellowship and a Hertz Fellowship. BB acknowledges the support of the NSF-Simons Center for Mathematical and Statistical Analysis of Biology at Harvard (award #1764269) and the Harvard Q-Bio Initiative. We thank Jacob Zavatone-Veth and Abdul Canatar for helpful discussions and feedback.
References
Appendix A Derivation of Equation 2
Given a training set of data points , the dynamics of the network training errors close in terms of a time-varying neural tangent kernel
which can easily be verified to solve with initial condition . Under the condition that commutes with , which is true in the settings of interest in this paper, specifically the setting discussed in Appendix B, we can simplify the Peano-Baker series into a simple matrix exponential
Thus, under the condition that commutes with we can exactly solve for the training error dynamics in terms of integrating factors
We expect this formula to hold approximately whenever the eigenvectors of are approximately equal to the eigenvectors of .
A.2 Test Point Predictions with Time Varying Kernel
Given access to the value of the function on training points, one can evaluate the function on test points. We have that the evolution of the function on a test point is given by
where . This gives the final value of to be
Appendix B Kernel Evolution in Scale Only
We consider the model of kernel evolution introduced in Section 2 where the kernel evolves only in scale for and is of small overall size for ,
Introducing operator norm of a matrix, , we will now bound the operator norm of the change in the transition matrix introduced in section A.1.
We begin by noting that, due to the triangle inequality,
where the final inequality follows from the fact that is positive semidefinite for all . Therefore we have shown that . Using this inequality, we find that
With the above Lemma 1, we can bound the discrepancy and , namely
This inequality must therefore hold entry-wise as well, so that
We will now establish how the training predictions evolve for the second interval .
Suppose that from that obeys the dynamics where is as in equation 26. Then, for all ,
The differential equation can be solved through eigendecomposition and integrating factors. Let represent the -th component of in the eigenbasis of which is static for . Let the corresponding eigenvalue of be . The scalar variable obeys the dynamics
This can be solved with integrating factors, noting that . This implies that . Written as a vector, . Since by Lemma 1 we have , we obtain the desired result. ∎
We will now combine the results of the previous two lemmas which analyze the evolution of the network predictions on the training set to give our main silent alignment result, which specifies what the neural network function predicts for an arbitrary test point .
Let the kernel have dynamics of Equation 18 where is a continuous, integrable function with . The function learned by the neural network is
Using Lemma 1 and 2, we know the full dynamics for training predictions . Using , we can solve for the final predictor by integrating dynamics .
We can now integrate the matrix exponential in the second term, using the fact that
Using the fact that from Lemma 1, we arrive at the desired result
We have now established that, given the kernel dynamics in Equation 18, converges to the kernel regression solution with final NTK as provided . This is generic in the settings we consider in this paper for networks with small initialization. In this small initialization setting, is also negligible so that itself is a kernel regression solution. For example, in a linear depth neural network with initial weight scale , the initial scale of the kernel is while the time to alignment scales as thus can be made arbitrarily small by taking . Lastly, the initial network outputs can also be made arbitrarily small.
Appendix C Phases of Learning at Small Initialization
We now present an analysis distinct from that of the previous subsection to go beyond the first step of gradient descent. The NTK for the two layer linear network has the form with . Our goal is to determine the eigendecomposition of . Introduce the variables and . These dynamics form a closed two dimensional linear system early in training
The variable represents the alignment of the NTK with the optimal direction while defines the alignment of the network with the teacher. We see that this alignment increases exponentially with timescale . While the above equations hold for early time and small initialization for any initial condition , we can further estimate these initial values under random initialization provided the input dimension is large. We stress that this limit is not necessary for the silent alignment, but allows for a nice simplification. For Gaussian initialization with large , we have
In the large limit, we have with high probability and thus and . Note that this gives the quantities early in training. Now consider a unit vector which is orthogonal to the solution . We find that the projection of along this direction evolves dynamically as:
We can conclude that is equal to up to an additive initialization constant. We see that this is evolving half as quickly as . Since is small compared to the exponentially growing , the only matrix that satisfies these two conditions must necessarily take the form
The first term, which is growing exponentially in will eventually overwhelm the randomly initialized kernel , which is .
C.2 Phase I: Error in the Leading Order Approximation
In solving the equations of the previous section, we truncated the full gradient descent equations at order . It is important to confirm that the error generated by this truncation remains bounded. We will argue by self-consistency. The full equations are
One can use these equations to solve for the dynamics of the variables:
The second equality in equation 38 comes from inserting a complete basis of states including and into the last term of the left-hand side. The second equality in equation 39 comes from writing . Note that the final term in brackets on the right hand side is a conserved quantity for linear networks, and so is always of order .
Assuming the solutions for are valid to order , we get that .
Further noting that because of the conservation law is also constant in time. This gives us that
We now note that both grow as a (-independent) constant times times . The correction to the dynamics of both equations is then bounded by a constant times . This will be less than as long as satisfies
For , the alignment time falls within this range and we are guaranteed alignment to the Ganguli-Saxe configuration.
The error of the full solution at time can be bounded by the integral of this error bound from to , namely a constant times . As long as , we are guaranteed that the error of the kernel is as given in equation 7.
C.3 Phase I: Two Layer Analysis with Unwhitened Data
We now study the same linearization around the initial fixed point used in the main text but for the two layer network with unwhitened data. In this case,
which holds asymptotically as . We introduce the following variables which form a closed linear dynamical system
Introduce the constants . Using the weight dynamics, it is straightforward to show that
This matrix has eigenvalues . Since there are only two positive eigenvalues , it suffices to consider evolution along those two eigendirections, where the kernel and neural network function will be amplified. Evolution along these direcions give
where are constants determined by intialization. At large time, the large eigenvalue mode will dominate. Decomposing we find that the only self consistent solution is . This implies that the kernel evolution will take the form
We see that the kernel evolves along the directions early in training for unwhitened data. We visualize the two stages of learning for unwhitened data in Figure C.1.
C.4 Phase II: Whitened Data
Consider a two layer network where balance has been achieved and . Once this balance condition is stable for fixed , we can calculate the time derivative of
Letting , we find that , which is the two layer dynamics derived in Saxe et al. (2014). This dynamics has solution .
C.5 Solutions to the full training dynamics of linear networks at small initialization
By combining the analyses of the subsection C.1 with the exact solutions discussed in C.4 we can match both solutions to obtain formulas for and for the entire network’s training path that are exact up to corrections. Up to we then have that
This yields that the initialization constant plays the effective role of in the Ganguli-Saxe solution for phase II. Equation 34 yields the expected value of this initialization constant. We have empirically verified that these exact equations hold to high accuracy across a variety network sizes, initialization scales, and whitened datasets. We illustrate some of these in figure C.3.
Appendix D Balancing of Weights in Deep Linear Networks
The second line follows from the first since the following quantities are identical
By inspection these two quantities are equal. Thus we have, for any loss function, a deep linear network has the following conservation laws
Appendix E NTK Formula for Deep Linear Networks
The neural tangent kernel for a linear network is defined as an inner-product over all gradients
𝐿2\sigma^{-L+2} as in the linear setting. (b) Deep networks with nonlinearities are seen to undergo the silent alignment effect early on in training. The dashed lines indicate when the kernel has grown to 10% of its final value. (c) The trained network outputs on test data match closely the kernel regression with the final learned kernel, but do not match regression the initial kernel. E.1 Deep Linear Network Dynamics Under Balance In this section, we will consider the dynamics of the variable once the balance condition is satisfied. Let . Then the dynamics for under the balancing assumption is
which implies the fact . Changing variables to we obtain
When is initialized to a very small value compared to we can
This implies a timescale to learn of .
We can approximate the timescale for the first layer’s singular vector to align to as well. Let be a vector orthogonal to .
This suggests that alignment in a deep network should also occur on a timescale of . While there is no strict separation of timescales in terms of the scaling of alignment and learning with , we find that alignment tends to precede a significant drop in the loss as we show in Figure 3.
E.2 Final NTK for Deep Linear Networks
Independent of the structure of the data, the first vector and the final NTK has the following form:
This formula is merely a consequence of the balancing condition and convergence to the optimum . We provide empirical support that final kernel alignment increases with depth in Figure E.3.
E.3 Formula for the Intermediate NNGP Kernels
Appendix F Generalization Error in Transfer Learning Task
The structure of the final kernel can alter the ability of the network to flexibly transfer to new tasks with a small amount of data. In this section, we examine how learned intermediate representations compares with the inductive bias of the original isotropic kernel . In particular, we study the offline generalization performance of kernels of the form
In Section 4 we showed that could be altered by changing the network depth. Concretely, our transfer learning problem consists of training a linear probe on one of the intermediate layers of the network (Alain & Bengio (2016); Cohen et al. (2020)). This would also produce a kernel regression solution for kernel with which depends on the chosen layer and the depth of the network. For simplicity, we assume that the data are generated according to a simple Gaussian distribution and that the target values are generated with a linear function . We decompose the new task vector where . The expected generalization error after training with samples can be computed with methods from the physics of disordered systems (Bordelon et al., 2020; Canatar et al., 2021; Loureiro et al., 2021). For any , the easiest transfer task is (). If , increasing the alignment strictly decreases the generalization error. This is illustrated in Figure F.1.
Once this integral eigenvalue problem is solved for eigenvalues and orthonormal eigenfunctions , the average case generalization error at training examples is
In our case, we are interested in the generalization performance of the linear kernel
Since is a linear kernel, its eigenfunctions should be linear functions . Assuming that the data distribution has identity covariance, we find
This implies that the vectors are eigenvectors of . The first eigenvector is with eigenvalue . The other eigenvectors can be chosen as any frame in the dimensional subspace orthogonal to . Each of these eigenvectors has eigenvalue . Using these results, and the fact that , we can calculate the expected generalization error.
By the result proven in Canatar et al. (2021), the lowest possible error for fixed occurs by maximizing the fraction of variance along the large eigenvalue direction, corresponding to .
Appendix G Linear NTKs During GD Learn The Same Function
which is the kernel regression solution for the initial kernel . This is unsurprising due to a symmetry argument: when , the only privileged point on the affine space is the point closest to the origin, which is precisely the solution above. Surprisingly, the final pseudo-inverse solution also minimizes the RKHS norms for any of the kernels throughout gradient descent. Up to an overall scale, the kernels throughout evolution take the form (see Section 4.1) which would induce the following kernel interpolation problems
The solution to this optimization problem is indeed the kernel regression solution with kernel since the learned function takes the form with . Using the Sherman-Morrison rule, we show that the solution to each of these problems gives the same result, namely the pseudo-inverse solution. This can be seen from the following
Now, we let , where . This is the general decomposition for the set of interpolators which have the property .
The solution is merely to set . Thus the optimal solution is therefore the same for any finite value of . However, the final RKHS norm of the learned function decreases with time, indicating that the kernel becomes more aligned with the pseudo-inverse direction as increases.
From the balance condition , we also have
Equating the two above expressions for and taking an inner product with from the left gives the following
Thus, if then so that the full dynamics of lie in the subspace spanned by the training data. At initialization, we have so the initial vector will indeed align with the span of the training data in the limit.
By the fact that , the learned linear coefficients must also be in so . These must also interpolate the data provided , giving the following condition
While anisotropy of the data makes no impact on what function is ultimately learned in the linear network case, the anisotropy can have a signficant influence on whether the preconditions for silent alignment are satisfied in a nonlinear network, which can prevent the final function from being a NTK regressor with final NTK.
Appendix H Final Kernel in Multi-Class Networks
For a network with output channel, balancing and alignment guarantee that the configuration of the network is orthogonal and balanced as in the setting of Saxe et al. (2014). One can then integrate each mode separately to obtain the final kernel as
Specifically, for a depth network, both the the alignment time and the time to learn a given singular value scale as , as shown in appendix E.1. The differences in alignment times for modes therefores scales as .
H.2 More Refined Balancing Analysis
We can derive corrections to our decomposition of the kernel by including the initial conditions in our derived conservation laws. In particular, we will consider balancing for large width networks. Note that the network does not need to be in the lazy regime. The balancing condition is
In the limit, we can solve that as before. However, we can now obtain leading order corrections (in ) which take the form of the form
Appendix I Laziness in Homogenous Networks
Here, denotes the operator norm of a matrix. We are interested in the ratio of the loss’ time derivative to the gradient’s time derivative. With an initialization scale of we find
where in the last step we approximated for small initialization scale since and . Now we will estimate the scale of each of the terms above. For a homogenous model and . Counting powers of in numerator and denominator, we find that this quantity of interest scales as
This result indicates that, from small initialization, the gradient NTK features and thus the kernel itself will evolve much more rapidly than the loss. This effect can be amplified by increasing depth and decreasing initialization scale.
Appendix J ResNet Experimental Details
Below, we provide the alignment and loss dynamics for wide resnet for CIFAR-10 with training points. Because the loss decreases significantly before the kernel reaches its final alignment value, the final NTK is not perfectly correlated with the final neural network function. The wide ResNet model is taken from Novak et al. (2020) and is based on the original architecture of Zagoruyko & Komodakis (2017) with a widening factor of and a single block per ResNet group , giving a final network with trainable layers. For both Figure 1 (d) and (e) as well as Figure J.1 use Adam with a learning rate of and initial weight scale of in standard parameterization for all intermediate blocks. For the first conv layer, we used . We find that small initial weight variance in the first layer gives rise to less stable learning and worse kernel alignment.
Below, in Figure J.2, we provide comprehensive results for different depths which we control by increasing the number of blocks per group , corresponding to WRNs with trainable conv layers.
6𝑏16b+1 total conv layers). (a) The deeper models train faster with Adam. (b) The final NTK norm increases with depth. (c) The alignment achieves close to its final value by the time the kernel norm reaches of its final value (dashed line), indicating successful silent alignment. (d) The neural network predictions are very close to the predictions of the final NTK but is not accurately predicted by the initial NTK. Figure J.3: Varying the ResNet widening parameter also alters the kernel and loss dynamics. (a) The loss curve for WRNs with widening factor . Wider networks train more quickly. (b) The kernel norm increases more rapidly for wider networks but changes by a smaller amount. (c) Alignment reaches close to its asymptote by the time the kernel norm grows to its final value (dashed). (d) The final kernel is a much better predictor of the NN function than the initial kernel. J.1 Adaptive Optimizers and the Relevant Kernel Many adaptive gradient methods compute updates to parameters according to
where are time-varying functions which are computed in terms of the history of gradient moments for parameter or in terms of its instantaneous gradient. The relevant kernel at time which governs instantaneous evolution of network predictions is
since . Though we do not calculate this kernel which is relevant to the adaptive learning rate scheme since it is not supported in Neural Tangents API, this could be a worthy future investigation.