Pathological spectra of the Fisher information metric and its variants in deep neural networks
Ryo Karakida, Shotaro Akaho, Shun-ichi Amari
Introduction
Deep neural networks (DNNs) have outperformed many standard machine-learning methods in practical applications . Despite their practical success, many theoretical aspects of DNNs remain to be uncovered, and there are still many heuristics used in deep learning. We need a solid theoretical foundation for elucidating how and under what conditions DNNs and their learning algorithms work well.
The Fisher information matrix (FIM) is a fundamental metric tensor that appears in statistics and machine learning. An empirical FIM is equivalent to the Hessian of the loss function around a certain global minimum, and it affects the performance of optimization in machine learning. In information geometry, the FIM defines the Riemannian metric tensor of the parameter manifold of a statistical model . The natural gradient method is a first-order gradient method in the Riemannian space where the FIM works as its Riemannian metric . The FIM also acts as a regularizer to prevent catastrophic forgetting ; a DNN trained on one dataset can learn another dataset without forgetting information if the parameter change is regularized with a diagonal FIM.
However, our understanding of the FIM for neural networks has so far been limited to empirical studies and theoretical analyses of simple networks. Numerical experiments empirically confirmed that the eigenvalue spectra of the FIM and those of the Hessian are highly distorted; that is, most eigenvalues are close to zero, while others take on large values . Focusing on shallow neural networks, Pennington and Worah theoretically analyzed the FIM’s eigenvalue spectra by using random matrix theory, and Fukumizu derived a condition under which the FIM becomes singular. Liang et al. have connected FIMs to the generalization ability of DNNs by using model complexity, but their results are restricted to linear networks. Thus, theoretical evaluations of deeply nonlinear cases seem to be difficult mainly because of iterated nonlinear transformations. To go one step further, it would be helpful if a framework that is widely applicable to various DNNs could be constructed.
Investigating DNNs with random weights has given promising results. When such DNNs are sufficiently wide, we can formulate their behavior by using simpler analytical equations through coarse-graining of the model parameters, as is discussed in mean field theory and random matrix theory . For example, Schoenholz et al. proposed a mean field theory for backpropagation in fully-connected DNNs. This theory characterizes the amplitudes of gradients by using specific quantities, i.e., order parameters in statistical physics, and enables us to quantitatively predict parameter regions that can avoid vanishing or explosive gradients. This theory is applicable to a wide class of DNNs with various non-linear activation functions and depths. Such DNNs with random weights are substantially connected to Gaussian process and kernel methods . Furthermore, the theory of the neural tangent kernel (NTK) explains that even trained parameters are close enough to the random initialization in sufficiently wide DNNs and the performance of trained DNNs is determined by the NTK on the initialization .
Karakida et al. focused on the FIM corresponding to the mean square error (MSE) loss and proposed a framework to express certain eigenvalue statistic by using order parameters. They revealed that when fully-connected networks with random initialization are sufficiently wide, the FIM’s eigenvalue spectrum asymptotically becomes pathologically distorted. As the network width increases, a small number of the eigenvalues asymptotically take on huge values and become outliers while the others are much smaller. The distorted shape of the eigenvalue spectrum is consistent with empirical reports . While LeCun et al. implied that such pathologically large eigenvalue might appear in multi-layered networks and affect the training dynamics, its theoretical elucidation has been limited to a data covariance matrix in a linear regression model. The results of can be regarded as a theoretical verification of this large eigenvalue suggested by . The obtained eigenvalue statistics have given insight into the convergence of gradient dynamics , mechanism of batch normalization to decrease the sharpness of the loss function , and generalization measure of DNNs based on the minimum description length .
In this paper, we extend the framework of the previous work and reveal that various types of FIMs and variants show pathological spectra. Our main contribution is the following:
FIM for classification tasks with softmax output: While the previous works analyzed the FIM for regression based on the MSE loss, we typically use the cross-entropy loss with softmax output in classification tasks. We analyze this FIM for classification tasks and reveal that its spectrum is pathologically distorted as well. While the FIM for regression tasks has unique and degenerated outliers in the infinite-width limit, the softmax output can make these outliers disperse and remove the degeneracy. Our theory shows that there are number-of-classes outlier eigenvalues, which is consistent with experimental reports . Experimental results demonstrate that the eigenvalue density has a tail of outliers spreading form the bulk.
Furthermore, we also give a unified perspective on the variants:
Diagonal Blocks of FIM: We give a detailed analysis of the diagonal block parts of the FIM for regression tasks. Natural gradient algorithms often use a block diagonal approximation of the FIM . We show that the diagonal blocks also suffer from pathological spectra.
Connection to NTK: The NTK and FIM inherently share the same non-zero eigenvalues. Paying attention to a specific re-scaling of the parameters assumed in studies of NTK, we clarify that NTK’s eigenvalue statistics become independent of the width scale. Instead, the gap between the average and maximum eigenvalues increases with the sample size. This suggests that, as the sample size increases, the training dynamics converge non-uniformly and that calculations with the NTK become ill-conditioned. We also demonstrate a simple normalization method to make eigenvalue statistics that are independent of both the width and the sample size.
Metric tensors for input and feature spaces: We consider metric tensors for input and feature spaces spanned by neurons in input and hidden layers. These metric tensors potentially enable us to evaluate the robustness of DNNs against perturbations in the input and feedforward propagated signals. We show that the spectrum is pathologically distorted, similar to FIMs, in the sense that the outlier of the spectrum is much far from most of the eigenvalues. The softmax output makes the outliers disperse as well.
In summary, this study sheds light into the asymptotical eigenvalue statistics common to various wide networks.
Preliminaries
We investigated the fully-connected feedforward neural network shown in Fig. 1. The network consists of one input layer, hidden layers (), and one output layer. It includes shallow nets () and arbitrary deep nets (). The network width is denoted by . The pre-activations and activations of units in the -th layer are defined recursively by
and consider the limiting case of a sufficiently large with constant coefficients . The number of output units is taken to be a constant , as is usually done in practice. We denote the linear output of the last layer by
We also investigate DNNs with softmax outputs in Section 3.3. The -dimensional softmax function is given by
To avoid complicating the notation, we will omit the index of the output unit, i.e., . To evaluate the above feedforward and backward signals, we assume the following conditions.
Random weights and biases: Suppose that the parameter set is an ensemble generated by
and thus is fixed, where denotes a Gaussian distribution with zero mean and variance . Treating the case in which different layers have different variances is straightforward. Note that the variances of the weights are scaled in the order of . In practice, the learning of DNNs usually starts from random initialization with this scaling .
Activation functions: Suppose the following two conditions: (i) the activation function has a polynomially bounded weak derivative. (ii) the network is non-centered, which means a DNN with bias terms () or activation functions satisfying a non-zero Gaussian mean. The definition of the non-zero Gaussian mean is . The notation means integration over the standard Gaussian density.
Condition (i) is used to obtain recurrence relations of backward order parameters . Condition (ii) plays an essential role in our evaluation of the FIM . The two conditions are valid in various realistic settings, because conventional networks include bias terms, and widely used activation functions, such as the sigmoid function and (leaky-) ReLUs, have bounded weak derivatives and non-zero Gaussian means. Different layers may have different activation functions.
2 Overview of metric tensors
We will analyze two types of metric tensors (metric matrices) that determine the responses of network outputs, i.e., the response to a local change in parameters and the response to a local change in the input and hidden neurons. They are summarized in Fig. 2. One can systematically understand these tensors from the perspective of perturbations of variables.
where we have abbreviated the network outputs as to avoid complicating the notation. This is an empirical FIM in the sense that the average is computed over empirical input. We can express it in the matrix form shown in Fig. 2(a). The Jacobian is a matrix whose each column corresponds to (, ). We investigate this type of empirical metric tensor for arbitrary . One can set as a constant value or make it increase depending on . The empirical FIM (10) converges to the expected FIM as . In addition, the FIM can be partitioned into layer-wise block matrices. We denote the -th block as (). We take a closer look at the eigenvalue statistics of diagonal blocks in Section 3.2.
In this paper, we also investigate another FIM denoted by which corresponds to classification tasks with cross-entropy loss. As shown in Fig. 2(a), one can represent as a modification of . A specific coefficient matrix is inserted between the Jacobian and its transpose. is defined by (25) and composed of nothing but softmax functions. Its mathematical definition is given in Section 3.3. One more interesting quantity is a left-to-right reversed product of described in Fig. 2 (b). This matrix is known as the neural tangent kernel (NTK). The FIM and NTK share the same non-zero eigenvalues by definition, although we need to be careful in the change of parameterization used in the studies of NTK. The details are shown in Section 4.
We refer to as the metric tensor for the input and feature spaces because each acts as the input to the next layer and corresponds to the features realized in the network. We can also deal with a layer-wise diagonal block of . Let us denote the -th block by (). In particular, the first diagonal block indicates the robustness of the network output against perturbation of the input:
Robustness against input noise has been investigated by similar (but different) quantities, such as sensitivity , and robustness against adversarial examples .
3 Order parameters for wide neural networks
where is the output of the -th layer generated by the -th input sample (). The variable describes the total activity in the -th layer, and the variable describes the overlap between the activities for different input samples and . These variables have been utilized to describe the depth to which signals can propagate from the perspective of order-to-chaos phase transitions . In the large limit, these variables can be recursively computed by integration over Gaussian distributions :
for . Because the input samples generated by Eq. (7) yield and for all and , in each layer takes the same value for all ; so does for all . A two-dimensional Gaussian integral is given by
with . One can represent this integral in a bit simpler form, i.e.,
Next, let us define the following variables for backpropagated signals:
These order parameters depend only on the type of activation function, depth, and the variance parameters and . The recurrence relations for the order parameters require iterations of one- and two-dimensional numerical integrals. Moreover, we can obtain explicit forms of the recurrence relations for some of the activation functions .
Eigenvalue statistics of FIMs
This section shows the asymptotic eigenvalue statistics of the FIMs. When we have an metric tensor whose eigenvalues are (), we compute the following quantities:
The obtained results are universal for any sample size which may depend on .
This subsection overviews the results obtained in the previous studies . The metric tensor is equivalent to the Fisher information matrix (FIM) , originally defined by
Basically, there are two types of FIM for supervised learning, depending on the definition of the statistical model. One type corresponds to the mean squared error (MSE) loss for regression tasks; the other corresponds to the cross-entropy loss for classification tasks. The latter is discussed in Section 3.3. Let us consider the following statistical model for the regression:
The previous studies uncovered the following eigenvalue statistics of the FIM (10):
When is sufficiently large, the eigenvalue statistics of can be asymptotically evaluated as
where , and positive constants and are obtained using order parameters,
The average of the eigenvalue spectrum asymptotically decreases in the order of , while the variance takes a value of and the largest eigenvalue takes a huge value of . It implies that most of the eigenvalues are asymptotically close to zero, while the number of large eigenvalues is limited. Thus, when the network is sufficiently wide, one can see that the shape of the spectrum asymptotically becomes pathologically distorted. This suggests that the parameter space of the DNNs is locally almost flat in most directions but highly distorted in a few specific directions.
In particular, regarding the large eigenvalues, we have
When is sufficiently large, the eigenspace corresponding to is spanned by eigenvectors,
When with a constant , under the gradient independence assumption, the second largest eigenvalue is bounded by
for non-negative constants and .
The subtraction (22) means eliminating the largest eigenvalues from . Numerical experiments confirmed that the largest eigenvalue of is of order 1. Note that when , the sample size is sufficiently large but the network satisfies and keeps overparameterized.
Figure 3 (a; left) shows a typical spectrum of the FIM. We computed the eigenvalues by using random Gaussian weights, biases, and inputs. We used deep Tanh networks with , , , and . The sample size was . The histograms were made from eigenvalues over different networks with different random seeds. The histogram had two populations. The red dashed histogram was made by eliminating the largest eigenvalues. It coincides with the smaller population. Thus, one can see that the larger population corresponds to the largest eigenvalues. The larger population in experiments can be distributed around because is large but finite.
Remark : Loss landscape and gradient methods. The empirical FIM (10) is equivalent to the Hessian of the loss, i.e., , around the global minimum with zero training loss. Karakida et al. referred to the steep shape of the local loss landscape caused by as pathological sharpness. The sharpness of the loss landscape is connected to an appropriate learning rate of gradient methods for convergence. The previous work empirically confirmed that a learning rate satisfying
is necessary for the steepest gradient method to converge . In fact, this acts as a boundary of neural tangent kernel regime . Because increases depending on the width and depth, we need to carefully choose an appropriately scaled learning rate to train the DNNs.
2 Diagonal blocks of FIM
It is helpful to investigate the relation between and its diagonal blocks when one considers the diagonal block approximation of . For example, use of a diagonal block approximation can decrease the computational cost of natural gradient algorithms . When a matrix is composed only of diagonal blocks, its eigenvalues are given by those of each diagonal block. approximated in this fashion has the same mean of the eigenvalues as the original and the largest eigenvalue , which is of . Thus, the diagonal block approximation also suffers from a pathological spectrum. Eigenvalues that are close to zero can make the inversion of the FIM in natural gradient unstable, whereas using a damping term seems to be an effective way of dealing with this instability .
3 FIM for multi-label classification tasks
The cross-entropy loss is typically used in multi-label classification tasks. It comes from the log-likelihood of the following statistical model:
The contribution of softmax output appears only in . is linked to through the matrix representation shown in Fig. 2(a). One can view as a matrix representation with , that is, the identity matrix. In contrast, corresponds to a block-diagonal whose -th block is given by the matrix . In a similar way to Eq. (8), we can see as the metric tensor for the parameter space. Using the softmax output , we have
where the square root is taken entry-wise.
We obtain the following result of ’s eigenvalue statistics:
When is sufficiently large, the eigenvalue statistics of are asymptotically evaluated as
where the constant coefficients are given by
The derivation is shown in Appendix B.1. We find that the eigenvalue spectrum shows the same width dependence as the FIM for regression tasks. Although the evaluation of in Theorem 3.3 is based on inequalities, one can see that linearly increases as the width or the depth increase. The softmax functions appear in the coefficients . It should be noted that the values of generally depend on the index of each softmax output. This is because the values of the softmax functions depend on the specific configuration of and .
If a relatively loose bound is acceptable, we can use the following simpler evaluation. Let us denote the eigenvalue statistics of shown in Theorem 3.1 by . They corresponds to the contribution of linear output before putting it into the softmax function. Taking into account the contribution of the softmax function, we have
Note that ’s eigenvalues satisfy (; ). The inequality (28) comes from .
Figure 4 shows that our theory predicts experimental results rather well for artificial data. We computed the eigenvalues of with random Gaussian weights, biases, and inputs. We set , , , and () = () in the tanh case, () in the ReLU case, and () in the linear case. The sample size was set to . The predictions of Theorem 3.3 coincided with the experimental results for sufficiently large widths.
Exhaustive experiments on the cross-entropy loss have recently confirmed that there are dominant large eigenvalues (so-called outliers) . Consistent with the results of this experimental study, we found that there are large eigenvalues:
has the first largest eigenvalues of .
The theorem is proved in Appendix B.2. These large eigenvalues are reminiscent of the largest eigenvalues of shown in Theorem 3.1.
These largest eigenvalues can act as outliers. Note that because , we have . denotes the -th largest eigenvalue of . Let us suppose that and the assumptions of Theorem 3.2 hold. In this case, is upper-bounded by (21) and then is of at most. This means that the first largest eigenvalues of can become outliers. It would be interesting to extend the above results and theoretically quantify more precise values of these outliers. One promising direction will be to analyze a hierarchical structure of empirically investigated by Papyan . It is also noteworthy that our outliers disappear under the mean subtraction of in the last layer mentioned in (22). This is because we have under the mean subtraction.
Figure 3(b; left) shows a typical spectrum of . We set and other settings were the same as in the case of . We found that compared to , had the largest eigenvalues which were widely spread from the bulk of the spectrum. Naively speaking, this is because the coefficient matrix has distributed eigenvalues as is shown in Fig. 3(b; right). Compared to the FIM for regression, which corresponds to , the distributed ’s eigenvalues can make ’s eigenvalues disperse. Note that the black histogram is obtained as the summation over trials (different random initializations). The largest eigenvalues in a single trial are shown as blue boxes. The diagonal block also has the same characteristics of the spectrum as the original in Fig. 3(b; middle).
Although this work mainly focuses on DNNs with random weights, it also gives some insight into the training of sufficiently wide neural networks. It is known that the whole training dynamics of gradient descent can be explained by NTK in sufficiently wide neural networks. See Section 4 for more details. In the case of the cross-entropy loss, its functional gradient is given by . is the NTK at random initialization described in (29). depends on the softmax function at each time step. One can numerically solve the training dynamics of and obtain the theoretical value of the training loss . In Fig. 5(left), we confirmed that the theoretical line coincided well with the experimental results of gradient descent training. We set random initialization by . We used artificial data with Gaussian inputs () and generated their labels by a teacher network whose architecture was the same as the trained network. In addition, Figure 5(right) shows the largest eigenvalues during the training. As is expected from NTK theory, kept unchanged during the training. In contrast, the ’s largest eigenvalue dynamically changed because depends on the softmax output which changes with the scale of . We calculated theoretical bounds of (blue lines) by substituting at each time step into Theorem 3.3. Note that the bounds obtained in Theorem 3.3 are available even in the NTK regime because the coefficients admit any value of and are not limited to the random initialization. These theoretical bounds explained well the experimental results of during the training.
Because the global minimum is given by , all entries of and approach zero after a large enough number of steps. In the MSE case, we can explain the critical learning rate for convergence as is remarked in (23). In contrast, it is challenging to estimate such a critical learning rate in the cross-entropy case since and dynamically change.
Connection to Neural Tangent Kernel
The empirical FIM (10) is essentially connected to a recently proposed Gram matrix, i.e., the Neural Tangent Kernel (NTK). Jacot et al. defined the NTK by
Note that the Jacobian is a matrix whose each column corresponds to (, ). Under certain conditions with sufficiently large , the NTK at random initialization governs the whole training process in the function space by
where the notation corresponds to the time step of the parameter update and represents the learning rate. Specifically, NTK’s eigenvalues determine the speed of convergence of the training dynamics. Moreover, one can predict the network output on the test samples by using the Gaussian process with the NTK .
The NTK and empirical FIM share essentially the same non-zero eigenvalues. It is easy to see that one can represent the empirical FIM (10) by . This means that the NTK (29) is the left-to-right reversal of up to the constant factor . Karakida et al. introduced , which is essentially the same as the NTK, and referred to as the dual of . They used to derive Theorem 3.1.
It should be noted that the studies of NTK typically suppose a special parameterization different from the usual setting. They consider DNNs with a parameter set which determines weights and biases by
This NTK parameterization changes the scaling of Jacobian. For instance, we have . This makes the eigenvalue statistics slightly change from those of Theorem 3.1. When is sufficiently large, the eigenvalue statistics of under the NTK parameterization are asymptotically evaluated as
The positive constants and are obtained using order parameters,
The derivation is given in Appendix C. The NTK parameterization makes the eigenvalue statistics independent of the width scale. This is because the NTK parameterization maintains the scale of the weights but changes the scale of the gradients with respect to the weights. It also makes () shift to (). This shift occurs because the NTK parameterization makes the order of the weight gradients comparable to that of the bias gradients . The second terms in and correspond to a non-negligible contribution from .
While is independent of the sample size , depends on it. This means that the NTK dynamics converge non-uniformly. Most of the eigenvalues are relatively small and the NTK dynamics converge more slowly in the corresponding eigenspace. In addition, a prediction made with the NTK requires the inverse of the NTK to be computed . When the sample size is large, the condition number of the NTK, i.e., , is also large and the computation with the inverse NTK is expected to be numerically inaccurate.
2 Scale-independent NTK
A natural question is under what condition do NTK’s eigenvalue statistics become independent of both the width and the sample size? As indicated in Eq. (22), the mean subtraction in the last layer with is a simple way to make the FIM’s largest eigenvalue independent of the width. Similarly, one can expect that the mean subtraction makes the NTK’s largest eigenvalue of disappear and the eigenvalue spectrum take a range of independent of the width and sample size.
Figure 5 empirically confirms this speculation. We set , , and used the Gaussian inputs and weights with . As shown in Fig. 5 (left), NTK’s eigenvalue spectrum becomes pathologically distorted as the sample size increases. To make an easier comparison of the spectra, the eigenvalues in this figure are normalized by . As the sample size increases, most of the eigenvalues concentrate close to zero while the largest eigenvalues become outliers. In contrast, Figure 5 (right) shows that the mean subtraction keeps NTK’s whole spectrum in the range of under the condition . The spectrum empirically converged to a fixed distribution in the large limit.
Metric tensor for input and feature spaces
When is sufficiently large, the eigenvalue statistics of are asymptotically evaluated as
Figure 7(a; left) shows typical spectra of and Figure 7(a; right) those of . We used deep Tanh networks with and . The other experimental settings are the same as those in Fig. 3(a). The pathological spectra appear as the theory predicts. Similarly, we show spectra of softmax output in Fig. 7(b). The softmax output made the outliers widely spread in the same manner as in .
Let us remark on some related works in the literature of deep learning. First, Pennington et al. investigated similar but different matrices. Briefly speaking, they used random matrix theory and obtained the eigenvalue spectrum of with , . They found that the isometry of the spectrum is helpful to solve the vanishing gradient problem. Second, DNNs are known to be vulnerable to a specific noise perturbation, i.e., the adversarial example . One can speculate that the eigenvector corresponding to may be related to adversarial attacks, although such a conclusion will require careful considerations.
Discussion
We evaluated the asymptotic eigenvalue statistics of the FIM and its variants in sufficiently wide DNNs. They have pathological spectra in the conventional setting of random initialization and activation functions. This suggests that we need to be careful about the eigenvalue statistics and their influence on the learning when we use large-scale deep networks in naive settings.
The current work focused on fully-connected neural networks and it will be interesting to explore the spectra of other architectures such as ResNets and CNN. It will be also fundamental to explore the eigenvalue statistics that the current study cannot capture. While our study captured some of the basic eigenvalue statistics, it remains to derive the whole spectrum analytically. In particular, after the normalization excludes large outliers, the bulk of the spectrum becomes dominant. Random matrix theory enables us to analyze the FIM’s eigenvalue spectrum in a shallow and centered network without bias terms . Extending the random matrix theory to more general cases including deep neural networks seems to be a prerequisite for further progress. Furthermore, we assumed a finite number of network output units. In order to deal with multi-label classifications with high dimensionality, it would be helpful to investigate eigenvalue statistics in the wide limit of both hidden and network output layers. Finally, although we focused on the finite depth and regarded order parameters as constants, they can exponentially explode on extremely deep networks in the chaotic regime . The NTK in such a regime has been investigated in .
It would also be interesting to explore further connections between the eigenvalue statistics and learning. Recent studies have yielded insights into the connection between the generalization performance of DNNs and the eigenvalues statistics of certain Gram matrices including FIM and NTK . We expect that the theoretical foundation of the metric tensors given in this paper will lead to a more sophisticated understanding and development of deep learning in the future.
Appendices
where represents the Kronecker product. The variables and are functions of , and the expectation is taken over . This expression is not easy to treat in the analysis, and we introduce a dual expression of the FIM, which is essentially the same as NTK, as follows.
First, we briefly overview the derivation of Theorem 3.1 shown in . The essential point is that a Gram matrix has the same non-zero eigenvalues as its dual. One can represent the empirical FIM (10) as
Its columns are the gradients on each input, i.e., . Let us refer to a matrix as the dual of FIM. Matrices and have the same non-zero eigenvalues by definition. This can be partitioned into block matrices. The -th block is given by
for . In the large limit, the previous study showed that asymptotically satisfies
where is the Kronecker delta. As is summarized in Lemma A.1 in , the second term of Eq. (A.4) is negligible in the large limit. In particular, it is reduced to under certain condition. The matrix has entries given by
Note that is positive by definition and is positive under the condition (ii) of activation functions. The other eigenvalues of are given by .
A.2 Diagonal blocks
We can immediately derive eigenvalue statistics of diagonal blocks in the same way as Theorem 3.1. We can represent the diagonal blocks as with
where the parameter set means all parameters in the -th layer. The matrix can be partitioned into block matrices whose -th block is given by
for . As one can see from the additivity of , the following evaluation is part of Eq. (A.4):
where is a matrix. One can rearrange the columns and rows of and partition it into block matrices whose entries are given by
for . Each block is a diagonal matrix. Note that the non-zero eigenvalues of are equivalent to those of . Since we have , we should investigate the eigenvalues of the following matrix:
The third line holds asymptotically, since the order of in Eq. (A.4) is higher than that of (). The fourth line comes from .
where means a summation over excluding the -th sample.
Finally, we derive the largest eigenvalue. Let us denote the eigenvectors of as
Because we have asymptotically , the lower bound is given by
Taking the index that maximizes the right-hand side, we obtain the lower bound of . The upper bound of immediately comes from a simple inequality for non-negative variables, i.e., . Thus, we obtain Theorem 3.3.
Note that we immediately have from (B.3). This means that ’s eigenvalues satisfy (). In addition, we have because and . Therefore, holds and we obtain inequalities (28).
B.2 Derivation of Theorem 3.4
for . Define to be a linear subspace spanned by eigenvectors of corresponding to , i.e., . The indices are chosen from without duplication.
where is of from Theorem 3.1. This holds for all of and we can say that there exist large eigenvalues of .
C Derivation of NTK’s eigenvalue statistics
The NTK is defined as under the NTK parameterization. In the same way as Eq. (A.4), the -th block of the NTK is asymptotically given by
for . In contrast to Eq. (A.4), the NTK parameterization makes multiplied by . The negligible term of is reduced to under the condition summarized in . The entries of are given by
In the same way as with the FIM, the trace of leads to , the Frobenius norm of leads to , and has the largest eigenvalue for arbitrary . The eigenspace of corresponding to is also the same as . It is spanned by eigenvectors ().
D Eigenvalue statistics of A𝐴A
The metric tensor can be represented by , where is an matrix and its columns are the gradients on each input, i.e., . Let us introduce the dual matrix of , i.e., . It has the same non-zero eigenvalues as by definition. Its -th entry is given by
in the large limit. Accordingly, we have
D.2 Diagonal blocks
In the same way as in the FIM, we can also evaluate the eigenvalue statistics of the diagonal blocks (). The metric tensor can be represented by . Consider its dual, i.e., . Note that we partitioned into block matrices whose -th block is expressed by an matrix:
In the large limit, we have asymptotically
The eigenvalue statistics of are asymptotically evaluated as