On the impact of activation and normalization in obtaining isometric embeddings at initialization

Amir Joudaki, Hadi Daneshmand, Francis Bach

Introduction

Optimization of deep neural networks is a challenging non-convex problem. Various components and optimization techniques have been developed over the last decades to make optimization feasible. Components such as activation functions (Hendrycks and Gimpel, 2016), normalization layers (Ioffe and Szegedy, 2015), and residual connections (He et al., 2016) have significantly influenced network training and have thus become the building blocks of neural networks. The practical success of these components has inspired extensive theoretical studies on the intricate role of weight initialization (Saxe et al., 2013; Daneshmand et al., 2021), normalization (Yang et al., 2019; Kohler et al., 2019; Daneshmand et al., 2021, 2023; Joudaki et al., 2023) and activation layers (Pennington et al., 2018; Joudaki et al., 2023), on neural network training. For example, the training of large language models hinges on carefully utilizing residual connections, normalization layers, and tailored activations (Vaswani et al., 2017; Radford et al., 2018). Noci et al. (2022a) highlight that the absence or improper utilization of these components can substantially slow training.

To delve deeper into the influence of normalization and activation layers on training, one line of research has studied neural networks at initialization (Pennington et al., 2018; de G. Matthews et al., 2018; Jacot et al., 2018; Yang et al., 2019; Li et al., 2022). Several studies have focused on the Gram matrix, which captures the inner products of intermediate representations for a batch of inputs, revealing that Gram matrices become degenerate as the network depth increases (Saxe et al., 2013; Daneshmand et al., 2021; Joudaki et al., 2023). These issue of degeneracy or rank deficiency has been observed in multilayer perceptrons (MLPs) (Saxe et al., 2013; Daneshmand et al., 2020), convolutional networks (Bjorck et al., 2018), and transformers (Dong et al., 2021), posing challenges to the training process (Noci et al., 2022b; Pennington et al., 2018; Xiao et al., 2018). Research indicates that normalization layers can act effectively to circumvent such Gram degeneracy, thereby improving training (Yang et al., 2019; Daneshmand et al., 2020, 2021; Bjorck et al., 2018).

Analyses based of neural networks in the mean-field, i.e., the infinite width regime, have revealed profound insights about initialization by characterizing the local solutions to the Gram dynamics Yang et al. (2019); Pennington et al. (2018). However, these results do not guarantee global convergence towards the mean-field solutions or depend on technical assumptions that are challenging to verify numerically (Daneshmand et al., 2021, 2020; Joudaki et al., 2023). Furthermore, all these theories primarily pertain to the network at initialization, where parameters are typically random and do not necessarily hold during or after the training. In this work, our objective is to bridge these existing gaps.

Building upon existing literature that elucidates the spectral properties of the Gram matrix, we introduce the concept of isometry, which quantifies the similarity of the Gram matrix to the identity. Our initial theoretical finding demonstrates that isometry does not decrease under conditions of (batch and layer) normalization. This finding illuminates the bias of normalization layers towards isometry at various stages, namely initialization, during, and post-training.

We subsequently extend our analysis to explore the impact of non-linear activations on the isometry of intermediate representations in MLPs. Within the mean-field regime, we establish that non-linear activations incline the intermediate representations towards isometry at an exponential rate in depth. Our principal contribution is quantifying this rate by utilizing the Hermit polynomial expansion of activations. Intriguingly, our empirical experiments unveil a correlation between this rate and the convergence of stochastic gradient descent in MLPs equipped with layer normalization and standard activations used in practice.

Related works

A line of research investigates the interplay between signal propagation of the network and training. The existing literature postulates that in order to ensure fast training (Schoenholz et al., 2017; Poole et al., 2016), the network output must be sensitive to input changes, quantified by the spectrum of input-input Jacobean. This hypothesis is employed by Xiao et al. (2018) to train a 10,000-layer CNN using proper weight initialization without stabilizing components such as skip connection or normalization layers. He et al. (2023) demonstrate the critical role of the Jacobean spectra in large language models. In this paper, we analyze the spectrum of Gram matrices that connect to the spectral properties of input-output Jacobean.

Mean-field theory has been extensively used to characterize Gram matrix dynamics in the limit of infinite width. In this setting, the Gram matrix is a fixed point of a recurrence equation that depends on the network architecture (Schoenholz et al., 2017; Yang et al., 2019; Pennington et al., 2018). This fixed-point analysis can provide insights into the structure and spectral properties of Gram matrices in deep neural networks, thereby shedding light on the degeneracy of Gram matrices in networks (Schoenholz et al., 2017; Yang et al., 2019). However, often fixed-points are not unique, and they can be degenerate or non-degenerate Yang et al. (2019). In this paper, we establish a convergence rate to a non-degenerate fixed-point for a family of MLPs.

Batch normalization (Ioffe and Szegedy, 2015) and layer normalization (Ba et al., 2016) layers are widely used in deep neural networks (DNNs) to improve training. Batch normalization ensures that each feature within a layer across a mini-batch has zero mean and unit variance. In contrast, layer normalization centers and divides the output of each layer by its standard deviation. There have been numerous theoretical studies on the effects of batch normalization due to its popularity Yang et al. (2019); Daneshmand et al. (2021); Joudaki et al. (2023). While layer normalization has been the subject of increasing interest due to its application in transformers Xiong et al. (2020), there are relatively fewer studies on its theoretical underpinnings. While we primarily focus on layer normalization, we define and characterize a property that is shared between batch and layer normalization.

A broad spectrum of activation functions such as ReLU (Fukushima, 1969), GeLU (Hendrycks and Gimpel, 2016), SeLU (Klambauer et al., 2017), and Hyperbolic Tangent, and Sigmoid, are used in DNNs. These functions have various computational and statistical consequences in deep learning. Despite this diversity, only the design of SeLU activation is theoretically motivated (Klambauer et al., 2017), while a broader theoretical understanding of activations remains elusive. To address this issue, we develop a theoretical framework to characterize the influence of a broad range of activations on intermediate representations in DNNs.

Preliminaries

Normalization layers.

MLP setup.

While the original ordering of layer normalization and activation is different (Ba et al., 2016), Xiong et al. (2020) show that the above ordering is more effective for large language models.

Gram matrices and isometry.

Let GG be an n×nn\times n positive semi-definite matrix. We define the isometry I(G)\mathcal{I}(G) of GG as the ratio of its normalized determinant to its normalized trace:

Remarkably, I(M)\mathcal{I}(M) has the following properties (see Lemma A.1 for formal statements and proofs):

Scaling-invariant: For all constants c>0c>0, we have I(G)=I(cG)\mathcal{I}(G)=\mathcal{I}(cG).

Range: I∈\mathcal{I}\in where the boundaries and 11 are achieved for to degenerate and identity matrices respectively.

We also define the isometry gap as negative logarithm of isometry −log⁡I(M).-\log\mathcal{I}(M). Based on these properties of isometry, isometry gap lies between and ∞,\infty, with and ∞\infty indicating the perfect isometry (identity matrix) and degenerate matrices respectively. Isometry allows us to establish the inherent bias of normalization layers in the following section.

Isometry bias of normalization

This section is devoted to discussing the remarkable property of isometry in the context of normalization. We present a theorem that formalizes this property, followed by its geometric interpretation and implications.

where ai:=∥xi∥,a_{i}\mathrel{\mathop{\ordinarycolon}}=\|x_{i}\|, and aˉ:=1n∑inai\bar{a}\mathrel{\mathop{\ordinarycolon}}=\frac{1}{n}\sum_{i}^{n}a_{i}.

Isometry can be considered a measurement of the “volume” of the parallelepiped formed by sample vectors, made scale and dimension-independent. The normalization process effectively equalizes the edge lengths of this parallelepiped, enhancing the overall "volume" or isometry, provided there is a variance in the sample norms. Thus, projection onto the unit sphere xi→xi/∥xi∥x_{i}\to x_{i}/\|x_{i}\| makes the edge lengths of the parallelepiped equal while leaving the angles between its edges intact. From this geometric perspective, Theorem 4.1 implies that among parallelepiped with similar angles between their edges and fixed total squared edge lengths, the one with equal edge lengths has the highest volume (and thereby isometry).

The proof of Theorem 4.1 is intuitive for the special case of vectors forming a cube, where max volume is realized when all edge lengths are equal. This fact that maximum volume is achieved when edge lengths are equal can be deduced from the arithmetic vs geometric mean inequality. Strikingly, the proof for the general case is nearly as simple as this special case. The high-level intuition behind the proof is that the determinant allows us to decouple the role of angles and edge lengths in volume formulation. This fact is evident for n=2n=2 in Figure 3.1. Since normalization does not modify the angles between edges, the remainder of the proof falls back onto the case where edges form a cube.

Theorem 4.1 further shows a subtle property of normalization: as long as there is some variation in the sample norms, i.e., ∥xi∥\|x_{i}\|’s are not all equal, the post-normalization Gram has strictly higher isometry than the pre-normalization Gram matrix. It further quantifies the improvement in isometry as a function of variation of norms. Intuitively, terms a‾\overline{a} and 1n∑in(ai−aˉ)2\frac{1}{n}\sum_{i}^{n}(a_{i}-\bar{a})^{2} can be interpreted as the average and variance of sample norms a1,…,an.a_{1},\dots,a_{n}. Thus, a higher variation in the norms aia_{i}’s leads to a larger increase in isometry after normalization.

1 Implications for layer (and batch) normalization

Theorem 4.1 reveals insights into the biases introduced by layer and batch normalization in neural networks, particularly highlighting the improvement in isometry not just limited to initialization but also persistent through the training process.

What makes the above result distinct from related studies (Daneshmand et al., 2021, 2020; Yang et al., 2019) is that the increase in isometry is not limited to random initialization. Thus, layer normalization increases the isometry even during and after training. This calls for future research on the role of this inherent bias in enhanced optimization and generalization performance with batch normalization (Ioffe and Szegedy, 2015; Yang et al., 2019; Lyu et al., 2022; Kohler et al., 2019).

Despite the seemingly vast differences between layer normalization and batch normalization (Lubana et al., 2021), the following corollary shows a link between these two different normalization techniques.

Gram matrices of networks with batch normalization have been the subject of many previous studies at network initialization: it has been postulated that BN prevents rank collapse issue (Daneshmand et al., 2020) and that it orthogonalizes the representations (Daneshmand et al., 2021), and that it imposes isometry (Yang et al., 2019). It is straightforward to verify that orthogonal matrices have the maximum isometry. Thus, the increase in isometry links to the orthogonalization of hidden representation characterized by Daneshmand et al. (2021). While all previous results heavily rely on Gaussian random weights to establish this inherent bias, Corollary 4.3 is not limited to random weights.

2 Empirical validation of Corollary 4.2 in an MLP setup

We can validate Corollary 4.2 by tracking the isometry of various layers of an MLP with layer normalization. Figure 4.1 shows the isometry of intermediate representations in an MLP with layer normalization and hyperbolic tangent on CIFAR10 dataset. Shades in the figure mark layers illustrate that the isometry of the Gram matrix is non-decreasing after each layer normalization layer. We can see in Figure 4.1 that both before (left) and after training (right), the normalization layers maintain or improve isometry. To highlight the fact that the claims of Corollary 4.2 holds at all times and not only for initialization, Figure 4.1 tracks isometry of various layers both at initialization (left) and after training (right). This can be verified by the fact that isometry in the normalization layers (shaded blue) is either stable or increased, which validates Corollary 4.2.

Isometry bias of non-linear activation functions

Analyzing Gram matrix dynamics for non-linear activation is challenging since even small modifications in the scale or shape of activations can lead to significant changes in the representations. A powerful tool to analyze activations is to express activations in the Hermite polynomial basis. Inspired by previous successful applications of Hermite polynomials in neural network analysis (Daniely et al., 2016; Yang, 2019), we explore their impact on the isometry of activation functions.

Hermite polynomial of degree k,k, denoted by Hek(x),\mathit{He}_{k}(x), is defined as

All square-integrable function σ\sigma with respect to the Gaussian kernel, which obeys ∫−∞∞σ(x)2e−x2/2dx<∞\int_{-\infty}^{\infty}\sigma(x)^{2}e^{-x^{2}/2}dx<\infty, can be expressed as a linear combination of Hermite polynomials as σ(x)=∑kckHek(x)\sigma(x)=\sum_{k}c_{k}\mathit{He}_{k}(x) with (see section A for more details):

The subsequent section will discuss how to leverage the Hermite expression of the activation to analyze the dynamics of Gram matrix isometry.

2 Non-linear activations bias Gram dynamics towards isometry

Given activation σ\sigma with Hermite expansion {ck}k≥0,\{c_{k}\}_{k\geq 0}, define its isometry strength βσ\beta_{\sigma} as:

Theorem 5.1 reveals the importance of non-linear Hermite coefficients (ck,k≥2c_{k},k\geq 2) in activation function to ensure βσ>1\beta_{\sigma}>1 and obtain isometry in depth. This connection between βσ\beta_{\sigma} and isometry is the rationale for referring to βσ\beta_{\sigma} as the isometry strength. This constant can be computed in closed form for various activations, as shown in Table 5.1, and for all other activations, it can be computed numerically by sampling.

Implications of our theory

In this section, we elucidate the implications of our theory, beginning with insights into layer normalization through the Hermit expansion of activation functions, followed by an examination of its impact on training.

Layer normalization primarily involves two steps: (i) centering and (ii) normalizing the norms. Through the Hermit expansion of activation functions, we unravel the underlying intricacies of these components and propose alternatives based on insights from the Hermit expansion.

1 Centering and Hermit expansion

2 Normalization and Hermit expansion

3 Isometry strength correlates with SGD convergence rate in shallow MLPs.

Besides the direct consequences of our theory, we observe a striking correlation between the convergence of SGD and isometry strength in a specific range of neural network hyper-parameters. Figure 6.4 shows the convergence of SGD is faster for activations with a significantly larger isometry strength βσ\beta_{\sigma} (see Definition 3) for shallow MLPs, e.g., with 1010 layers or less. We can speculate that this correlation reflects the input-output sensitivity of the networks with higher non-linearity. Surprisingly, this correlation does not extend to deeper networks. This discrepancy between shallow and deep networks regarding SGD convergence may be due to the issue of gradient explosion studied by Meterez et al. (2023). This finding suggests multiple avenues for future research.

Discussion

In this study, we explored the influence of layer normalization and nonlinear activation functions on the isometry of MLP representations. Our findings open up several avenues for future research.

It is worth investigating whether we can impose isometry without layer normalization. Our empirical observations suggest that certain activations, such as ReLU, require layer normalization to attain isometry. In contrast, other activations, which can be considered as “self-normalizing” (e.g., SeLU (Klambauer et al., 2017) and hyperbolic tangent), can achieve isometry with only offset and scale adjustments (see Figure 7.1). We experimentally show how we can replace centering and normalization by leveraging Hermit expansion of activation. Thus, we believe Hermit expansion provides a theoretical grounding to analyze the isometry of SeLU.

Impact of the ordering of normalization and activation layers on isometry.

Theorem 5.1 highlights that the ordering of activation and normalization layers has a critical impact on the isometry. Figure 7.1 demonstrates that a different ordering can lead to a non-isotropic Gram matrix. Remarkably, the structure analyzed in this paper is used in transformers (Vaswani et al., 2017).

Normalization’s role in stabilizing mean-field accuracy through depth.

Numerous theoretical studies conjecture that mean-field predictions may not be reliable for considerably deep neural networks (Li et al., 2021; Joudaki et al., 2023). Mean-field analysis incurs a O(1/width)O(1/\sqrt{\text{width}}) error per layer when the network width is finite. This error may accumulate with depth, making mean-field predictions increasingly inaccurate with an increase in depth. However, Figure 7.2 illustrates that layer normalization controls this error accumulation through depth. This might be attributable to the isometry bias induced by normalization, as proven in Theorem 4.1. Similarly, batch normalization also prevents error propagation with depth by imposing the same isometry (Joudaki et al., 2023). This observation calls for future research on the essential role normalization plays in ensuring the accuracy of mean-field predictions.

Acknowledgements

Amir Joudaki is funded through Swiss National Science Foundation Project Grant #200550 to Andre Kahles, and partially funded by ETH Core funding award to Gunnar Ratsch. Hadi Daneshmand acknowledges support from the NSF TRIPODS program (DMS-2022448).

References

Appendix outline

The appendix is partitioned into four main components, each serving its purpose as described:

Section A details the all the proofs, with most of it dedicated to proof of Theorem 5.1 alongside numerical confirmation of significant steps

Section A.1 elaborates on the basic properties of isometry.

Section A.2 provides an elaborate review of the mean-field Gram dynamics.

Section B outlines our rebuttal responses to the reviews that we chose to leave out of the main text.

Section B.1 gives additional details concerning the experiments reported in the main text and appendix.

Section B.2 explores the effect of gain on isometry, and links the rate to the associated isometry strength.

Section B.3 explores the effect of varying widths of hidden layers on the isometry.

Section B.4 explores the notion of isometry for representations in language models.

Appendix A Proofs

It is straightforward to check isometry obeys the following basic isometry-preserving properties:

For PSD matrix M,M, the isometry defined in (3) obeys the following properties: 1) scale-invariance I(cM)=I(M),\mathcal{I}(cM)=\mathcal{I}(M), 2) only takes value in the unit range I(M)∈\mathcal{I}(M)\in 3) it takes its maximum value if and only if MM is identity I(M)=1  ⟺  M=In,\mathcal{I}(M)=1\iff M=I_{n}, and 3) takes minimum value if and only if MM is degenerate I(M)=0.\mathcal{I}(M)=0.

A.2 Mean-field Gram Dynamics

Recall the mean-field Gram dynamics stated in equation (10):

Assuming that inputs are encoded as columns of XX, we can restate the MLP dynamics as follows

where centering and layer normalization are applied column-wise, as defined in the main text. Observe that Gram matrix of representations can be written as

where ∗ denotes the mean-field regime d→∞.d\to\infty. This concludes the connection between the mean-field Gram dynamics and infinitely wide Gram dynamics.

A.3 Introducing a potential

Here we will introduce a Lyapunov function that enables us to precisely quantify the isometry of activations in deep networks:

Remarkably, γ\gamma obey exhibits an geometric contraction under one MLP layer update, which is stated in the following theorem:

Let GG be PSD matrix with unit diagonals Gii=1,i=1,…,n.G_{ii}=1,i=1,\dots,n. It holds:

where G∗0G^{0}_{*} denotes the input Gram matrix.

In Figure A.1, we experimentally validated the above equation for MLPs with a finite width and activations {tanh ,relu, sigmoid}\{\text{tanh ,relu, sigmoid}\} where we observe that the mean-field analysis well predicates the decay in γ\gamma.

Interestingly, we can connect the Lyapunov γ\gamma to the isometry, by proving an upper and lower based on the determinant GG based on γ(G)\gamma(G), when GG is PSD and has unit diagonals:

For PSD matrix GG with unit diagonals holds:

where the lower bound holds if (n−1)∣Gij∣≤1.(n-1)|G_{ij}|\leq 1.

Now, we are ready to prove the main theorem.

First, note that we have the following transformation

By upper bound of Lemma A.4 we have max⁡i≠j∣Gij∣≤1−det⁡(G)≤1−det⁡(G)/2,\max_{i\neq j}|G_{ij}|\leq\sqrt{1-\det(G)}\leq 1-\det(G)/2, implying γ(G0)≤2det⁡(G0)−1,\gamma(G_{0})\leq 2\det(G_{0})^{-1}, which allows us to

Lower bound. The lower bound is a result of Gershgorin circle theorem [Gershgorin, 1931], which implies that every eigenvalue must be within the disc [1−(n−1)max⁡i≠j∣Gij∣,1+(n−1)max⁡i≠j∣Gij∣].[1-(n-1)\max_{i\neq j}|G_{ij}|,1+(n-1)\max_{i\neq j}|G_{ij}|]. Thus, the determinant is lower-bounded by (1−(n−1)max⁡i≠j∣Gij∣)n.(1-(n-1)\max_{i\neq j}|G_{ij}|)^{n}.

Upper bound. Since GG is PSD, we can write G=X⊤X,G=X^{\top}X, which implies that columns of XX are unit norm ∥xi∥=1,\|x_{i}\|=1, and GG encodes the angles between them cos⁡∠(xi,xj)=⟨xi,xj⟩=Gij.\cos\angle(x_{i},x_{j})=\langle x_{i},x_{j}\rangle=G_{ij}. Furthermore, we have det⁡(G)=vol(X)2,\det(G)=\text{vol}(X)^{2}, where the volume refers to the parallelepiped spanned by the columns of X.X. With this formulation, we write volume recursively as volume spanned by columns 11 up to n−1,n-1, times the projection distance of xnx_{n} from their span

where span refers to the space of all linear combinations of these vectors. Note that projection distance of xix_{i} onto xjx_{j} can be written as sin⁡∠(xi,xj)=1−⟨xi,xj⟩2=1−Gij2.\sin\angle(x_{i},x_{j})=\sqrt{1-\langle x_{i},x_{j}\rangle^{2}}=\sqrt{1-G_{ij}^{2}}. Since the projection distance onto the linear span cannot be greater than projection distance onto a single vector, which is bounded by 1−max⁡i≠jGij2\sqrt{1-\max_{i\neq j}G_{ij}^{2}}. This concludes the upper bound that det⁡(G)≤1−max⁡i≠jGij2\det(G)\leq 1-\max_{i\neq j}G_{ij}^{2}. ∎

A.4 Dual activation and proof of Thm. A.2

According to Daniely et al. , we leverage the notion dual activation associated with the activation function σ,\sigma, which is defined in following.

The following lemma connects the dual activation and its mean-reduced version with the Hermite expansion of σ\sigma:

Finally, we have the tools to prove the first main theorem.

The remainder proof relies on the following contractive property of Gram matrix potential:

Consider activation σ,\sigma, with normalized Hermite coefficients {ck}k≥0.\{c_{k}\}_{k\geq 0}. For all ρ∈(0,1),\rho\in(0,1), the mean-reduced dual activation σˉ\bar{\sigma} obeys

which the right hand-side is strictly larger if some nonlinear coefficient is nonzero ck≠0c_{k}\neq 0 for some k≥2.k\geq 2.

Thus we can apply Lemma A.7 on each element i≠ji\neq j to conclude that

since the inequality holds for any value of i≠j,i\neq j, we can take the maximum over i≠jji\neq jj to write:

Note the ratio is invariant to scaling of ∑k=1∞ck2\sum_{k=1}^{\infty}c_{k}^{2}. Hence, we assume ∑k=1∞ck2=1\sum_{k=1}^{\infty}c_{k}^{2}=1 without loss of generality. With this simplification, we have σˉ(1)=1.\bar{\sigma}(1)=1. For the positive range ρ∈\rho\in we have

By Jensen inequality for convex function x↦∣x∣x\mapsto|x| we have

Because x↦x/(1−x)x\mapsto x/(1-x) is monotonically increasing for x∈x\in, we have

Where we invoked the inequality that was proven for ρ∈,\rho\in, because ∣ρ∣∈.|\rho|\in. ∎

The following lemma, which is a consequence of Mehler’s formula, is at the hart of proof of Lemma A.6:

If X,Y∼N(0,1)X,Y\sim N(0,1) with covariance EXY=ρEXY=\rho we have

Let σ(x)=∑k=0ckHek(x),\sigma(x)=\sum_{k=0}c_{k}\mathit{He}_{k}(x), denote the Hermite expansion of σ.\sigma. Thus, we have

The property can be deduced from Mehler’s formula Mehler . The formula states that

where the m!m! factor difference is due to the definition of Hermite polynomials with an additional 1/m!1/\sqrt{m!} compared to the one used in Mehler’s kernel. Observe that the left hand side is equal to p(x,y)/p(x)p(y)p(x,y)/p(x)p(y), where p(x,y)p(x,y) is the joint PDF of (X,Y)(X,Y), and p(x),p(y)p(x),p(y) are PDF of XX and YY respectively. Therefore, we can take the expectation using the expansion

where in the last line we used the orthogonality property EX∼N(0,1)Hk(x)Hn(X)=δnkE_{X\sim N(0,1)}H_{k}(x)H_{n}(X)=\delta_{nk}. ∎

Appendix B Additional experiments

Experiments that did not require training a model were run on a AMD Ryzen 9 3900X 12-Core CPU, which takes about 5 minutes overall. All experiments that required training a model were trained a single NVIDIA GeForce RTX 3090 GPU, which takes about 1 minutes per each training task.

Figures.

The solid lines in all plots represent the average performance over multiple independent runs, and the shaded regions indicate the confidence intervals. Unless stated otherwise, each average is computed over #10 independent runs.

Codes and reproducibility.

We implemented our experiments in Python using the PyTorch framework Paszke et al. . All the figures are reproducible with the code attached in the supplementary.

Training procedure

For all training-related experiments, the isometry or isometry gap are computed per each batch by sampling a few of the batches randomly, and then averaged over. The epoch ii corresponds to the network at after ii steps of training on the training set of CIFAR10 (epoch means network is at initialization).

Pre-trained large language models

The pre-trained language models and their default configuration was downloaded from Huggingface Wolf et al. library.

B.2 Quantifying the influence of gain on isometry through non-linearity strength

The concept of gain in neural networks is vital and closely connected with the weights initialization. A neural network with properly initialized weights can learn faster, have a lesser chance of getting stuck at sub-optimal solutions, and provide better generalization. The impact of gain can be visualized through the lens of weight initialization strategies such as Xavier normalization [Glorot and Bengio, 2010], which has shown significant effectiveness in optimizing neural networks. These initialization strategies apply a gain value to the weights, which is a scaling factor, to ensure a good signal flow through many layers during the forward and backward passes. The gain value essentially determines the variance of the weights in the initialization stage.

As an extension to our prior investigations, we delve into understanding the influence of gain on isometry, predominantly through our calculated metric, the non-linearity strength, denoted as βσ\beta_{\sigma}, as a function of gain α.\alpha. For certain instances, such as ReLU, sine, and exponential activations, we are capable of deriving βσ(α)\beta_{\sigma}(\alpha) in a closed form. Table B.1 presents a few of these cases.

For a more extensive selection of activation functions, we have numerically computed the non-linearity strength as a function of gain α\alpha, as visualized in Figure B.1. Leveraging Theorem 5.1, these computed values provide an estimation for the isometry strength for various activations. This correlation proves to be remarkably predictive, as shown in Figure B.2. Remarkably, in the case of ReLU activation, its unique characteristics lead to both our closed form βσ\beta_{\sigma} (refer Table B.1) and rate towards isometry (refer Figure B.2) remaining consistent across different values of α\alpha.

Inspired by the results so far, we can compare mean-field centering and normalization to Xavier gain for activations. Figure B.3 demonstrates that all mean-field based gains improve the isometry when compared with Xavier initialization. However, this is markedly stronger for ReLU and leaky ReLU. We can explain this starker contrast by the fact that both activations have a significant offset term c0,c_{0}, which is not corrected by the Xavier initialization.

B.3 Varying width of hidden layers

While in our theoretical setup, we assume the network width is constant across the layers, this is only a choice to streamline our proof and notation. Since our primary result is derived from the mean-field regime, the only criterion for it to hold is for the width to be sufficiently large to approximate the mean-field regime. Our experiments in Figure B.4 substantiate this claim that the specific sizes of hidden layers, as long as they are large, will not impact our main results on the isometry. We empirically validate this for four different configurations and show that the decay of the isometry gap remains largely consistent across these configurations.

B.4 Isometry in pre-trained large language models

Since our theory for normalization is not limited to initialization, we can expand our search for isometry to other architectures. Figure B.5 shows the important role of normalization in the pre-trained GPT2 network. However, we need to adjust the notion of isometry with the architecture of layer norm in a transformer in mind. In fact, the mean and standard deviation are computed over features separately for each token. Thus, to adapt the notion of isometry, we can view each token as a sample and define Gram over different tokens. Thus, isometry here quantifies the similarity between various tokens within one sample. As can be seen in the figure below, LayerNorm layers (shaded in red) in the last six layers of the pre-trained GPT2 increase the isometry between tokens, which is consistent with our theory of layer normalization. It is crucial that our theory holds deterministically, which extends to the pre-trained model.

One caveat in interpreting our results is that, in practice, LayerNorm layers have learnable parameters that make them deviate from our theory. It would be fruitful to study the effects of learned parameters to discern it from the role of centering and normalization for a future study.

B.5 Tracking isometry at initialization and optimization for more activations

Here there are more numerical experiments related to to tracking isometry before and after training.

The following plot shows that isometry gap remains relatively stable.