The local relaxation flow approach to universality of the local statistics for random matrices

Laszlo Erdos, Benjamin Schlein, Horng-Tzer Yau, Jun Yin

Introduction

A central question concerning random matrices is the universality conjecture which states that local statistics of eigenvalues of large N×NN\times N square matrices HH are determined by the symmetry type of the ensembles but are otherwise independent of the details of the distributions. In particular they coincide with that of the corresponding Gaussian ensemble. The most commonly studied ensembles are

(i) hermitian, symmetric and quaternion self-dual matrices with identically distributed and centered entries that are independent (subject to the natural restriction of the symmetry);

(ii) sample covariance matrices of the form H=A∗AH=A^{*}A, where AA is an M×NM\times N matrix with centered real or complex i.i.d. entries.

There are two types of universalities: the edge universality and the bulk universality concerning energy levels near the spectral edges and in the interior of the spectrum, respectively. Since the works of Sinai and Soshnikov , the edge universality is commonly approached via the fairly robust moment method ; very recently an alternative approach was given in .

The bulk universality is a subtler problem. In the simplest case of the hermitian Wigner ensemble, it states that, independent of the distribution of the entries, the local kk-point correlation functions of the eigenvalues (see (2.3) for the precise definition later), after appropriate rescaling and in the N→∞N\to\infty limit, are given by the determinant of the sine kernel

Similar statement is expected to hold for all other ensembles mentioned above but the explicit formulas are somewhat more complicated. Detailed formulas for the different Wigner ensembles can be found e.g., in . The various sample covariance ensembles have the same local statistics for their singular values as the local eigenvalue statistics of the corresponding Wigner ensembles.

For ensembles of hermitian, symmetric or quaternion self-dual matrices that remain invariant under the transformations H→U∗HUH\to U^{*}HU for any unitary, orthogonal or symplectic matrix UU, respectively, the joint probability density function of all the NN eigenvalues can be explicitly computed. These ensembles are typically given by the probability density

where VV is a real function with sufficient growth at infinity and dH{\rm d}H is the flat Lebesgue measure on the corresponding symmetry class of matrices. The eigenvalues are strongly correlated and they are distributed according to a Gibbs measure with a long range logarithmic interaction potential. The joint probability density of the eigenvalues of HH with distribution (1.2) can be computed explicitly:

where β=1,2,4\beta=1,2,4 for hermitian, symmetric and symplectic ensembles, respectively, and const. is a normalization factor. The formula (1.3) defines a joint probability density of NN real random variables for any β≥1\beta\geq 1 even when there is no underlying matrix ensemble. This ensemble is called the invariant β\beta-ensemble. Quadratic VV corresponds to the Gaussian ensembles; we note that these are the only ensembles that are simultaneously invariant and have i.i.d. matrix entries. These are called the Gaussian Orthogonal, Unitary and Symplectic Ensembles (GOE, GUE, GSE for short) in case of β=1,2,4\beta=1,2,4, respectively. Somewhat different choices of VV lead to two other classical ensembles, the Laguerre and the Jacobi ensembles, that also have matrix interpretation for β=1,2,4\beta=1,2,4 (e.g., the Laguerre ensemble corresponds to the Gaussian sample covariance matrices which are also called Wishart matrices), see for more details. The local statistics can be obtained via a detailed analysis of orthogonal polynomials on the real line with respect to the weight function exp⁡(−V(x))\exp(-V(x)). This approach was originally applied to classical ensembles by Dyson , Mehta and Gaudin and Mehta that lead to classical orthogonal polynomials. Later general methods using orthogonal polynomials were developed to tackle a very general class of invariant ensembles by Deift et.al., see and references therein, and also by Bleher and Its and Pastur and Schcherbina .

Many natural matrix ensembles are typically not unitarily invariant; the most prominent examples are the Wigner matrices or the sample covariance matrices mentioned in (i) and (ii). For these ensembles, apart from the identically distributed Gaussian case, no explicit formula is available for the joint eigenvalue distribution. Thus the basic algebraic connection between eigenvalue ensembles and orthogonal polynomials is missing and completely new methods needed to be developed.

The bulk universality for hermitian Wigner ensembles has been established recently in , by Tao and Vu in and in . These works rely on the Wigner matrices with Gaussian divisible distribution, i.e., ensembles of the form

where H^\widehat{H} is a Wigner matrix, VV is an independent standard GUE matrix and ss is a positive constant. Johansson (see also Ben Arous and Péché and the recent paper ) proved the bulk universality for the eigenvalues of such matrices by an asymptotic analysis on an explicit formula for the correlation functions adapted from Brézin-Hikami . Unfortunately, the similar formula for symmetric or quaternion self-dual Wigner matrices, as well as for real sample covariance matrices, is not very explicit and the technique of cannot be extended to prove universality. Complex sample covariance matrices can however be handled with an analogous formula and universality without any Gaussian component is a work in progress .

A key observation of Dyson is that if the matrix H^+sV\widehat{H}+\sqrt{s}V is embedded into a stochastic matrix flow, i.e. one considers H^+V(s)\widehat{H}+V(s) where the matrix elements of V(s)V(s) are independent standard Brownian motions with variance s/Ns/N, then the evolution of the eigenvalues is given by a system of coupled stochastic differential equations (SDE), commonly called the Dyson Brownian motion (DBM) . If we replace the Brownian motions by the Ornstein-Uhlenbeck processes to keep the variance constant, then the resulting dynamics on the eigenvalues, which we still call DBM, has the GUE eigenvalue distribution as the invariant measure. Similar stochastic processes can be constructed for symmetric, quaternion self-dual and sample covariance type matrices, and, in fact, on the level of eigenvalue SDE they can be extended to other values of β\beta (see (5.5) and (5.8) for the precise formulas).

The result of can be interpreted as stating that the local statistics of GUE is reached via DBM for time of order one. In fact, by analyzing the dynamics of DBM with ideas from the hydrodynamical limit, we have extended Johansson’s result to s≫N−3/4s\gg N^{-3/4} . The key observation of is that the local statistics of eigenvalues depend exclusively on the approach to local equilibrium which in general is faster than reaching the global equilibrium. Unfortunately, the identification of local equilibria in still uses explicit representations of correlation functions by orthogonal polynomials (following e.g. ), and the extension to other ensembles is not a simple task.

In we introduced an approach based on a new stochastic flow, the local relaxation flow, which locally behaves like DBM, but has a faster decay to equilibrium. This method completely circumvented explicit formulas and it resulted in proving universality for symmetric Wigner matrices (the method applies to hermitian and quaternion self-dual Wigner matrices as well). As an input of this method, we needed a fairly detailed control on the local density of eigenvalues that could be obtained from our previous works on Wigner matrices .

In this paper we will prove a general theorem which states that as long as the eigenvalues are at most N−1/2−εN^{-1/2-\varepsilon} distance near their classical location on average, the local statistics is universal and in particular it coincides with the Gaussian case for which explicit formulas have been computed. To introduce this flow, denote by γj\gamma_{j} the location of the jj-th eigenvalue that will be defined in (2.12). We first define the pseudo equilibrium measure by

where μN\mu_{N} is the probability measure for the eigenvalue distribution of the corresponding Gaussian ensemble. In case of Wigner matrices, μN\mu_{N} is the measure for the general β\beta ensemble (β≥1\beta\geq 1 and β=2\beta=2 for GUE):

In this setting, it is natural to view eigenvalues as random points and their equilibrium measure as Gibbs measure with a Hamiltonian H{\mathcal{H}}. We will freely use the terminology of statistical mechanics. Note that the additional term WjW_{j} in ωN{\omega}_{N} confines the jj-th point xjx_{j} near its classical location, but the probability w.r.t. the equilibrium measure μN\mu_{N} of the event that xjx_{j} near its classical location will be shown to be very close to 11. Furthermore, we will prove that the local statistics of the measures ωN{\omega}_{N} and μN\mu_{N} are identical in the limit N→∞N\to\infty and this justifies the term pseudo equilibrium measure.

The local relaxation flow is defined to be the reversible flow (or the gradient flow) generated by the pseudo-equilibrium measure. The main advantage of the local relaxation flow is that it has a faster decay to global equilibrium (Theorem 4.2) compared with the DBM. The idea behind this construction can be related to the treatment of metastability in statistical physics. Imagine that we have a double well potential and we wish to treat the dynamics of a particle in one of the two wells. Up to a certain time, say t0t_{0}, the particle will be confined in the well where the particle initially located. However, the potential of this particle, given by the double well, is not convex. A naive idea is to regain the convexity before the time t0t_{0} is to modify the potential to be a single well! Now as long as we can prove that the particle was confined in the initial well up to t0t_{0}, there is no difference between these two dynamics. But the modified dynamics, being w.r.t. a convex potential, can be estimated much more precisely and this estimate can be carried over to the original dynamics up to the time t0t_{0}.

In our case, the convexity of the equilibrium measure μN\mu_{N} is rather weak and in fact, it comes from the quadratic confining potential βxi2/4\beta x_{i}^{2}/4 of (1.6). So the potential is convex, just not “convex enough”. There is no sharp transition like jumping from one metastable state to another as in the double well case. Instead, there are two time scales: in short time the local equilibrium is formed, on longer time, it approaches the global equilibrium. The approach to the local equilibrium is governed by a strong intrinsic convexity in certain directions due to the interactions (see (2.10) later for a precise formula). To reveal this additional convexity, in our previous paper we introduced a pseudo equilibrium measure where we replaced the long range part of the interaction by a mean-field potential term using the classical locations of far away particles. This potential term inherited the intrinsic convexity of the interaction and it could be directly used to enhance the decay to the local statistics. One technical difficulty with this approach was that we needed to handle the singular behavior of the logarithmic interaction potential. In this paper we show that the pseudo equilibrium measure can be defined by adding a Gaussian term. This simple modification turns out to be sufficient and is also model-independent. Since the Gaussian modification is regular, we no longer need to deal with singularities. The price to pay is that we need a slightly stronger local semicircle law which will be treated in Section 8.

The method of local relaxation flow itself proves universality for Wigner matrices with a small Gaussian component sV\sqrt{s}V (typically of variance s≥N−γs\geq N^{-\gamma} with some 0<γ<10<\gamma<1). In other words, we can prove universality for a Wigner ensemble whose single entry distribution (the distribution of its matrix elements) is given by etBu0e^{tB}u_{0}, where BB is the generator of the Ornstein-Uhlenbeck process and u0u_{0} is any initial distribution (We remark that in our approach of decay to equilibrium, the Brownian motion in the construction of DBM is always replaced by the Ornstein-Uhlenbeck process). To obtain universality for Wigner matrices without any Gaussian component, it remains to prove that for a given Wigner matrix ensemble with a single entry distribution ν\nu we can find u0u_{0} and tt such that the eigenvalue distributions of the ensembles given by ν\nu and etBu0e^{tB}u_{0} are very close to each other. By the method of reverse heat flow introduced in , we choose u0u_{0} to be an approximation of e−tBνe^{-tB}\nu. Although the Ornstein-Uhlenbeck evolution cannot be reversed, we can approximately reverse it provided that ν\nu is sufficient smooth and the time is short. This enables us to compare local statistics of Wigner ensembles with and without small Gaussian components assuming that the single entry distribution is sufficiently smooth (see Section 6).

As an application, we will use this method to prove the bulk universality of sample covariance ensembles. The necessary apriori control on the location of eigenvalues will be obtained by a local semicircle law. In addition to sample covariance ensembles, we will outline the modifications needed for proving the bulk universality of symplectic ensembles.

Universality for the local relaxation flow

We consider time dependent permutational symmetric probability measures with density ft(x)f_{t}({\bf{x}}), t≥0t\geq 0, with respect to the measure μ(dx)=μ(x)dx\mu({\rm d}{\bf{x}})=\mu({\bf{x}}){\rm d}{\bf{x}}. The dynamics is characterized by the forward equation

with a given permutation symmetric initial data f0f_{0}. Here the generator LL is defined via the Dirichlet form as

Formally, we have L=12NΔ−12(∇H)∇L=\frac{1}{2N}\Delta-\frac{1}{2}(\nabla{\mathcal{H}})\nabla. In Appendix A we will show that under general conditions on H{\mathcal{H}} the generator can be defined as a self-adjoint operator on an appropriate domain and the dynamics is well defined for any f0∈L1(dμ)f_{0}\in L^{1}({\rm d}\mu) initial data. Strictly speaking, we will consider a sequence of Hamiltonians HN{\mathcal{H}}_{N} and corresponding dynamics LNL_{N} and ft,Nf_{t,N} parametrized by NN, but the NN-dependence will be omitted. All results will concern the N→∞N\to\infty limit.

With a slight abuse of notations, we will sometimes also use μ\mu to denote the density of the measure μ\mu with respect to the Lebesgue measure. The correlation functions of the equilibrium measure are denoted by

We now list our main assumptions on the initial distribution f0f_{0} and on its evolution ftf_{t}. We first define the subdomain

of ordered sets of points x{\bf{x}}. In the application to the sample covariance matrices, we will use the subdomain

Assumption I. The Hamiltonian H{\mathcal{H}} of the equilibrium measure has the form

The condition U′′≥0U^{\prime\prime}\geq 0 can be relaxed to inf⁡U′′>−∞\inf U^{\prime\prime}>-\infty, see remark after (4.11).

Alternatively, in order to discuss the case of the sample covariance matrices, we will also consider the following modification of Assumption I.

Assumption I’. The Hamiltonian H{\mathcal{H}} of the equilibrium measure has the form

where β≥1\beta\geq 1 and cN≥1c_{N}\geq 1. The function UU satisfies the same conditions as in Assumption I.

It is easy to check that the condition (2.7) guarantees that the following bound holds for the normalization constant

with some exponent mm depending on δ\delta.

It follows from Assumption I (or I’) that the Hessian matrix of H{\mathcal{H}} satisfies the following bound:

This convexity bound is the key assumption; our method works for a broad class of general Hamiltonians as long as (2.10) holds. In particular, an arbitrary many-body potential function V(x)V({\bf{x}}) can be added to the Hamiltonians (2.6), (2.8), as long as VV is convex on the open sets ΣN\Sigma_{N} and ΣN+\Sigma_{N}^{+}, respectively. The argument in the proof of the main Theorem 2.1 remains unchanged, but the technical details of the regularization of the singular dynamics (Appendix B) becomes more involved. We do not pursue this direction here since we do not need it for the application for Wigner and sample covariance matrices.

Let γj=γj,N\gamma_{j}=\gamma_{j,N} denote the location of the jj-th point under the limiting density, i.e., γj\gamma_{j} is defined by

We will call γj\gamma_{j} the classical location of the jj-th point. Note that γj\gamma_{j} may not be uniquely defined if the support of ϱ\varrho is not connected but in this case the next Assumption III will not be satisfied anyway.

Assumption III. There exists an ε>0\varepsilon>0 such that

Under Assumption II, the typical spacing between neighboring points is of order 1/N1/N away from the spectral edges, i.e., in the vicinity of any energy EE with ϱ(E)>0\varrho(E)>0. Assumption III guarantees that typically the random points xjx_{j} remain in the N−1/2−εN^{-1/2-\varepsilon} vicinity of their classical location.

where ε\varepsilon is the exponent from Assumption III.

The main general theorem is the following:

for τ=N−2ε+δ\tau=N^{-2\varepsilon+\delta} where ε>0\varepsilon>0 is the exponent from Assumption III.

Suppose in addition to the Assumption I-IV, that there exists an A>0A>0 such that, for any c ′>0c\,^{\prime}>0

for some constants cc and CC only depending on c ′c\,^{\prime}. Then for τ=N−2ε+δ\tau=N^{-2\varepsilon+\delta} we have

This theorem shows that the local statistics of the points xjx_{j} in the bulk with respect to the time evolved distribution ftf_{t} coincides with the local statistics with respect to the equilibrium distribution μ\mu as long as t≫N−2εt\gg N^{-2\varepsilon}. In many applications, the local equilibrium statistics can be explicitly computed and in the b→0b\to 0 limit it becomes independent of EE, in particular this is the case for the classical matrix ensembles (see next section). The restriction on the time t≫N−2εt\gg N^{-2\varepsilon} will be removed by the reverse heat flow argument (see Section 6) for matrix ensembles.

Since the eigenvalues fluctuate at least on a scale 1/N1/N, the best possible exponent in Assumption III is 2ε∼12\varepsilon\sim 1, but we will only be able to prove it for some ε>0\varepsilon>0 for the ensembles considered in this paper. Similarly, the optimal exponent in (2.16) is A∼0A\sim 0. If we use these optimal estimates, 2ε∼12\varepsilon\sim 1, A∼0A\sim 0, and we choose δ=2ε∼1\delta=2\varepsilon\sim 1, thus τ∼1\tau\sim 1, then we can choose b∼N−1b\sim N^{-1}, i.e., we obtain the universality with essentially no averaging in EE. On the other hand, the error estimate is the strongest, of order ∼N−1/2\sim N^{-1/2}, for an averaging on an energy window of size b∼1b\sim 1. These errors become weaker if time τ\tau is reduced. These considerations are not important in this paper, but will be useful when good estimates on ε\varepsilon and AA can be obtained.

Convention: Throughout the paper the letters C,cC,c denote positive constants whose values may change from line to line and they are independent of the relevant parameters. Since we will always take the N→∞N\to\infty limit at the end, all estimates are understood for sufficiently large NN.

Universality for Matrix Ensembles

Now we specialize Theorem 2.1 to Wigner and sample covariance matrices with i.i.d. entries. In the next sections we give the precise definitions of these ensembles; formulas for the equilibrium measure and the dynamics will be deferred until Section 5.

In order to apply Theorem 2.1 to Wigner and sample covariance matrices, we need to check that Assumptions I-IV are satisfied for these ensembles. Assumptions I or I’ are satisfied by the definition of the Hamiltonian, the precise formulas are given in Section 5. Assumption II is satisfied since the density of eigenvalues is given by the Wigner semicircle law (3.7) for Wigner matrices . In case of the sample covariance matrices, the singular values of AA will play the role of xjx_{j}’s and their density is given by the Marchenko-Pastur law (3.14) after an obvious transformation (3.15) . In fact, in Section 8 we prove a local version of the Marchenko-Pastur law in analogy with our previous work on the local semicircle law for Wigner matrices . In Section 9 (Theorem 9.1) we will show that Assumption III is satisfied for these ensembles (more precisely, we will prove that Assumption III is satisfied for sample covariance matrices; the proof for Wigner matrices is analogous, and will not be given in details). Assumption IV will be proved in Lemma 8.1 for the sample covariance matrices, for Wigner matrices the proof was given, e.g., in Theorem 4.6 of . We remark that the assumption that the matrix entries are identically distributed, will only be used in checking Assumptions III and IV. Assumption II holds under much more general conditions on the matrix entries. Finally, the apriori estimate on the entropy Sμ(ft0)S_{\mu}(f_{t_{0}}) follows from the smoothing property of the OU-flow (see Section 5).

To fix the notation, we assume that in the case of real symmetric matrices, the matrix elements of HH are given by

Finally, for the quaternion self-dual case we assume that HH is a 2N2N by 2N2N complex matrix that can be viewed as an N×NN\times N matrix with elements consisting of 2×22\times 2 blocks of the form

The dual of the quaternion qq is defined to be q+:=a−bi−cj−dkq^{+}:=a-b{\bf i}-c{\bf j}-d{\bf k} which corresponds to the hermitian conjugate of the matrix (3.3).

where xkkx_{kk} has a law ν~\widetilde{\nu} with zero expectation and variance 12\frac{1}{2}. The spectrum of HH is doubly degenerate and we will neglect this degeneracy, i.e., we consider only NN real (typically distinct) eigenvalues, x1<x2<…<xNx_{1}<x_{2}<\ldots<x_{N}.

In particular, the typical spacing between neighboring eigenvalues is of order 1/N1/N in the bulk of the spectrum.

We will often need to assume that the distributions ν\nu and ν~\widetilde{\nu} have Gaussian decay, i.e., there exists δ0>0\delta_{0}>0 such that

In several statements we can relax this condition to assuming only subexponential decay, i.e., that there exists δ0>0\delta_{0}>0 and γ>0\gamma>0 such that

For some statements we will need to assume that the measures ν\nu, ν~\widetilde{\nu} satisfy the logarithmic Sobolev inequality, i.e., for any density h≥0h\geq 0 with ∫hdν=1\int h{\rm d}\nu=1 it holds that

and a similar bound holds for ν~\widetilde{\nu}. We remark that (3.10) implies (3.8), see, e.g. .

2 Sample Covariance Matrix

The real sample covariance matrix ensemble consists of symmetric N×NN\times N matrices of the form H=A∗AH=A^{*}A. Here AA is an M×NM\times N real matrix with d=N/Md=N/M fixed and we assume that 0<d<10<d<1. The elements of AA are given by

Moreover, analogously to (3.6), the empirical density of eigenvalues converges weakly in probability to the Marchenko-Pastur law

Most of the analysis will be done for the singular values of AA that are denoted by x=(x1,…,xN){\bf{x}}=(x_{1},\ldots,x_{N}). They are supported asymptotically in [λ−,λ+][\sqrt{\lambda_{-}},\sqrt{\lambda_{+}}] and therefore the typical spacing between neighboring singular values is of order 1/N1/N. Their empirical density converges to

We remark that the assumption that ν\nu is symmetric is used only at one technical step, namely when we refer to the large deviation result for the extreme eigenvalues of the sample covariance matrices in (see Lemma 9.2 below). The similar result for Wigner matrices has been proven without the symmetry condition, see Theorem 1.4 in .

3 Main Theorems

With the remarks at the beginning of Section 3, Theorem 2.1 applies directly to prove universality for Wigner and sample covariance ensembles with a small Gaussian component; we will not state these theorems separately. To remove the small time restriction from Theorem 2.1, we will apply the reverse heat flow argument. This will give our main result:

Here μ\mu denotes the probability measure of the eigenvalues of the appropriate Gaussian ensemble, i.e. GUE, GOE, GSE for the case of hermitian, symmetric, and, respectively, quaternion self-dual Wigner matrices; and the ensembles of real or complex sample covariance matrices with Gaussian entries (Wishart ensemble) in case of the covariance matrices A∗AA^{*}A. These measures are given in (1.6), with β=1,2,4\beta=1,2,4, for Wigner matrices, and, expressed in terms of singular values, in (5.6), with β=1,2\beta=1,2, for sample covariance matrices.

Remark 1.: In the case of symmetric and hermitian Wigner matrices, the condition (3.16) can be removed by applying the Four-moment theorem of Tao and Vu (Theorem 15 of ) as in the proof of Corollary 2.4 of . Similar remark applies to the sample covariance ensembles and to the quaternion self-dual Wigner ensemble provided the corresponding Four-moment theorem is established.

We also remark that a manuscript by Ben-Arous and Péché with a similar statement is in preparation for complex sample covariance matrices that holds for a fixed E′E^{\prime}, i.e., without averaging over the energy parameter in (3.17).

Remark 2.: After the first version of this manuscript was posted on the arxiv, the question that whether the four moment theorem for sample covariance matrices holds was settled in . In particular, gives an alternative proof of the universality of local statistics for the complex sample covariance ensemble when combined with the result of . For the real sample covariance ensemble the universality was established for distributions whose first four moments match the standard Gaussian variable. An important common ingredient to both our approach and that of is the local Marchenko-Pastur law, established in Proposition 8.1; a slightly different version suitable for the application to prove the four moment theorem is proved in .

The four moment theorem in compares the distributions of individual eigenvalues for two different ensembles. For our application to the correlation functions and gap distributions, an alternative approach is to use the recent Green function comparison theorem . This will also remove the smoothness and logarithmic Sobolev inequality restrictions in Theorem 3.1.

We now state our result concerning the eigenvalue gap distribution both for Wigner and sample covariance ensembles. For any s>0s>0 and EE with ρ(E)>0\rho(E)>0 we define the density of eigenvalue pairs with distance less than s/Nϱ(E)s/N\varrho(E) in the vicinity of EE by

Consider an N×NN\times N Wigner or sample covariance matrix as in Theorem 3.1 such that the probability measure dν=u0dx{\rm d}\nu=u_{0}{\rm d}x of the matrix elements satisfies the logarithmic Sobolev inequality (3.10) and, additionally, ν\nu is symmetric in the sample covariance matrix case. Suppose that the initial density u0u_{0} satisfies

where μ\mu is the probability measure of the eigenvalues of the appropriate Gaussian ensemble, as in Theorem 3.1.

Theorem 3.2 shows that, in particular, the probability to find no eigenvalue in the interval [E,E+α/(ϱ(E)N)][E,E+\alpha/(\varrho(E)N)] is asymptotically the same as in the corresponding classical Gaussian ensemble. Theorems 3.1 and 3.2 will follow from Theorem 2.1 and the reverse heat flow argument that we present in Section 6. We remark that the additional condition on the symmetry of ν\nu in the case of sample covariance matrices stems from using a result from on the lowest eigenvalue of these matrices, see Lemma 9.2.

Theorem 3.2 can be proven directly from Theorem 4.1 since the test functions of the form

determine the distribution of the random variable Λ(E;s)\Lambda(E;s) uniquely. Here we take JJ to be the set

Local Relaxation Flow

Suppose that the Hamiltonian H{\mathcal{H}} given in (2.6) satisfies the convexity bound (2.10) with β≥1\beta\geq 1. Let ftf_{t} be the solution of the forward equation (2.1) with an initial density f0f_{0}. Fix a positive ε>0\varepsilon>0, set t0=N−2εt_{0}=N^{-2\varepsilon} and define

Then for any sufficiently small ε′>0\varepsilon^{\prime}>0, there exist constants C,cC,c, depending only on ε′\varepsilon^{\prime} and GG such that for any J⊂{1,2,…,N−mn}J\subset\{1,2,\ldots,N-m_{n}\} and for any τ≥3t0=3N−2ε\tau\geq 3t_{0}=3N^{-2\varepsilon}, we have

where ∣J∣|J| is the number of the elements in JJ.

The proof of this theorem is similar but much simpler than that of Theorem 2.1 of . The estimate (4.3) improves slightly over the similar estimate in by a factor ∣J∣/N|J|/N due to the improvement in (4). Theorem 2.1 will follow from the fact that in case τ≥N−2ε+δ\tau\geq N^{-2\varepsilon+\delta}, the assumption (2.13) guarantees that

with the choice ε′=δ/3\varepsilon^{\prime}=\delta/3 and using ∣J∣≤N|J|\leq N. More precise error bound will be obtained by relating bb to ∣J∣|J|. Therefore the local statistics of observables involving eigenvalue differences coincide in the N→∞N\to\infty limit. To complete the proof of Theorem 2.1, we will have to show that the convergence of the observables Gi,m{\mathcal{G}}_{i,{\bf{m}}} is sufficient to identify the correlation functions of the xix_{i}’s in the sense prescribed in Theorem 2.1. The details will be given in Section 7.

Proof of Theorem 4.1. Without loss of generality we can assume in the sequel that f0∈L∞(dμ)f_{0}\in L^{\infty}({\rm d}\mu). To see this, note that any f0∈L1(dμ)f_{0}\in L^{1}({\rm d}\mu) can be approximated by a sequence of bounded functions f0(k)f_{0}^{(k)} in L1L^{1}-norm with arbitrary precision and the dynamics is a contraction in L1L^{1} (see Appendix A), thus fτf_{\tau} and fτ(k)f_{\tau}^{(k)} are arbitrarily close in L1L^{1}. Since GG is bounded on the left hand side of (4.3), this is sufficient to pass to the limit k→∞k\to\infty.

Every constant in this proof depends on ε′\varepsilon^{\prime} and GG, and we will not follow the precise dependence. We can assume that ε′<ε\varepsilon^{\prime}<\varepsilon. Given τ>0\tau>0, we define

Notice that the choice of RR depending on τ\tau which is the main reason that τ\tau appears in the denominator on the right hand side of (4.3). We now introduce the pseudo equilibrium measure, ωN=ω=ψμN{\omega}_{N}={\omega}=\psi\mu_{N}, defined by

where Z~\widetilde{Z} is chosen such that ω{\omega} is a probability measure, in particular ω=e−NH~/Z~\omega=e^{-N\widetilde{\mathcal{H}}}/\widetilde{Z} with

Note that the additional term WjW_{j} confines the jj-th point xjx_{j} near its classical location. We will prove that the probability w.r.t. the equilibrium measure μN\mu_{N} of the event that xjx_{j} near its classical location is very close to 11. Thus there is little difference between the two measures ωN{\omega}_{N} and μN\mu_{N} and in fact, we will prove that their local statistics are identical in the limit N→∞N\to\infty. The main advantage of the pseudo equilibrium measure comes from the fact that it has a faster decay to global equilibrium as shown in Theorem 4.2.

Since the additional potential WjW_{j} is uniformly convex with

Here we have used U′′≥0U^{\prime\prime}\geq 0 in the last estimate. If this assumption is replaced by

for some constant MM independent of NN, then there will be an extra term −M∥v∥2-M\|{\bf{v}}\|^{2} in (4.10). Assuming τ≤Nε′\tau\leq N^{\varepsilon^{\prime}}, we have R≤N−ε′/2R\leq N^{-\varepsilon^{\prime}/2}, then this extra term can be controlled by the R−2R^{-2} term and the same proof will go through. Since for the applications in this paper, the condition U′′≥0U^{\prime\prime}\geq 0 is satisfied, we will not use this remark here.

The R−2R^{-2} in the first term comes from the additional convexity of the local interaction and it enhances the “local Dirichlet form dissipation”. In particular we have the uniform lower bound

This guarantees that the relaxation time to equilibrium ω{\omega} for the L~\widetilde{L} dynamics is bounded above by CR2CR^{2}. We recall the definition of the relative entropy of ff with respect to any probability measure dλ{\rm d}\lambda

The first ingredient to prove Theorem 4.1 is the analysis of the local relaxation flow which satisfies the logarithmic Sobolev inequality and the following dissipation estimate. Its proof follows the standard argument in (used in this context in Section 5.1 of ). In Appendix B we will explain how to extend this argument onto the subdomain ΣN\Sigma_{N}. Here we only remark that the key inputs are the convexity bounds (4.10, 4.12) on the Hessian of H~\widetilde{\mathcal{H}} (4.10).

Suppose (4.10) holds. Consider the forward equation

with an initial condition q0q_{0} and with the reversible measure ω\omega. Assume that q0∈L∞(dω)q_{0}\in L^{\infty}({\rm d}{\omega}). Then we have the following estimates

with a universal constant CC. Thus the time to equilibrium is of order R2R^{2}:

The estimate (4.15) on the second term in (4.10) plays a key role in the next theorem.

Proof. For simplicity, we will consider the case when m=(1,2,…n){\bf{m}}=(1,2,\ldots n), the general case easily follows by appropriately redefining the function GG. Let qtq_{t} satisfy

with an initial condition qq. Thanks to the exponential decay of the entropy on time scale τ≫R2\tau\gg R^{2}, see (4.17), and the entropy bound on the initial state qq, the difference between the local statistics w.r.t. qτωq_{\tau}{\omega} and q∞ω=ωq_{\infty}{\omega}={\omega} is subexponentially small in NN,

giving the second term on the r.h.s. of (4.18). To compare qq with qτq_{\tau}, by differentiation, we have

Here we used the definition of L~\widetilde{L} from (4.7) and note that the 1/N1/N factor present in (4.7) cancels the factor NN from the argument of GG (4.2). From the Schwarz inequality and ∂q=2q∂q\partial q=2\sqrt{q}\partial\sqrt{q}, the last term is bounded by

since GG is smooth and compactly supported. This proves Theorem 4.3.

As a comparison to Theorem 4.3, we state the following result which can be proved in a similar way.

Notice that by exploiting the local Dirichlet form dissipation coming from the second term on the r.h.s. of (4.14), we have gained the crucial factor N−1/2N^{-1/2} in the estimate (4.18) compared with (4.20).

The final ingredient to prove Theorem 4.1 is the following entropy and Dirichlet form estimates.

Suppose that (2.10) holds and recall τ=R2Nε′≥3t0\tau=R^{2}N^{\varepsilon^{\prime}}\geq 3t_{0} with t0=N−2εt_{0}=N^{-2\varepsilon}. Let gt=ft/ψg_{t}=f_{t}/\psi so that Sμ(ft∣ψ)=Sω(gt)S_{\mu}(f_{t}|\psi)=S_{\omega}(g_{t}). Assume that Sμ(ft0)≤CNmS_{\mu}(f_{t_{0}})\leq CN^{m} with some fixed mm. Then the entropy and the Dirichlet form satisfy the estimates:

Proof. Recall that ∂tft=Lft\partial_{t}f_{t}=Lf_{t}. The standard estimate on the entropy of ftf_{t} with respect to the invariant measure is obtained by differentiating the entropy twice and using the logarithmic Sobolev inequality. The entropy and the Dirichlet form in (4.21) are, however, computed with respect to the measure ω{\omega}. This yields the additional second term in the following identity that holds for any probability density ψt\psi_{t}:

where gt=ft/ψtg_{t}=f_{t}/\psi_{t}. In our application we set ψt\psi_{t} to be time independent, ψt=ψ=ω/μ\psi_{t}=\psi={\omega}/\mu, hence we have

Since ω{\omega} is invariant, the middle term on the right hand side vanishes, and from the Schwarz inequality

To obtain the first inequality in (4.21), we integrate (4.24) from t0=N−2εt_{0}=N^{-2\varepsilon} to τ/2\tau/2, using that τ=R2Nε′\tau=R^{2}N^{\varepsilon^{\prime}} and Sω(gt0)≤CNm+N2QS_{\omega}(g_{t_{0}})\leq CN^{m}+N^{2}Q with some finite mm, depending on ε\varepsilon. This apriori bound follows from

where we used (2.9) and (4.6). The second inequality in (4.21) can be obtained from the first one by integrating (4.22) from t=τ/2t=\tau/2 to t=τt=\tau and using the monotonicity of the Dirichlet form in time.

Finally, we complete the proof of Theorem 4.1. Recall that τ=R2Nε′\tau=R^{2}N^{\varepsilon^{\prime}} and t0=N−2εt_{0}=N^{-2\varepsilon}. Choose qτ:=gτ=fτ/ψq_{\tau}:=g_{\tau}=f_{\tau}/\psi as density qq in Theorem 4.3. The condition qτ∈L∞q_{\tau}\in L^{\infty} can be guaranteed by the approximation argument from the beginning of the proof of Theorem 4.1. Then Theorem 4.5, Theorem 4.3 together with (4.25) and the fact that Λτ=Qτ−1N2ε′\Lambda\tau=Q\tau^{-1}N^{2\varepsilon^{\prime}} directly imply that

i.e., the local statistics of fτμf_{\tau}\mu and ω{\omega} can be compared. Clearly, equation (4.26) also holds for the special choice f0=1f_{0}=1 (for which fτ=1f_{\tau}=1), i.e., local statistics of μ\mu and ω{\omega} can also be compared. This completes the proof of Theorem 4.1.

Equilibrium measure and Dyson Brownian motion

where β≥1\beta\geq 1 is an arbitrary parameter, i.e., this corresponds to choosing U(x)=x2/4U(x)=x^{2}/4 in (2.6). With a slight abuse of notations we will use μ\mu for both the measure dμ{\rm d}\mu and its density e−NHβ/Zβe^{-N{\mathcal{H}}_{\beta}}/Z_{\beta} with respect to the Lebesgue measure. The specific value β=1,2,4\beta=1,2,4 correspond to the GUE, GOE and GSE ensembles, respectively.

acting on L2(μ)L^{2}(\mu). The measure μ\mu is invariant and reversible with respect to the dynamics generated by LL. Define the Dirichlet form and entropy by

Let ftdμf_{t}{\rm d}\mu denote the probability measure on the set ΣN\Sigma_{N} at the time tt with the given generator LL. Then ftf_{t} satisfies the forward equation

with initial condition f0f_{0}. This dynamics is the Dyson Brownian motion.

The Dyson Brownian motion is the corresponding system of stochastic differential equations for the vector x(t){\bf x}(t) that is given by

The treatment of the sample covariance ensembles is fully analogous, but the formulas change slightly. We use the convention in the sample covariance case that xix_{i} denotes the singular values of AA and λi=xi2\lambda_{i}=x_{i}^{2} are the eigenvalues of A∗AA^{*}A. Most of the formulas will be in terms of xix_{i}’s; in particular we consider the joint distribution function f0(x)f_{0}({\bf{x}}) of the singular values. The invariant measure for the singular values is given by (c.f. (5.1)):

where d=M/Nd=M/N and β=1\beta=1 when XX is a real matrix, β=2\beta=2 when XX is a complex matrix. This formula can be obtained by direct calculation (see also Proposition 2.16 of or Fig. 1 of after appropriate rescaling). Define the generator (c.f. (5.2))

Finally, the stochastic differential equation is given by (c.f. (5.5))

In applications to Wigner matrices (β=1,2,4\beta=1,2,4), f0dμf_{0}{\rm d}\mu will be the joint probability density of the eigenvalues of the initial hermitian, symmetric or quaternion self-dual Wigner matrix H^\widehat{H}. The limiting density is the Wigner semicircle law given in (3.7). The Dyson Brownian motion describes the eigenvalues of the matrix valued process

with H0=H^H_{0}=\widehat{H}. Here Bt{\bf B}_{t} is a symmetric, hermitian or quaternion self-dual matrix-valued process whose offdiagonal elements are standard real, complex or quaternion Brownian motions with variance one and the diagonal elements of are real Brownian motions with variance 2,12,1 and 12\frac{1}{2}, in case β=1,2,4\beta=1,2,4, respectively. More precisely, let utu_{t} denote the density function of the distribution of one real component of the (ij)(ij)-th entry of HtH_{t}, i<ji<j (there are two real components for the hermitian matrices and four for the quaternion matrices), then

Let γ(dx)=γ(x)dx:=(β/2π)1/2e−βx2/2dx\gamma({\rm d}x)=\gamma(x){\rm d}x:=(\beta/2\pi)^{1/2}e^{-\beta x^{2}/2}{\rm d}x denote the reversible measure for this process. The diagonal elements evolve according to an OU process with twice variance. For any t≥0t\geq 0, the solution to (5.9), HtH_{t}, has the same distribution as

The generator of the induced stochastic process on the eigenvalues is given by (5.2). The equilibrium measure μ\mu is the GUE, GOE or GSE eigenvalue distribution. Theorem 2.1 thus says in this case that the local eigenvalue statistics of a Wigner random matrix with a small Gaussian component coincides with the local statistics of the corresponding Gaussian ensemble. The entropy condition on Sμ(ft0)S_{\mu}(f_{t_{0}}) in Theorem 2.1 can be easily obtained by

In the real or complex sample covariance case (β=1,2\beta=1,2), the matrix elements of AA evolve according to the OU process (5.10), i.e. AtA_{t} has the same distribution as

where WW is an M×NM\times N matrix whose elements are i.i.d real or complex Gaussian variables with mean 0 and variance 1/β1/\beta.

Reverse heat flow

To remove the short time restriction from Theorem 2.1 in case of Wigner and sample covariance ensembles and to prove Theorems 3.1 and 3.2, we apply the reverse heat flow argument, presented first in and used also in Corollary 2.4 of .

For fixed β=1,2\beta=1,2 or 4, recall the Ornstein-Uhlenbeck process from (5.10) with the reversible Gaussian measure γ(dx)\gamma({\rm d}x). Let uu be a positive density with respect to γ\gamma, i.e., ∫udγ=1\int u{\rm d}\gamma=1 and we write u(x)=exp⁡(−V(x))u(x)=\exp(-V(x)). Suppose that for any KK fixed there are constants C1,C2C_{1},C_{2} depending on KK such that

and the measure dν=udγ{\rm d}\nu=u{\rm d}\gamma satisfies the subexponential decay condition. We will apply this for the initial distribution dν=u0(x)dx{\rm d}\nu=u_{0}(x){\rm d}x, so uu and u0u_{0} differ by a Gaussian factor.

Suppose that ν=uγ\nu=u\gamma satisfies the subexponential decay condition and (6.1) for some KK. Then there is a small constant αK\alpha_{K} depending on KK such that for t≤αKt\leq\alpha_{K} there exists a probability density gtg_{t} with mean zero and variance 12\frac{1}{2} such that

for some C>0C>0 depending on KK. Furthermore, gtg_{t} can be chosen such that if the logarithmic Sobolev inequality (3.10) holds for the measure ν=uγ\nu=u\gamma, then it holds for gtγg_{t}\gamma as well, with the logarithmic Sobolev constant changing by a factor of at most 22.

Furthermore, let B=B⊗n{\mathcal{B}}=B^{\otimes n}, F=u⊗nF=u^{\otimes n} with some n≤CN2n\leq CN^{2}. Denote by Gt=gt⊗nG_{t}=g_{t}^{\otimes n}. Then we also have

In order to prove Theorems 3.1 and 3.2, appropriate observables JJ need to be chosen that depend on the matrix elements via the eigenvalues to express the quantities in (3.17) and (3.20). It is easy to see that ∥J∥∞\|J\|_{\infty} may grow at most polynomially in NN. But we can always choose KK large enough to compensate for it with the choice t=N−2ε+δt=N^{-2\varepsilon+\delta} allowed in Theorem 2.1. Here the verifications of the Assumptions I-IV of Theorem 2.1 were explained at the beginning of Section 3. This completes the proof of our main theorems.

Proof of Proposition 6.1. Define θ(x)=θ0(tαx)\theta(x)=\theta_{0}(t^{\alpha}x) with some small positive α>0\alpha>0 depending on KK, where θ0\theta_{0} is a smooth cutoff function satisfying θ0(x)=1\theta_{0}(x)=1 for ∣x∣≤1|x|\leq 1 and θ0(x)=0\theta_{0}(x)=0 for ∣x∣≥2|x|\geq 2. Set

By assumption (6.1), hsh_{s} is positive and

for any s≤ts\leq t if tt is small enough. To see this, take, e.g., K=2K=2 and we have

where we have used α≪1\alpha\ll 1, s≤ts\leq t and the assumption (6.1).

Define vs=esBhsv_{s}=e^{sB}h_{s} and by definition, v0=uv_{0}=u. Then

Since the Ornstein-Uhlenbeck is a contraction in L1(dγ)L^{1}({\rm d}\gamma), together with (6.1), we have

for sufficiently small tt. To estimate the last two terms, we also used that on the support of θ−1\theta-1 the measure dγ{\rm d}\gamma decays subexponentially in tt.

Notice that hth_{t} may not be normalized as a probability density w.r.t. γ\gamma but this can be easily adjusted. To compute this normalization, take for example, K=1K=1 and we have, by using s≤tαs\leq t^{\alpha},

The last term is bounded by O(tM)O(t^{M}) for any M>0M>0 due to that u(x)γu(x)\gamma has a subexponential decay and using the assumption (6.1) on VV.

We have proved that there is a constant ct=1+O(tM)c_{t}=1+O(t^{M}), for any M>0M>0 positive, such that cthtc_{t}h_{t} is a probability density. Clearly,

and the same formulas hold if hth_{t} is replaced by vtv_{t} since the OU flow preserves expectation and variance. Let gtg_{t} be defined by

Then gtg_{t} is a probability density w.r.t. γ\gamma with zero mean and variance β−1\beta^{-1}. It is easy to check that the total variation norm of ht−gth_{t}-g_{t} is smaller than any power of tt. Using again the contraction property of etBe^{tB} and (6.5), we get

Now we check the LSI constant for gtg_{t}. Recall that gtg_{t} was obtained from hth_{t} by translation and dilation. By definition of the LSI constant, the translation does not change it. The dilation changes the constant, but since our dilation constant is nearly one, the change of LSI constant is also nearly one. So we only have to compare the LSI constants between dν=udγ{\rm d}\nu=u{\rm d}\gamma and cthtdγc_{t}h_{t}{\rm d}\gamma. From (6.4) and that ctc_{t} is nearly one, the LSI constant changes by a factor less than 22. This proves the claim on the LSI constant.

and this completes the proof of Proposition 6.1.

Proof of Theorem 2.1

We will set Yi,m=0Y_{i,{\bf{m}}}=0 if i+mn>Ni+m_{n}>N. Our goal is to estimate the difference

Let MM be an NN-dependent parameter chosen at the end of the proof, in fact, MM will be chosen as M=NcM=N^{c} with some small positive exponent c>0c>0, depending on nn. Let

and note that ∣Sn(M)∣≤Mn−1|S_{n}(M)|\leq M^{n-1}. We have the simple bound Θ≤ΘM(1)(τ)+ΘM(2)(τ)+ΘM(2)(∞)\Theta\leq\Theta_{M}^{(1)}(\tau)+\Theta_{M}^{(2)}(\tau)+\Theta_{M}^{(2)}(\infty) where

Note that ΘM(2)(∞)\Theta_{M}^{(2)}(\infty) is the same as ΘM(2)(τ)\Theta_{M}^{(2)}(\tau) but with fτf_{\tau} replaced by the constant 11, i.e., f∞dμf_{\infty}{\rm d}\mu is the equilibrium.

After performing the dE′{\rm d}E^{\prime} integration, we will eventually apply Theorem 4.1 to the function

For any EE and 0<ξ<b0<\xi<b define sets of integers J=JE,b,ξJ=J_{E,b,\xi} and J±=JE,b,ξ±J^{\pm}=J^{\pm}_{E,b,\xi} by

where γi\gamma_{i} was defined in (2.12). Clearly J−⊂J⊂J+J^{-}\subset J\subset J^{+}. With these notations, we have

The error term ΩJ,m+\Omega^{+}_{J,{\bf{m}}}, defined by (7.6) indirectly, comes from those i∉J+i\not\in J^{+} indices, for which xi∈[E−b+O(N−1),E+b+O(N−1)]x_{i}\in[E-b+O(N^{-1}),E+b+O(N^{-1})] since Yi,m(E′,x)=0Y_{i,{\bf{m}}}(E^{\prime},{\bf{x}})=0 unless ∣xi−E′∣≤C/N|x_{i}-E^{\prime}|\leq C/N, the constant depending on the support of OO. Thus

for any sufficiently large NN, assuming ξ≫1/N\xi\gg 1/N and using that OO is a bounded function. The additional N−1N^{-1} factor comes from the dE′{\rm d}E^{\prime} integration. Taking the expectation with respect to the measure fτdμf_{\tau}{\rm d}\mu, we get

using Assumption III (2.13). We can also estimate

where the error term ΞJ,m+\Xi^{+}_{J,{\bf{m}}}, defined by (7.9), comes from indices i∈J−i\in J^{-} such that xi∉[E−b,E+b]+O(1/N)x_{i}\not\in[E-b,E+b]+O(1/N). It satisfies the same bound (7.8) as ΩJ,m+\Omega^{+}_{J,{\bf{m}}}.

By the continuity of ϱ\varrho, the density of γi\gamma_{i}’s is bounded by CNCN, thus ∣J+∖J−∣≤CNξ|J^{+}\setminus J^{-}|\leq CN\xi and ∣J∖J−∣≤CNξ|J\setminus J^{-}|\leq CN\xi. Therefore, summing up the formula (7.5) for i∈Ji\in J, we obtain from (7.6) and (7.9)

for each m∈Sn{\bf{m}}\in S_{n}. A similar lower bound can be proved analogously and we obtain

Adding up (7.12) for all m∈Sn(M){\bf{m}}\in S_{n}(M), we get

and the same estimate holds for the equilibrium, i.e., if we set τ=∞\tau=\infty in (7.13). We now subtract the these two formulas and apply (4.3) from Theorem 4.1 to each summand on the second term in (7.13). Choosing ξ=N−(1+2ε)/3\xi=N^{-(1+2\varepsilon)/3} to minimize the two error terms involving ξ\xi, we conclude that

where we have used τ=N−2ε+δ\tau=N^{-2\varepsilon+\delta} and that ∣J∣≤CNb|J|\leq CNb.

Together with (7.14), we have thus proved that

Choosing MM such that Mn=Nε′M^{n}=N^{\varepsilon^{\prime}} and then choose kk large enough so that the last term CaMk−1\frac{C_{a}}{M^{k-1}} is smaller than, say, N−2N^{-2}. We have thus proved that

for τ=N−2ε+δ\tau=N^{-2\varepsilon+\delta} and this concludes (2.15).

For the proof of (2.17), we choose ξ≥2N−1+A\xi\geq 2N^{-1+A}, and then by using (2.16) we can estimate ΩJ,m+\Omega^{+}_{J,{\bf{m}}} directly as

for any K>0K>0, instead of (7.8). Therefore, the estimate on the right hand side of (7.12) and the subsequent estimates can be replaced by

provided ξ≥2N−1+A\xi\geq 2N^{-1+A}. Choosing ξ=2N−1+A\xi=2N^{-1+A} and following the same proof, we can improve the estimate (7.18) to

for τ=N−2ε+δ\tau=N^{-2\varepsilon+\delta}. This proves (2.17) and we have completed the proof of Theorem 2.1.

Local Marchenko-Pastur law

In this section we establish that the empirical density of eigenvalues for sample covariance matrices is close to the Marchenko-Pastur law even on short scale. We do this by controlling the difference of the Stieltjes transform, establishing results analogous to Theorem 4.1. and Proposition 4.2 of . In this section, we focus on 0<d=N/M<10<d=N/M<1, in particular the lower spectral edge λ−>0\lambda_{-}>0. The constants appearing in this subsection may depend on dd.

Before the detailed proof, we explain the main steps of the argument which is similar to the method we have successively developed in . The proof given here is somewhat complicated by fact that the matrix elements themselves are not independent but are generated as a quadratic expression of independent random variables. The first step, Lemma 8.1, is an apriori bound on the local density on short scales, η≫1/N\eta\gg 1/N, using resolvent expansion and a large deviation principle for quadratic forms. Expressing the resolvent of HH in terms the resolvents of its minors, we obtain a self-consistent equation (8.20) for the Stieltjes transform mNm_{N} of the eigenvalues. This equation is very close to the defining quadratic equation of the Stieltjes transform mWm_{W} of the Marchenko-Pastur law, see (8.7), with a perturbation term Y(z)Y(z). This term can be estimated by large deviation arguments and using the a-priori bound on the local density. Then in Lemma 8.3 we investigate the stability of the self-consistent equation for mWm_{W}. Although the perturbed equation has two solutions, only one of them can be close to mNm_{N}. To select the correct solution, we use a continuity argument in the spectral parameter zz. For z=z0z=z_{0} with a large imaginary part, say z0=10+5iz_{0}=10+5i, the explicit formula (8.27) for the solution can be directly analyzed. For zz approaching to the real axis, we prove that the two unperturbed solutions remain far away from each other (8.35). Since the perturbed solutions are also continuous in the spectral parameter, for a sufficiently small perturbation they must remain in the vicinity of the correct solution of the unperturbed equation.

Let 0<E<100<E<10 and 0<d<10<d<1. Consider the interval Iη=[E−η,E+η]I_{\eta}=[E-\eta,E+\eta]. Let NI{\mathcal{N}}_{I} denote the number of eigenvalues of H=A∗AH=A^{*}A in the interval IηI_{\eta}. Suppose that N−1+ε≤η≤E/2N^{-1+\varepsilon}\leq\eta\leq E/2, for some ε>0\varepsilon>0. Then there exist constants C,c>0C,c>0 such that

for all N,KN,K large enough (independent of EE).

We remark that the assumption on η\eta can be relaxed to CN−1≤η≤E/2CN^{-1}\leq\eta\leq E/2. But we do not need this result here. For details, one can refer to Theorem 5.1 of .

Proof of Lemma 8.1. We observe, first of all, that

where we defined z=E+iηz=E+i\eta. It follows that

Denoting by a1a_{1} the first column of AA and by BB the M×(N−1)M\times(N-1) matrix consisting of the last N−1N-1 column of AA, we have

Denote μα\mu_{\alpha}’s (α=1,…,N−1)(\alpha=1,\ldots,N-1) the eigenvalues of the (N−1)×(N−1)(N-1)\times(N-1) matrix B∗BB^{*}B. The μα\mu_{\alpha}’s are also the eigenvalues of M×MM\times M matrix BB∗BB^{*} and the other eigenvalues of BB∗BB^{*} are zeros. Then define vαv_{\alpha} (α=1,…,N−1)(\alpha=1,\ldots,N-1) as the normalized eigenvectors of BB∗BB^{*} associated with non-zero eigenvalues μα\mu_{\alpha}, i.e., the matrix elements of BB∗BB^{*} are given by

where we used the assumption that η<E/2\eta<E/2. Because the eigenvalues μα\mu_{\alpha}’s are the eigenvalues of the (N−1)×(N−1)(N-1)\times(N-1) matrix B∗BB^{*}B, they are interlaced with the eigenvalues of HH and ∣{α:∣μα−E∣≤η/2}∣≥NI−1|\{\alpha:|\mu_{\alpha}-E|\leq\eta/2\}|\geq{\mathcal{N}}_{I}-1. It follows from (8.2) that

where in the last step we used Lemma 4.7 from . The claim follows by the assumption that Nη≥NεN\eta\geq N^{\varepsilon} and that NN and KK are large enough.

Consider sample covariance matrices H=A∗AH=A^{*}A with AA an M×NM\times N matrix with independent and identically distributed complex entries. Let 0<d<10<d<1. Recall λ−\lambda_{-} and λ+\lambda_{+} in (3.13) and define κ\kappa as

We will often drop the argument EE from the notation of κ\kappa for brevity. Then for any EE, η\eta satisfying N−1+ε≤η≤12EN^{-1+\varepsilon}\leq\eta\leq\frac{1}{2}E, 12λ−≤E≤10\frac{1}{2}\lambda_{-}\leq E\leq 10, the Stieltjes transform,

of the empirical eigenvalue distribution of H=A∗AH=A^{*}A satisfies

for any δ\delta small enough (independent of EE and η\eta) and N≥2N\geq 2. Here mW(z)m_{W}(z) is the unique solution of

with positive imaginary part for all zz with Im z>0\text{Im }z>0.

Recall λ±=(1±d1/2)2\lambda_{\pm}=(1\pm d^{1/2})^{2} from (3.13). The function mWm_{W} defined in (8.7) depends on dd and can be written as

where   \sqrt{\,\,} denotes the square root on complex plane whose branch cut is the negative real line. Explicit calculation shows mW(z)m_{W}(z) is the Stieltjes transform of the Marchenko-Pastur density given in (3.14).

Using (8.6) and (3.14), we have the local Marchenko-Pastur law for the number of eigenvalues in a small interval:

Consider an interval I=[E−η,E+η]⊂[λ−,λ+]I=[E-\eta,E+\eta]\subset[\lambda_{-},\lambda_{+}] within the bulk spectrum. Let δ\delta be a suffiiciently small parameter. Suppose that EE and η\eta are chosen such that δ−2N−1+ε≤η≤C−1min⁡{κ,δ1/2κ3/4}\delta^{-2}N^{-1+\varepsilon}\leq\eta\leq C^{-1}\min\{\kappa,\delta^{1/2}\kappa^{3/4}\} with a large constant CC and with κ=κ(E)\kappa=\kappa(E) given in (8.5). Then we have the convergence of the counting function, i.e.,

where Nη(E)=∣{λα:∣λα−E∣≤η}∣\mathcal{N}_{\eta}(E)=|\{\lambda_{\alpha}:|\lambda_{\alpha}-E|\leq\eta\}| denotes the number of eigenvalues of H=A∗AH=A^{*}A in the interval I=[E−η,E+η]I=[E-\eta,E+\eta].

Proof of Corollary 8.2. The proof of (8.9) follows from the inequality (8.6) with a similar argument as the proof of Proposition 4.1 of .

We remark that, similarly to Theorem 3.1 in , the assumption on the lower bound η≥N−1+ε\eta\geq N^{-1+\varepsilon} can be relaxed to η≥KN−1\eta\geq KN^{-1} and obtain the local Marchenko-Pastur law on the shortest possible scale, at least away from the spectral edges.

Proof of Proposition 8.1. Let aja_{j} be the jj-th column of AA and let B(j)B^{(j)} be the remaining M×(N−1)M\times(N-1) matrix obtained from AA after removing the jj-th column aja_{j}. Let μα(j)\mu^{(j)}_{\alpha}, vα(j)v^{(j)}_{\alpha} be the non-zero eigenvalues and the eigenvectors of the matrix B(j)[B(j)]∗B^{(j)}[B^{(j)}]^{*} and we define ξα(j)=M∣aj⋅vα(j)∣2\xi_{\alpha}^{(j)}=M|a_{j}\cdot v^{(j)}_{\alpha}|^{2}. Then we have the formula

Note that the vector aja_{j} is independent of μα(j)\mu^{(j)}_{\alpha} and vα(j)v^{(j)}_{\alpha}. Therefore, we have

Define mN−1(j)(z)≡1N−1Tr([B(j)]∗B(j)−z)−1m_{N-1}^{(j)}(z)\equiv\frac{1}{N-1}{\rm Tr}([B^{(j)}]^{*}B^{(j)}-z)^{-1}, then

So with 12λ−≤E=Re(z)≤10\frac{1}{2}\lambda_{-}\leq E=Re(z)\leq 10 and N−1+ε≤η≤E/2N^{-1+\varepsilon}\leq\eta\leq E/2, we have

with some fixed large KK (using dyadic decomposition and (8.1), similarly to the argument in Lemma 4.2 of ) apart from an event of probability e−cNηe^{-c\sqrt{N\eta}}. Then with Proposition 4.5 of , we have

for sufficiently small δ>0\delta>0. Since the eigenvalues of B∗BB^{*}B are interlaced with the eigenvalues of H=A∗AH=A^{*}A, we have

Combining (8.14), (8.15), and (8.16), we find that

for sufficiently small δ>0\delta>0. With the assumption η≤Re(z)/2=E/2\eta\leq Re(z)/2=E/2 and Nη≥NεN\eta\geq N^{\varepsilon}, we obtain

On the other hand, with the definition of Y(j)Y^{(j)}, for any jj, zz, z′z^{\prime} such that, ∣z∣,∣z′∣≤10|z|,|z^{\prime}|\leq 10, \mboxIm(z),\mboxIm(z′)≥η\mbox{Im}(z),\mbox{Im}(z^{\prime})\geq\eta, we have

Together with (8.17), we obtain, for N−1+ε≤η≤12EN^{-1+\varepsilon}\leq\eta\leq\frac{1}{2}E, 12λ−≤E≤10\frac{1}{2}\lambda_{-}\leq E\leq 10 and sufficiently small δ\delta,

where L(z,P10)L(z,P_{10}) is the line segment connecting points zz and P10P_{10}. Then the Proposition 8.1 follows from the next lemma.

Assume HH is a N×NN\times N positive semidefinite matrix with ∥H∥≤5\|H\|\leq 5. For fixed 0<d<10<d<1, we recall the notation λ±=(1±d)2\lambda_{\pm}=(1\pm\sqrt{d})^{2}. Let z0=E+iηz_{0}=E+i\eta and N−1+ε≤η≤12EN^{-1+\varepsilon}\leq\eta\leq\frac{1}{2}E, 12λ−≤E≤10\frac{1}{2}\lambda_{-}\leq E\leq 10. Denote L(z0,P10)L(z_{0},P_{10}) the line segment connecting z0z_{0} and P10=10+5iP_{10}=10+5i. Suppose that for any z∈L(z0,P10)z\in L(z_{0},P_{10}), the Stieltjes transform mN(z)=1NTr(H−z)−1m_{N}(z)=\frac{1}{N}{\rm Tr}(H-z)^{-1} satisfies the following self-consistent relation:

for some Y(j)(z)Y^{(j)}(z)’s. Then there exists δ0>0\delta_{0}>0 depending only on dd, such that, whenever

with κ=κ(E):=∣(λ+−E)(E−λ−)∣\kappa=\kappa(E):=|(\lambda_{+}-E)(E-\lambda_{-})|.

Proof of Lemma 8.3. We begin with a special case: z0=P10z_{0}=P_{10}. In this case if z∈L(z0,P10)z\in L(z_{0},P_{10}) then z=z0=P10z=z_{0}=P_{10}. With the assumptions on HH and 0<d<10<d<1, it is easy to see that:

Insert it into (8.20), we obtain when z=P10z=P_{10},

by S±Δ(z)S^{\Delta}_{\pm}(z). Explicit calculation shows

we note mW(z)=S+(z)m_{W}(z)=S_{+}(z) (see (8.7)). Then the following lemma implies, with Im(mN(z))>0{\rm Im}(m_{N}(z))>0, ImS+(P10)>0{\rm Im}S_{+}(P_{10})>0, ImS−(P10)<0{\rm Im}S_{-}(P_{10})<0, that (8.22) holds for z0=P10z_{0}=P_{10} if δ\delta is small enough.

Let S±Δ(z)S^{\Delta}_{\pm}(z) be the solutions of (8.26). Let z=E+iηz=E+i\eta and 12λ−≤E≤10\frac{1}{2}\lambda_{-}\leq E\leq 10. For sufficiently small Δ\Delta, depending on dd,

Proof of Lemma 8.4. First, when Δ\Delta is small enough, an easy calculation shows

Let a=(z−λ−)(λ+−z)a=(z-\lambda_{-})(\lambda_{+}-z) and b=(z−λ−Δ)(λ+Δ−z)−(z−λ−)(λ+−z)b=(z-\lambda^{\Delta}_{-})(\lambda^{\Delta}_{+}-z)-(z-\lambda_{-})(\lambda_{+}-z). Note that ∣b∣≤CΔ~|b|\leq C\widetilde{\Delta} and therefore, by (8.31), ∣b∣≤CΔ|b|\leq C\Delta. Hence, (8.30) follows from ∣(λ+−z)(z−λ−)∣≥Cκ(E)|(\lambda_{+}-z)(z-\lambda_{-})|\geq C\kappa(E) and from the inequality

which holds for any complex number aa and bb.

Now we prove (8.22) for the case z0≠P10z_{0}\neq P_{10}. We first note that the two solutions of (8.20) are S±(z)S_{\pm}(z) when Y(j)=0Y^{(j)}=0. One can check that for z∈L(z0,P10)z\in L(z_{0},P_{10}), these two solutions are bounded by some constant C1C_{1}:

Furthermore, ∣S−(z0)−S+(z0)∣|S_{-}(z_{0})-S_{+}(z_{0})| can be bounded by ∣S−(z)−S+(z)∣|S_{-}(z)-S_{+}(z)| for any z∈L(z0,P10)z\in L(z_{0},P_{10}) as follows,

On the other hand, for any z∈L(z0,P10)z\in L(z_{0},P_{10}), we claim that if mN(z){m_{N}}(z) is close to S−(z)S_{-}(z) or S+(z)S_{+}(z), then it should be really close to S−(z)S_{-}(z) or S+(z)S_{+}(z), i.e., if

To see this, note that (8.37) together with (8.34) imply

Then with (8.20) and (8.21), we obtain that (8.25) holds for any z∈L(z0,P10)z\in L(z_{0},P_{10}). Using Lemma 8.4 again, we have (8.38).

We have seen that (8.22) and (8.37) (for small δ\delta) hold when z=P10z=P_{10}. Because mN(z){m_{N}}(z), S±(z)S_{\pm}(z) are continuous functions of zz, with (8.38) and (8.37), we can see that when δ\delta is small enough, (8.38) holds for every z∈L(z0,P10)z\in L(z_{0},P_{10}). This result shows that mN(z){m_{N}}(z) must be close to at least one of S+(z)S_{+}(z) and S−(z)S_{-}(z) and it is close to S+(z)S_{+}(z) when z=P10z=P_{10}.

Now we claim that if mN(z0){m_{N}}(z_{0}) were close to S−(z0)S_{-}(z_{0}), i.e.,

then mN(z0){m_{N}}(z_{0}) is also close to S+(z0)S_{+}(z_{0}), which implies that mN(z0){m_{N}}(z_{0}) is always close to S+(z0)S_{+}(z_{0}).

Again, with the continuity of mN(z){m_{N}}(z) and S±(z)S_{\pm}(z) and (8.38), if mN(z0){m_{N}}(z_{0}) is close to S−(z0)S_{-}(z_{0}) in the sense of (8.40), then there exists z∈L(z0,P10)z\in L(z_{0},P_{10}) such that mN(z)m_{N}(z) is close to both of S−(z)S_{-}(z) and S+(z)S_{+}(z), i.e.,

Together with (8.40), we obtain mN(z0)m_{N}(z_{0}) is still close to the S+(z0)S_{+}(z_{0}). It means that (8.22) holds for all z0z_{0}’s in our assumption, using the fact S+(z0)=mW(z0)S_{+}(z_{0})={m_{W}}(z_{0}). This completes the proof of Lemma 8.3.

The following lemma shows that the expectation value of mN(z){m_{N}}(z) is close to mW(z)m_{W}(z).

Let z=E+iηz=E+i\eta, such that N−1+ε≤η≤12EN^{-1+\varepsilon}\leq\eta\leq\frac{1}{2}E, 12λ−≤E≤10\frac{1}{2}\lambda_{-}\leq E\leq 10, for some ε>0\varepsilon>0. Then we have

for large enough NN depending on ε\varepsilon.

Proof of Lemma 8.5. Using (8.34) and the estimate (8.6) from Proposition 8.1, we have

uniformly in z=E+iηz=E+i\eta within the range N−1+ε≤η≤12EN^{-1+\varepsilon}\leq\eta\leq\frac{1}{2}E, 12λ−≤E≤10\frac{1}{2}\lambda_{-}\leq E\leq 10.

From (8.10), (8.11), (8.13), (8.15) and (8.17), we obtain

Using (8.6), (8.34) and Nηκ≥Nε/3N\eta\kappa\geq N^{\varepsilon/3}, we obtain that ∣B∣|B| is bounded from below by a constant C0C_{0}. Furthermore, for some δ>0\delta>0,

Combining this with (8.47), (8.48) and (8.49), we obtain:

Let z=E+ηiz=E+\eta i, N−1+ε≤η≤E/2N^{-1+\varepsilon}\leq\eta\leq E/2, λ−/2≤E≤10\lambda_{-}/2\leq E\leq 10 and ε>0\varepsilon>0. Suppose Nκη≥Nε′N\kappa\eta\geq N^{\varepsilon^{\prime}} for some ε′>0\varepsilon^{\prime}>0, we have

when NN is sufficiently large (depending on ε′\varepsilon^{\prime}).

Proof of Lemma 8.6. We only prove the case of the real sample covariance matrix. The case of the complex sample covariance matrix can be treated similarly.

Let λα\lambda_{\alpha} and uαu_{\alpha} be the eigenvalues and eigenvectors of H=A∗AH=A^{*}A. The derivative of λα\lambda_{\alpha} with respect to the (i,j)(i,j)-th matrix element AijA_{ij} is given by

Using ∑juα(j)uβ(j)=δα,β\sum_{j}u_{\alpha}(j)u_{\beta}(j)=\delta_{\alpha,\beta} and ∑i(Auα)(i)(Auβ)(i)=λαδα,β\sum_{i}(Au_{\alpha})(i)(Au_{\beta})(i)=\lambda_{\alpha}\delta_{\alpha,\beta}, one can obtain the following result, as in (3.3) of ,

Then with Lemma 8.1, as in (3.6) of , we obtain (8.57).

As in the proof of Lemma 8.5, with a continuity argument and Lemma 8.4, we can obtain (8.56) from (8.62) and complete the proof.

Verifying Assumption III

The following theorem gives the estimate (2.13) for the singular values of the sample covariance matrices.

Assume that the single site distribution dν{\rm d}\nu of the entries of MA\sqrt{M}A satisfies the logarithmic Sobolev inequality (3.10). Recall that ρW\rho_{W} in (3.14) denotes the density in Marchenko-Pastur law. Define γj∈[λ−,λ+]\gamma_{j}\in[\lambda_{-},\lambda_{+}], (j=1,2,…,N)(j=1,2,\ldots,N) with the relation

Denote xj=λj1/2x_{j}=\lambda_{j}^{1/2} the singular values of AA. Then there exists δ>0\delta>0, such that,

Remark. An analogous result holds for the eigenvalues of the Wigner matrices; the proof is similar and we will not give the details here. We only point out that the key ingredients of the argument below are: (i) apriori bound on the extreme eigenvalues (see Lemma 9.2 and the remark afterwards); (ii) concentration of the local density of states (used in Lemma 9.3). We also critically use the fact that the density of states (semicircle law or Marchenko-Pastur law) has a square root singularity at the edges.

The real numbers λ1≤λ2≤…≤λN−1\lambda_{1}\leq\lambda_{2}\leq\ldots\leq\lambda_{N-1} denote the eigenvalues of BB and the nonnegative real numbers ξα\xi_{\alpha} are given by

The proof of (9.3) is a straightforward computation. This identity is the key to extend our results on the local semicircle law for quaternion self-dual matrices without any further modifications.

Now we return to the the proof of Theorem 9.1 and we start with some preparatory lemmas. First, we recall the following result from .

then nλ(λ−−N−1/5)≤Ce−Nεn^{\lambda}(\lambda_{-}-N^{-1/5})\leq Ce^{-N^{\varepsilon}} and 1−nλ(λ++N−1/5)≤Ce−Nε1-n^{\lambda}(\lambda_{+}+N^{-1/5})\leq Ce^{-N^{\varepsilon}} for some ε>0\varepsilon>0. Therefore for any 1≤j≤N1\leq j\leq N,

Remark. In fact, the error term in is N−2/3+εN^{-2/3+\varepsilon} instead of N−1/5N^{-1/5} but we will use only the weaker bound (9.6) in order to indicate that our proof goes through for Wigner matrices with not necessarily symmetric distributions as well. In the latter case only N−1/4+εN^{-1/4+\varepsilon} has been proven by Vu in for compactly supported distribution ν\nu with an effective dependence of the constant CC on the support. This effective dependence is necessary to remove the compact support condition as in Lemma 4.1 of . Strictly speaking, the result in was stated only for symmetric Wigner matrices but it holds for hermitian and quaternion self-dual Wigner matrices as well, since the key estimate (equation (5) in ) is independent of the matrix ensemble.

Recall that ρW\rho_{W} in (3.14) denotes the density in Marchenko-Pastur law. Let

Proof of Lemma 9.3. To prove (9.8), with Lemma 9.2, one only needs to prove

To prove (9.9), for fixed EE, we can assume nλ(E)>nWλ(E)n^{\lambda}(E)>n_{W}^{\lambda}(E) and denote Δ=nλ(E)−nWλ(E)\Delta=n^{\lambda}(E)-n_{W}^{\lambda}(E). Because nλ(E)n^{\lambda}(E) is an increasing function and the derivative of nWλ(E)n_{W}^{\lambda}(E) is bounded by ∥ρW∥∞=1π(d−d2)−1/2\|\rho_{W}\|_{\infty}=\frac{1}{\pi}(d-d^{2})^{-1/2}, we have:

Using (9.8), it follows that Δ≤O(N−3/7)\Delta\leq O(N^{-3/7}).

Similarly to the calculation in Theorem 3.1 of , by using the logarithmic Sobolev inequality, we have

hold with an extremely high probability, where we introduced the notations j−≡N1−ε1j_{-}\equiv N^{1-\varepsilon_{1}} and j+≡N−N1−ε1j_{+}\equiv N-N^{1-\varepsilon_{1}}. Then for some ε>0\varepsilon>0, we have

Proof. By symmetry, we only need to prove that (9.16) holds for the sum on the indices jj with γj≤αj\gamma_{j}\leq\alpha_{j}. Introduce the notation

The estimate (9.13), with K=1K=1, implies that max⁡j∣λj−αj∣≤N−1/2+δ\max_{j}|\lambda_{j}-\alpha_{j}|\leq N^{-1/2+\delta} holds with an extremely high probability, for any positive δ\delta. Therefore, we can bound nα(E)n^{\alpha}(E) from above by (for any EE)

Similarly, we can obtain the lower bound. Putting them together, we have that:

holds for any EE. The assumption (9.14) implies that

holds with an extremely high probability, for any j≤j−j\leq j_{-} or j≥j+j\geq j_{+}. For the other jj’s, for which λj\lambda_{j} may appear in [λ−+N−ε2,λ+−N−ε2][\lambda_{-}+N^{-\varepsilon_{2}},\lambda_{+}-N^{-\varepsilon_{2}}], we use (9.15) and obtain the following improved bound on nα(E)n^{\alpha}(E): when λ−+N−ε2≤E≤λ+−N−ε2\lambda_{-}+N^{-\varepsilon_{2}}\leq E\leq\lambda_{+}-N^{-\varepsilon_{2}},

Let F(E)F(E) be a continuous and differentiable function, such that N−1/2−ε3≤F(E)≤N−1/2+δN^{-1/2-\varepsilon_{3}}\leq F(E)\leq N^{-1/2+\delta}, for 0≤δ≤110min⁡{ε1,ε2,ε3}0\leq\delta\leq\frac{1}{10}\min\{\varepsilon_{1},\varepsilon_{2},\varepsilon_{3}\}, F(E)=N−1/2−ε3F(E)=N^{-1/2-\varepsilon_{3}} for λ−+2N−ε2≤E≤λ+−2N−ε2\lambda_{-}+2N^{-\varepsilon_{2}}\leq E\leq\lambda_{+}-2N^{-\varepsilon_{2}}, F(E)=N−1/2+δF(E)=N^{-1/2+\delta} for E≤λ−+N−ε2E\leq\lambda_{-}+N^{-\varepsilon_{2}} or E≥λ+−N−ε2E\geq\lambda_{+}-N^{-\varepsilon_{2}} and ∣F′(E)∣≤N−δ|F^{\prime}(E)|\leq N^{-\delta}. Combining (9.21) and (9.19), we obtain

for any jj, such that, αj>γj\alpha_{j}>\gamma_{j}. Therefore we can write

where in the second line we used the fact that the difference between j/Nj/N and nα(E)n^{\alpha}(E) must be a multiple of N−1N^{-1}. Since max⁡jλj≤10\max_{j}\lambda_{j}\leq 10 holds with an extremely high probability, using (9.22), we can replace nα(E)n^{\alpha}(E) with nλ(E−F(E))n^{\lambda}(E-F(E)) in (9.24), i.e.,

Change the variable from EE to t(E)=E−F(E)t(E)=E-F(E). With ∣F′∣≤N−δ|F^{\prime}|\leq N^{-\delta}, we obtain F(t)=(1+O(N−δ))F(E)F(t)=(1+O(N^{-\delta}))F(E) and dt/dE=(1+O(N−δ)){\rm d}t/{\rm d}E=(1+O(N^{-\delta})). Thus

where E(t)E(t) is the inverse function of t(E)t(E). Note, when 1(⋯ )1(⋯ )=1{\bf 1}\left(\cdots\right){\bf 1}\left(\cdots\right)=1 in (9.26), we have nW(E′)≥j/N>nλ(t)n_{W}(E^{\prime})\geq j/N>n_{\lambda}(t). Define the inverse function of nWn_{W} as nW−1(1)=λ+n_{W}^{-1}(1)=\lambda_{+}, nW−1(0)=λ−n_{W}^{-1}(0)=\lambda_{-} and nW−1(nW(x))=xn_{W}^{-1}(n_{W}(x))=x for 0<nW(x)<10<n_{W}(x)<1. Then

Inserting this inequality into (9.26) and performing the dE′dE^{\prime} integration, we can see that

where we expanded (9.28) into four terms:

Since nW′(t)=ρW(t)≤Cn_{W}^{\prime}(t)=\rho_{W}(t)\leq C, F(t)=N−1/2−ε3F(t)=N^{-1/2-\varepsilon_{3}} when λ−+2N−ε2≤t≤λ+−2N−ε2\lambda_{-}+2N^{-\varepsilon_{2}}\leq t\leq\lambda_{+}-2N^{-\varepsilon_{2}} and F(t)≤N−1/2+δF(t)\leq N^{-1/2+\delta} for any EE, we obtain A1≤N−1−εA_{1}\leq N^{-1-\varepsilon}, for some ε>0\varepsilon>0. Next, from (9.8) and F(t)≤N−1/2+δF(t)\leq N^{-1/2+\delta} for any tt, we can see A2≤(N−6/7+N−1)N−1/2+δ≤N−1−εA_{2}\leq(N^{-6/7}+N^{-1})N^{-1/2+\delta}\leq N^{-1-\varepsilon}.

To prove A3≤N−1−εA_{3}\leq N^{-1-\varepsilon}, we start with writing A3A_{3} as

The first term on the r.h.s. of (9.29) is equal to zero, since nWn_{W} is constant outside [λ−,λ+][\lambda_{-},\lambda_{+}]. The second term can be bounded by N−1−εN^{-1-\varepsilon}, for some ε>0\varepsilon>0, using the facts F(t)≤N−1/2+δF(t)\leq N^{-1/2+\delta} and nW(λ−+E)≤CE3/2n_{W}(\lambda_{-}+E)\leq CE^{3/2}, i.e.,

Now we prove that the third and fourth term of (9.29) are less N−1−εN^{-1-\varepsilon}, for some ε>0\varepsilon>0.

From the explicit definition of nWn_{W}, an easy calculation shows that, for all t∈(λ−,λ+)t\in(\lambda_{-},\lambda_{+}),

Combining this with the fact ∣nW(t+2F(t))−nW(t)∣≤C∥nW′∥∞F(t)≤CN−1/2+δ|n_{W}(t+2F(t))-n_{W}(t)|\leq C\|n^{\prime}_{W}\|_{\infty}F(t)\leq CN^{-1/2+\delta}, we obtain that the third and fourth terms of (9.29) are less than CN−2/7N−1/2+δN−1/4≤N−1−εCN^{-2/7}N^{-1/2+\delta}N^{-1/4}\leq N^{-1-\varepsilon}, for some ε>0\varepsilon>0.

To bound the last term of (9.29), we use, once again the bound ∣nW(t+2F(t))−nW(t)∣≤CN−1/2+δ|n_{W}(t+2F(t))-n_{W}(t)|\leq CN^{-1/2+\delta}. From (9.31), we find therefore that

At last, we prove A4≤N−1−εA_{4}\leq N^{-1-\varepsilon}. We rewrite A4A_{4} as

where Σ(t)≡(∣nW(t)−nλ(t)∣+N−1)⋅∣t−nW−1(nλ(t))∣\Sigma(t)\equiv(|n_{W}(t)-n^{\lambda}(t)|+N^{-1})\cdot|t-n_{W}^{-1}(n^{\lambda}(t))|. When t∉(λ−,λ+)t\notin(\lambda_{-},\lambda_{+}), from (9.9) and (9.32), one can see that

Here we also used the fact that for large tt, ∣nλ(t)−1∣|n^{\lambda}(t)-1| decays exponentially fast to zero (see, for example, Lemma 7.3 of , which is stated for matrices with complex entries, but can be trivially extended to the case of real entries). Together with (9.8), we obtain that

When t∈(λ−,λ+)t\in(\lambda_{-},\lambda_{+}), with (9.32) and (9.8), we can see

Combining (9.35) and (9.38), we obtain A4≤N−1−εA_{4}\leq N^{-1-\varepsilon} for some ε>0\varepsilon>0. Together with (9.28), this compeletes the proof of Lemma 9.5.

Next, we show that the assumptions (9.14) and (9.15) in Lemma 9.5 always hold. First we prove (9.14) in the next Lemma 9.6 with an analogous proof as Lemma 9.5. Then in Lemma 9.7 we show that (9.15) holds when (9.14) holds.

There exist small positive numbers ε1\varepsilon_{1} and ε2\varepsilon_{2}, such that,

hold with an extremely high probability, where we recall the notations j−≡N1−ε1j_{-}\equiv N^{1-\varepsilon_{1}} and j+≡N−N1−ε1j_{+}\equiv N-N^{1-\varepsilon_{1}}.

Proof. As in (9.19), for any EE, δ>0\delta>0 and sufficiently large NN, we have

So without any other assumptions, one can obtain (9.28), if we set F(E)≡N−1/2+δF(E)\equiv N^{-1/2+\delta} instead of F(E)F(E) defined in the proof of Lemma 9.5. With a similar argument as in the proof of Lemma 9.5 but with this redefined F(E)F(E), we have

We prove this claim by contradiction; assume that for some j0j_{0} we have ∣αj0−γj0∣≥N−110\left|\alpha_{j_{0}}-\gamma_{j_{0}}\right|\geq N^{-\frac{1}{10}}. By symmetry we can assume that j0≤N/2j_{0}\leq N/2, the case j0≥N/2j_{0}\geq N/2 is analogous. We start with the case j0≤N1/2j_{0}\leq N^{1/2}. Then γj0≤λ−+CN−1/4\gamma_{j_{0}}\leq\lambda_{-}+CN^{-1/4} and in this case αj0\alpha_{j_{0}} must be larger than γj0\gamma_{j_{0}}, otherwise αj0≤γj0−N−110≤λ−−12N−110\alpha_{j_{0}}\leq\gamma_{j_{0}}-N^{-\frac{1}{10}}\leq\lambda_{-}-\frac{1}{2}N^{-\frac{1}{10}} would contradict to αj0∈[λ−−CN−1/5,λ++CN−1/5]\alpha_{j_{0}}\in[\lambda_{-}-CN^{-1/5},\lambda_{+}+CN^{-1/5}], see (9.6). Using

for any i,ji,j and that αj\alpha_{j} is monotone, we obtain that

for any jj such that j0≤j≤j0+N1/2j_{0}\leq j\leq j_{0}+N^{1/2}. Then

with some positive c>0c>0 which would contradict to (9.41). Now we consider the case j0≥N1/2j_{0}\geq N^{1/2}. The previous argument remains unchanged if αj0>γj0\alpha_{j_{0}}>\gamma_{j_{0}}. If αj0<γj0\alpha_{j_{0}}<\gamma_{j_{0}}, then we use

for any jj such that j0−N1/2≤j≤j0j_{0}-N^{1/2}\leq j\leq j_{0} and we obtain

which again contradicts to (9.41). This completes the proof of (9.42).

On the other hand, the estimate (9.13), with K=1K=1, implies max⁡j∣λj−αj∣≤N−1/2+δ\max_{j}|\lambda_{j}-\alpha_{j}|\leq N^{-1/2+\delta} holds with an extremely high probability. Combining (9.42) with this fact, we can see that for any small enough ε1\varepsilon_{1}, there exists ε2\varepsilon_{2} such that (9.39) holds, which completes the proof of Lemma 9.6.

The next Lemma guarantees the assumption (9.15) in Lemma 9.5, given (9.14).

If there exist sufficiently small positive numbers ε1\varepsilon_{1} and ε2\varepsilon_{2}, such that

holds with an extremely high probability, then there exists ε3>0\varepsilon_{3}>0 such that,

holds with an extremely high probability, where we recall the notations j−≡N1−ε1j_{-}\equiv N^{1-\varepsilon_{1}} and j+≡N−N1−ε1j_{+}\equiv N-N^{1-\varepsilon_{1}}.

Proof. For simplicity, we only prove the case of j≤N/2j\leq N/2, the case j>N/2j>N/2 is analogous. Using (9.13), for any N/2≥j>j−N/2\geq j>j_{-}, δ>0\delta>0, with K=N1/4K=N^{1/4}, we have

Now we claim that, for K=N1/4K=N^{1/4}, j−<j<j+j_{-}<j<j_{+},

holds with an extremely high probability, which implies

Suppose now that λj+K−λj≥N−5/8\lambda_{j+K}-\lambda_{j}\geq N^{-5/8}. With the assumption (9.45), we have that, for j−<j<N/2j_{-}<j<N/2,

with an extremely high probability. Divide this interval into small intervals with the length 12N−5/8\frac{1}{2}N^{-5/8}. By the local Marchenko-Pastur law, i.e., Corollary 8.2, the event that the number of the eigenvalues in each piece is larger than CN1−3ε2−5/8CN^{1-3\varepsilon_{2}-5/8} holds with an extremely high probability. On the other hand, if λj+K−λj≥N−5/8\lambda_{j+K}-\lambda_{j}\geq N^{-5/8}, then the total number of eigenvalues in at least one of these intervals is less than K=N1/4K=N^{1/4}, which implies that λj+K−λj≤N−5/8\lambda_{j+K}-\lambda_{j}\leq N^{-5/8} holds with an extremely high probability. Together with (9.50), we have (9.48). Then combining (9.48), (9.49) and (9.47), we obtain (9.46) and complete the proof.

Proof of Theorem 9.1. Note that the assumptions in Lemma 9.5 are proved in Lemma 9.6 and 9.7. Combining Lemma 9.5, 9.6 and 9.7, we obtain (9.16), i.e.,

holds with an extremely high probability. To see (9.53), first notice that (9.13), with K=1K=1, implies that, for any δ\delta and jj

holds with an extremely high probability. The estimate (9.46) shows that there exist ε1>0\varepsilon_{1}>0 and ε3>0\varepsilon_{3}>0 such that

holds with an extremely high probability. Combining this with (9.54) for the remaining indices j≤Nε1j\leq N^{\varepsilon_{1}} or j≥N−Nε1j\geq N-N^{\varepsilon_{1}}, we obtain (9.53). Together with (9.52), we have:

for some ε>0\varepsilon>0. Using the definition xj=λj1/2x_{j}=\lambda_{j}^{1/2}, one has

Inserting (9.57) into (9.56), we obtain (9.2) and complete the proof of Theorem 9.1.

Appendix A Existence and restriction of the dynamics

Moreover, by approximating ff by L2L^{2} functions and using that TtT_{t} is contraction in L1L^{1} (Section 1.5 in ), the differential equation holds even if the initial condition ff is only in L1L^{1}. In this case the convergence ft→ff_{t}\to f, as t→0+0t\to 0+0, holds only in L1L^{1}. We remark that TtT_{t} is also a contraction on L∞L^{\infty}, by duality.

To establish the relation between LL and L(Σ)L^{(\Sigma)}, we first define the symmetrized version of Σ\Sigma

Note that both LL with L(Σ)L^{(\Sigma)} are local operators and LL is symmetric with respect to the permutation of the variables. For any function ff defined on Σ\Sigma, we define its symmetric extension onto Σ~\widetilde{\Sigma} by f~\widetilde{f}. Clearly Lf~=L(Σ)f~L\widetilde{f}=\widetilde{L^{(\Sigma)}f} for any f∈C0∞(Σ)f\in C_{0}^{\infty}(\Sigma). Since the generator is uniquely determined by its action on its core, and the generator uniquely determines the dynamics, we see that for any f∈L1(Σ,dμ)f\in L^{1}(\Sigma,{\rm d}\mu), one can determine Tt(Σ)fT^{(\Sigma)}_{t}f by computing Ttf~T_{t}\widetilde{f} and restricting it to Σ\Sigma. In other words, the dynamics (2.1) is well defined when restricted to Σ=ΣN\Sigma=\Sigma_{N}.

as ε→0\varepsilon\to 0. It is sufficient to consider only the positive semi-axis, i.e., r>0r>0. Extending the functions to two dimensional radial functions, G(x):=g(∣x∣)G(x):=g(|x|), Gε(x)=gε(∣x∣)G_{\varepsilon}(x)=g_{\varepsilon}(|x|), this statement is equivalent to the fact that a point in two dimensions has zero capacity.

Appendix B Bakry-Emery argument on a subdomain

The estimate (4.14) in Theorem 4.2 is based on the Bakry-Emery argument for the dissipation of the Dirichlet form. This method uses a lower bound on the Hessian of H~\widetilde{\mathcal{H}} and an integration by parts. Since the dynamics is restricted to Σ=ΣN\Sigma=\Sigma_{N}, we need to check that the boundary term in the integration by parts vanishes.

In our application, this argument will be used for the Hamiltonian H~\widetilde{\mathcal{H}} (see (4.5)) and its generator L~=12NΔ−12(∇H~)∇\widetilde{L}=\frac{1}{2N}\Delta-\frac{1}{2}(\nabla\widetilde{\mathcal{H}})\nabla, but for simplicity, we omit the tilde from the notation below. With h=ht=qth=h_{t}=\sqrt{q_{t}} a standard calculation (see (5.8) of with somewhat different notations) shows that

assuming that the quantities in each step are well defined and that the boundary term

in the integration by parts in the third line vanishes. In we argued with a somewhat specific form of qq, an information not directly available here.

The rigorous proof in the general case uses a regularization and a cutoff argument. First we regularize the function q=qt∈D(L)q=q_{t}\in D(L), t>0t>0, by defining

for some ε>0\varepsilon>0. This has the advantage that the derivatives of hεh^{\varepsilon} can be bounded by those of qεq^{\varepsilon}. We consider a cutoff function θ∈C0∞(Σ)\theta\in C_{0}^{\infty}(\Sigma) to be specified later and we insert θ\theta in the calculation (B.1). Since LL is an elliptic operator with smooth coefficients away from the boundary ∂Σ\partial\Sigma, by standard parabolic regularity we know that qq and thus hh are smooth functions inside Σ\Sigma. Thus each step in the cutoff version of (B.1) is justified with an additional term coming from the derivative hitting θ\theta in the integration by parts. After repeating the steps in (B.1), we obtain

We now show that, by an appropriate choice of a sequence of cutoff functions, the second term in (B.3) vanishes. We first define the set of higher order coalescences where at least three point coincide as

We remark that in case of Assumption I’ we formally introduce x0=0x_{0}=0 to this definition, so that QQ will include also three point singularities of the type x1=x2=0x_{1}=x_{2}=0. For any δ>0\delta>0 we define the set

is the δ\delta-neighborhood of the three-point singularity set within Σ\Sigma. Introduce an additional small positive parameter η≪δ\eta\ll\delta. We now choose the cutoff function θ\theta of the form θ=θ1θ2\theta=\theta_{1}\theta_{2}, depending on the parameters δ\delta and η\eta, such that

(i) θ1(x)≡1\theta_{1}({\bf{x}})\equiv 1 if \mboxdist(x,∂Σ)≥2η\mbox{dist}({\bf{x}},\partial\Sigma)\geq 2\eta, θ1(x)≡0\theta_{1}({\bf{x}})\equiv 0 if \mboxdist(x,∂Σ)≤η\mbox{dist}({\bf{x}},\partial\Sigma)\leq\eta and ∣∇θ1∣≤O(η−1)|\nabla\theta_{1}|\leq O(\eta^{-1});

(ii) θ2(x)≡1\theta_{2}({\bf{x}})\equiv 1 if \mboxdist(x,Q)≥2δ\mbox{dist}({\bf{x}},Q)\geq 2\delta, θ2(x)≡0\theta_{2}({\bf{x}})\equiv 0 if \mboxdist(x,Q)≤δ\mbox{dist}({\bf{x}},Q)\leq\delta and ∣∇θ2∣≤O(δ−1)|\nabla\theta_{2}|\leq O(\delta^{-1}).

We state two estimates on the solution qtq_{t} of (4.13) that will be proven at the end of the section.

Assume that q0∈L∞q_{0}\in L^{\infty}. Then the solution qtq_{t} of (4.13) satisfies a uniform supremum bound on the closure of Σ\Sigma,

Furthermore, qtq_{t} is regular away from the higher order coalescence singularities with the estimate

where KK is a compact set and the constant depends only on the indicated parameters. In particular, qtq_{t} is regular up to the boundary ∂Σ∖Qδ\partial\Sigma\setminus Q_{\delta}, i.e., at the two-point coalescence points away from higher order coalescences.

Using this lemma, we can treat the second term on the r.h.s. of (B.3). We split the integration into two regimes. First we consider the regime where θ2∇θ1≠0\theta_{2}\nabla\theta_{1}\neq 0, i.e., an (2η)(2\eta)-neighborhood of ∂Σ∖Qδ\partial\Sigma\setminus Q_{\delta}. On this set we note the local density scales at as ηβ\eta^{\beta}, thanks to the term ∣xi−xj∣β|x_{i}-x_{j}|^{\beta} in e−NHe^{-N{\mathcal{H}}}. Thus the measure of the support of ∇θ\nabla\theta near ∂Σ∖Qδ\partial\Sigma\setminus Q_{\delta} scales as η1+β\eta^{1+\beta}, while ∣∇θ∣≤Cη−1|\nabla\theta|\leq C\eta^{-1} (assuming η≤δ\eta\leq\delta). Since (B.5) guarantees that the derivatives of hεh^{\varepsilon} remain locally bounded (with a bound depending on ε,δ,t\varepsilon,\delta,t and NN), the boundary term near ∂Σ∖Qδ\partial\Sigma\setminus Q_{\delta} vanishes as η→0\eta\to 0.

To estimate the integral on the support of ∇θ2\nabla\theta_{2}, i.e., on a subset of Q2δQ_{2\delta}, we use that θ\theta can be replaced with θ2\theta_{2} after taking the η→0\eta\to 0 limit. Since we have ∣∇θ2∣=O(δ−1)|\nabla\theta_{2}|=O(\delta^{-1}) and ∣∇khtε∣≤Cε∣∇kqtε∣≤Cε,t,Nδ−k|\nabla^{k}h_{t}^{\varepsilon}|\leq C_{\varepsilon}|\nabla^{k}q_{t}^{\varepsilon}|\leq C_{\varepsilon,t,N}\delta^{-k} with k=1,2k=1,2, the integrand scales at most δ−4\delta^{-4}. Since the local density scales at least as δ3β\delta^{3\beta} due to a factor of the type ∣xi−xi+1∣β∣xi+1−xi+2∣β∣xi−xi+2∣β|x_{i}-x_{i+1}|^{\beta}|x_{i+1}-x_{i+2}|^{\beta}|x_{i}-x_{i+2}|^{\beta}, the total measure of Q2δQ_{2\delta} is of order δ2+3β\delta^{2+3\beta}. Hence the integral on Q2δQ_{2\delta} scales at most as δ2+3β−4≤δ\delta^{2+3\beta-4}\leq\delta in the δ\delta parameter and therefore the contribution of the neighborhood of higher order singularities to the second term in (B.3) vanishes as δ→0\delta\to 0.

After having removed θ\theta and the second term from (B.3), we let ε→0\varepsilon\to 0 and this gives the desired result (4.14).

To complete the argument, finally, we need to prove Lemma B.1.

Proof of Lemma B.1. The bound (B.4) follows immediately, since q0∈L∞q_{0}\in L^{\infty} and the semigroup TtT_{t} a contraction in L∞L^{\infty} (see Appendix A).

The second statement of Lemma B.1 follows from a standard regularization argument for a typical two-point singularity at xi=xi+1x_{i}=x_{i+1} that was already outlined in . Fix a point x∗∈∂Σ∖Qδ{\bf{x}}^{*}\in\partial\Sigma\setminus Q_{\delta} and assume that xi∗=xi+1∗x_{i}^{*}=x_{i+1}^{*}, but for all other pairs ∣xj∗−xj+1∗∣≥δ|x_{j}^{*}-x_{j+1}^{*}|\geq\delta. We remark that the neighborhood of two (or more) independent singularities, e.g., xi=xi+1x_{i}=x_{i+1} and xj=xj+1x_{j}=x_{j+1}, ∣i−j∣≥2|i-j|\geq 2, can be treated similarly by applying the same regularization argument separately. We omit these details here.

where LregL_{reg} is an elliptic operator with second derivatives in the y{\bf{y}} variables and with coefficients regular on the scale δ\delta (since all other singularities are at least at a distance O(δ)O(\delta) away from Φ(B)\Phi(B)).

For the β=1\beta=1 case, by introducing a function q^t(a,b,y):=qt(a2+b2,y)\widehat{q}_{t}(a,b,{\bf{y}}):=q_{t}(\sqrt{a^{2}+b^{2}},{\bf{y}}) of N+1N+1 variables, we see that q^t\widehat{q}_{t} satisfies ∂tq^t=L^q^t\partial_{t}\widehat{q}_{t}=\widehat{L}\widehat{q}_{t}, where

i.e., LL becomes an elliptic operator L^\widehat{L} with bounded and regular coefficients in the new variables. A similar transformation is possible for any integer β≥1\beta\geq 1, where uu is considered as the radial part of a (β+1)(\beta+1)-dimensional variable.

We claim that the singular point u=0u=0 becomes a removable singularity in the variables (a,b)(a,b) around (0,0)(0,0). Note that the singular set is a codimension two subspace in the (a,b,y,t)(a,b,{\bf{y}},t) space-time coordinate system which becomes a line segment in the (a,b,t)(a,b,t) space-time system if we disregard the variable y{\bf{y}}. Note that y{\bf{y}} plays no role in this argument since every coefficient is regular in y{\bf{y}}. The parabolic equation ∂tq^t=L^q^t\partial_{t}\widehat{q}_{t}=\widehat{L}\widehat{q}_{t} holds in a strong sense away from the origin (a,b)=(0,0)(a,b)=(0,0) in these two variables, and moreover q^t\widehat{q}_{t} is bounded by (B.4). We can thus apply Theorem II of with p=2p=2, r=∞r=\infty to see that q^t\widehat{q}_{t} must coincide with the regular solution obtained by using the fundamental solution to the equation in a small space-time neighborhood of the singular set. This proves that q^t\widehat{q}_{t}, and hence qtq_{t}, is a smooth function up to the boundary ∂Σ∖Qδ\partial\Sigma\setminus Q_{\delta}.

To obtain the quantitative estimate (B.5), we consider the regularity of the coefficients of LregL_{reg}. Due to the special structure of H{\mathcal{H}}, every term in L=12NΔ−12(∇H)∇L=\frac{1}{2N}\Delta-\frac{1}{2}(\nabla{\mathcal{H}})\nabla is either regular on any small scales, or it scales as \mbox(length)−2\mbox{(length)}^{-2}. Since the neighborhood BB is at least at distance O(δ)O(\delta) away from the other singularities, the coefficients of LregL_{reg} are regular on scale δ\delta. Therefore the solution qtq_{t} is regular on scale δ\delta on BB and this gives the δ\delta-scaling of the estimate (B.5). This completes the proof of Lemma B.1.

Acknowledgement. The authors thank Alice Guionnet for pointing out some errors in the preliminary version of the manuscript.

References