Universality of local spectral statistics of random matrices

Laszlo Erdos, Horng-Tzer Yau

Introduction

What do the eigenvalues of a typical large matrix look like? Do we expect certain universal patterns of eigenvalue statistics to emerge? Although random matrices appeared already in a concrete statistical application by Wishart in 1928 , these natural questions were not raised until the pioneering work of E. Wigner in the 1950’s. To make the problem simpler, we restrict ourselves to either real symmetric or complex Hermitian matrices so that the eigenvalues are real. For definiteness, we consider N×NN\times N square matrices H=H(N)=(hij)H=H^{(N)}=(h_{ij}) with matrix elements having mean zero and variance 1/N1/N, i.e.,

The random variables hijh_{ij}, i,j=1,…,Ni,j=1,\ldots,N are real or complex independent random variables subject to the symmetry constraint hij=h‾jih_{ij}=\overline{h}_{ji}. These ensembles of random matrices are called Wigner matrices. We will always consider the limit as the matrix size goes to infinity, i.e., N→∞N\to\infty.

The first rigorous result about the spectrum of a random matrix of this type is the famous Wigner semicircle law which states that the empirical densities of the eigenvalues, λ1,λ2,…,λN\lambda_{1},\lambda_{2},\ldots,\lambda_{N}, of large symmetric or Hermitian matrices, after proper normalization such as (1.1), are given by

in the weak limit as N→∞N\to\infty. The limit density is independent of the details of the distribution of hijh_{ij}. The motivation for Wigner was to find a phenomenological model for the energy gap statistics of large atomic nuclei since the energy levels of large quantum systems are impossible to compute from first principles. After several attempts, Wigner was convinced that random matrices were the right models. Besides the semicircle law, he also predicted that the eigenvalue gap distribution in the bulk of the spectrum is given by the Wigner surmise, e.g. in the case of symmetric matrices,

where ϱ\varrho is the local density of eigenvalues (see for an overview).

In the Gaussian case, the joint probability density of the eigenvalues can be expressed explicitly as

The Vandermonde determinant structure allows one to compute the kk-point correlation functions in the large NN limit via Hermite polynomials that are the orthogonal polynomials with respect to the Gaussian weight function.

The result of Dyson, Gaudin and Mehta asserts that for any fixed energy EE in the bulk of the spectrum, i.e., ∣E∣<2|E|<2, the small scale behavior of pN(n)p^{(n)}_{N} is given explicitly by

Note that the limit in (1.5) is independent of the energy EE as long as it is in the bulk of the spectrum. The rescaling by a factor N−1N^{-1} of the correlation functions in (1.5) corresponds to the typical distance between consecutive eigenvalues and we will refer to the law under such scaling as local statistics. Similar but much more complicated formulas for symmetric matrices were also obtained. It is well-known that the eigenvalue gap distribution can be computed from the correlation functions via the inclusion-exclusion principle and thus (1.5) also yields a precise asymptotics for eigenvalue gap distributions. In a striking coincidence, the Wigner surmise, which was based on a 2×22\times 2 matrix computation, agrees with this sophisticated formula with a typical error of only a few percentage points. Note that the correlation functions do not factorize, i.e. the eigenvalues are strongly correlated despite that the matrix elements are independent. Eigenvalues of random matrices thus represent a strongly correlated point process obtained from independent random variables in a natural way.

The central thesis of Wigner is the belief that the eigenvalue gap distributions for large complicated quantum systems are universal in the sense that they depend only on the symmetry class of the physical system but not on other detailed structures. This thesis has never been proved for any truly interacting system and there is even no heuristically convincing argument for its correctness. Despite this, there is a general belief that the random matrix statistics and Poisson statistics represent two paradigms of energy level statistics for many-body quantum systems: Poisson for independent systems and random matrix for highly correlated systems. In fact, these paradigms extend even to certain one-body systems such as the quantization of the geodesic flow in a domain or on a manifold or random Schrödinger operators .

In retrospect, Wigner’s idea should have received even more attention. For centuries, the primary territory of probability theory was to model uncorrelated or weakly correlated systems via the law of large numbers or the central limit theorem. Random matrix statistics is essentially the first and only general computable pattern for complicated correlated systems and it is conjectured to be ubiquitous. We only mention here the spectacular result of Montgomery which proves a special case of the conjecture (under the assumption of the Riemann hypothesis) that the distribution of zeros of the Riemann zeta function on the critical line is given by a random matrix statistics.

The simplest class to test Wigner’s universality hypothesis upon is the random matrix ensemble itself. All calculations by Dyson, Gaudin and Mehta are for Gaussian ensembles, i.e., where the matrix elements hijh_{ij} are real or complex Gaussian random variables. These ensembles are called the Gaussian orthogonal ensemble (GOE) and Gaussian unitary ensemble (GUE). If Wigner’s universality hypothesis is correct, then the local eigenvalue statistics should be independent of the law of the matrix elements. This is generally referred to as the universality conjecture of random matrices and we will call it the Wigner-Dyson-Gaudin-Mehta conjecture due to the vision of Wigner and the pioneering work of these authors. It was first formulated in Mehta’s treatise on random matrices in 1967 and has remained a key question in the subject ever since. Our goal in this paper is to review the recent progress in this direction and sketch some of the important ideas.

Random matrices have been intensively studied in the last 15-20 years and we will not be able to present all aspects of this research. We refer the reader to recent comprehensive books .

The laws of random matrices can be generally divided into invariant and non-invariant ensembles. The invariant ensembles are characterized by a probability measure of the form Z−1e−NβTrV(H)/2dHZ^{-1}e^{-N\beta{\rm Tr}V(H)/2}{\rm d}H where NN is the size of the matrix, VV is a real valued potential and ZZ is the normalization constant. The parameter β>0\beta>0 is determined by the symmetry class of the model and dH{\rm d}H is the Lebesgue measure on matrices in the class. These ensembles are called invariant since the probability law depends only on the trace of a function of the matrix and thus is invariant under changes of coordinates. The matrix elements are in general correlated and they are independent if only if the model is Gaussian, i.e., VV is quadratic.

For invariant ensembles, the probability distribution of the eigenvalues λ=(λ1,…,λN)\lambda=(\lambda_{1},\dots,\lambda_{N}) with λ1≤⋯≤λN\lambda_{1}\leq\dots\leq\lambda_{N} for the measure e−NβTrV(H)/2/Ze^{-N\beta{\rm Tr}V(H)/2}/Z is given by the explicit formula (c.f. (1.4))

where the parameter β\beta is determined by the symmetry class: β=1\beta=1 for symmetric matrices, β=2\beta=2 for Hermitian matrices and β=4\beta=4 for self dual quaternion matrices. The key structural ingredient of this formula, the Vandermonde determinant, is the same as in the Gaussian case, (1.4). Thus all previous computations, developed for the Gaussian case, can be carried out for β=1,2,4\beta=1,2,4 provided that the Gaussian weight function for the orthogonal polynomials is replaced with the function e−βV(x)/2e^{-\beta V(x)/2}. Thus the analysis of the correlation functions depends critically on the the asymptotic properties of the corresponding orthogonal polynomials. In the pioneering work of Dyson, Gaudin and Mehta , the potential is the quadratic polynomial V(x)=x2/2V(x)=x^{2}/2 and the orthogonal polynomials are the Hermite polynomials whose asymptotic properties are well-known.

The extension of this approach to a general potential is a demanding task; important progress was made since the late 1990’s by Fokas-Its-Kitaev , Bleher-Its , Deift et. al. , Pastur-Shcherbina and more recently by Lubinsky . These results concern the simpler β=2\beta=2 case. For β=1,4\beta=1,4, the universality was established only quite recently for analytic VV with additional assumptions using earlier ideas of Widom . The final outcome of these sophisticated analyses is that universality holds for the measure (1.7) in the sense that the short scale behavior of the correlation functions is independent of the potential VV (with appropriate assumptions) provided that β\beta is one of the classical values, i.e., β∈{1,2,4}\beta\in\{1,2,4\}, that corresponds to an underlying matrix ensemble.

Notwithstanding matrix ensembles or orthogonal polynomials, the measure (1.7) is perfectly well defined for any β>0\beta>0 and it can be interpreted as the Gibbs measure for a system of particles with a logarithmic interaction (log-gas) at inverse temperature β\beta. It is therefore a natural question to extend universality to non-classical β\beta but the orthogonal polynomial methods are difficult to apply for this case. For all β>0\beta>0 the local statistics for the Gaussian case V(x)=x2/2V(x)=x^{2}/2 can, however, be characterized by the “Brownian carousel” which was derived from a tridiagonal matrix representation of Gaussian random matrices.

Apart from the invariant ensembles there are many natural non-invariant ensembles; the simplest and most important one being the Wigner ensemble for which the matrix elements are independent subject to a symmetry requirement, e.g. hij=hˉjih_{ij}=\bar{h}_{ji} in the Hermitian case. For non-invariant ensembles there is no explicit formula analogous to (1.7) for the joint distribution of the eigenvalues. Hence the methods for the invariant ensembles described above are not applicable. Until very recently, most rigorous results have been on the density of eigenvalues, i.e. the convergence to the the Wigner semicircle law (1.2) was established with certain error estimates, see e.g. the works by Bai et al and Guionnet and Zeitouni . The universality of the local statistics could only be established for Hermitian Wigner matrices with a substantial Gaussian component by Johansson and Ben Arous-Péché . All previous results on local universality have relied on explicitly computable algebraic formulae. These were provided by orthogonal polynomials in case of the invariant ensembles, and by a modification of the Harish-Chandra/Itzykson/Zuber integral in case of . Nevertheless, following Wigner’s thesis, universality is expected to hold for general Wigner matrices as well.

Having summarized the existing rigorous results that were available until 2008, we set the two main problems we wish to address in this article:

Problem 1: Prove the Wigner-Dyson-Gaudin-Mehta conjecture, i.e. the universality for Wigner matrices with a general distribution for the matrix elements.

Problem 2: Prove the universality of the local statistics for the log-gas (1.7) for all β>0\beta>0.

We were able to solve Problem 1 for a very general class of distributions. As for Problem 2, we solved it for the case of real analytic potentials VV assuming that the equilibrium measure is supported on a single interval, which, in particular, holds for any convex potential. We now state our results precisely.

[26, Theorem 7.2] Suppose that H=(hij)H=(h_{ij}) is a Hermitian (respectively, symmetric) Wigner matrix. Suppose that for some ε>0\varepsilon>0

Here ϱsc\varrho_{sc} is the semicircle law defined in (1.2), pN(n)p_{N}^{(n)} is the nn-point correlation function of the eigenvalue distribution of HH, and pG,N(n)p^{(n)}_{{\rm G},N} is the nn-point correlation function of an N×NN\times N GUE (respectively, GOE) matrix.

We remark that the convergence in this theorem is in weak sense, and it also involves averaging over a small energy interval E′∈[E−bN,E+bN]E^{\prime}\in[E-b_{N},E+b_{N}]. Stronger types of convergence may also be considered and we will comment on one possible such extension in Section 5. We believe that the issue of convergence types is of a technical nature and it is dwarfed by the challenge to prove universality for the largest possible family of matrix ensembles. The fundamental challenge in random matrix theory remains in answering the question of why random matrix law is ubiquitous for seemingly disparate ensembles and physical systems. We will present a few extensions in this direction in Sections 8 and 10.

In the case of invariant ensembles, it is well-known that for VV satisfying certain mild conditions the sequence of one-point correlation functions, or densities, associated with μ(N)\mu^{(N)} has a limit as N→∞N\to\infty and the limiting equilibrium density ϱ(s)\varrho(s) can be obtained as the unique minimizer of the functional

Moreover, for convex VV the support of ϱ\varrho is a single interval [A,B][A,B] and ϱ\varrho satisfies the equation

for any t∈(A,B)t\in(A,B). For the Gaussian case, V(x)=x2/2V(x)=x^{2}/2, the equilibrium density is given by the semicircle law ϱ=ϱsc\varrho=\varrho_{sc}, see (1.2).

i.e. the appropriately normalized correlation functions of the measure μβ,V(N)\mu_{\beta,V}^{(N)} at the level EE in the bulk of the limiting density asymptotically coincide with those of the Gaussian case and they are independent of the value of EE in the bulk.

We close this introduction with some short remarks concerning these two theorems. Theorem 1.1 holds for a much larger class of matrix ensembles with independent entries and we will review some of them in Sections 8 and 10. Although Theorem 1.1 in its current form was proved in , the key ideas have been developed through several important steps in . In particular, the Wigner-Dyson-Gaudin-Mehta (WDGM) conjecture for Hermitian matrices was first solved in in a joint work with the current authors and Péché, Ramírez and Schlein. This result holds whenever the distributions of matrix elements are smooth. The smoothness requirement was partially removed in and completely removed in a joint paper with Ramírez, Schlein, Tao and Vu . The WDGM conjecture for symmetric matrices was resolved in . In this paper, a novel idea based on Dyson Brownian motion was discovered. The most difficult case, the real symmetric Bernoulli matrices, was solved in where a “Fluctation Averaging Lemma” (Lemma 3.4 of the current paper) exploiting cancellation of matrix elements of the Green function was first introduced. We will give a more detailed historical review in Section 11.

The proof of Theorem 1.1 consists of the following three steps, discussed in Sections 3, 2 and 4, respectively. Our three-step strategy was first introduced in .

Step 1. Local semicircle law and delocalization of eigenvectors: It states that the density of eigenvalues is given by the semicircle law not only as a weak limit on macroscopic scales (1.2), but also in a strong sense and down to short scales containing only NεN^{\varepsilon} eigenvalues for all ε>0\varepsilon>0. This will imply the rigidity of eigenvalues, i.e., that the eigenvalues are near their classical location in the sense to be made clear in Section 2. We also obtain precise estimates on the matrix elements of the Green function which in particular imply complete delocalization of eigenvectors.

Step 2. universality for Gaussian divisible ensembles: The Gaussian divisible ensembles are matrices of the form Ht=e−t/2H0+1−e−tUH_{t}=e^{-t/2}H_{0}+\sqrt{1-e^{-t}}U, where H0H_{0} is a Wigner matrix and UU is an independent GUE matrix. The parametrization of HtH_{t} reflects that it is most conveniently obtained by an Ornstein-Uhlenbeck process. There are two methods and both methods imply the bulk universality of HtH_{t} for t=N−τt=N^{-\tau} for the entire range of 0<τ<10<\tau<1 with different estimates.

Proposition 3.1 of which uses an extension of Johansson’s formula .

Local ergodicity of the Dyson Brownian motion (DBM):

The approach in 2a yields a slightly stronger estimate than the approach in 2b, but it works only in the Hermitian case. In this review, we will focus on the Dyson Brownian approach.

Step 3. Approximation by Gaussian divisible ensembles: It is a simple density argument in the space of matrix ensembles which shows that for any probability distribution of the matrix elements there exists a Gaussian divisible distribution with a small Gaussian component, as in Step 2, such that the two associated Wigner ensembles have asymptotically identical local eigenvalue statistics. The first implementation of this approximation scheme was via a reverse heat flow argument ; it was later replaced by the Green function comparison theorem .

The proof of Theorem 1.2 consists of the following two steps that will be presented in Sections 6 and 7.

Step 1. Rigidity of eigenvalues. This establishes that the location of the eigenvalues are not too far from their classical locations determined by the equilibrium density ϱ(s)\varrho(s).

Step 2. Uniqueness of local Gibbs measures with logarithmic interactions. With the precision of eigenvalue location estimates from the Step 1 as an input, the eigenvalue spacing distributions are shown to be given by the corresponding Gaussian ones. (We will take the uniqueness of the spacing distributions as our definition of the uniqueness of Gibbs state.)

There are several similarities and differences between these two methods. Both start with rigidity estimates on eigenvalues and then establish that the local spacing distributions are the same as in the Gaussian cases. The Gaussian divisible ensembles, which play a key role in our theory for noninvariant ensembles, are completely absent for invariant ensembles. The key connection between the two methods, however, is the usage of DBM (or its analogue) in the Steps 2. In Section 2, we will first present this idea.

The method for the proof of Theorem 1.1 is extremely general. As of this writing, it has been applied to the generalized Wigner ensembles, the sample covariance ensembles and the Erdős-Rényi matrices for certain range of the sparseness parameter. It can also be extended to the edges of the spectrum, and it yields edge universality under more general conditions than were previously known. This will be reviewed in Section 9. Extensions to generalized Wigner matrices and Erdős-Rényi matrices will also be discussed in Sections 8 and 10. As the proof of Theorem 1.2 was just completed, we do not know how far this method can reach; currently we can generalize the result to the nonconvex case under the assumption that the equilibrium measure ρ\rho is supported on a single interval . The theory we have developed to prove Theorems 1.1 and 1.2 is purely analytic and we believe that it unveils the genuine mechanism of the Wigner-Dyson-Gaudin-Mehta universality. Finally, a short summary concerning the recent history of universality is given in Section 11.

Acknowledgement. The results in this review were obtained in collaboration with Benjamin Schlein, Jun Yin, Antti Knowles and Paul Bourgade and in some work, also with Jose Ramirez and Sandrine Péche. This article is to report the joint progress with these authors.

Dyson Brownian motion and the local relaxation flow

The Dyson Brownian motion (DBM) describes the evolution of the eigenvalues of a Wigner matrix as an interacting point process if each matrix element hijh_{ij} evolves according to independent (up to symmetry restriction) Brownian motions. We will slightly alter this definition by generating the dynamics of the matrix elements by an Ornstein-Uhlenbeck (OU) process which leaves the standard Gaussian distribution invariant. In the Hermitian case, the OU process for the rescaled matrix elements vij:=N1/2hijv_{ij}:=N^{1/2}h_{ij} is given by the stochastic differential equation

where βij\beta_{ij}, i<ji<j, are independent complex Brownian motions with variance one and βii\beta_{ii} are real Brownian motions of the same variance. Denote the distribution of the eigenvalues λ=(λ1,λ2,…,λN)\lambda=(\lambda_{1},\lambda_{2},\ldots,\lambda_{N}) of HtH_{t} at time tt by ft(λ)μG(dλ)f_{t}({\mathbf{\lambda}})\mu_{G}({\rm d}{\bf\lambda}) where μG\mu_{G} is given by (1.7) with the potential V(x)=x2/2V(x)=x^{2}/2.

The parameter β\beta is chosen as follows: β=2\beta=2 for complex Hermitian matrices and β=1\beta=1 for symmetric real matrices. Our formulation of the problem has already taken into account Dyson’s observation that the invariant measure for this dynamics is μG\mu_{G}. A natural question regarding the DBM is how fast the dynamics reaches equilibrium. Dyson had already posed this question in 1962:

Dyson’s conjecture : The global equilibrium of DBM is reached in time of order one and the local equilibrium (in the bulk) is reached in time of order 1/N1/N. Dyson further remarked,

“The picture of the gas coming into equilibrium in two well-separated stages, with microscopic and macroscopic time scales, is suggested with the help of physical intuition. A rigorous proof that this picture is accurate would require a much deeper mathematical analysis.”

We will prove that Dyson’s conjecture is correct if the initial data of the flow is a Wigner ensemble, which was Dyson’s original interest. Our result in fact is valid for DBM with much more general initial data that we now survey. Briefly, it will turn out that the global equilibrium is indeed reached within a time of order one, but local equilibrium is achieved much faster if an a-priori estimate on the location of the eigenvalues (also called points) is satisfied. To formulate this estimate, let γj=γj,N\gamma_{j}=\gamma_{j,N} denote the location of the jj-th point under the semicircle law, i.e., γj\gamma_{j} is defined by

We will call γj\gamma_{j} the classical location of the jj-th point.

A-priori Estimate: There exists an a>0{\mathfrak{a}}>0 such that

with a constant CC uniformly in NN. (This a-priori estimate was referred to as Assumption III in .)

The main result on the local ergodicity of Dyson Brownian motion states that if the a-priori estimate (2.5) is satisfied then the local correlation functions of the measure ftμGf_{t}\mu_{G} are the same as the corresponding ones for the Gaussian measure, μG=f∞μG\mu_{G}=f_{\infty}\mu_{G}, provided that tt is larger than N−2aN^{-2{\mathfrak{a}}}. The nn-point correlation functions of the probability measure ftdμGf_{t}{\rm d}\mu_{G} are defined, similarly to (1.3), by

Due to the convention that one can view the locations of eigenvalues as the coordinates of particles, we have used x{\bf{x}}, instead of λ{\boldsymbol{\lambda}}, in the last equation. From now on, we will use both conventions depending on which viewpoint we wish to emphasize. Notice that the probability distribution of the eigenvalues at the time tt, ftμGf_{t}\mu_{G}, is the same as that of the Gaussian divisible matrix:

where H0H_{0} is the initial Wigner matrix and UU is an independent standard GUE (or GOE) matrix. This establishes the universality of the Gaussian divisible ensembles. The precise statement is the following theorem:

We can choose b=bNb=b_{N} depending on NN. In explicit bounds on the speed of convergence and the optimal range of bb were also established. In particular, thanks to the optimal rigidity estimate , i.e., (2.5) with a=1/2{\mathfrak{a}}=1/2, the range of the energy averaging in (2.8) was reduced to bN≥N−1+ξb_{N}\geq N^{-1+\xi}, ξ>0\xi>0, but only for t≥N−ξ/8t\geq N^{-\xi/8} (Theorem 2.3 of ).

Theorem 2.1 is a consequence of the following theorem which identifies the gap distribution of the eigenvalues.

In particular, if the a-priori estimate (2.5) holds with some a>0{\mathfrak{a}}>0 and ∣J∣|J| is of order NN, then for any t>N−2a+3εt>N^{-2{\mathfrak{a}}+3\varepsilon} the right hand side converges to zero as N→∞N\to\infty, i.e. the gap distributions for ftdμGf_{t}{\rm d}\mu_{G} and dμG{\rm d}\mu_{G} coincide.

for any nn fixed which is needed to identify higher order correlation functions. In applications, JJ is chosen to be the indices of the eigenvalues in the interval [E−b,E+b][E-b,E+b] and thus ∣J∣∼Nb|J|\sim Nb. This identifies the gap distributions of eigenvalues completely and thus also identifies the correlation functions and concludes Theorem 2.1. Note that the input of this theorem, the apriori estimate (2.5), identifies the location of the eigenvalues only on a scale N−1/2−aN^{-1/2-{\mathfrak{a}}} which is much weaker than the 1/N1/N precision for the eigenvalue differences in (2.9).

By the rigidity estimates (see Corollary 3.2 below), the a-priori estimate (2.5) holds for any a<1/2{\mathfrak{a}}<1/2 if the initial data of the DBM is a Wigner ensemble. Therefore, Theorem 2.2 holds for any t≥N−1+εt\geq N^{-1+\varepsilon} for any ε>0\varepsilon>0 and this establishes Dyson’s conjecture.

2 Main ideas behind the proof of Theorem 2.2

The key method is to analyze the relaxation to equilibrium of the dynamics (2.2). This approach was first introduced in Section 5.1 of ; the presentation here follows .

and let L{\mathscr{L}} be the generator, symmetric with respect to the measure dμ{\rm d}\mu, defined by the associated Dirichlet form

Recall the relative entropy of two probability measures:

If dν=fdμ{\rm d}\nu=f{\rm d}\mu, then we will sometimes use the notation Sμ(f):=S(fμ∣μ)S_{\mu}(f):=S(f\mu|\mu). The entropy can be used to control the total variation norm via the well known inequality

Let ftf_{t} be the solution to the evolution equation

with a given initial condition f0f_{0}. The evolution of the entropy Sμ(ft)=S(ftμ∣μ)S_{\mu}(f_{t})=S(f_{t}\mu|\mu) satisfies

By Bakry and Émery , the evolution of the Dirichlet form satisfies the inequality

with some constant ϑ>0\vartheta>0, then we have

Integrating (2.15) and (2.18) back from infinity to 0, we obtain the logarithmic Sobolev inequality (LSI)

and the exponential relaxation of the entropy and Dirichlet form on time scale t∼1/ϑt\sim 1/\vartheta

As a consequence of the logarithmic Sobolev inequality, we also have the concentration inequality for any kk and a>0a>0

We will not use this inequality in this section, but it will become important in Section 6.

Returning to the classical ensembles, we assume from now on that H{\mathcal{H}} is given by (1.7) with V(x)=x2/2V(x)=x^{2}/2 and the equilibrium measure is the Gaussian one, μ=μG\mu=\mu_{G}. We then have the convexity inequality

This guarantees that μ\mu satisfies the LSI with ϑ=1/2\vartheta=1/2 and the relaxation time to equilibrium is of order one.

The key idea is that the relaxation time is in fact much shorter than order one for local observables that depend only on the eigenvalue differences. Equation (2.22) shows that the relaxation in the direction vi−vjv_{i}-v_{j} is much faster than order one provided that xi−xjx_{i}-x_{j} are close. However, this effect is hard to exploit directly due to that all modes of different wavelengths are coupled. Our idea is to add an auxiliary strongly convex potential W(x)W({\bf{x}}) to the Hamiltonian to “speed up” the convergence to local equilibrium. On the other hand, we will also show that the cost of this speeding up can be effectively controlled if the a-priori estimate (2.5) holds.

The auxiliary potential W(x)W({\bf{x}}) is defined by

i.e. it is a quadratic confinement on scale τ\sqrt{\tau} for each eigenvalue near its classical location, where the parameter τ>0\tau>0 will be chosen later. The total Hamiltonian is given by

where H{\mathcal{H}} is the Gaussian Hamiltonian given by (1.7). The measure with Hamiltonian H~\widetilde{\mathcal{H}},

will be called the local relaxation measure. This measure was named the pseudo-equilibrium measure in our previous papers.

The local relaxation flow is defined to be the flow with the generator characterized by the natural Dirichlet form w.r.t. ω{\omega}, explicitly, L~\widetilde{\mathscr{L}}:

We will typically choose τ≪1\tau\ll 1 so that the additional term WW substantially increases the lower bound (2.17) on the Hessian, hence speeding up the dynamics so that the relaxation time is at most τ\tau.

The idea of adding an artificial potential WW to speed up the convergence appears to be unnatural here. The current formulation is a streamlined version of a much more complicated approach that appeared in and which took ideas from the earlier work . Roughly speaking, in hydrodynamical limit, the short wavelength modes always have shorter relaxation times than the long wavelength modes. A direct implementation of this idea is extremely complicated due to the logarithmic interaction that couples short and long wavelength modes. Adding a strongly convex auxiliary potential W(x)W({\bf{x}}) shortens the relaxation time of the long wavelength modes, but it does not affect the short modes, i.e. the local statistics, which are our main interest. The analysis of the new system is much simpler since now the relaxation is faster, uniform for all modes. Finally, we need to compare the local statistics of the original system with those of the modified one. It turns out that the difference is governed by (∇W)2(\nabla W)^{2} which can be directly controlled by the a-priori estimate (2.5).

Our method for enhancing the convexity of H{\mathcal{H}} is reminiscent of a standard convexification idea concerning metastable states. To explain the similarity, consider a particle near one of the local minima of a double well potential separated by a local maximum, or energy barrier. Although the potential is not convex globally, one may still study a reference problem defined by convexifying the potential along with the well in which the particle initially resides. Before the particle reaches the energy barrier, there is no difference between these two problems. Thus questions concerning time scales shorter than the typical escape time can be conveniently answered by considering the convexified problem; in particular the escape time in the metastability problem itself can be estimated by using convex analysis. Our DBM problem is already convex, but not sufficiently convex. The modification by adding WW enhances convexity without altering the local statistics. This is similar to the convexification in the metastability problem which does not alter events before the escape time.

3 Some details on the proof of Theorem 2.2

The core of the proof is divided into three theorems. For the flow with generator L~\widetilde{\mathscr{L}}, we have the following estimates on the entropy and Dirichlet form.

with initial condition q0=qq_{0}=q and with the reversible measure ω\omega. Assume that ∫q0dω=1\int q_{0}{\rm d}{\omega}=1. Then we have the following estimates

with a universal constant CC. Thus the relaxation time to equilibrium is of order τ\tau:

Proof. Denote by h=qh=\sqrt{q} and we have the equation

In our case, (2.22) and (2.23) imply that the Hessian of H~\widetilde{\mathcal{H}} is bounded from below as

with some positive constant CC. This proves (2.27) and (2.28). The rest can be proved by straightforward arguments given in the earlier part of this section.

The estimate (2.28) plays a key role in the next theorem.

Proof. For simplicity, we assume that J={1,2,…,N−1}J=\{1,2,\ldots,N-1\}. Let qtq_{t} satisfy

with an initial condition q0=qq_{0}=q. We write

The second term can be estimated by (2.13), the decay of the entropy (2.30) and the boundedness of GG; this gives the second term in (2.33).

To estimate the first term in (2.34), by the evolution equation ∂qt=L~qt\partial q_{t}=\widetilde{\mathscr{L}}q_{t} and the definition of L~\widetilde{\mathscr{L}}:

From the Schwarz inequality and ∂q=2q∂q\partial q=2\sqrt{q}\partial\sqrt{q}, the last term is bounded by

where we have used (2.28) and that \Big{[}G^{\prime}(N(x_{i}-x_{i+1}))\Big{]}^{2}(x_{i}-x_{i+1})^{2}\leq CN^{-2} due to GG being smooth and compactly supported.

Alternatively, we could have directly estimated the left hand side of (2.33) by using the total variation norm between qωq{\omega} and ω{\omega}, which in turn could be estimated by the entropy (2.13) and the Dirichlet form using the logarithmic Sobolev inequality, i.e., by

However, compared with this simple bound, the estimate (2.33) gains an extra factor ∣J∣∼N|J|\sim N in the denominator, i.e. it is in terms of Dirichlet form per particle. The improvement is due to the observable in (2.33) being of special form and we exploit the term (2.28).

The final ingredient in proving Theorem 2.2 is the following entropy and Dirichlet form estimates.

Suppose that (2.22) holds. Let a>0{\mathfrak{a}}>0 be fixed and recall the definition of Q=QaQ=Q_{\mathfrak{a}} from (2.5). Fix a constant τ≥N−2a\tau\geq N^{-2{\mathfrak{a}}} and consider the local relaxation measure ω\omega with this τ\tau. Set ψ:=ω/μ\psi:={\omega}/\mu and let gt:=ft/ψg_{t}:=f_{t}/\psi. Suppose there is a constant mm such that

Then for any t≥τNεt\geq\tau N^{\varepsilon} the entropy and the Dirichlet form satisfy the estimates:

where the constants depend on ε\varepsilon and mm.

Proof. The evolution of the entropy S(ftμ∣ω)=Sω(gt)S(f_{t}\mu|{\omega})=S_{\omega}(g_{t}) can be computed explicitly by the formula

Since ω{\omega} is L~\widetilde{\mathscr{L}}-invariant and time independent, the middle term on the right hand side vanishes, and from the Schwarz inequality

Together with the logarithmic Sobolev inequality (2.29), we have

Integrating the last inequality from τ\tau to tt and using the assumption (2.37) and t≥τNεt\geq\tau N^{\varepsilon}, we have proved the first inequality of (2.38). Using this result and integrating (2.39), we have

By the convexity of the Hamiltonian, Dμ(ft)D_{\mu}(\sqrt{f_{t}}) is decreasing in tt. Since Dω(gs)≤CDμ(fs)+CN2Qτ−2D_{\omega}(\sqrt{g_{s}})\leq CD_{\mu}(\sqrt{f_{s}})+CN^{2}Q\tau^{-2}, this proves the second inequality of (2.38).

Finally, we complete the proof of Theorem 2.2. For any given t>0t>0 we now choose τ:=tN−ε\tau:=tN^{-\varepsilon} and we construct the local relaxation measure ω{\omega} with this τ\tau. Set ψ=ω/μ\psi={\omega}/\mu and let q:=gt=ft/ψq:=g_{t}=f_{t}/\psi be the density qq in Theorem 2.4. Then Theorem 2.5, Theorem 2.4 and an easy bound on the entropy Sω(q)≤CNmS_{\omega}(q)\leq CN^{m} imply that

i.e., the local statistics of ftμf_{t}\mu and ω{\omega} are the same for any initial data fτf_{\tau} for which (2.37) is satisfied. Applying the same argument to the Gaussian initial data, f0=fτ=1f_{0}=f_{\tau}=1, we can also compare μ\mu and ω{\omega}. We have thus proved (2.9) and hence the universality.

Local semicircle law via Green function

The Wigner semicircle law asserts that (1.2) is valid in a weak limit, i.e., for any smooth test function OO with compact support we have

For the rest of this paper, we will assume that the probability distribution of the matrix elements satisfy the following subexponential condition:

with some positive constants C0,ϑC_{0},\vartheta, where we set vij=Nhijv_{ij}=\sqrt{N}h_{ij}. This condition can be relaxed to (1.8) via a cutoff argument, but we will not discuss such technical details here.

the following estimates hold for any sufficiently large N≥N0(C0,ϑ)N\geq N_{0}(C_{0},\vartheta):

(i) The Stieltjes transform of the empirical eigenvalue distribution of HH satisfies

(ii) The individual matrix elements of the Green function satisfy

Theorem 3.1 is the strongest form of the local semicircle law that gives optimal error estimates (modulo logarithmic factors) on the smallest possible scale, which is valid uniformly in the spectrum including the edge, and which controls not only the Stieltjes transform but also individual matrix elements of the resolvent. This theorem is the final result of subsequent improvements in of our first local semicirle law in .

The local semicircle estimates imply that the jj-th eigenvalue, λj\lambda_{j}, is very close to its classical location γj\gamma_{j}, defined in (2.4):

[38, Theorem 2.2] Under the assumptions of Theorem 3.1 we have

for any sufficiently large N≥N0N\geq N_{0}.

This corollary in particular proves the a-priori estimate (2.5) for any a<1/2{\mathfrak{a}}<1/2.

Corollary 3.2 is a simple consequence of the Helffer-Sjöstrand formula which translates information on the Stieltjes transform of the empirical measure first to the counting function and then to the locations of eigenvalues. The formula yields the representation

We also mention that Theorem 3.1 immediately implies complete delocalization of each eigenvector of the Wigner matrix:

for any sufficiently large N≥N0N\geq N_{0}.

For the proof, notice that (3.7) implies the bound ∣Gjj(z)∣=O(1)|G_{jj}(z)|=O(1) with very high probability for any z∈SLz\in{\bf S}_{L}. Therefore,

The original proof of delocalization of eigenvectors was derived from the Stieltjes transform of the empirical measure , motivated by a question posed by T. Spencer.

Sketch of the proof of Theorem 3.1. For simplicity, we will assume here that E=Re zE={\mathfrak{Re}\,}z is away from the spectral edges. The starting point is the following well known formula. Let AA, BB, CC be n×nn\times n, m×nm\times n and m×mm\times m matrices and set

Applying this formula to the resolvent matrix G=(H−z)−1G=(H-z)^{-1}, we have

Here we used the interlacing property of eigenvalues between a matrix and its minors, which implies that

Defining vi:=Gii−mscv_{i}:=G_{ii}-m_{sc}, we thus have

Expanding the denominator, using the identity msc(z)+[msc(z)+z]−1=0m_{sc}(z)+[m_{sc}(z)+z]^{-1}=0 and neglecting the error terms hii+O(N−1)=O(N−1/2)h_{ii}+O(N^{-1})=O(N^{-1/2}), we have

Summing up ii and dividing by NN, we obtain, modulo negligible errors,

To estimate ZiZ_{i}, we compute its second moment

The second term in (3.20) can be estimated by a similar bound. These estimates confirm that the size of ZiZ_{i}, at least in the second moment sense, is roughly

Neglecting the [v]2{[v]}^{2} term in (3.18) and using that ∣1−msc2∣≥c|1-m_{sc}^{2}|\geq c away from the spectral edge for some positive cc, we thus have ∣m(z)−msc(z)∣≲C(Nη)−1/2|m(z)-m_{sc}(z)|\lesssim C(N\eta)^{-1/2}. A similar but more involved argument gives the same bound for individual viv_{i}’s, showing the estimate (3.7) for the diagonal elements GiiG_{ii}. The estimate for the off-diagonal terms, GijG_{ij}, i≠ji\neq j, is obtained from the identity G_{ij}=G_{jj}G_{ii}^{(j)}\big{[}Z_{ij}-h_{ij}\big{]} which can be proved using (3.12). Here ZijZ_{ij} is defined analogously to (3.14) as

where G(ij)G^{(ij)} is the resolvent of the (N−2)×(N−2)(N-2)\times(N-2) minor of HH after removing the ii-th and jj-th row and column. The bound (3.22) holds for ZijZ_{ij} as well.

The estimate for [v]=m−msc{[v]}=m-m_{sc}, the average of viv_{i}’s, is of order (Nη)−1(N\eta)^{-1} in (3.5), i.e. it is better than the (Nη)−1/2(N\eta)^{-1/2} estimate for the individual matrix elements in (3.7). The key mechanism for this improvement is the cancellation of the ZjZ_{j}’s in their average [Z]{[Z]}. If ZjZ_{j}’s were independent, we would gain a factor N−1/2N^{-1/2} by the central limit theorem. But ZjZ_{j}’s are correlated and the cancellation takes the following form:

With the notations of Theorem 3.1, for any ε>0\varepsilon>0 we have

Using this lemma and (3.18), we have proved the stronger estimate for [v]{[v]}. This completes the sketch of the proof of the local semicircle law, Theorem 3.1.

The Green function comparison theorems

We now state the Green function comparison theorem, Theorem 4.1. It will quickly lead to Theorem 4.2 stating that the correlation functions of eigenvalues of two matrix ensembles are identical on a scale smaller than 1/N1/N provided that the first four moments of all matrix elements of these two ensembles are almost the same. We will state a limited version for real Wigner matrices for simplicity of presentation.

[36, Theorem 2.3] Suppose that we have two N×NN\times N Wigner matrices, H(v)H^{(v)} and H(w)H^{(w)}, with matrix elements hijh_{ij} given by the random variables N−1/2vijN^{-1/2}v_{ij} and N−1/2wijN^{-1/2}w_{ij}, respectively, with vijv_{ij} and wijw_{ij} satisfying the uniform subexponential decay condition (3.3). We assume that the first four moments of vijv_{ij} and wijw_{ij} are close to each other in the sense that

holds for some δ>0\delta>0. Then there are positive constants C1C_{1} and ε\varepsilon, depending on ϑ\vartheta and C0C_{0} from (3.3) such that for any η\eta with N−1−ε≤η≤N−1N^{-1-\varepsilon}\leq\eta\leq N^{-1} and for any z1,z2z_{1},z_{2} with Im zj=±η{\mathfrak{Im}\,}z_{j}=\pm\eta, j=1,2j=1,2, we have

where G(v)G^{(v)} and G(w)G^{(w)} denotes the Green functions of H(v)H^{(v)} and H(w)H^{(w)}.

The matching condition (4.1) is essentially the same as the one appeared in . Here we formulated Theorem 4.1 for a product of two traces of the Green function, but the result holds for a large class of smooth functions depending on several individual matrix elements of the Green functions as well, see for the precise statement. (The matching condition (4.1) is slightly weaker than in , but the proof in without any change yields this slightly stronger version.) This general version of Theorem 4.1 implies the correlation functions of the two ensembles at the scale 1/N1/N are identical:

The basic idea for proving Theorem 4.1 is similar to Lindeberg’s proof of the central limit theorem, where the random variables are replaced one by one with a Gaussian one. We will replace the matrix elements vijv_{ij} with wijw_{ij} one by one and estimate the effect of this change on the resolvent by a resolvent expansion. The idea of applying Lindeberg’s method in random matrices was recently used by Chatterjee for comparing the traces of the Green functions; the idea was also used by Tao and Vu in the context of comparing individual eigenvalue distributions. There are two main differences between our method and the one that appeared in :

We compare the statistics of eigenvalues of two different ensembles near fixed energies while compared the statistics of the j1,j2,…jkj_{1},j_{2},\ldots j_{k}-th eigenvalues for fixed labels j1,j2,…jkj_{1},j_{2},\ldots j_{k}.

There is a serious difficulty in the approach concerning possible resonances of neighboring eigenvalues that may render the expansion unstable.

The Green function method eliminates this difficulty completely and Theorem 4.1 is a simple corollary of the Green function estimate Theorem 3.1.

For a sketch of the proof, fix a bijective ordering map on the index set of the independent matrix elements,

and denote by HγH_{\gamma} the Wigner matrix whose matrix elements hijh_{ij} follow the vv-distribution if ϕ(i,j)≤γ\phi(i,j)\leq\gamma and they follow the ww-distribution otherwise; in particular H(v)=H0H^{(v)}=H_{0} and H(w)=Hγ(N)H^{(w)}=H_{\gamma(N)}.

Consider the telescopic sum of differences of expectations (we present only one resolvent for simplicity of the presentation):

with a matrix QQ that has zero matrix element at the (i,j)(i,j) and (j,i)(j,i) positions.

and a similar expression holds for the resolvent SγS_{\gamma} of by HγH_{\gamma}. From the local semicircle law for individual matrix elements (3.7), the matrix elements of all Green functions RR, Sγ−1,SγS_{\gamma-1},S_{\gamma} are bounded by CNεCN^{\varepsilon} for any ε>0\varepsilon>0. By assumption (4.1), the difference between the expectation of matrix elements of Sγ−1S_{\gamma-1} and SγS_{\gamma} is of order N−2−δ+CεN^{-2-\delta+C\varepsilon}. Since the number of steps, γ(N)\gamma(N) is of order N2N^{2}, the difference in (4.4) is of order N2N−2−δ+Cε≪1N^{2}N^{-2-\delta+C\varepsilon}\ll 1, and this proves Theorem 4.1 for a single resolvent. It is very simple to turn this heuristic argument into a rigorous proof and to generalize it to the product of several resolvents. The real difficulty is the input that the local semicircle law holds for a general class of Wigner matrices.

Universality for Wigner matrices: putting it together

In this short section we put the previous information together to prove Theorem 1.1. We first focus on the case when bNb_{N} is independent of NN. Recall that Theorem 2.1 states that the correlation functions of the Gaussian divisible ensemble,

where H0H_{0} is the initial Wigner matrix and UU is an independent standard GUE (or GOE) matrix, are given by the corresponding GUE (or GOE) for t≥N−2a+εt\geq N^{-2{\mathfrak{a}}+\varepsilon} provided that the a-priori estimate (2.5) holds for the solution ftf_{t} of the forward equation (2.2) with some exponent a>0{\mathfrak{a}}>0. Since the rigidity of eigenvalues, Corollary 3.2, holds uniformly for all Wigner matrices, we have proved (2.5) for a=1/2−ε{\mathfrak{a}}=1/2-\varepsilon with any ε>0\varepsilon>0.

From the evolution of the OU process (2.1) for vij=N1/2hijv_{ij}=N^{1/2}h_{ij} we have

The argument for NN-dependent b=bNb=b_{N} in the range bN≥N−1+ξb_{N}\geq N^{-1+\xi}, ξ>0\xi>0, is slightly different. For such a small bNb_{N}, (2.8) could be established only for relatively large times, t≥N−ξ/8t\geq N^{-\xi/8}. We cannot therefore compare H0H_{0} with HtH_{t} directly, since the deviation of the third moments of vij(0)v_{ij}(0) and vij(t)v_{ij}(t) in (5.2) would not satisfy (4.1). Instead, we construct an auxiliary Wigner matrix H^0\widehat{H}_{0} such that up to the third moment its time evolution H^t\widehat{H}_{t} under the OU flow (5.1) matches exactly the original matrix H0H_{0} and the fourth moments are close even for tt of order N−ξ/8N^{-\xi/8} (see Lemma 3.4 of ). Theorem 2.1 will then be applied for H^t\widehat{H}_{t}, and Theorem 4.1 can be used to compare H^t\widehat{H}_{t} and H0H_{0}.

We finally discuss the extension of Theorem 1.1 without averaging in E′E^{\prime}. For Hermitian matrices, with the notations of Theorem 1.1, for any fixed ∣E∣<2|E|<2 we have that

This convergence was first proved in Theorem 1.1 of for matrices with distribution which is CnCn-times differentiable for some universal constant CC. For a general distribution it was stated as Theorem 5 in . Although the proof in took a slightly different path, this generalization is an immediate corollary of our previous results . Recall our three step approach reviewed in the introduction. If we substitute Step 2b with Step 2a, then all our results in the Hermitian case would need no time average. More precisely, Proposition 3.1 of asserts that the bulk universality in the Hermitian case holds at a fixed energy for the Gaussian convolution matrix HtH_{t} with t∼N−1+δt\sim N^{-1+\delta}. The first four moments of HtH_{t} and H0H_{0} are sufficiently close to apply directly the Green function comparison theorem for correlation functions (Theorem 4.2 in this article). This concludes the bulk universality of the original matrix H0H_{0} at a fixed energy, which is the Theorem 5 in . In fact, our theory implies the same result for generalized Hermitian matrices (defined in Section 8) with finite 4+ε4+\varepsilon moments.

Beta ensemble: Rigidity estimates

Similarly to (2.4) we again denote by γk\gamma_{k} the classical location of the kk-th point w.r.t. the limiting equilibrium density ϱ(s)\varrho(s), i.e. γk\gamma_{k} is defined by

[10, Theorem 3.1] Fix any α,ε>0\alpha,\varepsilon>0 and assume that (6.1) holds. Then there are constants δ,c1,c2>0\delta,c_{1},c_{2}>0 such that for any N≥1N\geq 1 and k∈⟦αN,(1−α)N⟧k\in\llbracket\alpha N,(1-\alpha)N\rrbracket,

The first ingredient to prove Theorem 6.1 is an analysis of the loop equation following Johansson and Shcherbina . The equilibrium density ϱ\varrho, for a convex potential VV, is given by

Let mˉN\bar{m}_{N} and mm be the Stieltjes transforms of the density pN(1)p_{N}^{(1)} and the equilibrium density ϱ\varrho, respectively. Notice that in Section 3 we have used m=mNm=m_{N} to denote the Stieltjes transform of the empirical measure (3.2); here mˉN\bar{m}_{N} denotes the ensemble average of the analogous quantity.

The equation used by Johansson (which can be obtained by a change of variables in (6.4) or by integration by parts ), is a variation of the loop equation (see, e.g., ) used in the physics literature and it takes the form

Equation (6.5) expresses the difference mˉN−m\bar{m}_{N}-m in terms of (mˉN−m)2(\bar{m}_{N}-m)^{2}, bNb_{N} and cNc_{N}. In the regime where ∣mˉN−m∣|\bar{m}_{N}-m| is small, we can neglect the quadratic term. The term bNb_{N} is of the same order as ∣mˉN−m∣|\bar{m}_{N}-m| and is difficult to treat. As observed in , for analytic VV, this term vanishes when we perform a contour integration. So we have roughly the relation

The key idea in this section is the observation that the accuracy information on the λ\lambda’s can be used to improve the local convexity of the measure μ\mu in the direction involving the differences of λ\lambda’s. To explain this idea, we compute the Hessian of the Hamiltonian of μ\mu:

The naive lower bound on ∇2H\nabla^{2}{\mathcal{H}} is ϑ\vartheta, but for a typical λ=(λ1,λ2,…,λN)\lambda=(\lambda_{1},\lambda_{2},\ldots,\lambda_{N}) it is in fact much better in most directions. To see this effect, suppose we know ∣λi−λj∣≲M/N|\lambda_{i}-\lambda_{j}|\lesssim M/N with some MM for any i,j∈IkMi,j\in I_{k}^{M}, where IkM:=⟦k−M,k+M⟧I_{k}^{M}:=\llbracket k-M,k+M\rrbracket. Then for v=(vk−M,…,vk+M){\bf{v}}=(v_{k-M},\ldots,v_{k+M}) with ∑jvj=0\sum_{j}v_{j}=0 we have

This improves the convexity of the Hessian to N/MN/M on the hyperplane ∑jvj=0\sum_{j}v_{j}=0. Let

denote the block average of the locations of particles and rewrite

as a telescopic sum with an appropriate sequence of M1=0M_{1}=0, M2,…M_{2},\ldots. We can now use the improved concentration on the hyperplane ∑jvj=0\sum_{j}v_{j}=0 to the variables λk[Mj]−λk[Mj+1]\lambda_{k}^{[M_{j}]}-\lambda_{k}^{[M_{j+1}]} to control the fluctuation of λk−λk[N1−ε]\lambda_{k}-\lambda_{k}^{[N^{1-\varepsilon}]}. Since the fluctuation of λk[N1−ε]\lambda_{k}^{[N^{1-\varepsilon}]} is very small for small ε\varepsilon, we finally arrive at the estimate

Beta ensemble: The local equilibrium measure

Having completed the first step, the rigidity estimate, we now focus on the second step, i.e. on the uniqueness of the local Gibbs measure. Let 0<κ<1/20<\kappa<1/2. Choose q∈[κ,1−κ]q\in[\kappa,1-\kappa] and set L=[Nq]L=[Nq] (the integer part). Fix an integer K=NkK=N^{k} with k<1k<1. We will study the local spacing statistics of KK consecutive particles

These particles are typically located near EqE_{q} determined by the relation

We will distinguish the inside and outside particles by renaming them as

all in increasing order, i.e. x∈Ξ(K){\bf{x}}\in\Xi^{(K)} and y∈Ξ(N−K){\bf{y}}\in\Xi^{(N-K)}. We will refer to the yy’s as external points and to the xx’s as internal points.

We will fix the external points (also called as boundary conditions) and study conditional measures on the internal points. We define the local equilibrium measure on x{\bf{x}} with fixed boundary condition y{\bf{y}} by

Note that for any fixed y∈Ξ(N−K){\bf{y}}\in\Xi^{(N-K)}, the measure μy\mu_{\bf{y}} is supported on configurations of KK points x={xj}j∈I{\bf{x}}=\{x_{j}\}_{j\in I} located in the interval [yL,yL+K+1][y_{L},y_{L+K+1}].

The Hamiltonian Hy{\mathcal{H}}_{\bf{y}} of the measure μy(dx)∼exp⁡(−NHy(x))dx\mu_{\bf{y}}({\rm d}{\bf{x}})\sim\exp(-N{\mathcal{H}}_{\bf{y}}({\bf{x}})){\rm d}{\bf{x}} is given by

We now define the set of good boundary configurations with a parameter δ=δ(N)>0\delta=\delta(N)>0

where κ\kappa is a small constant to cutoff points near the spectral edges. Some rather weak additional conditions for y{\bf{y}} near the spectral edges will also be needed, but we will neglect this issue here.

Let σ\sigma and μ\mu be two measures of the form (1.7) with potentials WW and VV and densities ϱ=ϱW\varrho=\varrho_{W} and ϱV\varrho_{V}, respectively. For our purpose W(x)=x2/2W(x)=x^{2}/2, i.e., σ\sigma is the Gaussian β\beta-ensemble and ϱW(t)=12π(4−t2)+1/2\varrho_{W}(t)=\frac{1}{2\pi}(4-t^{2})^{1/2}_{+} is the Wigner semicircle law. Let the sequence γj\gamma_{j} be the classical locations for μ\mu and the sequence θj\theta_{j} be the classical locations for σ\sigma. Similarly to the construction of the measure μy\mu_{{\bf{y}}}, for any positive integer L′∈⟦1,N−K⟧L^{\prime}\in\llbracket 1,N-K\rrbracket we can construct the measure σθ\sigma_{\boldsymbol{\theta}} conditioned that the particles outside are given by the classical locations θj\theta_{j} for j∉⟦L′,L′+K⟧j\notin\llbracket L^{\prime},L^{\prime}+K\rrbracket. More precisely, we define a reference local Gaussian measure σθ\sigma_{\theta} on the set [θL′,θL′+K+1][\theta_{L^{\prime}},\theta_{L^{\prime}+K+1}] via the Hamiltonian

where I′:=⟦L′+1,L′+K⟧I^{\prime}:=\llbracket L^{\prime}+1,L^{\prime}+K\rrbracket. Since L′L^{\prime} will not play an active role, we will abuse the notation and set L′=LL^{\prime}=L.

The measure μy\mu_{{\bf{y}}} lives on the interval [yL,yL+K+1][y_{L},y_{L+K+1}] while the measure σθ\sigma_{{\boldsymbol{\theta}}} lives on the interval [θL,θL+K+1][\theta_{L},\theta_{L+K+1}] and it is difficult to compare them. But after an appropriate translation and dilation, they will live on the same interval and from now on we assume that [yL,yL+K+1]=[θL,θL+K+1][y_{L},y_{L+K+1}]=[\theta_{L},\theta_{L+K+1}]. The parameter K=NkK=N^{k} has to be sufficiently small since ϱV\varrho_{V} and ϱW\varrho_{W} are not constant functions and we have to match these two densities quite precisely in the whole interval. There are some other subtle issues related to the rescaling, but we will neglect them here to concentrate on the main ideas. Our main result is the following theorem which is essentially a combination of Proposition 4.2 and Theorem 4.4 from .

Let 0<φ≤1380<\varphi\leq\frac{1}{38}. Fix K=NkK=N^{k}, δ=N−d\delta=N^{-d} with d=1−φd=1-\varphi and k=392φk=\frac{39}{2}\varphi. Then for y∈G{\bf{y}}\in{\mathcal{G}} we have

as N→∞N\to\infty for any smooth and compactly supported test function GG. A similar formula holds for more complicated observables of the form (2.10).

The basic idea for proving Theorem 7.1 is to use the Dirichlet form inequality (2.33). Although (2.33) was stated for an infinite volume measure, it holds for any measure with repulsive logarithmic interactions in a finite volume and with the parameter τ−1\tau^{-1} being the lower bound on the Hessian of the Hamiltonian. In our setting, we denote by τσ−1\tau_{\sigma}^{-1} the lower bound for ∇2Hσ\nabla^{2}{\mathcal{H}}_{\sigma}, and the Dirichlet form inequality becomes

Using the equilibrium relation (1.10) between the potentials VV, WW and the densities ϱV\varrho_{V}, ϱW\varrho_{W}, we have

Hence ZjZ_{j} is the sum of the error terms,

and there is a term similar to AjA_{j} with yjy_{j} replaced by θj\theta_{j} and ϱV\varrho_{V} replaced by ϱW\varrho_{W}.

With our convention, the total numbers of particles in the interval [yL+K+1,yL][y_{L+K+1},{y_{L}}] are equal and thus

Since the densities ρV\rho_{V} and ρW\rho_{W} are C1C^{1} functions away from the endpoints AA and BB and yL+K+1−yLy_{L+K+1}-{y_{L}} is small, ∣ρV−ρW∣|\rho_{V}-\rho_{W}| is small in the interval [yL+K+1,yL][y_{L+K+1},{y_{L}}] and thus BjB_{j} is small. For estimating AjA_{j}, we can replace the integral ∫−∞yLϱV(y)xj−ydy\int_{-\infty}^{y_{L}}\frac{\varrho_{V}(y)}{x_{j}-y}{\rm d}y by 1N∑k<L1xj−γk\frac{1}{N}\sum_{k<L}\frac{1}{x_{j}-\gamma_{k}} with negligible errors, at least for jj’s away from the edges, j∈⟦L+Nε,L+K−Nε⟧j\in\llbracket L+N^{\varepsilon},L+K-N^{\varepsilon}\rrbracket. Thus

and TjkT^{k}_{j} can be estimated by the assumption ∣yk−γk∣≤δ|y_{k}-\gamma_{k}|\leq\delta from y∈G{\bf{y}}\in{\mathcal{G}}. The same argument works if jj is close to the edge, but kk is away from the edges, i.e. k≤L−Nεk\leq L-N^{\varepsilon} or k≥L+K+Nεk\geq L+K+N^{\varepsilon}. The edge terms, TjkT^{k}_{j} for ∣j−k∣≤Nε|j-k|\leq N^{\varepsilon}, are difficult to estimate due to the singularity in the denominator and the event that many yky_{k}’s with k<Lk<L may pile up near yLy_{L}. To resolve this difficulty, we show that the averaged local statistics of the measure μy\mu_{\bf{y}} are insensitive to the change of the boundary conditions for y{\bf{y}} near the edges. This can be achieved by the simple inequality

for any two boundary conditions y{\bf{y}} and y′{\bf{y}}^{\prime}. Although we still have to estimate the entropy that includes a logarithmic singularity, this can be done much more easily. Therefore, we can replace the boundary condition yky_{k} with yk′=θky_{k}^{\prime}=\theta_{k} for ∣j−k∣≤Nε|j-k|\leq N^{\varepsilon} and then the most singular edge terms in (7.10) cancel out.

We note that we can perform this replacement only for a small number of index pairs (j,k)(j,k), since estimating the gap distribution by the total entropy, as noted in (2.36) in Section 2, is not as efficient as the estimate using the Dirichlet form per particle. Thus we can afford to use this argument only for the edge terms, ∣j−k∣≤Nε|j-k|\leq N^{\varepsilon}. For all other index pairs (j,k)(j,k) we still have to estimate TjkT_{j}^{k} by exploiting that y{\bf{y}} is a good configuration, i.e. yk−γky_{k}-\gamma_{k} is small.

Unfortunately, even with the optimal accuracy δ∼N−1+ε′\delta\sim N^{-1+\varepsilon^{\prime}} in (7.4) as an input, the relation (7.9) still cannot be satisfied for any choice of Ncε′≤K≤N1−cε′N^{c\varepsilon^{\prime}}\leq K\leq N^{1-c\varepsilon^{\prime}}. We do not know whether this is due to our handling of the edge terms or some other intrinsic reasons. To understand why this might occur, we remark that while the edge terms become a smaller percentage of the total terms in (7.14) as KK gets bigger, the relaxation time to equilibrium for σθ\sigma_{\boldsymbol{\theta}}, determined by the convexity of Hθ′′{\mathcal{H}}_{\boldsymbol{\theta}}^{\prime\prime}, increases at the same time. At the end of our calculation, there is no good regime for the choice of KK. Fortunately, this can be resolved by using the idea of the local relaxation measure , i.e., we add a quadratic term 12τ(xj−γj)2\frac{1}{2\tau}(x_{j}-\gamma_{j})^{2} to the measure μy\mu_{\bf{y}} and 12τσ(xj−θj)2\frac{1}{2\tau_{\sigma}}(x_{j}-\theta_{j})^{2} to the measure σθ\sigma_{\boldsymbol{\theta}}. With these ideas, we can complete the proof of Theorem 7.1.

More general classes of random matrices

for some fixed positive constants CinfC_{inf} and CsupC_{sup}. These ensembles are called generalized Wigner ensembles. In the special case σij2=1/N\sigma_{ij}^{2}=1/N, we recover the original Wigner ensemble. All our results concerning the bulk universality, delocalization of eigenvectors and local semicircle laws hold for generalized Wigner matrices as well.

There is another important class of random matrices, the band matrices, which are characterized by the property that σij2\sigma_{ij}^{2} is a function of ∣i−j∣|i-j| on scale WW, which is called the bandwidth, i.e.,

The significance of the random band matrices stems from the fact that they interpolate between discrete random Schrödinger operators with short range hoppings (Anderson model) and the Wigner matrices. In particular, random matrix spectral statistics are expected to hold in the presumed delocalization regime of the Anderson model in three or higher dimensions. For more details on this exciting connection, see .

Finally we mention the ensemble of sample covariance matrices that play a fundamental role in statistics. These are matrices of the form H=A∗AH=A^{*}A where AA is an M×NM\times N matrix with independent identically distributed entries. The semicircle law is replaced with the Marchenko-Pastur law, but most results listed in this review remain valid. For more details, see .

Edge universality

Denote by λN\lambda_{N} is the largest eigenvalue of a generalized Wigner matrix. The probability distribution functions of λN\lambda_{N} for the classical Gaussian ensembles are identified by Tracy and Widom to be

where the functions Fβ(s)F_{\beta}(s) can be computed in terms of Painlevé equations and β=1,2,4\beta=1,2,4 corresponds to the standard classical ensembles. The distribution of λN\lambda_{N} is believed to be universal and independent of the Gaussian structure.

The local semicircle law, Theorem 3.1, combined with a modification of the Green function comparison theorem, Theorem 4.1, implies the following version of universality of the extreme eigenvalues. Although it holds for correlation functions of finite number of eigenvalues, for simplicity we state it for the largest one and for the case of symmetric matrices only.

Then there is an ε>0\varepsilon>0 depending on ϑ\vartheta in (3.3) such that for any real parameter ss (may depend on NN) we have

for N≥N0N\geq N_{0} sufficiently large, where N0N_{0} is independent of ss.

Note that although Theorem 9.1 states that the edge distribution is universal for a fixed choice of the variances σij2\sigma_{ij}^{2}, it does not identify this distribution. In particular, we do not know if it coincides with the Tracy-Widom distribution apart from the Hermitian case, when the method of can be applied. The extension of Theorem 9.1 to eigenvectors was recently obtained by Knowles and Yin , i.e., under the assumption (9.2), the distributions for the largest eigenvectors coincide. Similar results hold for the joint distribution of eigenvectors near the edges.

Erdős-Rényi matrix

The Erdős-Rényi matrix is the adjacency matrix of the Erdős-Rényi random graph . Its entries are independent (up to the constraint that the matrix be symmetric) and are equal to 11 with probability pp and with probability 1−p1-p. We rescale the matrix in such a way that its bulk eigenvalues typically lie in an interval of size of order one. Thus we have a symmetric N×NN\times N matrix A=(aij)A=(a_{ij}) whose entries aija_{ij} are independent (up to the symmetry constraint aij=ajia_{ij}=a_{ji}) and each element is distributed according to

Here γ:=(1−q2/N)−1/2\gamma\mathrel{\mathop{:}}=(1-q^{2}/N)^{-1/2} is a scaling introduced for convenience to compare with Wigner matrices. We also assume that q=pN≥(log⁡N)Clog⁡log⁡Nq=\sqrt{pN}\geq(\log N)^{C\log\log N}, in particular the Erdős-Rényi graph is connected.

[25, Theorem 2.9] Let m(z)m(z) denote the Stieltjes transform of the empirical eigenvalue distribution of the matrix AA and let G(z)=(A−z)−1G(z)=(A-z)^{-1} be its resolvent. Assume that the spectral parameter z=E+iηz=E+i\eta satisfies ∣E∣≤5|E|\leq 5 and (log⁡N)LN−1≤η≤3(\log N)^{L}N^{-1}\leq\eta\leq 3 with a sufficiently large constant LL. Then we have the following two estimates:

(i) The Stieltjes transform of the empirical eigenvalue distribution of AA satisfies

(ii) The individual matrix elements of the Green function satisfy that

Compared with the local semicircle law, Theorem 3.1, there is an extra factor 1/q1/q appearing in the error estimates of Theorem 10.1. This extra error term affects the rigidity estimate of eigenvalues, and (3.2) becomes

for q=pN≫N1/3q=\sqrt{pN}\gg N^{1/3}. We also have an estimate for the regime q≤N1/3q\leq N^{1/3} but that is weaker. Moreover, under the assumption q≫N1/3q\gg N^{1/3}, both bulk and edge universality are proved (see Theorem 2.5 and 2.7 in ). It is well-known that the largest eigenvalue λN\lambda_{N} of AA satisfies

hence it is located far away from the bulk spectrum. Therefore the edge universality for Erdős-Rényi matrices refers to the second largest eigenvalue instead of the largest one. Since the matrix elements of AA have nonzero means, both the edge and bulk universality require substantial new ideas in addition to those we have sketched. We refer the interested readers to the original papers for more detailed explanations.

Historical Remarks

Finally we summarize the recent history related to the universality of local eigenvalue statistics of Wigner matrices. The three-step approach was first introduced in in the context of Hermitian Wigner matrices and it led to the first proof of the Wigner-Dyson-Gaudin-Mehta conjecture for Hermitian Wigner matrices. It works whenever the distributions of the matrix elements are smooth. This approach was followed by all later works on the bulk universalities. We now review the history of Steps 1-3 separately and we start with the history of Step 1, the local semicircle law.

The semicircle law was proved by Wigner for energy windows of order one. Various improvements were made to shrink the spectral windows; in particular, results down to scale N−1/2N^{-1/2} were obtained by and . The result at the optimal scale, N−1N^{-1}, referred to as the local semicircle law, was established for Wigner matrices in a series of papers . The method was based on a self-consistent equation for the Stieltjes transform of the eigenvalues, m(z)m(z), and the continuity in the imaginary part of the spectral parameter zz. As a by-product, the optimal eigenvector delocalization estimate was proved. In order to deal with the generalized Wigner matrices, we needed to consider the self-consistent equation of Gij(z)G_{ij}(z), the matrix elements of the Green function, since there is no closed equation for m(z)=N−1\mboxTr G(z)m(z)=N^{-1}\mbox{Tr\,}G(z) . In particular, this method implied the optimal rigidity estimate of eigenvalues in the bulk in and up to the edges in . The estimate on GiiG_{ii} provided a simple alternative proof of the eigenvector delocalization estimate. The extension of the local semicircle law to the Erdős-Rényi matrices was recently made in .

We now review the history of Step 2. Recall that Hermitian Gaussian divisible ensembles are matrices of the form e−t/2H0+(1−e−t)1/2Ue^{-t/2}H_{0}+(1-e^{-t})^{1/2}U, where UU is the GUE and H0H_{0} is a Wigner ensemble. The universality of this ensemble for a large class of H0H_{0} and for parameters tt of order one was proved by Johansson . It was extended to complex sample covariance matrices by Ben Arous and Péché . There were two major restrictions of this method: 1. The Gaussian component was fairly large, it was required to be of order one independent of NN; 2. The method relies on an explicit formula by Brézin-Hikami for the correlation functions of eigenvalues. This formula originates in the Harish-Chandra-Itzykson-Zuber integral and it is valid only for Gaussian divisible ensembles with unitary invariant Gaussian component. The size of the Gaussian component was reduced to N−1+εN^{-1+\varepsilon} in by using an improved formula for correlation functions and the local semicircle law from .

To eliminate the usage of an explicit formula, a conceptual approach for Step 2 via the local ergodicity of Dyson Brownian motion was initiated in . In this paper, the first version of the local relaxation flow was introduced, but it was rather complicated. In we found a much simpler way to enhance the convexity of the Dyson Brownian motion and we proved a general theorem for local ergodicity of DBM and related flow, i.e., Theorem 2.1. This theorem applies to all classical ensembles, i.e., real and complex Wigner matrices, real and complex sample covariance matrices and quaternion Wigner matrices. The local relaxation flow in the simple form (2.25) first appeared in . The relaxation time to local equilibrium proved in these two papers was not optimal; the optimal relaxation time, conjectured by Dyson, was obtained later in .

The third and final step is to approximate the local eigenvalue distribution of a general Wigner matrix by that of a Gaussian divisible one. The first approximation result was obtained via the reversal heat flow in which required some smoothness of the distribution of matrix elements. Shortly after, Tao and Vu , proved a comparison theorem with a four moment matching condition. Instead of using a Gaussian divisible ensemble with a small (N−1+εN^{-1+\varepsilon}) Gaussian component, they relied on Johansson’s result to provide Hermitian Gaussian divisible ensembles for comparison. This proved the universality of Hermitian Wigner matrices, provided that the distributions of matrix elements have vanishing third moment and are supported on at least three points. These conditions were removed in by combining the arguments of and .

Due to the lack of a Brézin-Hikami type formula for the symmetric matrices, there was no extension of Johansson’s result to this case and the universality for symmetric Wigner ensembles was much more difficult to prove. However, the result of implies that the local eigenvalue statistics of symmetric Wigner matrices and GOE are the same, but under the restriction that the first four moments of the matrix elements exactly match those of GOE. The resolution of the Wigner-Dyson-Gaudin-Mehta conjecture for symmetric matrices, i.e., Theorem 2.1 for real symmetric matrices, was obtained in . In these papers, two new ideas were introduced: the local relaxation flow and the Green function comparison theorem . Starting from the paper , the variances were allowed to vary and the universality was extended to generalized Wigner matrices. The real Bernoulli random matrices required a more refined argument . Finally, the technical condition assumed in all these papers, i.e., that the probability distributions of the matrix elements decay subexponentially, was reduced to the (4+ε)(4+\varepsilon)-moment assumption (1.8) by using the universality of Erdős-Rényi matrices .

The Green function comparison theorem, Theorem 4.1, uses the same four moment conditions which appeared earlier in , but it compares matrix elements of Green functions at a fixed energy and not just traces of Green functions which carry information on eigenvalues near a fixed energy. The result of , on the other hand, concerns individual eigenvalues with fixed labels. Both proofs used the local semicircle law and Lindeberg’s idea (introduced in his proof of the central limit theorem). Lindeberg’s idea in the context of random matrices appeared earlier in a proof of the Wigner semicircle law by Chatterjee . The approach requires additional difficult estimates due to singularities from neighboring eigenvalues, but the Green function comparison theorem follows directly from the local semicircle law in Step 1, i.e., Theorem 3.1, via standard resolvent expansions. The difficulties associated with the singularities of eigenvalue resonances are completely absent in the Green function comparison theorem. Finally, we mention that Green function comparison can also yield comparison of eigenvalues with fixed labels, see the recent work by Knowles and Yin .

The edge universality for Wigner matrices was first proved via the moment method by Soshnikov (see also the earlier work ) for Hermitian and symmetric ensembles with symmetric distributions. By combining the moment method and Chebyshev polynomials, Sodin proved edge universality of certain band matrices and some special class of sparse matrices with symmetric distribution. The symmetry assumption was partially removed in . The edge universality without any symmetry assumption was proved in under the condition that the distribution of matrix element is subexponential decay and the first three moments match those of a Gaussian distribution. The subexponential decay condition is not optimal for edge universality, in fact the finiteness of the fourth moment was conjectured to be sufficient. For Gaussian divisible Hermitian ensembles this was proved in . This is optimal, since on the other hand, the result by Auffinger, Ben Arous and Péché showed that the distribution of the largest eigenvalues converges to a Poisson process if the entries have at most 4−ε4-\varepsilon moments. For Wigner matrices with arbitrary symmetry class, the edge universality was proved under the sole assumption that the matrix entries have 12+ε12+\varepsilon moments . Finally, we mention that extension of universality to eigenvectors near the edge was obtained by Knowles and Yin under the two moment matching condition and with four moment matching condition in .

Although we have focused only on Wigner matrices and β\beta-ensembles, the ideas summarized in this review should be applicable to a wide class of matrix ensembles. We have already mentioned some natural open questions related to possible improvements of our results. These concern removing some technical conditions such as (i) the restriction q≫N1/3q\gg N^{1/3} in the bulk universality of the Erdős-Rényi matrix; (ii) the 12+ε12+\varepsilon moment condition for edge universality. A more ambitious goal would be to prove universality for systems with some spatial structure such as band matrices or related models that may open up a path towards universality for random Schrödinger operators and other realistic models of quantum chaos.

References