Bulk universality for Wigner hermitian matrices with subexponential decay

Laszlo Erdos, Jose Ramirez, Benjamin Schlein, Terence Tao, Van Vu, Horng-Tzer Yau

Introduction

This paper is concerned with the bulk eigenvalue statistics of large Hermitian Wigner random matrices, which we now pause to define.

Let λ1≤…≤λn\lambda_{1}\leq\ldots\leq\lambda_{n} be the eigenvalues of a Wigner Hermitian matrix HnH_{n}, and let pn(x1,…,xn)p_{n}(x_{1},\ldots,x_{n}) be the associated (symmetric) density function, normalized to have total mass 11. For any 1≤k≤n1\leq k\leq n, the kk-point correlation function pn(k)(x1,…,xk)p^{(k)}_{n}(x_{1},\ldots,x_{k}) is defined by the formula

Here we follow the convention that the correlation functions are L1L^{1}-normalized. Other conventions are also used in the literature. e.g. the correlation functions RkR_{k} used in are related to pn(k)p^{(k)}_{n} via Rk=n!(n−k)!pn(k)R_{k}=\frac{n!}{(n-k)!}p^{(k)}_{n}.

The main results of this note are as follows.

and ρsc(u)\rho_{sc}(u) is the Wigner semicircle distribution

Fix uu such that −2<u<2-2<u<2, choose k≥1k\geq 1 and s>0s>0, and let HH be a Wigner random matrix. Then the expectation value of the empirical gap distribution

where Kα{\mathcal{K}}_{\alpha} is the integral operator on L2((0,α))L^{2}((0,\alpha)) whose kernel is the Dyson sine kernel (2).

We now briefly recall the previous partial results towards these theorems. In the case of GUE, the proofs of Theorems 2, 3 can be found in . These proofs relied heavily on the explicit formula for the kk-point correlation functions in the GUE case. An important extension was made by Johansson , who established the above theorems in Gaussian divisible case when the matrix HH took the form H^+aV\hat{H}+aV, where H^\hat{H} is another Wigner matrix, VV is a GUE matrix independent of H^\hat{H} and a>0a>0 is an arbitrary fixed constant independent of nn. This proof also requires the use of an explicit formula, however, its analysis is much more delicate than in the GUE case. By a simple rescaling, it is clear that Johansson’s result extends to Wigner matrices of the form

where H^\hat{H} and VV are as above, and the time t>0t>0 is fixed independent of nn. One can view HH as the evolution of H^\hat{H} under an Ornstein-Uhlenbeck process of time tt.

In , the other two authors introduced a different way to study the local eigenvalue statistics of random matrices, based on a variant of Lindeberg replacement strategy. This enables one to compare the statistics of a general Wigner matrix to those of a Johansson matrix, or any random matrix where the desired statistics is known, given that some moment condition is satisfied. As an application, Theorems 2, 3 were established in under the additional hypothesis that the atom distribution xx had vanishing third moment and was supported on at least three points. For more details, we refer to [7, Section 1].

In this note, we observe that one can combine the arguments from with the arguments of to eliminate the regularity hypotheses from the former and the moment and support hypotheses from the latter. The basic idea is simple and can be described as follows. The main result of (see [7, Theorem 15]) allows one to compare the local statistics of two Wigner hermitian matrices whose atom variables ξ,ξ′\xi,\xi^{\prime} have the same third and fourth moments, namely Eξ3=Eξ′3{\mathbf{E}}\xi^{3}={\mathbf{E}}\xi^{\prime 3} and Eξ4=Eξ′4{\mathbf{E}}\xi^{4}={\mathbf{E}}\xi^{\prime 4}. A closer look reveals that it is enough to require that the differences ∣Eξ3−Eξ′3∣|{\mathbf{E}}\xi^{3}-{\mathbf{E}}\xi^{\prime 3}| and ∣Eξ4−Eξ′4∣|{\mathbf{E}}\xi^{4}-{\mathbf{E}}\xi^{\prime 4}| are sufficiently small (in term of nn). This allows us to compare the original matrix HH with a Johansson-type matrix (5) with t=n−1+δt=n^{-1+\delta} for some small δ\delta, instead of comparing with a Johansson matrix with t=Θ(1)t=\Theta(1) as done in . On the other hand, arguments from guarantee that such a Johansson matrix has the desired statistics.

One defect of our method is that the statistics need to be averaged over a spatial scale ε{\varepsilon} which is independent of nn before universality is established. This appears to be a largely technical condition, arising from the lack of good concentration results for individual eigenvalues λi\lambda_{i} for matrices of the form (5) (in particular, we do not know how to eliminate the vanishing third moment hypothesis from [7, Theorem 29]). However, if we reimpose the hypothesis that the third moment of the atom distribution vanishes (i.e., E x3=0{\mathbf{E}}\,x^{3}=0), we can remove the averaging in Theorem 2, thus refining [7, Theorem 11] by removing the condition that the support must have at least 3 points (see Section 3).

Proof of theorems

We now prove Theorem 2. We fix u,ε,f,Hu,{\varepsilon},f,H; by a limiting argument we may take ff to be smooth. We will always assume nn to be sufficiently large depending on these parameters. Using the exponential decay of xx and a standard truncation argument (see [7, Section 3.1]) we may assume that

almost surely. (Note that the Gaussian random variable N(0,1/2)N(0,1/2) does not quite obey this bound, but can be modified on a set of extremely small probability to do so whenever necessary, so we may effectively assume that (6) holds in this case also.)

where t:=n−1+δt:=n^{-1+\delta} and VV is a GUE matrix independent of HH. We will choose δ=0.01\delta=0.01 (say), but the proof below actually requires only δ<1/2\delta<1/2.

Observe that H′H^{\prime} is another Wigner random matrix, with atom distribution x′x^{\prime} given by

where g≡N(0,1/2)g\equiv N(0,1/2) is a Gaussian random variable independent of xx. In particular we have the moment comparision bounds

Let (pn′)(k)(p^{\prime}_{n})^{(k)} be the kk-point correlation function associated to H′H^{\prime}. The first step is to show that Theorem 2 is valid for H′H^{\prime}:

By the dominated convergence theorem, it suffices to show that

for every fixed −2<u<2-2<u<2. The dominated convergence theorem can be applied since the integral on the lhs of (9) can be bounded by Ck E NIkC_{k}\,{\mathbf{E}}\,N_{I}^{k}, where NIN_{I} is the number of eigenvalues in an interval II of size C/nC/n about uu. The boundedness of E NIk{\mathbf{E}}\,N_{I}^{k} follows from [2, Theorem 5.1]. Originally this result was stated for atom distribution with Gaussian decay, but the proof can be modified for random variables with subexponential decay and satisfying (6).

Eq. (9) is proven in Proposition 3.1 of . The proof is based on an explicit integral representation for the marginal (pn′)(k)(p^{\prime}_{n})^{(k)}. The only assumption used to prove (9) in is the asymptotic validity of the local semicircle law for the density of the eigenvalues of the Wigner matrix HH (with no Gaussian part) on scales of order n−1+δn^{-1+\delta}. More precisely, if

denote the Stieltjes transforms of the empirical eigenvalue distribution and of the semicircle law, respectively, then we need to know [3, Eq. (3.9)] that for any κ,δ>0\kappa,\delta>0 there exist constants c,C>0c,C>0 such that

for n−1+δ≤η≤1n^{-1+\delta}\leq\eta\leq 1 and for all ε>0{\varepsilon}>0 sufficiently small.

Eq. (10) has been established in under the assumption of Gaussian decay of the atom distribution, that is assuming that Eeβx2<∞{\mathbf{E}}e^{\beta x^{2}}<\infty, for some β>0\beta>0. To reduce this assumption to the desired sub-exponential decay assumption, we can use a result in (see the paragraph following [7, Lemma 60]) which proved (10) under this weaker assumption with a weaker probability bound exp⁡(−ω(log⁡n))\exp(-\omega(\log n)), where ω(log⁡n)\omega(\log n) denotes a quantity whose ratio with log⁡n\log n goes to infinity as n→∞n\to\infty. One can also follow the ideas in Section 5 of , where it is shown that the Gaussian decay can be relaxed to exponential decay (at the expense of a slightly weaker probability bound). These ideas can be further refined to deal with sub-exponential decay.

Following the proof of Proposition 3.1 in , it is clear that the argument holds using this weaker bound on the probability as well. This concludes the proof of Proposition 4. ∎

To conclude the proof of Theorem 2, it thus suffices to show that the expression (1) only changes by o(1)o(1) when we replace HH with H′H^{\prime}, where we use o(1)o(1) to denote an expression that goes to zero as n→∞n\to\infty.

From the definition of pn(k)p_{n}^{(k)}, we can rewrite (1) as

On the other hand, by [7, Proposition 61] (or [2, Theorem 5.1]), taking advantage of the hypothesis (6), we see that, with probability 1−O(n−100)1-O(n^{-100}) (say),

for all intervals I⊂[u−ε,u+ε]I\subset[u-{\varepsilon},u+{\varepsilon}] with ∣I∣≥log⁡Cn/n|I|\geq\log^{C}n/n, for some sufficiently large absolute constant CC, C~\widetilde{C}, where C~\widetilde{C} can depend on uu and ε{\varepsilon}. Because of this, and the compact support of ff, we can restrict the summation in (11) to those i1,…,iki_{1},\ldots,i_{k} for which ij=i1+O(log⁡O(1)n)i_{j}=i_{1}+O(\log^{O(1)}n) for all 1≤j≤k1\leq j\leq k, while only paying an acceptable error of o(1)o(1). In particular, there are now only O(nlog⁡O(1)n)O(n\log^{O(1)}n) summands in (11). Similarly, using the convergence to the Wigner semicircular law (see e.g. ) we may restrict i1,…,iki_{1},\ldots,i_{k} to the interval [δn,(1−δ)n][\delta n,(1-\delta)n] for some fixed δ>0\delta>0 independent of nn.

A routine computation shows that GG enjoys the derivative bounds

for all y1,…,yky_{1},\ldots,y_{k} and 0≤j≤50\leq j\leq 5, where the implied constant in the O()O() notation can depend on u,ε,fu,{\varepsilon},f.

Suppose first that that the random variables x,x′x,x^{\prime} matched to fourth order. One could then apply the four moment theorem in [7, Theorem 15] to conclude for each i1,…,ik∈[δn,(1−δ)n]i_{1},\ldots,i_{k}\in[\delta n,(1-\delta)n], the quantity EG(nλi1,…,nλik){\mathbf{E}}G(n\lambda_{i_{1}},\ldots,n\lambda_{i_{k}}) changes by O(n−c)O(n^{-c}) for some absolute constant c>0c>0, uniformly in i1,…,iki_{1},\ldots,i_{k}, when HH is replaced by H′H^{\prime}. Inserting this into (11), we obtain the desired claim.

In our situation, xx and x′x^{\prime} do not match exactly to fourth order, but only to second order. However, we have approximate matching to third and fourth order thanks to (8), and this is enough to recover the conclusions of the four moment theorem. Indeed, an inspection of the proof of [7, Theorem 15] in [7, Section 3.3] reveals that the matching moments condition is used in only one place, namely to show that the expectation of an expression of the form

The reader will observe that there is some room to spare here regarding powers of nn. Indeed, a more careful examination of the argument reveals that one can sharpen Theorem 2 by not keeping ε{\varepsilon} fixed with respect to nn, but instead letting it decay as ε=n−1/2+δ{\varepsilon}=n^{-1/2+\delta} for any 0<δ<1/20<\delta<1/2. This should not be the optimal result, however; the correct scale should be O(n−1+δ)O(n^{-1+\delta}) (and in fact one should not need to average in uu at all). Unfortunately, our methods so far can only achieve this by putting more moment or regularity hypotheses on the distribution xx, see for instance Section 3.

We now sketch the proof of Theorem 3. The first step is to prove that the empirical gap distribution Λn′\Lambda_{n}^{\prime} of the matrix H′H^{\prime} satisfies the equation (4). This follows from the same argument used in the proof of of Theorem 1.2 in . The only difference is that the parameter ε{\varepsilon} in (4) was replaced by n−1+δn^{-1+\delta} (with some 0<δ<10<\delta<1) in . Since the change of the density ϱsc(u′)\varrho_{sc}(u^{\prime}) vanishes as ε→0{\varepsilon}\to 0, the same proof applies to the current setting, using that the limiting function on the rhs of (4) is continuous in ss.

In the second step we deduce (4) for HH by using that the expected values of observables of the form (11) are asymptotically the same for HH and H′H^{\prime}. By the exclusion-inclusion principle, Λn\Lambda_{n} can be written as an alternating series of expressions in the form (11) with functions GG which are products of characteristic functions of intervals of length one or larger. These characteristic functions can be approximated by bounded smooth functions in L1L^{1}-sense with arbitrary precision κ\kappa. Controlling the alternating series by its last term, letting first n→∞n\to\infty then κ→0\kappa\to 0 and using the continuity in ss of the limiting function, we obtain (4) for HH as well.

Concluding remarks

If one assumes the additional third moment condition Ex3=0{\mathbf{E}}x^{3}=0 on the moment condition, then it was shown in [7, Theorem 29] that the eigenvalues λi\lambda_{i} with i∈[δn,(1−δ)n]i\in[\delta n,(1-\delta)n] obeyed the localization property

with probability 1−O(n−C)1-O(n^{-C}) for any C,c>0C,c>0 (and nn sufficiently large depending on C,c,δC,c,\delta), where t(a)t(a) was defined by the formula

Because of this localization, it is possible to eliminate the averaging over u′u^{\prime} in Theorem 2, and to reduce the averaging in Theorem 3 to an interval of O(nδ)O(n^{\delta}) eigenvalues rather than O(n)O(n), by combining the above the arguments with those used to prove [7, Theorems 9,11]. The net effect of this is to eliminate the hypothesis in those theorems that xx is supported on at least three points. Thus, for instance, we can take xx to be the Bernoulli variable that takes values ±1\pm 1 with an equal probability 1/21/2.

If the third moment condition from [7, Theorem 29] could be omitted, then one could similarly reduce the averaging required for Theorems 2, 3 even when the third moment was non-zero. To do so, however, would require either a relaxation of the moment conditions needed in [7, Theorem 15], or else an analogue of [7, Theorem 29] for the matrices H′H^{\prime} studied here. It seems plausible that at least one of these two approaches could eventually be made to work, but this does not appear to be easily deducible from the existing literature.

Another remark is that our theorems hold for a more general model of random matrices, where the real and imaginary parts of the entries are not necessarily independent. The reason is that [7, Theorem 15] (and its variant used here) do not require this assumption (see also the end of [7, Section 1]).

References