Universality for generalized Wigner matrices with Bernoulli distribution
László Erdos, Horng-Tzer Yau, Jun Yin
Introduction
The universality of local eigenvalue statistics in the bulk of the spectrum of random matrices has been traditionally considered only for invariant ensembles . For non-invariant ensembles, a new approach to prove the bulk universality was developed in . It consists of the following three steps:
Universality for Gaussian divisible ensembles.
Approximation by Gaussian divisible ensembles.
In Step 2, the universality of the local eigenvalue statistics for a large class of matrices, i.e., Gaussian divisible matrices, was established. Thus in order to prove the universality of a given ensemble, it remains to approximate the matrix elements in this ensemble by Gaussian divisible distribution in such a way that the local eigenvalue statistics are unchanged. This approximation is intrinsically a density theorem and it can be achieved by perturbative expansions in several different ways. In the most recent approach , the universality for Gaussian divisible ensembles was proved via the Dyson Brownian motion and the stability of eigenvalues in Step 3 was provided by the Green function comparison theorem. In Step 2 a technical tool, the logarithmic Sobolev inequality (LSI), was needed to estimate the fluctuations of eigenvalue distribution. This restriction could not be completely removed in Step 3 and thus the Bernoulli measures were excluded in . In this paper, we will improve the local semicircle law so that the LSI is no longer needed. This will enable us to prove the universality for generalized Wigner matrices with Bernoulli distributions. As a byproduct of the new stronger form of local semicircle law, we also obtain much stronger estimates on the eigenvalue density and on the matrix elements of the resolvent.
Recall the Stieltjes transform of the empirical measure of the eigenvalues is defined by
We have proved in that the difference between and , the Stieltjes transform of the semicircle law (2.9), is bounded by where . The main result of this paper states that the error can be improved to . The improvement of a factor resembles the usual factor in the central limit theorem and it results from a new estimate on the correlations of error terms. This estimate also implies that the error between the normalized empirical counting function of the eigenvalues and the one given by the semicircle law is less than in the bulk of the spectrum for any . This new input is sufficiently strong to replace the usage of the (LSI) in , see the discussion after Theorem 2.2 for more details.
Notice that this improvement of a factor and the removal of the LSI need a substantial amount of work. Our motivations to take on this endeavor are for the following two reasons: (1) The distributions of the Bernoulli random matrices are very singular while the Gaussian measures in GOE are very smooth. It is not a priori clear that the universality holds for such singular distributions. (2) The adjacency matrices for random graphs are natural examples of symmetric random matrices. The matrix elements of these matrices take the values or and thus they form Bernoulli random matrices. Our current results do not cover this case since we require the mean zero condition, but they represent the first step toward the universality of the adjacency matrices of random graphs.
Main results
Matrices with independent, zero mean entries and with the normalization condition (2.1) will be called universal Wigner matrices. The basic parameter of such matrices is the quantity
Note that corresponds to the standard Wigner matrices and the conditions define more general Wigner matrices with comparable variances.
We will also consider an even more general case when for different indices are not comparable. A special case is the band matrix, where for with some parameter .
Denote by the matrix of variances which is symmetric, doubly stochastic by (2.1), and in particular satisfies . Let the spectrum of be supported in
with some nonnegative constants . We will always have the following spectral assumption
The local semicircle law will be proven under this general condition, but the precision of the estimate near the spectral edge will also depend on in an explicit way. For the orientation of the reader, we mention two special cases that provided the main motivation for our work.
One important class of universal Wigner matrices is the generalized Wigner ensemble which is defined by the extra condition that
It is easy to check that (2.4) holds with
Another example is the band matrix ensemble whose variances are given by
Define the Stieltjes transform of the empirical eigenvalue distribution of by
Define as the unique solution of
with positive imaginary part for all with , i.e.,
Here the square root function is chosen with a branch cut in the segment $\sqrt{z^{2}-4}\sim zm_{sc}\text{Im }z>0$ and it is the Wigner semicircle distribution
The Wigner semicircle law states that for any fixed , i.e., provided that is independent of . We have proved a local version of this result for universal Wigner matrices and the main result can be stated as the following probability estimate:
with some constant . The accuracy of this estimate can be improved from to , which is the content of the next theorem. It summarizes the results of Theorems 4.1 and 5.1. Prior to our result in , a central limit theorem for the semicircle law on macroscopic scale for band matrices was established by Guionnet and Anderson and Zeitouni ; a semicircle law for Gaussian band matrices was proved by Disertori, Pinson and Spencer . For a review on band matrices, see the recent article by Spencer.
We consider universal Wigner matrices and its special class, the generalized Wigner matrices in parallel. The parameter will distinguish between the two cases; we set for universal Wigner matrices, and for generalized Wigner matrices, where the results will be stronger.
where \kappa:=\big{|}\,|E|-2\big{|}. Then there exist constants , , and , depending only on , and in (2.5), such that for any and the Stieltjes transform of the empirical eigenvalue distribution of satisfies
for sufficiently large . Furthermore, the diagonal matrix elements of the Green function satisfy that
and for the off-diagonal elements we have
The subexponential decay condition (2.11) can also be easily weakened if we are not aiming at error estimates faster than any power law of . This can be easily carried out and we will not pursue it in this paper.
Denote the eigenvalues of by and let be their (symmetric) probability density. For any the -point correlation function of the eigenvalues is defined by
We now state our main result concerning these correlation functions. The same result was proved in under the additional assumption (2.26).
where is the -point correlation function for the GUE ensemble. The same statement holds for symmetric matrices, with GOE replacing the GUE ensemble.
Remark. We can take for some small constant so that there is no double limit taken. This is because all our bounds have an effective error estimate . In case of hermitian matrices there is no need for averaging in the energy parameter . The limit (2.17) holds even for any fixed energy , with , since, instead of relying on the local relaxation flow of , we can use the result of for Gaussian divisible ensembles at a fixed energy.
It is well-known that the limiting correlation functions of the GUE ensemble are given by the sine kernel
and a similar universal formula is available for the limiting gap distribution. The formulas for the GOE cases are more complicated and we refer the reader to standard references such as .
We will prove Theorem 2.2 using the approach of . The logarithmic Sobolev inequality was an important tool in these papers and it was the main obstacle why the case of Bernoulli random matrices were not covered. We note that the Bernoulli distribution satisfies the discrete version of the LSI but it would not be sufficient for our purposes. To explain the necessity of LSI, we now review the three basic ingredients of the approach of .
Local semicircle law: It states that the density of eigenvalues is given by the semicircle law down to short scales containing only eigenvalues for all , where is the size of the matrix.
Local ergodicity of the Dyson Brownian motion: The Dyson Brownian motion is given by the flow
be the probability measure of the eigenvalues of the general ensemble, ( for the hermitian case and for the symmetric case). Denote the distribution of the eigenvalues of at time by . Then satisfies
We now recall the following theorem concerning the universality of the Dyson Brownian motion. Following the convention in , we label the assumptions as Assumptions II–IV since the Assumption I, a convexity property of the Hamiltonian for the invariant measure of the Dyson Brownian motions, is automatically satisfied for any ensembles.
where is the density of the semicircle law (2.10).
Let denote the location of the -th point under the semicircle law, i.e., is defined by
We will call the classical location of the -th point.
Assumption III. There exists an such that
where is the exponent from Assumption III and and are arbitrarily small numbers.
We have proved that Assumption IV follows from the local semicircle law and Assumption III also follows from the local semicircle law provided that a uniform LSI for the distributions of the matrix elements is assumed.
Green function comparison theorem: It asserts that the correlation functions of the eigenvalues of two matrix ensembles are identical up to the scale provided that the first four moments of the matrix elements of these two ensembles are almost identical. Given this theorem and the universality for the Dyson Brownian motion for , the universality for a matrix ensemble holds if we can find another matrix ensemble such that the first four moments of the matrix elements of and (given by (2.18)) are almost the same. Furthermore, is required to satisfy a uniform LSI so that the Assumption III can be verified. This is possible if the first four moments of satisfy
where is the -th moment of the matrix element in the symmetric case. In the hermitian case, the moments of the real and imaginary parts have to satisfy (2.26).
Combining these ingredients, the universality of local eigenvalue statistics in the bulk was proved for all generalized Wigner ensembles (see (2.6) for the definition) satisfying (2.26) and a subexponential decay technical condition. The restriction (2.26) was needed to guarantee the existence of a matching matrix ensemble whose matrix element distributions satisfy the LSI so that the Assumption III can be verified. The local semicircle estimates in Theorem 2.1 imply that the empirical counting function of the eigenvalues is close to the semicircle counting function (Theorem 6.3) and that the location of the eigenvalues are close to their classical location in mean square deviation sense (Theorem 7.1). This provides a direct proof to the Assumption III (2.24) and thus removes the usage of the LSI.
Finally we summarize the recent results related to the bulk universality of local eigenvalue statistics. The local semicircle law for Step 1 was first established for Wigner matrices in a series of papers . The method was based on a self-consistent equation for the Stieltjes transform of the eigenvalues and the continuity of the imaginary part of the spectral parameter in the Stieltjes transform. As a by-product, an eigenvector delocalization estimate was proved.
The universality for Gaussian divisible ensembles was proved by Johansson for hermitian Wigner ensembles. It was extended to complex sample covariance matrices by Ben Arous and Péché . There were two major restrictions of this method: 1. The Gaussian component was fairly large, it was required to be of order one independent of . 2. It relies on explicit formulas for the correlation functions of eigenvalues which are valid only for Gaussian divisible ensembles with unitary invariant Gaussian component. The size of the Gaussian component was reduced to in by using an improved formula for correlation functions and the local semicircle law from . The Gaussian component was then removed by a perturbation argument using the reverse heat flow. Thus the three step strategy to prove the universality was introduced and it led to the first proof of the bulk universality for hermitian Wigner ensembles. Due to the reverse heat flow argument used in Step 3, the universality class established in was restricted to matrices with smooth distributions for the matrix elements. Shortly after, Tao and Vu proved the four moment theorem which in particular removes the smoothness restriction in Step 3. It thus proved the universality for hermitian Wigner matrices whose matrix element distributions were supported on at least three points. The last condition was removed in by combining the arguments of . The result of also implies that the local statistics of symmetric Wigner matrices and GOE are the same, but under the restriction that the first four moments of the matrix elements match those of GOE. Thus the universality class for the local correlation functions established via the approach of combining and was broader for the hermitian ensembles than for the symmetric ones. This improvement was due to Johansson’s result , which provided the universality for Gaussian divisible ensembles in Step 2, was available only for hermitian ensembles.
A more general and conceptually very appealing approach for Step 2 is via the local ergodicity of Dyson Brownian motion. This approach, initiated in , was applied to prove the universality for symmetric Wigner matrices with the three point support condition. In , we formulated a general theorem for the bulk universality which applies to all classical ensembles, i.e., real and complex Wigner matrices, real and complex sample covariance matrices and quaternion Wigner matrices. Later on, Tao and Vu also extended their results to the sample covariance matrices with the three point support condition for complex covariance matrices and four moment matching conditions for real ones. Shortly after , Péché also extended the approach to the complex sample covariance matrices and proved the universality in the bulk.
Most recently, we introduced the Green function comparison theorem and extended the local semicircle law to include the matrix elements of the Green functions. This allows us to remove the smoothness restriction from the reverse heat flow argument in Step 3 of our approach. We remark that the comparison theorems in concern individual eigenvalues with a fixed index, while the Green function comparison theorem is at a fixed energy. On the other hand, in the variances of the matrix elements were allowed to vary, i.e., the matrices belonged to generalized Wigner ensembles. The three step strategy can thus be applied and the universality was proved for generalized Wigner ensembles with essentially only one class of measures, the Bernoulli measures, excluded due to the LSI used in verifying Assumption III in Step 2. Finally, in the current paper, Assumption III will be shown to be a consequence of a strong local semicircle law, which will be proved for all ensembles with a subexponential decay property. In particular, Bernoulli measures are now included in the universality class (in the sense of (2.17)) for both hermitian and symmetric generalized Wigner ensembles. We have thus removed all restrictions except the subexponential decay in our approach. A clear picture of the three step strategy emerges: Step 2 and 3 hold under very general conditions and are model independent. The main task of proving the universality is to establish a strong version of the local semicircle law—which can be model dependent. We believe that our method applies to generalized sample covariance matrices as well, but we will not pursue this direction in this paper.
Proof of Universality
We now prove the main universality theorem, Theorem 2.2.
Step 1. Universality for Dyson Brownian Motion: Under the Assumptions II–IV in the introduction, the universality for the Dyson Brownian Motion was proved in . We recall the statement in the following Theorem.
Notice that the assumption on the initial entropy is not needed as was remarked in .
Step 2 Universality for Gaussian divisible ensembles: The Dyson Brownian motion is generated by the matrix flow (2.18). Our task is to determine the initial ensemble so that the Assumptions II–IV of Theorem 3.1 can be proved for the flow. The Assumption IV is a direct consequence of the local semicircle law, i.e., Theorem 4.1. The Assumption III will be proved in Proposition 7.1. For the generalized Wigner matrices, the only assumption of Theorem 4.1 and Proposition 7.1 is the subexponential decay property of the distributions of the matrix elements. Since the evolution of the matrix element is given by an Ornstein-Uhlenbeck process, the subexponential property is preserved and we only have to check it for the initial data. We have thus proved the following theorem.
Suppose that the probability law for the initial matrix satisfies the assumptions of Theorem 2.2. Then there exists such that for any , the probability law for the eigenvalues of satisfies the universality equation (2.17).
Step 3 Green function comparison theorem: We have proved the universality for all ensembles with the matrix element at distributed by with
where are independent Gaussian random variables with mean and variance and . In order to prove Theorem 2.2, it remains to approximate all random variables with the subexponential property by . The only requirement of is the subexponential decay property and the mean zero and variance one normalization. Our tool is the following Green function comparison theorem from . It implies that the correlation functions of the eigenvalues of two matrix ensembles at a fixed energy are identical up to the scale provided that the first four moments of the matrix elements of these two ensembles are almost identical. Prior to this theorem, it was proved that the joint distribution of individual eigenvalues for Wigner ensembles is the same under the four moment assumption. Tao-Vu’s theorem addresses the distribution of individual eigenvaluesIn a recent preprint (appeared after the current preprint was first posted), it was pointed out that if the four moment condition is violated, then the differences between individual eigenvalues of the two ensembles are bigger than the eigenvalue spacing. Thus the four moment condition is also necessary for locating the individual eigenvalues. This is in contrast with the main theme of this paper that gap distribution and correlation functions are even independent of the second moments as long as they are nonzero. while Theorem 3.3 compares Green functions (and thus eigenvalues) at a fixed energy.
Suppose that we have two generalized Wigner matrices, and , with matrix elements given by the random variables and , respectively, with and satisfying the uniform subexponential decay condition (2.11). Fix a bijective ordering map on the index set of the independent matrix elements,
and denote by the generalized Wigner matrix whose matrix elements follow the -distribution if and they follow the -distribution otherwise; in particular and . Let be arbitrary and suppose that for any small parameter and for any we have the following estimate on the diagonal elements of the resolvent:
with some constants depending only on . Moreover, we assume that the first three moments of and are the same, i.e.
and the difference between the fourth moments of and is much less than 1, say
for some given . Let be arbitrary and choose an with . For any sequence of positive integers , set complex parameters , , with and with an arbitrary choice of the signs. Let be the resolvent and let be a function such that for any multi-index with and for any sufficiently small, we have
Then, there is a constant , depending on , and such that for any with and for any choices of the signs in the imaginary part of
where in the second term the arguments of are changed from the Green functions of to and all other parameters remain unchanged.
Given this theorem, for any matrix ensemble whose matrix element at are distributed according to , we need to find such that the first four moments of and are almost the same and has a subexponential decay. Since the real and imaginary parts are i.i.d., it is sufficient to match them individually. This is the content of the following lemma which is stated for real random variables normalized to variance one. With this lemma, we have proved Theorem 2.2. This lemma is essentially the same as Lemma 28 in .
Let and be two real numbers such that
for some positive constant . Let be a real Gaussian random variable with mean and variance . Then for any sufficient small (depending on ), there exists a real random variable with subexponential decay and independent of , such that the first four moments of
are , , and , and
for some positive constant depending on .
Proof. It is easy to see by an explicit construction that the following holds:
For any real random variable , independent of , and with the first 4 moments being , , and , the first 4 moments of
Using (3), we obtain that for any there exists a real random variable such that the first four moments are , ,
With , we have , thus
for some positive constant depending on . Hence with (3.10) and (3.11), we obtain that satisfies and (3.8). This completes the proof of Lemma 3.4.
Large Deviation of Local Semicircle Law
We first reprove the large deviation of local semicircle law given in . The result of this section is relevant only for .
Let . Then for all with
for sufficiently large N with positive some constants and that depend only and in (2.11) and in (2.4) and (2.5).
The theorem will be proved at the end of the section after collecting several lemmas. The first lemma describes the behavior of in the various regimes, its proof is elementary calculus. We use the notation for two positive functions in some domain if there is a positive universal constant such that holds for all .
We have for all with that
From now on, let with and . If , then we have
For the behavior of and we distinguish two cases.
Thus the control function has the following behavior
Note that the precise formula (4.1) for is not important, only its asymptotic behavior for small , and is relevant. The theorem remains valid if is replaced by with . In particular, can be chosen to be order one when is not near the edges of the spectrum. If we are only concerned with the generalized Wigner ensemble , then by (2.7) we can choose for any (). For universal Wigner matrices we have for , i.e., using the parameter introduced in Theorem 2.1, we have
Based upon these formulas, we also have, for any with ,
First, we introduce some notations. Recall that denotes the matrix element
These quantities depend on , but we mostly neglect this dependence in the notation.
The following two results were proved in our previous work (Lemma 4.2 and Corollary B.3 of ) and they will be our key inputs. We start with the self-consistent perturbation formulas.
for some constants depending on and in (2.11).
We start with determining a system of self-consistent equations for the diagonal matrix elements of the resolvent. We can write as follows,
We will estimate the following key quantities
Both quantities and will be typically small, eventually we will prove that their size is less than , modulo logarithmic corrections and a factor involving the distance to the edge. We thus define the exceptional event
We will always work in , and, in particular, we will have
since by (4.12). Define the set
for any with some universal constant . Here we estimated \big{|}|G_{ii}|-|m_{sc}|\big{|}\leq\Lambda_{d}, and we used from (4.6)–(4.7) that satisfies for .
Thus, a special case of (4.16) or (4.15),
together with (4.25) implies that for any and with a sufficiently large constant
with being the constant in (4.25) and we also used that . Similarly, with one more expansion step, we get
Using these estimates, the following lemma shows that and are small assuming is small and the ’s are not too large. The control parameter for the ’s is , defined below (4.32). These bounds hold uniformly in .
to be the set of all exceptional events. Then we have
Proof: Under the assumption of (2.11), we have
therefore we can work on the complement set . Define the event
Notice that the estimates (4.26)–(4.31) also hold on , maybe with different constants . We now prove that for any fixed , we have
To see (4.36), we apply the estimate (4.18) from the large deviation Lemma 4.4, and we obtain that
holds with a probability larger than for sufficiently large .
Denote by and () the eigenvectors and eigenvalues of . Let denote the -th coordinate of . Then, using and (4.28), we have
Here we defined for any matrix and we used (4.12) to estimate . Together with (4.38) we have proved (4.36) for a fixed .
For the offdiagonal estimate (4.37), for , we have from (4.19) that
holds with a probability larger than for sufficiently large . Similarly to (4), by using (4.31), we get
Now we start proving (4.34). First we choose an -net in the set , i.e., a collection of points, , such that for any there is such that . The net can be chosen such that . Then (4.36) and (4.37) imply that
Now let be arbitrary and choose such that . For any fixed , we have
By and , we have
In the last inequality, we used the assumption . Thus
Since for , we obtain
Moreover, by estimating in , we see that , , and are Lipschitz continuous functions in with a Lipschitz constant bounded by . Therefore can be replaced with in the lower bound on and obtained from (4.41), and, furthermore, using a trivial upper bound . Thus we get
Combining this with (4.35), we obtain (4.34) and thus Lemma 4.5.
Our goal is to show that is smaller than (modulo edge and logarithmic corrections) for any in the event . We will use a continuity argument. In Lemma 4.6 we show for any that if is smaller than , then it is actually also smaller than . In Lemma 4.9 we show that this input condition holds at least for . Then reducing , we show by a continuity argument that it holds for each .
Let and satisfy (4.2), in particular . Recall , and defined in (4.23) and (4.33). Then we have that, in the event , if
and we also have a stronger bound for the off-diagonal terms:
Proof of Lemma 4.6. First note that condition (4.43) is equivalent assuming the event and we have
so the event holds. We recall from (4.12) that
With the assumption (4.43) we have (see (4.25), (4.27))
We first estimate the offdiagonal term . From (4.14) we have
where we used (4.50) to show that the first term can be absorbed into the second. From the second inequality in (4.50) we also have
This proves the estimate (4.45). Using (4.47), we also see that (4.44) holds for the summand .
Now we estimate the diagonal terms. Recalling from (4.21), with (4.29), (4.50), (4.52) we have,
Again, the first term can be absorbed into the second, so we have proved
Using , and the fact that , so with (see in (4.49) and (4.54)), we can expand (4.55) as
Summing up this formula for all and recalling the definition yield
Introducing the notations , for simplicity, we have (using )
and the error terms for each sums up to zero. Therefore, with , we have
To estimate the norm of the resolvent, we recall the following elementary lemma (Lemma 5.3 in ).
Let be a given constant. Then there exist small real numbers and , depending only on , such that for any positive number , we have
Suppose that satisfies (2.4), i.e., . Then we have
with some constant depending on and with defined in (4.61)
with given in (4.60). By (4.60), we have
Choosing with a large , we have proved the Lemma.
We now return to the proof of Lemma 4.6, recall that we are in the set . First, inserting (4.5) and (4.62) into (4.59), and using , we obtain
By the assumption (4.43), we have , for large enough , therefore we get
Using the bound on in (4.54) and (4.50), we obtain
which, together with (4.52), completes the proof of (4.44).
recall the definitions of , and from (4.33) and define
Furthermore, in the set we have
Denote by and the eigenvectors and eigenvalues of . On the set all eigenvalues are bounded, . In this set we have, with ,
with some positive constant . We also have the upper bound and . In particular, for , we have
Similarly, the argument (4.51)–(4.52) shows that in the set , we have
and the argument (4.53)–(4.54) guarantees that
in . Finally, to control , we use that from the self consistent equation (4.55) and the definition of , we have
For , with (2.9), we have . Using and , we obtain
Using (4.69), together with and (4.71), we obtain that the absolute value of the r.h.s of (4.70) is less than
Taking the absolute value of (4.70) and maximizing over , we have
Since the denominator satisfies ,
follows from the last equation. Combining it with (4.68) and (4.69), we obtain (4.65), and this completes the proof of Lemma 4.9.
Proof of Theorem 4.1. Lemma 4.6 states that, in the event , if then with
By assumption (4.2) of Theorem 4.1, we have for any and these functions are continuous. Lemma 4.9 states that in the set the bound holds for .
Thus by a continuity argument, in the set as long as the condition (4.2) is satisfied. Finally, once is proven, we can use (in the domain ) and Lemma 4.6 once more to conclude the stronger bound on . This proves Theorem 4.1.
We record that combining the bound on with (4.54), we also proved that under the assumption (4.2) we have
Local semicircle law
In this section we strengthen the estimate of Theorem 4.1 for the Stieltjes transform . The key improvement is that will be estimated with a precision while the was controlled by a precision only (modulo logarithmic terms and terms expressing the deterioration of the estimate near the edge).
Assume the conditions of Theorem 4.1 and recall the notations \kappa=\kappa_{E}:=\big{|}|E|-2\big{|} and from (4.1). Define the domain
Then for any and there exists a constant such that
Proof of Theorem 5.1. We will work in the set , which has almost full probability by (4.34) and (4.64). Note that the set is included in the domain defined by (4.2), therefore we can use the estimates from Section 4.
As in (4.57), where , we have that
holds with a very high probability. Recall that and we mostly omit the argument from the notations. The quantities , and were defined in (4.23), (4.21) and (4.53). Then with (4.75) we have
holds with a very high probability for any small . Recall that . We have, from (4.20), (4.25) and ,
where we used (4.75) to bound and (4.47) to control the term.
holds with a very high probability. Since ’s are independent, applying the first estimate in the large deviation Lemma 4.4, we have
On the complement event, the estimate can be included in the last error term in (5.3). It only remains to bound
whose moment is bounded in the next lemma which will be proved in Sections 8 and 9.
For fixed in domain (5.1) and any even number , we have
Using Lemma 5.2, we have that for any and ,
for sufficiently large . Combining this with (5.4) and (5.3) and noting that , see (4.7), we obtain (5.2) and complete the proof of Theorem 5.1.
Empirical counting function
In this section we translate the information on the Stieltjes transform obtained in Theorem 5.1 to an asymptotic on the empirical counting function. The main ingredient for the first step is the following lemma based upon the Helffer-Sjöstrand formula. We will formulate this lemma for general signed measures, but we will apply it to the Stieltjes transform of the difference between the empirical density and the semicircle law. A similar statement was already proven in Lemma B.1 in and Lemma 7.7 in .
and in case of we additionally assume . Then
with some constant depending on and .
where is a smooth cutoff function with support in $\chi(y)=1|y|\leq 1/2$ and with bounded derivatives. The first term is estimated by, with (6.1),
For the second term in r.h.s of (6) we use that from (6.1) it follows for any that
With and
As in (B.17) and (B.19) in , we integrate the third term in (6) by parts first in , then in . Then we bound it with an absolute value by
The second term is bounded in (6.4). By using (6.1) and (6.6) in the first term and (6.1) in the third, we have
Let be the ordered eigenvalues of a universal Wigner matrix. We define the normalized empirical counting function by
be the distribution function of the semicircle law which is very close to the counting function of ’s, .
We will need some control on the spectral edge, we recall the Lemma 7.2 from .
(1) Let the universal Wigner matrix satisfy (2.1), (2.2) and (2.11) with . Then we have
(2) Let be a generalized Wigner matrix with subexponential decay, i.e., (2.1), (2.2), (2.6) and (2.11) hold. Then
for any small with an depending on . Furthermore, for ,
With these preliminary lemmas, we have the following theorem that we state for universal Wigner matrices and for their subclass, the generalized Wigner matrices in parallel.
Let for universal Wigner matrices and for generalized Wigner matrices. Suppose that the universal Wigner matrix ensemble satisfies (2.1), (2.2) and (2.11) with and the generalized Wigner matrix ensemble satisfies (2.1), (2.2), (2.6) and (2.11). We recall in the latter case. Then for any and there exists a constant such that
where the and were defined in (6.8) and (6.10) and \kappa_{E}=\big{|}|E|-2\big{|}.
Proof. For definiteness, we will consider the case of generalized Wigner matrices, i.e., . In this case , (see (2.7)) and thus for , see (4.10). For simplicity of the presentation, we assume that as overall constant factors do not matter (see the remark after (4.10)). We set , and apply Lemma 6.1 to the difference . Let , where is the normalized empirical counting measure of eigenvalues. First we check the conditions of Lemma 6.1. To check that (6.1) holds, set and for a fixed , let satisfy , so that . Clearly (6.1) holds for any with a very high probability by (5.2). In particular, we know that
Consider , set , and estimate
Now we use the fact that the functions and are monotone increasing for any since both are Stieltjes transforms of a positive measure. Therefore the integral in (6.15) can be bounded by
By the choice of and using that , we have
with a possible larger in the r.h.s. Thus (6.1) holds for the difference .
The application of Lemma 6.1 shows that for
Recall that the characteristic function of the interval , smoothed on scale at the edges. The additional in the denominator in the r.h.s. of (6.20) comes from the case when , are very small and the trivial estimate with gives a better bound than Lemma 6.1.
With the fact is monotone increasing for any , (6.19) implies a crude upper bound on the empirical density. Indeed, for any interval , with , we have
Choose arbitrary , then we have
from (6.21). Since is bounded, we also have
Subtracting (6.22) and (6.23) and using (6.20), we obtain that for any
with a very high probability, i.e., apart from a set of probability smaller than for any . The estimate (6.13) from Lemma 6.2 on the extreme eigenvalues shows that is supported in ${\mathfrak{n}}(-3)=n_{sc}(-3)=0{\mathfrak{n}}(3)=n_{sc}(3)=1$. Thus we obtain that
holds for any fixed with an overwhelming probability.
We now choose a fine grid of equidistant points with , then (6.24) holds simultaneously for every with an overwhelming probability. For any we can find an with and by (6.21) we obtain
This guarantees that (6.24) holds simultaneously for all . Since was arbitrary, this proves Theorem 6.3 for generalized Wigner matrices.
The proof for universal Wigner matrices is very similar, just replaces in the estimates, and instead of one uses which follows from (4.10). The main technical estimate (6.20) is modified to
Location of eigenvalues
In this section we estimate the mean square deviation of the eigenvalues from their classical location. The main input is Theorem 6.3, the estimate on the counting function. For simplicity, we consider only the case of generalized Wigner matrices. Similar, but weaker results can be obtained along the same lines for universal Wigner matrices.
Let be a generalized Wigner matrix with subexponential decay, i.e., assume that (2.1), (2.2), (2.6) and (2.11) hold. Let denote the eigenvalues of and be their classical location, defined by (2.23). Then for any and for any there exists a constant , depending on and , such that
Proof. The proof of (7.2) directly follows from (7.1) by using the estimates on the extreme eigenvalue (6.13) from Lemma 6.2. For the proof of (7.1), we can assume that since the complement event has a negligible probability by (6.12) and (6.13) of Lemma 6.2. From Theorem 6.3 we can also assume that
From the definition of it follows that for , i.e., ,
with some positive constants .
Choose . Consider first those -indices for which with a sufficiently large constant. We choose so that (7.4) would imply . We then claim that
We will show that , the upper bound is analogous. Suppose that were smaller than , then . On the other hand, and thus
with some positive constant . Therefore
where the second inequality follows from (7.3), but this contradicts to the choice .
Let satisfy ; the indices can be treated analogously. Note that by (7.3). Define to be index of the -point right below , i.e.,
By (7.5) we see that and from (7.3) and (7.4) it follows that
By the choice of we have , i.e., (7.6) implies . Using now (7.4), we have
using and hence and are comparable. In the last step we also used (7.4). Combining this with (7.7), we have
and the same estimate holds for and thus
by the choice of and similar estimate holds for the sum over the indices as well.
Now we consider the indices and . By a similar argument that proved (7.5), we can see that there is a constant such that , otherwise , but , which would contradict (7.3). It is easy to see that for all , therefore in this regime we estimate and thus
The indices and can be treated similarly.
Finally we deal with the extreme eigenvalues with index and we can assume that . For these indices and we can estimate
For any with , we have , thus we obtain from (7.3) that
The other extreme eigenvalues, , are treated analogously.
Combining (7.8), (7.9) and (7) and choosing sufficiently small in the definition of , we proved (7.1) with any .
Moment Estimates of Error Terms
In this section we prove the second and fourth moment estimates of Lemma 5.2; the general cases will be proved in Section 9.
Recall the definition of , which we rewrite as
We first prove a bound on the Green function .
and for some constant , independent of ,
Proof Consider first the case . Let denote the event inside the probability in the equation (4.3). The proof of Theorem 4.1 yields that . It is clear that (8.4) holds in the event and this proves (8.4) in in the case . Similarly, in the case of , we can prove (8.3) using the event in the equation (4.4). By definition of the domain , the right side of (8.4) is and this proves (8.5) in the case .
For the case and , using (4.15) and (4.16), we obtain that
Since in , (8.4) and (8.3) in the case follows from (8.6) and the case . Repeating this process, we prove (8.4) and (8.3) for any by induction on .
Now we return to the second and fourth moment estimates of Lemma 5.2.
Now we prove the special case of Lemma 5.2 for . The second moment of is given by
We start with estimating the first term of (8.7) for and . The basic idea is to rewrite as
with independent of , and independent of . The ’s have two upper indices. The first one refers to the fact that it comes from the minor (i.e. follows the upper index of ) and the second one indicates the additional independence.
To construct this decomposition for , by (4.15) or (4.16) we can rewrite as
The first term on the r.h.s is independent of . With Lemma 8.1, we have that the bound
Next we define for ().
Hence (8.8) holds and is independent of .
With this convention, we have the following expansion of
Since in , this lemma also implies that
Proof. First we rewrite as follows
By the large deviation estimate (4.19), we have
Similarly, from (4.17), using as , and keeping fixed, we have
By (4.35), holds with a very high probability. We can thus replace by in (8.19). The third term in (8.17) can be estimated in the same way, and the last term can be bounded by with very high probability.
Since , by the definition of in (5.6) we have
This inequality implies the desired inequality (8.14) except for the contribution from the exceptional set where (8.21) fails. Since all Green functions are bounded by , the contribution from the exceptional set is negligible and this proves (8.14). Finally, a similar proof yields (8.15).
Exchange the index and , we can define and and expand as
Here is independent of and ; is independent of . Combining (8.22) with (8.13), we have
The only non-vanishing term on the right-hand side is
By the Cauchy-Schwarz inequality and Lemma 8.2, we obtain
Similarly, Lemma 8.2 and (8.20) imply that
Since the indices and can be replaced by , together with (8.7) we have thus proved Lemma 5.2 for .
2 Proof of Lemma 5.2 for p=4𝑝4p=4
Now we prove the special case of Lemma 5.2 for :
Here means the permutation of the ordered indices and the complex conjugate operators. We are going to compute the first two terms in the r.h.s of (8.2). The other two terms can be treated analogously. By the permutation symmetry of the indices, we can assume that , , and . As in the estimate for the second moment, the key idea is to decompose in a suitable way:
There exist two decompositions of
Furthermore, the decompositions can be chosen in such a way that for all the following estimates hold:
We postpone the proof of this lemma and first finish the proof of Lemma 5.2 in the case of . It is clear that Lemma 8.3 holds for different index combinations. E.g. can be decomposed as
and ’s have the same properties (except for the exchange of and ) as in (8.29) and (8.31) . By this property, we can estimate the first term on the r.h.s. of (8.2) by
Using (8.31) and Schwarz inequality, we have thus proved that
We now estimate the second term in r.h.s of (8.2).
Using (8.30), (8.20) and a Schwarz inequality, we have
For the other terms in (8.2), we can just use Schwarz inequality and (8.16). We have thus proved the Lemma 5.2 for .
We now prove Lemma 8.3. First we prove the properties of ’s. Notice that the decomposition with ’s in (8.27) removes the dependence on rows . The starting point is an expansion of
In order to prove (8.30), we give another representation of the ’s. We begin by removing the dependence of the matrix element of the Green function on the index for . By (4.15) or (4.16), we can rewrite the first term of r.h.s of (8.9) as
This removes the dependence of on the index with the last term as the error term. For the last term on r.h.s of (8.9), using (4.15) and (4.16) again, we have
This removes the dependence on the index of both the Green functions and their inverse in the last term in (8.9). Inserting (8.37)–(8.39) into (8.9), we obtain that if
For and , , and .
For and , and .
For and , , and .
For and , and .
Since the probability of the exceptional set is extremely small, a simple argument which we repeated many times shows that it can be neglected in the estimate of the expectation in (8.30). Hence (8.30) follows from (8.45).
Using the same method we used for ’s, one can prove the properties of ’s in Lemma 8.3. The details will be omitted since we will prove the general cases in the next section.
General case
then (9.3) follows from (8.14), (8.15) and (8.30).
To achieve the decomposition (9.1), as in (8.36) in Section 8.2, we start with a decomposition on .
From this definition one can easily check that
With these definitions, we can decompose as follows.
Proof. Using the definition (9.5), we have
for sufficiently large depending only on .
For the precise argument, we start with the cases:
where for some . Similarly, we define the set of diagonal matrix elements:
With these notations, the equation (8.2) asserts that for , there is a function such that
Proof of Lemma 9.4: By symmetry, we only need to prove the cases that
with a finite sum of elements of the form
where is the minor of with the -th row and -th column removed.
We can remove the dependence on the -th row by the procedure in (8.37)-(8.39). Using (4.15), (4.16) and the notation:
Inserting (9.25) and (9.26) into (9.21) and expanding it, we obtain that (9.21) is equal to
By (9.22) and (9.23), we can set which is in and we have thus proved Lemma 9.4 by induction.
Then (9.12) in this special case follows from Lemma 8.1.
where depends on . We have thus proved (9.12) for the Case 3 and this completes the proof of Lemma 9.3.
Proof of Lemma 5.2. We first introduce the following notations which will be useful for the expansion of the -th moment of in (5.5).
Let be a dimensional vector such that or for .
Let be a dimensional vector such that for .
where is the complex conjugate operator.
where we sum up and under the conditions in Definition 9.3. Lemma 5.2 is now a simple consequence of the following estimate on .
for sufficiently large depending only on .
Using (8.20), i.e., , we have
By definition, . Similarly, we define to be the number of times that appears in , i.e.,
Without loss of generality, we assume that . Then using (9.44), we know
Therefore, under the assumption (9.41) we have
Combining this identity with (9.40), we obtain (9.36) and thus conclude Lemma 9.5.
Acknowledgement. The authors thank Terry Tao for his helpful comments on an earlier version of this manuscript.