Anisotropic local laws for random matrices
Antti Knowles, Jun Yin
Introduction
for large and with high probability. Here is the Stieltjes transform of the asymptotic eigenvalue density, which we denote by . We call an estimate of the form (1.1) an averaged law.
As may be easily seen by taking the imaginary part of (1.1), control of the convergence of yields control of an order eigenvalues around the point . A local law is an estimate of the form (1.1) for all . Note that the approximation (1.1) cannot be correct at or below the scale , at which the behaviour of the left-hand side of (1.1) is governed by the fluctuations of individual eigenvalues. Such local laws have become a cornerstone of random matrix theory, starting from the work ESY2 where a local law was first established for Wigner matrices. Well known corollaries of a local law include bounds on the eigenvalue counting function as well as eigenvalue rigidity. Moreover, local laws constitute the main tool needed to analyse (a) the distribution of eigenvalues (including the universality of the local spectral statistics), (b) eigenvector delocalization, (c) the distribution of eigenvectors, and (d) finite-rank deformations of .
In fact, for all of the applications (a)–(d), the averaged local law from (1.1) is not sufficient. One has to control not only the normalized trace of but the matrix itself, by showing that is close to some deterministic matrix depending on , provided that . Such control was first obtained for Wigner matrices in EYY1 , where the closeness was established in the sense of individual matrix entries: . We call such an estimate an entrywise local law. More generally, in KY2 ; BEKYY this closeness was established in the sense of generalized matrix entries:
Analogous results for uncorrelated sample covariance matrices were obtained in PY ; BEKYY . The estimate (1.2) states that for large the resolvent is approximately isotropic (i.e. proportional to the identity matrix), and we accordingly call an estimate of the form (1.2) an isotropic local law. We remark that the basis-independent control in (1.2) is crucial for many applications, including the distribution of eigenvectors and the study of finite-rank deformations of .
Unlike in the case of Wigner matrices and uncorrelated sample covariance matrices mentioned above, the resolvent is in general not close to a multiple of the identity matrix, but rather to some general deterministic matrix . In that case (1.2) is to be replaced with
We call an estimate of the form (1.3) an anisotropic local law. The goal of this paper is to develop a method yielding (anisotropic) local laws for many matrix models built from sums and products of deterministic or independent random matrices. Applications include all the four (a)–(d) listed above, some of which we illustrate in this paper.
2. Sample covariance matrices
3. Outline of results
For simplicity, we focus on the matrix , bearing in mind that similar results also apply for (see Section 11.2). We assume that the three matrix dimensions are comparable, and that the entries of possess a sufficient number of bounded moments. Moreover, we assume that is bounded, and that the spectrum of satisfies certain regularity conditions, given Definition 2.7 below, which essentially state that the connected components of the support of are separated by some positive constant, and that the density of has square root decay near its edges in . Note that we do not assume that is square, and in particular may have many vanishing singular values; this allows us to cover e.g. the general linear model of multivariate statistics.
Our main result is the anisotropic local law for . Roughly, it states that (1.3) holds with
where is the Stieltjes transform of the asymptotic density . In fact, we prove a more general anisotropic local law that is more useful in applications. Its formulation is most transparent under the additional assumption that , although this assumption is a mere convenience and may be easily removed (see Section 11.1). We prove an anisotropic local law of the form
A simple application of Schur’s complement formula to (1.5) yields the anisotropic local law for the resolvent of and a similar result for the resolvent of the companion matrix . The estimate (1.5) holds with high probability, and we give an explicit and optimal error bound. We remark that the anisotropic local law holds under very general assumptions on the distribution of , the dimensions of and , and the spectrum of . In particular, we make no assumptions on the singular vectors of . We remark that, previously, an anisotropic global law, valid for , was derived in HLNP for a different matrix model.
As an application of the anisotropic local law, we prove the edge universality of the eigenvalues near the soft spectral edges, whereby the joint distribution of the eigenvalues is asymptotically governed by the Tracy-Widom-Airy statistics of random matrix theory. More precisely, we prove that the asymptotic distribution of the eigenvalues near the soft edges depends only on the nonzero spectrum of . This may be regarded as a universality result in both the distribution of the entries of and the (left and right) singular vectors of , including their dimensions. We then conclude that the Tracy-Widom-Airy statistics hold near the soft edges by noting that they have been previously established HHN ; ElK2 ; Ona ; SL for Gaussian and diagonal .
We also prove the rigidity of eigenvalues, as well as the complete delocalization of the eigenvectors with respect to an arbitrary deterministic basis. Further applications of the anisotropic local law, such as the distribution of the eigenvectors and an analysis of the outliers and BBP-type phase transitions in finite-rank deformations, will appear elsewhere.
Finally, we also apply our method to deformed Wigner matrices of the form , where is a Wigner matrix and a bounded Hermitian matrix. This model describes Wigner matrices whose entries may have arbitrary expectations. As for , we establish the Tracy-Widom-Airy statistics near the spectral edges of . More precisely, we prove that the asymptotic distribution of the eigenvalues near the edges depends only on the asymptotic spectrum of , which may be regarded as a universality result in the distribution of and the eigenvectors of . We then conclude that the Tracy-Widom-Airy statistics hold near the edges by noting that they have been previously established LSY for diagonal .
4. Overview of the proof
We conclude this section by outlining some ideas of the proof of the anisotropic local law. Roughly, the proof proceeds in three steps: (A) the entrywise local law for Gaussian and diagonal , (B) the anisotropic local law for Gaussian and general , and (C) the anisotropic local law for general and general . Steps (A) and (B) may be performed by adapting the methods of BEKYY , and we do not comment on them any further.
The main argument, and the bulk of the proof, is Step (C). Its core is a self-consistent comparison method, which yields the anisotropic local law for general assuming it has been proved for Gaussian . Up to now, all interpolation or Lindeberg replacement arguments in random matrix theory have crucially relied on a local law as input. Since the local law is exactly what we are trying to prove, a different approach is clearly needed – one that does not need a local law to work. Our method yields a new way of deriving local laws for general random matrices expressed as polynomials of deterministic and random matrices whose entries are independent. It relies on two key novel ideas: (a) continous Green function comparison argument where the errors are controlled self-consistently, and (b) a bootstrapping of the Green function comparison on the spectral scale .
In the remainder of this subsection we give a few details of our method, and in particular explain the ideas (a) and (b) in more detail. We construct a continuous family of matrices, whereby is a Gaussian ensemble and is the ensemble in which we are interested. Then we operate in the -plane and perform simultaneously a continuous interpolation in and a discrete bootstrapping in . (See Figure 1.1 below.)
Interpolation methods have been extensively used in random matrix theory, in the form of both discrete Lindeberg-type replacement schemes Chat ; TV1 ; EYY1 ; EYY3 and continuous interpolations SL ; LSY ; LP . Moreover, Dyson Brownian motion, as used e.g. in SL ; LSY ; ESY4 ; ESYY , may be regarded as a form of continuous interpolation for the case that is Gaussian. The usual choice of interpolation, as used in SL ; LP and also resulting from Dyson Brownian motion, is . In this paper, we instead interpolate using i.i.d. Bernoulli random variables:
where is defined through the formal power series
Next, we explain the basic strategy behind the self-consistent comparison method mentioned in (a) above. We choose a function
The next section is devoted to the definition of the model and the statement of results. We give an outline of the structure of the paper in Section 3.5 below.
Conventions
The fundamental large parameter is . All quantities that are not explicitly constant may depend on ; we almost always omit the argument from our notation.
We use in various assumptions to denote a positive constant that may be chosen arbitrarily small. A smaller value of corresponds to a weaker assumption. All constants may depend on , and we neither indicate nor track this dependence.
Model
In this section we define our model, list our key assumptions, and explain the basic structure of the asymptotic eigenvalue density .
We consider the matrix , where is a deterministic matrix and a random matrix. We regard as the fundamental parameter and and as depending on . Here, and throughout the following, we omit the index from our notation, bearing in mind that all quantities that are not explicitly constant (such as the constant ) may depend on . For simplicity, we always make the assumption that
(This assumption may relaxed to with some extra work; see KYY . We do not pursue this direction here.) We introduce the dimensional ratio
provided is chosen small enough.
We assume that the entries of the matrix are independent (but not necessarily identically distributed) real-valued random variables satisfying
The population covariance matrix is defined as
denote the empirical spectral density of . We suppose that
This latter assumption means that the spectrum of cannot be concentrated at zero.
Sometimes it will be convenient to make the following stronger assumption on :
The assumption (2.9) will frequently simplify the presentation and the proofs. Thanks to our general assumptions (2.7) and (2.8), it will always be relatively easy to remove (2.9). In particular, we emphasize that the assumption is purely qualitative in nature, and is made in order to simplify expressions involving the inverse of . The case of may always be easily obtained by considering and then taking at fixed . We refer to Section 11.1 below for the details on how to remove the assumption (2.9).
To avoid repetition, we summarize our basic assumptions for future reference.
We suppose that (2.1), (2.4), (2.5), (2.7), and (2.8) hold.
2. Asymptotic eigenvalue density
Moreover, is the Stieltjes transform of a probability measure with bounded support in .
The rest of this subsection is devoted to a discussion of the basic properties of the asymptotic eigenvalue density . Much this discussion is well known; see e.g. BS2 ; SC ; HHN . The reader interested only in the principal components of , i.e. its top eigenvalues, can skip this subsection and proceed directly to the results in Theorem 3.14 (with ) and Corollary 3.19 (i).
Let be the number of distinct nonzero eigenvalues of , and write
In Figures 2.1 and 2.2, we illustrate the graph of for the cases and respectively. In Figure 2.3 we plot the density of for the examples from Figures 2.1 and 2.2. The behaviour of may be entirely understood by an elementary analysis of .
We have and for .
We deduce from Lemma 2.4 that is even. We denote by the critical points in , and by the unique critical point in . For we define the critical values .
The following result gives the basic structure of .
Our assumptions on (i.e. on the spectrum of ) take the form of the following regularity conditions.
We say that the edge is regular if
The edge regularity condition from Definition 2.7 (i) has previously appeared (in a slightly different form) in several works on sample covariance matrices. For the rightmost edge , it was introduced in ElK2 and was subsequently used in the works SL ; Ona ; BaoPanZhou on the distribution of eigenvalues near the top edge . For general , it was introduced in HHN . The second condition of (2.14) states that the gap in the spectrum of adjacent to the edge does not close for large ; the third condition of (2.14) ensures a regular square-root behaviour of the spectral density near , and in particular rules out outliers.
We conclude this subsection with a couple of examples verifying the regularity conditions of Definition 2.7.
We suppose that is fixed, and that and all converge in as . We suppose that all critical points of are nondegenerate, and that for . Then it is easy to check that, for small enough , all edges and bulk components are regular in the sense of Definition 2.7.
Results
In this section we state our main results for the sample covariance matrix defined in Section 2. To ease readability, we state the anisotropic local law under three different sets of increasingly weak assumptions. In Section 3.2, we assume that all edges and bulk components are regular in the sense of Definition 2.7. In Section 3.3, we assume only regularity of a single edge or bulk component, and state the anisotropic local law in the vicinity of the corresponding edge or bulk component. As an application, we prove the edge universality near a regular edge. Finally, in Section 3.4 we give a general anisotropic law, where the concrete regularity assumptions from Definition 2.7 are replaced with a more general but abstract stability condition. The reader interested only in the principal components of can proceed directly to the results in Theorem 3.14 (with ) and Corollary 3.19 (i).
In this preliminary subsection we introduce some basic notations and definitions. Our main results take the form of local laws, which relate the resolvents
to the Stieltjes transform of the asymptotic density . These local laws may be formulated in a simple, unified fashion under the assumption (2.9) using an block matrix, which is a linear function of .
We consistently use the letters , , and . We label the indices of the matrices according to
The motivation behind this definition is that, assuming (2.9), a control of immediately yields control of the resolvents and via the identities
for . Both of these identities may be easily checked using Schur’s complement formula.
Next, we introduce a deterministic matrix , which we shall prove is close to with high probability (in the sense of Definition 3.4 (iii) below).
We also extend to an matrix
The following notion of a high-probability bound was introduced in EKY2 , and has been subsequently used in a number of works on random matrix theory. It provides a simple way of systematizing and making precise statements of the form “ is bounded with high probability by up to small powers of ”.
be two families of nonnegative random variables, where is a possibly -dependent parameter set.
We say that is stochastically dominated by , uniformly in , if for all (small) and (large) we have
If is stochastically dominated by , uniformly in , we use the notation . Moreover, if for some complex family we have we also write .
Because of (2.1), all (or some) factors of in Definition 3.4 could be replaced with or without changing the definition of stochastic domination.
2. The full spectrum
As explained at the beginning of this section, we first state the anisotropic local law for all arguments under the assumption that all edges and bulk components are regular in the sense of Definition 2.7. In the next subsection, we relax these assumptions by restricting the domain of to the vicinity of an edge or bulk component, and only requiring the regularity of the corresponding edge or bulk component.
Our main result is the following anisotropic local law. We introduce the fundamental control parameter
and the Stieltjes transform of the empirical eigenvalue density of ,
Fix . Suppose that (2.9) and Assumption 2.1 hold. Suppose moreover that every edge satisfying and every bulk component is regular in the sense of Definition 2.7. Then
(Note that the presence of the factors in (3.10) strengthens the result; by (2.7), they can be trivially dropped to obtain a weaker estimate. They ensure that the control of the error is stronger in directions where the covariance is small.)
Outside the support of the asymptotic spectrum, one has stronger control all the way down to the real axis.
Fix . Suppose that (2.9) and Assumption 2.1 hold. Suppose moreover that every edge satisfying is regular in the sense of Definition 2.7 (i). Then
uniformly in satisfying .
An explicit expression for the error term in (3.12) may be obtained from (A.7) below.
Theorem 3.7 can be used to obtain a complete picture of the outlier eigenvalues of in the case where a bounded number of eigenvalues of are changed to some arbitrary values (in particular possibly violating the regularity assumption from Definition 2.7 (i)). We note that the outliers may also lie between bulk components of . The analysis is similar to the one performed in KYY for the case ; we omit the details.
In the remainder of this subsection, we state several corollaries of Theorem 3.6 where the assumption (2.9) is removed. From Theorem 3.6 it is not hard to deduce the following result on the resolvents and , defined in (3.1).
Fix . Suppose that Assumption 2.1 holds. Suppose moreover that every edge satisfying and every bulk component is regular in the sense of Definition 2.7. Then
Theorem 3.7 has an analogous corollary for and for satisfying , whereby the right-hand sides of (3.13) and (3.14) are replaced with the right-hand side of (3.12); we omit the precise statement.
The following consistency check may be applied to the deterministic matrices on the left-hand sides of (3.13) and (3.14). If, in the identity , we replace and with the corresponding deterministic matrices from the left-hand sides of (3.13) and (3.14), we recover (2.11).
We conclude this subsection with another consequence of Theorem 3.6 – eigenvalue rigidity. We denote by
the nontrivial eigenvalues of , and by the classical eigenvalue locations defined through
If , it is convenient to relabel and separately for each bulk component . To that end, for we define the classical number of eigenvalues in the -th bulk component through
Fix . Suppose that Assumption 2.1 holds. Suppose moreover that every edge satisfying and every bulk component is regular in the sense of Definition 2.7. Then we have for all and satisfying that
As in (BEKYY, , Theorem 2.8), Corollary 3.9 and Theorem 3.12 imply the complete delocalization, with respect to an arbitrary deterministic basis, of the eigenvectors of and associated with eigenvalues satisfying .
Note that Theorem 3.12 in particular implies an exact separation of the eigenvalues into connected components, whereby the number of eigenvalues in the -th connected component is with high probability equal to the deterministic number . This phenomenon of exact separation was first established in BS2 ; BS3 .
3. Individual spectral regions and edge universality
In this subsection we elaborate on the results of Subsection 3.2 by only requiring the regularity of the edge or bulk component near which the spectral parameter lies. To that end, for fixed we define the subdomains
Fix . Suppose that Assumption 2.1 holds. Suppose that the edge is regular in the sense of Definition 2.7 (i). Then there exists a constant , depending only on , such that the following holds.
Let be the bulk component to which the edge belongs. Then for all satisfying we have
Fix . Suppose that Assumption 2.1 holds. Suppose that the bulk component is regular in the sense of Definition 2.7 (ii).
Suppose that at least one of the two edges and is regular in the sense of Definition 2.7 (i). Then for all satisfying we have .
Fix . Suppose that Assumption 2.1 holds.
As in Section 3.2, if then the error parameters on the right-hand sides of (3.10), (3.13), and (3.14) in (i) and (ii) of Theorems 3.14 and 3.16 can be replaced with the smaller quantity and the lower bound relaxed to . (See Theorem 3.7.) We omit the detailed statement.
Like Theorem 3.18, Corollary 3.19 (ii) also holds for the joint distribution at several regular edges. In particular, for the case and under the assumption that the top and bottom edges of are regular, we obtain the universality of the condition number of .
Corollary 3.19 (i) for was previously established in SL under the assumption that is diagonal, corresponding to uncorrelated population entries. Before that, Corollary 3.19 (i) for and diagonal was established in BaoPanZhou , following the results of ElK2 ; Ona in the complex Gaussian case. Corollary 3.19 (ii) was recently established for Gaussian in HHN .
4. The general anisotropic local law
In this subsection we conclude the statement of our results with a general anisotropic local law, which takes the form of a black box yielding the anisotropic local law for assuming it has been established for a much simpler matrix. This latter result may be proved independently. In particular, this black box formulation may be used to establish the anisotropic local law in cases where the regularity assumptions from Definition 2.7 fail; we do not pursue such generalizations here. Aside from its great generality, this black box formulation also makes precise the three Steps (A) – (C) mentioned in the introduction, which constitute the basic strategy of our proof.
We begin by introducing some basic terminology.
The main conclusion of this paper is that the anisotropic local law holds for general and provided that the entrywise local law holds for Gaussian and diagonal . This latter case may be established independently, as we illustrate in Section 5 and Appendix A.
Aside from Assumption 2.1, the only assumption that we shall need is
This assumption holds for instance under the regularity assumptions of Definition 2.7 (see Lemmas A.4, A.6, and A.8 below). Clearly, we always have (see (2.11)), and (3.20) is a uniform version of this bound. Generally, the assumption (3.20) is necessary to guarantee that the generalized matrix entries of (or, alternatively, of ) remain bounded. Indeed, in Corollary 3.9 we saw that the generalized entries of are close to those of .
The next theorem shows that the hypotheses in (i) and (ii) of Theorem 3.21 may be verified under a stability condition on the spectrum of , made precise in Definition 5.4 below.
5. Outline of the paper
Having completed the proof of the anisotropic local law, we prove the averaged local law (Theorem 3.21 (ii)) in Section 9. This will conclude the proof of Theorems 3.21 and 3.22. At the end of Section 9, we explain how to deduce Theorem 3.6, Corollary 3.9, and Theorems 3.14 (i)–(ii), 3.15 (i)–(ii), and 3.16.
In Section 10 we prove the rigidity of the eigenvalues (Theorems 3.12, 3.14 (iii), and 3.15 (iii)) and the universality of their joint distribution near the edges (Theorem 3.18). Next, in Section 11 we explain how to remove the assumption (2.9) and how to extend all of our results from the matrix to the matrix .
In Section 12, as a further illustration of the self-consistent comparison method, we present and prove analogous results for deformed Wigner matrices.
Basic tools
The rest of this paper is devoted to the proofs. In this preliminary section we collect various identities from linear algebra and simple estimates that we shall use throughout the paper.
We always use the following convention for matrix multiplication.
Suppose (2.9). Define the matrices
as well as the matrix
and the matrix
Throughout the following we frequently omit the argument from our notation.
Since and are only defined under the assumption (2.9), we shall always tacitly assume (2.9) whenever we use them. Note that under the assumption (2.9) we have .
For we define the minor . We also write . The matrices and are defined similarly. We abbreviate and .
Suppose that is diagonal. Then for we have
For and we have
In addition, if is diagonal, we have
For and we have
All of the identities from (i)–(v) hold for instead of if or and is diagonal.
The identities (4.3), (4.4), and (4.6) follow from Schur’s complement formula. The remaining identities follow easily from resolvent identities that have been previously derived in EYY1 ; EKYY2 ; they are summarized e.g. in (EKYY4, , Lemma 4.5). ∎
Next, we introduce the spectral decomposition of . We use the notation
for the singular value decomposition of , where
Moreover, for and we have
Finally, the estimates (4.16)–(4.20) remain true for instead of if or and is diagonal.
In order to prove (4.18), we use (4.3) to write
In order to prove (4.19), we use (4.3) and (4.12) to get
Finally, the same estimates for instead of follow using a trivial modification of the above argument. ∎
The following result may be used to estimate the factors in Lemma 4.6 with high probability. It follows from (BEKYY, , Theorem 2.10).
Under the assumptions (2.1), (2.4), and (2.5), there exists a constant such that with high probability.
Using Lemma 4.8, we observe that we may improve (4.16) provided we settle for a high-probability instead of a deterministic statement.
The claim is an easy consequence of the first identity of (4.3) combined with (4.2). ∎
We conclude this section with the following basic properties of , which can be proved as in BaoPanZhou and the references therein.
Fix and suppose that (2.3), (2.7), and (2.8) hold. Then there exists a constant such that
In particular, from the upper bound in (4.21) we deduce that has a bounded density on .
Entrywise local law for diagonal ΣΣ\Sigma
In this section we prove Theorem 3.22, hence performing Step (A) of the proof mentioned in the introduction. The proof of Theorem 3.22 is similar to previous proofs of local entrywise laws, such as PY ; BEKYY . We follow the basic approach of (BEKYY, , Section 4), and only give the details where the argument departs significantly from that of BEKYY .
The main novel observation of this section is that the equation (2.11) arises very easily from the random matrix model by a double application of Schur’s complement formula. Heuristically, this may be seen using the identities (4.4) and (4.6). Indeed, suppose that for . We ignore the random fluctuations in (4.4) to get
Similarly, ignoring the random fluctuations in (4.6), we get
Plugging (5.2) into (5.1) yields (2.11). In this section we give a rigorous justification of these approximations.
In this subsection we establish the following weaker version of Theorem 3.22. It is analogous to (BEKYY, , Proposition 4.2).
The rest of this subsection is devoted to the proof of Proposition 5.1. For each we define
Recalling (2.11), we find that the functions and satisfy
Next, we define the random control parameters
We extend the definitions of and for by setting and for . We may therefore write
Moreover, we define the averaged control parameters
For we introduce the conditional expectation
Using (4.6) we get for
and using (4.4) we get for
In analogy to (BEKYY, , Section 4), we define the -dependent event \Xi\mathrel{\vbox{\hbox{.}\hbox{.}}}=\bigl{\{}{\Lambda\leqslant(\log N)^{-1}}\bigr{\}} and the control parameter
The following estimate is analogous to (BEKYY, , Lemma 4.4).
The proof relies on the identities from Lemma 4.4 and large deviation estimates, like that of (BEKYY, , Lemma 4.4) and (PY, , Theorems 6.8 and 6.9). Note first that (5.4), combined with (4.21) and (5.3), yields
Using (4.11) and a simple induction argument, it is not hard to conclude that
for any and satisfying .
Let us first estimate in (5.9). We shall in fact prove that
Let us start with for . Using (4.7), (5.12), and a large deviation estimate (see (BEKYY, , Lemma 3.1)), we find
where in the first step we used (4.11) and (5.12). This yields (5.13) for .
An analogous argument for (see e.g. (EKYY4, , Lemma 5.2)) completes the proof of (5.9).
In order to prove (5.10), we proceed similarly. For , we proceed as above to get , where we used that by (4.16). Similarly, as in (5.14) we get
where in the last step we used that , by (4.16). Moreover, from (4.1) and (4.3) we get immediately that , and a similar argument for implies that . This concludes the proof. ∎
As in (BEKYY, , Lemma 4.7), it is easy to derive from (5.16) and Lemma 5.2 that
Plugging this and (5.15) into (5.16) yields
Suppose that the assumptions of Theorem 3.22 hold. Define
If suppose also that
2. Fluctuation averaging and proof of Theorem 3.22
The estimate (5.21) is a trivial extension of (BEKYY, , Lemma 4.9) and (EKYY4, , Theorem 4.7). The estimate (5.22) may be proved using exactly the same method, explained in (EKYY4, , Appendix B). The only complication in the proof is that the coefficients are random and depend on . Using (4.11), this is dealt with in the proof of (EKYY4, , Appendix B) by writing, for any ,
Anisotropic local law for Gaussian X𝑋X
We now begin the proof of Theorem 3.21, which consists of Sections 6–9. In this section we perform the first step of the proof, by establishing Theorem 3.21 for the special case that is Gaussian. This corresponds to Step (B) of the proof mentioned in the introduction.
Theorem 3.21 holds if is Gaussian.
The rest of this section is devoted to the proof of Proposition 6.1. We shall in fact prove the following result.
In this proof we abbreviate . Using (6.1) with and , we find that it suffices to prove
for all . In components, this reads
for and .
The estimate (6.4) is trivial by assumption. What remains is the proof of (6.2) and (6.3). It is based on the polynomialization method developed in (BEKYY, , Section 5). The argument is very similar to that of BEKYY , and we only outline the differences.
Let us begin with (6.2). By the assumption and orthogonality of , we have
for (which follows from (4.7)), (4.6), and
as follows from (4.6), (5.4), and (which may itself be deduced from (4.6)). We omit further details.
The rest of the argument is the same as before. ∎
Self-consistent comparison I: the main argument
In this section we establish Theorem 3.21 (i) under the additional assumption that the third moment of all entries of is zero.
Suppose that the assumptions of Theorem 3.21 hold. Suppose moreover that satisfies the additional condition
The rest of this section is devoted to the proof of Proposition 7.1. This is the heart of the proof of the anisotropic local law.
Computing the derivatives on the left-hand side of (7.3) leads to a product of terms of the form
and their complex conjugates (here we omit the argument ). Terms of type (i) are simply kept as they are; they will be put into the second term on the right-hand side of (7.3). Terms of type (ii) are the key to the gain that allows us to compensate losses from several other terms, such as the terms of type (iii) (see blow). The gain is obtained in combination with the summation over and on the left-hand side of (7.3), according to estimates of the form
which allows us to estimate the right-hand side of (7.5) by . The a priori bound (7.6) will be obtained from a bootstrapping, as explained below.
The preceding discussion was rather cavalier with the various constants in the factors . In fact, some care has to be taken to ensure that the factors of arising from the use of the a priori bounds (7.6) and (7.7) are compensated by a sufficiently large number of factors of the form (7.5). Moreover, in order to be able to iterate this argument from large to small scales, as explained in the next paragraph, we need an explicit bound of the constant in (7.3) in terms of the constant in (7.6).
Since is fixed, the bootstrapping consists of steps. The resulting estimates contain an extra factor . Since can be made arbitrarily small, the claim will follow on all scales.
2. Bootstrapping on the spectral scale
The proof consists of a bootstrap argument from larger scales to smaller scales in multiplicative increments of . Here
is fixed, where is a universal constant that will be chosen large enough in the proof ( will work). For any we define , where
The bootstrapping is started by the following result.
What remains is the proof of Lemma 7.4. We shall estimate the random variable
3. Interpolation
We use the interpolation outlined in Section 7.1.
Introduce the notations and . For , , and , denote by the law of . For we define the law
Let be a triple of independent random matrices, where for the matrix has law
We shall prove Lemma 7.5 by interpolation between the ensembles and . The bound for the Gaussian case is given by the following result.
Lemma 7.5 holds if is replaced with .
The basic interpolation formula is given by the following lemma, which follows from the fundamental theorem of calculus.
In order to prove Lemma 7.10, we compare the ensembles and via . Clearly, it suffices to prove the following result.
where we recall the definitions of and from (7.10) and (7.11).
It suffices to estimate the first term. Setting and , we define the subsets of indices
and treat each separately. For we find
where in the second step we used that , as follows from (2.7) and Lemma 4.8. This concludes the proof. ∎
where in the second step we used (3.20) and the definition of from (3.5). The estimate (7.19) now follows easily using Lemma 7.14.
4. Expansion
Let and . Define the matrix through
The following result provides a priori bounds for the entries of .
Suppose that is a random variable satisfying . Then
for all and .
We use (7.21) with , , and , so that . By (3.20), we have
To simplify notation, we introduce the function
The following result is easy to deduce from (7.21) and Lemma 7.14.
where we used that has vanishing first and third moments, by (7.1), and its variance is equal to . Recalling our goal (7.18), we therefore find that we only have to prove
for . Here we used the bounds (2.5).
holds for . Then (7.27) holds for .
To simplify notation, we abbreviate and . The proof consists of a repeated application of the identity
for , which follows from Lemmas 7.7 and 7.15. Fix . Using (7.29) we get
The claim now follows easily using (2.5). ∎
What therefore remains is to prove (7.28). Since it only involves the matrix ensemble , for the remainder of the proof we abbreviate . Recalling the notation (7.24), we find from Lemma 7.16 that it suffices to prove the following result.
We conclude this subsection with a general digression on the algebra of the Bernoulli interpolation method underlying the proof of Proposition 7.1, and in particular establish the formula (1.8) from the introduction. Let and be arbitrary independent finite families of independent random variables. Define as in (1.6). Then we get
where denotes the family obtained from by replacing with . Fix and abbreviate f(x)\mathrel{\vbox{\hbox{.}\hbox{.}}}=F\bigl{(}{X^{\theta,x}_{(\alpha)}}\bigr{)}, , , and . We may assume that , , and are independent. We want to compute the difference
Since we are only interested in the algebra of the interpolation, we assume for simplicity that all random variables have finite exponential moments and that is analytic. (Otherwise, as in the computations above, the expansions have to be truncated.) Repeating the steps of the proof of Lemma 7.16, we find
5. Introduction of words and conclusion of the proof
In order to prove (7.30), we shall have to exploit the detailed structure of the derivatives in on the left-hand side of (7.30). The following definition introduces the basic algebraic objects that we shall use.
Next, we assign to each letter its value through
(Our choice of the names of the two letters is suggestive of their value. Note, however, that it is important to distinguish the abstract letter from its value, which is an index in and may be used as a summation index.)
If with as in (7.31) we define
for any . This may be easily deduced from (7.21). Using Leibnitz’s rule we conclude that
To prove (7.30), it therefore suffices to prove that
for and words satisfying . (The proof of (7.33) is the same with slightly heaver notation.) Treating words with separately, we find that it suffices to prove
for , , and words satisfying , , and for . Note that we also have the bound .
In order to estimate (7.35), we introduce the quantity
For we have the rough bound
Finally, for we have the sharper bound
In addition to the high-probability bounds from Lemma 7.20, we have the rough estimate
for some and all satisfying ; this follows easily from Definition 7.19 and Lemma 4.9.
By pigeonholing on the words appearing on the left-hand side of (7.35), if then there exist at least two words satisfying . Using Lemma 7.20 we therefore get
From Lemma 4.6 combined with Lemma 4.8 we get
where in the second step we used (7.20), and in the last step the definition of . Inserting (7.41) into (7.40), we get using Lemma 7.7 and (7.39) that the left-hand side of (7.35) is bounded by
Using that , we find that the left-hand side of (7.35) is bounded by
where we used that and . Choose . Then, by assumption (7.9) on we have . Moreover, if and then . We conclude that the left-hand side of (7.35) is bounded by
Now (7.35) follows using Young’s inequality. This concludes the proof of (7.30), and hence of Lemma 7.11. The proof of Proposition 7.1 is therefore complete.
Self-consistent comparison II: general X𝑋X
In this section and the next we prove the following result, which is Proposition 7.1 without the condition (7.1).
The proof of Proposition 8.1 builds on that of Proposition 7.1. Throughout this section, we take over the notations of Section 7 without further comment. In particular, as in (7.9), is some large enough constant (which may depend only on ) and is fixed and small, satisfying (7.9).
The assumption (7.1) was used in the proof of Proposition 7.1 only in (7.26), where it ensured that the summation over starts from instead of . Without the assumption (7.1), we in addition have to estimate the term in (7.26). It therefore suffices to prove the following result.
As in Lemma 7.16, one may easily replace the matrix in the definition of with . As in Lemma 7.17, from now on we abbreviate . Thus, we find that in order to prove Lemma 8.2 it suffices to prove the following result, which complements Lemma 7.17.
Next, recall the definition of words from Definition 7.19. As in (7.33)–(7.35), Lemma 8.3 is proved provided we can show the following result, which is analogous to (7.35).
Beyond the factor of , the need to obtain the bound from (7.34) that is strong enough to close the self-consistent estimate presents significant difficulties, which we outline briefly. These difficulties are roughly of two types. (a) The off-diagonal entries of are in general not small; only entries of are small. (b) A priori, using the estimate (7.19), the entries of are not bounded by but by . This would yield a bound on the coefficients of the polynomial in that is a high power of , which may become too large to conclude the proof.
2. Tagged words
Lemma 8.4 holds provided it holds with replaced by in (7.34).
Next, from Definition 7.19 we conclude, analogously to the proof of Lemma 7.20, that if for and then
where in the last step we used (7.41). We therefore conclude using Lemma 7.7 that
If (i.e. is the empty word) we define
where we introduced the matrix \Pi^{(\mu)}\mathrel{\vbox{\hbox{.}\hbox{.}}}=\bigl{(}{\Pi_{st}\mathrel{\vbox{\hbox{.}\hbox{.}}}s,t\in\mathcal{I}\setminus\{\mu\}}\bigr{)}.
From (8.4), (8.3), and Definition 8.7 we find that for we have
Recalling Lemma 8.5, we conclude that, in order to prove Lemma 8.4, it suffices to prove the following result.
We shall require the following rough bounds, which refine those of Lemma 7.20.
Suppose that (8.1) holds. Let be a tagged word. Then
Using a large deviation estimate (see (BEKYY, , Lemma 3.1)) combined with Lemmas 4.6 and 4.8 we get
where in the second step we used (8.1). Therefore,
for all . This concludes the proof. ∎
Using the large deviation estimates from (BEKYY, , Lemma 3.1), we find . Since by (8.1), we therefore deduce using that
Hence there exists a constant such that
Plugging in the definition of , we find
Now we replace the factors on the left-hand side of (8.7) with the leading term in (8.15). From (8.15), (7.19), (8.16), and Lemma 7.7 we get
where the coefficients are -measurable and satisfy the bound
Moreover, for any tagged word we split
4. Degree counting
For the following we choose , , and as in Lemma 8.8. We introduce the abbreviation
Here the constant accounts for the immaterial constants depending on arising from the combinatorics of all partitions. For , , and as above, we may estimate
where we used (2.5). Now we use the estimate
which follows from (8.11), (8.12), and (8.1). Hence we get
where the maxima are subject to the same conditions as above.
Next, from and we deduce that
Together with , this gives
where in the last step we used that , and . For the case , we also used that .
We conclude that the left-hand side of (8.21) is bounded by
Plugging (8.33) and (8.34) into (8.32), and recalling Lemma 7.7, yields
5. Power counting: proof of (8.36)
as may be easily checked from Definition 8.7. (Recall that for .) We conclude that (8.36) holds provided we can show that
The proof of (8.21), and hence of Lemma 8.8, is therefore complete. This concludes the proof of Lemma 8.4.
We have hence proved Proposition 8.1. Recalling Proposition 6.1, we conclude the proof of Theorem 3.21 (i).
Averaged local law
for some constant . In particular, we recover Definition 3.20 (iii) by setting . In this section we prove the following result, which also completes the proof of Theorem 3.21.
The proof is similar to that of Propositions 7.1 and 8.1, and we only explain the differences. Note that now there is no bootstrapping, since the necessary a priori bounds are obtained from the anisotropic local law. In analogy to (7.15), we define
Following the argument leading up to (7.30) and Lemma 8.3 to the letter, we find that it suffices to prove that
For , the claim (9.2) easily follows from (9.3) and an application of Young’s inequality to the factors satisfying and ; we omit the details. What remains, therefore, is to verify (9.2) for .
which may be regarded as an improvement of Lemma 8.11. The proof of (9.5) is similar to that of Lemma 8.11, and we merely give a sketch. We have to estimate an expression of the form
where and . (This is simply the general form of a polynomial in of the correct degree.) The proof of (9.5) relies on the crucial observations that, by Theorem 3.21 (i),
Next, let satisfy , so that the left-hand side of (9.5) does not depend on . There are such indices , and for each such we find
As in the argument following Lemma 8.11, we conclude that the claim holds provided that
This concludes the proof of Theorem 3.21. We conclude this section by drawing consequences from Theorems 3.21 and 3.22.
Similarly, Theorems 3.15 (i) and 3.16 (i) follow from Lemmas A.6 and A.8 respectively, combined with Theorems 3.21 and 3.22 ∎
These follow immediately from Theorem 3.6 and Theorems 3.14 (i), 3.15 (i), and 3.16 (i) respectively, combined with (3.3), (3.4), and (3.5). ∎
How to remove the assumption (2.9) in the above proof is explained in Section 11.1.
Eigenvalue rigidity and edge universality
In the first part of this section we establish eigenvalue rigidity (Theorems 3.12, 3.14 (iii), and 3.15 (iii)). As a consequence of Theorem 3.6 we prove Theorem 3.7. In the second part of this section we establish the edge universality from Theorem 3.18. We assume (2.9) throughout this section; how to remove it is explained in Section 11.1 below.
For simplicity, we first prove Theorem 3.12, i.e. we assume that all edges and bulk components are regular. The proof consists of three steps. First, we prove that with high probability there are no eigenvalues at a distance greater than from the support of . Second, we prove that a neighbourhood of the -th bulk component of contains with high probability exactly eigenvalues (recall the definition (3.16)). Third, we use the averaged local law from Theorem 3.6 together with the first two steps to complete the proof.
We begin with the first step. We define the distance to the spectral edge through
Fix . Under (2.9) and the assumptions of Theorem 3.12 we have
with high probability (recall Definition 4.7).
The argument is similar to that of the previous works EYY3 ; PY , and we only explain how to adapt it. From (BEKYY, , Theorem 2.10), we find that with high probability for some large enough constant depending on . From (2.7) we therefore deduce that with high probability for some constant depending on . It therefore suffices to prove that, after a possible decrease of ,
In order to prove (10.2), it suffices to prove that . By (A.7) and Lemmas A.4, A.6, and A.8, we have
It is easy to check that the control parameter on the right-hand side of (10.4) satisfies (9.1), so that by Proposition 9.1 and Theorem 3.6 it suffices to prove (10.4) for diagonal . The proof is identical to that of Section 5, except that we use the strong stability of (2.11) from Definition A.2, which is established in Lemmas A.5, A.6, and A.8. This concludes the proof of (10.4), and hence of (10.2).
With applications to the deformed Wigner matrices in Section 12 in mind, we give another proof of (10.2), which is based on (BS2, , Section 6). Using a partial fraction decomposition, one easily finds that there exist universal constants such that
Similarly, we have (with the same and )
for all . Setting , we get from (10.5) and (10.6) that
where in the last step we used that for . This immediately implies that with high probability there is no eigenvalue in . ∎
The second step represents most of the work. It is a counting argument, based on a continuous deformation of the matrix to another matrix for which the claim is obvious. Since the eigenvalues depend continuously on the deformation parameter and each intermediate matrix satisfies a gap condition from Lemma 10.1, we shall be able to conclude that the number of eigenvalues in a neighbourhood of the -th component does not change under the deformation. We shall in fact need two deformations: one which deforms the original matrix to a Gaussian one, , with the same expectation as , and another which deforms the Gaussian matrix to another Gaussian matrix where some eigenvalues of have been increased.
For , we introduce the number of eigenvalues to the right of the -th gap,
Under (2.9) and the assumptions of Theorem 3.12 we have with high probability.
As explained above, the first step in the proof of Proposition 10.2 is a deformation of the general matrix to a Gaussian one.
Let be general and a Gaussian matrix of the same dimensions. Suppose that (2.9) and the assumptions of Theorem 3.12 hold. If with high probability under the law of , then with high probability under the law of .
Let and be independent. For define and denote by the eigenvalues of . We also write for defined in terms of . Note that Lemmas 4.8 and 10.1 hold for uniformly in . Recalling (2.7), we deduce that there exists a constant such that
with high probability, uniformly in and . The claim now follows by considering in the lattice , where is chosen large enough that and is the constant from (10.8). Indeed, suppose with high probability. Then we use (10.8) and the gap from Lemma 10.1 to deduce that with high probability. The claim for therefore follows by induction. ∎
Using Lemma 10.3, in order to prove Proposition 10.2 it suffices to prove the following result.
Suppose that (2.9) and the assumptions of Theorem 3.12 hold, that is Gaussian, and that is diagonal. Then with high probability.
Abbreviate . We write the diagonal matrix in the block form , where contains the top eigenvalues of . We introduce the deformed covariance matrix . In particular, . The idea is to increase until the claim for replaced by may be deduced from simple linear algebra. Then we use a continuity argument to compare to .
We add the argument to quantities to indicate that they are defined in terms of instead of . Using (A.5) (see also Figure 2.1), it is not hard to find that the gap is increasing in , and that
for some constants . Using that , we deduce from Lemma 4.8 that with high probability. Now let be a fixed large time to be chosen later. By a continuity argument using Lemma 10.1 we therefore find that there exists a constant such that for any we have
It therefore remains to show that there exists a large enough , depending only on , such that for the left-hand side of (10.10) is equal to . Writing as a block matrix, we find
From (A.4) we deduce that , by the assumption of Theorem 3.12. From (BEKYY, , Theorem 2.10), we therefore deduce that with high probability for some positive constants . Moreover, from Lemma 4.8 we deduce that with high probability. We conclude that for large enough (depending on the constants and above) the matrix (10.11) has with high probability exactly eigenvalues of order , and all other eigenvalues are of order . Since \operatorname{spec}(Q(t))\cap\bigl{[}{a_{2k+1}(t)+N^{-2/3+\tau},a_{2k}-N^{-2/3+\tau}}\bigr{]}=\emptyset with high probability by Lemma 10.1, we conclude that for large enough the left-hand side of (10.10) is equal to with high probability. This concludes the proof. ∎
Proposition 10.2 follows immediately from Lemmas 10.3 and 10.4. This concludes the second step outlined above.
Finally, the third step – the conclusion of the proof of Theorem 3.12 – follows from Theorem 3.6, Lemma 10.1, and Proposition 10.2 by repeating the analysis of EYY3 ; PY with merely cosmetic changes, as explained in the proof of Theorem 3.12. This concludes the proof of Theorem 3.12 under the assumption (2.9). Moreover, as explained in Section 11.1, the assumption (2.9) may be easily removed.
Next, we note that the proof Theorem 3.7 under the assumption (2.9) is an easy consequence of Theorems 3.6 and 3.12, following the proof of (BEKYY, , Theorem 3.12). The key tool is the spectral decomposition (4.15). We omit further details.
Finally, the proof of Theorems 3.14 (iii) and 3.15 (iii) under the assumption (2.9) is exactly the same as that of Theorem 3.12. Indeed, one can check that in the proof of Theorem 3.12 only the weaker assumptions of Theorems 3.14 (iii) and 3.15 (iii) are needed. We omit the details.
2. Edge universality assuming (2.9)
Finally, we prove the edge universality from Theorem 3.18 under the assumption (2.9).
We may therefore without loss of generality assume that is Gaussian. By orthogonal / unitary invariance of the law of , we may furthermore assume that is diagonal. This concludes the proof. ∎
General matrices: removing (2.9) and extension to Q˙˙𝑄\dot{Q}
In this section we explain how our results, proved under the assumption (2.9) and for the matrix , may be generalized to hold without the assumption (2.9) and for the matrix from (1.4) as well.
For simplicity, throughout the proofs up to now we made the assumption (2.9). As advertised, this assumption is not necessary. In this section we explain how to dispense with it. The argument relies on simple approximation and linear algebra. Roughly, if has a zero eigenvalue, we consider instead and let ; if is not square, we augment it to a square matrix by padding it out with zeros. While this extension is simple, we emphasize that it relies crucially on the fact we do not assume that has a lower bound (the assumption (2.9) only requires the qualitative bound ).
We distinguish the cases and . Suppose first that . We extend to an matrix by setting . Define the matrices
By polar decomposition, we have , where is orthogonal. Therefore . Moreover, from (2.11) we get , which we use tacitly in the following.
We define the matrix
where we used (4.3). Note that has the block form
Using this representation of as a block of , it is easy to drop the assumption (2.9). For example, suppose that Theorem 3.6 has been proved under the assumption (2.9). Applying it to and using a simple approximation argument in , we find the following generalization of Theorem 3.6.
Fix . Suppose that (2.9) and Assumption 2.1 hold. Suppose moreover that every edge satisfying and every bulk component is regular in the sense of Definition 2.7. Then (3.10) and (3.11) hold for defined in (11.1) and defined in (3.1).
In particular, Corollary 3.9 and Theorems 3.12, 3.14, 3.15, 3.16, and 3.18 follow easily from their counterparts proved under the assumption (2.9). This concludes the discussion for the case .
Finally, we consider the case . We set and , where is an matrix, independent of , with independent entries satisfying (2.4) and (2.5). Hence, is and is . Now we have , and we have reduced the problem to the case , which was dealt with above.
2. Extension to Q˙˙𝑄\dot{Q}
In this section we explain how our results have to be modified for from (1.4). For simplicity of presentation, we make the assumption (2.9); it may be easily removed as explained in Section 11.1.
As before it is easy to obtain and from .
The local law for reads as follows.
To simplify notation, we omit the factor in the definition of . It may easily be put back by scaling the argument . We prove the anisotropic local law by estimating the four blocks of individually. Using simple linear algebra and (4.1) and (4.3), we get
Since satisfies the anisotropic local law by assumption, it is easy to deduce from Lemma 4.4 and (4.21) that
From (4.3) we get , so that, using the anisotropic local law for , we conclude
Finally, using (4.3) for as well as (11.4), we get
Using the anisotropic law for and Lemma 4.4, we therefore get
We also obtain eigenvalue rigidity for the eigenvalues of . For instance, Theorem 3.12 has the following counterpart.
Theorem 3.12 remains valid if is replaced with .
The proof follows that of Theorem 3.12 to the letter, using Theorem 11.2 as input. In fact, as explained around (10.4), we need a stronger bound than (11.3) outside of the spectrum; this stronger bound follows easily from (11.5) and the analogous stronger bound for established in (10.4). ∎
Finally, we obtain edge universality for . The following result is proved exactly like Theorem 3.18, using Theorems 11.2 and 11.3 as input.
Theorem 3.18 remains valid if is replaced with .
Deformed Wigner matrices
In this section we apply our method to deformed Wigner matrices as a further illustration of its applicability. Since the statements and arguments are similar to those of the previous sections, we keep the presentation concise.
Let be an Wigner matrix whose upper-triangular entries are independent and satisfy the same conditions (2.4) and (2.5) as . Let be a deterministic matrix satisfying . For definiteness, we suppose that and are real symmetric matrices, remarking that similar results also hold for complex Hermitian matrices.
The main result of this section is the anisotropic local law for the deformed Wigner matrix , analogous to Theorem 3.21. As an application, we establish the edge universality of . We remark that the entrywise local law and edge universality were previously established in LSY under the assumption that is diagonal. Before that, edge universality was established in CP1 ; Joh2 ; Shc under the assumption that is a GUE matrix. A somewhat different direction was pursued in BES15 ; BG15 ; Kar15 , where local laws for the sum of two matrices that are invariant under unitary or orthogonal conjugations were analysed.
We denote the eigenvalues of by . Moreover, we define the resolvent , as well as
The following definition is the analogue of Definition 3.20 for deformed Wigner matrices.
The following result is the analogue of Theorem 3.21. Throughout the following we denote by a GOE matrix and by the diagonalization of .
The following result guarantees the rigidity of the extreme eigenvalues of . As explained below, the rigidity of all eigenvalues will be a simple consequence of Theorems 12.2 (ii) and 12.4.
We say the extreme eigenvalues of are rigid if \bigl{[}{\lambda_{1}(W+A)-L_{+}}\bigr{]}_{+}\prec N^{-2/3} and \bigl{[}{\lambda_{N}(W+A)-L_{-}}\bigr{]}_{-}\prec N^{-2/3}.
Together, Theorems 12.2 (ii) and 12.4 easily yield the rigidity of the eigenvalues, as explained in the proof of (EYY3, , Theorem 2.2). We may apply Theorem 12.2 to extend the results of LSY to arbitrary non-diagonal . For instance, we obtain the following edge universality result.
A similar result holds for the extreme eigenvalues near the left edge.
In LSY , the assumptions of Theorem 12.4 were verified for a large class of diagonal matrices . Moreover, for such matrices , it was proved that the limit of the second term on the left-hand side of (12.1) is governed by the Tracy-Widom-Airy statistics. Theorem 12.5 therefore provides an extension of (LSY, , Theorem 2.8) to non-diagonal matrices . We refer to LSY for the detailed statements about the distribution of the eigenvalues of .
Further applications of Theorem 12.2 include a study of the eigenvectors of and the outliers of finite-rank perturbations of . We do not pursue these questions here.
2. Proof of Theorem 12.2 (i)
Here is a constant that may depend only on , and satisfies (7.9).
Making minor adjustments to the argument of Sections 7.3 and 7.4, we find that it suffices to prove the following result, which is analogous to Lemma 7.30.
for any . Moreover, if in addition then (12.5) holds also for .
The rest of this subsection is devoted to the proof of Lemma 12.7. We note that the word structure describing the derivatives on the left-hand side of (12.5) is very similar to that of (7.30) (given in Definition 7.19). Hence the proof of Lemma 12.7 is similar to that of Lemma 7.17.
The proof of (12.5) for is a trivial modification of the argument given in Section 7.5, whose details we omit. What remains therefore is the proof of (12.5) for under the assumption . From now on we omit the arguments .
The main new ingredient of the proof for deformed Wigner matrices is a further iteration step at fixed . Suppose that
for some . Since , the estimate (12.6) is stronger than the first estimate of (12.4). Note that, by assumption (12.4), the estimate (12.6) holds for . Assuming (12.6), we shall prove a self-improving bound of the form
Once (12.7) is proved, we may use it iteratively to obtain increasingly accurate bounds for the left-hand side of (12.3). After each step, we obtain an improved a priori bound (12.6), whereby is reduced by powers of . After iterations, (12.5) for follows.
It therefore suffices to prove (12.7) under the assumptions (12.4) and (12.6). As in Section 7.5, it suffices to prove
We discuss the three cases separately.
or a term obtained from one of these two by exchanging and ; here . We deal first with the first expression; the others are deal with analogously. We split it according to
We deal with the summation over using the estimate
where in the last step we used (12.4); a similar estimate holds for the summation over . Using (12.12), (12.6), and , we may estimate the second term on the right-hand side of (12.11) as
provided is chosen small enough, depending on and . The third and fourth terms on the right-hand side of (12.11) are estimated in exactly the same way.
The case q=2𝑞2q=2
or an expression obtained from one of these two by exchanging and . The contribution of the first expression of (12.13) is estimated using (12.4) and (12.12) by
provided is chosen small enough, depending on and .
Next, in order to estimate the contribution of the second expression of (12.13), we split
where we used (12.4), (12.6), and (12.12). Similarly, using (12.4) and (12.12) we get
for small enough . Putting (12.14) and (12.15) together, it is easy to deduce (12.9).
The case q=3𝑞3q=3
3. Proof of Theorem 12.2 (ii)
The proof is similar to that from Section 9. As in Section 12.2, it suffices to prove the following result.
for any . Moreover, if in addition then (12.16) holds also for .
The case can be easily proved as in covariance case. We therefore focus on the case in Lemma 12.8. The proof is similar to the discussion below (12.8). The main difference is that for each we have some extra averaging , and we need to extract an extra factor (or, alternatively, ) from this average. We take over the notations from Sections 7.5 and 12.2 without further comment. We consider the three cases separately, and tacitly use the anisotropic local law from Theorem 12.2 (i).
Consider first the case . We estimate
as desired. Next, in the case we estimate
Finally, in the case we estimate
Using a similar bound for the sum over , we find . All other terms are obtained from these three by exchanging and .
The case q=2𝑞2q=2
In this case is of the form
or an expression obtained from one of these two by exchanging and . These may be written as
We estimate the contribution of the first expression by
Next, we split the contribution of the second expression of (12.18) as
Using the anisotropic local law, it is easy to prove that and \bigl{\lvert}\sum_{j}\Pi_{jj}(G^{2})_{ji}\bigr{\rvert}\prec N^{3/2}\Psi^{2}. Therefore
where we estimate as above. This concludes the proof in the case .
The case q=3𝑞3q=3
In this case is of the form , or an expression obtained by exchanging and in some of the three factors. We estimate its contribution by
4. Proof of Theorem 12.4
The proof is analogous to that of Theorem 3.12 in Section 10. Define the domain
From the assumptions on , it is not hard to deduce the estimate
5. Proof of Theorem 12.5
Analogously to the proof of Theorem 3.18, the proof is a routine application of the Green function comparison method near the edge (EYY3, , Section 6). The key technical inputs are Theorems 12.2 and 12.4. Note that Theorem 12.2 is applicable since the Green function comparison argument only involves satisfying with some small constant . Hence the assumption (b) from Theorem 12.2 is satisfied.
Appendix A Properties of ϱitalic-ϱ\varrho and Stability of (2.11)
This appendix is devoted the proofs of the basic properties of and the stability of (2.11) in the sense of Definition 5.4. In this appendix we abbreviate .
In this subsection we establish the basic properties of and prove Lemmas 2.4–2.6.
From (A.1) we find that on . Therefore for there is at most one point such that . We conclude that has at most two critical points of . Using the boundary conditions of on , we conclude the proof of for .
From (A.1) we also find that for . We conclude that there exists at most one point such that . Using the boundary conditions of on , we deduce that .
Finally, if we have as . From the boundary conditions of on we therefore deduce that . Moreover, if we find from (A.1) that for all . This concludes the proof. ∎
By multiplying both sides of the equation in (2.11) with the product of all denominators on the right-hand side of (2.12), we find that may be also written as , where is a polynomial of degree , whose coefficients are affine linear functions of . (Here we used that all are distinct and that .) This polynomial characterization of is useful for the proofs of Lemmas 2.5 and 2.6.
For define the subset . The graph of restricted to is depicted in red in Figures 2.1 and 2.2. From Lemma 2.4 we deduce that if then if and only if contains two distinct critical points of , in which case is an interval. Moreover, we always have .
Next, we observe that for any we have . Indeed, if there were then we would have (see Figure 2.1). Since is equivalent to and has degree , we arrive at the desired contradiction. We conclude that the sets , , may be strictly ordered.
The claim may now be reformulated as
We focus on the first inequality of (A.2); the proof of the second one is similar. We have to show that whenever and . First, we claim that for we have
To prove (A.3), we remark that if then is equivalent to containing two distinct critical points. Moreover, in , from which we deduce that the number of distinct critical points in each , , does not decrease as decreases. Recalling that , we deduce (A.3).
Next, suppose that there exist satisfying and . From (A.3) we deduce that for all . Moreover, by a simple continuity argument and the fact that for each we have either or , we conclude that for all . As explained after (A.2), this is impossible for small enough . This concludes the proof of (A.2), and hence also of the first sentence of Lemma 2.5.
Finally, in order to prove the third sentence of Lemma 2.5 we have to show that and . It is easy to see that if then there is an with , which is impossible since and . Moreover, the estimate is easy to deduce from the definition of and the estimates and . ∎
For the following counting argument, in order to avoid extraneous complications we assume that for all . Recall the definition of from (3.16).
Suppose that for all . Then and
where in the last step we used the definition of from (2.12). Hence (A.5) follows. Note that we proved (A.5) also for . Now (A.4) follows easily using for all . Since (recall Lemma 2.6), we get . Therefore summing over in (A.5) yields , as desired. This concludes the proof in the case .
Next, suppose that . Then we deduce directly from (2.12) that (see Figure 2.2). Hence, (A.4) follows. Finally, exactly as above, we get (A.5) for . This concludes the proof. ∎
A.2. Stability near a regular edge
The rest of this appendix is devoted to the proofs of the assumption (3.20) and the stability of (2.11) in the sense of Definition 5.4. In fact, we shall prove the following stronger notation of stability, which will be needed to establish eigenvalue rigidity. Recall the definition of from (10.1).
That the notion of stability from Definition A.2 it is stronger than that of Definition 5.4 follows from the estimate
In this subsection we deal with a regular edge; the case of a regular bulk component is dealt with in the next subsection. We begin with basic estimates on the behaviour of near a critical point.
Fix . Suppose that the edge is regular in the sense of Definition 2.7 (i). Then there exists a , depending only on , such that
From Lemma 2.5 we get . Since by assumption for some constant , we find from Lemma 4.10 that . The lower bound follows from and (A.1), using that . This proves the first estimate of (A.9).
Next, from and (A.1) we find
Using the first estimate of (A.9) and (2.14) we find from (A.10) that . Moreover for (i.e. ) all terms in (A.10) have the same sign, and it is easy to deduce that .
Next, we prove the lower bound for . Suppose for definiteness that is odd. (The case of even is handled in exactly the same way.) For we have and
where in the second step we used that as follows from (A.1), and in the last step the estimate . We therefore find
Since by the assumption in Definition 2.7 (i) and , we deduce the lower bound .
Finally, the last estimate of (A.9) follows easily from the assumption (2.14) and by differentiating (A.1). ∎
We record the following easy consequence of (A.9). Recalling that and , we find, using (A.9) and the fact that is continuous in a neighbourhood of , that after choosing small enough (depending on ) we have
We first prove (A.7). From (A.11) we find that there exists a such that for we have
Finally, (A.8) follows from the assumption (2.14) and (A.11). ∎
We take over the notation of Definition 5.4. Abbreviate , so that .
Suppose first that for some constant . Then, from the assumption we get , and therefore from the uniqueness of (2.11) we find . Hence,
where we used the trivial bound . This yields (A.6) at .
What remains therefore is the case , which we assume for the rest of the proof. We drop the arguments and write the equation as
Note that and depend on , and also depends on . We suppose that
Then we claim that for small enough we have
In order to prove (A.16), we note that, by Lemma A.3, the statement of (A.11) holds under the assumption (A.15) provided is chosen small enough. Using (A.15), (A.8) (see Lemma A.3) , and (A.10) we get
Using (A.9) and (A.11) we conclude that for small enough we have . This concludes the proof of the first estimate of (A.16).
In order to prove the second estimate of (A.16), we note that , so that for small enough we get from (A.11) and Lemma A.3 that
Using (A.11) and Lemmas A.3 and (4.10), we conclude for small enough that
This concludes the proof of the second estimate of (A.16).
A.3. Stability in a regular bulk component
The estimates (A.7) and (A.8) follow trivially from the assumption in Definition 2.7 (ii), using as follows from Lemma 2.5.
Note that Definition 2.7 (ii) immediately implies that for some constant . The upper bound easily follows from the definition (A.14) combined with Lemma 4.10. What remains is the proof of the lower bound . To that end, we take the imaginary part of (2.11) to get
Using that , a simple analysis of the arguments of the expressions on the left-hand side of (A.17) yields
where in the last step we used (A.17). Recalling the definition of from (A.14), we conclude that . This concludes the proof. ∎
The bulk regularity condition from Definition 2.7 (ii) is stable under perturbation of . To see this, define the shifted empirical density , and the associated Stieltjes transform and function . Differentiating in yields . Using that
by (A.8), and , we find from that for . A simple extension of this argument shows that if Definition 2.7 (ii) holds for then it holds for all with in some ball of fixed radius around zero.
A.4. Stability outside of the spectrum
Since is the Stieltjes transform of a measure with bounded density (see Lemma 4.10), we find that
Next, we note that if for some fixed then (A.8) is trivially true by (A.18). On the other hand, if then we get
Acknowledgements
A.K. was partially supported by Swiss National Science Foundation grant 144662 and the SwissMAP NCCR grant. J.Y. was partially supported by NSF Grant DMS-1207961. We are very grateful to the Institute for Advanced Study, Thomas Spencer, and Horng-Tzer Yau for their kind hospitality during the academic year 2013-2014. We also thank the Institute for Mathematical Research (FIM) at ETH Zürich for its generous support of J.Y.’s visit in the summer of 2014. We are indebted to Jamal Najim for stimulating discussions.