Anisotropic local laws for random matrices

Antti Knowles, Jun Yin

Introduction

for large MM and with high probability. Here m(z)m(z) is the Stieltjes transform of the asymptotic eigenvalue density, which we denote by ϱ\varrho. We call an estimate of the form (1.1) an averaged law.

As may be easily seen by taking the imaginary part of (1.1), control of the convergence of M−1Tr⁡R(z)M^{-1}\operatorname{Tr}R(z) yields control of an order ηM\eta M eigenvalues around the point EE. A local law is an estimate of the form (1.1) for all η≫M−1\eta\gg M^{-1}. Note that the approximation (1.1) cannot be correct at or below the scale η≍M−1\eta\asymp M^{-1}, at which the behaviour of the left-hand side of (1.1) is governed by the fluctuations of individual eigenvalues. Such local laws have become a cornerstone of random matrix theory, starting from the work ESY2 where a local law was first established for Wigner matrices. Well known corollaries of a local law include bounds on the eigenvalue counting function as well as eigenvalue rigidity. Moreover, local laws constitute the main tool needed to analyse (a) the distribution of eigenvalues (including the universality of the local spectral statistics), (b) eigenvector delocalization, (c) the distribution of eigenvectors, and (d) finite-rank deformations of QQ.

In fact, for all of the applications (a)–(d), the averaged local law from (1.1) is not sufficient. One has to control not only the normalized trace of R(z)R(z) but the matrix R(z)R(z) itself, by showing that R(z)R(z) is close to some deterministic matrix depending on zz, provided that η≫M−1\eta\gg M^{-1}. Such control was first obtained for Wigner matrices in EYY1 , where the closeness was established in the sense of individual matrix entries: Rij(z)≈m(z)δijR_{ij}(z)\approx m(z)\delta_{ij}. We call such an estimate an entrywise local law. More generally, in KY2 ; BEKYY this closeness was established in the sense of generalized matrix entries:

Analogous results for uncorrelated sample covariance matrices were obtained in PY ; BEKYY . The estimate (1.2) states that for large MM the resolvent R(z)R(z) is approximately isotropic (i.e. proportional to the identity matrix), and we accordingly call an estimate of the form (1.2) an isotropic local law. We remark that the basis-independent control in (1.2) is crucial for many applications, including the distribution of eigenvectors and the study of finite-rank deformations of QQ.

Unlike in the case of Wigner matrices and uncorrelated sample covariance matrices mentioned above, the resolvent R(z)R(z) is in general not close to a multiple of the identity matrix, but rather to some general deterministic matrix P(z)P(z). In that case (1.2) is to be replaced with

We call an estimate of the form (1.3) an anisotropic local law. The goal of this paper is to develop a method yielding (anisotropic) local laws for many matrix models built from sums and products of deterministic or independent random matrices. Applications include all the four (a)–(d) listed above, some of which we illustrate in this paper.

2. Sample covariance matrices

3. Outline of results

For simplicity, we focus on the matrix QQ, bearing in mind that similar results also apply for Q˙\dot{Q} (see Section 11.2). We assume that the three matrix dimensions M,M^,NM,\widehat{M},N are comparable, and that the entries of NX\sqrt{N}X possess a sufficient number of bounded moments. Moreover, we assume that ∥Σ∥\lVert\Sigma\rVert is bounded, and that the spectrum of Σ\Sigma satisfies certain regularity conditions, given Definition 2.7 below, which essentially state that the connected components of the support of ϱ\varrho are separated by some positive constant, and that the density of ϱ\varrho has square root decay near its edges in (0,∞)(0,\infty). Note that we do not assume that TT is square, and in particular TT may have many vanishing singular values; this allows us to cover e.g. the general linear model of multivariate statistics.

Our main result is the anisotropic local law for QQ. Roughly, it states that (1.3) holds with

where m(z)m(z) is the Stieltjes transform of the asymptotic density ϱ\varrho. In fact, we prove a more general anisotropic local law that is more useful in applications. Its formulation is most transparent under the additional assumption that T=T∗=Σ1/2T=T^{*}=\Sigma^{1/2}, although this assumption is a mere convenience and may be easily removed (see Section 11.1). We prove an anisotropic local law of the form

A simple application of Schur’s complement formula to (1.5) yields the anisotropic local law for the resolvent of Q=TXX∗T∗Q=TXX^{*}T^{*} and a similar result for the resolvent of the companion matrix X∗T∗TXX^{*}T^{*}TX. The estimate (1.5) holds with high probability, and we give an explicit and optimal error bound. We remark that the anisotropic local law holds under very general assumptions on the distribution of XX, the dimensions of XX and TT, and the spectrum of TT∗TT^{*}. In particular, we make no assumptions on the singular vectors of TT. We remark that, previously, an anisotropic global law, valid for η≍1\eta\asymp 1, was derived in HLNP for a different matrix model.

As an application of the anisotropic local law, we prove the edge universality of the eigenvalues near the soft spectral edges, whereby the joint distribution of the eigenvalues is asymptotically governed by the Tracy-Widom-Airy statistics of random matrix theory. More precisely, we prove that the asymptotic distribution of the eigenvalues near the soft edges depends only on the nonzero spectrum of TT∗TT^{*}. This may be regarded as a universality result in both the distribution of the entries of XX and the (left and right) singular vectors of TT, including their dimensions. We then conclude that the Tracy-Widom-Airy statistics hold near the soft edges by noting that they have been previously established HHN ; ElK2 ; Ona ; SL for Gaussian XX and diagonal TT.

We also prove the rigidity of eigenvalues, as well as the complete delocalization of the eigenvectors with respect to an arbitrary deterministic basis. Further applications of the anisotropic local law, such as the distribution of the eigenvectors and an analysis of the outliers and BBP-type phase transitions in finite-rank deformations, will appear elsewhere.

Finally, we also apply our method to deformed Wigner matrices of the form W+AW+A, where WW is a Wigner matrix and AA a bounded Hermitian matrix. This model describes Wigner matrices whose entries may have arbitrary expectations. As for QQ, we establish the Tracy-Widom-Airy statistics near the spectral edges of W+AW+A. More precisely, we prove that the asymptotic distribution of the eigenvalues near the edges depends only on the asymptotic spectrum of AA, which may be regarded as a universality result in the distribution of WW and the eigenvectors of AA. We then conclude that the Tracy-Widom-Airy statistics hold near the edges by noting that they have been previously established LSY for diagonal AA.

4. Overview of the proof

We conclude this section by outlining some ideas of the proof of the anisotropic local law. Roughly, the proof proceeds in three steps: (A) the entrywise local law for Gaussian XX and diagonal Σ\Sigma, (B) the anisotropic local law for Gaussian XX and general Σ\Sigma, and (C) the anisotropic local law for general XX and general Σ\Sigma. Steps (A) and (B) may be performed by adapting the methods of BEKYY , and we do not comment on them any further.

The main argument, and the bulk of the proof, is Step (C). Its core is a self-consistent comparison method, which yields the anisotropic local law for general XX assuming it has been proved for Gaussian XX. Up to now, all interpolation or Lindeberg replacement arguments in random matrix theory have crucially relied on a local law as input. Since the local law is exactly what we are trying to prove, a different approach is clearly needed – one that does not need a local law to work. Our method yields a new way of deriving local laws for general random matrices expressed as polynomials of deterministic and random matrices whose entries are independent. It relies on two key novel ideas: (a) continous Green function comparison argument where the errors are controlled self-consistently, and (b) a bootstrapping of the Green function comparison on the spectral scale η\eta.

In the remainder of this subsection we give a few details of our method, and in particular explain the ideas (a) and (b) in more detail. We construct a continuous family (Xθ)θ∈(X^{\theta})_{\theta\in} of matrices, whereby X0X^{0} is a Gaussian ensemble and X1X^{1} is the ensemble in which we are interested. Then we operate in the (θ,η)(\theta,\eta)-plane and perform simultaneously a continuous interpolation in θ\theta and a discrete bootstrapping in η\eta. (See Figure 1.1 below.)

Interpolation methods have been extensively used in random matrix theory, in the form of both discrete Lindeberg-type replacement schemes Chat ; TV1 ; EYY1 ; EYY3 and continuous interpolations SL ; LSY ; LP . Moreover, Dyson Brownian motion, as used e.g. in SL ; LSY ; ESY4 ; ESYY , may be regarded as a form of continuous interpolation for the case that X0X^{0} is Gaussian. The usual choice of interpolation, as used in SL ; LP and also resulting from Dyson Brownian motion, is Xθ=1−θ X0+θ X1X^{\theta}=\sqrt{1-\theta}\,X^{0}+\sqrt{\theta}\,X^{1}. In this paper, we instead interpolate using i.i.d. Bernoulli random variables:

where Kn,αθK_{n,\alpha}^{\theta} is defined through the formal power series

Next, we explain the basic strategy behind the self-consistent comparison method mentioned in (a) above. We choose a function

The next section is devoted to the definition of the model and the statement of results. We give an outline of the structure of the paper in Section 3.5 below.

Conventions

The fundamental large parameter is NN. All quantities that are not explicitly constant may depend on NN; we almost always omit the argument NN from our notation.

We use τ>0\tau>0 in various assumptions to denote a positive constant that may be chosen arbitrarily small. A smaller value of τ\tau corresponds to a weaker assumption. All constants may C,cC,c depend on τ\tau, and we neither indicate nor track this dependence.

Model

In this section we define our model, list our key assumptions, and explain the basic structure of the asymptotic eigenvalue density ϱ\varrho.

We consider the M×MM\times M matrix Q\vbox..=TXX∗T∗Q\mathrel{\vbox{\hbox{.}\hbox{.}}}=TXX^{*}T^{*}, where TT is a deterministic M×M^M\times\widehat{M} matrix and XX a random M^×N\widehat{M}\times N matrix. We regard NN as the fundamental parameter and M≡MNM\equiv M_{N} and M^≡M^N\widehat{M}\equiv\widehat{M}_{N} as depending on NN. Here, and throughout the following, we omit the index NN from our notation, bearing in mind that all quantities that are not explicitly constant (such as the constant τ\tau) may depend on NN. For simplicity, we always make the assumption that

(This assumption may relaxed to log⁡M≍log⁡M^≍log⁡N\log M\asymp\log\widehat{M}\asymp\log N with some extra work; see KYY . We do not pursue this direction here.) We introduce the dimensional ratio

provided τ>0\tau>0 is chosen small enough.

We assume that the entries XiμX_{i\mu} of the M^×N\widehat{M}\times N matrix XX are independent (but not necessarily identically distributed) real-valued random variables satisfying

The population covariance matrix is defined as

denote the empirical spectral density of Σ\Sigma. We suppose that

This latter assumption means that the spectrum of Σ\Sigma cannot be concentrated at zero.

Sometimes it will be convenient to make the following stronger assumption on TT:

The assumption (2.9) will frequently simplify the presentation and the proofs. Thanks to our general assumptions (2.7) and (2.8), it will always be relatively easy to remove (2.9). In particular, we emphasize that the assumption Σ>0\Sigma>0 is purely qualitative in nature, and is made in order to simplify expressions involving the inverse of Σ\Sigma. The case of Σ⩾0\Sigma\geqslant 0 may always be easily obtained by considering Σ+εIM\Sigma+\varepsilon I_{M} and then taking ε↓0\varepsilon\downarrow 0 at fixed NN. We refer to Section 11.1 below for the details on how to remove the assumption (2.9).

To avoid repetition, we summarize our basic assumptions for future reference.

We suppose that (2.1), (2.4), (2.5), (2.7), and (2.8) hold.

2. Asymptotic eigenvalue density

Moreover, m(z)m(z) is the Stieltjes transform of a probability measure ϱ\varrho with bounded support in [0,∞)[0,\infty).

The rest of this subsection is devoted to a discussion of the basic properties of the asymptotic eigenvalue density ϱ\varrho. Much this discussion is well known; see e.g. BS2 ; SC ; HHN . The reader interested only in the principal components of QQ, i.e. its top eigenvalues, can skip this subsection and proceed directly to the results in Theorem 3.14 (with k=1k=1) and Corollary 3.19 (i).

Let n\vbox..=∣supp⁡π∖{0}∣n\mathrel{\vbox{\hbox{.}\hbox{.}}}=\lvert\operatorname{supp}\pi\setminus\{0\}\rvert be the number of distinct nonzero eigenvalues of Σ\Sigma, and write

In Figures 2.1 and 2.2, we illustrate the graph of ff for the cases ϕ<1\phi<1 and ϕ>1\phi>1 respectively. In Figure 2.3 we plot the density of ϱ\varrho for the examples from Figures 2.1 and 2.2. The behaviour of ϱ\varrho may be entirely understood by an elementary analysis of ff.

We have ∣C∩I0∣=∣C∩I1∣=1\lvert\mathcal{C}\cap I_{0}\rvert=\lvert\mathcal{C}\cap I_{1}\rvert=1 and ∣C∩Ii∣∈{0,2}\lvert\mathcal{C}\cap I_{i}\rvert\in\{0,2\} for i=2,…,ni=2,\dots,n.

We deduce from Lemma 2.4 that ∣C∣=2p\lvert\mathcal{C}\rvert=2p is even. We denote by x1⩾x2⩾⋯⩾x2p−1x_{1}\geqslant x_{2}\geqslant\cdots\geqslant x_{2p-1} the 2p−12p-1 critical points in I1∪⋯∪InI_{1}\cup\dots\cup I_{n}, and by x2px_{2p} the unique critical point in I0I_{0}. For k=1,…,2pk=1,\dots,2p we define the critical values ak\vbox..=f(xk)a_{k}\mathrel{\vbox{\hbox{.}\hbox{.}}}=f(x_{k}).

The following result gives the basic structure of ϱ\varrho.

Our assumptions on π\pi (i.e. on the spectrum of Σ\Sigma) take the form of the following regularity conditions.

We say that the edge k=1,…,2pk=1,\dots,2p is regular if

The edge regularity condition from Definition 2.7 (i) has previously appeared (in a slightly different form) in several works on sample covariance matrices. For the rightmost edge k=1k=1, it was introduced in ElK2 and was subsequently used in the works SL ; Ona ; BaoPanZhou on the distribution of eigenvalues near the top edge a1a_{1}. For general kk, it was introduced in HHN . The second condition of (2.14) states that the gap in the spectrum of ϱ\varrho adjacent to the edge aka_{k} does not close for large NN; the third condition of (2.14) ensures a regular square-root behaviour of the spectral density near aka_{k}, and in particular rules out outliers.

We conclude this subsection with a couple of examples verifying the regularity conditions of Definition 2.7.

We suppose that nn is fixed, and that s1,…,sns_{1},\dots,s_{n} and ϕπ({s1}),…,ϕπ({sn})\phi\pi(\{s_{1}\}),\dots,\phi\pi(\{s_{n}\}) all converge in (0,∞)(0,\infty) as N→∞N\to\infty. We suppose that all critical points of lim⁡Nf\lim_{N}f are nondegenerate, and that lim⁡Nai>lim⁡Nai+1\lim_{N}a_{i}>\lim_{N}a_{i+1} for i=1,…,2pi=1,\dots,2p. Then it is easy to check that, for small enough τ\tau, all edges k=1,…,2pk=1,\dots,2p and bulk components k=1,…,pk=1,\dots,p are regular in the sense of Definition 2.7.

Results

In this section we state our main results for the sample covariance matrix QQ defined in Section 2. To ease readability, we state the anisotropic local law under three different sets of increasingly weak assumptions. In Section 3.2, we assume that all edges and bulk components are regular in the sense of Definition 2.7. In Section 3.3, we assume only regularity of a single edge or bulk component, and state the anisotropic local law in the vicinity of the corresponding edge or bulk component. As an application, we prove the edge universality near a regular edge. Finally, in Section 3.4 we give a general anisotropic law, where the concrete regularity assumptions from Definition 2.7 are replaced with a more general but abstract stability condition. The reader interested only in the principal components of QQ can proceed directly to the results in Theorem 3.14 (with k=1k=1) and Corollary 3.19 (i).

In this preliminary subsection we introduce some basic notations and definitions. Our main results take the form of local laws, which relate the resolvents

to the Stieltjes transform mm of the asymptotic density ϱ\varrho. These local laws may be formulated in a simple, unified fashion under the assumption (2.9) using an (M+N)×(M+N)(M+N)\times(M+N) block matrix, which is a linear function of XX.

We consistently use the letters i,j∈IMi,j\in\mathcal{I}_{M}, μ,ν∈IN\mu,\nu\in\mathcal{I}_{N}, and s,t∈Is,t\in\mathcal{I}. We label the indices of the matrices according to

The motivation behind this definition is that, assuming (2.9), a control of GG immediately yields control of the resolvents RNR_{N} and RMR_{M} via the identities

for μ,ν∈IN\mu,\nu\in\mathcal{I}_{N}. Both of these identities may be easily checked using Schur’s complement formula.

Next, we introduce a deterministic matrix Π\Pi, which we shall prove is close to GG with high probability (in the sense of Definition 3.4 (iii) below).

We also extend Σ\Sigma to an I×I\mathcal{I}\times\mathcal{I} matrix

The following notion of a high-probability bound was introduced in EKY2 , and has been subsequently used in a number of works on random matrix theory. It provides a simple way of systematizing and making precise statements of the form “ξ\xi is bounded with high probability by ζ\zeta up to small powers of NN”.

be two families of nonnegative random variables, where U(N)U^{(N)} is a possibly NN-dependent parameter set.

We say that ξ\xi is stochastically dominated by ζ\zeta, uniformly in uu, if for all (small) ε>0\varepsilon>0 and (large) D>0D>0 we have

If ξ\xi is stochastically dominated by ζ\zeta, uniformly in uu, we use the notation ξ≺ζ\xi\prec\zeta. Moreover, if for some complex family ξ\xi we have ∣ξ∣≺ζ\lvert\xi\rvert\prec\zeta we also write ξ=O≺(ζ)\xi=O_{\prec}(\zeta).

Because of (2.1), all (or some) factors of NN in Definition 3.4 could be replaced with MM or M^\widehat{M} without changing the definition of stochastic domination.

2. The full spectrum

As explained at the beginning of this section, we first state the anisotropic local law for all arguments zz under the assumption that all edges and bulk components are regular in the sense of Definition 2.7. In the next subsection, we relax these assumptions by restricting the domain of zz to the vicinity of an edge or bulk component, and only requiring the regularity of the corresponding edge or bulk component.

Our main result is the following anisotropic local law. We introduce the fundamental control parameter

and the Stieltjes transform of the empirical eigenvalue density of X∗T∗TXX^{*}T^{*}TX,

Fix τ>0\tau>0. Suppose that (2.9) and Assumption 2.1 hold. Suppose moreover that every edge k=1,…,2pk=1,\dots,2p satisfying ak⩾τa_{k}\geqslant\tau and every bulk component k=1,…,pk=1,\dots,p is regular in the sense of Definition 2.7. Then

(Note that the presence of the factors Σ‾ ⁣ \underline{\Sigma}\!\, in (3.10) strengthens the result; by (2.7), they can be trivially dropped to obtain a weaker estimate. They ensure that the control of the error is stronger in directions where the covariance Σ\Sigma is small.)

Outside the support of the asymptotic spectrum, one has stronger control all the way down to the real axis.

Fix τ>0\tau>0. Suppose that (2.9) and Assumption 2.1 hold. Suppose moreover that every edge k=1,…,2pk=1,\dots,2p satisfying ak⩾τa_{k}\geqslant\tau is regular in the sense of Definition 2.7 (i). Then

uniformly in z∈[τ,τ−1]×(0,τ−1]z\in[\tau,\tau^{-1}]\times(0,\tau^{-1}] satisfying dist⁡(E,supp⁡ϱ)⩾N−2/3+τ\operatorname{dist}(E,\operatorname{supp}\varrho)\geqslant N^{-2/3+\tau}.

An explicit expression for the error term in (3.12) may be obtained from (A.7) below.

Theorem 3.7 can be used to obtain a complete picture of the outlier eigenvalues of QQ in the case where a bounded number of eigenvalues of Σ\Sigma are changed to some arbitrary values (in particular possibly violating the regularity assumption from Definition 2.7 (i)). We note that the outliers may also lie between bulk components of ϱ\varrho. The analysis is similar to the one performed in KYY for the case Σ=IM\Sigma=I_{M}; we omit the details.

In the remainder of this subsection, we state several corollaries of Theorem 3.6 where the assumption (2.9) is removed. From Theorem 3.6 it is not hard to deduce the following result on the resolvents RNR_{N} and RMR_{M}, defined in (3.1).

Fix τ>0\tau>0. Suppose that Assumption 2.1 holds. Suppose moreover that every edge k=1,…,2pk=1,\dots,2p satisfying ak⩾τa_{k}\geqslant\tau and every bulk component k=1,…,pk=1,\dots,p is regular in the sense of Definition 2.7. Then

Theorem 3.7 has an analogous corollary for RNR_{N} and RMR_{M} for zz satisfying dist⁡(E,supp⁡ϱ)⩾N−2/3+τ\operatorname{dist}(E,\operatorname{supp}\varrho)\geqslant N^{-2/3+\tau}, whereby the right-hand sides of (3.13) and (3.14) are replaced with the right-hand side of (3.12); we omit the precise statement.

The following consistency check may be applied to the deterministic matrices on the left-hand sides of (3.13) and (3.14). If, in the identity 1MTr⁡RM=1MTr⁡RN+ϕ−1ϕ1z\frac{1}{M}\operatorname{Tr}R_{M}=\frac{1}{M}\operatorname{Tr}R_{N}+\frac{\phi-1}{\phi}\frac{1}{z}, we replace Tr⁡RM\operatorname{Tr}R_{M} and Tr⁡RN\operatorname{Tr}R_{N} with the corresponding deterministic matrices from the left-hand sides of (3.13) and (3.14), we recover (2.11).

We conclude this subsection with another consequence of Theorem 3.6 – eigenvalue rigidity. We denote by

the nontrivial eigenvalues of QQ, and by γ1⩾γ2⩾⋯⩾γM∧N\gamma_{1}\geqslant\gamma_{2}\geqslant\cdots\geqslant\gamma_{M\wedge N} the classical eigenvalue locations defined through

If p⩾1p\geqslant 1, it is convenient to relabel λi\lambda_{i} and γi\gamma_{i} separately for each bulk component k=1,…,pk=1,\dots,p. To that end, for k=1,…,pk=1,\dots,p we define the classical number of eigenvalues in the kk-th bulk component through

Fix τ>0\tau>0. Suppose that Assumption 2.1 holds. Suppose moreover that every edge k=1,…,2pk=1,\dots,2p satisfying ak⩾τa_{k}\geqslant\tau and every bulk component k=1,…,pk=1,\dots,p is regular in the sense of Definition 2.7. Then we have for all k=1,…,pk=1,\dots,p and i=1,…,Nki=1,\dots,N_{k} satisfying γk,i⩾τ\gamma_{k,i}\geqslant\tau that

As in (BEKYY, , Theorem 2.8), Corollary 3.9 and Theorem 3.12 imply the complete delocalization, with respect to an arbitrary deterministic basis, of the eigenvectors of TXX∗T∗TXX^{*}T^{*} and X∗T∗TXX^{*}T^{*}TX associated with eigenvalues λi\lambda_{i} satisfying γi⩾τ\gamma_{i}\geqslant\tau.

Note that Theorem 3.12 in particular implies an exact separation of the eigenvalues into connected components, whereby the number of eigenvalues in the kk-th connected component is with high probability equal to the deterministic number NiN_{i}. This phenomenon of exact separation was first established in BS2 ; BS3 .

3. Individual spectral regions and edge universality

In this subsection we elaborate on the results of Subsection 3.2 by only requiring the regularity of the edge or bulk component near which the spectral parameter zz lies. To that end, for fixed τ,τ′>0\tau,\tau^{\prime}>0 we define the subdomains

Fix τ>0\tau>0. Suppose that Assumption 2.1 holds. Suppose that the edge k=1,…,2pk=1,\dots,2p is regular in the sense of Definition 2.7 (i). Then there exists a constant τ′>0\tau^{\prime}>0, depending only on τ\tau, such that the following holds.

Let k^\vbox..=⌊(k+1)/2⌋\hat{k}\mathrel{\vbox{\hbox{.}\hbox{.}}}=\lfloor{(k+1)/2}\rfloor be the bulk component to which the edge kk belongs. Then for all i=1,…,Nk^i=1,\dots,N_{\hat{k}} satisfying γk^,i∈[ak−τ′,ak+τ′]\gamma_{\hat{k},i}\in[a_{k}-\tau^{\prime},a_{k}+\tau^{\prime}] we have

Fix τ,τ′>0\tau,\tau^{\prime}>0. Suppose that Assumption 2.1 holds. Suppose that the bulk component k=1,…,2pk=1,\dots,2p is regular in the sense of Definition 2.7 (ii).

Suppose that at least one of the two edges 2k2k and 2k−12k-1 is regular in the sense of Definition 2.7 (i). Then for all i=1,…,Nki=1,\dots,N_{k} satisfying γk,i∈[a2k+τ′,a2k−1−τ′]\gamma_{k,i}\in[a_{2k}+\tau^{\prime},a_{2k-1}-\tau^{\prime}] we have ∣λk,i−γk,i∣≺N−1\lvert\lambda_{k,i}-\gamma_{k,i}\rvert\prec N^{-1}.

Fix τ,τ′>0\tau,\tau^{\prime}>0. Suppose that Assumption 2.1 holds.

As in Section 3.2, if dist⁡(E,supp⁡ϱ)⩾N−2/3+τ\operatorname{dist}(E,\operatorname{supp}\varrho)\geqslant N^{-2/3+\tau} then the error parameters Ψ(z)\Psi(z) on the right-hand sides of (3.10), (3.13), and (3.14) in (i) and (ii) of Theorems 3.14 and 3.16 can be replaced with the smaller quantity Im⁡m(z)/(Nη)\sqrt{\operatorname{Im}m(z)/(N\eta)} and the lower bound η⩾N−1+τ\eta\geqslant N^{-1+\tau} relaxed to η>0\eta>0. (See Theorem 3.7.) We omit the detailed statement.

Like Theorem 3.18, Corollary 3.19 (ii) also holds for the joint distribution at several regular edges. In particular, for the case β=2\beta=2 and under the assumption that the top and bottom edges of ϱ\varrho are regular, we obtain the universality of the condition number of QQ.

Corollary 3.19 (i) for β=1,2\beta=1,2 was previously established in SL under the assumption that Σ\Sigma is diagonal, corresponding to uncorrelated population entries. Before that, Corollary 3.19 (i) for β=2\beta=2 and diagonal Σ\Sigma was established in BaoPanZhou , following the results of ElK2 ; Ona in the complex Gaussian case. Corollary 3.19 (ii) was recently established for Gaussian XX in HHN .

4. The general anisotropic local law

In this subsection we conclude the statement of our results with a general anisotropic local law, which takes the form of a black box yielding the anisotropic local law for QQ assuming it has been established for a much simpler matrix. This latter result may be proved independently. In particular, this black box formulation may be used to establish the anisotropic local law in cases where the regularity assumptions from Definition 2.7 fail; we do not pursue such generalizations here. Aside from its great generality, this black box formulation also makes precise the three Steps (A) – (C) mentioned in the introduction, which constitute the basic strategy of our proof.

We begin by introducing some basic terminology.

The main conclusion of this paper is that the anisotropic local law holds for general XX and TT provided that the entrywise local law holds for Gaussian XX and diagonal TT. This latter case may be established independently, as we illustrate in Section 5 and Appendix A.

Aside from Assumption 2.1, the only assumption that we shall need is

This assumption holds for instance under the regularity assumptions of Definition 2.7 (see Lemmas A.4, A.6, and A.8 below). Clearly, we always have ∣1+m(z)σi∣>0\lvert 1+m(z)\sigma_{i}\rvert>0 (see (2.11)), and (3.20) is a uniform version of this bound. Generally, the assumption (3.20) is necessary to guarantee that the generalized matrix entries of (Q−z)−1(Q-z)^{-1} (or, alternatively, of G(z)G(z)) remain bounded. Indeed, in Corollary 3.9 we saw that the generalized entries of (Q−z)−1(Q-z)^{-1} are close to those of −z−1(1+m(z)Σ)−1-z^{-1}(1+m(z)\Sigma)^{-1}.

The next theorem shows that the hypotheses in (i) and (ii) of Theorem 3.21 may be verified under a stability condition on the spectrum of Σ\Sigma, made precise in Definition 5.4 below.

5. Outline of the paper

Having completed the proof of the anisotropic local law, we prove the averaged local law (Theorem 3.21 (ii)) in Section 9. This will conclude the proof of Theorems 3.21 and 3.22. At the end of Section 9, we explain how to deduce Theorem 3.6, Corollary 3.9, and Theorems 3.14 (i)–(ii), 3.15 (i)–(ii), and 3.16.

In Section 10 we prove the rigidity of the eigenvalues (Theorems 3.12, 3.14 (iii), and 3.15 (iii)) and the universality of their joint distribution near the edges (Theorem 3.18). Next, in Section 11 we explain how to remove the assumption (2.9) and how to extend all of our results from the matrix QQ to the matrix Q˙\dot{Q}.

In Section 12, as a further illustration of the self-consistent comparison method, we present and prove analogous results for deformed Wigner matrices.

Basic tools

The rest of this paper is devoted to the proofs. In this preliminary section we collect various identities from linear algebra and simple estimates that we shall use throughout the paper.

We always use the following convention for matrix multiplication.

Suppose (2.9). Define the I×I\mathcal{I}\times\mathcal{I} matrices

as well as the IM×IM\mathcal{I}_{M}\times\mathcal{I}_{M} matrix

and the IN×IN\mathcal{I}_{N}\times\mathcal{I}_{N} matrix

Throughout the following we frequently omit the argument zz from our notation.

Since H(z)H(z) and G(z)G(z) are only defined under the assumption (2.9), we shall always tacitly assume (2.9) whenever we use them. Note that under the assumption (2.9) we have RN=GNR_{N}=G_{N}.

For S⊂IS\subset\mathcal{I} we define the minor H(S)\vbox..=(Hst\vbox..s,t∈I∖S)H^{(S)}\mathrel{\vbox{\hbox{.}\hbox{.}}}=(H_{st}\mathrel{\vbox{\hbox{.}\hbox{.}}}s,t\in\mathcal{I}\setminus S). We also write G(S)\vbox..=(H(S))−1G^{(S)}\mathrel{\vbox{\hbox{.}\hbox{.}}}=(H^{(S)})^{-1}. The matrices GN(S)G_{N}^{(S)} and GM(S)G_{M}^{(S)} are defined similarly. We abbreviate ({s})≡(s)(\{s\})\equiv(s) and ({s,t})≡(st)(\{s,t\})\equiv(st).

Suppose that Σ\Sigma is diagonal. Then for i∈IMi\in\mathcal{I}_{M} we have

For i∈IMi\in\mathcal{I}_{M} and μ∈IN\mu\in\mathcal{I}_{N} we have

In addition, if Σ\Sigma is diagonal, we have

For r∈Ir\in\mathcal{I} and s,t∈I∖{k}s,t\in\mathcal{I}\setminus\{k\} we have

All of the identities from (i)–(v) hold for G(S)G^{(S)} instead of GG if S⊂INS\subset\mathcal{I}_{N} or S⊂IS\subset\mathcal{I} and Σ\Sigma is diagonal.

The identities (4.3), (4.4), and (4.6) follow from Schur’s complement formula. The remaining identities follow easily from resolvent identities that have been previously derived in EYY1 ; EKYY2 ; they are summarized e.g. in (EKYY4, , Lemma 4.5). ∎

Next, we introduce the spectral decomposition of GG. We use the notation

for the singular value decomposition of Σ1/2X\Sigma^{1/2}X, where

Moreover, for i∈IMi\in\mathcal{I}_{M} and μ∈IN\mu\in\mathcal{I}_{N} we have

Finally, the estimates (4.16)–(4.20) remain true for G(S)G^{(S)} instead of GG if S⊂INS\subset\mathcal{I}_{N} or S⊂IS\subset\mathcal{I} and Σ\Sigma is diagonal.

In order to prove (4.18), we use (4.3) to write

In order to prove (4.19), we use (4.3) and (4.12) to get

Finally, the same estimates for G(S)G^{(S)} instead of GG follow using a trivial modification of the above argument. ∎

The following result may be used to estimate the factors ∥X∗X∥\lVert X^{*}X\rVert in Lemma 4.6 with high probability. It follows from (BEKYY, , Theorem 2.10).

Under the assumptions (2.1), (2.4), and (2.5), there exists a constant C>0C>0 such that ∥X∗X∥⩽C\lVert X^{*}X\rVert\leqslant C with high probability.

Using Lemma 4.8, we observe that we may improve (4.16) provided we settle for a high-probability instead of a deterministic statement.

The claim is an easy consequence of the first identity of (4.3) combined with (4.2). ∎

We conclude this section with the following basic properties of mm, which can be proved as in BaoPanZhou and the references therein.

Fix τ>0\tau>0 and suppose that (2.3), (2.7), and (2.8) hold. Then there exists a constant C>0C>0 such that

In particular, from the upper bound in (4.21) we deduce that ϱ\varrho has a bounded density on [τ,∞)[\tau,\infty).

Entrywise local law for diagonal ΣΣ\Sigma

In this section we prove Theorem 3.22, hence performing Step (A) of the proof mentioned in the introduction. The proof of Theorem 3.22 is similar to previous proofs of local entrywise laws, such as PY ; BEKYY . We follow the basic approach of (BEKYY, , Section 4), and only give the details where the argument departs significantly from that of BEKYY .

The main novel observation of this section is that the equation (2.11) arises very easily from the random matrix model by a double application of Schur’s complement formula. Heuristically, this may be seen using the identities (4.4) and (4.6). Indeed, suppose that Gμμ≈mG_{\mu\mu}\approx m for μ∈IN\mu\in\mathcal{I}_{N}. We ignore the random fluctuations in (4.4) to get

Similarly, ignoring the random fluctuations in (4.6), we get

Plugging (5.2) into (5.1) yields (2.11). In this section we give a rigorous justification of these approximations.

In this subsection we establish the following weaker version of Theorem 3.22. It is analogous to (BEKYY, , Proposition 4.2).

The rest of this subsection is devoted to the proof of Proposition 5.1. For each i∈IMi\in\mathcal{I}_{M} we define

Recalling (2.11), we find that the functions mm and mim_{i} satisfy

Next, we define the random control parameters

We extend the definitions of σi\sigma_{i} and mim_{i} for i∈IMi\in\mathcal{I}_{M} by setting σμ\vbox..=1\sigma_{\mu}\mathrel{\vbox{\hbox{.}\hbox{.}}}=1 and mμ\vbox..=mm_{\mu}\mathrel{\vbox{\hbox{.}\hbox{.}}}=m for μ∈IN\mu\in\mathcal{I}_{N}. We may therefore write

Moreover, we define the averaged control parameters

For s∈Is\in\mathcal{I} we introduce the conditional expectation

Using (4.6) we get for i∈IMi\in\mathcal{I}_{M}

and using (4.4) we get for μ∈IN\mu\in\mathcal{I}_{N}

In analogy to (BEKYY, , Section 4), we define the zz-dependent event \Xi\mathrel{\vbox{\hbox{.}\hbox{.}}}=\bigl{\{}{\Lambda\leqslant(\log N)^{-1}}\bigr{\}} and the control parameter

The following estimate is analogous to (BEKYY, , Lemma 4.4).

The proof relies on the identities from Lemma 4.4 and large deviation estimates, like that of (BEKYY, , Lemma 4.4) and (PY, , Theorems 6.8 and 6.9). Note first that (5.4), combined with (4.21) and (5.3), yields

Using (4.11) and a simple induction argument, it is not hard to conclude that

for any S⊂IS\subset\mathcal{I} and t∈I∖St\in\mathcal{I}\setminus S satisfying ∣S∣⩽C\lvert S\rvert\leqslant C.

Let us first estimate Λo\Lambda_{o} in (5.9). We shall in fact prove that

Let us start with GijG_{ij} for i≠j∈IMi\neq j\in\mathcal{I}_{M}. Using (4.7), (5.12), and a large deviation estimate (see (BEKYY, , Lemma 3.1)), we find

where in the first step we used (4.11) and (5.12). This yields (5.13) for s,t∈IMs,t\in\mathcal{I}_{M}.

An analogous argument for ZsZ_{s} (see e.g. (EKYY4, , Lemma 5.2)) completes the proof of (5.9).

In order to prove (5.10), we proceed similarly. For η⩾1\eta\geqslant 1, we proceed as above to get ∣Zs∣≺N−1/2\lvert Z_{s}\rvert\prec N^{-1/2}, where we used that ∥G(s)∥⩽C\lVert G^{(s)}\rVert\leqslant C by (4.16). Similarly, as in (5.14) we get

where in the last step we used that Im⁡Gμμ(ij)⩽C\operatorname{Im}G_{\mu\mu}^{(ij)}\leqslant C, by (4.16). Moreover, from (4.1) and (4.3) we get immediately that ∣Gii∣⩽Cσi\lvert G_{ii}\rvert\leqslant C\sigma_{i}, and a similar argument for G(i)G^{(i)} implies that ∣Gjj(i)∣⩽Cσj\lvert G_{jj}^{(i)}\rvert\leqslant C\sigma_{j}. This concludes the proof. ∎

As in (BEKYY, , Lemma 4.7), it is easy to derive from (5.16) and Lemma 5.2 that

Plugging this and (5.15) into (5.16) yields

Suppose that the assumptions of Theorem 3.22 hold. Define

If Im⁡z<1\operatorname{Im}z<1 suppose also that

2. Fluctuation averaging and proof of Theorem 3.22

The estimate (5.21) is a trivial extension of (BEKYY, , Lemma 4.9) and (EKYY4, , Theorem 4.7). The estimate (5.22) may be proved using exactly the same method, explained in (EKYY4, , Appendix B). The only complication in the proof is that the coefficients σi2(1+mNσi)2\frac{\sigma_{i}^{2}}{(1+m_{N}\sigma_{i})^{2}} are random and depend on ii. Using (4.11), this is dealt with in the proof of (EKYY4, , Appendix B) by writing, for any j∈INj\in\mathcal{I}_{N},

Anisotropic local law for Gaussian X𝑋X

We now begin the proof of Theorem 3.21, which consists of Sections 6–9. In this section we perform the first step of the proof, by establishing Theorem 3.21 for the special case that X=XGaussX=X^{\text{\rm{Gauss}}} is Gaussian. This corresponds to Step (B) of the proof mentioned in the introduction.

Theorem 3.21 holds if X=XGaussX=X^{\text{\rm{Gauss}}} is Gaussian.

The rest of this section is devoted to the proof of Proposition 6.1. We shall in fact prove the following result.

In this proof we abbreviate GD≡GG^{D}\equiv G. Using (6.1) with OM=U∗O_{M}=U^{*} and ON=1O_{N}=1, we find that it suffices to prove

for all s,t∈Is,t\in\mathcal{I}. In components, this reads

for i,j∈IMi,j\in\mathcal{I}_{M} and μ,ν∈IN\mu,\nu\in\mathcal{I}_{N}.

The estimate (6.4) is trivial by assumption. What remains is the proof of (6.2) and (6.3). It is based on the polynomialization method developed in (BEKYY, , Section 5). The argument is very similar to that of BEKYY , and we only outline the differences.

Let us begin with (6.2). By the assumption ∣Gkk−mk∣≺Ψσk2\lvert G_{kk}-m_{k}\rvert\prec\Psi\sigma_{k}^{2} and orthogonality of UU, we have

for T⊂IMT\subset\mathcal{I}_{M} (which follows from (4.7)), (4.6), and

as follows from (4.6), (5.4), and (XG(iT)X∗)ii−m=O≺(σi2Ψ)(XG^{(iT)}X^{*})_{ii}-m=O_{\prec}(\sigma_{i}^{2}\Psi) (which may itself be deduced from (4.6)). We omit further details.

The rest of the argument is the same as before. ∎

Self-consistent comparison I: the main argument

In this section we establish Theorem 3.21 (i) under the additional assumption that the third moment of all entries of XX is zero.

Suppose that the assumptions of Theorem 3.21 hold. Suppose moreover that XX satisfies the additional condition

The rest of this section is devoted to the proof of Proposition 7.1. This is the heart of the proof of the anisotropic local law.

Computing the derivatives on the left-hand side of (7.3) leads to a product of terms of the form

and their complex conjugates (here we omit the argument XθX^{\theta}). Terms of type (i) are simply kept as they are; they will be put into the second term on the right-hand side of (7.3). Terms of type (ii) are the key to the gain that allows us to compensate losses from several other terms, such as the terms of type (iii) (see blow). The gain is obtained in combination with the summation over ii and μ\mu on the left-hand side of (7.3), according to estimates of the form

which allows us to estimate the right-hand side of (7.5) by NCδΨ2N^{C\delta}\Psi^{2}. The a priori bound (7.6) will be obtained from a bootstrapping, as explained below.

The preceding discussion was rather cavalier with the various constants CC in the factors NCδN^{C\delta}. In fact, some care has to be taken to ensure that the factors of NCδN^{C\delta} arising from the use of the a priori bounds (7.6) and (7.7) are compensated by a sufficiently large number of factors of the form (7.5). Moreover, in order to be able to iterate this argument from large to small scales, as explained in the next paragraph, we need an explicit bound of the constant CC in (7.3) in terms of the constant CC in (7.6).

Since δ>0\delta>0 is fixed, the bootstrapping consists of O(δ−1)O(\delta^{-1}) steps. The resulting estimates contain an extra factor NCδN^{C\delta}. Since δ>0\delta>0 can be made arbitrarily small, the claim will follow on all scales.

2. Bootstrapping on the spectral scale

The proof consists of a bootstrap argument from larger scales to smaller scales in multiplicative increments of N−δN^{-\delta}. Here

is fixed, where C0>0C_{0}>0 is a universal constant that will be chosen large enough in the proof (C0=25C_{0}=25 will work). For any η⩾N−1\eta\geqslant N^{-1} we define η0⩽η1⩽⋯⩽ηL\eta_{0}\leqslant\eta_{1}\leqslant\dots\leqslant\eta_{L}, where

The bootstrapping is started by the following result.

What remains is the proof of Lemma 7.4. We shall estimate the random variable

3. Interpolation

We use the interpolation outlined in Section 7.1.

Introduce the notations X0\vbox..=GGaussX^{0}\mathrel{\vbox{\hbox{.}\hbox{.}}}=G^{\text{\rm{Gauss}}} and X1\vbox..=XX^{1}\mathrel{\vbox{\hbox{.}\hbox{.}}}=X. For u∈{0,1}u\in\{0,1\}, i∈IMi\in\mathcal{I}_{M}, and μ∈IN\mu\in\mathcal{I}_{N}, denote by ρiμu\rho^{u}_{i\mu} the law of XiμuX^{u}_{i\mu}. For θ∈\theta\in we define the law

Let (X0,Xθ,X1)(X^{0},X^{\theta},X^{1}) be a triple of independent IM×IN\mathcal{I}_{M}\times\mathcal{I}_{N} random matrices, where for u∈{0,θ,1}u\in\{0,\theta,1\} the matrix Xu=(Xiμu)X^{u}=(X^{u}_{i\mu}) has law

We shall prove Lemma 7.5 by interpolation between the ensembles X0X^{0} and X1X^{1}. The bound for the Gaussian case X0X^{0} is given by the following result.

Lemma 7.5 holds if XX is replaced with X0X^{0}.

The basic interpolation formula is given by the following lemma, which follows from the fundamental theorem of calculus.

In order to prove Lemma 7.10, we compare the ensembles X(iμ)θ,Xiμ0X^{\theta,X^{0}_{i\mu}}_{(i\mu)} and X(iμ)θ,Xiμ1X^{\theta,X^{1}_{i\mu}}_{(i\mu)} via X(iμ)θ,0X^{\theta,0}_{(i\mu)}. Clearly, it suffices to prove the following result.

where we recall the definitions of L(η)L(\eta) and ηl\eta_{l} from (7.10) and (7.11).

It suffices to estimate the first term. Setting η−1\vbox..=0\eta_{-1}\mathrel{\vbox{\hbox{.}\hbox{.}}}=0 and ηL+1\vbox..=∞\eta_{L+1}\mathrel{\vbox{\hbox{.}\hbox{.}}}=\infty, we define the subsets of indices

and treat each ll separately. For l=1,…,Ll=1,\dots,L we find

where in the second step we used that λk≺1\lambda_{k}\prec 1, as follows from (2.7) and Lemma 4.8. This concludes the proof. ∎

where in the second step we used (3.20) and the definition of Π\Pi from (3.5). The estimate (7.19) now follows easily using Lemma 7.14.

4. Expansion

Let i∈IMi\in\mathcal{I}_{M} and μ∈IN\mu\in\mathcal{I}_{N}. Define the I×I\mathcal{I}\times\mathcal{I} matrix Δ(iμ)λ\Delta^{\lambda}_{(i\mu)} through

The following result provides a priori bounds for the entries of G(iμ)θ,λG^{\theta,\lambda}_{(i\mu)}.

Suppose that yy is a random variable satisfying ∣y∣≺N−1/2\lvert y\rvert\prec N^{-1/2}. Then

for all i∈IMi\in\mathcal{I}_{M} and μ∈IN\mu\in\mathcal{I}_{N}.

We use (7.21) with K\vbox..=10K\mathrel{\vbox{\hbox{.}\hbox{.}}}=10, λ′\vbox..=y\lambda^{\prime}\mathrel{\vbox{\hbox{.}\hbox{.}}}=y, and λ\vbox..=Xiμθ\lambda\mathrel{\vbox{\hbox{.}\hbox{.}}}=X_{i\mu}^{\theta}, so that G(iμ)θ,λ=GθG_{(i\mu)}^{\theta,\lambda}=G^{\theta}. By (3.20), we have

To simplify notation, we introduce the function

The following result is easy to deduce from (7.21) and Lemma 7.14.

where we used that Xiμ1X_{i\mu}^{1} has vanishing first and third moments, by (7.1), and its variance is equal to 1/N1/N. Recalling our goal (7.18), we therefore find that we only have to prove

for n=4,…,4pn=4,\dots,4p. Here we used the bounds (2.5).

holds for n=4,…,4pn=4,\dots,4p. Then (7.27) holds for n=4,…,4pn=4,\dots,4p.

To simplify notation, we abbreviate f(iμ)≡ff_{(i\mu)}\equiv f and Xiμθ≡ξX^{\theta}_{i\mu}\equiv\xi. The proof consists of a repeated application of the identity

for l⩽4pl\leqslant 4p, which follows from Lemmas 7.7 and 7.15. Fix n=4,…,4pn=4,\dots,4p. Using (7.29) we get

The claim now follows easily using (2.5). ∎

What therefore remains is to prove (7.28). Since it only involves the matrix ensemble XθX^{\theta}, for the remainder of the proof we abbreviate Xθ≡XX^{\theta}\equiv X. Recalling the notation (7.24), we find from Lemma 7.16 that it suffices to prove the following result.

We conclude this subsection with a general digression on the algebra of the Bernoulli interpolation method underlying the proof of Proposition 7.1, and in particular establish the formula (1.8) from the introduction. Let (Xα0)α(X_{\alpha}^{0})_{\alpha} and (Xα1)α(X_{\alpha}^{1})_{\alpha} be arbitrary independent finite families of independent random variables. Define Xθ=(Xαθ)αX^{\theta}=(X^{\theta}_{\alpha})_{\alpha} as in (1.6). Then we get

where X(α)θ,xX^{\theta,x}_{(\alpha)} denotes the family obtained from XθX^{\theta} by replacing XαθX_{\alpha}^{\theta} with xx. Fix α\alpha and abbreviate f(x)\mathrel{\vbox{\hbox{.}\hbox{.}}}=F\bigl{(}{X^{\theta,x}_{(\alpha)}}\bigr{)}, ζ\vbox..=Xα1\zeta\mathrel{\vbox{\hbox{.}\hbox{.}}}=X_{\alpha}^{1}, ζ′\vbox..=Xα0\zeta^{\prime}\mathrel{\vbox{\hbox{.}\hbox{.}}}=X_{\alpha}^{0}, and ξ\vbox..=Xαθ\xi\mathrel{\vbox{\hbox{.}\hbox{.}}}=X_{\alpha}^{\theta}. We may assume that ζ\zeta, ζ′\zeta^{\prime}, and ξ\xi are independent. We want to compute the difference

Since we are only interested in the algebra of the interpolation, we assume for simplicity that all random variables have finite exponential moments and that ff is analytic. (Otherwise, as in the computations above, the expansions have to be truncated.) Repeating the steps of the proof of Lemma 7.16, we find

5. Introduction of words and conclusion of the proof

In order to prove (7.30), we shall have to exploit the detailed structure of the derivatives in on the left-hand side of (7.30). The following definition introduces the basic algebraic objects that we shall use.

Next, we assign to each letter ∗* its value [∗]≡[∗]i,μ∈I[*]\equiv[*]_{i,\mu}\in\mathcal{I} through

(Our choice of the names of the two letters is suggestive of their value. Note, however, that it is important to distinguish the abstract letter from its value, which is an index in I\mathcal{I} and may be used as a summation index.)

If n(w)⩾1n(w)\geqslant 1 with ww as in (7.31) we define

for any n=0,1,2,…n=0,1,2,\dots. This may be easily deduced from (7.21). Using Leibnitz’s rule we conclude that

To prove (7.30), it therefore suffices to prove that

for 4⩽n⩽4p4\leqslant n\leqslant 4p and words wr∈Ww_{r}\in\mathcal{W} satisfying ∑rn(wr)=n\sum_{r}n(w_{r})=n. (The proof of (7.33) is the same with slightly heaver notation.) Treating words wrw_{r} with n(wr)=0n(w_{r})=0 separately, we find that it suffices to prove

for 4⩽n⩽4p4\leqslant n\leqslant 4p, 1⩽q⩽p1\leqslant q\leqslant p, and words wr∈Ww_{r}\in\mathcal{W} satisfying ∑rn(wr)=n\sum_{r}n(w_{r})=n, n(w0)=0n(w_{0})=0, and n(wr)⩾1n(w_{r})\geqslant 1 for r⩾1r\geqslant 1. Note that we also have the bound q⩽nq\leqslant n.

In order to estimate (7.35), we introduce the quantity

For w∈Ww\in\mathcal{W} we have the rough bound

Finally, for n(w)=1n(w)=1 we have the sharper bound

In addition to the high-probability bounds from Lemma 7.20, we have the rough estimate

for some C>0C>0 and all w∈Ww\in\mathcal{W} satisfying n(w)⩽4pn(w)\leqslant 4p; this follows easily from Definition 7.19 and Lemma 4.9.

By pigeonholing on the words appearing on the left-hand side of (7.35), if n⩽2q−2n\leqslant 2q-2 then there exist at least two words wrw_{r} satisfying n(wr)=1n(w_{r})=1. Using Lemma 7.20 we therefore get

From Lemma 4.6 combined with Lemma 4.8 we get

where in the second step we used (7.20), and in the last step the definition of Ψ\Psi. Inserting (7.41) into (7.40), we get using Lemma 7.7 and (7.39) that the left-hand side of (7.35) is bounded by

Using that Ψ⩾cN−1/2\Psi\geqslant cN^{-1/2}, we find that the left-hand side of (7.35) is bounded by

where we used that q⩽nq\leqslant n and n⩾4n\geqslant 4. Choose C0⩾25C_{0}\geqslant 25. Then, by assumption (7.9) on δ\delta we have N(C0/2+12)δΨ⩽1N^{(C_{0}/2+12)\delta}\Psi\leqslant 1. Moreover, if n⩾4n\geqslant 4 and n⩾2q−1n\geqslant 2q-1 then n⩾q+2n\geqslant q+2. We conclude that the left-hand side of (7.35) is bounded by

Now (7.35) follows using Young’s inequality. This concludes the proof of (7.30), and hence of Lemma 7.11. The proof of Proposition 7.1 is therefore complete.

Self-consistent comparison II: general X𝑋X

In this section and the next we prove the following result, which is Proposition 7.1 without the condition (7.1).

The proof of Proposition 8.1 builds on that of Proposition 7.1. Throughout this section, we take over the notations of Section 7 without further comment. In particular, as in (7.9), C0C_{0} is some large enough constant (which may depend only on τ\tau) and δ>0\delta>0 is fixed and small, satisfying (7.9).

The assumption (7.1) was used in the proof of Proposition 7.1 only in (7.26), where it ensured that the summation over nn starts from 44 instead of 33. Without the assumption (7.1), we in addition have to estimate the term n=3n=3 in (7.26). It therefore suffices to prove the following result.

As in Lemma 7.16, one may easily replace the matrix Xiμθ,0X^{\theta,0}_{i\mu} in the definition of fiμ(3)(0)f^{(3)}_{i\mu}(0) with XθX^{\theta}. As in Lemma 7.17, from now on we abbreviate Xθ≡XX^{\theta}\equiv X. Thus, we find that in order to prove Lemma 8.2 it suffices to prove the following result, which complements Lemma 7.17.

Next, recall the definition of words from Definition 7.19. As in (7.33)–(7.35), Lemma 8.3 is proved provided we can show the following result, which is analogous to (7.35).

Beyond the factor of N−1/2N^{-1/2}, the need to obtain the bound from (7.34) that is strong enough to close the self-consistent estimate presents significant difficulties, which we outline briefly. These difficulties are roughly of two types. (a) The off-diagonal entries of G(μ)G^{(\mu)} are in general not small; only entries of G(μ)−ΠG^{(\mu)}-\Pi are small. (b) A priori, using the estimate (7.19), the entries of G(μ)G^{(\mu)} are not bounded by O≺(1)O_{\prec}(1) but by O≺(N2δ)O_{\prec}(N^{2\delta}). This would yield a bound on the coefficients of the polynomial in XμX_{\mu} that is a high power of N2δN^{2\delta}, which may become too large to conclude the proof.

2. Tagged words

Lemma 8.4 holds provided it holds with AA replaced by A^\widehat{A} in (7.34).

Next, from Definition 7.19 we conclude, analogously to the proof of Lemma 7.20, that if n(wr)⩾1n(w_{r})\geqslant 1 for 1⩽r⩽q1\leqslant r\leqslant q and ∑r=1qn(wr)=3\sum_{r=1}^{q}n(w_{r})=3 then

where in the last step we used (7.41). We therefore conclude using Lemma 7.7 that

If n=0n=0 (i.e. ww is the empty word) we define

where we introduced the matrix \Pi^{(\mu)}\mathrel{\vbox{\hbox{.}\hbox{.}}}=\bigl{(}{\Pi_{st}\mathrel{\vbox{\hbox{.}\hbox{.}}}s,t\in\mathcal{I}\setminus\{\mu\}}\bigr{)}.

From (8.4), (8.3), and Definition 8.7 we find that for w∈Wnw\in\mathcal{W}_{n} we have

Recalling Lemma 8.5, we conclude that, in order to prove Lemma 8.4, it suffices to prove the following result.

We shall require the following rough bounds, which refine those of Lemma 7.20.

Suppose that (8.1) holds. Let (w,σ)(w,\sigma) be a tagged word. Then

Using a large deviation estimate (see (BEKYY, , Lemma 3.1)) combined with Lemmas 4.6 and 4.8 we get

where in the second step we used (8.1). Therefore,

for all σ=0,1,2\sigma=0,1,2. This concludes the proof. ∎

Using the large deviation estimates from (BEKYY, , Lemma 3.1), we find ∣Zμ∣≺N−τ/2+2δ\lvert Z_{\mu}\rvert\prec N^{-\tau/2+2\delta}. Since ∣Gμμ∣≺N2δ\lvert G_{\mu\mu}\rvert\prec N^{2\delta} by (8.1), we therefore deduce using −τ/2+2δ<−2δ-\tau/2+2\delta<-2\delta that

Hence there exists a constant K≡K(τ)K\equiv K(\tau) such that

Plugging in the definition of ZμZ_{\mu}, we find

Now we replace the factors GμμG_{\mu\mu} on the left-hand side of (8.7) with the leading term in (8.15). From (8.15), (7.19), (8.16), and Lemma 7.7 we get

where the coefficients Zμ,k\mathcal{Z}_{\mu,k} are X(μ)X^{(\mu)}-measurable and satisfy the bound

Moreover, for any tagged word (w,σ)(w,\sigma) we split

4. Degree counting

For the following we choose qq, w1,…,wpw_{1},\dots,w_{p}, and σ1,…,σr\sigma_{1},\dots,\sigma_{r} as in Lemma 8.8. We introduce the abbreviation

Here the constant CpC_{p} accounts for the immaterial constants depending on pp arising from the combinatorics of all partitions. For LL, {dl}\{d_{l}\}, and {kl}\{k_{l}\} as above, we may estimate

where we used (2.5). Now we use the estimate

which follows from (8.11), (8.12), and (8.1). Hence we get

where the maxima are subject to the same conditions as above.

Next, from ∑l(dl+kl)=2k+d\sum_{l}(d_{l}+k_{l})=2k+d and dl+kl⩾2d_{l}+k_{l}\geqslant 2 we deduce that

Together with ∑l[dl−2]+⩽∑ldl=d\sum_{l}[d_{l}-2]_{+}\leqslant\sum_{l}d_{l}=d, this gives

where in the last step we used that [dl+kl−2]+⩾[dl−2]+[d_{l}+k_{l}-2]_{+}\geqslant[d_{l}-2]_{+}, and ∑l(2∧dl)+∑l[dl−2]+=d\sum_{l}(2\wedge d_{l})+\sum_{l}[d_{l}-2]_{+}=d. For the case d=1d=1, we also used that ∑l(2∧dl)+∑l[dl+kl−2]+−1⩾1\sum_{l}(2\wedge d_{l})+\sum_{l}[d_{l}+k_{l}-2]_{+}-1\geqslant 1.

We conclude that the left-hand side of (8.21) is bounded by

Plugging (8.33) and (8.34) into (8.32), and recalling Lemma 7.7, yields

5. Power counting: proof of (8.36)

as may be easily checked from Definition 8.7. (Recall that n(wr)=0n(w_{r})=0 for r⩾q+1r\geqslant q+1.) We conclude that (8.36) holds provided we can show that

The proof of (8.21), and hence of Lemma 8.8, is therefore complete. This concludes the proof of Lemma 8.4.

We have hence proved Proposition 8.1. Recalling Proposition 6.1, we conclude the proof of Theorem 3.21 (i).

Averaged local law

for some constant c>0c>0. In particular, we recover Definition 3.20 (iii) by setting Φ=(Nη)−1\Phi=(N\eta)^{-1}. In this section we prove the following result, which also completes the proof of Theorem 3.21.

The proof is similar to that of Propositions 7.1 and 8.1, and we only explain the differences. Note that now there is no bootstrapping, since the necessary a priori bounds are obtained from the anisotropic local law. In analogy to (7.15), we define

Following the argument leading up to (7.30) and Lemma 8.3 to the letter, we find that it suffices to prove that

For n⩾4n\geqslant 4, the claim (9.2) easily follows from (9.3) and an application of Young’s inequality to the factors rr satisfying n(wr)=0n(w_{r})=0 and n(wr)⩾1n(w_{r})\geqslant 1; we omit the details. What remains, therefore, is to verify (9.2) for n=3n=3.

which may be regarded as an improvement of Lemma 8.11. The proof of (9.5) is similar to that of Lemma 8.11, and we merely give a sketch. We have to estimate an expression of the form

where sl∈{i,ν1,…,νp}s_{l}\in\{i,\nu_{1},\dots,\nu_{p}\} and bi∈{0,1}b_{i}\in\{0,1\}. (This is simply the general form of a polynomial in XX of the correct degree.) The proof of (9.5) relies on the crucial observations that, by Theorem 3.21 (i),

Next, let r⩾q+1r\geqslant q+1 satisfy σr=0\sigma_{r}=0, so that the left-hand side of (9.5) does not depend on νr\nu_{r}. There are p−q−∑s=q+1p∣σs∣p-q-\sum_{s=q+1}^{p}\lvert\sigma_{s}\rvert such indices rr, and for each such rr we find

As in the argument following Lemma 8.11, we conclude that the claim holds provided that

This concludes the proof of Theorem 3.21. We conclude this section by drawing consequences from Theorems 3.21 and 3.22.

Similarly, Theorems 3.15 (i) and 3.16 (i) follow from Lemmas A.6 and A.8 respectively, combined with Theorems 3.21 and 3.22 ∎

These follow immediately from Theorem 3.6 and Theorems 3.14 (i), 3.15 (i), and 3.16 (i) respectively, combined with (3.3), (3.4), and (3.5). ∎

How to remove the assumption (2.9) in the above proof is explained in Section 11.1.

Eigenvalue rigidity and edge universality

In the first part of this section we establish eigenvalue rigidity (Theorems 3.12, 3.14 (iii), and 3.15 (iii)). As a consequence of Theorem 3.6 we prove Theorem 3.7. In the second part of this section we establish the edge universality from Theorem 3.18. We assume (2.9) throughout this section; how to remove it is explained in Section 11.1 below.

For simplicity, we first prove Theorem 3.12, i.e. we assume that all edges and bulk components are regular. The proof consists of three steps. First, we prove that with high probability there are no eigenvalues at a distance greater than N−2/3+εN^{-2/3+\varepsilon} from the support of ϱ\varrho. Second, we prove that a neighbourhood of the kk-th bulk component of ϱ\varrho contains with high probability exactly NkN_{k} eigenvalues (recall the definition (3.16)). Third, we use the averaged local law from Theorem 3.6 together with the first two steps to complete the proof.

We begin with the first step. We define the distance to the spectral edge through

Fix ε∈(0,2/3)\varepsilon\in(0,2/3). Under (2.9) and the assumptions of Theorem 3.12 we have

with high probability (recall Definition 4.7).

The argument is similar to that of the previous works EYY3 ; PY , and we only explain how to adapt it. From (BEKYY, , Theorem 2.10), we find that ∥X∗X∥⩽C0\lVert X^{*}X\rVert\leqslant C_{0} with high probability for some large enough constant C0>0C_{0}>0 depending on τ\tau. From (2.7) we therefore deduce that λ1⩽C\lambda_{1}\leqslant C with high probability for some constant C>0C>0 depending on τ\tau. It therefore suffices to prove that, after a possible decrease of τ\tau,

In order to prove (10.2), it suffices to prove that Im⁡mN≺N−ε/2(Nη)−1\operatorname{Im}m_{N}\prec N^{-\varepsilon/2}(N\eta)^{-1}. By (A.7) and Lemmas A.4, A.6, and A.8, we have

It is easy to check that the control parameter on the right-hand side of (10.4) satisfies (9.1), so that by Proposition 9.1 and Theorem 3.6 it suffices to prove (10.4) for diagonal Σ\Sigma. The proof is identical to that of Section 5, except that we use the strong stability of (2.11) from Definition A.2, which is established in Lemmas A.5, A.6, and A.8. This concludes the proof of (10.4), and hence of (10.2).

With applications to the deformed Wigner matrices in Section 12 in mind, we give another proof of (10.2), which is based on (BS2, , Section 6). Using a partial fraction decomposition, one easily finds that there exist universal constants Cn,kC_{n,k} such that

Similarly, we have (with the same Cn,kC_{n,k} and zkz_{k})

for all k=1,…,nk=1,\dots,n. Setting n\vbox..=⌈2/ε⌉n\mathrel{\vbox{\hbox{.}\hbox{.}}}=\lceil{2/\varepsilon}\rceil, we get from (10.5) and (10.6) that

where in the last step we used that ∣x−E∣⩾κ\lvert x-E\rvert\geqslant\kappa for x∈supp⁡ϱx\in\operatorname{supp}\varrho. This immediately implies that with high probability there is no eigenvalue in [E−η,E+η][E-\eta,E+\eta]. ∎

The second step represents most of the work. It is a counting argument, based on a continuous deformation of the matrix QQ to another matrix for which the claim is obvious. Since the eigenvalues depend continuously on the deformation parameter and each intermediate matrix satisfies a gap condition from Lemma 10.1, we shall be able to conclude that the number of eigenvalues in a neighbourhood of the ii-th component does not change under the deformation. We shall in fact need two deformations: one which deforms the original matrix QQ to a Gaussian one, QGaussQ^{\text{\rm{Gauss}}}, with the same expectation Σ\Sigma as QQ, and another which deforms the Gaussian matrix QGaussQ^{\text{\rm{Gauss}}} to another Gaussian matrix where some eigenvalues of Σ\Sigma have been increased.

For k=1,…,p−1k=1,\dots,p-1, we introduce the number of eigenvalues to the right of the kk-th gap,

Under (2.9) and the assumptions of Theorem 3.12 we have Υk=∑l⩽kNl\Upsilon_{k}=\sum_{l\leqslant k}N_{l} with high probability.

As explained above, the first step in the proof of Proposition 10.2 is a deformation of the general matrix XX to a Gaussian one.

Let XX be general and XGaussX^{\text{\rm{Gauss}}} a Gaussian matrix of the same dimensions. Suppose that (2.9) and the assumptions of Theorem 3.12 hold. If Υk=∑l⩽kNl\Upsilon_{k}=\sum_{l\leqslant k}N_{l} with high probability under the law of XGaussX^{\text{\rm{Gauss}}}, then Υk=∑l⩽kNl\Upsilon_{k}=\sum_{l\leqslant k}N_{l} with high probability under the law of XX.

Let X1\vbox..=XX_{1}\mathrel{\vbox{\hbox{.}\hbox{.}}}=X and X0\vbox..=XGaussX_{0}\mathrel{\vbox{\hbox{.}\hbox{.}}}=X^{\text{\rm{Gauss}}} be independent. For t∈t\in define X(t)\vbox..=tX1+1−tX0X(t)\mathrel{\vbox{\hbox{.}\hbox{.}}}=\sqrt{t}X_{1}+\sqrt{1-t}X_{0} and denote by λi(t)\lambda_{i}(t) the eigenvalues of X(t)ΣX(t)∗X(t)\Sigma X(t)^{*}. We also write Υk(t)\Upsilon_{k}(t) for Υk\Upsilon_{k} defined in terms of λi(t)\lambda_{i}(t). Note that Lemmas 4.8 and 10.1 hold for X(t)X(t) uniformly in t∈t\in. Recalling (2.7), we deduce that there exists a constant C>0C>0 such that

with high probability, uniformly in t,s∈t,s\in and ii. The claim now follows by considering tt in the lattice 0,1/K,2/K,…,10,1/K,2/K,\dots,1, where KK is chosen large enough that CK−1/2⩽(a2k−a2k+1)/10CK^{-1/2}\leqslant(a_{2k}-a_{2k+1})/10 and CC is the constant from (10.8). Indeed, suppose Υk(t)=∑l⩽kNl\Upsilon_{k}(t)=\sum_{l\leqslant k}N_{l} with high probability. Then we use (10.8) and the gap from Lemma 10.1 to deduce that Υk(t+K−1)=∑l⩽kNl\Upsilon_{k}(t+K^{-1})=\sum_{l\leqslant k}N_{l} with high probability. The claim for t=1=K/Kt=1=K/K therefore follows by induction. ∎

Using Lemma 10.3, in order to prove Proposition 10.2 it suffices to prove the following result.

Suppose that (2.9) and the assumptions of Theorem 3.12 hold, that XX is Gaussian, and that Σ\Sigma is diagonal. Then Υk=∑l⩽kNl\Upsilon_{k}=\sum_{l\leqslant k}N_{l} with high probability.

Abbreviate dk\vbox..=∑l⩽kNld_{k}\mathrel{\vbox{\hbox{.}\hbox{.}}}=\sum_{l\leqslant k}N_{l}. We write the diagonal matrix Σ=diag⁡(σ1,σ2,…,σM)\Sigma=\operatorname{diag}(\sigma_{1},\sigma_{2},\dots,\sigma_{M}) in the block form Σ=diag⁡(Σ1,Σ2)\Sigma=\operatorname{diag}(\Sigma_{1},\Sigma_{2}), where Σ1\Sigma_{1} contains the dkd_{k} top eigenvalues of Σ\Sigma. We introduce the deformed covariance matrix Σ(t)\vbox..=diag⁡(tΣ1,Σ2)\Sigma(t)\mathrel{\vbox{\hbox{.}\hbox{.}}}=\operatorname{diag}(t\Sigma_{1},\Sigma_{2}). In particular, Σ(1)=Σ\Sigma(1)=\Sigma. The idea is to increase tt until the claim for Σ\Sigma replaced by Σ(t)\Sigma(t) may be deduced from simple linear algebra. Then we use a continuity argument to compare Σ(t)\Sigma(t) to Σ\Sigma.

We add the argument tt to quantities to indicate that they are defined in terms of Σ(t)\Sigma(t) instead of Σ\Sigma. Using (A.5) (see also Figure 2.1), it is not hard to find that the gap a2k(t)−a2k+1(t)a_{2k}(t)-a_{2k+1}(t) is increasing in tt, and that

for some constants c,C>0c,C>0. Using that ∥Σ(t)−Σ(s)∥⩽C∣t−s∣\lVert\Sigma(t)-\Sigma(s)\rVert\leqslant C\lvert t-s\rvert, we deduce from Lemma 4.8 that ∥Q(t)−Q(s)∥⩽C∣t−s∣\lVert Q(t)-Q(s)\rVert\leqslant C\lvert t-s\rvert with high probability. Now let TT be a fixed large time to be chosen later. By a continuity argument using Lemma 10.1 we therefore find that there exists a constant K≡K(T)K\equiv K(T) such that for any t∈[1,T]t\in[1,T] we have

It therefore remains to show that there exists a large enough TT, depending only on τ\tau, such that for t=Tt=T the left-hand side of (10.10) is equal to dkd_{k}. Writing XX∗=(E11E12E21E22)XX^{*}=\begin{pmatrix}E_{11}&E_{12}\\ E_{21}&E_{22}\end{pmatrix} as a block matrix, we find

From (A.4) we deduce that dk⩽N−Np⩽(1−c)Nd_{k}\leqslant N-N_{p}\leqslant(1-c)N, by the assumption of Theorem 3.12. From (BEKYY, , Theorem 2.10), we therefore deduce that c⩽E11⩽Cc\leqslant E_{11}\leqslant C with high probability for some positive constants c,Cc,C. Moreover, from Lemma 4.8 we deduce that ∥E12∥+∥E21∥+∥E22∥⩽C\lVert E_{12}\rVert+\lVert E_{21}\rVert+\lVert E_{22}\rVert\leqslant C with high probability. We conclude that for large enough tt (depending on the constants cc and CC above) the matrix (10.11) has with high probability exactly dkd_{k} eigenvalues of order ≍t\asymp t, and all other eigenvalues are of order O(t)O(\sqrt{t}). Since \operatorname{spec}(Q(t))\cap\bigl{[}{a_{2k+1}(t)+N^{-2/3+\tau},a_{2k}-N^{-2/3+\tau}}\bigr{]}=\emptyset with high probability by Lemma 10.1, we conclude that for large enough tt the left-hand side of (10.10) is equal to dkd_{k} with high probability. This concludes the proof. ∎

Proposition 10.2 follows immediately from Lemmas 10.3 and 10.4. This concludes the second step outlined above.

Finally, the third step – the conclusion of the proof of Theorem 3.12 – follows from Theorem 3.6, Lemma 10.1, and Proposition 10.2 by repeating the analysis of EYY3 ; PY with merely cosmetic changes, as explained in the proof of Theorem 3.12. This concludes the proof of Theorem 3.12 under the assumption (2.9). Moreover, as explained in Section 11.1, the assumption (2.9) may be easily removed.

Next, we note that the proof Theorem 3.7 under the assumption (2.9) is an easy consequence of Theorems 3.6 and 3.12, following the proof of (BEKYY, , Theorem 3.12). The key tool is the spectral decomposition (4.15). We omit further details.

Finally, the proof of Theorems 3.14 (iii) and 3.15 (iii) under the assumption (2.9) is exactly the same as that of Theorem 3.12. Indeed, one can check that in the proof of Theorem 3.12 only the weaker assumptions of Theorems 3.14 (iii) and 3.15 (iii) are needed. We omit the details.

2. Edge universality assuming (2.9)

Finally, we prove the edge universality from Theorem 3.18 under the assumption (2.9).

We may therefore without loss of generality assume that XX is Gaussian. By orthogonal / unitary invariance of the law of XX, we may furthermore assume that Σ\Sigma is diagonal. This concludes the proof. ∎

General matrices: removing (2.9) and extension to Q˙˙𝑄\dot{Q}

In this section we explain how our results, proved under the assumption (2.9) and for the matrix QQ, may be generalized to hold without the assumption (2.9) and for the matrix Q˙\dot{Q} from (1.4) as well.

For simplicity, throughout the proofs up to now we made the assumption (2.9). As advertised, this assumption is not necessary. In this section we explain how to dispense with it. The argument relies on simple approximation and linear algebra. Roughly, if Σ\Sigma has a zero eigenvalue, we consider Σ+ε\Sigma+\varepsilon instead and let ε↓0\varepsilon\downarrow 0; if TT is not square, we augment it to a square matrix by padding it out with zeros. While this extension is simple, we emphasize that it relies crucially on the fact we do not assume that Σ\Sigma has a lower bound (the assumption (2.9) only requires the qualitative bound Σ>0\Sigma>0).

We distinguish the cases M^⩾M\widehat{M}\geqslant M and M^<M\widehat{M}<M. Suppose first that M^⩾M\widehat{M}\geqslant M. We extend TT to an M^×M^\widehat{M}\times\widehat{M} matrix by setting T^\vbox..=(0T)\widehat{T}\mathrel{\vbox{\hbox{.}\hbox{.}}}=\binom{0}{T}. Define the M^×M^\widehat{M}\times\widehat{M} matrices

By polar decomposition, we have T^=U^S^1/2\widehat{T}=\widehat{U}\widehat{S}^{1/2}, where U^\widehat{U} is orthogonal. Therefore Σ^=U^S^U^∗\widehat{\Sigma}=\widehat{U}\widehat{S}\widehat{U}^{*}. Moreover, from (2.11) we get m≡mΣ,N=mΣ^,N=mS^,Nm\equiv m_{\Sigma,N}=m_{\widehat{\Sigma},N}=m_{\widehat{S},N}, which we use tacitly in the following.

We define the (M^+N)×(M^+N)(\widehat{M}+N)\times(\widehat{M}+N) matrix

where we used (4.3). Note that G^\widehat{G} has the block form

Using this representation of GG as a block of G^\widehat{G}, it is easy to drop the assumption (2.9). For example, suppose that Theorem 3.6 has been proved under the assumption (2.9). Applying it to G^\widehat{G} and using a simple approximation argument in ε\varepsilon, we find the following generalization of Theorem 3.6.

Fix τ>0\tau>0. Suppose that (2.9) and Assumption 2.1 hold. Suppose moreover that every edge k=1,…,2pk=1,\dots,2p satisfying ak⩾τa_{k}\geqslant\tau and every bulk component k=1,…,pk=1,\dots,p is regular in the sense of Definition 2.7. Then (3.10) and (3.11) hold for GG defined in (11.1) and RNR_{N} defined in (3.1).

In particular, Corollary 3.9 and Theorems 3.12, 3.14, 3.15, 3.16, and 3.18 follow easily from their counterparts proved under the assumption (2.9). This concludes the discussion for the case M^⩾M\widehat{M}\geqslant M.

Finally, we consider the case M^<M\widehat{M}<M. We set T~\vbox..=(T,0)\widetilde{T}\mathrel{\vbox{\hbox{.}\hbox{.}}}=(T,0) and X~\vbox..=(XY)\widetilde{X}\mathrel{\vbox{\hbox{.}\hbox{.}}}=\binom{X}{Y}, where YY is an (M−M^)×N(M-\widehat{M})\times N matrix, independent of XX, with independent entries satisfying (2.4) and (2.5). Hence, T~\widetilde{T} is M×MM\times M and X~\widetilde{X} is M×NM\times N. Now we have TX=T~X~TX=\widetilde{T}\widetilde{X}, and we have reduced the problem to the case M^=M\widehat{M}=M, which was dealt with above.

2. Extension to Q˙˙𝑄\dot{Q}

In this section we explain how our results have to be modified for Q˙\dot{Q} from (1.4). For simplicity of presentation, we make the assumption (2.9); it may be easily removed as explained in Section 11.1.

As before it is easy to obtain R˙M\dot{R}_{M} and R˙N\dot{R}_{N} from G˙\dot{G}.

The local law for G˙\dot{G} reads as follows.

To simplify notation, we omit the factor NN−1\frac{N}{N-1} in the definition of Q˙\dot{Q}. It may easily be put back by scaling the argument zz. We prove the anisotropic local law by estimating the four blocks of G˙\dot{G} individually. Using simple linear algebra and (4.1) and (4.3), we get

Since GG satisfies the anisotropic local law by assumption, it is easy to deduce from Lemma 4.4 and (4.21) that

From (4.3) we get X∗GMX=z2GN+zX^{*}G_{M}X=z^{2}G_{N}+z, so that, using the anisotropic local law for GG, we conclude

Finally, using (4.3) for G˙\dot{G} as well as (11.4), we get

Using the anisotropic law for GG and Lemma 4.4, we therefore get

We also obtain eigenvalue rigidity for the eigenvalues λ˙1⩾λ˙2⩾⋯⩾λ˙M\dot{\lambda}_{1}\geqslant\dot{\lambda}_{2}\geqslant\cdots\geqslant\dot{\lambda}_{M} of Q˙\dot{Q}. For instance, Theorem 3.12 has the following counterpart.

Theorem 3.12 remains valid if λk\lambda_{k} is replaced with λ˙k\dot{\lambda}_{k}.

The proof follows that of Theorem 3.12 to the letter, using Theorem 11.2 as input. In fact, as explained around (10.4), we need a stronger bound than (11.3) outside of the spectrum; this stronger bound follows easily from (11.5) and the analogous stronger bound for GNG_{N} established in (10.4). ∎

Finally, we obtain edge universality for Q˙\dot{Q}. The following result is proved exactly like Theorem 3.18, using Theorems 11.2 and 11.3 as input.

Theorem 3.18 remains valid if λi\lambda_{i} is replaced with λ˙i\dot{\lambda}_{i}.

Deformed Wigner matrices

In this section we apply our method to deformed Wigner matrices as a further illustration of its applicability. Since the statements and arguments are similar to those of the previous sections, we keep the presentation concise.

Let W=W∗W=W^{*} be an N×NN\times N Wigner matrix whose upper-triangular entries (Wij\vbox..1⩽i⩽j⩽N)(W_{ij}\mathrel{\vbox{\hbox{.}\hbox{.}}}1\leqslant i\leqslant j\leqslant N) are independent and satisfy the same conditions (2.4) and (2.5) as XiμX_{i\mu}. Let A=A∗A=A^{*} be a deterministic N×NN\times N matrix satisfying ∥A∥⩽τ−1\lVert A\rVert\leqslant\tau^{-1}. For definiteness, we suppose that WW and AA are real symmetric matrices, remarking that similar results also hold for complex Hermitian matrices.

The main result of this section is the anisotropic local law for the deformed Wigner matrix W+AW+A, analogous to Theorem 3.21. As an application, we establish the edge universality of W+AW+A. We remark that the entrywise local law and edge universality were previously established in LSY under the assumption that AA is diagonal. Before that, edge universality was established in CP1 ; Joh2 ; Shc under the assumption that WW is a GUE matrix. A somewhat different direction was pursued in BES15 ; BG15 ; Kar15 , where local laws for the sum of two matrices that are invariant under unitary or orthogonal conjugations were analysed.

We denote the eigenvalues of W+AW+A by λ1(W+A)⩾λ2(W+A)⩾⋯⩾λN(W+A)\lambda_{1}(W+A)\geqslant\lambda_{2}(W+A)\geqslant\cdots\geqslant\lambda_{N}(W+A). Moreover, we define the resolvent GW(z)\vbox..=(W+A−z)−1G^{W}(z)\mathrel{\vbox{\hbox{.}\hbox{.}}}=(W+A-z)^{-1}, as well as

The following definition is the analogue of Definition 3.20 for deformed Wigner matrices.

The following result is the analogue of Theorem 3.21. Throughout the following we denote by WGaussW^{\text{\rm{Gauss}}} a GOE matrix and by D≡DAD\equiv D_{A} the diagonalization of AA.

The following result guarantees the rigidity of the extreme eigenvalues of H+AH+A. As explained below, the rigidity of all eigenvalues will be a simple consequence of Theorems 12.2 (ii) and 12.4.

We say the extreme eigenvalues of W+AW+A are rigid if \bigl{[}{\lambda_{1}(W+A)-L_{+}}\bigr{]}_{+}\prec N^{-2/3} and \bigl{[}{\lambda_{N}(W+A)-L_{-}}\bigr{]}_{-}\prec N^{-2/3}.

Together, Theorems 12.2 (ii) and 12.4 easily yield the rigidity of the eigenvalues, as explained in the proof of (EYY3, , Theorem 2.2). We may apply Theorem 12.2 to extend the results of LSY to arbitrary non-diagonal AA. For instance, we obtain the following edge universality result.

A similar result holds for the extreme eigenvalues near the left edge.

In LSY , the assumptions of Theorem 12.4 were verified for a large class of diagonal matrices DD. Moreover, for such matrices DD, it was proved that the limit of the second term on the left-hand side of (12.1) is governed by the Tracy-Widom-Airy statistics. Theorem 12.5 therefore provides an extension of (LSY, , Theorem 2.8) to non-diagonal matrices AA. We refer to LSY for the detailed statements about the distribution of the eigenvalues of WGauss+DW^{\text{\rm{Gauss}}}+D.

Further applications of Theorem 12.2 include a study of the eigenvectors of W+AW+A and the outliers of finite-rank perturbations of W+AW+A. We do not pursue these questions here.

2. Proof of Theorem 12.2 (i)

Here C0C_{0} is a constant that may depend only on τ\tau, and δ\delta satisfies (7.9).

Making minor adjustments to the argument of Sections 7.3 and 7.4, we find that it suffices to prove the following result, which is analogous to Lemma 7.30.

for any n=4,5,…,4pn=4,5,\dots,4p. Moreover, if in addition Ψ(z)⩽N−1/4−c0\Psi(z)\leqslant N^{-1/4-c_{0}} then (12.5) holds also for n=3n=3.

The rest of this subsection is devoted to the proof of Lemma 12.7. We note that the word structure describing the derivatives on the left-hand side of (12.5) is very similar to that of (7.30) (given in Definition 7.19). Hence the proof of Lemma 12.7 is similar to that of Lemma 7.17.

The proof of (12.5) for n⩾4n\geqslant 4 is a trivial modification of the argument given in Section 7.5, whose details we omit. What remains therefore is the proof of (12.5) for n=3n=3 under the assumption Ψ⩽N−1/4−c0\Psi\leqslant N^{-1/4-c_{0}}. From now on we omit the arguments zz.

The main new ingredient of the proof for deformed Wigner matrices is a further iteration step at fixed zz. Suppose that

for some Φ⩽1\Phi\leqslant 1. Since ∥Π∥⩽C\lVert\Pi\rVert\leqslant C, the estimate (12.6) is stronger than the first estimate of (12.4). Note that, by assumption (12.4), the estimate (12.6) holds for Φ=1\Phi=1. Assuming (12.6), we shall prove a self-improving bound of the form

Once (12.7) is proved, we may use it iteratively to obtain increasingly accurate bounds for the left-hand side of (12.3). After each step, we obtain an improved a priori bound (12.6), whereby Φ\Phi is reduced by powers of N−c0/2N^{-c_{0}/2}. After O(1/c0)O(1/c_{0}) iterations, (12.5) for n=3n=3 follows.

It therefore suffices to prove (12.7) under the assumptions (12.4) and (12.6). As in Section 7.5, it suffices to prove

We discuss the three cases q=1,2,3q=1,2,3 separately.

or a term obtained from one of these two by exchanging ii and jj; here a,b,c,d∈{i,j}a,b,c,d\in\{i,j\}. We deal first with the first expression; the others are deal with analogously. We split it according to

We deal with the summation over ii using the estimate

where in the last step we used (12.4); a similar estimate holds for the summation over jj. Using (12.12), (12.6), and ∥Π∥⩽τ−1\lVert\Pi\rVert\leqslant\tau^{-1}, we may estimate the second term on the right-hand side of (12.11) as

provided δ\delta is chosen small enough, depending on τ\tau and c0c_{0}. The third and fourth terms on the right-hand side of (12.11) are estimated in exactly the same way.

The case q=2𝑞2q=2

or an expression obtained from one of these two by exchanging ii and jj. The contribution of the first expression of (12.13) is estimated using (12.4) and (12.12) by

provided δ\delta is chosen small enough, depending on τ\tau and c0c_{0}.

Next, in order to estimate the contribution of the second expression of (12.13), we split

where we used (12.4), (12.6), and (12.12). Similarly, using (12.4) and (12.12) we get

for small enough δ\delta. Putting (12.14) and (12.15) together, it is easy to deduce (12.9).

The case q=3𝑞3q=3

3. Proof of Theorem 12.2 (ii)

The proof is similar to that from Section 9. As in Section 12.2, it suffices to prove the following result.

for any n=4,5,…,4pn=4,5,\dots,4p. Moreover, if in addition Ψ(z)⩽N−1/4−c0\Psi(z)\leqslant N^{-1/4-c_{0}} then (12.16) holds also for n=3n=3.

The case n⩾4n\geqslant 4 can be easily proved as in covariance case. We therefore focus on the case n=3n=3 in Lemma 12.8. The proof is similar to the discussion below (12.8). The main difference is that for each qq we have some extra averaging N−q∑s1,…,sq( ⋅ )N^{-q}\sum_{s_{1},\dots,s_{q}}(\,\cdot\,), and we need to extract an extra factor Ψq\Psi^{q} (or, alternatively, (N−1−c0/2η−1Ψ−1)q(N^{-1-c_{0}/2}\eta^{-1}\Psi^{-1})^{q}) from this average. We take over the notations from Sections 7.5 and 12.2 without further comment. We consider the three cases q=1,2,3q=1,2,3 separately, and tacitly use the anisotropic local law from Theorem 12.2 (i).

Consider first the case As,s,i,j(w1)=GsiGjjGijGisA_{s,s,i,j}(w_{1})=G_{si}G_{jj}G_{ij}G_{is}. We estimate

as desired. Next, in the case As,s,i,j(w1)=GsiGjiGjiGjsA_{s,s,i,j}(w_{1})=G_{si}G_{ji}G_{ji}G_{js} we estimate

Finally, in the case As,s,i,j(w1)=GsiGjjGiiGjsA_{s,s,i,j}(w_{1})=G_{si}G_{jj}G_{ii}G_{js} we estimate

Using a similar bound for the sum over jj, we find 1N∑s∑i,jGsiGjjGiiGjs=O≺(N2Ψ4)=O≺(N3/2Ψ2)\frac{1}{N}\sum_{s}\sum_{i,j}G_{si}G_{jj}G_{ii}G_{js}=O_{\prec}(N^{2}\Psi^{4})=O_{\prec}(N^{3/2}\Psi^{2}). All other terms are obtained from these three by exchanging ii and jj.

The case q=2𝑞2q=2

In this case N−1∑s1,s2∏r=12Asr,sr,i,j(wr)N^{-1}\sum_{s_{1},s_{2}}\prod_{r=1}^{2}A_{s_{r},s_{r},i,j}(w_{r}) is of the form

or an expression obtained from one of these two by exchanging ii and jj. These may be written as

We estimate the contribution of the first expression by

Next, we split the contribution of the second expression of (12.18) as

Using the anisotropic local law, it is easy to prove that ∣(G2)ij∣≺NΨ2\lvert(G^{2})_{ij}\rvert\prec N\Psi^{2} and \bigl{\lvert}\sum_{j}\Pi_{jj}(G^{2})_{ji}\bigr{\rvert}\prec N^{3/2}\Psi^{2}. Therefore

where we estimate Tr⁡∣G∣4\operatorname{Tr}\lvert G\rvert^{4} as above. This concludes the proof in the case q=2q=2.

The case q=3𝑞3q=3

In this case ∑s1,s2,s3∏r=13Asr,sr,i,j(wr)\sum_{s_{1},s_{2},s_{3}}\prod_{r=1}^{3}A_{s_{r},s_{r},i,j}(w_{r}) is of the form (G2)ij3(G^{2})_{ij}^{3}, or an expression obtained by exchanging ii and jj in some of the three factors. We estimate its contribution by

4. Proof of Theorem 12.4

The proof is analogous to that of Theorem 3.12 in Section 10. Define the domain

From the assumptions on WGauss+DW^{\text{\rm{Gauss}}}+D, it is not hard to deduce the estimate

5. Proof of Theorem 12.5

Analogously to the proof of Theorem 3.18, the proof is a routine application of the Green function comparison method near the edge (EYY3, , Section 6). The key technical inputs are Theorems 12.2 and 12.4. Note that Theorem 12.2 is applicable since the Green function comparison argument only involves zz satisfying ∣Ψ(z)∣⩽N−1/3−c\lvert\Psi(z)\rvert\leqslant N^{-1/3-c} with some small constant c>0c>0. Hence the assumption (b) from Theorem 12.2 is satisfied.

Appendix A Properties of ϱitalic-ϱ\varrho and Stability of (2.11)

This appendix is devoted the proofs of the basic properties of ϱ\varrho and the stability of (2.11) in the sense of Definition 5.4. In this appendix we abbreviate ri\vbox..=ϕπ({si})r_{i}\mathrel{\vbox{\hbox{.}\hbox{.}}}=\phi\pi(\{s_{i}\}).

In this subsection we establish the basic properties of ϱ\varrho and prove Lemmas 2.4–2.6.

From (A.1) we find that (x3f′′(x))′>0(x^{3}f^{\prime\prime}(x))^{\prime}>0 on I0∪⋯∪InI_{0}\cup\cdots\cup I_{n}. Therefore for i=2,…,ni=2,\dots,n there is at most one point x∈Iix\in I_{i} such that f′′(x)=0f^{\prime\prime}(x)=0. We conclude that IiI_{i} has at most two critical points of ff. Using the boundary conditions of ff on ∂Ii\partial I_{i}, we conclude the proof of ∣C∩Ii∣∈{0,2}\lvert\mathcal{C}\cap I_{i}\rvert\in\{0,2\} for i=2,…ni=2,\dots n.

From (A.1) we also find that (x2f′(x))′<0(x^{2}f^{\prime}(x))^{\prime}<0 for x∈I1x\in I_{1}. We conclude that there exists at most one point x∈I1x\in I_{1} such that f′(x)=0f^{\prime}(x)=0. Using the boundary conditions of f′f^{\prime} on ∂I1\partial I_{1}, we deduce that ∣C∩I1∣=1\lvert\mathcal{C}\cap I_{1}\rvert=1.

Finally, if ϕ≠1\phi\neq 1 we have f(x)=(ϕ−1)x−1+O(x−2)f(x)=(\phi-1)x^{-1}+O(x^{-2}) as x→∞x\to\infty. From the boundary conditions of ff on ∂I0\partial I_{0} we therefore deduce that ∣C∩I0∣=1\lvert\mathcal{C}\cap I_{0}\rvert=1. Moreover, if ϕ=1\phi=1 we find from (A.1) that f′(x)≠0f^{\prime}(x)\neq 0 for all x∈I0∖{∞}x\in I_{0}\setminus\{\infty\}. This concludes the proof. ∎

By multiplying both sides of the equation z=f(m)z=f(m) in (2.11) with the product of all denominators on the right-hand side of (2.12), we find that z=f(m)z=f(m) may be also written as Pz(m)=0P_{z}(m)=0, where PzP_{z} is a polynomial of degree n+1n+1, whose coefficients are affine linear functions of zz. (Here we used that all sis_{i} are distinct and that m+si−1≠0m+s_{i}^{-1}\neq 0.) This polynomial characterization of mm is useful for the proofs of Lemmas 2.5 and 2.6.

For i=0,…,ni=0,\dots,n define the subset Ji\vbox..={x∈Ii\vbox..f′(x)>0}J_{i}\mathrel{\vbox{\hbox{.}\hbox{.}}}=\{{x\in I_{i}\mathrel{\vbox{\hbox{.}\hbox{.}}}f^{\prime}(x)>0}\}. The graph of ff restricted to J0∪⋯∪JnJ_{0}\cup\cdots\cup J_{n} is depicted in red in Figures 2.1 and 2.2. From Lemma 2.4 we deduce that if i=2,…,ni=2,\dots,n then Ji≠∅J_{i}\neq\emptyset if and only if IiI_{i} contains two distinct critical points of ff, in which case JiJ_{i} is an interval. Moreover, we always have J1≠∅≠J0J_{1}\neq\emptyset\neq J_{0}.

Next, we observe that for any 0⩽i<j⩽n0\leqslant i<j\leqslant n we have f(Ji)∩f(Jj)=∅f(J_{i})\cap f(J_{j})=\emptyset. Indeed, if there were E∈f(Ji)∩f(Jj)E\in f(J_{i})\cap f(J_{j}) then we would have ∣{x\vbox..f(x)=E}∣>n+1\lvert\{{x\mathrel{\vbox{\hbox{.}\hbox{.}}}f(x)=E}\}\rvert>n+1 (see Figure 2.1). Since f(x)=Ef(x)=E is equivalent to PE(x)=0P_{E}(x)=0 and PEP_{E} has degree n+1n+1, we arrive at the desired contradiction. We conclude that the sets f(Ji)f(J_{i}), 0⩽i⩽n0\leqslant i\leqslant n, may be strictly ordered.

The claim a1⩾a2⩾⋯⩾a2pa_{1}\geqslant a_{2}\geqslant\cdots\geqslant a_{2p} may now be reformulated as

We focus on the first inequality of (A.2); the proof of the second one is similar. We have to show that f(Ji)>f(Jj)f(J_{i})>f(J_{j}) whenever 1⩽i<j⩽n1\leqslant i<j\leqslant n and Ji≠∅≠JjJ_{i}\neq\emptyset\neq J_{j}. First, we claim that for 1⩽i⩽n1\leqslant i\leqslant n we have

To prove (A.3), we remark that if 2⩽i⩽n2\leqslant i\leqslant n then Jit≠∅J_{i}^{t}\neq\emptyset is equivalent to IiI_{i} containing two distinct critical points. Moreover, ∂t∂xft(x)<0\partial_{t}\partial_{x}f^{t}(x)<0 in I2∪⋯∪InI_{2}\cup\cdots\cup I_{n}, from which we deduce that the number of distinct critical points in each IiI_{i}, i=2,…,ni=2,\dots,n, does not decrease as tt decreases. Recalling that I1t≠∅I_{1}^{t}\neq\emptyset, we deduce (A.3).

Next, suppose that there exist 1⩽i<j⩽n1\leqslant i<j\leqslant n satisfying Ji≠∅≠JjJ_{i}\neq\emptyset\neq J_{j} and f(Ji)<f(Jj)f(J_{i})<f(J_{j}). From (A.3) we deduce that Jit≠∅≠JjtJ_{i}^{t}\neq\emptyset\neq J_{j}^{t} for all t∈(0,1]t\in(0,1]. Moreover, by a simple continuity argument and the fact that for each tt we have either ft(Jit)<ft(Jjt)f^{t}(J_{i}^{t})<f^{t}(J_{j}^{t}) or ft(Jit)>ft(Jjt)f^{t}(J_{i}^{t})>f^{t}(J_{j}^{t}), we conclude that ft(Jit)<ft(Jjt)f^{t}(J_{i}^{t})<f^{t}(J_{j}^{t}) for all t∈(0,1]t\in(0,1]. As explained after (A.2), this is impossible for small enough t>0t>0. This concludes the proof of (A.2), and hence also of the first sentence of Lemma 2.5.

Finally, in order to prove the third sentence of Lemma 2.5 we have to show that a2p⩾0a_{2p}\geqslant 0 and a1⩽Ca_{1}\leqslant C. It is easy to see that if a2p<0a_{2p}<0 then there is an E<0E<0 with ∣{x\vbox..f(x)=E}∣>n+1\lvert\{x\mathrel{\vbox{\hbox{.}\hbox{.}}}f(x)=E\}\rvert>n+1, which is impossible since {x\vbox..f(x)=E}={x\vbox..PE(x)=0}\{x\mathrel{\vbox{\hbox{.}\hbox{.}}}f(x)=E\}=\{x\mathrel{\vbox{\hbox{.}\hbox{.}}}P_{E}(x)=0\} and deg⁡PE=n+1\deg P_{E}=n+1. Moreover, the estimate a1⩽Ca_{1}\leqslant C is easy to deduce from the definition of ff and the estimates ϕ⩽τ−1\phi\leqslant\tau^{-1} and s1−1⩾τs_{1}^{-1}\geqslant\tau. ∎

For the following counting argument, in order to avoid extraneous complications we assume that σi>0\sigma_{i}>0 for all ii. Recall the definition of NkN_{k} from (3.16).

Suppose that σj>0\sigma_{j}>0 for all jj. Then Nϱ({0})=(N−M)+N\varrho(\{0\})=(N-M)_{+} and

where in the last step we used the definition of ff from (2.12). Hence (A.5) follows. Note that we proved (A.5) also for k=pk=p. Now (A.4) follows easily using σi>0\sigma_{i}>0 for all ii. Since supp⁡ϱ⊂{0}∪[a2p,a2p−1]∪⋯∪[a2,a1]\operatorname{supp}\varrho\subset\{0\}\cup[a_{2p},a_{2p-1}]\cup\cdots\cup[a_{2},a_{1}] (recall Lemma 2.6), we get N=Nϱ({0})+∑k=1pNkN=N\varrho(\{0\})+\sum_{k=1}^{p}N_{k}. Therefore summing over k=1,…,pk=1,\dots,p in (A.5) yields Nϱ({0})=N−MN\varrho(\{0\})=N-M, as desired. This concludes the proof in the case N>MN>M.

Next, suppose that N⩽MN\leqslant M. Then we deduce directly from (2.12) that ϱ({0})=0\varrho(\{0\})=0 (see Figure 2.2). Hence, (A.4) follows. Finally, exactly as above, we get (A.5) for i=1,…,p−1i=1,\dots,p-1. This concludes the proof. ∎

A.2. Stability near a regular edge

The rest of this appendix is devoted to the proofs of the assumption (3.20) and the stability of (2.11) in the sense of Definition 5.4. In fact, we shall prove the following stronger notation of stability, which will be needed to establish eigenvalue rigidity. Recall the definition of κ\kappa from (10.1).

That the notion of stability from Definition A.2 it is stronger than that of Definition 5.4 follows from the estimate

In this subsection we deal with a regular edge; the case of a regular bulk component is dealt with in the next subsection. We begin with basic estimates on the behaviour of ff near a critical point.

Fix τ>0\tau>0. Suppose that the edge kk is regular in the sense of Definition 2.7 (i). Then there exists a τ′>0\tau^{\prime}>0, depending only on τ\tau, such that

From Lemma 2.5 we get xk=m(ak)x_{k}=m(a_{k}). Since by assumption ak∈[τ,C]a_{k}\in[\tau,C] for some constant CC, we find from Lemma 4.10 that ∣xk∣⩽C\lvert x_{k}\rvert\leqslant C. The lower bound ∣xi∣⩾c\lvert x_{i}\rvert\geqslant c follows from f′(xk)=0f^{\prime}(x_{k})=0 and (A.1), using that ∣si∣⩽C\lvert s_{i}\rvert\leqslant C. This proves the first estimate of (A.9).

Next, from f′(xk)f^{\prime}(x_{k}) and (A.1) we find

Using the first estimate of (A.9) and (2.14) we find from (A.10) that ∣f′′(xk)∣⩽C\lvert f^{\prime\prime}(x_{k})\rvert\leqslant C. Moreover for k∈{1,2p}k\in\{1,2p\} (i.e. xk∈I0∪I1x_{k}\in I_{0}\cup I_{1}) all terms in (A.10) have the same sign, and it is easy to deduce that ∣f′′(xk)∣⩾c\lvert f^{\prime\prime}(x_{k})\rvert\geqslant c.

Next, we prove the lower bound ∣f′′(xk)∣⩾c\lvert f^{\prime\prime}(x_{k})\rvert\geqslant c for k=2,…,2p−1k=2,\dots,2p-1. Suppose for definiteness that kk is odd. (The case of even kk is handled in exactly the same way.) For x∈[xk,xk−1]x\in[x_{k},x_{k-1}] we have f′(x)⩾0f^{\prime}(x)\geqslant 0 and

where in the second step we used that (−x3f′′(x))′<0(-x^{3}f^{\prime\prime}(x))^{\prime}<0 as follows from (A.1), and in the last step the estimate ∣xk∣≍∣xk−1∣≍1\lvert x_{k}\rvert\asymp\lvert x_{k-1}\rvert\asymp 1. We therefore find

Since ak−1−ak⩾τa_{k-1}-a_{k}\geqslant\tau by the assumption in Definition 2.7 (i) and xk−1−xk⩽∣xk−1∣+∣xk∣⩽Cx_{k-1}-x_{k}\leqslant\lvert x_{k-1}\rvert+\lvert x_{k}\rvert\leqslant C, we deduce the lower bound f′′(xk)⩾cf^{\prime\prime}(x_{k})\geqslant c.

Finally, the last estimate of (A.9) follows easily from the assumption (2.14) and by differentiating (A.1). ∎

We record the following easy consequence of (A.9). Recalling that f(xk)=akf(x_{k})=a_{k} and f′(xk)=0f^{\prime}(x_{k})=0, we find, using (A.9) and the fact that ff is continuous in a neighbourhood of xkx_{k}, that after choosing τ′>0\tau^{\prime}>0 small enough (depending on τ\tau) we have

We first prove (A.7). From (A.11) we find that there exists a τ′>0\tau^{\prime}>0 such that for ∣E−ak∣⩽τ′\lvert E-a_{k}\rvert\leqslant\tau^{\prime} we have

Finally, (A.8) follows from the assumption (2.14) and (A.11). ∎

We take over the notation of Definition 5.4. Abbreviate w(z)\vbox..=f(u(z))w(z)\mathrel{\vbox{\hbox{.}\hbox{.}}}=f(u(z)), so that ∣w(z)−z∣⩽δ(z)\lvert w(z)-z\rvert\leqslant\delta(z).

Suppose first that Im⁡z⩾τ′\operatorname{Im}z\geqslant\tau^{\prime} for some constant τ′>0\tau^{\prime}>0. Then, from the assumption δ(z)⩽(log⁡N)−1\delta(z)\leqslant(\log N)^{-1} we get Im⁡w(z)⩾τ′/2\operatorname{Im}w(z)\geqslant\tau^{\prime}/2, and therefore from the uniqueness of (2.11) we find u(z)=m(w(z))u(z)=m(w(z)). Hence,

where we used the trivial bound ∣m′(ζ)∣⩽(Im⁡ζ)−2\lvert m^{\prime}(\zeta)\rvert\leqslant(\operatorname{Im}\zeta)^{-2}. This yields (A.6) at zz.

What remains therefore is the case ∣z−ak∣⩽2τ′\lvert z-a_{k}\rvert\leqslant 2\tau^{\prime}, which we assume for the rest of the proof. We drop the arguments zz and write the equation f(u)−f(m)=w−zf(u)-f(m)=w-z as

Note that α\alpha and β\beta depend on zz, and α\alpha also depends on uu. We suppose that

Then we claim that for small enough τ′\tau^{\prime} we have

In order to prove (A.16), we note that, by Lemma A.3, the statement of (A.11) holds under the assumption (A.15) provided τ′>0\tau^{\prime}>0 is chosen small enough. Using (A.15), (A.8) (see Lemma A.3) , and (A.10) we get

Using (A.9) and (A.11) we conclude that for small enough τ′>0\tau^{\prime}>0 we have ∣α∣≍1\lvert\alpha\rvert\asymp 1. This concludes the proof of the first estimate of (A.16).

In order to prove the second estimate of (A.16), we note that β=m2f′(m)\beta=m^{2}f^{\prime}(m), so that for small enough τ′>0\tau^{\prime}>0 we get from (A.11) and Lemma A.3 that

Using (A.11) and Lemmas A.3 and (4.10), we conclude for small enough τ′>0\tau^{\prime}>0 that

This concludes the proof of the second estimate of (A.16).

A.3. Stability in a regular bulk component

The estimates (A.7) and (A.8) follow trivially from the assumption in Definition 2.7 (ii), using κ⩽C\kappa\leqslant C as follows from Lemma 2.5.

Note that Definition 2.7 (ii) immediately implies that Im⁡m⩾c\operatorname{Im}m\geqslant c for some constant c>0c>0. The upper bound ∣α∣+∣β∣⩽C\lvert\alpha\rvert+\lvert\beta\rvert\leqslant C easily follows from the definition (A.14) combined with Lemma 4.10. What remains is the proof of the lower bound ∣β∣⩾c\lvert\beta\rvert\geqslant c. To that end, we take the imaginary part of (2.11) to get

Using that Im⁡m⩾c\operatorname{Im}m\geqslant c, a simple analysis of the arguments of the expressions (m+si−1)2(m+s_{i}^{-1})^{2} on the left-hand side of (A.17) yields

where in the last step we used (A.17). Recalling the definition of β\beta from (A.14), we conclude that ∣β∣⩾c\lvert\beta\rvert\geqslant c. This concludes the proof. ∎

The bulk regularity condition from Definition 2.7 (ii) is stable under perturbation of π\pi. To see this, define the shifted empirical density πt\vbox..=M−1∑i=1Mδσi+t\pi_{t}\mathrel{\vbox{\hbox{.}\hbox{.}}}=M^{-1}\sum_{i=1}^{M}\delta_{\sigma_{i}+t}, and the associated Stieltjes transform mt(E)m_{t}(E) and function ft(x)f_{t}(x). Differentiating ft(mt(E))=Ef_{t}(m_{t}(E))=E in tt yields (∂tm)t(E)=−(∂tf)t(mt(E))/(∂xf)t(mt(E))(\partial_{t}m)_{t}(E)=-(\partial_{t}f)_{t}(m_{t}(E))/(\partial_{x}f)_{t}(m_{t}(E)). Using that

by (A.8), and (∂xf)t(mt(E))=m(E)−2β(E)(\partial_{x}f)_{t}(m_{t}(E))=m(E)^{-2}\beta(E), we find from ∣β(E)∣≍∣m(E)∣≍1\lvert\beta(E)\rvert\asymp\lvert m(E)\rvert\asymp 1 that (∂tm)0(E)=O(1)(\partial_{t}m)_{0}(E)=O(1) for E∈[a2k+τ′,a2k−1−τ′]E\in[a_{2k}+\tau^{\prime},a_{2k-1}-\tau^{\prime}]. A simple extension of this argument shows that if Definition 2.7 (ii) holds for π\pi then it holds for all πt\pi_{t} with tt in some ball of fixed radius around zero.

A.4. Stability outside of the spectrum

Since mm is the Stieltjes transform of a measure ϱ\varrho with bounded density (see Lemma 4.10), we find that

Next, we note that if η⩾ε\eta\geqslant\varepsilon for some fixed ε>0\varepsilon>0 then (A.8) is trivially true by (A.18). On the other hand, if Im⁡m⩽ε\operatorname{Im}m\leqslant\varepsilon then we get

Acknowledgements

A.K. was partially supported by Swiss National Science Foundation grant 144662 and the SwissMAP NCCR grant. J.Y. was partially supported by NSF Grant DMS-1207961. We are very grateful to the Institute for Advanced Study, Thomas Spencer, and Horng-Tzer Yau for their kind hospitality during the academic year 2013-2014. We also thank the Institute for Mathematical Research (FIM) at ETH Zürich for its generous support of J.Y.’s visit in the summer of 2014. We are indebted to Jamal Najim for stimulating discussions.

References