Universality for general Wigner-type matrices

Oskari Ajanki, Laszlo Erdos, Torben Krüger

Introduction

In the seminal paper Wigner introduced random self-adjoint matrices, H=H∗\mathbf{H}=\mathbf{H}^{*}, with centered, identically distributed and independent entries (subject to the symmetry constraint). He proved that the empirical density of the eigenvalues converges to the semicircle distribution. Wigner also conjectured that the distribution of the distance between consecutive eigenvalues (gap statistics) is universal, hence it is the same as in the Gaussian model. His revolutionary observation was that these universality phenomena hold for much larger classes of physical systems and only the basic symmetry type determines local spectral statistics. It is generally believed, but mathematically unproven, that random matrix theory (RMT), among many other examples, also describes the local statistics of random Schrödinger operators in the delocalized regime and quantization of chaotic classical Hamiltonians.

become not only deterministic but also independent of ii as the the matrix size NN goes to infinity. They asymptotically satisfy a system of self-consistent equations

that reduces to a particularly simple scalar equation

for the common value m≈Giim\approx G_{ii} for all ii as N→∞N\to\infty. The solution m=m(z)m=m(z) of (1.3) is the Stieltjes transform of the Wigner semicircle law.

In the context of random matrices importance of this equation has been realized by Girko , Shlyakhtenko , Khorunzhy and Pastur , see also Guionnet , as well as Anderson and Zeitouni , but no detailed study has been initiated. In the companion paper we analyzed (1.4) in full detail. See also Section 3 of for how the QVE is related to other random matrix models. We showed that ⟨\mspace1.0mum⟩:=N−1∑imi\langle\mspace{1.0mu}m\rangle:=N^{-1}\sum_{i}m_{i} is the Stieltjes transform of a probability density ρ\rho that is supported on a finite number of intervals, inside of which it is a real analytic function. We also described the behavior of ρ\rho near the edges of its support; it features only square root or cubic root (cusp) singularities and an explicit one parameter family of profiles interpolating between them as a gap in the support closes.

The main result of the current paper is the universality of the local eigenvalue statistics in the bulk for Wigner-type matrices with a general variance matrix (cf. Theorem 1.16). This extends Wigner’s vision towards full universality by considering a much larger class of matrix ensembles than previously studied. In particular, we demonstrate that local statistics, as expected, are fully independent of the global density. This fact has already been established for very general β\beta-ensembles in (see also and ) and for additively deformed Wigner ensembles having a density with a single interval support . Our class admits a general variance matrix and allows for densities with several intervals (we do not, however, consider non-centered distributions here; an extension to matrices with non-centered entries on the diagonal may be incorporated in our analysis with additional technical effort).

In a separate paper we apply the results of this work and to treat Gaussian random matrices with correlated entries. Assuming translation invariance of the correlation structure in these Gaussian matrix ensembles we prove an optimal local law, bulk universality and non-trivial decay of off-diagonal resolvent entries.

Acknowledgement. We thank the anonymous referee for several useful comments and suggestions. We are grateful to Johannes Alt for pointing out several typos.

The dependence of H\mathbf{H} and other quantities on the dimension NN will be suppressed in our notation. The matrix of variances, S=(sij)i,j=1N\mathbf{S}=(s_{ij})_{i,j=1}^{N}, is defined through

It is symmetric with non-negative entries. In it was shown that for every such matrix S\mathbf{S} the quadratic vector equation (QVE),

For all NN the matrix S\mathbf{S} is flat, i.e.,

For all NN the matrix S\mathbf{S} is uniformly primitive, i.e.,

For all NN the matrix S\mathbf{S} induces a bounded solution of the QVE, i.e., the unique solution m\mathbf{m} of (1.7) corresponding to S\mathbf{S} is bounded,

The assumption on the boundedness of m\mathbf{m} is an implicit condition in the sense that it can be checked only after solving (1.7). In Theorem 6.1 of we list sufficient, explicitly checkable conditions on S\mathbf{S}, which ensure (1.10). We also remark that the assumption (1.8) can be replaced by sij≤C/Ns_{ij}\leq C/N for some positive constant CC. This will lead to a rescaling (cf. Remark 2.2 of ) of m\mathbf{m}. We pick the normalization C=1C=1 just for convenience.

The primitivity condition (1.9) excludes some important models, e.g. matrices of the form

In addition to the assumptions on the variances of hijh_{ij}, we also require uniform boundedness of higher moments. This leads to another basic model parameter, μ‾=(μ1,μ2,… )\underline{\mu}=(\mu_{1},\mu_{2},\dots), which is a sequence of non-negative real numbers.

For all NN the entries hijh_{ij} of the random matrix H\mathbf{H} have bounded moments,

In order to state our main result, in the next corollary we collect a few facts about the solution of the QVE that are proven in . Although these properties are sufficient for the formulation of our results, for their proofs we will need much more precise information about the solution of the QVE. Theorems 4.1 and 4.2 summarize everything that is needed from besides the existence and uniqueness of the solution of the QVE. In particular, the statement of Corollary 1.3 follows easily from Theorem 4.1 below.

is a probability density. Its support is contained in $$ and is a union of closed disjoint intervals

There exists a positive constant δ∗\delta_{*}, depending only on the model parameters pp, PP and LL, such that the sizes of these intervals are bounded from below by

Note that (1.14) provides a lower bound on the length of the intervals that constitute supp⁡ρ\operatorname{supp}\rho, while the length of the gaps, αk+1−βk\alpha_{k+1}-\beta_{k}, between neighboring intervals can be arbitrarily small. Figure 1.1 shows a shape that the density of states typically might have. In particular, ρ\rho may have gaps in its support and may have additional zeros (cusps) in the interior of supp⁡ρ\operatorname{supp}\rho. However, the behavior of ρ\rho on the domain ρ≤ε\rho\leq\varepsilon, for some sufficiently small ε>0\varepsilon>0, can be completely characterized by universal shape functions. For more details see Theorem 2.6 of .

The function ρ\rho defined in (1.12) is called the density of states. Its harmonic extension to the upper half plane

is again denoted by ρ\rho. With a slight abuse of notation we still write supp⁡ρ\operatorname{supp}\rho, as in (1.13), for the support of the density of states as a function on the real line.

Let αk\alpha_{k} and βk\beta_{k} be the edges of the support of the density of states (cf. (1.13)) and δ∗\delta_{*} the constant introduced in Corollary 1.3. Then for any δ∈[0,δ∗)\delta\in[0,\delta_{*}) we set

For a compact statement of the main theorem we define the notion of stochastic domination, introduced in and . This notion is designed to compare sequences of random variables in the large NN limit up to small powers of NN on high probability sets.

In this case we write φ≺ψ\varphi\prec\psi.

Basic properties of the stochastic domination that are used extensively in this paper are listed in Lemma A.1. The threshold N0(ε,D)=N0(ε,D;P,p,L,μ‾,γ)N_{0}(\varepsilon,D)=N_{0}(\varepsilon,D;P,p,L,\underline{\mu},\gamma) will always be an explicit function whose value will be increased throughout the paper, though we will not follow its form. This will happen only finitely many times, ensuring that N0N_{0} stays finite. The threshold is uniform in all other parameters, e.g. in the spectral parameter zz, as well as in the indices i,j,…i,j,\dots of the matrix entries, that the sequences φ\varphi and ψ\psi may depend on. Typically, we will not mention the existence of N0N_{0} - it is implicit in the notation φ≺ψ\varphi\prec\psi. As an example, we see that the bounded moment condition, (D), implies

Actually, the function N0N_{0} depends only on finitely many moment parameters (μ1,…,μM)(\mu_{1},\dots,\mu_{M}) instead of the entire sequence μ‾\underline{\mu}, where now the number of required moments MM =M(ε,D;P,p,L,γ)=M(\varepsilon,D;P,p,L,\gamma), is an explicit function.

Now we are ready to state our main result on the local law. Suppose H=H(N)\mathbf{H}=\mathbf{H}^{(N)} is a sequence of self-adjoint random matrices with the corresponding sequence S=S(N)\mathbf{S}=\mathbf{S}^{(N)} of variance matrices and ρ=ρ(N)\rho=\rho^{(N)} the induced sequence of densities of state. Recall that δ∗\delta_{*} is the positive constant, depending only on pp, PP and LL, introduced in Corollary 1.3 and Δδ\Delta_{\delta} is defined as in Definition 1.5.

In particular, for wi=1w_{i}=1 this implies

The function κ\kappa may be chosen to be

where Δ=Δδ\Delta=\Delta_{\delta}, with some δ∈(0,δ∗)\delta\in(0,\delta_{*}) that depends only on the model parameters pp, PP and LL.

In the regime, where zz is not too close to the support of the density of states in the sense that

where ρ∗\rho_{\ast} is considered as an additional model parameter.

Theorem 1.7 generalizes the previous local laws for stochastic variance matrices S\mathbf{S} (see and references therein). It is valid for densities ρ\rho with an edge behavior different from the square root growth that is known from Wigner’s semicircular law. In particular, singularities that interpolate between a square root and a cubic root are possible. In the bulk of the support of the density of states, i.e., where ρ\rho is bounded away from zero, the function κ\kappa is bounded. The same is true near the edges, unless the nearby gap is small. The bound deteriorates near small gaps in the support of ρ\rho.

In applications, the sequence S=S(N)\mathbf{S}=\mathbf{S}^{(N)} satisfying (A)–(C) may be constructed by discretizing a piecewise 1/21/2-Hölder continuous limit function (cf. Remark 6.2 in ). As a particularly simple example, suppose ff is a smooth, non-negative, symmetric, f(x,y)=f(y,x)f(x,y)=f(y,x), function on 2^{2} with a positive diagonal, f(x,x)>0f(x,x)>0. Then the sequence of variance matrices,

satisfies conditions (A)–(C). The validity of (C) can be verified by using the general criteria (cf. Theorem 2.10 and Theorem 6.1 of ) for uniform boundedness. In this case the solution, m=(m1,…,mN)\mathbf{m}=(m_{1},\dots,m_{N}), of the QVE converges to a limit in the sense that

The continuous QVE such as this one fall into the class of general QVEs thoroughly analyzed in the companion paper . In particular, the stability analysis applies and the density of states converges to a limit

We introduce a notion for expressing that events hold with high probability in the limit as NN tends to infinity.

We denote by λ1≤⋯≤λN\lambda_{1}\leq\dots\leq\lambda_{N} the eigenvalues of the random matrix H\mathbf{H}. The following corollary shows that the eigenvalue distribution converges to the density of states as NN tends to infinity.

Furthermore, for an arbitrary tolerance exponent γ∈(0,1)\gamma\in(0,1) there are no eigenvalues away from the support of the density of states,

where we interpret β0:=−∞\beta_{0}:=-\infty, αK+1:=+∞\alpha_{K+1}:=+\infty and δk\delta_{k} is defined as δ0:=δK:=Nγ−2/3\delta_{0}:=\delta_{K}:=N^{\gamma-2/3}, as well as

Based on (1.16) we define the index, i(τ)i(\tau), of an eigenvalue that we expect to be located close to the spectral parameter τ\tau by

Assume (A)-(D), and let γ∈(0,1)\gamma\in(0,1) be an arbitrary tolerance exponent. Denote

and ε0:=εK:=Nγ−2/3\varepsilon_{0}:=\varepsilon_{K}:=N^{\gamma-2/3}. Then uniformly for every

Furthermore, if τ\tau is close to the extreme edge, τ∈(α1,α1+ε0)\tau\in(\alpha_{1},\alpha_{1}+\varepsilon_{0}) or τ∈(βK−εK,βK\mspace1.0mu]\tau\in(\beta_{K}-\varepsilon_{K},\beta_{K}\mspace{1.0mu}], then

Finally, if τ∈(βk−εk,αk+1+εk)\tau\in(\beta_{k}-\varepsilon_{k},\alpha_{k+1}+\varepsilon_{k}) for some 1≤k≤K−11\leq k\leq K-1, then the corresponding eigenvalue is close to an internal edge in the sense that

The statements (1.35) and (1.36) are an immediate consequence of (1.34) and (1.29). They simply express the fact that the small number of O(Nε)\mathcal{O}(N^{\varepsilon}) eigenvalues, very close to the edges, are found in the space that is left for them by the other eigenvalues for which the rigidity statement (1.34) applies. For an illustration see Figure 1.2. We also note that results of this type date back to at least (in the sample covariance context).

where κ\kappa is the function from Theorem 1.7.

In particular, the eigenvectors are completely delocalized, i.e., ∥\mspace1.0muu(i)∥∞=max⁡j∣\mspace1.0muuj(i)∣≺N−1/2\lVert\mspace{1.0mu}\mathbf{u}^{(i)}\rVert_{\infty}=\max_{j}\lvert\mspace{1.0mu}u^{(i)}_{j}\rvert\prec N^{-1/2}.

We say that H\mathbf{H} is qq-full for some q>0q>0 (independent of NN) if either of the following applies:

H\mathbf{H} is complex hermitian and for all i,j=1,…,Ni,j=1,\dots,N the real symmetric 2×22\times 2-matrix,

We use the convention that every positive constant with a lower star index, such as δ∗\delta_{*}, c∗c_{*} and λ∗\lambda_{*}, explicitly depends only on the model parameters PP, pp and LL from (B)–(D). These dependencies can be reconstructed from the proofs, but we will not follow them. Constants c,c1,c2,…,C,C1,C2,…c,c_{1},c_{2},\dots,C,C_{1},C_{2},\dots also depend only on PP, pp and LL. They will have a local meaning within a specific proof.

For two non-negative functions φ\varphi and ψ\psi depending on a set of parameters u∈Uu\in U, we use the comparison relation

if there exists a positive constant cc, depending explicitly on PP, pp and LL such that φ(u)≥c\mspace2.0muψ(u)\varphi(u)\geq c\mspace{2.0mu}\psi(u) for all u∈Uu\in U. The notation ψ∼φ\psi\sim\varphi means that both ψ≲φ\psi\lesssim\varphi and ψ≳φ\psi\gtrsim\varphi hold true. In this case we say that ψ\psi and φ\varphi are comparable. We also write ψ=φ+O(ϑ)\psi=\varphi+\mathcal{O}(\vartheta), if ∣ψ−φ∣≲ϑ|\psi-\varphi|\lesssim\vartheta.

Bound on the random perturbation of the QVE

We will make the following standing assumptions for the rest of this paper,

The assumptions (A)–(D) hold true and an arbitrary tolerance exponent γ∈(0,1)\gamma\in(0,1) is fixed;

which are always assumed to hold unless explicitly otherwise stated.

We introduce the notation G(V)\mathbf{G}^{(V)} for the resolvent of the matrix H(V)\mathbf{H}^{(V)}, which is identical to H\mathbf{H} except for the removal of the rows and columns corresponding to the indices V⊆{1,…,N}V\subseteq\{1,\dots,N\}. The enumeration of the indices is kept, even though G(V)\mathbf{G}^{(V)} has a lower dimension.

The diagonal elements of the resolvent, g:=(G11,…,GNN)\mathbf{g}:=(G_{11},\dots,G_{NN}), satisfy the perturbed quadratic vector equation

Here and in the following, the upper indices on the sums indicate which indices are not summed over. For the proof of this simple identity as well as (2.3) below via the Schur complement formula we refer to . As in (2.2) we will often omit the dependence on the spectral parameter zz in our notation, i.e., Gij=Gij(z)G_{ij}=G_{ij}(z), dk=dk(z)d_{k}=d_{k}(z), etc..

We will now derive an upper bound on ∥d∥∞=max⁡i∣di∣\lVert\mathbf{d}\rVert_{\infty}=\max_{i}|d_{i}|, provided ∣gi−mi∣|g_{i}-m_{i}| is bounded by a small constant. At the same time we will control the off-diagonal elements GklG_{kl} of the resolvent. These satisfy the identity

for k≠lk\neq l. The strategy in what follows below is that (2.2) and (2.3) are used to improve a rough bound on the entries of the resolvent G\mathbf{G} to get the correct bounds on the random perturbation and the off-diagonal resolvent elements. Later, in Section 3, the stability of the QVE under the small perturbation, d\mathbf{d}, will provide the improved bound on the diagonal elements, Gii−mi=gi−miG_{ii}-m_{i}=g_{i}-m_{i}.

We introduce a short notation for the difference between g\mathbf{g} and the solution m\mathbf{m} of the unperturbed equation (1.7),

The following lemma is analogous to Lemma 5.2 in with minor modifications. For the completeness of this paper, we repeat these arguments. One small modification is that our estimates also deal with the regime where ∣z∣|z| is large. To keep the formulas short we denote

For the proof of this lemma we will need an additional property of the solution of the QVE that is a corollary of Theorem 4.1, where all properties of m\mathbf{m} taken from are summarized.

The absolute value of the solution of the QVE satisfies

Here we use the three large deviation estimates,

Since G(V)\mathbf{G}^{(V)} is independent of the rows and columns of H\mathbf{H} with indices in VV, these estimates follow directly from the large deviation bounds in Appendix C of . Furthermore, we use

where latter the inequality is just assumption (1.8) and the bound on hijh_{ij} follows from (1.11). We remark that the stochastic domination in (2.7) and (2.8) is uniform in k,lk,l and i,ji,j, respectively, i.e., the threshold function N0N_{0} in Definition 1.6 does not depend on i,j,k,li,j,k,l.

We will now show that the removal of a few rows and columns in H\mathbf{H} will only have a small effect on the entries of the resolvent. The general resolvent identity,

In the inequality we used that ∣mk(z)∣∼[z]−1|m_{k}(z)|\sim[z]^{-1} (cf. Corollary 2.2), ∣gk∣=∣mk∣+O(Λ)|g_{k}|=|m_{k}|+\mathcal{O}(\Lambda) and that λ∗\lambda_{*} is chosen to be small enough. We use (2.10) in a similar calculation for Gij(l)G_{ij}^{(l)} and find that on the event where Λ≤λ∗/[z]\Lambda\leq\lambda_{*}/[z],

Again using (2.10) and that the denominator of the last expression is comparable to [z]−1[z]^{-1}, we conclude

We have now collected all necessary ingredients and use them to estimate all the terms in (2.2) one by one. We start with the first summand. By (2.7a) we find

With the help of (2.10) we remove the upper index from Gij(k)G_{ij}^{(k)} and get

For the second summand in (2.2) we use the large deviation bound for the diagonal, (2.7c), and find that

By removing the upper index again we estimate

We use this in (2.15) and for sufficiently small λ∗\lambda_{*} we arrive at

The third summand in (2.2) is estimated directly by

We combine the estimates for the individual terms (2.14), (2.17), (2.18) and (2.8). Altogether we conclude that

Here, we applied the large deviation bound (2.7b). Using the Ward identity for the resolvent G(kl)\mathbf{G}^{(kl)},

and (2.10) for removing the upper index of Gll(k)G_{ll}^{(k)} we get

We remove the upper indices from Gii(kl)G_{ii}^{(kl)} and end up with

Local law away from local minima

In this section we will use the stability of the QVE to establish the main result away from the local minima of the density of states inside its own support, i.e. away from the set

First we consider the regime η≥δ∗/2\eta\geq\delta_{\ast}/2. Using Δ1/3+ρ≲1\Delta^{1/3}+\rho\lesssim 1 we see that B0≳1B_{0}\gtrsim 1. Similarly, we get B1≳η\mspace2.0mu[z]−1B_{1}\gtrsim\eta\mspace{2.0mu}[z]^{-1}. Since (Nη)−1Bk(N\eta)^{-1}B_{k}, k=0,1k=0,1, are both bigger than the right hand side of (3.3), we obtain (1.21) for η≥δ∗/2\eta\geq\delta_{\ast}/2.

The proof of Proposition 3.1 uses a continuity argument in zz. In particular, continuity of the solution of the QVE is needed. The statement of the following corollary is part of the properties of m\mathbf{m} listed in Theorem 4.1 below.

The solution of the QVE is uniformly Hölder-continuous,

Since we will estimate the difference, g−m\mathbf{g}-\mathbf{m}, we start by deriving an equation for this quantity. Using the QVE for m\mathbf{m} and the perturbed equation (2.1) for g\mathbf{g} we find

Using the bound inside the indicator function from (3.8) and ∣z∣≥10|z|\geq 10 the assertion (3.8) of the lemma follows.

For the proof of Proposition 3.1 we use the stability of (3.7) also close to supp⁡ρ\operatorname{supp}\rho. This requires more care and is carried out in detail in . The result of that analysis is Theorem 4.2. Here we will only need the following consequence of that theorem and (4.5a).

Furthermore, the following fluctuation averaging result is needed. It was first established for generalized Wigner matrices with Bernoulli distributed entries in .

In the setting where H\mathbf{H} is a generalized Wigner matrix and ∣z∣≤10|z|\leq 10 this bound is precisely the content of Theorem 4.7 from .

The a priori bound used in the proof of that theorem is replaced by

for any V⊆{1,…,N}V\subseteq\{1,\dots,N\} with NN-independent size. This bound is proven in the same way as (2.19). Here, the N0N_{0} hidden in the stochastic domination depends on the size ∣V∣|V| of the index set. Following the proof of Theorem 4.7 given in with (3.17) and tracking the zz-dependence,

yields the fluctuation averaging, Theorem 3.5. ∎

From this and the definition of d\mathbf{d} in (2.2) we read off the a priori bound,

Here, we used the general resolvent identity (2.9) in the form GikGki=gk(gi−Gii(k))G_{ik}G_{ki}=g_{k}(g_{i}-G_{ii}^{(k)}). Since g\mathbf{g} satisfies the perturbed QVE (2.1) and ∣∑j=1Nsij\mspace1.0mugj(z)+di(z)∣≺N2−γ|\sum_{j=1}^{N}s_{ij}\mspace{1.0mu}g_{j}(z)+d_{i}(z)|\prec N^{2-\gamma} from (3.19) and (3.18) we conclude that uniformly for ∣z∣≥N\mspace1.0mu2|z|\geq N^{\mspace{1.0mu}2} we have

for k≠lk\neq l. With the bound (3.20) we conclude that

For the bound on the off-diagonal error term we plug this result into (3.24) and get

according to (3.8) in Lemma 3.3 (for ∣z∣≥10|z|\geq 10) and (3.13) from Corollary 3.4 (for ∣z∣≤10|z|\leq 10), where λ∗\lambda_{*} is a sufficiently small positive constant.

Using again the weighted Cauchy-Schwarz inequality in the second term yields

In particular, we combine (3.27) and (3.28) to establish a gap in the values that Λ\Lambda can take,

Here we used η≥Nγ−1\eta\geq N^{\gamma-1}. This shows that either Λ≥λ∗/[z]\Lambda\geq\lambda_{\ast}/[z] or Λ≤N−γ/4/[z]\Lambda\leq N^{-\gamma/4}/[z] a.w.o.p.

Now we apply Lemma A.2 on the connected domain

The continuity condition (A.1) of the lemma for these two functions follows from the Hölder-continuity, (3.5), of the solution of the QVE and the weak continuity of the resolvent elements,

We will now sketch the proof of Corollary 1.8. The set-up in this corollary differs slightly from the one used in the rest of this paper, because the uniform bound (assumption (C)) on the solution of (1.7) is not assumed. We therefore use additional information from about m\mathbf{m} in this more general setting.

Since the boundedness assumption (C) on the solution of the QVE is dropped in this corollary, its proof starts by showing that nevertheless for some constant P>0P>0 we have

Local law close to local minima

The probability densities are comparable,

The size of the harmonic extension (1.15) of ρ\rho, up to constant factors, is given by explicit functions as follows. Let η∈[0,δ∗]\eta\in[0,\delta_{*}].

At an internal edge: At the edges αi,βi−1\alpha_{i},\beta_{i-1} with i=2,…,Ki=2,\dots,K in the direction where the support of the density of states continues the size of ρ\rho is

Inside a gap: Between two neighboring edges βi−1\beta_{i-1} and αi\alpha_{i} with i=2,…,Ki=2,\dots,K, the function ρ\rho satisfies

for all ω∈[0,(αi−βi−1)/2]\omega\in[0,(\alpha_{i}-\beta_{i-1})/2].

Around an extreme edge: At the extreme points α1\alpha_{1} and βK\beta_{K} of supp⁡ρ\operatorname{supp}\rho the density of states grows like a square root ,

Away from the support: Away from the interval in which supp⁡ρ\operatorname{supp}\rho is contained

The next theorem shows that the QVE is stable under small perturbations, d\mathbf{d}, in the sense that once a solution of the perturbed QVE (4.6) is sufficiently close to m\mathbf{m}, then the difference between the two can be estimated in terms of ∥d∥∞\lVert\mathbf{d}\rVert_{\infty}. In it is stated as Proposition 10.1.

in the following two ways. On the whole complex upper half plane

Furthermore, the function σ\sigma is related to the density of states by

2 Coefficients of the cubic equation

There exist δ∗,c∗∼1\delta_{\ast},c_{\ast}\sim 1 such that for all η∈[0,δ∗]\eta\in[0,\delta_{*}] the coefficients, π1\pi_{1} and π2\pi_{2}, of the cubic equation (4.10) satisfy the following bounds.

Around an internal edge: At the edges αi,βi−1\alpha_{i},\beta_{i-1} of the gap with length Δ:=αi−βi−1\Delta:=\alpha_{i}-\beta_{i-1} for i=2,…,Ki=2,\dots,K, we have

Well inside a gap: Between two neighboring edges βi−1\beta_{i-1} and αi\alpha_{i} of the gap with length Δ:=αi−βi−1\Delta:=\alpha_{i}-\beta_{i-1} for i=2,…,Ki=2,\dots,K, the first coefficient, π1\pi_{1}, satisfies

The second coefficient, π2\pi_{2}, satisfies the upper bounds,

Around an extreme edge: Around the extreme points α1\alpha_{1} and βK\beta_{K} of supp⁡ρ\operatorname{supp}\rho, we have

In the last relation we used the behavior (4.5e) of ρ\rho from Theorem 4.1. By (4.11) we conclude that inside the δ∗\delta_{*}-neighborhood of τ0\tau_{0},

Using the upper and lower bounds on ρ(z)\rho(z) again, gives the desired result, (4.15e). Around an internal edge: First we prove the bounds on ∣π2∣|\pi_{2}|, starting from (4.11). The upper bound simply uses the 1/31/3-Hölder-continuity and the behavior at the edge points of σ\sigma,

where τ0\tau_{0} is one of the edge points αi\alpha_{i} or βi−1\beta_{i-1}. The claim follows from plugging in the size of ρ\rho from the two corresponding domains in Theorem 4.1, i.e., the domain close to an edge, (4.5b), and the domain inside a gap, (4.5c).

For the lower bound we consider two different regimes. In the first case zz is close to the edge point, ∣z−τ0∣≤c\mspace1.0muΔ|z-\tau_{0}|\leq c\mspace{1.0mu}\Delta, for some small positive constant cc, depending only on the model parameters pp, PP and LL. We find

provided cc is small enough. This bound coincides with the lower bound on π2\pi_{2} in (4.15a), once the size of ρ\rho from (4.5b) is used.

This finishes the proof of the upper and lower bound on ∣π2∣|\pi_{2}| on this domain. For the claim about ∣π1∣|\pi_{1}| we plug the result about ∣π2∣|\pi_{2}| and the size of ρ\rho into

3 Rough bound on ΛΛ\Lambda close to local minima

Now we start the detailed proof from the fact that Θ\Theta satisfies the cubic equation (4.10), whose right hand side is bounded by C∥d∥∞C\lVert\mathbf{d}\rVert_{\infty} for some constant CC, depending only on the model parameters. Note that ∥d∥∞≲1\lVert\mathbf{d}\rVert_{\infty}\lesssim 1 as long as Λ≤λ∗\Lambda\leq\lambda_{*} because in this case ∣mi∣∼1|m_{i}|\sim 1, ∣gi∣∼1|g_{i}|\sim 1 and g\mathbf{g} satisfies the perturbed QVE with perturbation d\mathbf{d}. From the definition of Θ\Theta in (4.7) and the uniform bound on s\mathbf{s} from (4.13), we get Θ≲Λ\Theta\lesssim\Lambda. Since the coefficient ∣π2∣|\pi_{2}| is uniformly bounded (cf. (4.11)), the cubic equation for Θ\Theta implies the three bounds

In order to satisfy the constraint of (4.10) we have also used ε2≤λ∗\varepsilon_{2}\leq\lambda_{\ast}. This together with (4.22) yields (4.21b).

The function Δ\Delta is from Definition 1.5 and its value is simply the length of the gap at the point τ0∈∂supp⁡ρ\tau_{0}\in\partial\operatorname{supp}\rho where it is evaluated. We also define the δ\delta-neighborhoods of these subsets,

As an immediate consequence of the upper and lower bounds on the coefficients, π1\pi_{1} and π2\pi_{2}, presented in Proposition 4.3, we see that

Now we make a choice for the two constants ε1\varepsilon_{1} and ε2\varepsilon_{2}. We express them in terms of δ\delta as

We pair the bounds on Θ\Theta from (4.21) with the corresponding bounds from (4.23) on the coefficients of the cubic equation. For small enough δ\delta the conditions on π1\pi_{1} in (4.21a) and π2\pi_{2} in (4.21b) are automatically satisfied by the choice of ε1\varepsilon_{1} and ε2\varepsilon_{2}, as well as the upper and lower bounds from (4.23a) and (4.23b). Thus, for small enough δ\delta we end up with

4 Proof of Theorem 1.7

When τ0∈∂supp⁡ρ\tau_{0}\in\partial\operatorname{supp}\rho we denote by θ=θ(τ0)∈{±1}\theta=\theta(\tau_{0})\in\{{\pm 1}\} the direction that points towards the gap in supp⁡ρ\operatorname{supp}\rho at τ0\tau_{0}. In case τ0∉∂supp⁡ρ\tau_{0}\notin\partial\operatorname{supp}\rho we make the arbitrary choice θ:=+1\theta:=+1, i.e.,

where η∈(0,δ∗]\eta\in(0,\delta_{*}] and ω∈[−δ∗,δ∗]\omega\in[-\delta_{*},\delta_{*}]. We will then prove the local law in the form

where the positive error function E:[−δ∗,δ∗]×(0,δ∗]→(0,∞)\mathcal{E}:[-\delta_{*},\delta_{*}]\times(0,\delta_{*}]\to(0,\infty) is given as the unique solution of an explicit cubic equation in (4.30) below.

To define E\mathcal{E} we introduce explicit auxiliary functions π~1\widetilde{\pi}_{1}, π~2\widetilde{\pi}_{2} and ρ~\widetilde{\rho} that are comparable in size to the corresponding functions π1\pi_{1}, π2\pi_{2} and ρ\rho. The reason for using these auxiliary quantities for the definition of E\mathcal{E} instead of the original ones is twofold. Firstly, in this way E\mathcal{E} will be an explicit function instead of one that is implicitly defined through the solution of the QVE. The function E\mathcal{E} is explicit in the sense that there is a formula for the solution of the cubic equation that defines it and the coefficients are given by the explicit functions π~1\widetilde{\pi}_{1}, π~2\widetilde{\pi}_{2} and ρ~\widetilde{\rho}. Secondly, E\mathcal{E} will be monotonic of its second variable, η\eta. This property will be used later. The definition of the three auxiliary functions will be different, depending on whether τ0\tau_{0} is in the boundary of the support of the density of states or not. Recall the definition (1.17) of Δδ(τ)\Delta_{\delta}(\tau).

Edge: If τ0∈∂supp⁡ρ\tau_{0}\in\partial\operatorname{supp}\rho, i.e. τ0\tau_{0} is an edge of a gap of size Δ:=Δ0(τ0)\Delta:=\Delta_{0}(\tau_{0}) in the support of the density of states or an extreme edge. Then we define the three explicit functions

Here, c∗∼1c_{*}\sim 1 is the constant from Proposition 4.3.

By design (cf. Proposition 4.3 and Theorem 4.1) these functions satisfy

except in one special case where the second bound does not hold, namely when k=2k=2, τ0∈∂supp⁡ρ\tau_{0}\in\partial\operatorname{supp}\rho and ω∈[c∗Δ,Δ/2]\omega\in[c_{*}\Delta,\Delta/2]. In this case only the direction ∣π2∣≲π~2|\pi_{2}|\lesssim\widetilde{\pi}_{2} is true (cf. (4.15c)).

We fix a positive constant ε~∈(0,γ/16)\widetilde{\varepsilon}\in(0,\gamma/16). The value of the function E\mathcal{E} at (ω,η)(\omega,\eta) is then defined to be the unique positive solution of the cubic equation

With the choices (1.23) and (1.25) for κ=κ(z)\kappa=\kappa(z) we have

for any N≥N0N\geq N_{0}, where the threshold N0N_{0} here depends on ε~\widetilde{\varepsilon} in addition to pp, PP, LL, μ‾\underline{\mu} and γ\gamma. The inequality (4.31) is verified by plugging its right hand side into (4.30) in place of E\mathcal{E} and checking that on each regime the resulting expression on the right hand side of (4.30) is smaller than the resulting expression on the left hand side of (4.30). The factor of N9\mspace1.0muε~N^{9\mspace{1.0mu}\widetilde{\varepsilon}} in (4.31) can be absorbed in the stochastic domination in (4.26). Thus (4.26) becomes (1.20) and (1.21) of Theorem 1.7.

Before we start the proof of the local law (4.26), let us motivate the definition of E\mathcal{E}. As a consequence of Lemma 4.4 the indicator function equals one a.w.o.p. in the statement of Lemma 2.1. Thus, uniformly in the δ∗\delta_{*}-neighborhood of τ0\tau_{0} we have

Up to the technical factor of N8\mspace1.0muεN^{8\mspace{1.0mu}\varepsilon} the right hand side coincides with the right hand side of the cubic equation defining E\mathcal{E}. On the other hand, the right hand side of the cubic equation (4.10) for the quantity Θ\Theta from Theorem 4.2 is of the same form as the left hand side of (4.33). Therefore, we infer

We will argue that on appropriately chosen domains out of the three summands in the cubic expression in Θ\Theta always one is the biggest by far. Therefore, the error function E\mathcal{E}, defined by (4.30), is essentially the best bound on Θ\Theta that one may hope to deduce from (4.34). Indeed, since Θ\Theta is by definition an average of g−m\mathbf{g}-\mathbf{m}, we expect Θ≺E\Theta\prec\mathcal{E}.

We will now prove (4.26). To this end we gradually improve the bound on Θ\Theta. Fix some ε∈(0,ε~)\varepsilon\in(0,\widetilde{\varepsilon}). The sequence of deterministic bounds on this quantity is defined as

From here on until the end of this section the threshold function N0N_{0} from the definition of the stochastic domination (cf. Definition 1.6) as well as the definition of ’a.w.o.p.’ (cf. Definition 1.9) may depend on ε\varepsilon in addition to pp, PP, LL, μ‾\underline{\mu} and γ\gamma. At the end of the proof we will remove this dependence. The following lemma is essential for doing one step in the upcoming iteration.

Then Θ(z)≺\mspace2.0muΦk+1(ω,η)\Theta(z)\prec\mspace{2.0mu}\Phi_{k+1}(\omega,\eta).

where N−1N^{-1} from (3.15) has been neglected since ρ≳η\rho\gtrsim\eta. In this way we see that the hypothesis (4.36) of Lemma 4.5 is satisfied. Using the lemma the bound on Θ\Theta is improved to

where w~=Tw\widetilde{\mathbf{w}}=\mathbf{T}\mathbf{w} is a bounded, ∥w~∥∞≲1\lVert\widetilde{\mathbf{w}}\rVert_{\infty}\lesssim 1, deterministic vector. Together with the bound (4.37) we apply the fluctuation averaging (Theorem 3.5) again,

We repeat this step finitely many times and each time improve Φk\Phi_{k} by a factor of N−εN^{-\varepsilon} until it reaches its target value N9εEN^{9\varepsilon}\mathcal{E} and is not improved anymore. Note that all constants in our estimates, explicit and hidden, depend only on the model parameters and ε\varepsilon. In particular, the number of steps needed is uniform in (ω,η)(\omega,\eta). At that stage we have

Finally, with the help of (4.9), (4.41) and the fluctuation averaging, we prove the bound on averages of g−m\mathbf{g}-\mathbf{m} against any bounded, ∥w∥∞≤1\lVert\mathbf{w}\rVert_{\infty}\leq 1, deterministic vector,

This finishes the proof of Theorem 1.7 apart from the proof of Lemma 4.5 which we will tackle now.

Edge: If τ0∈∂supp⁡ρ\tau_{0}\in\partial\operatorname{supp}\rho is an edge of a gap of size Δ:=Δ0(τ0)\Delta:=\Delta_{0}(\tau_{0}), then we define

If any of the two regimes Il(ω)I_{l}(\omega) with l=2,3l=2,3 consists of a single point only, then we set Il(ω):=∅I_{l}(\omega):=\emptyset.

If I3(ω)I_{3}(\omega) consists of a single point only, then we set I3(ω):=∅I_{3}(\omega):=\emptyset.

In the cubic equation (4.30), used to define the error function E\mathcal{E}, the coefficients π~1\widetilde{\pi}_{1} and π~2\widetilde{\pi}_{2} on the left hand side are monotonously increasing functions of η\eta. The linear and the constant coefficient of E\mathcal{E} on the right hand side are monotonously decreasing in η\eta. Thus, E\mathcal{E} itself is a monotonously decreasing function of η\eta. From this fact and the definition of the regimes I1I_{1}, I2I_{2} and I3I_{3} we see that I1=[η1,δ∗]I_{1}=[\eta_{1},\delta_{*}], I2=[η2,η1]I_{2}=[\eta_{2},\eta_{1}] and I3=[Nγ−1,η2]I_{3}=[N^{\gamma-1},\eta_{2}] for some η1,η2∈[Nγ−1,δ∗]\eta_{1},\eta_{2}\in[N^{\gamma-1},\delta_{*}]. Here, we interpret I2=∅I_{2}=\emptyset if η1≤η2\eta_{1}\leq\eta_{2} and I3=∅I_{3}=\emptyset if η2≤Nγ−1\eta_{2}\leq N^{\gamma-1}.

Now we define a zz-dependent indicator function

This function fixes the values of Θ\Theta to a small interval just below the deterministic control parameter Φk\Phi_{k}. We will prove that Θ\Theta cannot take these values, i.e. χ=0\chi=0 a.w.o.p.. Figure 4.1 illustrates this argument. Compared to Figure 6.1 in we see that instead of two there are now three domains, I1(ω)I_{1}(\omega), I2(ω)I_{2}(\omega) and I3(ω)I_{3}(\omega), to be distinguished. The reason for this extra complication is that (4.10) is cubic in Θ\Theta, compared to the quadratic equation for [v][v] that appeared in the proof of Lemma 6.2 in . To see that χ=0\chi=0, first note that the choice of the domains, IlI_{l}, ensures that there is always one summand on the left hand side of the cubic equation (4.10) for Θ\Theta which dominates the two others by a factor NεN^{\varepsilon}, whenever χ\chi does not vanish. In fact, by construction we have: Claim: The random functions Θ\Theta and χ\chi satisfy a.w.o.p.

We will verify this fact at the end of the proof of this lemma. Now we will simply use it. First we combine the assumption (4.36) of the lemma and (4.43) to obtain

Here we also gave up a factor of NεN^{\varepsilon} to get an inequality instead of the stochastic domination, and replaced ρ\rho by the comparable quantity ρ~\widetilde{\rho}. By the definition of the indicator function χ\chi we have Θχ≥N−7\mspace1.0muεΦk\Theta\chi\geq N^{-7\mspace{1.0mu}\varepsilon}\Phi_{k}. Using this to bound the left hand side, and that ε≤ε~\varepsilon\leq\widetilde{\varepsilon}, we obtain

Comparing this with the defining equation (4.30) for E\mathcal{E} we conclude that a.w.o.p. N−8\mspace1.0muεΦk\mspace2.0muχ≤EN^{-8\mspace{1.0mu}\varepsilon}\Phi_{k}\mspace{2.0mu}\chi\leq\mathcal{E}.

On the other hand, by the definition of Φk\Phi_{k} in (4.35) we know that Φk>N8\mspace1.0muεE\Phi_{k}>N^{8\mspace{1.0mu}\varepsilon}\mathcal{E}. These two inequalities yield

where as explained after the definition of I1I_{1}, I2I_{2} and I3I_{3} above we have I1=[η1,δ∗]I_{1}=[\eta_{1},\delta_{*}], I2=[η2,η1]I_{2}=[\eta_{2},\eta_{1}] and I3=[Nγ−1,η2]I_{3}=[N^{\gamma-1},\eta_{2}]. The condition (A.1) of the lemma is satisfied by the definition of Θ\Theta in (4.7), the Hölder-continuity of the solution of the QVE, the weak Lipschitz-continuity of g\mathbf{g} with Lipschitz-constant N\mspace1.0mu2N^{\mspace{1.0mu}2} and the Hölder-continuity of s\mathbf{s} from (4.12). The gap condition, (A.2), holds because of (4.44) and the definition of χ\chi and Φ\Phi for an appropriate choice of the exponent D3D_{3}.

This finishes the proof of Lemma 4.5 up to verifying the claim (4.43). Proof of the claim: For the proof of (4.43) one verifies case by case that on I1I_{1} the term π~1Θ∼∣π1∣Θ\widetilde{\pi}_{1}\Theta\sim|\pi_{1}|\Theta is bigger than the two other terms, π~2Θ2\widetilde{\pi}_{2}\Theta^{2} and Θ3\Theta^{3} by a factor of NεN^{\varepsilon}. If I3I_{3} is not empty then the term Θ3\Theta^{3} is the biggest. If I2I_{2} is not empty, then ∣π2∣∼π~2|\pi_{2}|\sim\widetilde{\pi}_{2} and π~2Θ2\widetilde{\pi}_{2}\Theta^{2} is the biggest term by a factor of NεN^{\varepsilon}. More specifically, when η∈Ij\eta\in I_{j} and χ=χ(ω,η)=1\chi=\chi(\omega,\eta)=1 we show

where π3=π~3:=1\pi_{3}=\widetilde{\pi}_{3}:=1. As an example we demonstrate these relations in a few cases:

Well inside a gap: If τ0∈∂supp⁡ρ\tau_{0}\in\partial\operatorname{supp}\rho and ω∈[c∗Δ,Δ/2]\omega\in[c_{*}\Delta,\Delta/2] then I2(ω)=∅I_{2}(\omega)=\emptyset. We now check that on I1(ω)I_{1}(\omega) the linear term in Θ\Theta is the biggest while on I3(ω)I_{3}(\omega) the cubic term dominates. First, let η∈I1(ω)\eta\in I_{1}(\omega). Then the following chain of inequalities hold,

Here, we used (4.29), (4.15b), the definition of I1(ω)I_{1}(\omega) and (4.27c) in the form π~2∼(Δ+η)1/3\widetilde{\pi}_{2}\sim(\Delta+\eta)^{1/3}. Now we can use χ\chi to replace Φk\Phi_{k} by Θ\Theta. By definition of χ\chi and since π~k≳∣πk∣\widetilde{\pi}_{k}\gtrsim|\pi_{k}| for k=1,2k=1,2 we also get

We conclude that on I1(ω)I_{1}(\omega) the linear term in Θ\Theta dominates the others,

Suppose now that η∈I3(ω)\eta\in I_{3}(\omega). In this case, using the choice of the indicator function χ\chi,

By definition of I3(ω)I_{3}(\omega) and (4.27c) we find that

Altogether we find that the cubic term dominates the two others,

Inside a gap close to an edge on I2I_{2}: If τ0∈∂supp⁡ρ\tau_{0}\in\partial\operatorname{supp}\rho, ω∈[0,c∗Δ]\omega\in[0,c_{*}\Delta] and η∈I2(ω)\eta\in I_{2}(\omega), then we will show the quadratic term in Θ\Theta dominates the two other terms. We have

where in the inequality we used the definition of I2(ω)I_{2}(\omega). The choice of χ\chi guarantees that Φk\mspace2.0muχ≥N3\mspace1.0muε\mspace1.0muΘ\mspace2.0muχ\Phi_{k}\mspace{2.0mu}\chi\geq N^{3\mspace{1.0mu}\varepsilon}\mspace{1.0mu}\Theta\mspace{2.0mu}\chi. Thus, the quadratic term is larger than the cubic term by a factor of NεN^{\varepsilon}. On the other hand

Here, in the first inequality we used the indicator function χ\chi and in the second inequality the definition of I2(ω)I_{2}(\omega). Altogether, we arrive at

Here, we used (4.29) and the definitions of π~1\widetilde{\pi}_{1} and I1(ω)I_{1}(\omega), respectively. Since Φk\mspace2.0muχ≥N6\mspace1.0muεΘ\mspace2.0muχ\Phi_{k}\mspace{2.0mu}\chi\geq N^{6\mspace{1.0mu}\varepsilon}\Theta\mspace{2.0mu}\chi and by the definition of π~2\widetilde{\pi}_{2} this shows that the linear term is larger than the quadratic term by a factor of N4\mspace1.0muεN^{4\mspace{1.0mu}\varepsilon}. In order to compare the linear with the cubic term we estimate further. By definition of I1(ω)I_{1}(\omega),

Again we use the lower bound on Φk\mspace2.0muχ\Phi_{k}\mspace{2.0mu}\chi and get

Thus we showed that on the domain I1(ω)I_{1}(\omega)

The other cases are proven similarly. This completes the proof of (4.43). ∎

Rigidity and delocalization of eigenvectors

Here we explain how the local law, Theorem 1.7, is used to estimate the difference between the cumulative density of states and the eigenvalue distribution function of the random matrix H\mathbf{H}. The following auxiliary result shows that the difference between two probability measures can be estimated in terms of the difference of their respective Stieltjes transforms. For completeness the proof is given in the appendix. It uses a Cauchy-integral formula that was also applied in the construction of the Helffer-Sjöstrand functional calculus (cf. ) and it appeared in different variants in , and .

Here, the three contributions to the error, J1J_{1}, J2J_{2} and J3J_{3}, are defined as

where mνm_{\nu} denotes the Stieltjes transform of ν\nu for any signed measure ν\nu.

We will now apply this lemma to prove Corollary 1.10 with the choices of the measures

As a first step we show that a.w.o.p. there are no eigenvalues with an absolute value larger or equal than 1010, i.e.,

We focus on the eigenvalues λi≥10\lambda_{i}\geq 10. The ones with λi≤−10\lambda_{i}\leq-10 are treated in the same way. We will show first that there are no eigenvalues in a small interval around τ\tau with τ≥10\tau\geq 10. In fact, we prove that for γ∈(0,1/3)\gamma\in(0,1/3),

For this we apply Lemma 5.1 with the same choices of the measures ν1\nu_{1} and ν2\nu_{2} as in (5.3) and with

Plugging this bound into the definitions of J1J_{1}, J2J_{2} and J3J_{3} from (5.2) and using (5.1) and the fact that ρ=0\rho=0 in this regime shows the validity of (5.5).

We conclude that a.w.o.p. there are no eigenvalues in an interval of length N−1N^{-1} to the right of τ\tau. By using a union bound this implies that

The eigenvalues larger than NN are treated by the following simple argument,

Now we apply Lemma 5.1 to prove (1.28). In case ∣τ∣≥10|\tau|\geq 10 the bound (1.28) follows because a.w.o.p. there are no eigenvalues of H\mathbf{H} with absolute value larger or equal than 1010. Thus, we fix τ∈(−10,10)\tau\in(-10,10) and make the choices

Again we use (1.21) from Theorem 1.7, the Lipschitz-continuity of ⟨g⟩\langle\mathbf{g}\rangle and the Hölder-continuity of ⟨m⟩\langle\mathbf{m}\rangle to see that uniformly for all η≥Nγ−1\eta\geq N^{\gamma-1},

Here we evaluated Δ(τ1)=1\Delta(\tau_{1})=1 and thus κ≲η+(Nη)−1\kappa\lesssim\eta+(N\eta)^{-1}. With J1J_{1} defined as in (5.2) we infer J1≺N−1J_{1}\prec N^{-1}. Theorem 1.7 also implies the bound

since in this regime κ≲1\kappa\lesssim 1, thus showing that J3≺N−1J_{3}\prec N^{-1}. We are left with estimating the three terms constituting J2J_{2}. The first and second of these terms are estimated trivially by using the boundedness of their integrands. Therefore, we conclude that

This expression is derived by using the bound (1.23) on κ\kappa for the integrand of the third contribution to J2J_{2}.

for any ω∈[0,Nγ−1]\omega\in[0,N^{\gamma-1}] and η∈[Nγ−1,2]\eta\in[N^{\gamma-1},2]. With this the size of RR is given by

The bound (5.11) follows by performing the integration over η\eta.

This finishes the proof of (5.11). We insert this bound into (5.9) and use that γ\gamma was arbitrary. Thus, we find

This finishes the proof of (1.28) since there are no eigenvalues below −10-10.

where δk\delta_{k} are defined as in (1.30) and δ0=Nγ−2/3\delta_{0}=N^{\gamma-2/3}. Note that there is nothing to show if k>1k>1 and the size of the gap, αk−βk−1\alpha_{k}-\beta_{k-1}, is smaller than 2\mspace1.0muδk2\mspace{1.0mu}\delta_{k}, i.e., if such a τ\tau does not exist. In particular, we have αk−βk−1=Δ(τ)≳N−1/2\alpha_{k}-\beta_{k-1}=\Delta(\tau)\gtrsim N^{-1/2}. We will show that a.w.o.p. there are no eigenvalues in an interval of length N−2/3N^{-2/3} to the right of τ\tau, i.e.

We apply Lemma 5.1 with the same choices of the measures ν1\nu_{1} and ν2\nu_{2} as in (5.3). Additionally, we set

We use the local law, Theorem 1.7, to estimate the differences between the Stieltjes transforms of the two measures for the integrands in the definition of the three error terms, J1J_{1}, J2J_{2} and J3J_{3} from (5.2). By the definition of δk\delta_{k} the condition (1.24) is satisfied inside the integrals and we use the improved bound, (1.25), on κ\kappa. Indeed, we find

where the supremum is taken over ω∈[τ−N−2/3,τ+2N−2/3]\omega\in[\tau-N^{-2/3},\tau+2N^{-2/3}] and η∈[N−2/3,2N−2/3]\eta\in[N^{-2/3},2N^{-2/3}]. With this, the definition of δk\delta_{k} and the size of ρ\rho from (4.5c) and (4.5d) we infer

From this (5.12) follows. The claim, (1.29), is now a consequence of a simple union bound taken over the events in (5.12) with different choices of τ\tau. This finishes the proof of Corollary 1.10. ∎

2 Proof of Corollary 1.11

Here we show how we get the rigidity, Corollary 1.11, from Corollary 1.10. Fix a τ∈[α1,βK]\tau\in[\alpha_{1},\beta_{K}]. We define the random fluctuation to the left, δ−\delta_{-}, and to the right, δ+\delta_{+}, of the eigenvalue λi(τ)\lambda_{i(\tau)} as

We start with the upper bound on λi(τ)\lambda_{i(\tau)}. By the definition of i(τ)i(\tau) we find the inequality

The definition of δ+=δ+(τ)\delta_{+}=\delta_{+}(\tau) implies that

By monotonicity of the cumulative eigenvalue distribution, we conclude that λi(τ)≤τ+δ+\lambda_{i(\tau)}\leq\tau+\delta_{+}. Thus, the upper bound is proven.

Now we show the lower bound. We start similarly,

Here the lim inf⁡\liminf is necessary, since the cumulative eigenvalue distribution is not continuous from the left. We conclude that λi(τ)≥τ−δ−−ε\lambda_{i(\tau)}\geq\tau-\delta_{-}-\varepsilon for all ε>0\varepsilon>0 and therefore the lower bound is proven.

Now we start with the proof of (1.34). For this we show that for any τ\tau that is well inside the support of the density of states, i.e., that satisfies (1.33), we have

We apply (1.28) from Corollary 1.10 and, using (4.5e) again, we get

We will now verify that for large enough NN,

We distinguish three cases. First let us consider the regime where ρ(τ0)3+∣τ−τ0∣≤N−3/5\rho(\tau_{0})^{3}+|\tau-\tau_{0}|\leq N^{-3/5}. Then we have δ=N−3/5\delta=N^{-3/5} and

Now we treat the situation where, N−3/5<ρ(τ0)3+∣τ−τ0∣≤N3ε/2−3/5N^{-3/5}<\rho(\tau_{0})^{3}+|\tau-\tau_{0}|\leq N^{3\varepsilon/2-3/5}. In this case

Finally, we consider ρ(τ0)3+∣τ−τ0∣>N3ε/2−3/5\rho(\tau_{0})^{3}+|\tau-\tau_{0}|>N^{3\varepsilon/2-3/5}. Then for large enough NN we find on the one hand

Thus, (5.19) holds true and since ε\varepsilon was arbitrary, we infer from (5.17) and (5.18) that δ+(τ)≺δ\delta_{+}(\tau)\prec\delta. Along the same lines we prove δ−(τ)≺δ\delta_{-}(\tau)\prec\delta. Thus (5.16) and with it (1.34) are proven.

The statement about the fluctuation of the eigenvalues at the leftmost edge, (1.35) follows directly from (1.34) and (1.29) in Corollary 1.10. Indeed, for τ∈[α1,α1+ε0)\tau\in[\alpha_{1},\alpha_{1}+\varepsilon_{0}) we have λi(τ)≤λi(α1+ε0)\lambda_{i(\tau)}\leq\lambda_{i(\alpha_{1}+\varepsilon_{0})} and from (1.34) with Δ(τ)=1\Delta(\tau)=1, as well as ρ(α1+ε0)∼ε01/2\rho(\alpha_{1}+\varepsilon_{0})\sim\varepsilon_{0}^{1/2}, and from the definition of ε0\varepsilon_{0} we see that

On the other hand, (1.29) shows that a.w.o.p. λi(τ)≥α1−Nγ−2/3\lambda_{i(\tau)}\geq\alpha_{1}-N^{\gamma-2/3}. Since γ\gamma was arbitrary, (1.35) follows. The rigidity at the rightmost edge is proven along the same lines.

The claim, (1.36), about the remaining eigenvalues follows from a similar argument. For τ∈(βk−εk,αk+1+εk)\tau\in(\beta_{k}-\varepsilon_{k},\alpha_{k+1}+\varepsilon_{k}), as a consequence of (1.29), we have

From (1.34) and the definition of εk\varepsilon_{k} we infer λi(βk−εk)≥βk−2\mspace1.0muεk\lambda_{i(\beta_{k}-\varepsilon_{k})}\geq\beta_{k}-2\mspace{1.0mu}\varepsilon_{k} a.w.o.p., as well as λi(αk+1+εk)≤αk+1+2\mspace1.0muεk\lambda_{i(\alpha_{k+1}+\varepsilon_{k})}\leq\alpha_{k+1}+2\mspace{1.0mu}\varepsilon_{k} a.w.o.p., which finishes the proof of (1.36). ∎

3 Proof of Corollary 1.14

The delocalization of eigenvectors is a simple consequence of the anisotropic local law Theorem 1.13 using the argument from . Expressing the resolvent in the eigenbasis, we have

by keeping only a single summand i=ki=k from (5.20). As γ>0\gamma>0 was arbitrary we conclude that

Anisotropic law and universality

The first term containing the diagonal elements GiiG_{ii} is clearly bounded by the right hand side of (1.37) by Theorem 1.7. This is the first instant where the nontrivial ii-dependence of mim_{i} is used.

The main technical part of the proof in is then to control Z\mathcal{Z}, the contribution of the off diagonal terms. We can follow this proof in our case to the letter; the nontrivial ii-dependence of mim_{i} requires a slight modification only at one point. To see this, we recall the main structure of the proof. For any even pp, the moment

This formula replaces (5.41) from . Taking the inverse of this formula and expanding around the leading term mbm_{b}, we get a geometric series representation for Gbb(B\b)G^{(B\backslash b)}_{bb} in terms of powers of the last three term in (6.2). The resulting formula is analogous to (5.42) in . The geometric series converges because the last three term on the right hand side of (6.2) are much smaller than ∣1/mb∣∼1\lvert 1/m_{b}\rvert\sim 1 a.w.o.p.. Indeed, the last two terms in (6.2) are of size N−1N^{-1} and N−1/2+cN^{-1/2+c} a.w.o.p., respectively. The double sum in (6.2) is small by using the large deviation estimates (2.7a)–(2.7c), similarly as in the proof of Lemma 2.1. When estimating the diagonal sum i=ji=j, we note that ∣Gii(B)−mi∣\lvert G^{(B)}_{ii}-m_{i}\rvert is small by first estimating ∣Gii(B)−Gii∣\lvert G^{(B)}_{ii}-G_{ii}\rvert similarly to (2.12), and then we use the local law Theorem 1.7 to see that also ∣Gii−mi∣\lvert G_{ii}-m_{i}\rvert is small.

The proof in did not use the specific form of the subtracted term sbimiδijs_{bi}m_{i}\delta_{ij} in (6.2), just the fact that the subtraction made (2.7c) applicable for the double summation in (6.2). After this slight modification, the rest of the proof in goes through without any further changes. ∎

2 Proof of Theorem 1.16

For the proof of Theorem 1.16 we follow the method developed in . Theorem 2.1 from was designed for proving universality for a random matrix with a small independent Gaussian component and densities of state that may differ from Wigner’s semicircle law. The main theorem in asserts that if local laws hold in a sufficiently strong sense then bulk universality holds locally for matrices with a small Gaussian component. We remark that a similar approach was independently developed in that can also be easily used to conclude bulk universality from Theorem 1.7, but here we follow . In Section 2.5 of a recipe was given how to use this theorem to establish universality for a quite general class of random matrix models even without the Gaussian component, as long as uniform local laws on the optimal scale are known and the matrix satisfies the appropriate qq-fullness condition (cf. Definition 1.15) that allows for an application of the moment matching (Lemma 6.5 in ) and the Green’s function comparison theorem (Theorem 2.3 in ).

Appendix A Appendix

The relation ≺\prec is transitive and it satisfies the following arithmetic rules:

If ϕ≺Nδψ\phi\prec N^{\delta}\psi, for every δ>0\delta>0, then ϕ≺ψ \phi\prec\psi\,;

If ϕ≺N−δϕ+ψ\phi\prec N^{-\delta}\phi+\psi, for some δ>0\delta>0, then ϕ≺ψ \phi\prec\psi\,.

These properties follow directly from the definition (Definition 1.6) of stochastic domination. For further details see .

Then the sequence φ\varphi satisfies the bound

For i=0i=0 this follows from (A.3) and (A.2). For all other ii it follows by induction using the continuity condition (A.1), which implies ∣φ(zi+1)−φ(zi)∣+∣Φ(zi+1)−Φ(zi)∣≤N−D3/2|\varphi(z_{i+1})-\varphi(z_{i})|+|\Phi(z_{i+1})-\Phi(z_{i})|\leq N^{-D_{3}}/2. This shows that if φ(zi)≤Φ(zi)−N−D3\varphi(z_{i})\leq\Phi(z_{i})-N^{-D_{3}}, then φ(zi+1)≤Φ(zi+1)\varphi(z_{i+1})\leq\Phi(z_{i+1}) and with (A.2) even that φ(zi+1)≤Φ(zi+1)−N−D3\varphi(z_{i+1})\leq\Phi(z_{i+1})-N^{-D_{3}}. In particular, φ(z)≤Φ(z)−N−D3\varphi(z)\leq\Phi(z)-N^{-D_{3}} a.w.o.p..

where the three integrals L1L_{1}, L2L_{2} and L3L_{3} are given as

and mνm_{\nu} is the Stieltjes transform of ν\nu.

We split the integral, L1L_{1}, into the contributions,

For the treatment of L1,1,>L_{1,1,>} we integrate by parts, first in σ\sigma and then in η\eta,

We use max⁡η∣χ(η)+η\mspace1.0muχ′(η)∣≲1\max_{\eta}|\chi(\eta)+\eta\mspace{1.0mu}\chi^{\prime}(\eta)|\lesssim 1 and max⁡σ∈[\mspace1.0mu0\mspace1.0mu,\mspace2.0muη1]∣f′(τ1−σ)∣≲η1−1\max_{\sigma\in[\mspace{1.0mu}0\mspace{1.0mu},\mspace{2.0mu}\eta_{1}]}|f^{\prime}(\tau_{1}-\sigma)|\lesssim\eta_{1}^{-1}. In this way we estimate for ν=ν1−ν2\nu=\nu_{1}-\nu_{2},

Going through the same steps we also arrive at

We continue by estimating L2L_{2} from above.

Finally we derive a bound for L3L_{3}. We split the integral into two components,

We combine this with the estimates from (A.5), (A.6), (A.7), (A.8) and (A.9). Altogether we have

where the three terms on the right hand side are given by

Now we use this bound for the smoothed out indicator function to derive a bound on the difference of number of eigenvalues in the interval [τ1,τ2][\tau_{1},\tau_{2}] and the predicted number, given by the integral over the density of states. We use

References