On the principal components of sample covariance matrices

Alex Bloemendal, Antti Knowles, Horng-Tzer Yau, Jun Yin

Introduction

In this paper we investigate M×MM\times M sample covariance matrices of the form

Correspondingly, one has to subtract from AiμA_{i\mu} the empirical mean of the ii-th row of AA, which we denote by [A]i:=1N∑μ=1NAiμ[A]_{i}\mathrel{\mathop{:}}=\frac{1}{N}\sum_{\mu=1}^{N}A_{i\mu}. Hence, we replace (1.1) with

By the law of large numbers, if MM is fixed and NN taken to infinity, the sample covariance matrix Q\mathcal{Q} converges almost surely to the population covariance matrix Σ\Sigma. In many modern applications, however, the population size MM is very large and obtaining samples is costly. Thus, one is typically interested in the regime where MM is of the same order as NN, or even larger. In this case, as it turns out, the behaviour of Q\mathcal{Q} changes dramatically and the problem becomes much more difficult. In principal component analysis, one seeks to understand the correlations by considering the principal components, i.e. the top eigenvalues and associated eigenvectors, of Q\mathcal{Q}. These provide an effective low-dimensional projection of the high-dimensional data set AA, in which the significant trends and correlations are revealed by discarding superfluous data.

the empirical eigenvalue density of the rescaled matrix Q=ϕ−1/2QQ=\phi^{-1/2}\mathcal{Q} has the same asymptotics for large MM and NN as

2. Examples and outline of the model

The population covariance matrix is Σ=TT∗=IM+UU∗\Sigma=TT^{*}=I_{M}+UU^{*}.

Below we shall refer to these examples as Examples (1) and (2) respectively. Motivated by them, we now outline our model. Let BB be an (M+r)×N(M+r)\times N matrix whose entries are independent with zero mean and unit variance. We choose a deterministic M×(M+r)M\times(M+r) matrix TT, and set Q=1NTBB∗T∗\mathcal{Q}=\frac{1}{N}TBB^{*}T^{*}. We stress that we do not assume that the underlying randomness is Gaussian. Our key assumptions are (i) rr is bounded; (ii) Σ−IM\Sigma-I_{M} has bounded rank; (iii) log⁡N\log N is comparable to log⁡M\log M; (iv) the entries of BB are independent, with zero mean and unit variance, and have a sufficient number of bounded moments. The precise assumptions are given in Section 1.3 below. We emphasize that everything apart from rr and the rank of Σ−IM\Sigma-I_{M} is allowed to depend on NN in an arbitrary fashion.

3. Definition of model

In this section we give the precise definition of our model and introduce some basic notations. For convenience, we always work with the rescaled sample covariance matrix

The motivation behind this rescaling is that, as observed in (1.6), it ensures that the bulk spectrum of QQ has asymptotically a fixed diameter, 44, for arbitrary NN and MM.

We always regard NN as the fundamental large parameter, and write M≡MNM\equiv M_{N}. Here, and throughout the following, in order to unclutter notation we omit the argument NN in quantities, such as MM, that depend on it. In other words, every symbol that is not explicitly a constant is in fact a sequence indexed by NN. We assume that MM and NN satisfy the bounds

Fix a constant r=0,1,2,3,…r=0,1,2,3,\dots. Let XX be an (M+r)×N(M+r)\times N random matrix and TT an M×(M+r)M\times(M+r) deterministic matrix. For definiteness, and bearing the motivation of sample covariance matrices in mind, we assume that the entries of XX and TT are real. However, our method also trivially applies to complex-valued XX and TT, with merely cosmetic changes to the proofs. We consider the M×MM\times M matrix

Since TXTX is an M×NM\times N matrix, we find that QQ has

We define the population covariance matrix

for the eigenvalues σi\sigma_{i}. We always order the values did_{i} such that

We suppose that Σ\Sigma is positive definite, so that each did_{i} lies in the interval

Moreover, we suppose that Σ−IM\Sigma-I_{M} has bounded rank, i.e.

We assume that the entries XiμX_{i\mu} of XX are independent (but not necessarily identically distributed) random variables satisfying

Our results concern the eigenvalues of QQ, denoted by

and the associated unit eigenvectors of QQ, denoted by

4. Sketch of behaviour of the principal components of Q𝑄Q

To guide the reader, we now give a heuristic description of the behaviour of principal components of QQ. The spectrum of QQ consists of a bulk spectrum and of outliers—eigenvalues separated from the bulk. The bulk contains an order KK eigenvalues, which are distributed on large scales according to the Marchenko-Pastur law (1.6). In addition, if ϕ>1\phi>1 there are M−KM-K trivial eigenvalues at zero. Each did_{i} satisfying ∣di∣>1\lvert d_{i}\rvert>1 gives rise to an outlier located near its classical location

Any did_{i} satisfying ∣di∣<1\lvert d_{i}\rvert<1 does not result in an outlier. We summarize this picture in Figure 1.1.

The creation or annihilation of an outlier as a did_{i} crosses ±1\pm 1 is known as the BBP phase transition BBP . It takes place on the scaleWe use the symbol ≍\asymp to denote quantities of comparable size; see “Conventions” at the end of this section for a precise definition. \bigl{\lvert}\lvert d_{i}\rvert-1\bigr{\rvert}\asymp K^{-1/3}. This scale has a simple heuristic explanation (we focus on the right edge of the spectrum). Suppose that d1∈(0,1)d_{1}\in(0,1) and all other did_{i}’s are zero. Then the top eigenvalue μ1\mu_{1} exhibits universality, and fluctuates on the scale K−2/3K^{-2/3} around γ+\gamma_{+} (see Theorem 8.3 and Remark 8.7 below). Increasing d1d_{1} beyond the critical value 11, we therefore expect μ1\mu_{1} to become an outlier when its classical location θ(d1)\theta(d_{1}) is located at a distance greater than K−2/3K^{-2/3} from γ+\gamma_{+}. By a simple Taylor expansion of θ\theta, the condition θ(d1)−γ+≫K−2/3\theta(d_{1})-\gamma_{+}\gg K^{-2/3} becomes d1−1≫K−1/3d_{1}-1\gg K^{-1/3}.

for di>1d_{i}>1. The function uu determines the aperture 2arccos⁡u(di)2\arccos\sqrt{u(d_{i})} of the cone. Note that u(di)∈(0,1)u(d_{i})\in(0,1) and u(di)u(d_{i}) converges to 11 as di→∞d_{i}\to\infty. See Figure 1.2.

5. Summary of previous related results

There is an extensive literature on spiked covariance matrices. So far most of the results have focused on the outlier eigenvalues of Example (1), with the nonzero did_{i} independent of NN and ϕ\phi fixed. Eigenvectors and the non-outlier eigenvalues have seen far less attention.

For the uncorrelated case Σ=IM\Sigma=I_{M} and Gaussian XX in (1.10) with fixed ϕ\phi, it was proved in Joh1 for the complex case and in Johnstone for the real case that the top eigenvalue, rescaled as K2/3(μ1−γ+)K^{2/3}(\mu_{1}-\gamma_{+}), is asymptotically distributed according the Tracy-Widom law of the appropriate symmetry class TW1 ; TW2 . Subsequently, these results were shown to be universal, i.e. independent of the distribution of the entries of XX, in Sosh2 ; PY . The assumption that ϕ\phi be fixed was relaxed in Pe1 ; ElK1 .

The study of covariance matrices with nontrivial population covariance matrix Σ≠IM\Sigma\neq I_{M} goes back to the seminal paper of Johnstone Johnstone , where the Gaussian spiked model was introduced. The BBP phase transition was established by Baik, Ben Arous, and Péché BBP for complex Gaussian XX, fixed rank of Σ−IM\Sigma-I_{M}, and fixed ϕ\phi. Subsequently, the results of BBP were extended to the other Gaussian symmetry classes, such as real covariance matrices, in BV1 ; BV2 . The proofs of BBP ; Pec use an asymptotic analysis of Fredholm determinants, while those of BV1 ; BV2 use an explicit tridiagonal representation of XX∗XX^{*}; both of these approaches rely heavily on the Gaussian nature of XX. See also PeBor for a generalization of the BBP phase transition.

For the model from Example (1) with fixed nonzero {di}\{d_{i}\} and ϕ\phi, the almost sure convergence of the outliers was established in BS . It was also shown in BS that if ∣di∣<1\lvert d_{i}\rvert<1 for all ii, the top eigenvalue μ1\mu_{1} converges to γ+\gamma_{+}. For this model, a central limit theorem of the outliers was proved in BY1 . In BY2 , the almost sure convergence of the outliers was proved for a generalized spiked model whose population covariance matrix is of the block diagonal form Σ=diag⁡(A,T)\Sigma=\operatorname{diag}(A,T), where AA is a fixed r×rr\times r matrix and TT is chosen so that the associated sample covariance matrix has no outliers.

In BGN , the almost sure convergence of the projection of the outlier eigenvectors onto the finite-dimensional spike subspace was established, under the assumption that ϕ\phi and the nonzero did_{i} are fixed, and that BB and TT are both random and one of them is orthogonally invariant. In particular, the cone concentration from (1.18) was established in BGN . In Paul , under the assumption that XX is Gaussian and ϕ\phi and the nonzero did_{i} are fixed, a central limit theorem for a certain observable, the so-called sample vector, of the outlier eigenvectors was established. The result of Paul was extended to non-Gaussian entries for a special class of Σ\Sigma in Shi .

Moreover, in BGN2 ; NA1 result analogous to those of BGN were obtained for the model from Example (2). Finally, a related class of models, so-called deformed Wigner matrices, have been the subject of much attention in recent years; we refer to SoshPert ; SoshPert2 ; KY2 ; KY3 for more details; in particular, the joint distribution of all outliers was derived in KY3 .

6. Overview of results

In this subsection we give an informal overview of our results.

Our results on the eigenvalues of QQ consist of two parts. First, we derive large deviation bounds on the locations of the outliers (Theorem 2.3). Second, we prove eigenvalue sticking for the non-outliers (Theorem 2.7), whereby each non-outlier “sticks” with high probability and very accurately to the eigenvalues of a related covariance matrix satisfying Σ=IM\Sigma=I_{M} and whose top eigenvalues exhibit universality. As a corollary (Remark 8.7), we prove that the top non-outlier eigenvalue of QQ has asymptotically the Tracy-Widom-1 distribution. This sticking is very accurate if all did_{i}’s are separated from the critical point 11, and becomes less accurate if a did_{i} is in the vicinity of 11. Eventually, it breaks down precisely on the BBP transition scale ∣di−1∣≍K−1/3\lvert d_{i}-1\rvert\asymp K^{-1/3}, at which the Tracy-Widom-1 distribution is known not to hold for the top non-outlier eigenvalue. These results generalize those from (KY2, , Theorem 2.7).

If the outlier μi\mu_{i} approaches the bulk spectrum or another outlier, the cone concentration becomes less accurate. For the case of two nearby outlier eigenvalues, for instance, the cone concentration (1.18) of the eigenvectors breaks down when the distributions of the outlier eigenvalues have a nontrivial overlap. In order to understand this behaviour in more detail, we introduce the deterministic projection

Finally, the proofs of universality of the non-outlier eigenvalues and eigenvectors require the universality of QQ for the uncorrelated case Σ=IM\Sigma=I_{M} as input. This universality result is given in Theorem 8.3, which is also of some independent interest. It establishes the joint, fixed-index, universality of the eigenvalues and eigenvectors of QQ (and hence, as a special case, the quantum unique ergodicity of the eigenvectors of QQ mentioned in Section 1.1). It works for all eigenvalue indices ii satisfying i⩽K1−τi\leqslant K^{1-\tau} for any fixed τ>0\tau>0.

We conclude this subsection by outlining the key novelties of our work.

We introduce the general models QQ from (1.10) and Q˙\dot{Q} from (2.23) below, which subsume and generalize several models considered previously in the literatureIn particular, the current paper is the first to study the principal components of a realistic sample covariance matrix (1.3) instead of the zero mean case (1.1).. We allow the entries of XX to be arbitrary random variables (up to a technical assumption on their tails). All quantities except rr and the rank of Σ−IM\Sigma-I_{M} may depend on NN. We make no assumption on TT beyond the bounded-rank condition of TT∗−IMTT^{*}-I_{M}. The dimensions MM and NN may be wildly different, and are only subject to the technical condition (1.9).

We obtain quantitative bounds (i.e. rates of convergence) on the outlier eigenvalues and the generalized components of the eigenvectors. We believe these bounds to be optimal.

We establish the joint, fixed-index, universality of the eigenvalues and eigenvectors for the case Σ=IM\Sigma=I_{M}. This result holds for any eigenvalue indices ii satisfying i⩽K1−τi\leqslant K^{1-\tau} for an arbitrary τ>0\tau>0. Note that previous works KY1 ; TV3 (established in the context of Wigner matrices) required either the much stronger condition i⩽(log⁡K)Clog⁡log⁡Ki\leqslant(\log K)^{C\log\log K} or a four-moment matching condition.

We remark that the large deviation bounds derived in this paper also allow one to derive the joint distribution of the generalized components of the outlier eigenvectors; this will be the subject of future work.

Conventions

The fundamental large parameter is NN. All quantities that are not explicitly constant may depend on NN; we almost always omit the argument NN from our notation.

We use τ>0\tau>0 in various assumptions to denote a positive constant that may be chosen arbitrarily small. A smaller value of τ\tau corresponds to a weaker assumption. All of our estimates depend on τ\tau, and we neither indicate nor track this dependence.

Results

In this section we state our main results. The following notion of a high-probability bound was introduced in EKY2 , and has been subsequently used in a number of works on random matrix theory. It provides a simple way of systematizing and making precise statements of the form “AA is bounded with high probability by BB up to small powers of NN”.

be two families of nonnegative random variables, where U(N)U^{(N)} is a possibly NN-dependent parameter set. We say that AA is stochastically dominated by BB, uniformly in uu, if for all (small) ε>0\varepsilon>0 and (large) D>0D>0 we have

for large enough N⩾N0(ε,D)N\geqslant N_{0}(\varepsilon,D). Throughout this paper the stochastic domination will always be uniform in all parameters (such as matrix indices) that are not explicitly fixed. Note that N0(ε,D)N_{0}(\varepsilon,D) may depend on the constants from (1.9) and (1.16) as well as any constants fixed in the assumptions of our main results. If AA is stochastically dominated by BB, uniformly in uu, we use the notation A≺BA\prec B. Moreover, if for some complex family AA we have ∣A∣≺B\lvert A\rvert\prec B we also write A=O≺(B)A=O_{\prec}(B).

Because of (1.9), all (or some) factors of NN in Definition (2.1) could be replaced with MM without changing the definition of stochastic domination.

We begin with results on the locations of the eigenvalues of QQ. These results will also serve as a fundamental input for the proofs of the results on eigenvectors presented in Sections 2.2 and 2.3.

Recall that QQ has M−KM-K zero eigenvalues. We shall therefore focus on the KK nontrivial eigenvalues μ1⩾⋯⩾μK\mu_{1}\geqslant\cdots\geqslant\mu_{K} of QQ. On the global scale, the eigenvalues of QQ are distributed according to the Marchenko-Pastur law (1.6). This may be easily inferred from the fact that (1.6) gives the global density of the eigenvalues for the uncorrelated case Σ=IM\Sigma=I_{M}, combined with eigenvalue interlacing (see Lemma 4.1 below). In this section we focus on local eigenvalue information.

As explained in Section 1.4, each i∈Oi\in\mathcal{O} gives rise to an outlier of QQ near the classical location θ(di)\theta(d_{i}) defined in (1.17). In the definition (2.2), the lower bound 1+K−1/31+K^{-1/3} is chosen for definiteness; it could be replaced with 1+aK−1/31+aK^{-1/3} for any fixed a>0a>0. We denote by

the number of outliers to the left (s−s_{-}) and right (s+s_{+}) of the bulk spectrum.

The function Δ(d)\Delta(d) will be used to give an upper bound on the magnitude of the fluctuations of an outlier associated with dd. We give such a precise expression for Δ\Delta in order to obtain sharp large deviation bounds for all d∈D∖d\in\mathcal{D}\setminus. (Note that the discontinuity of Δ\Delta at d=2d=2 is immaterial since Δ\Delta is used as an upper bound with respect to ≺\prec. The ratio of the right- and left-sided limits at 22 of Δ\Delta lies in $$.)

Our result on the outlier eigenvalues is the following.

Fix τ>0\tau>0. Then for i∈Oi\in\mathcal{O} we have the estimate

provided that di>0d_{i}>0 or ∣ϕ−1∣⩾τ\lvert\phi-1\rvert\geqslant\tau.

Furthermore, the extremal non-outliers μs++1\mu_{s_{+}+1} and μK−s−\mu_{K-s_{-}} satisfy

and, assuming in addition that ∣ϕ−1∣⩾τ\lvert\phi-1\rvert\geqslant\tau,

Theorem 2.3 gives large deviation bounds for the locations of the outliers to the right of the bulk. Since τ>0\tau>0 may be arbitrarily small, Theorem 2.3 also gives the full information about the outliers to the left of the bulk except in the case 1>ϕ=1+o(1)1>\phi=1+o(1). Although our methods may be extended to this case as well, we exclude it here to avoid extraneous complications.

By definition of s−s_{-} and D\mathcal{D}, if ϕ>1\phi>1 then s−=0s_{-}=0. Hence, by (2.6), if ϕ>1\phi>1 there are no outliers on the left of the bulk spectrum.

Previously, the model from Example (1) in Section 1.2 with fixed nonzero {di}\{d_{i}\} and ϕ\phi was investigated in BS ; BY1 . In BS , it was proved that each outlier eigenvalue μi\mu_{i} with i∈Oi\in\mathcal{O} convergences almost surely to θ(di)\theta(d_{i}). Moreover, a central limit theorem for μi\mu_{i} was established in BY1 .

The locations of the non-outlier eigenvalues μi\mu_{i}, i∉Oi\notin\mathcal{O}, are governed by eigenvalue sticking, whereby the eigenvalues of QQ “stick” with high probability to eigenvalues of a reference matrix which has a trivial population covariance matrix. The reference matrix is QQ from (1.10) with uncorrelated entries. More precisely, we set

Fix τ>0\tau>0. Then we have for all i∈[ ⁣[1,(1−τ)K] ⁣]i\in[\![{1,(1-\tau)K}]\!]

Similarly, if ∣ϕ−1∣⩾τ\lvert\phi-1\rvert\geqslant\tau then we have for all i∈[ ⁣[τK,K] ⁣]i\in[\![{\tau K,K}]\!]

As outlined above, in Theorem 8.3 below we prove that the asymptotic joint distribution of the non-bulk eigenvalues of HH is universal, i.e. it coincides with that of the Wishart matrix HWish=XX∗H_{\text{\rm{Wish}}}=XX^{*} with r=0r=0 and XX Gaussian. As an immediate corollary of Theorems 2.7 and 8.3, we obtain the universality of the non-outlier eigenvalues of QQ with index i⩽K1−τα+3i\leqslant K^{1-\tau}\alpha_{+}^{3}. This condition states simply that the right-hand side of (2.9) is much smaller than the scale on which the eigenvalue λi\lambda_{i} fluctuates, which is K−2/3i−1/3K^{-2/3}i^{-1/3}. See Remark 8.7 below for a precise statement.

Theorem 2.7 is analogous to Theorem 2.7 of KY2 , where sticking was first established for Wigner matrices. Previously, eigenvalue sticking was established for a certain class of random perturbations of Wigner matrices in BGGM1 ; BGGM2 . We refer to (KY2, , Remark 2.8) for a more detailed discussion.

Aside from holding for general covariance matrices of the form (1.10), Theorem 2.7 is stronger than its counterpart from KY2 because it holds much further into the bulk: in (KY2, , Theorem 2.7), sticking was established under the assumption that i⩽(log⁡K)Clog⁡log⁡Ki\leqslant(\log K)^{C\log\log K}.

The edge universality following from Theorem 2.7 (as explained in Remark 2.8) generalizes the recent result BaoPanZhou2 . There, for the model from Example (1) in Section 1.2 with fixed nonzero {di}\{d_{i}\} and ϕ\phi, it was proved that if di<1d_{i}<1 for all ii and Σ\Sigma is diagonal, then μ1\mu_{1} converges (after a suitable affine transformation) in distribution to the Tracy-Widom-1 distribution.

2. Outlier eigenvectors

For i∈[ ⁣[1,M] ⁣]i\in[\![{1,M}]\!] we define νi⩾0\nu_{i}\geqslant 0 through

For definiteness, we only state our results for the outliers on the right-hand side of the bulk spectrum. Analogous results hold for the outliers on the left-hand side. Since the behaviour of the fluctuating error term is different in the regimes μi−γ+≪1\mu_{i}-\gamma_{+}\ll 1 (near the bulk) and μi−γ+≫1\mu_{i}-\gamma_{+}\gg 1 (far from the bulk), we split these two cases into separate theorems.

Fix τ>0\tau>0. Suppose that A⊂OA\subset\mathcal{O} satisfies 1+K−1/3⩽di⩽τ−11+K^{-1/3}\leqslant d_{i}\leqslant\tau^{-1} for all i∈Ai\in A. Define the deterministic positive quadratic form

We emphasize that the set AA in Theorem 2.11 may be chosen at will. If all outliers are well-separated, then the choice A={i}A=\{i\} gives the most precise information. However, as explained at the beginning of this subsection, the indices of outliers that are close to each other should be included in the same set AA. Thus, the freedom to chose ∣A∣⩾2\lvert A\rvert\geqslant 2 is meant for degenerate or almost degenerate outliers. (In fact, as explained after (2.14) below, the correct notion of closeness of outliers is that of overlapping.)

This gives a precise version of the cone concentration from (1.18). Note that the cone concentration holds provided the error is much smaller than the main term u(di)u(d_{i}), which leads to the conditions

here we used that di≍1d_{i}\asymp 1 and M≍(1+ϕ)KM\asymp(1+\phi)K.

We claim that both conditions in (2.14) are natural and necessary. The first condition of (2.14) simply means that μi\mu_{i} is an outlier. The second condition of (2.14) is a non-overlapping condition. To understand it, recall from (2.4) that μi\mu_{i} fluctuates on the scale (di−1)1/2K−1/2(d_{i}-1)^{1/2}K^{-1/2}. Then μi\mu_{i} is a non-overlapping outlier if all other outliers are located with high probability at a distance greater than this scale from μi\mu_{i}. Recalling the definition of the classical location θ(di)\theta(d_{i}) of μi\mu_{i}, the non-overlapping condition becomes

After a simple estimate using the definition of θ\theta, we find that this is precisely the second condition of (2.14). The degeneracy or almost degeneracy of outliers discussed at the beginning of this subsection is hence to be interpreted more precisely in terms of overlapping of outliers.

Provided μi\mu_{i} is well-separated from both the bulk spectrum and the other outliers, we find that the error in (2.13) is of order M−1/2M^{-1/2}.

As djd_{j} approaches did_{i} the delocalization bound from (2.16) deteriorates, and eventually when μi\mu_{i} and μj\mu_{j} start overlapping, i.e. the second condition of (2.14) is violated, the right-hand side of (2.16) has the same size as the leading term of (2.13). This is again a manifestation of the fact that the individual eigenspaces of overlapping outliers cannot be distinguished.

Suppose that we have an ∣A∣\lvert A\rvert-fold degenerate outlier, i.e. di=djd_{i}=d_{j} for all i,j∈Ai,j\in A. Then from Theorem 2.11 and Remark 2.12 (see the estimate (5.2)) we get, for all i,j∈Ai,j\in A,

This is the correct generalization of (2.13) from Example 2.13 to the degenerate case. The error term is the same as in (2.13), and its size and relation to the main term is exactly the same as in Example 2.13. Hence the discussion following (2.13) may be take over verbatim to this case.

We conclude this example by remarking that a similar discussion also holds for a group of outliers that is not degenerate, but nearly degenerate, i.e. ∣di−dj∣≪∣di−dk∣\lvert d_{i}-d_{j}\rvert\ll\lvert d_{i}-d_{k}\rvert for all i,j∈Ai,j\in A and k∉Ak\notin A. We omit the details.

The next result is the analogue of Theorem 2.11 for outliers far from the bulk.

We leave the discussion on the interpretation of the error in (2.17) to the reader; it is similar to that of Examples 2.13, 2.14, and 2.15.

3. Non-outlier eigenvectors

This quantity should be interpreted as a deterministic version of ∣μa−γ−∣∧∣μa−γ+∣\lvert\mu_{a}-\gamma_{-}\rvert\wedge\lvert\mu_{a}-\gamma_{+}\rvert for a∉Oa\notin\mathcal{O}; see Theorem 3.5 below.

Suppose that a⩽Ca\leqslant C (μa\mu_{a} is near the edge), which implies that κa≍K−2/3\kappa_{a}\asymp K^{-2/3}. Suppose moreover that did_{i} is near the transition point 11. Then we get

An analogous statement holds near the left spectral edge provided ∣ϕ−1∣⩾τ\lvert\phi-1\rvert\geqslant\tau; we omit the details.

More generally, our method also yields the asymptotic joint distribution of the family

(after a suitable affine rescaling of the variables, as in Theorem 8.3 below), where a1,…,ak,b1,…,bk∈[ ⁣[1,K1−τα+3] ⁣]∖Oa_{1},\dots,a_{k},b_{1},\dots,b_{k}\in[\![{1,K^{1-\tau}\alpha_{+}^{3}}]\!]\setminus\mathcal{O}. We omit the precise statement, which is a universality result: it says essentially that the asymptotic distribution of (2.22) coincides with that under the standard Wishart ensemble (i.e. an uncorrelated Gaussian sample covariance matrix). The proof is a simple corollary of Theorem 2.7, Proposition 6.2, Proposition 6.3, and Theorem 8.3.

Finally, instead of QQ defined in (1.10), we may also consider

Preliminaries

The rest of this paper is devoted to the proofs of the results from Sections 2.1–2.3. To clarify the presentation of the main ideas of the proofs, we shall first assume that

We make the assumption (3.1) throughout Sections 3–7. The additional arguments required to relax the assumption (3.1) are presented in Section 8. Under the assumption (3.1) we have

Moreover, the extension of our results from QQ to Q˙\dot{Q}, and hence the proof of Theorem 2.23, is given in Section 9.

In this section we collect the key tool of our analysis: the isotropic Marchenko-Pastur law from BEKYY .

It is well known that the empirical distribution of the eigenvalues of the N×NN\times N matrix X∗XX^{*}X has the same asymptotics as the Marchenko-Pastur law

where we recall the edges γ±\gamma_{\pm} of the limiting spectrum defined in (1.7). Similarly, as noted in (1.6), the empirical distribution of the eigenvalues of the M×MM\times M matrix XX∗XX^{*} has the same asymptotics as ϱϕ−1\varrho_{\phi^{-1}}.

Note that (3.3) is normalized so that its integral is equal to one. The Stieltjes transform of the Marchenko-Pastur law (3.3) is

where the square root is chosen so that mϕm_{\phi} is holomorphic in the upper half-plane and satisfies mϕ(z)→0m_{\phi}(z)\to 0 as z→∞z\to\infty. The function mϕ=mϕ(z)m_{\phi}=m_{\phi}(z) is also characterized as the unique solution of the equation

satisfying Im⁡m(z)>0\operatorname{Im}m(z)>0 for Im⁡z>0\operatorname{Im}z>0. The formulas (3.3)–(3.5) were originally derived for the case when ϕ=M/N\phi=M/N is independent of NN (or, more precisely, when ϕ\phi has a limit in (0,∞)(0,\infty) as N→∞N\to\infty). Our results allow ϕ\phi to depend on NN under the constraint (1.9), so that mϕm_{\phi} and ϱϕ\varrho_{\phi} may also depend on NN through ϕ\phi.

Throughout the following we use a spectral parameter

with η>0\eta>0, as the argument of Stieltjes transforms and resolvents. Define the resolvent

Throughout the following we regard the quantities E(z)E(z), η(z)\eta(z), and κ(z)\kappa(z) as functions of zz and usually omit the argument unless it is needed to avoid confusion.

Sometimes we shall need the following notion of high probability.

Fix a (small) ω∈(0,1)\omega\in(0,1) and define the domain

Beyond the support of the limiting spectrum, one has stronger control all the way down to the real axis. For fixed (small) ω>0\omega>0 define the region

of spectral parameters separated from the asymptotic spectrum by K−2/3+ωK^{-2/3+\omega}, which may have an arbitrarily small positive imaginary part η\eta. Throughout the following we regard ω\omega as fixed once and for all, and do not track the dependence of constants on ω\omega.

Suppose that (1.15), (1.9), and (1.16) hold. Then

for all ε>0\varepsilon>0, D>0D>0, and N⩾N0(ε,D)N\geqslant N_{0}(\varepsilon,D). See (BEKYY, , Remark 2.6).

The next results are on the nontrivial (i.e. nonzero) eigenvalues of H:=XX∗H\mathrel{\mathop{:}}=XX^{*} as well as the corresponding eigenvectors. The matrix HH has KK nontrivial eigenvalues, which we order according to

(The remaining M−KM-K eigenvalues of HH are zero.) Moreover, we denote by

the unit eigenvectors of HH associated with the nontrivial eigenvalues λ1⩾λ2⩾⋯⩾λK\lambda_{1}\geqslant\lambda_{2}\geqslant\dots\geqslant\lambda_{K}.

Fix τ>0\tau>0, and suppose that (1.15), (1.9), and (1.16) hold. Then for i∈[ ⁣[1,K] ⁣]i\in[\![{1,K}]\!] we have

if either i⩽(1−τ)Ki\leqslant(1-\tau)K or ∣ϕ−1∣⩾τ\lvert\phi-1\rvert\geqslant\tau.

The following result is on the rigidity of the nontrivial eigenvalues of HH. Let γ1⩾γ2⩾⋯⩾γK\gamma_{1}\geqslant\gamma_{2}\geqslant\cdots\geqslant\gamma_{K} be the classical eigenvalue locations according to ϱϕ\varrho_{\phi} (see (3.3)), defined through

Fix τ>0\tau>0, and suppose that (1.15), (1.9), and (1.16) hold. Then for i∈[ ⁣[1,M] ⁣]i\in[\![{1,M}]\!] we have

if i⩽(1−τ)Ki\leqslant(1-\tau)K or ∣ϕ−1∣⩾τ\lvert\phi-1\rvert\geqslant\tau.

2. Link to the semicircle law

Note that this is nothing but Wigner’s semicircle law centred at ϕ1/2+ϕ−1/2\phi^{1/2}+\phi^{-1/2}. Thus,

where in the last step we used (3.4). Note that

The estimates (3.19) and (3.20) follow from the explicit expressions in (3.4) and (3.17). In fact, these estimates have already appeared in previous works. Indeed, for mϕm_{\phi} the estimates (3.19) and (3.20) were proved in (BEKYY, , Lemma 3.3). In order to prove them for wϕw_{\phi}, we observe that the estimates (3.19) and (3.20) follow from the corresponding ones for the semicircle law, which were proved in (EKYY4, , Lemma 4.3). The estimates (3.21) follow from (3.20) and the elementary identity

which can be derived from (3.18); the estimates for mϕm_{\phi} are derived similarly. Finally, (3.22) follows easily from

which may itself be derived from (3.5). ∎

In analogy to wϕw_{\phi} (see (3.17)), we define the matrix-valued function

Theorem 3.2 has the following analogue, which compares FF with mϕm_{\phi}.

Suppose that (1.15), (1.9), and (1.16) hold. Then

3. Extension of the spectral domain

In this section we extend the spectral domain on which Theorem 3.2 and Lemma 3.7 hold. The argument relies on the Helffer-Sjöstrand functional calculus Davies . Define the domains

The basic idea of the proof is to apply the Helffer-Sjöstrand formula to the function

where x0x_{0} is chosen below. To that end, we need a smooth compactly supported cutoff function χ\chi on the complex plane satisfying χ(w)∈\chi(w)\in and ∣∂wˉχ(w)∣⩽C(ω,τ)\lvert\partial_{\bar{w}}\chi(w)\rvert\leqslant C(\omega,\tau). We distinguish the three cases ϕ<1−τ\phi<1-\tau, ∣ϕ−1∣⩽τ\lvert\phi-1\rvert\leqslant\tau, and ϕ>1+τ\phi>1+\tau.

Let us first focus on the case ϕ<1−τ\phi<1-\tau. Set x0:=ϕ1/2+ϕ−1/2x_{0}\mathrel{\mathop{:}}=\phi^{1/2}+\phi^{-1/2} and choose a constant ω′=ω′(ω,τ)∈(0,ω)\omega^{\prime}=\omega^{\prime}(\omega,\tau)\in(0,\omega) small enough that γ−⩾4ω′\gamma_{-}\geqslant 4\omega^{\prime}. We require that χ\chi be equal to 11 in the ω′\omega^{\prime}-neighbourhood of [γ−,γ+][\gamma_{-},\gamma_{+}] and outside of the 2ω′2\omega^{\prime}-neighbourhood of [γ−,γ+][\gamma_{-},\gamma_{+}]. By Theorem 3.5 we have supp⁡ρΔ⊂{χ=1}\operatorname{supp}\rho^{\Delta}\subset\{\chi=1\} with high probability. Now choose zz satisfying dist⁡(z,[γ−,γ+])⩾3ω′\operatorname{dist}(z,[\gamma_{-},\gamma_{+}])\geqslant 3\omega^{\prime}. Then the Helffer-Sjöstrand formula Davies yields, for x∈supp⁡ρΔx\in\operatorname{supp}\rho^{\Delta},

Finally, suppose that ϕ>1+τ\phi>1+\tau. Now we set x0:=0x_{0}\mathrel{\mathop{:}}=0. We choose the same ω′\omega^{\prime} and cutoff function χ\chi as in the case ϕ<1−τ\phi<1-\tau above. Suppose that dist⁡(z,[γ−,γ+])⩾3ω′\operatorname{dist}(z,[\gamma_{-},\gamma_{+}])\geqslant 3\omega^{\prime} and z≠0z\neq 0. Thus, (3.30) holds with high probability for x∈supp⁡ρΔ∖{0}x\in\operatorname{supp}\rho^{\Delta}\setminus\{0\}. Since fw(0)=0f_{w}(0)=0, we therefore find that (3.31) holds. As above, we find that for w∈{∂wˉχ≠0}w\in\{\partial_{\bar{w}}\chi\neq 0\} we have

and ∣mΔ(w)∣≺ϕ−1K−1/2\lvert m^{\Delta}(w)\rvert\prec\phi^{-1}K^{-1/2}. Recalling (3.10), we find that (3.28) follows easily. ∎

Proposition 3.8 yields the following result for FF defined in (3.24).

4. Identities for the resolvent and eigenvalues

In this section we derive the identities on which our analysis of the eigenvalues and eigenvectors relies. Recall the definition of the set R\mathcal{R} from (1.14). We write the population covariance matrix Σ\Sigma from (1.12) as

where H=XX∗H=XX^{*} and Q=Σ1/2HΣ1/2Q=\Sigma^{1/2}H\Sigma^{1/2}. We introduce the ∣R∣×∣R∣\lvert\mathcal{R}\rvert\times\lvert\mathcal{R}\rvert matrix

We also denote by σ(A)\sigma(A) the spectrum of a square matrix AA.

The following lemma collects the basic identities for analysing σ(Q)\sigma(Q) and G~\widetilde{G}. We remark that versions of its part (i) have already appeared in several previous works on finite-rank deformations of random matrix ensembles BGGM1 ; BY1 ; KY2 ; SoshPert .

Suppose that μ∉σ(H)\mu\notin\sigma(H). Then μ∈σ(Q)\mu\in\sigma(Q) if and only if

To prove (i), we write the condition μ∈σ(Q)\mu\in\sigma(Q) as

where we used that μ∉σ(H)\mu\notin\sigma(H). Using

the matrix identity det⁡(1+XY)=det⁡(1+YX)\det(1+XY)=\det(1+YX), and det⁡(Σ)≠0\det(\Sigma)\neq 0, we find

with A=H−zA=H-z, B=D(ϕ−1/2+D)−1B=D(\phi^{-1/2}+D)^{-1}, S=VS=V, and T=zV∗T=zV^{*}. ∎

The result (3.35), when restricted to the range of VV, has an alternative form (3.37) which is often easier to work with, since it collects all of the randomness in the single quantity W(z)W(z) on its right-hand side.

Eigenvalue locations

In this section we prove Theorems 2.3 and 2.7. The arguments are similar to those of (KY2, , Section 6), and we therefore only sketch the proofs. The proof of (KY2, , Section 6) relies on three main steps: (i) establishing a forbidden region which contains with high probability no eigenvalues of QQ; (ii) a counting estimate for the special case where DD does not depend on NN, which ensures that each connected component of the allowed region (complement of the forbidden region) contains exactly the right number of eigenvalues of QQ; and (iii) a continuity argument where the counting result of (ii) is extended to arbitrary NN-dependent DD using the gaps established in (i) and the continuity of the eigenvalues as functions of the matrix entries. The steps (ii) and (iii) are exactly the same as in KY2 , and will not be repeated here. The step (i) differs slightly from that of KY2 , and in the proofs below we explain these differences.

Let ∣R∣=1\lvert\mathcal{R}\rvert=1 and D=d∈DD=d\in\mathcal{D}. For d>0d>0 we have

Writing this in spectral decomposition yields

As above, a simple perturbation argument implies that we may without loss of generality assume that all scalar products in (4.1) are nonzero. Now take z∈(0,∞)z\in(0,\infty). Note that b(z)b(z) and dd have the same sign.

To conclude the proof, we observe that the left-hand side of (4.1) defines a function of z∈(0,∞)z\in(0,\infty) with M−1M-1 singularities and MM zeros, which is smooth and decreasing away from the singularities. Moreover, its zeros are the eigenvalues λ1,…,λM\lambda_{1},\dots,\lambda_{M}. The interlacing property now follows from the fact that zz is an eigenvalue of QQ if and only if the left-hand side of (4.1) is equal to −b(z)-b(z). ∎

For the rank-∣R∣\lvert\mathcal{R}\rvert model (1.10) we have

with the convention that λi=0\lambda_{i}=0 for i>Ki>K and λi=∞\lambda_{i}=\infty for i<1i<1.

Throughout the following we shall make use of the subsets of outliers

for τ⩾0\tau\geqslant 0. Note that O  =  O0+∪O0−\mathcal{O}\;=\;\mathcal{O}_{0}^{+}\cup\mathcal{O}_{0}^{-}.

The proof of Proposition 2.3 is similar to that of (KY2, , Equation (2.20)). We focus first on the outliers to the right of the bulk spectrum. Let ε>0\varepsilon>0. We shall prove that there exists an event Ξ\Xi of high probability (see Definition 3.1) such that for all i∈O4ε+i\in\mathcal{O}_{4\varepsilon}^{+} we have

and for i∈[ ⁣[∣O4ε+∣+1,∣O4ε+∣+r] ⁣]i\in[\![{\lvert\mathcal{O}_{4\varepsilon}^{+}\rvert+1,\lvert\mathcal{O}_{4\varepsilon}^{+}\rvert+r}]\!] we have

Before proving (4.3) and (4.4), we show how they imply (2.4) for di>0d_{i}>0 and (2.5). From (4.4) we get for ii satisfying K−1/3⩽di−1⩽K−1/3+4εK^{-1/3}\leqslant d_{i}-1\leqslant K^{-1/3+4\varepsilon}

Since ε>0\varepsilon>0 was arbitrary, (2.4) for di>0d_{i}>0 and (2.5) follow from (4.3) and (4.5).

What remains is the proof of (4.3) and (4.4). As in (KY2, , Proposition 6.5), the first step is to prove that with high probability there are no eigenvalues outside a neighbourhood of the classical outlier locations θ(di)\theta(d_{i}). To that end, we define for each i∈Oε+i\in\mathcal{O}_{\varepsilon}^{+} the interval

Moreover, we set I0:=[0,θ(1+K−1/3+2ε)]I_{0}\mathrel{\mathop{:}}=[0,\theta(1+K^{-1/3+2\varepsilon})].

We now claim that with high probability the complement of the set I(D):=I0∪⋃i∈Oε+Ii(D)I(D)\mathrel{\mathop{:}}=I_{0}\cup\bigcup_{i\in\mathcal{O}_{\varepsilon}^{+}}I_{i}(D) contains no eigenvalues of QQ. Indeed, from Theorem 3.5 and Corollary 3.9 combined with Remark 3.3 (with small enough ω≡ω(ε)\omega\equiv\omega(\varepsilon)), we find that there exists an event Ξ\Xi of high probability such that ∣λi−γ+∣⩽K−2/3+ε\lvert\lambda_{i}-\gamma_{+}\rvert\leqslant K^{-2/3+\varepsilon} for i∈[ ⁣[1,2r] ⁣]i\in[\![{1,2r}]\!] and

for all x∉I0x\notin I_{0}, where we defined

is singular. Since −di−1=wϕ(θ(di))-d_{i}^{-1}=w_{\phi}(\theta(d_{i})) for i∈Oε+i\in\mathcal{O}_{\varepsilon}^{+}, we conclude from the definition of I(D)I(D) that it suffices to show that if x∉I(D)x\notin I(D) then

We prove (4.6) using the two following observations. First, wϕw_{\phi} is monotone increasing on (γ+,∞)(\gamma_{+},\infty) and

We omit further details, which may be found e.g. in (KY2, , Section 6). Thus we conclude that on the event Ξ\Xi the complement of I(D)I(D) contains no eigenvalues of QQ.

The next step of the proof consists in making sure that the allowed neighbourhoods Ii(D)I_{i}(D) contain exactly the right number of outliers; the counting argument (sketched in the steps (ii) and (iii) at the beginning of this section) follows that of (KY2, , Section 6). First we consider the case D=D(0)D=D(0) where for all i≠j∈Oε+i\neq j\in\mathcal{O}_{\varepsilon}^{+} we have di(0),dj(0)⩾2d_{i}(0),d_{j}(0)\geqslant 2 and ∣di(0)−dj(0)∣⩾1\lvert d_{i}(0)-d_{j}(0)\rvert\geqslant 1, and show that each interval {Ii(D(0)):i∈Oε+}\{I_{i}(D(0))\mathrel{\mathop{:}}i\in\mathcal{O}_{\varepsilon}^{+}\} contains exactly one eigenvalue of QQ (see (KY2, , Proposition 6.6)). We then deduce the general case by a continuity argument, by choosing an appropriate continuous path (D(t))t∈(D(t))_{t\in} joining the initial configuration D(0)D(0) to the desired final configuration D=D(1)D=D(1). The continuity argument requires the existence of a gap in the set I(D)I(D) to the left of ⋃i∈O4ε+Ii(D)\bigcup_{i\in\mathcal{O}_{4\varepsilon}^{+}}I_{i}(D). The existence of such a gap follows easily from the definition of I(D)I(D) and the fact that ∣R∣\lvert\mathcal{R}\rvert is bounded. The details are the same as in (KY2, , Section 6.5). Hence (4.3) follows. Moreover, (4.4) follows from the same argument combined with Corollary 4.2 for a lower bound on μi\mu_{i}. This concludes the analysis of the outliers to the right of the bulk spectrum.

The case of outliers to the left of the bulk spectrum is analogous. Here we assume that ϕ<1−τ\phi<1-\tau. The argument is exactly the same as for di>0d_{i}>0, except that we use the bound (3.32) to the left of the bulk spectrum as well as ∣λi−γ−∣⩽K−2/3+ε\lvert\lambda_{i}-\gamma_{-}\rvert\leqslant K^{-2/3+\varepsilon} for i∈[ ⁣[K−2r,K] ⁣]i\in[\![{K-2r,K}]\!] with high probability. ∎

We only give the proof of (2.9); the proof of (2.10) is analogous. Fix ε>0\varepsilon>0. By Theorem 2.3, Theorem 3.5, Theorem 3.2, Lemma 3.7, and Remark 3.3, there exists a high-probability event Ξ≡ΞN(ε)\Xi\equiv\Xi_{N}(\varepsilon) satisfying the following conditions.

For the following we fix a realization H∈ΞH\in\Xi. We suppose first that

and define η:=K−1+2εα+−1\eta\mathrel{\mathop{:}}=K^{-1+2\varepsilon}\alpha_{+}^{-1}. Now suppose that xx satisfies

We shall show, using (3.34), that any xx satisfying (4.11) cannot be an eigenvalue of QQ. First we deduce from (4.8) that

The estimate (4.12) follows by spectral decomposition of F(⋅)F(\cdot) together with the estimate 2∣λi−x∣⩾(λi−x)2+η22\lvert\lambda_{i}-x\rvert\geqslant\sqrt{(\lambda_{i}-x)^{2}+\eta^{2}} for all ii. We get from (4.12) and Lemma 3.6 that

where we use the notation A=B+O(t)A=B+O(t) to mean ∥A−B∥⩽Ct\lVert A-B\rVert\leqslant Ct. Recalling (3.34), we conclude that on the event Ξ\Xi the value xx is not an eigenvalue of QQ provided

It is easy to check that this condition is satisfied if

where we used (4.10). Recalling (4.7), we therefore conclude that for i⩽K1−2εα+3i\leqslant K^{1-2\varepsilon}\alpha_{+}^{3} the set

The next step of the proof is a counting argument (sketched in the steps (ii) and (iii) at the beginning of this section), which uses the eigenvalue interlacing from Lemma 4.1. They details are the same as in (KY2, , Section 6), and hence omitted here. The counting argument implies that for i⩽K1−2εα+3i\leqslant K^{1-2\varepsilon}\alpha_{+}^{3} and assuming (4.10) we have

What remains is to check (4.13) for the cases α+<K−1/3+ε\alpha_{+}<K^{-1/3+\varepsilon} and i>K1−2εα+3i>K^{1-2\varepsilon}\alpha_{+}^{3}.

Suppose first that α+<K−1/3+ε\alpha_{+}<K^{-1/3+\varepsilon}. Then using the rigidity from (4.7) and interlacing from Corollary 4.2 we find

where we used the trivial bound i⩾1i\geqslant 1. Similarly, if i>K1−2εα+3i>K^{1-2\varepsilon}\alpha_{+}^{3} satisfies i⩽(1−τ)Ki\leqslant(1-\tau)K, we may repeat the same estimate.

We conclude that (4.13) under the sole assumption that i⩽(1−τ)Ki\leqslant(1-\tau)K. Since ε>0\varepsilon>0 was arbitrary, (2.9) follows. ∎

Outlier eigenvectors

In this section we focus on the outlier eigenvectors ξa\xi_{a}, a∈Oa\in\mathcal{O}. Here we in fact prove Theorem 2.11 under the stronger assumption

instead of 1+K−1/3⩽di⩽τ−11+K^{-1/3}\leqslant d_{i}\leqslant\tau^{-1}. How to improve the lower bound from 1+K−1/3+τ1+K^{-1/3+\tau} to the claimed K−1/3K^{-1/3} requires a completely different approach, relying on eigenvector delocalization bounds, and is presented in Section 6 in conjunction with results for the non-outlier eigenvectors ξa\xi_{a}, a∉Oa\notin\mathcal{O}.

The proof of Theorem 2.16 is similar to that of Theorem 2.11; one has to adapt the proof to cover the range di∈[1+τ,∞)d_{i}\in[1+\tau,\infty) instead of di∈[1+K−1/3,τ−1]d_{i}\in[1+K^{-1/3},\tau^{-1}]. The key input is the extension of the spectral domain from Corollary 3.9. For the sake of brevity we omit the details of the proof of Theorem 2.16, and focus solely on Theorem 2.11.

The following proposition is the main result of this section.

Fix τ>0\tau>0. Suppose that AA satisfies (5.1). Then for all i,j=1,…,Mi,j=1,\dots,M we have

where the symbol (i↔j)(i\leftrightarrow j) denotes the preceding terms with ii and jj interchanged.

Note that, under the assumption (5.1), Theorem 2.11 is an easy consequence of Proposition 5.1. As explained above, the proof of Theorem 2.11 in full generality is given in Section 6, where we give the additional argument required to relax (5.1).

The rest of this section is devoted to the proof of Proposition 5.1.

We first prove a slightly stronger version of (5.2) under the additional non-overlapping condition

for all i∈Ai\in A, where δ>0\delta>0 is a positive constant. This is a precise version of the second condition of (2.14), whose interpretation was given below (2.14): an outlier indexed by AA cannot overlap with an outlier indexed by AcA^{c}. Note, however, that there is no restriction on the outliers indexed by AA overlapping among themselves. The assumption (5.3) will be removed in Section 5.2. The main estimate for non-overlapping outliers is the following.

Fix τ>0\tau>0 and δ>0\delta>0. Suppose that AA satisfies (5.1) and (5.3) for all i∈Ai\in A. Then for all i,j=1,…,Mi,j=1,\dots,M we have

The rest of this subsection is devoted to the proof of Proposition 5.2. We begin by defining ω:=τ/2\omega\mathrel{\mathop{:}}=\tau/2 and letting ε<min⁡{τ/3,δ}\varepsilon<\min\{{\tau/3,\delta}\} be a positive constant to be determined later. We choose a high-probability event Ξ≡ΞN(ε,τ)\Xi\equiv\Xi_{N}(\varepsilon,\tau) (see Definition 3.1) satisfying the following conditions.

for i,j∈Ri,j\in\mathcal{R}, large enough KK, and all zz in the set

For all ii satisfying 1+K−1/3⩽di⩽ω−11+K^{-1/3}\leqslant d_{i}\leqslant\omega^{-1} we have

Note that such an event Ξ\Xi exists. Indeed, (5.7) and (5.8) may be satisfied using Theorem 2.3, and (5.5) using Theorem 3.7 combined with Remark 3.3.

For the sequel we fix a realization H∈ΞH\in\Xi satisfying the conditions (i)–(iii) above. Hence, the rest of the proof of Proposition 5.2 is entirely deterministic, and the randomness only enters in ensuring that Ξ\Xi has high probability. Our starting point is a contour integral representation of the projection PAP_{A}. In order to construct the contour, we define for each i∈Ai\in A the radius

We define the contour Γ:=∂Υ\Gamma\mathrel{\mathop{:}}=\partial\Upsilon as the boundary of the union of discs Υ:=⋃i∈ABρi(di)\Upsilon\mathrel{\mathop{:}}=\bigcup_{i\in A}B_{\rho_{i}}(d_{i}), where Bρ(d)B_{\rho}(d) is the open disc of radius ρ\rho around dd. We shall sometimes need the decomposition Γ=⋃i∈AΓi\Gamma=\bigcup_{i\in A}\Gamma_{i}, where Γi:=Γ∩∂Bρi(di)\Gamma_{i}\mathrel{\mathop{:}}=\Gamma\cap\partial B_{\rho_{i}}(d_{i}). See Figure 5.1 for an illustration of Γ\Gamma.

We shall have to use the estimate (5.5) on the set θ(Υ)‾ ⁣ \overline{\theta(\Upsilon)}\!\,. Its applicability is an immediate consequence of the following lemma.

The set θ(Υ)‾ ⁣ \overline{\theta(\Upsilon)}\!\, lies in (5.6).

It is easy to check that θ(ζ)⩽ω−1\theta(\zeta)\leqslant\omega^{-1} for all ζ∈Υ\zeta\in\Upsilon. In order to check the lower bound on Re⁡θ(ζ)\operatorname{Re}\theta(\zeta), we note that for any α∈(0,1)\alpha\in(0,1) there exists a constant c≡c(α,τ)c\equiv c(\alpha,\tau) such that

for Re⁡ζ⩾1\operatorname{Re}\zeta\geqslant 1, ∣Im⁡ζ∣⩽α(Re⁡ζ−1)\lvert\operatorname{Im}\zeta\rvert\leqslant\alpha(\operatorname{Re}\zeta-1), and ∣ζ∣⩽τ−1\lvert\zeta\rvert\leqslant\tau^{-1}. Now the claim follows easily from Re⁡ζ⩾1+K−1/3+τ/2\operatorname{Re}\zeta\geqslant 1+K^{-1/3+\tau}/2 for all ζ∈Υ\zeta\in\Upsilon, by choosing α=1/3\alpha=1/\sqrt{3}. ∎

Each outlier {μi}i∈A\{\mu_{i}\}_{i\in A} lies in θ(Υ)\theta(\Upsilon), and all other eigenvalues of QQ lie in the complement of θ(Υ)‾ ⁣ \overline{\theta(\Upsilon)}\!\,.

It suffices to prove that (a) for each i∈Ai\in A we have μi∈θ(Bρi(di))\mu_{i}\in\theta(B_{\rho_{i}}(d_{i})) and (b) all the other eigenvalues μj\mu_{j} satisfy μj∉θ(Bρi(di))\mu_{j}\notin\theta(B_{\rho_{i}}(d_{i})) for all i∈Ai\in A.

for i∈Ai\in A, as follows from (5.3) and (5.1). Using

it is then not hard to get (a) from (5.10) and (5.7).

In order to prove (b), we consider the two cases (i) 1+K−1/3⩽dj⩽ω−11+K^{-1/3}\leqslant d_{j}\leqslant\omega^{-1} with j∉Aj\notin A, and (ii) and j⩾s++1j\geqslant s_{+}+1. In the case (i), the claim (b) follows using (5.7), (5.11), and (5.3). In the case (ii), the claim (b) follows from (5.8) and the estimate

Using the spectral decomposition of G~(z)\widetilde{G}(z), Lemma 5.5, and the residue theorem, we may write the projection PAP_{A} as

This is the desired integral representation of PAP_{A}.

We now perform a resolvent expansion on the denominator

where we used Cauchy’s theorem, (4.2), and the fact that did_{i} lies in Υ\Upsilon if and only if i∈Ai\in A.

using the fact that fijf_{ij} is holomorphic inside Γ\Gamma and satisfies the bounds

The first bound of (5.21) follows from (5.5), (5.11), and (5.12). The second bound of (5.21) follows by plugging the first one into

where the contour C\mathcal{C} is the circle of radius ∣ζ−1∣/2\lvert\zeta-1\rvert/2 centred at ζ\zeta. (By assumptions on ε\varepsilon and ω\omega, the function fijf_{ij} is holomorphic in a neighbourhood of the closed interior of C\mathcal{C}.)

In order to estimate (5.20), we consider the three cases (i) i,j∈Ai,j\in A, (ii) i∈Ai\in A, j∉Aj\notin A, (iii) i∉Ai\notin A, j∈Aj\in A. Note that (5.20) vanishes if i,j∉Ai,j\notin A. We start with the case (i). Suppose first that i≠ji\neq j and di≠djd_{i}\neq d_{j}. Then we find

A simple limiting argument shows that this bound is also valid for di=djd_{i}=d_{j} and i=ji=j. Next, in the case (ii) we get from (5.21)

A similar estimate holds for the case (iii). Putting all three cases together, we find

What remains is the estimate of Sij(2)S_{ij}^{(2)}. Here residue calculations are unavailable, and the precise choice of the contour Γ\Gamma is crucial. We use the following basic estimate to control the integral.

For k∈Ak\in A, l∈Rl\in\mathcal{R}, and ζ∈Γk\zeta\in\Gamma_{k} we have

The upper bound ∣ζ−dl∣⩽ρk+∣dk−dl∣\lvert\zeta-d_{l}\rvert\leqslant\rho_{k}+\lvert d_{k}-d_{l}\rvert is trivial, so that we only focus on the lower bound. Suppose first that l∉Al\notin A. Then we get ∣ζ−dl∣⩾∣dk−dl∣−ρk\lvert\zeta-d_{l}\rvert\geqslant\lvert d_{k}-d_{l}\rvert-\rho_{k}, from which the claim follows since ∣dk−dl∣⩾2ρk\lvert d_{k}-d_{l}\rvert\geqslant 2\rho_{k} by (5.9).

For the remainder of the proof we may therefore suppose that l∈Al\in A. Define δ:=∣dk−dl∣−ρk−ρl\delta\mathrel{\mathop{:}}=\lvert d_{k}-d_{l}\rvert-\rho_{k}-\rho_{l}, the distance between the discs Dρk(dk)D_{\rho_{k}}(d_{k}) and Dρl(dl)D_{\rho_{l}}(d_{l}) (see Figure 5.1). We consider the two cases 4δ⩽∣dk−dl∣4\delta\leqslant\lvert d_{k}-d_{l}\rvert and 4δ>∣dk−dl∣4\delta>\lvert d_{k}-d_{l}\rvert separately.

Suppose first that 4δ⩽∣dk−dl∣4\delta\leqslant\lvert d_{k}-d_{l}\rvert. Then by definition of δ\delta we have ∣dk−dl∣⩽43(ρk+ρl)\lvert d_{k}-d_{l}\rvert\leqslant\frac{4}{3}(\rho_{k}+\rho_{l}). Now a simple estimate using the definition of ρi\rho_{i} yields ρk/5⩽ρl⩽5ρk\rho_{k}/5\leqslant\rho_{l}\leqslant 5\rho_{k}, from which we conclude ∣dk−dl∣⩽8ρk\lvert d_{k}-d_{l}\rvert\leqslant 8\rho_{k}. The claim now follows from the bound ∣ζ−dl∣⩾ρl\lvert\zeta-d_{l}\rvert\geqslant\rho_{l}.

Suppose now that 4δ>∣dk−dl∣4\delta>\lvert d_{k}-d_{l}\rvert. Hence ρk+ρl⩽34∣dk−dl∣\rho_{k}+\rho_{l}\leqslant\frac{3}{4}\lvert d_{k}-d_{l}\rvert, so that in particular ρk⩽∣dk−dl∣\rho_{k}\leqslant\lvert d_{k}-d_{l}\rvert. Thus we get

From (5.18), (5.5), (5.11), and (5.12) we get

where we also used the estimate ∣θ(ζ)∣≍ϕ−1/2(1+ϕ)\lvert\theta(\zeta)\rvert\asymp\phi^{-1/2}(1+\phi) for ζ∈Γ\zeta\in\Gamma.

In order to estimate the matrix norm, we observe that for ζ∈Γk\zeta\in\Gamma_{k} we have on the one hand

for any l∈Rl\in\mathcal{R}, where in the last step we used (5.10). Since ε<δ\varepsilon<\delta, these estimates combined with a resolvent expansion give the bound

for ζ∈Γk\zeta\in\Gamma_{k}. Decomposing the integration contour in (5.23) as Γ=⋃k∈AΓk\Gamma=\bigcup_{k\in A}\Gamma_{k}, and recalling that Γk\Gamma_{k} has length bounded by 2πρk2\pi\rho_{k}, we get from Lemma 5.6

We estimate the right-hand side using Cauchy-Schwarz. For i∉Ai\notin A we find, using (5.9),

For i∈Ai\in A we use (5.9) and the estimate ρk+∣di−dk∣⩾ρi\rho_{k}+\lvert d_{i}-d_{k}\rvert\geqslant\rho_{i} for all k∈Ak\in A to get

Recall that M≍(1+ϕ)KM\asymp(1+\phi)K. Hence, plugging (5.19), (5.22), and (5.27) into (5.15), we find

We have proved (5.28) under the assumption that i,j∈Ri,j\in\mathcal{R}. The general case is an easy corollary. For general i,j∈[ ⁣[1,M] ⁣]i,j\in[\![{1,M}]\!], we define R^:=R∪{i,j}\widehat{\mathcal{R}}\mathrel{\mathop{:}}=\mathcal{R}\cup\{i,j\} and consider

where d^k:=dk\widehat{d}_{k}\mathrel{\mathop{:}}=d_{k} for k∈Rk\in\mathcal{R} and d^k∈(0,1/2)\widehat{d}_{k}\in(0,1/2) for k∈R^∖Rk\in\widehat{\mathcal{R}}\setminus\mathcal{R}. Since ∣R^∣⩽r+2\lvert\widehat{\mathcal{R}}\rvert\leqslant r+2 and D^\widehat{D} is invertible, we may apply the result (5.28) to this modified model. Now taking the limit d^k→0\widehat{d}_{k}\to 0 for k∈R^∖Rk\in\widehat{\mathcal{R}}\setminus\mathcal{R} in (5.28) concludes the proof in the general case. Now Proposition 5.2 follows since ε\varepsilon may be chosen arbitrarily small. This concludes the proof of Proposition 5.2.

2. Removing the non-overlapping assumption

In this subsection we complete the proof of Proposition 5.1 by extending Proposition 5.2 to the case where (5.3) does not hold.

Let δ<τ/4\delta<\tau/4. We say that i,j∈Oτ/2+i,j\in\mathcal{O}_{\tau/2}^{+} overlap if ∣di−dj∣⩽(di−1)−1/2K−1/2+δ\lvert d_{i}-d_{j}\rvert\leqslant(d_{i}-1)^{-1/2}K^{-1/2+\delta} or ∣di−dj∣⩽(dj−1)−1/2K−1/2+δ\lvert d_{i}-d_{j}\rvert\leqslant(d_{j}-1)^{-1/2}K^{-1/2+\delta}. For A⊂Oτ+A\subset\mathcal{O}_{\tau}^{+} we introduce sets S(A),L(A)⊂Oτ/2+S(A),L(A)\subset\mathcal{O}_{\tau/2}^{+} satisfying S(A)⊂A⊂L(A)S(A)\subset A\subset L(A). Informally, S(A)⊂AS(A)\subset A is the largest subset of indices of AA that do not overlap with its complement. It is by definition constructed by successively choosing k∈Ak\in A, such that kk overlaps with an index of AcA^{c}, and removing kk from AA; this process is repeated until no such kk exists. One can check that the result is independent of the choice of kk at each step. Note that S(A)S(A) may be empty.

Informally, L(A)⊃AL(A)\supset A is the smallest subset of indices in Oτ/2+\mathcal{O}_{\tau/2}^{+} that do not overlap with its complement. It is by definition constructed by successively choosing k∈Oτ/2+∖Ak\in\mathcal{O}_{\tau/2}^{+}\setminus A, such that kk overlaps with an index of AA, and adding kk to AA; this process is repeated until no such kk exists. One can check that the result is independent of the choice of kk at each step. See Figure 5.2 for an illustration of S(A)S(A) and L(A)L(A). Throughout the following we shall repeatedly make use of the fact that, for any A⊂Oτ+A\subset\mathcal{O}_{\tau}^{+}, Proposition 5.2 is applicable with (τ,A)(\tau,A) replaced by (τ/2,S(A))(\tau/2,S(A)) or (τ/2,L(A))(\tau/2,L(A)).

After these preparations, we move on to the proof of (5.2). We divide the argument into four steps.

We consider two cases, i∉L(A)i\notin L(A) and i∈L(A)i\in L(A). Suppose first that i∉L(A)i\notin L(A). Using that ∣R∣\lvert\mathcal{R}\rvert is bounded, it is not hard to see that νi(A)≍νi(L(A))\nu_{i}(A)\asymp\nu_{i}(L(A)). We now invoke Proposition 5.2 and get

In the complementary case, i∈L(A)i\in L(A), a simple argument yields

as well as σi≍1+ϕ1/2\sigma_{i}\asymp 1+\phi^{1/2}. From Proposition 5.2 we therefore get

where we used that M≍(1+ϕ)KM\asymp(1+\phi)K. Recalling (5.29), we conclude

(b) i=j∈A𝑖𝑗𝐴i=j\in A

We consider the two cases i∈S(A)i\in S(A) and i∉S(A)i\notin S(A). Suppose first that i∈S(A)i\in S(A). We write

We compute the first term of (5.32) using Proposition 5.2 and the observation that νi(A)≍νi(S(A))\nu_{i}(A)\asymp\nu_{i}(S(A)):

In order to estimate the second term of (5.32), we note that νi(A)≍νi(A∖S(A))\nu_{i}(A)\asymp\nu_{i}(A\setminus S(A)). We therefore apply (5.31) with AA replaced by A∖S(A)A\setminus S(A) to get

Going back to (5.32), we have therefore proved that

Next, we consider the case i∉S(A)i\notin S(A). Now we have (5.30), so that Proposition 5.2 yields

By (5.30) and M≍(1+ϕ)KM\asymp(1+\phi)K, we have

from which we deduce (5.33) also in the case i∉S(A)i\notin S(A).

(c) i≠j𝑖𝑗i\neq j and i∉A𝑖𝐴i\notin A or j∉A𝑗𝐴j\notin A

From cases (a) and (b) (i.e. (5.31) and (5.33)), combined with the estimate

we find, assuming i∉Ai\notin A or j∉Aj\notin A, that (5.2) holds with an additional factor K2δK^{2\delta} multiplying the right-hand side.

(d) i≠j𝑖𝑗i\neq j and i,j∈A𝑖𝑗𝐴i,j\in A

We now deal with the last remaining case by using the splitting

Note that here σi≍σj≍1+ϕ1/2\sigma_{i}\asymp\sigma_{j}\asymp 1+\phi^{1/2}. We consider the four cases (i) i,j∈S(A)i,j\in S(A), (ii) i∈S(A)i\in S(A) and j∉S(A)j\notin S(A), (iii) i∉S(A)i\notin S(A) and j∈S(A)j\in S(A), and (iv) i,j∉S(A)i,j\notin S(A).

Next, consider the case (ii). For the first term of (5.34) we use the estimates

In order to estimate the last term, we first assume that dj⩽did_{j}\leqslant d_{i} and di−1⩽2∣di−dj∣d_{i}-1\leqslant 2\lvert d_{i}-d_{j}\rvert. Then we find

Conversely, if di⩽djd_{i}\leqslant d_{j} or di−1⩾2∣di−dj∣d_{i}-1\geqslant 2\lvert d_{i}-d_{j}\rvert, we have di−1⩽2(dj−1)d_{i}-1\leqslant 2(d_{j}-1). Therefore, using (5.36) and the estimate M≍(1+ϕ)KM\asymp(1+\phi)K, we get

Putting (5.37), (5.38), and (5.39) together, we may estimate the first term of (5.34) in the case (ii) as

For the second term of (5.34) in the case (ii) we use the estimates

where the last term is bounded by Kδ1+ϕ1/2Mνi(A)νj(A)K^{\delta}\frac{1+\phi^{1/2}}{M\nu_{i}(A)\nu_{j}(A)}. Recalling (5.40), we find (5.35) in the case (ii). The case (iii) is dealt with in the same way.

What remains therefore is case (iv). For the first term of (5.34) we use the estimates

For the second term of (5.34) we use the estimates

which is (5.35). This concludes the analysis of case (iv), and hence of case (d).

Conclusion of the proof

Putting the cases (a)–(d) together, we have proved that (5.2) holds for arbitrary i,ji,j with an additional factor K2δK^{2\delta} multiplying the error term on the right-hand side. Since δ>0\delta>0 can be chosen arbitrarily small, (5.2) follows. This concludes the proof of Proposition 5.1. ∎

Non-outlier eigenvectors

We first consider eigenvectors near the right edge of the bulk spectrum. Recall the typical distance from μa\mu_{a} to the spectral edges, denoted by κa\kappa_{a} and defined in (2.18).

Fix τ∈(0,1/3)\tau\in(0,1/3). For a∈[ ⁣[s++1,(1−τ)K] ⁣]a\in[\![{s_{+}+1,(1-\tau)K}]\!] we have

Moreover, if a∈[ ⁣[1,s+] ⁣]a\in[\![{1,s_{+}}]\!] satisfies da⩽1+K−1/3+τd_{a}\leqslant 1+K^{-1/3+\tau} then

Proposition 6.1 has a close analogue for the left edge of the bulk spectrum, which holds under the additional condition ∣ϕ−1∣⩾τ\lvert\phi-1\rvert\geqslant\tau; we omit its detailed statement.

Suppose first that i∈Ri\in\mathcal{R}. Let ε>0\varepsilon>0 and set ω:=ε/2\omega\mathrel{\mathop{:}}=\varepsilon/2. Using (3.25), Remark 3.3, Theorem 2.3, and (3.15), we choose a high-probability event Ξ\Xi satisfying (4.7), (4.8), and

Abbreviating κ≡κ(μa)\kappa\equiv\kappa(\mu_{a}), we find from (3.20) that

which follows easily by spectral decomposition. Since i∈Ri\in\mathcal{R}, we get from (3.37), omitting the arguments zz for brevity,

where the last step follows from a resolvent expansion as in (5.14). We estimate the error terms using

where we used the definition of η\eta and (6.4). Hence a resolvent expansion yields

Next, we claim that for any fixed δ∈[0,1/3−ε)\delta\in[0,1/3-\varepsilon) we have the lower bound

whenever μa∈[θ(0),θ(1+K−1/3+δ+ε)]\mu_{a}\in[\theta(0),\theta(1+K^{-1/3+\delta+\varepsilon})]. To prove (6.10), suppose first that ∣d−1∣⩾1/2\lvert d-1\rvert\geqslant 1/2. By (3.21), there exists a constant c0>0c_{0}>0 such that for κ⩽c0\kappa\leqslant c_{0} we have ∣Re⁡wϕ+1∣⩽1/4\lvert\operatorname{Re}w_{\phi}+1\rvert\leqslant 1/4. Thus we get, for κ⩽c0\kappa\leqslant c_{0},

where we used that Im⁡wϕ⩽C\operatorname{Im}w_{\phi}\leqslant C by (3.19). Moreover, if κ⩾c0\kappa\geqslant c_{0} we find from (3.20) that Im⁡wϕ⩾c\operatorname{Im}w_{\phi}\geqslant c, from which we get

where in the second step we used ∣Re⁡wϕ∣⩽C\lvert\operatorname{Re}w_{\phi}\rvert\leqslant C as follows from (3.19). This concludes the proof of (6.10) for the case ∣d−1∣⩾1/2\lvert d-1\rvert\geqslant 1/2.

Suppose now that ∣d−1∣⩽1/2\lvert d-1\rvert\leqslant 1/2. Then we get

We shall estimate this using the elementary bound

For μa∈[θ(0),θ(1)]\mu_{a}\in[\theta(0),\theta(1)] we get from (6.11) with M=CM=C, recalling (3.20) and (3.21), that ∣1+dwϕ∣⩾c(∣d−1∣+Im⁡wϕ)\lvert 1+dw_{\phi}\rvert\geqslant c(\lvert d-1\rvert+\operatorname{Im}w_{\phi}). By a similar argument, for μa∈[θ(1),θ(1+K−1/3+δ+ε)]\mu_{a}\in[\theta(1),\theta(1+K^{-1/3+\delta+\varepsilon})] we set M=K2δM=K^{2\delta} and get (6.10) using (6.5) and (6.6). This concludes the proof of (6.10).

Using ∣z∣≍μa≍ϕ−1/2+ϕ1/2\lvert z\rvert\asymp\mu_{a}\asymp\phi^{-1/2}+\phi^{1/2} and (6.10), we estimate the first term on the right-hand side of (6.12) as

where in the last step we used that η⩽K−2/3+4ε+δ\eta\leqslant K^{-2/3+4\varepsilon+\delta}, as follows from (6.5) and (6.6).

Next, we estimate the second term of (6.12) as

Putting all three estimates together, we conclude that

in which case we get by choosing δ=0\delta=0 in (6.10) that

Next, if a⩽s+a\leqslant s_{+} satisfies da⩽K−1/3+τd_{a}\leqslant K^{-1/3+\tau} we get from (6.5), (6.6), and (3.20) that

In this case we have μa⩽θ(1+K−1/3+τ+ε)\mu_{a}\leqslant\theta(1+K^{-1/3+\tau+\varepsilon}) by (6.3), so that setting δ=τ\delta=\tau in (6.10) yields

Since ε>0\varepsilon>0 was arbitrary, (6.1) and (6.2) follow from (6.14) and (6.15) respectively. This concludes the proof of Proposition 6.1 in the case i∈Ri\in\mathcal{R}.

Finally, the case i∉Ri\notin\mathcal{R} is handled by replacing R\mathcal{R} with R^:=R∪{i}\widehat{\mathcal{R}}\mathrel{\mathop{:}}=\mathcal{R}\cup\{i\} and using a limiting argument, exactly as after (5.28). ∎

2. Proof of Theorems 2.11 and 2.17

We now have all the ingredients needed to prove Theorems 2.11 and 2.17.

The estimate (2.19) is an immediate corollary of (6.1) from Proposition 6.1. The estimate (2.20) is proved similarly (see also the remark following Proposition 6.1). ∎

We prove Theorem 2.11 using Propositions 5.1, 5.2, and 6.1. First we remark that it suffices to prove that (5.2) holds for A⊂OA\subset\mathcal{O} satisfying 1+K−1/3⩽dk⩽τ−11+K^{-1/3}\leqslant d_{k}\leqslant\tau^{-1} for all k∈Ak\in A. Indeed, supposing this is done, we get the estimate

from which Theorem 2.11 follows by noting that the second error term may be absorbed into the first, recalling that σi≍1+ϕ1/2\sigma_{i}\asymp 1+\phi^{1/2} for i∈Ai\in A, that M≍(1+ϕ)KM\asymp(1+\phi)K, and that di−1⩾K−1/3d_{i}-1\geqslant K^{-1/3}.

Fix ε>0\varepsilon>0. Note that there exists some s∈[1,∣R∣]s\in[1,\lvert\mathcal{R}\rvert] satisfying the following gap condition: for all kk such that dk>1+sK−1/3+εd_{k}>1+sK^{-1/3+\varepsilon} we have dk⩾1+(s+1)K−1/3+εd_{k}\geqslant 1+(s+1)K^{-1/3+\varepsilon}. The idea of the proof is to split A=A0⊔A1A=A_{0}\sqcup A_{1}, such that dk⩽1+sK−1/3+εd_{k}\leqslant 1+sK^{-1/3+\varepsilon} for k∈A0k\in A_{0} and dk⩾1+(s+1)K−1/3+εd_{k}\geqslant 1+(s+1)K^{-1/3+\varepsilon} for k∈A1k\in A_{1}. Note that such a splitting exists by the above gap property. Without loss of generality, we assume that A0≠∅A_{0}\neq\emptyset (for otherwise the claim follows from Proposition 5.1).

It suffices to consider the six cases (a) i,j∈A0i,j\in A_{0}, (b) i∈A0i\in A_{0} and j∈A1j\in A_{1}, (c) i∈A0i\in A_{0} and j∉Aj\notin A, (d) i,j∈A1i,j\in A_{1}, (e) i∈A1i\in A_{1} and j∉Aj\notin A, (f) i,j∉Ai,j\notin A.

We apply Cauchy-Schwarz and Proposition 6.1 to the first term, and Proposition 5.1 to the second term. Using the above gap condition, we find

where the last step follows from di−1⩽CK−1/3+εd_{i}-1\leqslant CK^{-1/3+\varepsilon}.

For this case it is crucial to use the stronger bound (5.4) and not (5.2). Hence, we need the non-overlapping condition (5.3). To that end, we assume first that (5.3) holds with δ:=ε\delta\mathrel{\mathop{:}}=\varepsilon. Thus, by the above gap assumption (5.3) also holds for A1A_{1}. In this case we get from (6.16) and Propositions 5.2 and 6.1 that

Clearly, the first two terms are bounded by the right-hand side of (5.2) times K3εK^{3\varepsilon}. The last term is estimated as

where we used that di−1⩽dj−1⩽C∣di−dj∣d_{i}-1\leqslant d_{j}-1\leqslant C\lvert d_{i}-d_{j}\rvert be the above gap condition. This concludes the proof in the case where the non-overlapping condition (5.3) holds.

(c), (e), (f) j∉A𝑗𝐴j\notin A

We use the splitting (6.16) and apply Cauchy-Schwarz and Proposition 6.1 to the first term, and Proposition 5.1 to the second term. Since νj(A1)⩽∣dj−1∣\nu_{j}(A_{1})\leqslant\lvert d_{j}-1\rvert in all cases, it is easy to prove that (6.16) is bounded by K3εK^{3\varepsilon} times the right-hand side of (5.2).

From (6.16) and Propositions 6.1 and 5.1 we get

From which we get (5.2) with the error term multiplied by K3εK^{3\varepsilon}.

Conclusion of the proof

We have proved that, for all i,j∈[ ⁣[1,M] ⁣]i,j\in[\![{1,M}]\!] and AA satisfying the assumptions of Theorem 2.11, the estimate (5.2) holds with an additional factor K3εK^{3\varepsilon} multiplying the error term. Since ε\varepsilon was arbitrary, we get (5.2). This concludes the proof. ∎

3. The law of the non-outlier eigenvectors

the typical distance between λa+1\lambda_{a+1} and λa\lambda_{a}. More precisely, the classical locations γa\gamma_{a} defined in (3.14) satisfy γa−γa+1≍Δa\gamma_{a}-\gamma_{a+1}\asymp\Delta_{a} for a⩽K/2a\leqslant K/2.

Let s++1⩽a⩽K1−τα+3s_{+}+1\leqslant a\leqslant K^{1-\tau}\alpha_{+}^{3} and define b:=a−s+b\mathrel{\mathop{:}}=a-s_{+}. Define the event

Informally, Proposition 6.2 expresses generalized components of the eigenvectors of QQ in terms of generalized components of eigenvectors of HH, under the assumption that Ω\Omega has high probability. We first show how Proposition 6.2 implies Theorem 2.20. This argument requires two key tools. The first one is level repulsion, which, together with the eigenvalue sticking from Theorem 2.7, will imply that Ω\Omega indeed has high probability. The second tool is quantum unique ergodicity (See Section 1.1) of the eigenvectors of HH, which establishes the law of the generalized components of the eigenvectors of HH.

The precise statement of level repulsion sufficient for our needs is as follows.

Fix τ∈(0,1)\tau\in(0,1). For any ε>0\varepsilon>0 there exists a δ>0\delta>0 such that for all a⩽K1−τa\leqslant K^{1-\tau} we have

The proof of Proposition 6.3 consists of two steps: (i) establishing (6.19) for the case of Gaussian XX and (ii) a comparison argument showing that if X(1)X^{(1)} and X(2)X^{(2)} are two matrix ensembles satisfying (1.15) and (1.16), and if (6.19) holds for X(1)X^{(1)}, then (6.19) also holds for X(2)X^{(2)}. Both steps have already appeared, in a somewhat different form, in the literature. Step (i) is performed in Lemma 6.4 below, and step (ii) in Lemma 6.5 below. Together, Lemmas 6.4 and 6.5 immediately yield Proposition 6.3.

Proposition 6.3 holds if XX is Gaussian.

We mimic the proof of Theorem 3.2 in BEY3 . Indeed, the proof from (BEY3, , Appendix D) carries over almost verbatim. The key input is the eigenvalue rigidity from Theorem 3.5, which for the model of BEY3 was established using a different method than Theorem 3.5. As in BEY3 , we condition on the eigenvalues {λi:i>K1−τ}\{\lambda_{i}\mathrel{\mathop{:}}i>K^{1-\tau}\}. On the conditioned measure, level repulsion follows as in BEY3 . Finally, thanks to Theorem 3.5 we know that the frozen eigenvalues {λi:i>K1−τ}\{\lambda_{i}\mathrel{\mathop{:}}i>K^{1-\tau}\} are with high probability near their classical locations. Note that for ϕ≈1\phi\approx 1, the rigidity estimate (3.15) only holds for indices i⩽(1−τ)Ki\leqslant(1-\tau)K; however, this is enough for the argument of (BEY3, , Appendix D), which is insensitive to the locations of eigenvalues at a distance of order one from the right edge γ+\gamma_{+}. We omit the full details. ∎

Let X(1)X^{(1)} and X(2)X^{(2)} be two matrix ensembles satisfying (1.15) and (1.16). Suppose that Proposition 6.3 holds for X(1)X^{(1)}. Then Proposition 6.3 also holds for X(2)X^{(2)}.

The proof of Lemma 6.5 relies on Green function comparison, and is given in Section 7.4.

For simplicity, and bearing the application to Theorem 2.20 in mind, in Proposition 6.6 we establish the convergence of a single generalized component of a single eigenvector. However, our method may be easily extended to yield

The proof of this generalization of Proposition 6.6 follows that of Proposition 6.6 presented in Section 7, requiring only heavier notation. In fact, our method may also be used to prove the universality of the joint eigenvalue-eigenvector distribution for any matrix QQ of the form (1.10) with Σ=TT∗=IM\Sigma=TT^{*}=I_{M}; see Theorem 8.3 below for a precise statement.

The proof of Proposition 6.6 is postponed to Section 7.

Supposing Proposition 6.2 holds, together with Propositions 6.3 and 6.6, we may complete the proof of Theorem 2.20.

Abbreviating b:=a−s+b\mathrel{\mathop{:}}=a-s_{+} and

Then, by assumption on aa, we may rewrite (6.18) as

The remainder of this section is devoted to the proof of Proposition 6.2.

We define the contour Γa\Gamma_{a} as the positively oriented circle of radius K−τ/5ΔaK^{-\tau/5}\Delta_{a} with centre λb\lambda_{b}. Let ε>0\varepsilon>0 and τ:=1/2\tau\mathrel{\mathop{:}}=1/2, and choose a high-probability event Ξ\Xi such that (4.7), (4.8), and (4.9) hold. For the following we fix a realization H∈Ω∩ΞH\in\Omega\cap\Xi. Define

By the residue theorem and the definition of Ω\Omega, we find

In order to compute (6.22), we need precise estimates for WW on Γa\Gamma_{a}. Because the contour Γa\Gamma_{a} crosses the branch cut of wϕw_{\phi}, we should not compare W(z)W(z) to wϕ(z)w_{\phi}(z) for z∈Γaz\in\Gamma_{a}. Instead, we compare W(z)W(z) to wϕ(z0)w_{\phi}(z_{0}), where

for all z∈Γaz\in\Gamma_{a}. To see this, we split

We estimate the first term of (6.24) by spectral decomposition, using that dist⁡(z,σ(H))⩾cη\operatorname{dist}(z,\sigma(H))\geqslant c\eta, similarly to (4.12). The result is

where we used (4.7), (4.9), and Lemma 3.6. Moreover, we estimate the second term of (6.24) using (4.8) as

The proof of (6.25) is analogous to that of (6.10), using (4.7) and the assumption on aa; we omit the details.

Armed with (6.23) and (6.25), we may analyse (6.22). A resolvent expansion in the matrix wϕ(z0)−W(z)w_{\phi}(z_{0})-W(z) yields

We estimate the third term using the bound

To prove (6.27), we note first that by (6.25) we have

By (6.23) and assumption on aa, it is easy to check that

We may now return to (6.26). The first term vanishes, the second is computed by spectral decomposition of WW, and the third is estimated using (6.27). This gives

Recalling (6.21) and (4.7), we therefore get

where we used ϕ1/2μa≍1+ϕ\phi^{1/2}\mu_{a}\asymp 1+\phi.

In order to simplify the leading term, we use

where we used Lemma 3.6. Moreover, we use that

Using that Ξ\Xi has high probability for all ε>0\varepsilon>0 and recalling the isotropic delocalization bound (3.13), we therefore get for any random HH that

We proved (6.28) under the assumption that i∈Ri\in\mathcal{R}, but a continuity argument analogous to that given after (5.28) implies that (6.28) holds for all i∈[ ⁣[1,M] ⁣]i\in[\![{1,M}]\!]. The above argument may be repeated verbatim to yield

Quantum unique ergodicity near the soft edge of H𝐻H

This section is devoted to the proof of Proposition 6.6.

Fix τ∈(0,1)\tau\in(0,1). Let hh be a smooth function satisfying

for some positive constant CC. Let a⩽K1−τa\leqslant K^{1-\tau} and suppose that λa\lambda_{a} satisfies (6.19) with some constants ε\varepsilon and δ\delta. Then for small enough δ1=δ1(ε,δ)\delta_{1}=\delta_{1}(\varepsilon,\delta) and δ2=δ2(ε,δ,δ1)\delta_{2}=\delta_{2}(\varepsilon,\delta,\delta_{1}) the following holds. Defining

By the assumption (7.1) on hh, rigidity (3.15), and delocalization (3.13), we can write

where we defined λa±:=λa±Kδ1η\lambda_{a}^{\pm}\mathrel{\mathop{:}}=\lambda_{a}\pm K^{\delta_{1}}\eta. For the following we choose

In order to obtain (7.3), we have to rewrite the integrand on the right-hand side of (7.5) in terms of

Hence (7.5) and (7.6) combined with the mean value theorem imply that the left-hand side of (7.3) is bounded by

for any fixed δ2∈(0,δ1)\delta_{2}\in(0,\delta_{1}). When applying the mean value theorem, we estimated the value of θ′(⋅)\theta^{\prime}(\cdot) using (7.1), the fact that all terms on the right-hand side of (7.6) are nonnegative, and the estimate

Next, using the eigenvalue rigidity from (3.15), it is not hard to see that there exists a constant C1C_{1} such that the contribution of ∣b−a∣⩾KC1δ2|b-a|\geqslant K^{C_{1}\delta_{2}} to (7.7) is bounded by K−δ2K^{-\delta_{2}}. In order to prove (7.3), therefore, it suffices to prove

which is the right-hand side of (7.9) provided δ2\delta_{2} is chosen small enough. Here in the first step we replaced λb\lambda_{b} with λa+1\lambda_{a+1} using the estimates λb⩽λa+1⩽E−Kδ1η\lambda_{b}\leqslant\lambda_{a+1}\leqslant E-K^{\delta_{1}}\eta valid for b>ab>a and EE in the support of χ\chi.

For b<ab<a, we partition I=I1∪I2I=I_{1}\cup I_{2} with I1∩I2=∅I_{1}\cap I_{2}=\emptyset and

Let us therefore consider the integral over I1I_{1}. One readily finds, for λa⩽λa−1⩽λb\lambda_{a}\leqslant\lambda_{a-1}\leqslant\lambda_{b}, that

Using delocalization (3.13) we therefore find that

In the next step, stated in Lemma 7.2 below, we replace the sharp cutoff function χ\chi in (7.3) with a smooth function of HH. Note first that from Lemma 7.1 and the rigidity (3.15), we get

The following result is the appropriate smoothed version of (7.11). It is a simple extension of Lemma 3.2 and Equation (5.8) from KY1 , and its proof is omitted.

We may now conclude the proof of Proposition 6.6.

For the comparison argument, we use the Green function comparison method applied to the Helffer-Sjöstrand representation of f(H)f(H). Using Lemma 7.2 it suffices to estimate

Thus we get the functional calculus, with G(z)=(H−z)−1G(z)=(H-z)^{-1},

As in Lemma 5.1 of KY1 , one can easily extend (3.9) to η\eta satisfying the lower bound η>0\eta>0 instead of η⩾K−1+ω\eta\geqslant K^{-1+\omega} in (3.7); the proof is identical to that of (KY1, , Lemma 5.1). Thus we have, for e∈[γ+−1,γ++1]e\in[\gamma_{+}-1,\gamma_{+}+1] and σ∈(0,1)\sigma\in(0,1),

Therefore, by the trivial symmetry σ↦−σ\sigma\mapsto-\sigma combined with complex conjugation, the third term on the right-hand side of (7) is bounded by

Recalling (7.1) and using the mean value theorem, we find from (7.13), (7.17), and (7.18) that for large enough dd, in order to estimate (7.14), and hence prove (6.20), it suffices to prove the following lemma. Note that in it we choose X(1)X^{(1)} to be the original ensemble and X(2)X^{(2)} to be the Gaussian ensemble. ∎

Suppose that the two M×NM\times N matrix ensembles X(1)X^{(1)} and X(2)X^{(2)} satisfy (1.15) and (1.16). Suppose that the assumptions of Lemma 7.1 hold, and recall the notations

Then for any d>1d>1 and for small enough ε≡ε(τ,d)>0\varepsilon\equiv\varepsilon(\tau,d)>0 and δ2≡δ2(ε,δ,δ1)\delta_{2}\equiv\delta_{2}(\varepsilon,\delta,\delta_{1}) we have

The rest of this section is devoted to the proof of Lemma 7.3.

We shall use the Green function comparison method EYY1 ; EYY3 ; KY1 to prove Lemma 7.3. For definiteness, we assume throughout the remainder of Section 7 that ϕ⩾1\phi\geqslant 1. The case ϕ<1\phi<1 is dealt with similarly, and we omit the details.

We first collect some basic identities and estimates that serve as a starting point for the Green function comparison argument. We work on the product probability space of the ensembles X(1)X^{(1)} and X(2)X^{(2)}. We fix a bijective ordering map Φ\Phi on the index set of the matrix entries,

and define the interpolating matrix XγX_{\gamma}, γ∈[ ⁣[1,MN] ⁣]\gamma\in[\![{1,MN}]\!], through

In particular, X0=X(1)X_{0}=X^{(1)} and XMN=X(2)X_{MN}=X^{(2)}. Hence we have the telescopic sum

Let us now fix a γ\gamma and let (b,β)(b,\beta) be determined by Φ(b,β)=γ\Phi(b,\beta)=\gamma. Throughout the following we consider b,βb,\beta to be arbitrary but fixed and often omit dependence on them from the notation. Our strategy is to compare Xγ−1X_{\gamma-1} with XγX_{\gamma} for each γ\gamma. In the end we shall sum up the differences in the telescopic sum (7.24).

Note that Xγ−1X_{\gamma-1} and XγX_{\gamma} differ only in the matrix entry indexed by (b,β)(b,\beta). Thus we may write

Here Xˉ\bar{X} is the matrix obtained from XγX_{\gamma} (or, equivalently, from Xγ−1X_{\gamma-1}) by setting the entry indexed by (b,β)(b,\beta) to zero. Next, we define the resolvents

where A\mathcal{A} is polynomial of degree two in UbβU_{b\beta} whose coefficients are Xˉ\bar{X}-measurable.

The rest of this section is therefore devoted to the proof of (7.27). Recall that we assume throughout that ϕ⩾1\phi\geqslant 1 for definiteness; in particular, K=NK=N.

We begin by collecting some basic identities from linear algebra. In addition to G(z):=(XX∗−z)−1G(z)\mathrel{\mathop{:}}=(XX^{*}-z)^{-1} we introduce the auxiliary resolvent R(z):=(X∗X−z)−1R(z)\mathrel{\mathop{:}}=(X^{*}X-z)^{-1}. Moreover, for μ∈[ ⁣[1,M] ⁣]\mu\in[\![{1,M}]\!] we split

We also define the resolvent G[μ]:=(X[μ](X[μ])∗−z)−1G^{[\mu]}\mathrel{\mathop{:}}=(X^{[\mu]}(X^{[\mu]})^{*}-z)^{-1}. A simple Neumann series yields the identity

Moreover, from (BEKYY, , Equation (3.11)), we find

Throughout the following we shall make use of the fundamental error parameter

which is analogous to the right-hand side of (3.9) and will play a similar role. We record the following estimate, which is analogous to Theorem 3.2.

This result is a generalization of (5.22) in PY . The key identity is (7.30). Since G[μ]G^{[\mu]} is independent of (Xiμ)i=1N(X_{i\mu})_{i=1}^{N}, we may apply the large deviation estimate (BEKYY, , Lemma 3.1) to G[μ]X[μ]G^{[\mu]}X_{[\mu]}. Moreover, ∣Rμμ∣≺1\lvert R_{\mu\mu}\rvert\prec 1, as follows from Theorem 3.2 applied to X∗X^{*}, and Lemma 3.6. Thus we get

where the second step follows by spectral decomposition, the third step from Theorem 3.2 applied to X[μ]X^{[\mu]} as well as (3.22), and the last step by definition of Ψ\Psi. This concludes the proof of (7.32).

Finally, (7.33) follows easily from Theorem 3.2 applied to the identity X∗GX=1+zRX^{*}GX=1+zR. ∎

Using (7.35), we may extend these estimates to analogous ones on Xˉ\bar{X} and TT instead of XγX_{\gamma} and SS. Indeed, using the facts ∥R∥⩽η−1\lVert R\rVert\leqslant\eta^{-1}, Ψ⩾N−1/2\Psi\geqslant N^{-1/2}, and ∣Ubβ∣≺ϕ−1/4N−1/2\lvert U_{b\beta}\rvert\prec\phi^{-1/4}N^{-1/2} (which are easily derived from the definitions of the objects on the left-hand sides) combined with (7.35), we get the following result.

For A∈{S,T}A\in\{S,T\} and B∈{Xγ,Xˉ}B\in\{X_{\gamma},\bar{X}\} we have

The final tool that we shall need is the following lemma, which collects basic algebraic properties of stochastic domination ≺\prec. We shall use them tacitly throughout the following. Their proof is an elementary exercise using union bounds and Cauchy-Schwarz. See (BEKYY, , Lemma 3.2) for a more general statement.

Suppose that A(v)≺B(v)A(v)\prec B(v) uniformly in v∈Vv\in V. If ∣V∣⩽NC\lvert V\rvert\leqslant N^{C} for some constant CC then ∑v∈VA(v)≺∑v∈VB(v)\sum_{v\in V}A(v)\prec\sum_{v\in V}B(v).

Suppose that A1≺B1A_{1}\prec B_{1} and A2≺B2A_{2}\prec B_{2}. Then A1A2≺B1B2A_{1}A_{2}\prec B_{1}B_{2}.

If the above random variables depend on an additional parameter uu and all hypotheses are uniform in uu then so are the conclusions.

2. Proof of Lemma 7.3 II: the main expansion

Lemma 7.5 contains the a-priori estimates needed to control the resolvent expansion (7.34). The precise form that we shall need is contained in the following lemma, which is our main expansion. Define the control parameter

The following results hold for E∈IE\in I. (Recall the definition (7.20). For brevity, we omit EE from our notation.)

where xlx_{l} is a polynomial, with constant number of terms, in the variables

where JlJ_{l} is a polynomial, with constant number of terms, in the variables

In each term of JlJ_{l}, T2T^{2} appears exactly once, while the indices bb and β\beta each appear exactly ll times.

Here all constants CC depend on the fixed parameter dd.

The proof is an application of the resolvent expansion (7.34) with m=3m=3 to the definitions of xx and yy.

and the same estimates hold if TT is replaced by SS. Note that in (7.42) we used the bound ∣mϕ−1∣≍ϕ−1/2\lvert m_{\phi^{-1}}\rvert\asymp\phi^{-1/2}, which follows from the identity

and Lemma 3.6. Using (7.42), it is not hard to conclude the proof of part (i).

What remains is to prove the bounds in part (iii). To that end, we integrate by parts, first in ee and then in σ\sigma, in the term containing fE′′(e)f_{E}^{\prime\prime}(e), and obtain

The same argument yields the error bound in (7.40). This concludes the proof. ∎

Next, using lemma 7.7 and NΔa≍κa−1/2N\Delta_{a}\asymp\kappa_{a}^{-1/2}, we find

Using (7.44) and (7.46), we may do a Taylor expansion of hh on the left-hand side of (7.27). This yields

For the general case, we still have to estimate the last line of (7.48). The terms that we need to analyse are

These terms are dealt with in the following lemma.

Let YY denote any term of (7.49). Then there is a constant c>0c>0 such that

3. Proof of Lemma 7.3 III: the terms of order three and proof of Lemma 7.8

Recall that we assume ϕ⩾1\phi\geqslant 1, i.e. K=NK=N; the case ϕ⩽1\phi\leqslant 1 is dealt with analogously, and we omit the details.

We first remark that using the bounds (7.46) we find

Comparing this to (7.50), we see that we need to gain an additional factor N−1/2N^{-1/2}. How to do so is the content of this subsection.

Let us outline the rough idea of the parity argument. We use the notations Xˉ=Xˉ[β]+Xˉ[β]\bar{X}=\bar{X}_{[\beta]}+\bar{X}^{[\beta]} and T[β](z):=(Xˉ[β](Xˉ[β])∗−z)−1T^{[\beta]}(z)\mathrel{\mathop{:}}=(\bar{X}^{[\beta]}(\bar{X}^{[\beta]})^{*}-z)^{-1}, in analogy to those introduced before (7.28). A simple example of a polynomial is

The need to gain an additional factor N−1/2N^{-1/2} from odd polynomials imposes nontrivial constraints on the polynomial coefficients, which are carefully stated in Definitions 7.10–7.12; they have been tailored to the class of polynomials generated by the terms (7.49).

We now move on to the proof of Lemma 7.8. We recall that we assume throughout that ϕ⩾1\phi\geqslant 1. We first introduce a family of graded polynomials suitable for our purposes. It depends on a constant C0C_{0}, which we shall fix during the proof to be some large but fixed number.

Let ϱ=(ϱi:i∈[ ⁣[1,ϕN] ⁣])\varrho=(\varrho_{i}\mathrel{\mathop{:}}i\in[\![{1,\phi N}]\!]) be a family of deterministic nonnegative weights. We say that ϱ\varrho is an admissible weight if

be a polynomial in Xˉ\bar{X}. Analogously to the notation O≺(⋅)O_{\prec}(\cdot) introduced in Definition 2.1, we write P=O≺,d(A)\mathcal{P}=O_{\prec,d}(A) if the following conditions are satisfied.

AA is deterministic and Vi1⋯idV_{i_{1}\cdots i_{d}} is Xˉ[β]\bar{X}^{[\beta]}-measurable.

There exist admissible weights ϱ(1),…,ϱ(d)\varrho^{(1)},\dots,\varrho^{(d)} such that

We have the deterministic bound ∣Vi1⋯id∣⩽NC0\lvert V_{i_{1}\cdots i_{d}}\rvert\leqslant N^{C_{0}}.

The above definition extends trivially to the case d=0d=0, where P=V\mathcal{P}=V is Xˉ[β]\bar{X}^{[\beta]}-measurable.

Let P\mathcal{P} be a polynomial of the form

We write P  =  O≺,⋄(A)\mathcal{P}\;=\;O_{\prec,\diamond}(A) if ViV_{i} is Xˉ[β]\bar{X}^{[\beta]}-measurable, ∣Vi∣⩽NC0\lvert V_{i}\rvert\leqslant N^{C_{0}}, and ∣Vi∣≺A\lvert V_{i}\rvert\prec A for some deterministic AA.

We write P=O≺,even(A)\mathcal{P}=O_{\prec,\text{\rm{even}}}(A) if P\mathcal{P} is a sum of at most C0C_{0} terms of the form

where n,m⩽C0n,m\leqslant C_{0} and AA is deterministic.

Moreover, we write P=O≺,odd(A)\mathcal{P}=O_{\prec,\text{\rm{odd}}}(A) if P=P^ Peven\mathcal{P}=\widehat{\mathcal{P}}\,\mathcal{P}_{\text{\rm{even}}}, where P^=O≺,1(1)\widehat{\mathcal{P}}=O_{\prec,1}(1) and Peven=O≺,even(A)\mathcal{P}_{\text{\rm{even}}}=O_{\prec,\text{\rm{even}}}(A).

Definitions 7.10–7.12 refine Definition 2.1 in the sense that

Indeed, let P=O≺,d(A)\mathcal{P}=O_{\prec,d}(A) be of the form (7.53). Then a simple large deviation estimate (e.g. a trivial extension of (EKYY3, , Theorem B.1 (iii))) yields

where the last step follows from the definition of admissible weights. Similarly, if P=O≺,⋄(A)\mathcal{P}=O_{\prec,\diamond}(A) is of the form (7.55), a large deviation estimate (e.g.(EKYY3, , Theorem B.1 (i))) yields

after possibly increasing C0C_{0}. (As with the standard big O notation, such expressions are to be read from left to right.) We stress that such operations may be performed an arbitrary, but bounded, number of times. It is a triviality that all of the following arguments will involve at most C0C_{0} such algebraic operations on graded polynomials, for large enough C0C_{0}.

The point of the graded polynomials is that bounds of the form (7.56) are improved if dd is odd and we take the expectation. The precise statement is the following.

Let P=O≺,odd(A)\mathcal{P}=O_{\prec,\text{\rm{odd}}}(A) for some deterministic A⩽NCA\leqslant N^{C}. Then for any fixed D>0D>0 we have

It suffices to set A=1A=1 and consider P=P^P0∏s=1mPi\mathcal{P}=\widehat{\mathcal{P}}\mathcal{P}_{0}\prod_{s=1}^{m}\mathcal{P}_{i}, where P^\widehat{\mathcal{P}}, P0\mathcal{P}_{0}, and Pi\mathcal{P}_{i} are as in Definition 7.12. By linearity, it suffices to consider

where d=2nd=2n is even. We suppose that ∣Wi0∣≺ϱi0(0)\lvert W_{i_{0}}\rvert\prec\varrho^{(0)}_{i_{0}}, ∣Vi1⋯id∣≺ϱi1(1)⋯ϱid(d)\lvert V_{i_{1}\cdots i_{d}}\rvert\prec\varrho^{(1)}_{i_{1}}\cdots\varrho^{(d)}_{i_{d}}, and ∣Vid+l(d+l)∣≺1\lvert V^{(d+l)}_{i_{d+l}}\rvert\prec 1 for l=d+1,…,d+ml=d+1,\dots,d+m. Here ϱik(k)\varrho^{(k)}_{i_{k}} denotes an admissible weight (see Definition 7.9). Thus we have

The expectation imposes that each summation index i0,…,id+mi_{0},\dots,i_{d+m} coincide with at least one other summation index. Thus we get

Hence the summation over i1,…,id+mi_{1},\dots,i_{d+m} factors into a product over the blocks of PP. We shall show that the contribution of each block is at most one, and that there is a block whose contribution is at most N−1/2N^{-1/2}.

Fix p∈Pp\in P and denote by SpS_{p} the contribution of the block pp to the summation in the main term of (7.57). Define s:=∣p∩[ ⁣[0,d] ⁣]∣s\mathrel{\mathop{:}}=\lvert p\cap[\![{0,d}]\!]\rvert and t:=∣p∩[ ⁣[d+1,d+m] ⁣]∣t\mathrel{\mathop{:}}=\lvert p\cap[\![{d+1,d+m}]\!]\rvert. By definition of PP, we have s+t⩾2s+t\geqslant 2. By the inequality of arithmetic and geometric means, we have

Moreover, since dd is even, at least one block of PP satisfies (s,t)≠(2,0)(s,t)\neq(2,0).

Since d+m⩽2C0d+m\leqslant 2C_{0}, the proof is complete. ∎

We begin by noting that (3.9) applied to X[μ]X^{[\mu]} and (3.22) combined with a large deviation estimate (see (EKYY3, , Theorem B.1)) yields

Using (7.35) and Lemma 7.5, it is not hard to deduce that

where in the second step we used (3.5) and (3.23). Now we split

where in the second step we used the estimates ∣Tij[β]−δijmϕ−1∣≺ϕ−1Ψ\lvert T_{ij}^{[\beta]}-\delta_{ij}m_{\phi^{-1}}\rvert\prec\phi^{-1}\Psi and ∣mϕ−1∣⩽Cϕ−1/2\lvert m_{\phi^{-1}}\rvert\leqslant C\phi^{-1/2}. Since ∣zmϕ∣⩽Cϕ1/2\lvert zm_{\phi}\rvert\leqslant C\phi^{1/2}, we therefore conclude that

From (7.31) and the definition of η\eta, we readily find that Ψ⩽N−cτ\Psi\leqslant N^{-c\tau} for some constant cc. Therefore choosing n≡n(τ,D)n\equiv n(\tau,D) large enough yields

Having established (7.64), the remainder of the proof is relatively straightforward. From (7.29) and (7.30) we get

Now (7.62) follows easily from (7.65) and (7.64).

Moreover, (7.59) and (7.61) follow from (7.28) combined with (7.65) and (7.64). For (7.61) we estimate the second term in (7.28) by

where in the last step we used that Ψ⩾cN−1/2\Psi\geqslant cN^{-1/2}. Moreover, (7.60) is a trivial consequence of (7.59). Finally, (7.63) follows from (7.28) and (7.64) combined with

In each application of Lemma 7.14, we shall verify one of the conditions of (7.66). The first condition is verified for η⩽N−2/3\eta\leqslant N^{-2/3}, which always holds for the coefficients of x1x_{1}, x2x_{2}, and x3x_{3} (recall (7.2)).

The second condition of (7.66) will be verified when computing the coefficients of y1y_{1}, y2y_{2}, and y3y_{3}. To that end, we make use of the freedom of the choice of basis when computing the trace in the definition of J1J_{1}, J2J_{2}, and J3J_{3}. We shall choose a basis that is completely delocalized. The following simple result guarantees the existence of such a basis.

Fix D>0D>0. Then there exists C0=C0(D)C_{0}=C_{0}(D) such that

where the parity of JiJ_{i} follows easily from its definition.

Lemmas 7.14 and 7.17 are the key estimates of the coefficients appearing in (7.49). We claim that all estimates of Lemma 7.7, along with (7.42), remain valid, in the sense that an estimate of the form ∣u∣≺v\lvert u\rvert\prec v is to be replaced with

where the parity of xix_{i} may be easily deduced from their definitions. Moreover, for the estimates (7.41) we use (7.72) to get

Note that, thanks to Lemmas 7.14 and 7.17, we have obtained exactly the same upper bounds on the coefficients xix_{i} and yiy_{i} as the ones obtained in Lemma 7.7, but we have in addition expressed them, up to a negligible error, as graded polynomials, to which Lemma 7.13 is applicable.

which may be derived from (7.68), combined with a Taylor expansion of q(m)q^{(m)}. Similarly, we find that

We may now put everything together. Noting that the degree of the polynomializations of the expressions (7.49) is always odd, we obtain, in analogy to (7.51) that

for YY being any term of (7.49). Hence Lemma 7.8 follows from Lemma 7.13 and Young’s inequality.

4. Stability of level repulsion: proof of Lemma 6.5

This is a Green function comparison argument, using the machinery introduced in Section 7.1. A similar comparison argument was given in Propositions 2.4 and 2.5 of KY1 . The details in the sample covariance case and for indices aa satisfying a⩽K1−τa\leqslant K^{1-\tau} follow an argument very similar to (in fact simpler than) the one from Sections 7.1–7.3. As in the proofs of Propositions 2.4 and 2.5 of KY1 , one writes the level repulsion condition in terms of resolvents. In our case, one uses the representation (7) as the starting point. Then the machinery of Sections 7.1–7.3 may be applied with minor modifications. We omit the details.

Extension to general T𝑇T and universality for the uncorrelated case

In this section we relax the assumption (3.1), and hence extend all arguments of Sections 3–7 to cover general TT. We also prove the fixed-index joint eigenvector-eigenvalue universality of the matrix HH defined in (2.7), for indices bounded by K1−τK^{1-\tau} for some τ>0\tau>0.

Bearing the applications in the current paper in mind, we state the results of this section for the matrix HH from (2.7), but it is a triviality that all results and their proofs carry over to case of arbitrary QQ from (1.10) provided that Σ=TT∗=IM\Sigma=TT^{*}=I_{M}.

We start with the singular value decomposition of TT, which we write as

where H:=YY∗H\mathrel{\mathop{:}}=YY^{*} and Y:=(IM,0)OXY\mathrel{\mathop{:}}=(I_{M},0)OX were defined in (2.7). Comparing this to (3.2), we find that to relax the assumption (3.1) we have to generalize the arguments of Sections 3–7 by replacing XX∗XX^{*} with H=YY∗H=YY^{*}.

The generalization of G=(XX∗−z)−1G=(XX^{*}-z)^{-1} is the resolvent of YY∗YY^{*},

where we recall the definition (7.31) of Ψ\Psi. In fact, from Lemma 3.6 and (3.23) we find that Φ/∣mϕ−1∣⩽N−c\Phi/\lvert m_{\phi^{-1}}\rvert\leqslant N^{-c} for some positive constant cc depending on τ\tau. Hence (3.9) yields

Having established Theorem 8.1, all arguments from Sections 3–6 that use it as input may be taken over verbatim, after replacing GG by G^\widehat{G}. More precisely, all results from Sections 3–6 remain valid for a general QQ, with the exception of Proposition 6.3, Lemmas 6.4 and 6.5, and Proposition 6.6. Therefore we have completed the proofs of all of our main results except Theorem 2.20.

In order to prove Theorem 2.20, we still have to prove Lemmas 6.4 and 6.5 and Proposition 6.6 for YY∗YY^{*} instead of XX∗XX^{*}. Lemma 6.4 is easy: for Gaussian XX we have Y=d(IM,0)X=X~Y\overset{d}{=}(I_{M},0)X=\widetilde{X}, where X~\widetilde{X} is the M×NM\times N matrix obtained from XX by deleting its bottom rr rows.

The proofs of Lemma 6.5 and Proposition 6.6 rely on Green function comparison. What remains, therefore, is to extend the argument of Section 7 from H=XX∗H=XX^{*} to H=YY∗H=YY^{*}.

Lemma 7.3 remains valid if x(E)x(E) and y(E)y(E) are replaced with x^(E)\widehat{x}(E) and y^(E)\widehat{y}(E), obtained from the definitions (7.22) and (7.23) by replacing GG with G^\widehat{G}.

Recalling (3.9) and (7.31), we find that the second term is stochastically dominated by

where in the second step we used that Ψ⩽C(Nη)−1\Psi\leqslant C(N\eta)^{-1}, as follows from Lemma 3.6 and the definition of η\eta in (7.2). Recalling the definitions from (7.2), we therefore conclude that for small enough ε≡ε(τ)\varepsilon\equiv\varepsilon(\tau) we have

for some positive constant cc depending on τ\tau.

for some positive constant cc depending on τ\tau. Plugging (8.5) into the definition of y^(E)\widehat{y}(E) and estimating the error term using integration by parts, as in (7.43), we get

Using the mean value theorem and the bound ∣y(E)∣≺1\lvert y(E)\rvert\prec 1, we therefore get

Combined with (7.21), this concludes the proof. ∎

This concludes the proof of Theorem 2.20 for the case of general TT.

In this section we observe that the technology developed in Section 7 allows us to establish the universality of the joint eigenvalue-eigenvector distribution of QQ provided that Σ=IM\Sigma=I_{M}. Without loss of generality, we consider the case where QQ is given by H=YY∗H=YY^{*} defined in (2.7). This result applies to arbitrary eigenvalue and eigenvector indices which are bounded by K1−τK^{1-\tau}, and does in particular not need to invoke eigenvalue correlation functions.

This result generalizes the quantum unique ergodicity from Proposition 6.6 and its extension from Remark 6.7 by also including the distribution of the eigenvalues. The universality of both the eigenvalues and the eigenvectors is formulated in the sense of fixed indices. A result in a similar spirit was given in (KY1, , Theorem 1.6), except that the upper bound on the eigenvalue and eigenvector indices (log⁡K)Clog⁡log⁡N(\log K)^{C\log\log N} from KY1 is improved all the way to K1−τK^{1-\tau}, for any τ>0\tau>0. A result covering all eigenvalue and eigenvector indices, i.e. with an index upper bound KK, was given in (KY1, , Theorem 1.10) and (TV3, , Theorem 8), but under the assumption of a four-moment matching assumption. Theorem 8.3 is a true universality result in that it does not require any moment matching assumptions, but it does require an index upper bound of K1−τK^{1-\tau} instead of KK on the eigenvalue and eigenvector indices.

for some constant c≡c(τ,k,r,h)>0c\equiv c(\tau,k,r,h)>0.

The proof is a Green function comparison argument, a minor modification of that developed in Section 7. We write the distribution of λa−γa\lambda_{a}-\gamma_{a} in terms of the resolvent G^\widehat{G}, starting from the Helffer-Sjöstrand representation (7), exactly as in (KY1, , Sections 4 and 5). We omit further details. ∎

We note that even this fixed-index universality of eigenvalues is a new result, having previously only been established under the four-moment matching condition KY1 ; TV3 (in the context of Wigner matrices).

We formulated Theorem 8.3 for the real symmetric covariance matrices of the form (2.7), but it and its proof remain valid for complex Hermitian covariance matrices, as well as Wigner matrices (both real symmetric and complex Hermitian).

Assuming ∣ϕ−1∣>τ\lvert\phi-1\rvert>\tau, the condition a⩽K1−τa\leqslant K^{1-\tau} on the indices in Theorem 8.3 may be replaced with a∉[ ⁣[K1−τ,K−K1−τ] ⁣]a\notin[\![{K^{1-\tau},K-K^{1-\tau}}]\!].

Extension to Q˙˙𝑄\dot{Q} and proof of Theorem 2.23

In this section we explain how to extend our analysis from QQ defined in (1.10) to Q˙\dot{Q} defined in (2.23), hence proving Theorem 2.23. We define the resolvent

which will replace G(z)=(XX∗−z)−1G(z)=(XX^{*}-z)^{-1} when analysing with Q˙\dot{Q} instead of QQ. We begin by noting that the isotropic local laws hold for also for G˙\dot{G}.

As in the proof of Theorem 8.1, we only prove (3.9) for G˙\dot{G}. Using the identity (3.36) we get

Using (3.9), the proof will be complete provided we can show that

Lemma 7.3 remains valid if x(E)x(E) and y(E)y(E) are replaced with x˙(E)\dot{x}(E) and y˙(E)\dot{y}(E), obtained from the definitions (7.22) and (7.23) by replacing GG with G˙\dot{G}.

The proof mirrors closely that of Lemma 8.2, using the identity (9.1) instead of (8.2) as input. We omit the details. ∎

All that remains is the proof of the following estimate, which generalizes (7.32).

here we used the first identity of (7.30). A similar argument was given in (BEKYY, , Section 5). The basic idea is to make all resolvents on the right-hand side of (9.4) independent of the columns of XX indexed by {μ1,…,μp}\{\mu_{1},\dots,\mu_{p}\} (see (BEKYY, , Definition 3.7)). As in (BEKYY, , Section 5), we do this using the identities from (BEKYY, , Lemma 3.8) for the entries of RR. In addition, for the entries of GG we use the identity (in the notation of (BEKYY, , Definition 3.7))

Appendix A A few remarks on applications to statistics

In this appendix we give a few remarks on what our results imply for applications to statistics. We assume throughout that the population covariance matrix satisfies

We consider the following simple model problem. Suppose there is some (unknown) set S⊂{1,…,M}S\subset\{1,\dots,M\} whose associated variables (ak)k∈S(a_{k})_{k\in S} are strongly correlated. For simplicity, let us assume that the correlations are given by a single spike in Σ\Sigma, i.e.

The most naive way to proceed is to compare the entries of Q\mathcal{Q} with those of Σ\Sigma. Using (A.1) it is not hard to conclude that

We look at the off-diagonal terms of Q\mathcal{Q} and infer that kk belongs to SS if there exists an index ll such that Qkl\mathcal{Q}_{kl} is much larger than N−1/2N^{-1/2}. For this approach to work, we require that ∣Σkl∣≫N−1/2\lvert\Sigma_{kl}\rvert\gg N^{-1/2}, which reads (σ−1)v(k)v(l)≫N−1/2(\sigma-1)v(k)v(l)\gg N^{-1/2}. We conclude that this naive entrywise approach works provided that

The principal component analysis works and the naive componentwise approach does not if (A.4) holds and (A.3) does not. These conditions may be written as

Hence the principal component analysis for this example is very effective when the family of correlated variables is quite large, ∣S∣≫ϕ1/2M\lvert S\rvert\gg\phi^{1/2}\sqrt{M}.

More generally, the principal component approach may work in the regime

and cannot work for smaller ∣S∣\lvert S\rvert. Indeed, the assumption (A.1) is satisfied for σ⩽∣S∣\sigma\leqslant\lvert S\rvert, so that in the case of the strongest possible correlations, σ≍∣S∣\sigma\asymp\lvert S\rvert, the condition (A.4) reduces to (A.5). On the other hand, if ∣S∣≪ϕ1/2\lvert S\rvert\ll\phi^{1/2}, then from (A.2) and the assumption (A.1) we find σ−1⩽∣S∣≪ϕ1/2\sigma-1\leqslant\lvert S\rvert\ll\phi^{1/2}, in contradiction to (A.4).

Clearly, by (A.4), for the purposes of statistical inference it is desirable to make ϕ\phi as small as possible. This means that the number of samples per variables is as large as possible. It is therefore natural to attempt to make MM smaller so as to reduce ϕ\phi. Obviously, if we know a priori that SS is contained in some subset YY of {1,…,M}\{1,\dots,M\} of size M/2M/2, then we simply discard all variables indexed by YcY^{c} and consider the correlations restricted to (ai)i∈Y(a_{i})_{i\in Y}; we have halved ϕ\phi in the process.

However, if we have no such a priori knowledge about SS, discarding half of the variables aia_{i} is a bad idea. In this case, the best one can do is to choose YY at random. Thus, suppose that SS is uniformly distributed among the subsets of {1,…,M}\{1,\dots,M\} of size ∣S∣\lvert S\rvert. We cut the sample space {1,…,M}\{1,\dots,M\} in half by keeping only the M/2M/2 first elements. We therefore obtain a new family of variables with dimensional parameters

with high probability, we find that v~(k)≈2v(k)\widetilde{v}(k)\approx\sqrt{2}v(k) for k∈Sk\in S. Picking an entry Σkl=Σ~kl\Sigma_{kl}=\widetilde{\Sigma}_{kl} for k,l∈S{1,…,M/2}k,l\in S\mathcal{\{}1,\dots,M/2\}, we obtain

We conclude that detecting spikes in the new problem is more difficult than in the original problem, and the halving of sample space is therefore counterproductive unless one has some good a priori information about Σ\Sigma.

Using v(k)2=∣S∣−1v(k)^{2}=\lvert S\rvert^{-1} for k∈Sk\in S, we therefore get the condition (1+ϕ)1/2≫∣S∣(1+\phi)^{1/2}\gg\lvert S\rvert. This however cannot hold, since we have by assumption (A.1) on Σ\Sigma that

References