On the principal components of sample covariance matrices
Alex Bloemendal, Antti Knowles, Horng-Tzer Yau, Jun Yin
Introduction
In this paper we investigate sample covariance matrices of the form
Correspondingly, one has to subtract from the empirical mean of the -th row of , which we denote by . Hence, we replace (1.1) with
By the law of large numbers, if is fixed and taken to infinity, the sample covariance matrix converges almost surely to the population covariance matrix . In many modern applications, however, the population size is very large and obtaining samples is costly. Thus, one is typically interested in the regime where is of the same order as , or even larger. In this case, as it turns out, the behaviour of changes dramatically and the problem becomes much more difficult. In principal component analysis, one seeks to understand the correlations by considering the principal components, i.e. the top eigenvalues and associated eigenvectors, of . These provide an effective low-dimensional projection of the high-dimensional data set , in which the significant trends and correlations are revealed by discarding superfluous data.
the empirical eigenvalue density of the rescaled matrix has the same asymptotics for large and as
2. Examples and outline of the model
The population covariance matrix is .
Below we shall refer to these examples as Examples (1) and (2) respectively. Motivated by them, we now outline our model. Let be an matrix whose entries are independent with zero mean and unit variance. We choose a deterministic matrix , and set . We stress that we do not assume that the underlying randomness is Gaussian. Our key assumptions are (i) is bounded; (ii) has bounded rank; (iii) is comparable to ; (iv) the entries of are independent, with zero mean and unit variance, and have a sufficient number of bounded moments. The precise assumptions are given in Section 1.3 below. We emphasize that everything apart from and the rank of is allowed to depend on in an arbitrary fashion.
3. Definition of model
In this section we give the precise definition of our model and introduce some basic notations. For convenience, we always work with the rescaled sample covariance matrix
The motivation behind this rescaling is that, as observed in (1.6), it ensures that the bulk spectrum of has asymptotically a fixed diameter, , for arbitrary and .
We always regard as the fundamental large parameter, and write . Here, and throughout the following, in order to unclutter notation we omit the argument in quantities, such as , that depend on it. In other words, every symbol that is not explicitly a constant is in fact a sequence indexed by . We assume that and satisfy the bounds
Fix a constant . Let be an random matrix and an deterministic matrix. For definiteness, and bearing the motivation of sample covariance matrices in mind, we assume that the entries of and are real. However, our method also trivially applies to complex-valued and , with merely cosmetic changes to the proofs. We consider the matrix
Since is an matrix, we find that has
We define the population covariance matrix
for the eigenvalues . We always order the values such that
We suppose that is positive definite, so that each lies in the interval
Moreover, we suppose that has bounded rank, i.e.
We assume that the entries of are independent (but not necessarily identically distributed) random variables satisfying
Our results concern the eigenvalues of , denoted by
and the associated unit eigenvectors of , denoted by
4. Sketch of behaviour of the principal components of Q𝑄Q
To guide the reader, we now give a heuristic description of the behaviour of principal components of . The spectrum of consists of a bulk spectrum and of outliers—eigenvalues separated from the bulk. The bulk contains an order eigenvalues, which are distributed on large scales according to the Marchenko-Pastur law (1.6). In addition, if there are trivial eigenvalues at zero. Each satisfying gives rise to an outlier located near its classical location
Any satisfying does not result in an outlier. We summarize this picture in Figure 1.1.
The creation or annihilation of an outlier as a crosses is known as the BBP phase transition BBP . It takes place on the scaleWe use the symbol to denote quantities of comparable size; see “Conventions” at the end of this section for a precise definition. \bigl{\lvert}\lvert d_{i}\rvert-1\bigr{\rvert}\asymp K^{-1/3}. This scale has a simple heuristic explanation (we focus on the right edge of the spectrum). Suppose that and all other ’s are zero. Then the top eigenvalue exhibits universality, and fluctuates on the scale around (see Theorem 8.3 and Remark 8.7 below). Increasing beyond the critical value , we therefore expect to become an outlier when its classical location is located at a distance greater than from . By a simple Taylor expansion of , the condition becomes .
for . The function determines the aperture of the cone. Note that and converges to as . See Figure 1.2.
5. Summary of previous related results
There is an extensive literature on spiked covariance matrices. So far most of the results have focused on the outlier eigenvalues of Example (1), with the nonzero independent of and fixed. Eigenvectors and the non-outlier eigenvalues have seen far less attention.
For the uncorrelated case and Gaussian in (1.10) with fixed , it was proved in Joh1 for the complex case and in Johnstone for the real case that the top eigenvalue, rescaled as , is asymptotically distributed according the Tracy-Widom law of the appropriate symmetry class TW1 ; TW2 . Subsequently, these results were shown to be universal, i.e. independent of the distribution of the entries of , in Sosh2 ; PY . The assumption that be fixed was relaxed in Pe1 ; ElK1 .
The study of covariance matrices with nontrivial population covariance matrix goes back to the seminal paper of Johnstone Johnstone , where the Gaussian spiked model was introduced. The BBP phase transition was established by Baik, Ben Arous, and Péché BBP for complex Gaussian , fixed rank of , and fixed . Subsequently, the results of BBP were extended to the other Gaussian symmetry classes, such as real covariance matrices, in BV1 ; BV2 . The proofs of BBP ; Pec use an asymptotic analysis of Fredholm determinants, while those of BV1 ; BV2 use an explicit tridiagonal representation of ; both of these approaches rely heavily on the Gaussian nature of . See also PeBor for a generalization of the BBP phase transition.
For the model from Example (1) with fixed nonzero and , the almost sure convergence of the outliers was established in BS . It was also shown in BS that if for all , the top eigenvalue converges to . For this model, a central limit theorem of the outliers was proved in BY1 . In BY2 , the almost sure convergence of the outliers was proved for a generalized spiked model whose population covariance matrix is of the block diagonal form , where is a fixed matrix and is chosen so that the associated sample covariance matrix has no outliers.
In BGN , the almost sure convergence of the projection of the outlier eigenvectors onto the finite-dimensional spike subspace was established, under the assumption that and the nonzero are fixed, and that and are both random and one of them is orthogonally invariant. In particular, the cone concentration from (1.18) was established in BGN . In Paul , under the assumption that is Gaussian and and the nonzero are fixed, a central limit theorem for a certain observable, the so-called sample vector, of the outlier eigenvectors was established. The result of Paul was extended to non-Gaussian entries for a special class of in Shi .
Moreover, in BGN2 ; NA1 result analogous to those of BGN were obtained for the model from Example (2). Finally, a related class of models, so-called deformed Wigner matrices, have been the subject of much attention in recent years; we refer to SoshPert ; SoshPert2 ; KY2 ; KY3 for more details; in particular, the joint distribution of all outliers was derived in KY3 .
6. Overview of results
In this subsection we give an informal overview of our results.
Our results on the eigenvalues of consist of two parts. First, we derive large deviation bounds on the locations of the outliers (Theorem 2.3). Second, we prove eigenvalue sticking for the non-outliers (Theorem 2.7), whereby each non-outlier “sticks” with high probability and very accurately to the eigenvalues of a related covariance matrix satisfying and whose top eigenvalues exhibit universality. As a corollary (Remark 8.7), we prove that the top non-outlier eigenvalue of has asymptotically the Tracy-Widom-1 distribution. This sticking is very accurate if all ’s are separated from the critical point , and becomes less accurate if a is in the vicinity of . Eventually, it breaks down precisely on the BBP transition scale , at which the Tracy-Widom-1 distribution is known not to hold for the top non-outlier eigenvalue. These results generalize those from (KY2, , Theorem 2.7).
If the outlier approaches the bulk spectrum or another outlier, the cone concentration becomes less accurate. For the case of two nearby outlier eigenvalues, for instance, the cone concentration (1.18) of the eigenvectors breaks down when the distributions of the outlier eigenvalues have a nontrivial overlap. In order to understand this behaviour in more detail, we introduce the deterministic projection
Finally, the proofs of universality of the non-outlier eigenvalues and eigenvectors require the universality of for the uncorrelated case as input. This universality result is given in Theorem 8.3, which is also of some independent interest. It establishes the joint, fixed-index, universality of the eigenvalues and eigenvectors of (and hence, as a special case, the quantum unique ergodicity of the eigenvectors of mentioned in Section 1.1). It works for all eigenvalue indices satisfying for any fixed .
We conclude this subsection by outlining the key novelties of our work.
We introduce the general models from (1.10) and from (2.23) below, which subsume and generalize several models considered previously in the literatureIn particular, the current paper is the first to study the principal components of a realistic sample covariance matrix (1.3) instead of the zero mean case (1.1).. We allow the entries of to be arbitrary random variables (up to a technical assumption on their tails). All quantities except and the rank of may depend on . We make no assumption on beyond the bounded-rank condition of . The dimensions and may be wildly different, and are only subject to the technical condition (1.9).
We obtain quantitative bounds (i.e. rates of convergence) on the outlier eigenvalues and the generalized components of the eigenvectors. We believe these bounds to be optimal.
We establish the joint, fixed-index, universality of the eigenvalues and eigenvectors for the case . This result holds for any eigenvalue indices satisfying for an arbitrary . Note that previous works KY1 ; TV3 (established in the context of Wigner matrices) required either the much stronger condition or a four-moment matching condition.
We remark that the large deviation bounds derived in this paper also allow one to derive the joint distribution of the generalized components of the outlier eigenvectors; this will be the subject of future work.
Conventions
The fundamental large parameter is . All quantities that are not explicitly constant may depend on ; we almost always omit the argument from our notation.
We use in various assumptions to denote a positive constant that may be chosen arbitrarily small. A smaller value of corresponds to a weaker assumption. All of our estimates depend on , and we neither indicate nor track this dependence.
Results
In this section we state our main results. The following notion of a high-probability bound was introduced in EKY2 , and has been subsequently used in a number of works on random matrix theory. It provides a simple way of systematizing and making precise statements of the form “ is bounded with high probability by up to small powers of ”.
be two families of nonnegative random variables, where is a possibly -dependent parameter set. We say that is stochastically dominated by , uniformly in , if for all (small) and (large) we have
for large enough . Throughout this paper the stochastic domination will always be uniform in all parameters (such as matrix indices) that are not explicitly fixed. Note that may depend on the constants from (1.9) and (1.16) as well as any constants fixed in the assumptions of our main results. If is stochastically dominated by , uniformly in , we use the notation . Moreover, if for some complex family we have we also write .
Because of (1.9), all (or some) factors of in Definition (2.1) could be replaced with without changing the definition of stochastic domination.
We begin with results on the locations of the eigenvalues of . These results will also serve as a fundamental input for the proofs of the results on eigenvectors presented in Sections 2.2 and 2.3.
Recall that has zero eigenvalues. We shall therefore focus on the nontrivial eigenvalues of . On the global scale, the eigenvalues of are distributed according to the Marchenko-Pastur law (1.6). This may be easily inferred from the fact that (1.6) gives the global density of the eigenvalues for the uncorrelated case , combined with eigenvalue interlacing (see Lemma 4.1 below). In this section we focus on local eigenvalue information.
As explained in Section 1.4, each gives rise to an outlier of near the classical location defined in (1.17). In the definition (2.2), the lower bound is chosen for definiteness; it could be replaced with for any fixed . We denote by
the number of outliers to the left () and right () of the bulk spectrum.
The function will be used to give an upper bound on the magnitude of the fluctuations of an outlier associated with . We give such a precise expression for in order to obtain sharp large deviation bounds for all . (Note that the discontinuity of at is immaterial since is used as an upper bound with respect to . The ratio of the right- and left-sided limits at of lies in $$.)
Our result on the outlier eigenvalues is the following.
Fix . Then for we have the estimate
provided that or .
Furthermore, the extremal non-outliers and satisfy
and, assuming in addition that ,
Theorem 2.3 gives large deviation bounds for the locations of the outliers to the right of the bulk. Since may be arbitrarily small, Theorem 2.3 also gives the full information about the outliers to the left of the bulk except in the case . Although our methods may be extended to this case as well, we exclude it here to avoid extraneous complications.
By definition of and , if then . Hence, by (2.6), if there are no outliers on the left of the bulk spectrum.
Previously, the model from Example (1) in Section 1.2 with fixed nonzero and was investigated in BS ; BY1 . In BS , it was proved that each outlier eigenvalue with convergences almost surely to . Moreover, a central limit theorem for was established in BY1 .
The locations of the non-outlier eigenvalues , , are governed by eigenvalue sticking, whereby the eigenvalues of “stick” with high probability to eigenvalues of a reference matrix which has a trivial population covariance matrix. The reference matrix is from (1.10) with uncorrelated entries. More precisely, we set
Fix . Then we have for all
Similarly, if then we have for all
As outlined above, in Theorem 8.3 below we prove that the asymptotic joint distribution of the non-bulk eigenvalues of is universal, i.e. it coincides with that of the Wishart matrix with and Gaussian. As an immediate corollary of Theorems 2.7 and 8.3, we obtain the universality of the non-outlier eigenvalues of with index . This condition states simply that the right-hand side of (2.9) is much smaller than the scale on which the eigenvalue fluctuates, which is . See Remark 8.7 below for a precise statement.
Theorem 2.7 is analogous to Theorem 2.7 of KY2 , where sticking was first established for Wigner matrices. Previously, eigenvalue sticking was established for a certain class of random perturbations of Wigner matrices in BGGM1 ; BGGM2 . We refer to (KY2, , Remark 2.8) for a more detailed discussion.
Aside from holding for general covariance matrices of the form (1.10), Theorem 2.7 is stronger than its counterpart from KY2 because it holds much further into the bulk: in (KY2, , Theorem 2.7), sticking was established under the assumption that .
The edge universality following from Theorem 2.7 (as explained in Remark 2.8) generalizes the recent result BaoPanZhou2 . There, for the model from Example (1) in Section 1.2 with fixed nonzero and , it was proved that if for all and is diagonal, then converges (after a suitable affine transformation) in distribution to the Tracy-Widom-1 distribution.
2. Outlier eigenvectors
For we define through
For definiteness, we only state our results for the outliers on the right-hand side of the bulk spectrum. Analogous results hold for the outliers on the left-hand side. Since the behaviour of the fluctuating error term is different in the regimes (near the bulk) and (far from the bulk), we split these two cases into separate theorems.
Fix . Suppose that satisfies for all . Define the deterministic positive quadratic form
We emphasize that the set in Theorem 2.11 may be chosen at will. If all outliers are well-separated, then the choice gives the most precise information. However, as explained at the beginning of this subsection, the indices of outliers that are close to each other should be included in the same set . Thus, the freedom to chose is meant for degenerate or almost degenerate outliers. (In fact, as explained after (2.14) below, the correct notion of closeness of outliers is that of overlapping.)
This gives a precise version of the cone concentration from (1.18). Note that the cone concentration holds provided the error is much smaller than the main term , which leads to the conditions
here we used that and .
We claim that both conditions in (2.14) are natural and necessary. The first condition of (2.14) simply means that is an outlier. The second condition of (2.14) is a non-overlapping condition. To understand it, recall from (2.4) that fluctuates on the scale . Then is a non-overlapping outlier if all other outliers are located with high probability at a distance greater than this scale from . Recalling the definition of the classical location of , the non-overlapping condition becomes
After a simple estimate using the definition of , we find that this is precisely the second condition of (2.14). The degeneracy or almost degeneracy of outliers discussed at the beginning of this subsection is hence to be interpreted more precisely in terms of overlapping of outliers.
Provided is well-separated from both the bulk spectrum and the other outliers, we find that the error in (2.13) is of order .
As approaches the delocalization bound from (2.16) deteriorates, and eventually when and start overlapping, i.e. the second condition of (2.14) is violated, the right-hand side of (2.16) has the same size as the leading term of (2.13). This is again a manifestation of the fact that the individual eigenspaces of overlapping outliers cannot be distinguished.
Suppose that we have an -fold degenerate outlier, i.e. for all . Then from Theorem 2.11 and Remark 2.12 (see the estimate (5.2)) we get, for all ,
This is the correct generalization of (2.13) from Example 2.13 to the degenerate case. The error term is the same as in (2.13), and its size and relation to the main term is exactly the same as in Example 2.13. Hence the discussion following (2.13) may be take over verbatim to this case.
We conclude this example by remarking that a similar discussion also holds for a group of outliers that is not degenerate, but nearly degenerate, i.e. for all and . We omit the details.
The next result is the analogue of Theorem 2.11 for outliers far from the bulk.
We leave the discussion on the interpretation of the error in (2.17) to the reader; it is similar to that of Examples 2.13, 2.14, and 2.15.
3. Non-outlier eigenvectors
This quantity should be interpreted as a deterministic version of for ; see Theorem 3.5 below.
Suppose that ( is near the edge), which implies that . Suppose moreover that is near the transition point . Then we get
An analogous statement holds near the left spectral edge provided ; we omit the details.
More generally, our method also yields the asymptotic joint distribution of the family
(after a suitable affine rescaling of the variables, as in Theorem 8.3 below), where . We omit the precise statement, which is a universality result: it says essentially that the asymptotic distribution of (2.22) coincides with that under the standard Wishart ensemble (i.e. an uncorrelated Gaussian sample covariance matrix). The proof is a simple corollary of Theorem 2.7, Proposition 6.2, Proposition 6.3, and Theorem 8.3.
Finally, instead of defined in (1.10), we may also consider
Preliminaries
The rest of this paper is devoted to the proofs of the results from Sections 2.1–2.3. To clarify the presentation of the main ideas of the proofs, we shall first assume that
We make the assumption (3.1) throughout Sections 3–7. The additional arguments required to relax the assumption (3.1) are presented in Section 8. Under the assumption (3.1) we have
Moreover, the extension of our results from to , and hence the proof of Theorem 2.23, is given in Section 9.
In this section we collect the key tool of our analysis: the isotropic Marchenko-Pastur law from BEKYY .
It is well known that the empirical distribution of the eigenvalues of the matrix has the same asymptotics as the Marchenko-Pastur law
where we recall the edges of the limiting spectrum defined in (1.7). Similarly, as noted in (1.6), the empirical distribution of the eigenvalues of the matrix has the same asymptotics as .
Note that (3.3) is normalized so that its integral is equal to one. The Stieltjes transform of the Marchenko-Pastur law (3.3) is
where the square root is chosen so that is holomorphic in the upper half-plane and satisfies as . The function is also characterized as the unique solution of the equation
satisfying for . The formulas (3.3)–(3.5) were originally derived for the case when is independent of (or, more precisely, when has a limit in as ). Our results allow to depend on under the constraint (1.9), so that and may also depend on through .
Throughout the following we use a spectral parameter
with , as the argument of Stieltjes transforms and resolvents. Define the resolvent
Throughout the following we regard the quantities , , and as functions of and usually omit the argument unless it is needed to avoid confusion.
Sometimes we shall need the following notion of high probability.
Fix a (small) and define the domain
Beyond the support of the limiting spectrum, one has stronger control all the way down to the real axis. For fixed (small) define the region
of spectral parameters separated from the asymptotic spectrum by , which may have an arbitrarily small positive imaginary part . Throughout the following we regard as fixed once and for all, and do not track the dependence of constants on .
Suppose that (1.15), (1.9), and (1.16) hold. Then
for all , , and . See (BEKYY, , Remark 2.6).
The next results are on the nontrivial (i.e. nonzero) eigenvalues of as well as the corresponding eigenvectors. The matrix has nontrivial eigenvalues, which we order according to
(The remaining eigenvalues of are zero.) Moreover, we denote by
the unit eigenvectors of associated with the nontrivial eigenvalues .
Fix , and suppose that (1.15), (1.9), and (1.16) hold. Then for we have
if either or .
The following result is on the rigidity of the nontrivial eigenvalues of . Let be the classical eigenvalue locations according to (see (3.3)), defined through
Fix , and suppose that (1.15), (1.9), and (1.16) hold. Then for we have
if or .
2. Link to the semicircle law
Note that this is nothing but Wigner’s semicircle law centred at . Thus,
where in the last step we used (3.4). Note that
The estimates (3.19) and (3.20) follow from the explicit expressions in (3.4) and (3.17). In fact, these estimates have already appeared in previous works. Indeed, for the estimates (3.19) and (3.20) were proved in (BEKYY, , Lemma 3.3). In order to prove them for , we observe that the estimates (3.19) and (3.20) follow from the corresponding ones for the semicircle law, which were proved in (EKYY4, , Lemma 4.3). The estimates (3.21) follow from (3.20) and the elementary identity
which can be derived from (3.18); the estimates for are derived similarly. Finally, (3.22) follows easily from
which may itself be derived from (3.5). ∎
In analogy to (see (3.17)), we define the matrix-valued function
Theorem 3.2 has the following analogue, which compares with .
Suppose that (1.15), (1.9), and (1.16) hold. Then
3. Extension of the spectral domain
In this section we extend the spectral domain on which Theorem 3.2 and Lemma 3.7 hold. The argument relies on the Helffer-Sjöstrand functional calculus Davies . Define the domains
The basic idea of the proof is to apply the Helffer-Sjöstrand formula to the function
where is chosen below. To that end, we need a smooth compactly supported cutoff function on the complex plane satisfying and . We distinguish the three cases , , and .
Let us first focus on the case . Set and choose a constant small enough that . We require that be equal to in the -neighbourhood of and outside of the -neighbourhood of . By Theorem 3.5 we have with high probability. Now choose satisfying . Then the Helffer-Sjöstrand formula Davies yields, for ,
Finally, suppose that . Now we set . We choose the same and cutoff function as in the case above. Suppose that and . Thus, (3.30) holds with high probability for . Since , we therefore find that (3.31) holds. As above, we find that for we have
and . Recalling (3.10), we find that (3.28) follows easily. ∎
Proposition 3.8 yields the following result for defined in (3.24).
4. Identities for the resolvent and eigenvalues
In this section we derive the identities on which our analysis of the eigenvalues and eigenvectors relies. Recall the definition of the set from (1.14). We write the population covariance matrix from (1.12) as
where and . We introduce the matrix
We also denote by the spectrum of a square matrix .
The following lemma collects the basic identities for analysing and . We remark that versions of its part (i) have already appeared in several previous works on finite-rank deformations of random matrix ensembles BGGM1 ; BY1 ; KY2 ; SoshPert .
Suppose that . Then if and only if
To prove (i), we write the condition as
where we used that . Using
the matrix identity , and , we find
with , , , and . ∎
The result (3.35), when restricted to the range of , has an alternative form (3.37) which is often easier to work with, since it collects all of the randomness in the single quantity on its right-hand side.
Eigenvalue locations
In this section we prove Theorems 2.3 and 2.7. The arguments are similar to those of (KY2, , Section 6), and we therefore only sketch the proofs. The proof of (KY2, , Section 6) relies on three main steps: (i) establishing a forbidden region which contains with high probability no eigenvalues of ; (ii) a counting estimate for the special case where does not depend on , which ensures that each connected component of the allowed region (complement of the forbidden region) contains exactly the right number of eigenvalues of ; and (iii) a continuity argument where the counting result of (ii) is extended to arbitrary -dependent using the gaps established in (i) and the continuity of the eigenvalues as functions of the matrix entries. The steps (ii) and (iii) are exactly the same as in KY2 , and will not be repeated here. The step (i) differs slightly from that of KY2 , and in the proofs below we explain these differences.
Let and . For we have
Writing this in spectral decomposition yields
As above, a simple perturbation argument implies that we may without loss of generality assume that all scalar products in (4.1) are nonzero. Now take . Note that and have the same sign.
To conclude the proof, we observe that the left-hand side of (4.1) defines a function of with singularities and zeros, which is smooth and decreasing away from the singularities. Moreover, its zeros are the eigenvalues . The interlacing property now follows from the fact that is an eigenvalue of if and only if the left-hand side of (4.1) is equal to . ∎
For the rank- model (1.10) we have
with the convention that for and for .
Throughout the following we shall make use of the subsets of outliers
for . Note that .
The proof of Proposition 2.3 is similar to that of (KY2, , Equation (2.20)). We focus first on the outliers to the right of the bulk spectrum. Let . We shall prove that there exists an event of high probability (see Definition 3.1) such that for all we have
and for we have
Before proving (4.3) and (4.4), we show how they imply (2.4) for and (2.5). From (4.4) we get for satisfying
Since was arbitrary, (2.4) for and (2.5) follow from (4.3) and (4.5).
What remains is the proof of (4.3) and (4.4). As in (KY2, , Proposition 6.5), the first step is to prove that with high probability there are no eigenvalues outside a neighbourhood of the classical outlier locations . To that end, we define for each the interval
Moreover, we set .
We now claim that with high probability the complement of the set contains no eigenvalues of . Indeed, from Theorem 3.5 and Corollary 3.9 combined with Remark 3.3 (with small enough ), we find that there exists an event of high probability such that for and
for all , where we defined
is singular. Since for , we conclude from the definition of that it suffices to show that if then
We prove (4.6) using the two following observations. First, is monotone increasing on and
We omit further details, which may be found e.g. in (KY2, , Section 6). Thus we conclude that on the event the complement of contains no eigenvalues of .
The next step of the proof consists in making sure that the allowed neighbourhoods contain exactly the right number of outliers; the counting argument (sketched in the steps (ii) and (iii) at the beginning of this section) follows that of (KY2, , Section 6). First we consider the case where for all we have and , and show that each interval contains exactly one eigenvalue of (see (KY2, , Proposition 6.6)). We then deduce the general case by a continuity argument, by choosing an appropriate continuous path joining the initial configuration to the desired final configuration . The continuity argument requires the existence of a gap in the set to the left of . The existence of such a gap follows easily from the definition of and the fact that is bounded. The details are the same as in (KY2, , Section 6.5). Hence (4.3) follows. Moreover, (4.4) follows from the same argument combined with Corollary 4.2 for a lower bound on . This concludes the analysis of the outliers to the right of the bulk spectrum.
The case of outliers to the left of the bulk spectrum is analogous. Here we assume that . The argument is exactly the same as for , except that we use the bound (3.32) to the left of the bulk spectrum as well as for with high probability. ∎
We only give the proof of (2.9); the proof of (2.10) is analogous. Fix . By Theorem 2.3, Theorem 3.5, Theorem 3.2, Lemma 3.7, and Remark 3.3, there exists a high-probability event satisfying the following conditions.
For the following we fix a realization . We suppose first that
and define . Now suppose that satisfies
We shall show, using (3.34), that any satisfying (4.11) cannot be an eigenvalue of . First we deduce from (4.8) that
The estimate (4.12) follows by spectral decomposition of together with the estimate for all . We get from (4.12) and Lemma 3.6 that
where we use the notation to mean . Recalling (3.34), we conclude that on the event the value is not an eigenvalue of provided
It is easy to check that this condition is satisfied if
where we used (4.10). Recalling (4.7), we therefore conclude that for the set
The next step of the proof is a counting argument (sketched in the steps (ii) and (iii) at the beginning of this section), which uses the eigenvalue interlacing from Lemma 4.1. They details are the same as in (KY2, , Section 6), and hence omitted here. The counting argument implies that for and assuming (4.10) we have
What remains is to check (4.13) for the cases and .
Suppose first that . Then using the rigidity from (4.7) and interlacing from Corollary 4.2 we find
where we used the trivial bound . Similarly, if satisfies , we may repeat the same estimate.
We conclude that (4.13) under the sole assumption that . Since was arbitrary, (2.9) follows. ∎
Outlier eigenvectors
In this section we focus on the outlier eigenvectors , . Here we in fact prove Theorem 2.11 under the stronger assumption
instead of . How to improve the lower bound from to the claimed requires a completely different approach, relying on eigenvector delocalization bounds, and is presented in Section 6 in conjunction with results for the non-outlier eigenvectors , .
The proof of Theorem 2.16 is similar to that of Theorem 2.11; one has to adapt the proof to cover the range instead of . The key input is the extension of the spectral domain from Corollary 3.9. For the sake of brevity we omit the details of the proof of Theorem 2.16, and focus solely on Theorem 2.11.
The following proposition is the main result of this section.
Fix . Suppose that satisfies (5.1). Then for all we have
where the symbol denotes the preceding terms with and interchanged.
Note that, under the assumption (5.1), Theorem 2.11 is an easy consequence of Proposition 5.1. As explained above, the proof of Theorem 2.11 in full generality is given in Section 6, where we give the additional argument required to relax (5.1).
The rest of this section is devoted to the proof of Proposition 5.1.
We first prove a slightly stronger version of (5.2) under the additional non-overlapping condition
for all , where is a positive constant. This is a precise version of the second condition of (2.14), whose interpretation was given below (2.14): an outlier indexed by cannot overlap with an outlier indexed by . Note, however, that there is no restriction on the outliers indexed by overlapping among themselves. The assumption (5.3) will be removed in Section 5.2. The main estimate for non-overlapping outliers is the following.
Fix and . Suppose that satisfies (5.1) and (5.3) for all . Then for all we have
The rest of this subsection is devoted to the proof of Proposition 5.2. We begin by defining and letting be a positive constant to be determined later. We choose a high-probability event (see Definition 3.1) satisfying the following conditions.
for , large enough , and all in the set
For all satisfying we have
Note that such an event exists. Indeed, (5.7) and (5.8) may be satisfied using Theorem 2.3, and (5.5) using Theorem 3.7 combined with Remark 3.3.
For the sequel we fix a realization satisfying the conditions (i)–(iii) above. Hence, the rest of the proof of Proposition 5.2 is entirely deterministic, and the randomness only enters in ensuring that has high probability. Our starting point is a contour integral representation of the projection . In order to construct the contour, we define for each the radius
We define the contour as the boundary of the union of discs , where is the open disc of radius around . We shall sometimes need the decomposition , where . See Figure 5.1 for an illustration of .
We shall have to use the estimate (5.5) on the set . Its applicability is an immediate consequence of the following lemma.
The set lies in (5.6).
It is easy to check that for all . In order to check the lower bound on , we note that for any there exists a constant such that
for , , and . Now the claim follows easily from for all , by choosing . ∎
Each outlier lies in , and all other eigenvalues of lie in the complement of .
It suffices to prove that (a) for each we have and (b) all the other eigenvalues satisfy for all .
for , as follows from (5.3) and (5.1). Using
it is then not hard to get (a) from (5.10) and (5.7).
In order to prove (b), we consider the two cases (i) with , and (ii) and . In the case (i), the claim (b) follows using (5.7), (5.11), and (5.3). In the case (ii), the claim (b) follows from (5.8) and the estimate
Using the spectral decomposition of , Lemma 5.5, and the residue theorem, we may write the projection as
This is the desired integral representation of .
We now perform a resolvent expansion on the denominator
where we used Cauchy’s theorem, (4.2), and the fact that lies in if and only if .
using the fact that is holomorphic inside and satisfies the bounds
The first bound of (5.21) follows from (5.5), (5.11), and (5.12). The second bound of (5.21) follows by plugging the first one into
where the contour is the circle of radius centred at . (By assumptions on and , the function is holomorphic in a neighbourhood of the closed interior of .)
In order to estimate (5.20), we consider the three cases (i) , (ii) , , (iii) , . Note that (5.20) vanishes if . We start with the case (i). Suppose first that and . Then we find
A simple limiting argument shows that this bound is also valid for and . Next, in the case (ii) we get from (5.21)
A similar estimate holds for the case (iii). Putting all three cases together, we find
What remains is the estimate of . Here residue calculations are unavailable, and the precise choice of the contour is crucial. We use the following basic estimate to control the integral.
For , , and we have
The upper bound is trivial, so that we only focus on the lower bound. Suppose first that . Then we get , from which the claim follows since by (5.9).
For the remainder of the proof we may therefore suppose that . Define , the distance between the discs and (see Figure 5.1). We consider the two cases and separately.
Suppose first that . Then by definition of we have . Now a simple estimate using the definition of yields , from which we conclude . The claim now follows from the bound .
Suppose now that . Hence , so that in particular . Thus we get
From (5.18), (5.5), (5.11), and (5.12) we get
where we also used the estimate for .
In order to estimate the matrix norm, we observe that for we have on the one hand
for any , where in the last step we used (5.10). Since , these estimates combined with a resolvent expansion give the bound
for . Decomposing the integration contour in (5.23) as , and recalling that has length bounded by , we get from Lemma 5.6
We estimate the right-hand side using Cauchy-Schwarz. For we find, using (5.9),
For we use (5.9) and the estimate for all to get
Recall that . Hence, plugging (5.19), (5.22), and (5.27) into (5.15), we find
We have proved (5.28) under the assumption that . The general case is an easy corollary. For general , we define and consider
where for and for . Since and is invertible, we may apply the result (5.28) to this modified model. Now taking the limit for in (5.28) concludes the proof in the general case. Now Proposition 5.2 follows since may be chosen arbitrarily small. This concludes the proof of Proposition 5.2.
2. Removing the non-overlapping assumption
In this subsection we complete the proof of Proposition 5.1 by extending Proposition 5.2 to the case where (5.3) does not hold.
Let . We say that overlap if or . For we introduce sets satisfying . Informally, is the largest subset of indices of that do not overlap with its complement. It is by definition constructed by successively choosing , such that overlaps with an index of , and removing from ; this process is repeated until no such exists. One can check that the result is independent of the choice of at each step. Note that may be empty.
Informally, is the smallest subset of indices in that do not overlap with its complement. It is by definition constructed by successively choosing , such that overlaps with an index of , and adding to ; this process is repeated until no such exists. One can check that the result is independent of the choice of at each step. See Figure 5.2 for an illustration of and . Throughout the following we shall repeatedly make use of the fact that, for any , Proposition 5.2 is applicable with replaced by or .
After these preparations, we move on to the proof of (5.2). We divide the argument into four steps.
We consider two cases, and . Suppose first that . Using that is bounded, it is not hard to see that . We now invoke Proposition 5.2 and get
In the complementary case, , a simple argument yields
as well as . From Proposition 5.2 we therefore get
where we used that . Recalling (5.29), we conclude
(b) i=j∈A𝑖𝑗𝐴i=j\in A
We consider the two cases and . Suppose first that . We write
We compute the first term of (5.32) using Proposition 5.2 and the observation that :
In order to estimate the second term of (5.32), we note that . We therefore apply (5.31) with replaced by to get
Going back to (5.32), we have therefore proved that
Next, we consider the case . Now we have (5.30), so that Proposition 5.2 yields
By (5.30) and , we have
from which we deduce (5.33) also in the case .
(c) i≠j𝑖𝑗i\neq j and i∉A𝑖𝐴i\notin A or j∉A𝑗𝐴j\notin A
From cases (a) and (b) (i.e. (5.31) and (5.33)), combined with the estimate
we find, assuming or , that (5.2) holds with an additional factor multiplying the right-hand side.
(d) i≠j𝑖𝑗i\neq j and i,j∈A𝑖𝑗𝐴i,j\in A
We now deal with the last remaining case by using the splitting
Note that here . We consider the four cases (i) , (ii) and , (iii) and , and (iv) .
Next, consider the case (ii). For the first term of (5.34) we use the estimates
In order to estimate the last term, we first assume that and . Then we find
Conversely, if or , we have . Therefore, using (5.36) and the estimate , we get
Putting (5.37), (5.38), and (5.39) together, we may estimate the first term of (5.34) in the case (ii) as
For the second term of (5.34) in the case (ii) we use the estimates
where the last term is bounded by . Recalling (5.40), we find (5.35) in the case (ii). The case (iii) is dealt with in the same way.
What remains therefore is case (iv). For the first term of (5.34) we use the estimates
For the second term of (5.34) we use the estimates
which is (5.35). This concludes the analysis of case (iv), and hence of case (d).
Conclusion of the proof
Putting the cases (a)–(d) together, we have proved that (5.2) holds for arbitrary with an additional factor multiplying the error term on the right-hand side. Since can be chosen arbitrarily small, (5.2) follows. This concludes the proof of Proposition 5.1. ∎
Non-outlier eigenvectors
We first consider eigenvectors near the right edge of the bulk spectrum. Recall the typical distance from to the spectral edges, denoted by and defined in (2.18).
Fix . For we have
Moreover, if satisfies then
Proposition 6.1 has a close analogue for the left edge of the bulk spectrum, which holds under the additional condition ; we omit its detailed statement.
Suppose first that . Let and set . Using (3.25), Remark 3.3, Theorem 2.3, and (3.15), we choose a high-probability event satisfying (4.7), (4.8), and
Abbreviating , we find from (3.20) that
which follows easily by spectral decomposition. Since , we get from (3.37), omitting the arguments for brevity,
where the last step follows from a resolvent expansion as in (5.14). We estimate the error terms using
where we used the definition of and (6.4). Hence a resolvent expansion yields
Next, we claim that for any fixed we have the lower bound
whenever . To prove (6.10), suppose first that . By (3.21), there exists a constant such that for we have . Thus we get, for ,
where we used that by (3.19). Moreover, if we find from (3.20) that , from which we get
where in the second step we used as follows from (3.19). This concludes the proof of (6.10) for the case .
Suppose now that . Then we get
We shall estimate this using the elementary bound
For we get from (6.11) with , recalling (3.20) and (3.21), that . By a similar argument, for we set and get (6.10) using (6.5) and (6.6). This concludes the proof of (6.10).
Using and (6.10), we estimate the first term on the right-hand side of (6.12) as
where in the last step we used that , as follows from (6.5) and (6.6).
Next, we estimate the second term of (6.12) as
Putting all three estimates together, we conclude that
in which case we get by choosing in (6.10) that
Next, if satisfies we get from (6.5), (6.6), and (3.20) that
In this case we have by (6.3), so that setting in (6.10) yields
Since was arbitrary, (6.1) and (6.2) follow from (6.14) and (6.15) respectively. This concludes the proof of Proposition 6.1 in the case .
Finally, the case is handled by replacing with and using a limiting argument, exactly as after (5.28). ∎
2. Proof of Theorems 2.11 and 2.17
We now have all the ingredients needed to prove Theorems 2.11 and 2.17.
The estimate (2.19) is an immediate corollary of (6.1) from Proposition 6.1. The estimate (2.20) is proved similarly (see also the remark following Proposition 6.1). ∎
We prove Theorem 2.11 using Propositions 5.1, 5.2, and 6.1. First we remark that it suffices to prove that (5.2) holds for satisfying for all . Indeed, supposing this is done, we get the estimate
from which Theorem 2.11 follows by noting that the second error term may be absorbed into the first, recalling that for , that , and that .
Fix . Note that there exists some satisfying the following gap condition: for all such that we have . The idea of the proof is to split , such that for and for . Note that such a splitting exists by the above gap property. Without loss of generality, we assume that (for otherwise the claim follows from Proposition 5.1).
It suffices to consider the six cases (a) , (b) and , (c) and , (d) , (e) and , (f) .
We apply Cauchy-Schwarz and Proposition 6.1 to the first term, and Proposition 5.1 to the second term. Using the above gap condition, we find
where the last step follows from .
For this case it is crucial to use the stronger bound (5.4) and not (5.2). Hence, we need the non-overlapping condition (5.3). To that end, we assume first that (5.3) holds with . Thus, by the above gap assumption (5.3) also holds for . In this case we get from (6.16) and Propositions 5.2 and 6.1 that
Clearly, the first two terms are bounded by the right-hand side of (5.2) times . The last term is estimated as
where we used that be the above gap condition. This concludes the proof in the case where the non-overlapping condition (5.3) holds.
(c), (e), (f) j∉A𝑗𝐴j\notin A
We use the splitting (6.16) and apply Cauchy-Schwarz and Proposition 6.1 to the first term, and Proposition 5.1 to the second term. Since in all cases, it is easy to prove that (6.16) is bounded by times the right-hand side of (5.2).
From (6.16) and Propositions 6.1 and 5.1 we get
From which we get (5.2) with the error term multiplied by .
Conclusion of the proof
We have proved that, for all and satisfying the assumptions of Theorem 2.11, the estimate (5.2) holds with an additional factor multiplying the error term. Since was arbitrary, we get (5.2). This concludes the proof. ∎
3. The law of the non-outlier eigenvectors
the typical distance between and . More precisely, the classical locations defined in (3.14) satisfy for .
Let and define . Define the event
Informally, Proposition 6.2 expresses generalized components of the eigenvectors of in terms of generalized components of eigenvectors of , under the assumption that has high probability. We first show how Proposition 6.2 implies Theorem 2.20. This argument requires two key tools. The first one is level repulsion, which, together with the eigenvalue sticking from Theorem 2.7, will imply that indeed has high probability. The second tool is quantum unique ergodicity (See Section 1.1) of the eigenvectors of , which establishes the law of the generalized components of the eigenvectors of .
The precise statement of level repulsion sufficient for our needs is as follows.
Fix . For any there exists a such that for all we have
The proof of Proposition 6.3 consists of two steps: (i) establishing (6.19) for the case of Gaussian and (ii) a comparison argument showing that if and are two matrix ensembles satisfying (1.15) and (1.16), and if (6.19) holds for , then (6.19) also holds for . Both steps have already appeared, in a somewhat different form, in the literature. Step (i) is performed in Lemma 6.4 below, and step (ii) in Lemma 6.5 below. Together, Lemmas 6.4 and 6.5 immediately yield Proposition 6.3.
Proposition 6.3 holds if is Gaussian.
We mimic the proof of Theorem 3.2 in BEY3 . Indeed, the proof from (BEY3, , Appendix D) carries over almost verbatim. The key input is the eigenvalue rigidity from Theorem 3.5, which for the model of BEY3 was established using a different method than Theorem 3.5. As in BEY3 , we condition on the eigenvalues . On the conditioned measure, level repulsion follows as in BEY3 . Finally, thanks to Theorem 3.5 we know that the frozen eigenvalues are with high probability near their classical locations. Note that for , the rigidity estimate (3.15) only holds for indices ; however, this is enough for the argument of (BEY3, , Appendix D), which is insensitive to the locations of eigenvalues at a distance of order one from the right edge . We omit the full details. ∎
Let and be two matrix ensembles satisfying (1.15) and (1.16). Suppose that Proposition 6.3 holds for . Then Proposition 6.3 also holds for .
The proof of Lemma 6.5 relies on Green function comparison, and is given in Section 7.4.
For simplicity, and bearing the application to Theorem 2.20 in mind, in Proposition 6.6 we establish the convergence of a single generalized component of a single eigenvector. However, our method may be easily extended to yield
The proof of this generalization of Proposition 6.6 follows that of Proposition 6.6 presented in Section 7, requiring only heavier notation. In fact, our method may also be used to prove the universality of the joint eigenvalue-eigenvector distribution for any matrix of the form (1.10) with ; see Theorem 8.3 below for a precise statement.
The proof of Proposition 6.6 is postponed to Section 7.
Supposing Proposition 6.2 holds, together with Propositions 6.3 and 6.6, we may complete the proof of Theorem 2.20.
Abbreviating and
Then, by assumption on , we may rewrite (6.18) as
The remainder of this section is devoted to the proof of Proposition 6.2.
We define the contour as the positively oriented circle of radius with centre . Let and , and choose a high-probability event such that (4.7), (4.8), and (4.9) hold. For the following we fix a realization . Define
By the residue theorem and the definition of , we find
In order to compute (6.22), we need precise estimates for on . Because the contour crosses the branch cut of , we should not compare to for . Instead, we compare to , where
for all . To see this, we split
We estimate the first term of (6.24) by spectral decomposition, using that , similarly to (4.12). The result is
where we used (4.7), (4.9), and Lemma 3.6. Moreover, we estimate the second term of (6.24) using (4.8) as
The proof of (6.25) is analogous to that of (6.10), using (4.7) and the assumption on ; we omit the details.
Armed with (6.23) and (6.25), we may analyse (6.22). A resolvent expansion in the matrix yields
We estimate the third term using the bound
To prove (6.27), we note first that by (6.25) we have
By (6.23) and assumption on , it is easy to check that
We may now return to (6.26). The first term vanishes, the second is computed by spectral decomposition of , and the third is estimated using (6.27). This gives
Recalling (6.21) and (4.7), we therefore get
where we used .
In order to simplify the leading term, we use
where we used Lemma 3.6. Moreover, we use that
Using that has high probability for all and recalling the isotropic delocalization bound (3.13), we therefore get for any random that
We proved (6.28) under the assumption that , but a continuity argument analogous to that given after (5.28) implies that (6.28) holds for all . The above argument may be repeated verbatim to yield
Quantum unique ergodicity near the soft edge of H𝐻H
This section is devoted to the proof of Proposition 6.6.
Fix . Let be a smooth function satisfying
for some positive constant . Let and suppose that satisfies (6.19) with some constants and . Then for small enough and the following holds. Defining
By the assumption (7.1) on , rigidity (3.15), and delocalization (3.13), we can write
where we defined . For the following we choose
In order to obtain (7.3), we have to rewrite the integrand on the right-hand side of (7.5) in terms of
Hence (7.5) and (7.6) combined with the mean value theorem imply that the left-hand side of (7.3) is bounded by
for any fixed . When applying the mean value theorem, we estimated the value of using (7.1), the fact that all terms on the right-hand side of (7.6) are nonnegative, and the estimate
Next, using the eigenvalue rigidity from (3.15), it is not hard to see that there exists a constant such that the contribution of to (7.7) is bounded by . In order to prove (7.3), therefore, it suffices to prove
which is the right-hand side of (7.9) provided is chosen small enough. Here in the first step we replaced with using the estimates valid for and in the support of .
For , we partition with and
Let us therefore consider the integral over . One readily finds, for , that
Using delocalization (3.13) we therefore find that
In the next step, stated in Lemma 7.2 below, we replace the sharp cutoff function in (7.3) with a smooth function of . Note first that from Lemma 7.1 and the rigidity (3.15), we get
The following result is the appropriate smoothed version of (7.11). It is a simple extension of Lemma 3.2 and Equation (5.8) from KY1 , and its proof is omitted.
We may now conclude the proof of Proposition 6.6.
For the comparison argument, we use the Green function comparison method applied to the Helffer-Sjöstrand representation of . Using Lemma 7.2 it suffices to estimate
Thus we get the functional calculus, with ,
As in Lemma 5.1 of KY1 , one can easily extend (3.9) to satisfying the lower bound instead of in (3.7); the proof is identical to that of (KY1, , Lemma 5.1). Thus we have, for and ,
Therefore, by the trivial symmetry combined with complex conjugation, the third term on the right-hand side of (7) is bounded by
Recalling (7.1) and using the mean value theorem, we find from (7.13), (7.17), and (7.18) that for large enough , in order to estimate (7.14), and hence prove (6.20), it suffices to prove the following lemma. Note that in it we choose to be the original ensemble and to be the Gaussian ensemble. ∎
Suppose that the two matrix ensembles and satisfy (1.15) and (1.16). Suppose that the assumptions of Lemma 7.1 hold, and recall the notations
Then for any and for small enough and we have
The rest of this section is devoted to the proof of Lemma 7.3.
We shall use the Green function comparison method EYY1 ; EYY3 ; KY1 to prove Lemma 7.3. For definiteness, we assume throughout the remainder of Section 7 that . The case is dealt with similarly, and we omit the details.
We first collect some basic identities and estimates that serve as a starting point for the Green function comparison argument. We work on the product probability space of the ensembles and . We fix a bijective ordering map on the index set of the matrix entries,
and define the interpolating matrix , , through
In particular, and . Hence we have the telescopic sum
Let us now fix a and let be determined by . Throughout the following we consider to be arbitrary but fixed and often omit dependence on them from the notation. Our strategy is to compare with for each . In the end we shall sum up the differences in the telescopic sum (7.24).
Note that and differ only in the matrix entry indexed by . Thus we may write
Here is the matrix obtained from (or, equivalently, from ) by setting the entry indexed by to zero. Next, we define the resolvents
where is polynomial of degree two in whose coefficients are -measurable.
The rest of this section is therefore devoted to the proof of (7.27). Recall that we assume throughout that for definiteness; in particular, .
We begin by collecting some basic identities from linear algebra. In addition to we introduce the auxiliary resolvent . Moreover, for we split
We also define the resolvent . A simple Neumann series yields the identity
Moreover, from (BEKYY, , Equation (3.11)), we find
Throughout the following we shall make use of the fundamental error parameter
which is analogous to the right-hand side of (3.9) and will play a similar role. We record the following estimate, which is analogous to Theorem 3.2.
This result is a generalization of (5.22) in PY . The key identity is (7.30). Since is independent of , we may apply the large deviation estimate (BEKYY, , Lemma 3.1) to . Moreover, , as follows from Theorem 3.2 applied to , and Lemma 3.6. Thus we get
where the second step follows by spectral decomposition, the third step from Theorem 3.2 applied to as well as (3.22), and the last step by definition of . This concludes the proof of (7.32).
Finally, (7.33) follows easily from Theorem 3.2 applied to the identity . ∎
Using (7.35), we may extend these estimates to analogous ones on and instead of and . Indeed, using the facts , , and (which are easily derived from the definitions of the objects on the left-hand sides) combined with (7.35), we get the following result.
For and we have
The final tool that we shall need is the following lemma, which collects basic algebraic properties of stochastic domination . We shall use them tacitly throughout the following. Their proof is an elementary exercise using union bounds and Cauchy-Schwarz. See (BEKYY, , Lemma 3.2) for a more general statement.
Suppose that uniformly in . If for some constant then .
Suppose that and . Then .
If the above random variables depend on an additional parameter and all hypotheses are uniform in then so are the conclusions.
2. Proof of Lemma 7.3 II: the main expansion
Lemma 7.5 contains the a-priori estimates needed to control the resolvent expansion (7.34). The precise form that we shall need is contained in the following lemma, which is our main expansion. Define the control parameter
The following results hold for . (Recall the definition (7.20). For brevity, we omit from our notation.)
where is a polynomial, with constant number of terms, in the variables
where is a polynomial, with constant number of terms, in the variables
In each term of , appears exactly once, while the indices and each appear exactly times.
Here all constants depend on the fixed parameter .
The proof is an application of the resolvent expansion (7.34) with to the definitions of and .
and the same estimates hold if is replaced by . Note that in (7.42) we used the bound , which follows from the identity
and Lemma 3.6. Using (7.42), it is not hard to conclude the proof of part (i).
What remains is to prove the bounds in part (iii). To that end, we integrate by parts, first in and then in , in the term containing , and obtain
The same argument yields the error bound in (7.40). This concludes the proof. ∎
Next, using lemma 7.7 and , we find
Using (7.44) and (7.46), we may do a Taylor expansion of on the left-hand side of (7.27). This yields
For the general case, we still have to estimate the last line of (7.48). The terms that we need to analyse are
These terms are dealt with in the following lemma.
Let denote any term of (7.49). Then there is a constant such that
3. Proof of Lemma 7.3 III: the terms of order three and proof of Lemma 7.8
Recall that we assume , i.e. ; the case is dealt with analogously, and we omit the details.
We first remark that using the bounds (7.46) we find
Comparing this to (7.50), we see that we need to gain an additional factor . How to do so is the content of this subsection.
Let us outline the rough idea of the parity argument. We use the notations and , in analogy to those introduced before (7.28). A simple example of a polynomial is
The need to gain an additional factor from odd polynomials imposes nontrivial constraints on the polynomial coefficients, which are carefully stated in Definitions 7.10–7.12; they have been tailored to the class of polynomials generated by the terms (7.49).
We now move on to the proof of Lemma 7.8. We recall that we assume throughout that . We first introduce a family of graded polynomials suitable for our purposes. It depends on a constant , which we shall fix during the proof to be some large but fixed number.
Let be a family of deterministic nonnegative weights. We say that is an admissible weight if
be a polynomial in . Analogously to the notation introduced in Definition 2.1, we write if the following conditions are satisfied.
is deterministic and is -measurable.
There exist admissible weights such that
We have the deterministic bound .
The above definition extends trivially to the case , where is -measurable.
Let be a polynomial of the form
We write if is -measurable, , and for some deterministic .
We write if is a sum of at most terms of the form
where and is deterministic.
Moreover, we write if , where and .
Definitions 7.10–7.12 refine Definition 2.1 in the sense that
Indeed, let be of the form (7.53). Then a simple large deviation estimate (e.g. a trivial extension of (EKYY3, , Theorem B.1 (iii))) yields
where the last step follows from the definition of admissible weights. Similarly, if is of the form (7.55), a large deviation estimate (e.g.(EKYY3, , Theorem B.1 (i))) yields
after possibly increasing . (As with the standard big O notation, such expressions are to be read from left to right.) We stress that such operations may be performed an arbitrary, but bounded, number of times. It is a triviality that all of the following arguments will involve at most such algebraic operations on graded polynomials, for large enough .
The point of the graded polynomials is that bounds of the form (7.56) are improved if is odd and we take the expectation. The precise statement is the following.
Let for some deterministic . Then for any fixed we have
It suffices to set and consider , where , , and are as in Definition 7.12. By linearity, it suffices to consider
where is even. We suppose that , , and for . Here denotes an admissible weight (see Definition 7.9). Thus we have
The expectation imposes that each summation index coincide with at least one other summation index. Thus we get
Hence the summation over factors into a product over the blocks of . We shall show that the contribution of each block is at most one, and that there is a block whose contribution is at most .
Fix and denote by the contribution of the block to the summation in the main term of (7.57). Define and . By definition of , we have . By the inequality of arithmetic and geometric means, we have
Moreover, since is even, at least one block of satisfies .
Since , the proof is complete. ∎
We begin by noting that (3.9) applied to and (3.22) combined with a large deviation estimate (see (EKYY3, , Theorem B.1)) yields
Using (7.35) and Lemma 7.5, it is not hard to deduce that
where in the second step we used (3.5) and (3.23). Now we split
where in the second step we used the estimates and . Since , we therefore conclude that
From (7.31) and the definition of , we readily find that for some constant . Therefore choosing large enough yields
Having established (7.64), the remainder of the proof is relatively straightforward. From (7.29) and (7.30) we get
Now (7.62) follows easily from (7.65) and (7.64).
Moreover, (7.59) and (7.61) follow from (7.28) combined with (7.65) and (7.64). For (7.61) we estimate the second term in (7.28) by
where in the last step we used that . Moreover, (7.60) is a trivial consequence of (7.59). Finally, (7.63) follows from (7.28) and (7.64) combined with
In each application of Lemma 7.14, we shall verify one of the conditions of (7.66). The first condition is verified for , which always holds for the coefficients of , , and (recall (7.2)).
The second condition of (7.66) will be verified when computing the coefficients of , , and . To that end, we make use of the freedom of the choice of basis when computing the trace in the definition of , , and . We shall choose a basis that is completely delocalized. The following simple result guarantees the existence of such a basis.
Fix . Then there exists such that
where the parity of follows easily from its definition.
Lemmas 7.14 and 7.17 are the key estimates of the coefficients appearing in (7.49). We claim that all estimates of Lemma 7.7, along with (7.42), remain valid, in the sense that an estimate of the form is to be replaced with
where the parity of may be easily deduced from their definitions. Moreover, for the estimates (7.41) we use (7.72) to get
Note that, thanks to Lemmas 7.14 and 7.17, we have obtained exactly the same upper bounds on the coefficients and as the ones obtained in Lemma 7.7, but we have in addition expressed them, up to a negligible error, as graded polynomials, to which Lemma 7.13 is applicable.
which may be derived from (7.68), combined with a Taylor expansion of . Similarly, we find that
We may now put everything together. Noting that the degree of the polynomializations of the expressions (7.49) is always odd, we obtain, in analogy to (7.51) that
for being any term of (7.49). Hence Lemma 7.8 follows from Lemma 7.13 and Young’s inequality.
4. Stability of level repulsion: proof of Lemma 6.5
This is a Green function comparison argument, using the machinery introduced in Section 7.1. A similar comparison argument was given in Propositions 2.4 and 2.5 of KY1 . The details in the sample covariance case and for indices satisfying follow an argument very similar to (in fact simpler than) the one from Sections 7.1–7.3. As in the proofs of Propositions 2.4 and 2.5 of KY1 , one writes the level repulsion condition in terms of resolvents. In our case, one uses the representation (7) as the starting point. Then the machinery of Sections 7.1–7.3 may be applied with minor modifications. We omit the details.
Extension to general T𝑇T and universality for the uncorrelated case
In this section we relax the assumption (3.1), and hence extend all arguments of Sections 3–7 to cover general . We also prove the fixed-index joint eigenvector-eigenvalue universality of the matrix defined in (2.7), for indices bounded by for some .
Bearing the applications in the current paper in mind, we state the results of this section for the matrix from (2.7), but it is a triviality that all results and their proofs carry over to case of arbitrary from (1.10) provided that .
We start with the singular value decomposition of , which we write as
where and were defined in (2.7). Comparing this to (3.2), we find that to relax the assumption (3.1) we have to generalize the arguments of Sections 3–7 by replacing with .
The generalization of is the resolvent of ,
where we recall the definition (7.31) of . In fact, from Lemma 3.6 and (3.23) we find that for some positive constant depending on . Hence (3.9) yields
Having established Theorem 8.1, all arguments from Sections 3–6 that use it as input may be taken over verbatim, after replacing by . More precisely, all results from Sections 3–6 remain valid for a general , with the exception of Proposition 6.3, Lemmas 6.4 and 6.5, and Proposition 6.6. Therefore we have completed the proofs of all of our main results except Theorem 2.20.
In order to prove Theorem 2.20, we still have to prove Lemmas 6.4 and 6.5 and Proposition 6.6 for instead of . Lemma 6.4 is easy: for Gaussian we have , where is the matrix obtained from by deleting its bottom rows.
The proofs of Lemma 6.5 and Proposition 6.6 rely on Green function comparison. What remains, therefore, is to extend the argument of Section 7 from to .
Lemma 7.3 remains valid if and are replaced with and , obtained from the definitions (7.22) and (7.23) by replacing with .
Recalling (3.9) and (7.31), we find that the second term is stochastically dominated by
where in the second step we used that , as follows from Lemma 3.6 and the definition of in (7.2). Recalling the definitions from (7.2), we therefore conclude that for small enough we have
for some positive constant depending on .
for some positive constant depending on . Plugging (8.5) into the definition of and estimating the error term using integration by parts, as in (7.43), we get
Using the mean value theorem and the bound , we therefore get
Combined with (7.21), this concludes the proof. ∎
This concludes the proof of Theorem 2.20 for the case of general .
In this section we observe that the technology developed in Section 7 allows us to establish the universality of the joint eigenvalue-eigenvector distribution of provided that . Without loss of generality, we consider the case where is given by defined in (2.7). This result applies to arbitrary eigenvalue and eigenvector indices which are bounded by , and does in particular not need to invoke eigenvalue correlation functions.
This result generalizes the quantum unique ergodicity from Proposition 6.6 and its extension from Remark 6.7 by also including the distribution of the eigenvalues. The universality of both the eigenvalues and the eigenvectors is formulated in the sense of fixed indices. A result in a similar spirit was given in (KY1, , Theorem 1.6), except that the upper bound on the eigenvalue and eigenvector indices from KY1 is improved all the way to , for any . A result covering all eigenvalue and eigenvector indices, i.e. with an index upper bound , was given in (KY1, , Theorem 1.10) and (TV3, , Theorem 8), but under the assumption of a four-moment matching assumption. Theorem 8.3 is a true universality result in that it does not require any moment matching assumptions, but it does require an index upper bound of instead of on the eigenvalue and eigenvector indices.
for some constant .
The proof is a Green function comparison argument, a minor modification of that developed in Section 7. We write the distribution of in terms of the resolvent , starting from the Helffer-Sjöstrand representation (7), exactly as in (KY1, , Sections 4 and 5). We omit further details. ∎
We note that even this fixed-index universality of eigenvalues is a new result, having previously only been established under the four-moment matching condition KY1 ; TV3 (in the context of Wigner matrices).
We formulated Theorem 8.3 for the real symmetric covariance matrices of the form (2.7), but it and its proof remain valid for complex Hermitian covariance matrices, as well as Wigner matrices (both real symmetric and complex Hermitian).
Assuming , the condition on the indices in Theorem 8.3 may be replaced with .
Extension to Q˙˙𝑄\dot{Q} and proof of Theorem 2.23
In this section we explain how to extend our analysis from defined in (1.10) to defined in (2.23), hence proving Theorem 2.23. We define the resolvent
which will replace when analysing with instead of . We begin by noting that the isotropic local laws hold for also for .
As in the proof of Theorem 8.1, we only prove (3.9) for . Using the identity (3.36) we get
Using (3.9), the proof will be complete provided we can show that
Lemma 7.3 remains valid if and are replaced with and , obtained from the definitions (7.22) and (7.23) by replacing with .
The proof mirrors closely that of Lemma 8.2, using the identity (9.1) instead of (8.2) as input. We omit the details. ∎
All that remains is the proof of the following estimate, which generalizes (7.32).
here we used the first identity of (7.30). A similar argument was given in (BEKYY, , Section 5). The basic idea is to make all resolvents on the right-hand side of (9.4) independent of the columns of indexed by (see (BEKYY, , Definition 3.7)). As in (BEKYY, , Section 5), we do this using the identities from (BEKYY, , Lemma 3.8) for the entries of . In addition, for the entries of we use the identity (in the notation of (BEKYY, , Definition 3.7))
Appendix A A few remarks on applications to statistics
In this appendix we give a few remarks on what our results imply for applications to statistics. We assume throughout that the population covariance matrix satisfies
We consider the following simple model problem. Suppose there is some (unknown) set whose associated variables are strongly correlated. For simplicity, let us assume that the correlations are given by a single spike in , i.e.
The most naive way to proceed is to compare the entries of with those of . Using (A.1) it is not hard to conclude that
We look at the off-diagonal terms of and infer that belongs to if there exists an index such that is much larger than . For this approach to work, we require that , which reads . We conclude that this naive entrywise approach works provided that
The principal component analysis works and the naive componentwise approach does not if (A.4) holds and (A.3) does not. These conditions may be written as
Hence the principal component analysis for this example is very effective when the family of correlated variables is quite large, .
More generally, the principal component approach may work in the regime
and cannot work for smaller . Indeed, the assumption (A.1) is satisfied for , so that in the case of the strongest possible correlations, , the condition (A.4) reduces to (A.5). On the other hand, if , then from (A.2) and the assumption (A.1) we find , in contradiction to (A.4).
Clearly, by (A.4), for the purposes of statistical inference it is desirable to make as small as possible. This means that the number of samples per variables is as large as possible. It is therefore natural to attempt to make smaller so as to reduce . Obviously, if we know a priori that is contained in some subset of of size , then we simply discard all variables indexed by and consider the correlations restricted to ; we have halved in the process.
However, if we have no such a priori knowledge about , discarding half of the variables is a bad idea. In this case, the best one can do is to choose at random. Thus, suppose that is uniformly distributed among the subsets of of size . We cut the sample space in half by keeping only the first elements. We therefore obtain a new family of variables with dimensional parameters
with high probability, we find that for . Picking an entry for , we obtain
We conclude that detecting spikes in the new problem is more difficult than in the original problem, and the halving of sample space is therefore counterproductive unless one has some good a priori information about .
Using for , we therefore get the condition . This however cannot hold, since we have by assumption (A.1) on that