Central limit theorems for eigenvalues of deformations of Wigner matrices
Mireille Capitaine, Catherine Donati-Martin, Delphine Féral
Introduction
is a deterministic Hermitian matrix of fixed finite rank and whose spectrum does not depend on . The matrix is a Wigner Hermitian matrix such that the random variables , , are independent identically distributed with a centered distribution of variance .
As the rank of the ’s is assumed to be finite, the Wigner Theorem is still satisfied for the Deformed Wigner model (cf. Lemma 2.2 of [B]): the spectral measure of converges a.s. towards the semicircle law whose density is given by
When , it is well-known that once has a finite fourth moment, the first largest (resp. last smallest) eigenvalues of the rescaled Wigner matrix tend almost surely to the right (resp. left)-endpoint (resp. ) of the semicircle support (cf. [B]). The corresponding fluctuations, which have been first obtained by Tracy and Widom [T-W] in the Gaussian case and then extended by Soshnikov [So] for any symmetric probability measure having subgaussian moments, are governed by the so-called Tracy-Widom distributions. Note that the exponential decay condition (with symmetry assumption) has been replaced by a finite number of moments in [R], [K]. Under the subexponential decay assumption, the symmetry assumption on in [So] was replaced in [T-V] by the vanishing third moment condition and very recently, Erdös, Yau and Yin [E-Y-Y] proved the edge universality under the subexponential decay assumption alone. Let us describe how the asymptotic behavior of the extremal eigenvalues of the perturbed Wigner matrix may be affected by the perturbation by considering the particular case of a rank one perturbation with non-null eigenvalue . For a large class of probability measures , it turns out that the largest eigenvalue still tends to the right-endpoint if whereas jumps above the bulk to if . This was proved by Péché in her pionnering work [Pe] when is gaussian, extended in [Fe-Pe1] when is symmetric and has subgaussian moments but in the particular case of the full rank one deformation given by
and finally established in [C-D-F] for general when is symmetric and satisfies a Poincaré inequality.
Moreover, considering the perturbation matrix defined by (1.3), Féral and Péché [Fe-Pe1] proved that the fluctuations of are the same as in the gaussian setting investigated in [Pe] and in this sense are universal. Here is their result when :
If is symmetric and has subgaussian moments
where .
The proof of this result relies on the computations of moments of of high order (depending on ) and the knowledge of the fluctuations in the Gaussian case, established by Péché [Pe].
On the other hand, for the strongly localized perturbation matrix of rank 1 given by
with , we proved in [C-D-F] that the fluctuations of vary with the particular distribution of the entries of the Wigner matrix so that this phenomenon can be seen as an example of a non universal behavior :
Let be symmetric and satisfy a Poincaré inequality. Define
In the present paper, we consider perturbations of higher rank of Wigner matrices associated to some symmetric probability measure satisfying a Poincaré inequality. The a.s. convergence of the extreme eigenvalues has already been described in [C-D-F] (see Theorem 3.1 below). Whenever the largest eigenvalues of are extracted away from the bulk, we describe their fluctuations which depend on the localization of the eigenvectors of , as already seen in the above rank 1 examples. We investigate two quite general situations for which we exhibit a phenomenon of different nature. To explain this, let us focus on the largest eigenvalue of . We assume that so that the largest eigenvalues of converges a.s towards . First, when the eigenvectors associated to the largest eigenvalue of are localized, we establish that the limiting distribution in the fluctuations of , around is not universal and we give it explicitely in terms of these eigenvectors and of the distribution of the entries of the Wigner matrix, see Theorem 3.2. Secondly, if the eigenvectors are sufficiently delocalized, we establish the universality of the fluctuations of , see Theorem 3.3 . Actually, in the rank one case, this study allows us to exibit a necessary and sufficient condition on a normalized eigenvector of associated to the largest eigenvalue for the universality of the fluctuations (see Theorem 3.4 below). Moreover if such an eigenvector of is not localized but does not satisfy the criteria of universality, the largest eigenvalue of may fluctuate according to a mixture of and normal distributions generalizing (1.5). We will describe some of such intermediate situations. We detail the definition of localization/delocalization and these results in the Section 3.
The Deformed Wigner matrix model may be seen as the additive analogue of the spiked population models. These are random sample covariance matrices defined by where is a complex (resp. real) matrix (with and of the same order as ) whose entries satisfy first four moments conditions; the sample column vectors are assumed to be i.i.d, centered and of covariance matrix a deterministic Hermitian (resp. symmetric) matrix having all but finitely many eigenvalues equal to one. In their pioneering article on that topic [Bk-B-P], Baik-Ben Arous-Péché pointed out a phase transition phenomenon for the fluctuations of the largest eigenvalue of according to the largest eigenvalue of , in the complex Gaussian setting; their results were extended in [P] to the real case when the largest eigenvalue of is simple and sufficiently larger than 1 and in [O] to singular Wishart matrices. In the non Gaussian case, the fluctuations of the extreme eigenvalues have been recently studied by Bai-Yao [B-Ya2] and Féral-Péché [Fe-Pe2].
The paper is organized as follows. In Section 2, we present the matricial models under study and the notations that will be used throughout the paper. In Section 3, we present the main results of this paper. We give a summary of our approach in Section 4. Section 5 is devoted to the proof of Theorem 3.2, Theorem 3.3 and Theorem 3.4. Finally, we recall some basic facts on matrices, a CLT for random sesquilinear forms and prove some technical results in an Appendix.
Model and notations
where the matrices and are defined as follows:
Let us now fix such that and let be a unitary matrix of size such that
where is an Hermitian matrix with eigenvalues strictly smaller than .
namely is the upper left corner of of size . It satisfies
All along the paper, the parameter is such that (resp. ) in the real (resp. complex) setting and we let .
Given an arbitrary Hermitian or symmetric matrix of size , we will denote by its ordered eigenvalues.
Main results
We first recall the a.s. convergence of the extreme eigenvalues. Define
Observe that (resp. ) when (resp. ) (and if ). For definiteness, we set if . In [C-D-F], we have established the following universal convergence result.
(a.s. behaviour) Let (resp. ) be the number of j’s such that (resp. ).
¿From Theorem 3.1, for all , converges to a.s.. We shall describe their fluctuations in the extreme two cases:
Case a) localization of the eigenvectors associated to : The sequence is bounded,
Case b) delocalization of the eigenvectors associated to : when and satisfies
The main results of our paper are the following two theorems. Let be defined by
In Case a) (which includes the particular setting of Proposition 1.2), the fluctuations of the corresponding rescaled largest eigenvalues of are not universal.
In Case a): the -dimensional vector
Then, is the matrix defined by
In Case b): the -dimensional vector
converges in distribution to where the matrix is distributed as the GU(O)E().
Note that since is symmetric, analogue results can be deduced from Theorem 3.2 and Theorem 3.3 dealing with the lowest eigenvalues of and the such that .
where is a matrix of size defined by , with , . Then , , , . For , we are in Case a) if is bounded and in Case b) if . For , we are in Case a).
Dealing with a spike with multiplicity 1, it turns out that case b) is actually the unique situation where universality holds since we establish the following.
If , , then the fluctuations of are universal, namely
for any , the distribution of is ;
is a centered gaussian variable with variance
Sketch of the approach
Before we proceed to the proof of Theorems 3.2 and 3.3, let us give the sketch of our approach which are similar in both cases. To this aim, we define for any random variable ,
with given by (3.3). We also set with the convention that . The reasoning made in the setting of Proposition 1.2 (for which ) relies (following ideas previously developed in [P] and [B-B-P]) on the writing of the rescaled eigenvalue in terms of the resolvent of an underlying non-Deformed Wigner matrix. The conclusion then essentially follows from a CLT on random sesquilinear forms established by J. Baik and J. Silverstein in the Appendix of [C-D-F] (which corresponds to the following Theorem 6.2 in the scalar case). In the general case, to prove the convergence in distribution of the vector \big{(}\xi_{N}(\lambda_{\hat{k}_{j-1}+i}({\bf M}_{N}));i=1,\ldots,k_{j}\big{)}, we will extend, as [B-Ya2], the previous approach in the following sense. We will show that each of these rescaled eigenvalues is an eigenvalue of a random matrix which may be expressed in terms of the resolvent of a Deformed Wigner matrix whose eigenvalues do not jump asymptotically outside ; then, the matrix will arise from a multidimensional CLT on random sesquilinear forms. Nevertheless, due to the multidimensional situation to be considered now, additional considerations are required. Let us give more details. Consider an arbitrary random variable which converges in probability towards . Then, applying factorizations of type (6.1), we prove that is an eigenvalue of iff is (on some event having probability going to 1 as ) an eigenvalue of a matrix of the form
Note that the authors do not develop this difficulty in [B-Ya2] (pp. 464-465). Hence, in the last step of the proof (Step 4 in Section 5), we detail the additional arguments which are needed to get (4.3) when .
Our approach will cover Cases a) and b) and we will handle both cases once this will be possible. In fact, the main difference appears in the proof of the convergence in distribution of the matrix which gives rise to the ”occurrence or non-occurrence” of the distribution in the limiting fluctuations and then justifies the non-universality (resp. universality) in Case a) (resp. b)).
The proof is organized in four steps as follows. In Steps 1 and 2, we explain how to obtain (4.2): we exhibit the matrix and bring its leading term to light in Step 2. We establish the convergence in distribution of the matrix in Step 3. Step 4 is devoted to the concluding arguments of the proof.
Proofs of Theorem 3.2, Theorem 3.3 and Theorem 3.4
For a matrix (or ) and some integers and , we denote respectively by , , and the upper left, upper right, lower left and lower right corner of size of the matrix . If , we will often replace the indices by for convenience. Moreover if , we may replace or by and or by . Similarly if , we may replace or by and or by .
For simplicity in the writing we will define the , resp. , resp. matrix , resp. , resp. , by setting
Note also that since is a submatrix of , all its eigenvalues are strictly smaller than . Let . For any random variable , define the events
On , neither nor are eigenvalues of , thus the resolvent of is well defined at and .
Let us now introduce on some auxiliary matrices that will be of basic use to the proofs.
Note that we will justify that is well defined in the course of the proof of Proposition 5.1 below. Finally, set
STEP 1: We show that an eigenvalue of is an eigenvalue of a matrix of size . More, precisely, we have:
Proof: Let be a random variable. On ,
Now, note that we have also from (6.1) that
Moreover on , one can see using (6.1) that if is an eigenvalue of then is an eigenvalue of
Using oncemore (6.1), we get that on , is an eigenvalue of if and only if it is an eigenvalue of or equivalently if and only if is an eigenvalue of
one can replace by and get the following writing
The proposition (adding an extra matrix for future computations) readily follows.
Throughout Steps 2 and 3, denotes any random sequence converging in probability towards . The aim of these two steps is to study the limiting behavior of the matrix (defined by (5.16)) as goes to infinity.
STEP 2: We first focus on the negligible terms in and establish the following.
Assume that . For any random sequence converging in probability towards , on ,
with defined by (5.6)
The proof of this proposition is quite long and is divided in several lemmas. Although our final result in the case infinite holds only for , we will give some estimates for once this is possible.
Let . Then, on ,
Proof of Lemma 5.1: and are respectively defined by (5.10), (5.7), (5.8) and (5.9) .
where we used that for or . Therefore,
It follows from Lemma 6.3 in the Appendix that
Since , the fourth moment of is uniformly bounded.
We skip the proof of this lemma which follows from straightforward computations using the independence of the entries of and the fact that is unitary. Then, according to Theorem 6.1 and using (5.1),
Besides for , using the independence between and , we have:
where we denote by the matrix for simplicity. ¿From Lemma 5.2, for , the only terms giving a non null expectation in the above equation are those for which:
, and . In this case,
. In this case, using (5.23), there is a constant such that
The convergence in probability of towards zero readily follows by Tchebychev inequality. Lemma 5.1 is established.
To get Proposition 5.2, it remains to prove that if ,
Once , we readily have that
Assume that . Let and be defined as (5.13) and (5.15). On ,
For the proof, we use the following decomposition ( and being defined by (5.11) and (5.12)):
and we replaced by . We will prove the following lemma on .
Proof of Lemma 5.4: To prove (5.29), we use the decomposition
so that, for and using Tchebychev inequality, we can deduce that
Thus (5.30) and Lemma 5.4 are proved.
one can readily notice that Lemma 5.4 leads to
Proof of Lemma 5.5 : We will show that, on , for any ,
One can readily see that this leads to the announced result combining Lemma 5.4, (5.31) and (5.33). First, using the fact that is independent of and that for any , the random vector has independent centered entries with variance , one has that
where we denote as before . Thus letting ,
We are now in position to conclude the proof of Lemma 5.3. Indeed, writing
which gives (5.27) and completes the proof of Lemma 5.3.
Combining all the preceding, we have established Proposition 5.2. We now prove that provided it converges in distribution, with a probability going to one as goes to infinity, is actually an eigenvalue of a matrix of size .
Proof: Straightforward computations lead to the existence of some constant such that
The convergence of in probability towards zero readily follows by Tchebychev inequality. Following the proof in Lemma 5.1 of the convergence in probability of towards zero, one can get that
and the convergence in probability towards zero of the term inside the above expectation follows by Tchebychev inequality. Since moreover according to Lemma 6.3,
The proof of Lemma 5.7 is complete.
Let be an arbitrary random matrix. If converges in distribution, then, with a probability going to one as goes to infinity, it is an eigenvalue of iff is an eigenvalue of a matrix of size , satisfying
where is the element in the block decomposition of defined by (5.6); namely
with and defined respectively by (2.3) and (5.5).
Since converges in distribution, we can write the matrix given by (5.20) as
We first show that is not an eigenvalue of . Let . Since,
if is an eigenvalue of , then
Hence cannot be an eigenvalue of . Therefore, we can define
This follows from the previous computations showing that (for some constant )
combined with the definition of and Lemma 5.7. The statement of the proposition then follows from (6.1).
STEP 3: We now examine the convergence of the matrix
The matrix converges in distribution to a GU(O)E( if and only if converges to zero when goes to infinity.
Proof Assume that converges to zero when goes to infinity. We decompose the proof of the convergence of in distribution to a GU(O)E( into the two following lemmas.
If converges to zero when goes to infinity then the matrix converges in distribution to a GU(O)E(.
where is the Hermitian matrix defined by
Then (5.36) readily follows. In the following, we let . Since , converges to zero for each . Thus we can deduce from Janson’s theorem [J] that converges to a centered gaussian distribution with variance and the proof of Lemma 5.8 is complete in the complex case.
Dealing with symmetric matrices, one needs to consider the random variable
for any real numbers . One can similarly prove that converges to a centered gaussian distribution with variance
Note that Lemma 5.8 is true under the assumption of the existence of a fourth moment. This can be shown by using a Taylor development of the Fourier transform of .
Under the assumption that converges to zero when goes to infinity, the last term in the r.h.s of the two above equations tends to 0.
It can be seen that the proof of Theorem 7.1 still holds in this case once we verify that for and for or , for any ,
We postpone the proof of (5.37) to the end of the proof. Assuming that (5.37) holds true, we obtain the CLT theorem 7.1 ([B-Ya2]): the Hermitian matrix of size defined by
where the matrix is given by: with
and the coefficients are defined in Theorem 6.2. Here so that and (see the Appendix). ¿From Lemma 5.2,
Moreover in the complex case, and in the real case,
It follows that is a diagonal matrix given by:
Therefore, .
Assume now that the matrix converges in distribution towards a GU(O)E( whereas does not converge to zero when goes to infinity. There exists such that does not converge to zero. Let be such that . Now we have
where is a random variable which is independent with . One can find a subsequence such that converges in distribution towards where and is -distributed. This leads to a contradiction using Cramer-Lévy’s Theorem since converges towards a gaussian variable. The proof of Proposition 5.4 is complete.
In the case a), condition of Proposition 5.4 are obviously not satisfied and we have the following asymptotic result.
Then, is the matrix defined by
The proof follows from Theorem 6.2 and is omitted since we have detailed the similar proof of Lemma 5.9.
STEP 4: We are now in position to prove that
To prove (5.40), our strategy will be indirect: we start from the matrix and its eigenvalues and we will reverse the previous reasoning to raise to the normalized eigenvalues . This approach works in both Cases a) and b) as we now explain.
First, for any , we define such that
The following lines hold on . By using Weyl’s inequalities (Lemma 6.1), one has for all that
Now, to get (5.40), it is sufficient to prove that
Indeed, one can notice that on the event the following equality holds true
where denoting by the Haar measure on the unitary (resp. orthogonal) group. Thus, we deduce that the eigenvalues of are distinct (with probability one). Using Portmanteau’s Lemma with (5.42) then implies that the event
According to Theorem 3.3, in order to establish Theorem 3.4, we only need to prove that the condition (3.6) is actually necessary for universality of the fluctuations. Hence assume that . Proposition 5.1 and Proposition 5.3 lead to
It follows that converges towards the gaussian distribution and then according to Proposition 5.4, converges to zero when goes to infinity.
Let such that and . Let us prove now the description given in subsection 3.2 of the fluctuations of for some intermediate situations between Case a) and Case b). Let be a fixed integer number. Assume that for any is independent of , whereas when goes to infinity. Following the proofs of Lemma 5.8 and Lemma 5.9, one can check that converges in distribution towards in the complex case, in the real case, where are independent random variables such that
for any , the distribution of is ;
is a centered gaussian variable with variance
Now, following the lines of Step 4 (using the results of Steps 1 and 2), we can conclude that converges in distribution towards the mixture of -distributed or gaussian random variables in the complex case, in the real case.
Appendix
In this section, we recall some basic facts on matrices and some results on random sesquilinear forms needed for the proofs of Theorems 3.2 and 3.3.
For Hermitian matrices, denoting by the decreasing ordered eigenvalues, we have the Weyl’s inequalities:
(cf. Theorem 4.3.7 of [H-J]) Let B and C be two Hermitian matrices. For any pair of integers such that and , we have
For any pair of integers such that and , we have
In the computation of determinants, we shall use the following formula.
2 CLT for random sesquilinear forms
(Lemma 2.7 [B-S1]) Let be a Hermitian matrix and be a vector of size which contains i.i.d standardized entries with bounded fourth moment. Then there is a constant such that
,
,
.
Then the -dimensional random vector \frac{1}{\sqrt{N}}\Big{(}X(l)^{*}AY(l)-\rho(l){\operatorname{Tr}}A\Big{)} converges in distribution to a Gaussian complex-valued vector with mean zero. The Laplace transform of is given by
where the matrix is given by with:
3 CLT for the empirical distribution of a Wigner matrix and applications
(Theorem 1.1 in [B-Ya1]) Let be an analytic function on an open set of the complex plane including . If the entries of a general Wigner matrix of variance satisfy the conditions
then N\Big{(}\operatorname{tr}_{N}(f(\frac{1}{\sqrt{N}}{\bf W}_{N}))-\int fd\mu_{sc}\Big{)} converges in distribution towards a Gaussian variable, where is the semicircle distribution of variance .
We now prove some convergence results of the resolvent used in the previous proofs. Let and such that .
Each of the following convergence holds in probability as :
,
,
.
Proof of Lemma 6.3: We denote by the resolvent of the non-Deformed Wigner matrix . i) By Theorem 6.3, one knows that converges in probability towards 0. Now, we have (see [H-P] p. 94). It is thus enough to show that
Let then (resp. ) be a unitary (resp. diagonal) matrix such that . Then, one has
ii) It is sufficient to show that in probability since, by Theorem 6.3, one knows that converges in probability towards . Using the fact that , it is not hard to see that
Acknowledgments We would like to thank the anonymous referees for their pertinent comments which led to an overall improvement of the paper.