Central limit theorems for linear statistics of heavy tailed random matrices
Florent Benaych-Georges, Alice Guionnet, Camille Male
Introduction and statement of results
Recall that a Wigner matrix is a symmetric random matrix such that
the sub-diagonal entries of are independent and identically distributed (i.i.d.),
the random variables are distributed according to a measure that does not depend on and have all moments finite.
This model was introduced in 1956 by Wigner who proved the convergence of the moments
when is centered with unit variance. Moments can be easily replaced by bounded continuous functions in the above convergence and this convergence holds almost surely. Assumption 2 can also be weakened to assume only that the second moment is finite. The fluctuations around this limit or around the expectation were first studied by Jonsson in the (slightly different) Wishart model, then by Pastur et al. in , Sinai and Soshnikov with possibly going to infinity with . Since then, a long list of further-reaching results have been obtained: the central limit theorem was extended to so-called matrix models where the entries interact via a potential in , the set of test functions was extended and the assumptions on the entries of the Wigner matrices weakened in , a more general model of band matrices was considered in (see also for general covariance matrices), unitary matrices where considered in , and Chatterjee developed a general approach to these questions in , under the condition that the law can be written as a transport of the Gaussian law. Finally, but this is not really our concern here, the fluctuations of the trace of words in several random matrices were studied in . It turns out that in these cases
converges towards a Gaussian variable whose covariance depends on the first four moments of . Moments can also be replaced by regular enough functions and Assumption 2 can be weakened to assume that the fourth moment only is finite. The latter condition is however necessary as the covariance for the limiting Gaussian depends on it. The absence of normalization by shows that the eigenvalues of fluctuate very little, as precisely studied by Erdös, Schlein, Yau, Tao, Vu and their co-authors, who analyzed their rigidity in e.g. .
In this article, we extend these results for a variation of the Wigner matrix model where Assumption 2 is removed: some entries of the matrix can be very large, e.g. when does not have any second moment or when it depends on , with moments growing with . Then, Wigner’s convergence theorem (1) does not hold, even when moments are replaced by smooth bounded functions. The analogue of the convergence (1) was studied when the common law of the entries of belongs to the domain of attraction of an -stable law or depends on and has moments blowing up with . Although technical, the model introduced in Hypothesis 1.1 below has the advantage of containing these two examples (for some sequences depending implicitly on , means that as ).
Let, for each , be an real symmetric random matrix whose sub-diagonal entries are some i.i.d. copies of a random variable (depending implicitly on ) such that: The random variable can be decomposed into such that as ,
Moreover, if the ’s are independent copies of ,
For any independent of , the random variable can be decomposed into such that
Examples of random matrices satisfying Hypothesis 1.1 are defined as follows.
Let be a random symmetric matrix with i.i.d. sub-diagonal entries.
We say that is a Lévy matrix of parameter in when where the entries of have absolute values in the domain of attraction of -stable distribution, more precisely
We say that is a Wigner matrix with exploding moments with parameter whenever the entries of are centered, and for any
Both Lévy matrices and Wigner matrices with exploding moments satisfy Hypothesis 1.1. For Lévy matrices, the function is given by formula
The proof of this lemma, and of Lemmas 1.8 and 1.12, which show that our hypotheses hold for both Lévy matrices and Wigner matrices, are given in Section 6.
One can easily see that our results also apply to complex Hermitian matrices: in this case, one only needs to require Hypothesis 1.1 to be satisfied by the absolute value of non diagonal entries and to have going to zero as .
A Lévy matrix whose entries are truncated in an appropriate way is a Wigner matrix with exploding moments . The recentered versionThe recentering has in fact asymptotically no effect on the spectral measure as it is a rank one perturbation. of the adjacency matrix of an Erdös-Rényi graph, i.e. of a matrix such that
is also an exploding moments Wigner matrix, with (the measure of Lemma 1.3 is ). In this case the fluctuations were already studied in . The method of can be adapted to study the fluctuations of linear statistics of Wigner matrices with exploding moments. Nevertheless, since we actually use Wigner matrices with exploding moments to study of the fluctuations of Lévy matrices, it is worthwhile to study these ensembles together.
The weak convergence of the empirical eigenvalues distribution of a Lévy matrix has been established in (see also ) where it was shown that for any bounded continuous function ,
where is a heavy tailed probability measure which depends only on . Moreover, converges towards the semicircle law as goes to .
The convergence in moments, in expectation and in probability, of the empirical eigenvalues distribution of a Wigner matrix with exploding moments has been established by Zakharevich in . In that case, moments are well defined and for any continuous bounded function ,
where is a probability measure which depends only on the sequence .
We shall first state the fluctuations of moments of a Wigner matrix with exploding moments around their limit, namely prove the following theorem.
Let be a Wigner matrix with exploding moments with parameter . Then the process
converges in distribution to a centered Gaussian process.
This theorem has been established for the slightly more restrictive model of adjacency matrices of weighted Erdös-Rényi graphs in . Our proof is based on the moment method, the covariance of the process is of combinatorial nature, given in Section 2, Formula (40) and Theorem 2.2.
Note that the speed of the central limit theorem is as for independent integrable random variables, but differently from what happens for standard Wigner’s matrices. This phenomenon has also already been observed for adjacency matrix of random graphs and we will see below that it also holds for Lévy matrices. It suggests that the repulsive interactions exhibited by the eigenvalues of most models of random matrices with lighter tails than heavy tailed matrices no longer work here.
By the previous results, for both Lévy and exploding moments Wigner matrices, converges in probability to a deterministic limit as the parameter tends to infinity. We study the associated fluctuations.
In fact, even in the case of Wigner matrices with exploding moments, the CLT for moments does not imply a priori the CLT for Stieltjes functions even though concentration inequalities hold on the right scale, see . Indeed, one cannot approximate smooth bounded functions by polynomials for the total variation norm unless one can restrict oneself to compact subsets, a point which is not clear in this heavy tail setting. However, with additional arguments based on martingale technology, we shall prove the following result, valid for both Lévy matrices and Wigner matrices with exploding moments.
Under Hypothesis 1.1, where we assume additionally that for , there exists such that, for any , one has
Under the hypotheses of Theorem 1.5, the process
where is a centered gaussian process with covariance given in Theorem 1.5 and the function is given by
where is a smooth compactly supported function equal to one in the neighborhood of the of the origin.
The function in Theorem 1.5 is given by , where
where for the matrix obtained from by deleting the -th row and column, and for a copy of where the entries for or are independent of and the others are those of . Also, we assume and in (19).
The existence of the limit (19) is a consequence of a generalized convergence in moments, namely the convergence in distribution of traffics, of stated in , see Lemma 3.2. However, under stronger assumptions, an independent proof of this convergence and an intrinsic characterization of are provided in Theorem 1.13 below.
Let us first mention that the map is the almost sure point-wise limit
This fact was known in the Lévy case and is proved in greater generality in Corollary 3.3. Let us first give a characterization of .
The function of (7) admits the decomposition
where is a function such that for some constants , we have
Both Lévy matrices and Wigner matrices with exploding moments satisfy Hypothesis 1.7. For Lévy matrices, the function is , with , whereas for Wigner matrices with exploding moments, the function is
Above, the second point in Hypotheses 1.1 is not required anymore, as it served mainly to prove convergence, which is now insured by the uniqueness of limit points.
Note that the asymptotics of Wigner matrices with bounded moments is also described by (25). In this case , so and , which leads, by formula (21), to the classical quadratic equation
Let us now give a fixed point characterization for the function of (19).
The function of (7) either has the form
or admits the decomposition, for non zero:
Unfortunately does not satisfy (29) as its density blows up at : we shall treat both case separately.
The following lemma insures us that our two main examples satisfy Hypothesis 1.10.
Both Lévy matrices and Wigner matrices with exploding moments satisfy Hypothesis 1.10. For Wigner matrices with exploding moments, the measure is given by
Under Hypotheses 1.1, 1.7 and 1.10, the conclusions of Theorem 1.5 and Corollary 1.6 hold and the parameter of (18) is given by
where are analytic functions on and uniformly continuous on compacts in the variables (and - Hölder for in the Lévy matrices case, see Lemma 5.1), given by (1.13) as far as is concerned and unique solution, among such functions, of the following fixed point equation (1.13) as far as is concerned:
with the notations , and the measures , defined by Hypothesis 1.10 and Remark 1.11.
Let us conclude this introduction with three remarks.
Let be a Lévy matrix as defined at Definition 1.2 but with . Then using Example c) p. 44. of instead of the hypothesis made at Equation (7), one can prove that as , the spectral measure of converges almost surely to the semi-circle law with support $P(|X_{ij}|\geq u)\sim cu^{-2}c>0X\frac{X}{\sqrt{cN\log(N)}}$.
Our results also have an application to standard Wigner matrices (i.e. symmetric random matrices of the form , with having centered i.i.d. sub-diagonal entries with variance one and not depending on . In this case, the function of (7) is linear, which implies that for all , so that the covariance is null, (25) is the self-consistent equation satisfied by the Stieltjes transform of the semi-circle law, namely (26), and Corollary 1.6 only means that for functions , we have, for the convergence in probability,
This result is new for Wigner matrices whose entries have a second but not a fourth moment, (36) brings new information. Indeed, for such matrices, which could be called “semi heavy-tailed random matrices”, the convergence to the semi circle law holds (see or the remark right above that one) but the largest eigenvalues do not tend to the upper-bound of the support of the semi-circle law, are asymptotically in the scale (with as in Equation (8) when such an exponent exists) and distributed according to a Poisson process (see ), and it is not clear what the rate of convergence to the semi-circle law will be. Equation (36) shows that this rate is .
but otherwise a non trivial recentering should occur. See the end of Section 2.
CLT for the moments of Wigner matrices with exploding moments
The goal of this section is to prove Theorem 1.4. In order to prove the CLT for the moments of the empirical eigenvalues distribution of , we use a modification of the method of moments inspired by which consists of studying more general functionals of the entries of the matrix (the so-called injective moments) than only its moments. We describe this approach below.
Let be a Wigner matrix with exploding moments. Let be an integer. The normalized trace of the -th power of can be expanded in the following way.
where is the set of partitions of and is the set of multi-indices in such that .
We interpret as a functional on graphs instead of partitions. Let be a partition of . Let be the undirected graph (with possibly multiple edges and loops) whose set of vertices is and with multi-set of edges given by: there is one edge between two blocks and of for each in such that and (with notation modulo ). Then, one has
where and for any edge we have denoted A\big{(}\phi(e)\big{)}=A\big{(}\phi(i),\phi(j)\big{)}. There is no ambiguity in the previous definition since the matrix is symmetric.
to a Gaussian process, it is sufficient to prove the convergence of
Before giving the proof of this fact, we recall a result from , namely the convergence of for any partition . These limits are involved in our computation of the covariance of the limiting process of \big{(}Z_{N}(T^{\pi})\big{)}_{\pi\in\cup_{K}\mathcal{P}(K)}, and this convergence will be useful in the proof of the CLT for Stieltjes transforms latter.
Let be a Wigner matrix with exploding moments with parameter . For any partition in , with defined in (37),
where a fat tree is a graph that becomes a tree when the multiplicity of the edges is forgotten, and for such a graph we have denoted the number of edges of with multiplicity .
Let be a Wigner matrix with exploding moments. Then, the process \big{(}Z_{N}(T^{\pi})\big{)}_{\pi\in\cup_{K}\mathcal{P}(K)} defined by (39) converges to a centered Gaussian process \big{(}z(T^{\pi})\big{)}_{\pi\in\cup_{K}\mathcal{P}(K)} whose covariance is given by: for any ,
where is given by Proposition 2.1 and is the set of graphs obtained by considering disjoint copies of the graphs and and gluing them by requiring that they have at least one edge (and therefore two “adjacent” vertices) in common.
We show the convergence of joint moments of \big{(}Z_{N}(T^{\pi})\big{)}. Gaussian distribution being characterized by its moments, this will prove the theorem. Let be finite undirected graphs, each of them being of the form for a partition . We first write
is the set of partitions of the disjoint union of whose blocks contain at most one element of each ,
is the set of families of injective maps, , such that for any , one has .
First, it should be noticed that by invariance in law of by conjugacy by permutation matrices, for any in and in , the quantity depends only on . We then denote . Moreover, choosing a partition in is equivalent to merge certain vertices of different graphs among . We equip with the edges of and say that two vertices are adjacent if there is an edge between them. We denote by the subset of such that any graph has two adjacent vertices that are merged to two adjacent vertices of an other graph. By the independence of the entries of and the centering of the components in , for any in one has . Hence, since the cardinal of is , we get
Let in . We now analyze the term . We first expand its product.
where is the number of edges of once the multiplicity and the orientation of edges are forgotten. Recall assumption (9): for any ,
By the Cauchy-Schwarz inequality, for any ,
Hence, since the measure is centered, the quantity is bounded. Denote the graph obtained by merging the vertices of that belong to a same block of , its number of edges when orientation and multiplicity is forgotten, and by its number of components. We obtain from (42) and (43)
A partition induces a partition of : if and only if and belong to a same connected component of . Denote by the set of pair partitions of . One has
Secondly, one has with equality if and only if , so that
Moreover, by [28, Lemma 1.1] is the number of cycles of , the graph obtained from by forgetting the multiplicity and the orientation of its edges. Hence,
By (44), (45) and (46), if we denote by we get
where we have used the independence of the entries of to split . The case gives
The general case in (48) gives the Wick formula
which characterizes the Gaussian distribution. ∎
and therefore we obtain the same CLT if we recenter with the limit or the expectation, as noticed in the introduction.
CLT for Stieltjes transform and the method of martingales
Hence one can suppose that in Hypothesis 1.1, is centered.
and their complex transposes converges in distribution to a complex Gaussian variable. Since , it is enough to fix a linear combination
Notice first that (50) implies (51). Let us now prove (52). The proof of (50) will then be the main difficulty of the proof of Theorem 1.5.
Hence by (107) of Lemma 7.5 in the appendix, there is such that for all and all ,
We now pass to the main part of the proof of Theorem 1.5, namely the proof of (50). It is divided into several steps.
We can get rid of the linear combination in (49) and assume . As , both and are linear combinations of terms of the form
2. Removing the off-diagonal terms
In this section, we prove that we can replace in (53) by
Then for defined as in (53), as , we have the convergence
the ’s such that or are i.i.d. copies of (modulo the fact that is symmetric), independent of ,
for all other pairs , ,
where is defined out of in the same way as is defined out of in (60) (note that the -th column is the same in and in ). It follows that
We shall in the sequel prove that as tends to infinity, regardless to the value of , we have the almost sure convergences
and for any , as so that , we have
where . The convergences of (64) and (65) are based on an abstract convergence result stated in next section, where we use the convergence of generalized moments of Proposition 2.1. They are stated in Lemmas 3.4 and 3.5 respectively.
so once (64) and (65) proved, we will have the convergence in
Thus by (63), we will have proved the convergence of (61), hence completing the proof of the theorem.
3. An abstract convergence result
Remember that is the square matrix of size obtained by removing the -th row and the -th column of , that and that is a copy of where the entries for or are independent of and the other are those of . We denote .
By e.g. Theorem C.8 of , it is enough to prove that for any bounded and Lipschitz function , converges almost surely to . Moreover, adapting the proof of Lemma 7.4, one can easily see that for any such ,
converges almost surely to zero. The only modification of the proof is to complete the resolvent identity by noticing that
As is small, ,
and the Claim is proved, by Hypothesis (5).
We consider a polynomial and remark that
where denotes the Hadamard (entry-wise) product of matrices and
Such a family exists by Proposition 7.10. By [38, Proposition 3.10], the couple of random matrices satisfies the so-called convergence in distribution of traffics, so that the RHS converges. For the reader’s convenience, we give the limiting value of (68), even if we do not use it later. It is obtained by applying the rule of the so-called traffic freeness in .
we have considered the graph whose edges are labelled by indeterminates and , obtained by
identifying the vertex of each of these graphs (we get a connected graph, bouquet of cycles),
we have denoted by the set of vertices of , is the set of partitions of and denotes the graph obtained by identifying the vertices of that belong to a same block of ,
the quantity is as in Proposition 2.1,
where the sum is over all partitions of the set of vertices of and is the number of edges between adjacent vertices of , .
Formula (69) could also be derived by the same techniques than those developed in Section 2. The random variables and are distributed according to the limiting eigenvalues distribution of . Recall that we assume that the sequence defined in (9) satisfies for a constant . Then, following the proof of [49, Proposition 10], the exponential power series of the limiting eigenvalues distribution of has a positive radius of convergence. So, by a generalization of [12, Theorem 30.1] to the multi-dimensional case, we get that the distribution of is characterized by its moments. Then, we get that converges weakly to a family of random variables . We set . We then obtain the convergence
This convergence could also be proven without Proposition 7.10 but with appropriate bounds on the growth of moments. We have the following Corollary.
We used the exponential decay to switch the integral and the expectation. This also allows to truncate the integral: for any ,
where goes to zero as tends to infinity, uniformly on and on the randomness. Remember that by assumption (7), we have
where converges almost surely to zero as goes to infinity. By Corollary 3.3,
where converges to zero almost surely. Hence, we deduce the almost sure convergence
As has non positive real part, we conclude by dominated convergence theorem and by getting rid of the truncation of the integral. ∎
where and is defined in Corollary 3.3.
Hence, by (72), since by (108) the sign of the imaginary part of is , the random variable can be written
where and \eta_{m}(z)=\frac{\partial_{z}\lambda_{N}(z)}{\lambda_{N}(z)}\big{(}1-e^{i\operatorname{sgn}_{z}m\lambda_{N}(z)}\big{)}.
We next show that for all there exists so that for , large enough
By (110) and since the sign of the imaginary part of is , one has . More precisely, for any , we find
By (4), we deduce that for any we can choose small enough so that for large enough
so that for large enough and ,
where is arbitrary small as is large, uniformly on and on the randomness. Moreover, by (79) one has
Remind that in (74) we have shown in the proof of the last item of Corollary 3.3
As in the proof of Corollary 3.3, since the left hand side is analytic and uniformly bounded, we deduce by Montel’s theorem that its convergence entails the convergence of its derivatives. We then get by (81), for all so that , the almost sure convergence
We then obtain by dominated convergence (remember the integrant is uniformly bounded by (80) and (79)),
so we can let going to infinity to obtain the almost sure convergence
In the Lévy case, one has , and in the exploding moments case, \big{|}\frac{1}{t}\partial_{z}e^{itz+\rho_{z}(t)}\big{|}\leq 1+\frac{C_{2}}{\Im z^{2}}, so that the integral converges at zero and we obtain
where is defined in Corollary 3.3.
We shall start again by using Formula (3.4) for , and its analogue for
where is defined as , replacing by and the matrix by the matrix defined by (62), which gives
The upper bound (79) allows us to bound the first term uniformly by and to truncate the integrals for . Therefore, up to a small error uniform for large, small and provided , we have
As in the previous case, the upper bound (79) allows us to write
By assumption (7) and Corollary 3.3, we have the following almost sure convergence as goes to infinity and goes to in
Hence we have proved the convergences (64) and (65). This completes the proof of Theorem 1.5.
Proof of Corollary 1.6
Applying this to the empirical measure of eigenvalues and its expectation, we deduce that
Note here that is supported in a compact set and is bounded by for a finite constant . Hence, the integral is well converging. We next show that we can approximate it by
where . The random variable is only a finite sum of and therefore it converges in law by Theorem 1.5, towards
We finally show the convergence in probability of and to zero by bounding their norms
But, by Lemma 7.3, there exists a finite constant such that for any
Taking we deduce that there exists a finite constant such that
Fixed point characterizations
In this section, we provide characterizations of the functions and involved in the covariance of the limiting process of Theorem 1.5 as fixed points of certain functions. In fact, we also give an independent proof of the existence of such limits, and of Corollary 3.3.
where we have proved that this convergence holds almost surely in Corollary 3.3. Note however that under the assumptions of Theorem 1.9, the arguments below provide another proof of this convergence where we do not have to assume (14).
Then we use Hypothesis made at (22) to get for
Using the definition of and the fact that we assumed that it is bounded on every compact subset (since has non positive real part we can cut the integral to keep bounded up to a small error, as in the previous sections), we have
We now find a fixed point system of equations for the non random function of Corollary 3.3. For , we set
where we recall that and are as in (19). To simplify the notations below, as in the previous section, we denote instead of , even though their distribution depends on .
In the sequel we fix, as in Section 3.3, a number and will give limits in the regime where , and . We shall then prove that, under the hypotheses of Theorem 1.13 that we assume throughout this section, converges almost surely and that its limit satisfies a fixed point system of equations which has a unique analytic solution with non positive real part. The convergence could be shown with minor modifications of Lemma 3.2 under assumption (14), but we do not need this since we work with stronger assumptions. Using the concentration lemma 7.4 (note that is not Lipschitz but can be approximated by Lipschitz functions uniformly on compacts), it is sufficient to prove the fixed point equation for the expectation of these parameters. Moreover, by exchangeability of the first entries and last entries
by Lemma 7.2 (by continuity of , it can be approximated by Lipschitz functions). Hence we find that the limit points of satisfy (1.13) and (1.13). Moreover,
where for , we have set
We put where denote the constant function equal to one. We get after integrating both sides
with .
Proving that the limit points of are -Hölder maps for a allows to conclude the proof of Theorem 1.13. This is the content of the following lemma.
We next show that for any matrix model so that , for any
where we have used Fubini for non negative functions. But the above integral is well converging at infinity and we know that converges uniformly on for all finite; hence there exists a finite constant so that for large enough
Applying this with and gives
so that we deduce that for all
Equation (93) completes the proof by taking . ∎
We now prove the uniqueness of the solution of (1.13) on the set of functions described above. After some change of variables, the equation is equivalent to the following:
where we have denoted in short , , , and stand for the scalar products.
After the change of variables , , we can rewrite this system of equations as
with, if when ,
where we have denoted and
for a constant . The desired uniqueness follows from the following lemma.
To study the Lipschitz property of as a function of in for the norm , we first set
We next bound, for given in , in and with either in or (which allows to treat in one time two parts of )
Moreover, notice that for in , one has and and thus
so that . At last, with for , straightforward uses of the -norm and the inequalities , \big{|}|a_{1}|-|a_{2}|\big{|}\leq v|T_{1}-T_{2}|, if , \big{|}\frac{a_{1}}{|a_{1}|}-\frac{a_{2}}{|a_{2}|}\big{|}\leq(v\vee 1)^{-2}|T_{1}-T_{2}|v(1+v), and (94) gives the estimates
where f_{\beta}(v)=v^{\beta}\Big{(}(1+v)^{\frac{\alpha}{2}+\beta}(v\vee 1)^{-2\beta}+1\Big{)}. Using this series of estimates, we find that,
with if is large enough. ∎
Taking two solutions of (1.13) and (1.13) in , we deduce that they are equal when is large enough, and thus everywhere by analyticity.
Proofs of Lemmas 1.3, 1.8 and 1.12
Let us first treat the case of Wigner matrices with exploding moments. First and second parts of Hypothesis 1.1, as well as (4), are satisfied for . Let us then define to be the law of and to be the measure with density with respect to , so that for any test function , we have
Hence (5) is satisfied if is chosen large enough. For the convergence of the truncated even moments, see [38, Sect. 1.2.1]. Moreover, (7) follows from the results of e.g. Section 8.1.3 of . At last, (4) is satisfied for Lévy matrices by e.g. Section 10 of .
2. Proof of Lemma 1.8
In the case of Lévy matrices, the expression
which is proved in the following way (using (97)):
It follows that for the measure introduced in the proof of Lemma 1.3 above, we have
for . As converges weakly to and is continuous and bounded, we have
3. Proof of Lemma 1.12
The case of Lévy matrices is obvious. To treat the case of Wigner matrices with exploding moment, first note that by (98), writing
Then one concludes as for the proof of Lemma 1.8.
Appendix
This section is mostly a reminder of results from and .
The next lemma is an easy consequence of Cauchy-Weyl interlacing Theorem. It is an ingredient of the proof of Lemma 7.3.
where and , both supremums running over the elements of .
Moreover, let and note that and have rank one so that has rank one. On the other hand it is bounded uniformly by . Hence, we can write with a unit vector and bounded by . Therefore, since is Lipschitz,
Then, for , , we have for all , ,
The first point is proved as in [14, Lemma C.3]. We outline the proof of the second point which is very similar to [14, Lemma C.3]. We concentrate on , the other case being similar. By Azuma-Hoefding inequality, it is sufficient to show that
Let be the submatrix of obtained by removing its -th row and its -th column and set . Let also be the -th column of where the -th entry has been removed. Then
For (106), see [4, Th. A.5]. For (107), see Lemma 7.1. ∎
With the notation introduced above the previous lemma, for each ,
Let us now prove (110). By (108), we know that and have the same sign, so
Hence it remains only to prove (110). This is a direct consequence of (108) and (109) which imply the second and last inequality
2. Vanishing of non diagonal terms in certain quadratic sums of random vectors
Let denote the operator norm of a complex matrix with respect to the canonical Hermitian norms.
For each , let be a family of i.i.d. copies of an random variable such that can be decomposed into with such that is centered and (2), (3) of Hypothesis 1.1 are satisfied. Let also be a non random matrix such that is bounded. Then we have the convergence in probability
3. CLT for martingales
Let be a filtration such that and let be a square-integrable complex-valued martingale starting at zero with respect to this filtration. For , we define the random variables
Let now everything depend on a parameter , so that
Then we have the convergence in distribution
4. Extension of CLTs for random matrices
The following lemma is borrowed from the paper of Shcherbina and Tirozzi , except that we do not require, in the hypotheses here, to be continuous, which is very useful in our case.
where is a sequence of real numbers. We make the following hypotheses :
Then is continuous on , can be (uniquely) continuously extended to and (114) is true for any .
This is exactly Proposition 4 of , except that in , the hypotheses include the continuity of . Let us prove that the hypotheses made here imply that is continuous on . For any , is the second moment of the limit law of . Hence
This proves that the quadratic form is continuous.∎
5. On the Hadamard product of Hermitian matrices
Let be by Hermitian random matrices whose entries have all their moments. Then, there exists a family of random variables whose joint distribution is given by:
where denotes the Hadamard (entry-wise) product.
Step 1. We first assume that the matrices are deterministic and have distinct eigenvalues. By the spectral decomposition, for , we have where is the family of eigenvalues of arranged in increasing order, and is the family of associated eigenvectors. For any , one has
For , we set the empirical eigenvalues distributions of . We denote F_{\Lambda_{j}}(t)=\mu_{\Lambda_{j}}\big{(}(-\infty,t]\big{)} the cumulative function of . Since the eigenvalues of the matrices are distinct, one has for any , . Hence, we have
where f_{N}\big{(}(\lambda_{j},\Lambda_{j},U_{j})_{j=1,\ldots,p}\big{)}=\bigg{(}N^{p-1}\sum_{k=1}^{N}\prod_{j=1}^{p}\big{|}u_{j,(NF_{\Lambda_{j}}(\lambda_{j}))}(k)\big{|}^{2}\bigg{)}. Hence, a family of random variables as in the proposition exists and its joint distribution has density f_{N}\big{(}(\cdot,\Lambda_{j},U_{j})_{j=1,\ldots,p}\big{)} with respect to . Step 2. We now assume that are random and that their joint distributions have a density with respect to the Lesbegue measure on , where is the space of Hermitian matrices of size . In particular, the matrices have almost surely distinct eigenvalues, see . The spectral decompositions of the previous step are measurable (see [21, Section 5.3]) and, with the above notations for eigenvalues and eigenvectors of , we can write the joint distribution of in the form g_{N}\big{(}(\Lambda_{j},U_{j})_{j=1,\ldots,p}\big{)}\textrm{d}\mu_{\Delta_{N}}(\Lambda_{1})\dots\textrm{d}\mu_{\Delta_{N}}(\Lambda_{p})\textrm{d}\mu_{\mathcal{U}_{N}}(U_{1})\dots\textrm{d}\mu_{\mathcal{U}_{N}}(U_{p}). The symbol denotes the Lebesgue measure on and is the Haar measure on the set of unitary matrices of size . For any , one has
where is as in the previous step and
For any , we have
where is obtained by integrating with respect to the variables for . We finally obtain