The largest eigenvalues of finite rank deformation of large Wigner matrices: convergence and nonuniversality of the fluctuations

Mireille Capitaine, Catherine Donati-Martin, Delphine Féral

Introduction

This paper lies in the lineage of recent works studying the influence of some perturbations on the asymptotic spectrum of classical random matrix models. Such questions come from statistics (cf. John) and appeared in the framework of empirical covariance matrices, also called nonwhite Wishart matrices or spiked population models, considered by Baik, Ben Arous and Péché BBP and by Baik and Silverstein BS3. The work BBP deals with random sample covariance matrices (SN)N(S_{N})_{N} defined by

where YNY_{N} is a p×Np\times N complex matrix whose sample column vectors are i.i.d., centered, Gaussian and of covariance matrix a deterministic Hermitian matrix Σp{\Sigma}_{p} having all but finitely many eigenvalues equal to 1. Besides, the size of the samples NN and the size of the population p=pNp=p_{N} are assumed of the same order (as N→∞N\to\infty). The authors of BBP first noticed that, as in the classical case (known as the Wishart model) where Σp=Ip\Sigma_{p}=I_{p} is the identity matrix, the global limiting behavior of the spectrum of SNS_{N} is not affected by the matrix Σp\Sigma_{p}. Thus, the limiting spectral measure is the well-known Marchenko–Pastur law. On the other hand, they pointed out a phase transition phenomenon for the fluctuations of the largest eigenvalue according to the value of the largest eigenvalue(s) of Σp\Sigma_{p}. The approach of BBP does not extend to the real Gaussian setting and the whole analogue of their result is still an open question. Nevertheless, Paul was able to establish in Pa the Gaussian fluctuations of the largest eigenvalue of the real Gaussian matrix SNS_{N} when the largest eigenvalue of Σp\Sigma_{p} is simple and sufficiently larger than 1. More recently, Baik and Silverstein investigated in BS3 the almost sure limiting behavior of the extremal eigenvalues of complex or real nonnecessarily Gaussian matrices. Under assumptions on the first four moments of the entries of YNY_{N}, they showed in particular that when exactly kk eigenvalues of Σp{\Sigma}_{p} are far from 1, the kk first eigenvalues of SNS_{N} are almost surely outside the limiting Marchenko–Pastur support. Fluctuations of the eigenvalues that jump are universal and have been recently found by Bai and Yao in BY2 (we refer the reader to BY2 for the precise restrictions made on the definition of the covariance matrix Σp\Sigma_{p}). Note that the problem of the fluctuations in the very general setting of BS3 is still open.

Our purpose here is to investigate the asymptotic behavior of the first extremal eigenvalues of some complex or real Deformed Wigner matrices. These models can be seen as the additive analogue of the spiked population models and are defined by a sequence (MN)N(M_{N})_{N} given by

where WNW_{N} is a Wigner matrix such that the common distribution of its entries satisfies some technical conditions [given in (i) below] and ANA_{N} is a deterministic matrix of finite rank. We establish the analogue of the main result of BS3, namely that, once ANA_{N} has exactly kk (fixed) eigenvalues far enough from zero, the kk first eigenvalues of MNM_{N} jump almost surely outside the limiting semicircle support. This result is universal (as the one of BS3) since the corresponding limits only involve the variance of the entries of WNW_{N}. On the other hand, at the level of the fluctuations, we exhibit a striking phenomenon in the particular case where ANA_{N} is diagonal with a sole simple nonnull eigenvalue large enough. Indeed, we find that in this case, the fluctuations of the largest eigenvalue of MNM_{N} are not universal and strongly depend on the particular law of the entries of WNW_{N}. More precisely, we prove that the limiting distribution of the (properly rescaled) largest eigenvalue of MNM_{N} is the convolution of the distribution of the entries of WNW_{N} with a Gaussian law. In particular, if the entries of WNW_{N} are not Gaussian, the fluctuations of the largest eigenvalue of MNM_{N} are not Gaussian.

In the following section, we first give the precise definition of the Deformed Wigner matrices (2) considered in this paper and we recall the known results on their asymptotic spectrum. Then, we present our results and sketch the proofs. We also outline the organization of the paper.

Model and results

Throughout this paper, we consider complex or real Deformed Wigner matrices (MN)N(M_{N})_{N} of the form (2) where the matrices WNW_{N} and ANA_{N} are defined as follows:

WNW_{N} is an N×NN\times N Wigner Hermitian (resp., symmetric) matrix such that the N2N^{2} random variables (WN)ii(W_{N})_{ii}, 2ℜe((WN)ij)i<j\sqrt{2}\Re e((W_{N})_{ij})_{i<j}, 2ℑm((WN)ij)i<j\sqrt{2}\Im m((W_{N})_{ij})_{i<j} [resp., the N(N+1)/2N(N+1)/2 random variables 12(WN)ii\frac{1}{\sqrt{2}}(W_{N})_{ii}, (WN)ij(W_{N})_{ij}, i<ji<j] are independent identically distributed with a symmetric distribution μ\mu of variance σ2\sigma^{2} and satisfying a Poincaré inequality (see Section 3).

ANA_{N} is a deterministic Hermitian (resp., symmetric) matrix of fixed finite rank rr and built from a family of JJ fixed real numbers θ1>⋯>θJ\theta_{1}>\cdots>\theta_{J} independent of NN with some j0j_{0} such that θj0=0\theta_{j_{0}}=0. We assume that the nonnull eigenvalues θj\theta_{j} of ANA_{N} are of fixed multiplicity kjk_{j} (with ∑j≠j0kj=r\sum_{j\not=j_{0}}k_{j}=r), that is, ANA_{N} is similar to the diagonal matrix

Furthermore, note that this condition implies that μ\mu has moments of any order (cf. Corollary 3.2 and Proposition 1.10 in L).

Let us now introduce some notations. When the entries of WNW_{N} are further assumed to be Gaussian, that is, in the complex (resp., real) setting when WNW_{N} is of the so-called GUE (resp., GOE), we will write WNGW_{N}^{G} instead of WNW_{N}. Then XNG:=WNG/NX_{N}^{G}:=W_{N}^{G}/\sqrt{N} will be said to be of the GU(O)E(N,σ2NN,\frac{\sigma^{2}}{N}) and we will let MNG=XNG+ANM_{N}^{G}=X_{N}^{G}+A_{N} be the corresponding Deformed GU(O)E model.

In the following, given an arbitrary Hermitian matrix BB of order NN, we will denote by λ1(B)≥⋯≥λN(B)\lambda_{1}(B)\geq\cdots\geq\lambda_{N}(B) its NN ordered eigenvalues and by μB=1N∑i=1Nδλi(B)\mu_{B}=\frac{1}{N}\sum_{i=1}^{N}\delta_{\lambda_{i}(B)} its empirical measure. Spect⁡(B)\operatorname{Spect}(B) will denote the spectrum of BB. For notational convenience, we will also set λ0(B)=+∞\lambda_{0}(B)=+\infty and λN+1(B)=−∞\lambda_{N+1}(B)=-\infty.

The Deformed Wigner model is built in such a way that the Wigner theorem is still satisfied. Thus, as in the classical Wigner model (AN≡0A_{N}\equiv 0), the spectral measure (μMN)(\mu_{M_{N}}) converges a.s. toward the semicircle law μsc\mu_{sc} whose density is given by

This result follows from Lemma 2.2 of Bai. Note that it only relies on the two first moment assumptions on the entries of WNW_{N} and the fact that the ANA_{N}’s are of finite rank.

On the other hand, the asymptotic behavior of the extremal eigenvalues may be affected by the perturbation ANA_{N}. Recently, Péché studied in Pe the Deformed GUE under a finite rank perturbation ANA_{N} defined by (ii). Following the method of BBP, she highlighted the effects of the nonnull eigenvalues of ANA_{N} at the level of the fluctuations of the largest eigenvalue of MNGM_{N}^{G}. To explain this in more detail, let us recall that when AN≡0A_{N}\equiv 0, it was established in TW that as N→∞N\rightarrow\infty,

where F2F_{2} is the well-known GUE Tracy–Widom distribution (see TW for the precise definition). Dealing with the Deformed GUE MNGM_{N}^{G}, it appears that this result is modified as soon as the first largest eigenvalue(s) of ANA_{N} is (are) quite far from zero. In the particular case of a rank-1 perturbation ANA_{N} having a fixed nonnull eigenvalue θ>0\theta>0, Pe proved that the fluctuations of the largest eigenvalue of MNGM_{N}^{G} are still given by (5) when θ\theta is small enough and precisely when θ<σ\theta<\sigma. The limiting law is changed when θ=σ\theta=\sigma. As soon as θ>σ\theta>\sigma, Pe established that the largest eigenvalue λ1(MNG)\lambda_{1}(M_{N}^{G}) fluctuates around

(which is >2σ>2\sigma since θ>σ\theta>\sigma) as

Similar results are conjectured for the Deformed GOE but Péché emphasized that her approach fails in the real framework. Indeed, it is based on the explicit Fredholm determinantal representation for the distribution of the largest eigenvalue(s) that is specific to the complex setting. Nevertheless, Maïda Ma obtained a large deviation principle for the largest eigenvalue of the Deformed GOE MNGM_{N}^{G} under a rank-1 deformation ANA_{N}; from this result she could deduce the almost sure limit with respect to the nonnull eigenvalue of ANA_{N}. Thus, under a rank-1 perturbation ANA_{N} such that DN=diag⁡(θ,0,…,0)D_{N}=\operatorname{diag}(\theta,0,\ldots,0) where θ>0\theta>0, Ma showed that

Note that the approach of Ma extends with minor modifications to the Deformed GUE. Following the investigations of BS3 in the context of general spiked population models, one can conjecture that such a phenomenon holds in a more general and nonnecessarily Gaussian setting. The first result of our paper, namely the following Theorem 2.1, is related to this question. Before being more explicit, let us recall that when AN≡0A_{N}\equiv 0, the whole spectrum of the rescaled complex or real Wigner matrix XN=WN/NX_{N}=W_{N}/{\sqrt{N}} belongs almost surely to the semicircle support [−2σ,2σ][-2\sigma,2\sigma] as NN goes to infinity and that (cf. BYi or Theorem 2.12 in Bai)

Note that this last result holds true in a more general setting than the one considered here (see BYi for details) and in particular only requires the finiteness of the fourth moment of the law μ\mu. Moreover, one can readily extend the previous limits to the first extremal eigenvalues of XNX_{N}, that is,

Here, we prove that, under the assumptions (i)–(ii), (12) fails when some of the θj\theta_{j}’s are sufficiently far from zero: as soon as some of the first largest (resp., last smallest) nonnull eigenvalues θj\theta_{j} of ANA_{N} are taken strictly larger than σ\sigma (resp., strictly smaller than −σ-\sigma), the same part of the spectrum of MNM_{N} almost surely exits the semicircle support [−2σ,2σ][-2\sigma,2\sigma] as N→∞N\to\infty and the new limits are the ρθj\rho_{\theta_{j}}’s defined by

Observe that ρθj\rho_{\theta_{j}} is >2σ>2\sigma (resp., <−2σ<-2\sigma) when θj>σ\theta_{j}>\sigma (resp., <−σ<-\sigma) (and ρθj=±2σ\rho_{\theta_{j}}=\pm 2\sigma if θj=±σ\theta_{j}=\pm\sigma).

Here is the precise formulation of our result. For definiteness, we set k1+⋯+kj−1:=0k_{1}+\cdots+k_{j-1}:=0 if j=1j=1.

Let J+σJ_{+\sigma} (resp., J−σJ_{-\sigma}) be the number of j’s such that θj>σ\theta_{j}>\sigma (resp., θj<−σ\theta_{j}<-\sigma).

∀1≤j≤J+σ,∀1≤i≤kj,λk1+⋯+kj−1+i(MN)⟶ρθj\mboxa.s.\forall 1\leq j\leq J_{+\sigma},\forall 1\leq i\leq k_{j},\lambda_{k_{1}+\cdots+k_{j-1}+i}(M_{N})\longrightarrow\rho_{\theta_{j}}\mbox{ a.s.}

λk1+⋯+kJ+σ+1(MN)⟶2σ\mboxa.s.\lambda_{k_{1}+\cdots+k_{J_{+\sigma}}+1}(M_{N})\longrightarrow 2\sigma\mbox{ a.s.}

λk1+⋯+kJ−J−σ(MN)⟶−2σ\mboxa.s.\lambda_{k_{1}+\cdots+k_{J-J_{-\sigma}}}(M_{N})\longrightarrow-2\sigma\mbox{ a.s.}

∀j≥J−J−σ+1,∀1≤i≤kj,λk1+⋯+kj−1+i(MN)⟶ρθj\mboxa.s.\forall j\geq J-J_{-\sigma}+1,\forall 1\leq i\leq k_{j},\lambda_{k_{1}+\cdots+k_{j-1}+i}(M_{N})\longrightarrow\rho_{\theta_{j}}\mbox{ a.s}.

Following BS3, one can expect that this theorem holds true in a more general setting than the one considered here, namely one that would only require four first moment conditions on the law μ\mu of the Wigner entries. As we will explain in the following, the assumption that μ\mu satisfies a Poincaré inequality is actually fundamental in our reasoning since we will need several variance estimates.

This theorem will be proved in Section 4. The second part of this work is devoted to the study of the particular rank-1 diagonal deformation AN=diag⁡(θ,0,…,0)A_{N}=\operatorname{diag}(\theta,0,\ldots,0) such that θ>σ\theta>\sigma. We investigate the fluctuations of the largest eigenvalue of any real or complex Deformed model MNM_{N} satisfying (i) around its limit ρθ\rho_{\theta}. We obtain the following result.

Let AN=diag⁡(θ,0,…,0)A_{N}=\operatorname{diag}(\theta,0,\ldots,0) with θ>σ\theta>\sigma. Define

where t=4t=4 (resp., t=2t=2) when WNW_{N} is real (resp., complex) and m4:=∫x4 dμ(x)m_{4}:=\int x^{4}\,d\mu(x). Then

Note that when m4=3σ4m_{4}=3\sigma^{4} as in the Gaussian case, the variance of the limiting distribution of N(λ1(MN)−ρθ)\sqrt{N}(\lambda_{1}(M_{N})-\rho_{\theta}) is equal to σθ2\sigma_{\theta}^{2} (resp., 2σθ22\sigma_{\theta}^{2}) in the complex (resp., real) setting [with σθ\sigma_{\theta} given by (8)].

Since μ\mu is symmetric, it readily follows from Theorem 2.2 that when AN=diag⁡(θ,0,…,0)A_{N}=\operatorname{diag}(\theta,0,\ldots,0) and θ<−σ\theta<-\sigma, the smallest eigenvalue of MNM_{N} fluctuates as N(1−σ2/θ2)−1(λN(MN)−ρθ)⟶Lμ∗N(0,vθ).\sqrt{N}(1-\sigma^{2}/\theta^{2})^{-1}(\lambda_{N}(M_{N})-\rho_{\theta})\stackrel{{\scriptstyle\mathcal{L}}}{{\longrightarrow}}\mu\ast\mathcal{N}(0,v_{\theta}).

In particular, one derives the analogue of (7) for the Deformed GOE:

Let ANA_{N} be an arbitrary deterministic symmetric matrix of rank 1 having a nonnull eigenvalue θ\theta such that θ>σ\theta>\sigma. Then the largest eigenvalue of the Deformed GOE fluctuates as

Obviously, thanks to the orthogonal invariance of the GOE, this result is a direct consequence of Theorem 2.2.

It is worth noticing that, according to the Cramér–Lévy theorem (cf. Fel, Theorem 1, page 525), the limiting distribution (15) is not Gaussian if μ\mu is not Gaussian. Thus, (15) depends on the particular law μ\mu of the entries of the Wigner matrix WNW_{N} which implies the nonuniversality of the fluctuations of the largest eigenvalue of rank-1 diagonal deformation of symmetric or Hermitian Wigner matrices (as conjectured in Remark 1.7 of FePe).

The latter also shows that in the non-Gaussian setting, the fluctuations of the largest eigenvalue depend, not only on the spectrum of the deformation ANA_{N}, but also on the particular definition of the matrix ANA_{N}. Indeed, in collaboration with S. Péché, the third author of the present article has recently stated in FePe the universality of the fluctuations of some Deformed Wigner models under a full deformation ANA_{N} defined by (AN)ij=θ/N(A_{N})_{ij}={\theta}/{N} for all 1≤i,j≤N1\leq i,j\leq N (see also FK). Before giving some details on this work, we have to specify that FePe considered Deformed models such that the entries of the Wigner matrix WNW_{N} have sub-Gaussian moments. Nevertheless, thanks to the analysis made in Ru, one can observe that the assumptions of FePe can be reduced and that it is, for example, sufficient to assume that the Wi,jW_{i,j}’s have moments of any order. Thus, the conclusions of FePe apply to the setting considered in our paper. The main result of FePe establishes the universality of the fluctuations of the largest eigenvalue of the complex Deformed model MNM_{N} associated to a full deformation ANA_{N} and for any value of the parameter θ\theta. In particular, when θ>σ\theta>\sigma, it is proved therein the universality of the Gaussian fluctuations (7). The approach of FePe is mainly based on a combinatorial method inspired by the work So (which handles the non-Deformed Wigner model) and some results of Pe on the Deformed GUE. The combinatorial arguments of FePe also work (with minor modifications) in the real framework and yield the universality of the fluctuations if θ<σ\theta<\sigma. In the case where θ>σ\theta>\sigma which is of particular interest here, the analysis made in FePe reduces the universality problem in the real setting to the knowledge of the particular Deformed GOE model (this remark is also valid in the case where θ=σ\theta=\sigma). Here, we will prove the needed results on the Deformed GOE which, thanks to the analysis of FePe and Ru, allow us to claim the following universality.

Then the largest eigenvalue of the Deformed model MNM_{N} has the Gaussian fluctuations (16).

To be complete, let us notice that the previous result still holds when we allow the distribution ν\nu of the diagonal entries of WNW_{N} to be different from μ\mu provided that ν\nu is symmetric and has moments of any order.

and the Stieltjes transform of the expectation of the empirical measure of the eigenvalues of MNM_{N} by

the Stieltjes transform Note that in some papers to which we make reference, the Stieltjes transform is defined with the opposite sign. of a variable ss with semicircular distribution μsc\mu_{sc}.

Theorem 2.1 is the analogue of the main statement of BS3 established in the context of general spiked population models. The conclusion of BS3 requires numerous results obtained previously by Silverstein and co-authors in CS, BS1 and BS2 (a summary of all this literature can be found in Bai, pages 671–675). From very clever and tedious manipulations of some Stieltjes transforms and the use of the matricial representation (1), these works highlight a very close link between the spectra of the Wishart matrices and the covariance matrix (for quite general covariance matrix which includes the spiked population model). Our approach mimics the one of BS3. Thus, using the fact that the Deformed Wigner model is the additive analogue of the spiked population model, several arguments can be quite easily adapted here (this point has been explained in Chapter 4 of the Ph.D. thesis Fe). Actually, the main point in the proof consists in establishing that for any ε>0\varepsilon>0, almost surely,

Dealing with the particular diagonal perturbation AN=diag⁡(θ,0,…,0)A_{N}=\operatorname{diag}(\theta,0,\ldots,0) such that θ>σ\theta>\sigma, we obtain the fluctuations of the largest eigenvalue λ1(MN)\lambda_{1}(M_{N}) (Theorem 2.2) by an approach close to the one of Pa and the ideas of BBPbis. The reasoning relies on the writing of the rescaled variable N(λ1(MN)−ρθ)\sqrt{N}(\lambda_{1}(M_{N})-\rho_{\theta}) in terms of the resolvent of a non-Deformed Wigner matrix. Then, to complete the analysis of FePe and justify Theorem 2.4, we focus on the particular Deformed GOE model and improve the previous convergence at the level of Laplace transform.

The paper is organized as follows. In Section 3, we introduce preliminary lemmas which will be of basic use later on. Section 4 is devoted to the proof of Theorem 2.1. We first establish an equation (called master equation or master inequality) satisfied by gNg_{N} up to some correction of order 1N2\frac{1}{N^{2}} (see

Section 4.1). Then we explain how this master equation gives rise to an estimation of type (18) and thus to the inclusion (17) of the spectrum of MNM_{N} in Kσε(θ1,…,θJ)K^{\varepsilon}_{\sigma}(\theta_{1},\ldots,\theta_{J}) (see Sections 4.2 and 4.3). In Section 4.4, we use this inclusion to relate the asymptotic spectra of ANA_{N} and MNM_{N} and then deduce Theorem 2.1. Section 5 deals with the fluctuations results. The proof of Theorem 2.2 is given in Section 5.2; Theorem 2.4 is justified in Section 5.3.

Basic lemmas

Note that even if the result in CD is stated in the Hermitian case, the proof is valid and the result still holds in the symmetric case. Now (19) follows putting g(xij;i≤j):=f(xij+(AN)ij;i≤j)g(x_{ij};i\leq j):=f(x_{ij}+(A_{N})_{ij};i\leq j) in (20) and noticing that the (AN)ij(A_{N})_{ij} are uniformly bounded in i,j,Ni,j,N.

This lemma will be useful to estimate many variances. Now, we recall some useful properties of the resolvent (see KKP; CD).

∥G(z)∥≤∣ℑm(z)∣−1\|G(z)\|\leq|\Im m(z)|^{-1} where ∥⋅∥\|\cdot\| denotes the operator norm.

∣G(z)ij∣≤∣ℑm(z)∣−1|G(z)_{ij}|\leq|\Im m(z)|^{-1} for all i,j=1,…,Ni,j=1,\ldots,N.

The derivative with respect to MM of the resolvent G(z)G(z) satisfies

We just mention that (v) comes readily noticing that the eigenvalues of the normal matrix G(z)G(z) are the 1z−λi(M)\frac{1}{z-\lambda_{i}(M)}, i=1,…,N.i=1,\ldots,N.

We will also need the following estimations on the Stieltjes transform gσg_{\sigma} of the semicircular distribution μsc\mu_{sc}.

For (21), we refer the reader to Section 3.1 of Bai. Equation (25) is a consequence of ℑm(gσ(z))ℑm(z)<0\Im m(g_{\sigma}(z))\Im m(z)<0. Other inequalities derive from (21) and the definition of gσ.g_{\sigma}.

Almost sure convergence of the first extremal eigenvalues

for some kk and some ll and we give the precise majoration in the statements of the theorems or propositions.

Section 4.4 explains how to deduce Theorem 2.1 from the inclusion (17).

The goal of Sections 4.1 and 4.2 is to establish Proposition 4.4 below which is fundamental in the proof of the inclusion (17). Before describing rigorously the different ideas of these two sections, let us help the reader’s intuition by a heuristic understanding of the approach. Assume that we can establish that gN(z)g_{N}(z) satisfied the rough quadratic equation (also called master inequality):

Then, for any suitable zz, divided by gN(z)g_{N}(z) the last approximation would provide us an estimation of ΛN(z)−z\Lambda_{N}(z)-z where ΛN(z)=zσ(gN(z))\Lambda_{N}(z)=z_{\sigma}(g_{N}(z)) with zσ(g)=1g+σ2gz_{\sigma}(g)=\frac{1}{g}+\sigma^{2}g being the inverse function of gσg_{\sigma} (see Lemma 4.4 below). Then, intuitively, a Taylor expansion of gσg_{\sigma} between ΛN(z)\Lambda_{N}(z) and zz would lead to an estimation of the type

This intuitive process may throw light on the expression (48) of Lσ(z)L_{\sigma}(z) in Proposition 4.4 below.

Let us recall the integration by parts formula for the Gaussian distribution.

for any Hermitian matrix HH, or by linearity for H=EjkH=E_{jk}, 1≤j,k≤N1\leq j,k\leq N where EjkE_{jk}, 1≤j,k≤N1\leq j,k\leq N is the canonical basis of the complex space of N×NN\times N matrices.

Now, we consider the normalized sum 1N2∑ij\frac{1}{N^{2}}\sum_{ij} of the previous identities to obtain

Now, it is well known (see CD; HT and Lemma 3.1) that

Thus, in the case where XN=XNGX_{N}=X_{N}^{G} we obtain:

The Stieltjes transform gNg_{N} satisfies the following inequality:

We now explain how to obtain the corresponding (30) in the Wigner case. Since the computations are the same as in CD This paper treats the case of several independent non-Deformed Wigner matrices. and KKP, The authors considered a non-Deformed Wigner matrix in the symmetric real setting. we just give some hints of the proof.

The integration by parts formula for the Gaussian distribution is replaced by the following tool:

We apply this lemma with the function ϕ(ξ)\phi(\xi) given, as before, by ϕ(ξ)=Gij\phi(\xi)=G_{ij} and ξ\xi is now one of the variables ℜe((XN)kl)\Re e((X_{N})_{kl}), ℑm((XN)kl)\Im m((X_{N})_{kl}). Note that, since the above random variables are symmetric, only the odd derivatives in (31) give a nonnull term. Moreover, as we are concerned by estimation of order 1N2\frac{1}{N^{2}} of gNg_{N}, we only need to consider (31) up to the third derivative (see CD). The computation of the first derivative will provide the same term as in the Gaussian case.

We refer to CD or KKP for a detailed study of the third derivative. Using some bounds on GNG_{N} (see Lemma 3.2), we can prove that the only term arising from the third derivative in the master equation, giving a contribution of order 1N\frac{1}{N}, is

In conclusion, the first master equation in the Wigner case reads as follows:

where κ4\kappa_{4} is the fourth cumulant of the distribution μ\mu.

1.2 Estimation of |gN−gσ||g_{N}-g_{\sigma}|

To estimate ∣gN−gσ∣|g_{N}-g_{\sigma}| from (21) and (33), we follow the method initiated in HT and S. We do not develop it here since it follows exactly the lines of Section 3.4 in CD but we briefly recall the main arguments and results which will be useful later on. We define the open connected set

One can prove that for any zz in ON′\mathcal{O}^{\prime}_{N}:

writing (21) at the point ΛN(z)\Lambda_{N}(z), we easily get that

on the nonempty open subset ON′′={z∈ON′,ℑm(z)>2σ}\mathcal{O}^{\prime\prime}_{N}=\{z\in\mathcal{O}^{\prime}_{N},\Im m(z)>\sqrt{2}\sigma\} and then on ON′\mathcal{O}^{\prime}_{N} by the principle of uniqueness of continuation.

this allows us to get an estimation of ∣gN(z)−gσ(z)∣|g_{N}(z)-g_{\sigma}(z)| on ON′\mathcal{O}^{\prime}_{N} and then to deduce:

From now on and until the end of Section 4.1, we denote by γ1,…,γr\gamma_{1},\ldots,\gamma_{r} the nonnull eigenvalues of ANA_{N} (γi=θj\gamma_{i}=\theta_{j} for some j≠j0j\not=j_{0}) in order to simplify the writing. Let UN:=UU_{N}:=U be a unitary matrix such that AN=U∗ΔUA_{N}=U^{*}\Delta U where Δ\Delta is the diagonal matrix with entries Δii=γi, i≤r;Δii=0, i>r\Delta_{ii}=\gamma_{i},\ i\leq r;\Delta_{ii}=0,\ i>r. We set

Our aim is to express hN(z)h_{N}(z) in terms of the Stieltjes transform gN(z)g_{N}(z) for NN large, using the integration by parts formula. Note that since we want an estimation of order O(N−2)O(N^{-2}) in the master inequality (4.1), we only need an estimation of hN(z)h_{N}(z) of order O(N−1)O(N^{-1}). As in the previous subsection, we first write the equation in the Gaussian case and then study the additional term (third derivative) in the Wigner case.

(a) Gaussian case. Apply (29) to Φ(XN)=Gjl\Phi(X_{N})=G_{jl} and H=EilH=E_{il} to get

Expressing GXNGX_{N} in terms of GANGA_{N}, we obtain

Now, we consider the sum ∑i,jUik∗UkjIji\sum_{i,j}U^{*}_{ik}U_{kj}I_{ji}, k=1,…,rk=1,\ldots,r fixed and we denote αk=∑i,jUik∗UkjGji=(UGU∗)kk\alpha_{k}=\sum_{i,j}U^{*}_{ik}U_{kj}G_{ji}=(UGU^{*})_{kk}. Then, we have the following equality, using that UU is unitary:

(b) The general Wigner case. We shall prove that (42) still holds. We now rely on Lemma 4.2 to obtain the analogue of (41):

The term Ai,j,lA_{i,j,l} is a fixed linear combination of the third derivative of Φ:=Gjl\Phi:=G_{jl} with respect to Re(XN)il\mathcal{R}e(X_{N})_{il} (i.e., in the direction eil=Eil+Eli)e_{il}=E_{il}+E_{li}) and ℑm(XN)il\Im m(X_{N})_{il} [i.e., in the direction fil:=−1(Eil−Eli)f_{il}:=\sqrt{-1}(E_{il}-E_{li})]. We do not need to write the exact form of this term since we just want to show that this term will give a contribution of order O(N−1)O(N^{-1}) in the equation for hN(z)h_{N}(z). Let us write the derivative in the direction eile_{il}:

which is the sum of eight terms of the form

where if i2q+1=ii_{2q+1}=i (resp., ll), then i2q+2=li_{2q+2}=l (resp., ii), q=0,1,2q=0,1,2.

F(N)F(N) is the sum of eight terms corresponding to (45). Let us write, for example, the term corresponding to i1=ii_{1}=i, i3=ii_{3}=i, i5=ii_{5}=i:

We give the majoration for the term corresponding to i1=li_{1}=l, i3=li_{3}=l, i5=li_{5}=l:

As in the Gaussian case, we now consider the sum ∑i,jUik∗UkjJji\sum_{i,j}U^{*}_{ik}U_{kj}J_{ji}. From Lemma 4.3 and the bound (using the Cauchy–Schwarz inequality)

we still get (42) and thus (43). More precisely, we proved:

We now study the last term in the master inequality of Theorem 4.1. For the non-Deformed Wigner matrices, it is shown in KKP that

Moreover, Proposition 3.2 in CD, in the more general setting of several independent Wigner matrices, gives an estimate of ∣RN(z)−gσ4(z)∣|R_{N}(z)-g_{\sigma}^{4}(z)|. The above convergence holds true in the Deformed case. We just give some hints of the proof of the estimate of ∣RN(z)−gσ4(z)∣|R_{N}(z)-g_{\sigma}^{4}(z)| since the computations are almost the same as in the non-Deformed case. Let us set

For the last term, we apply an integration by parts formula (Lemma 4.2) to obtain (see KKP; CD)

It remains to see that the additional term due to ANA_{N} is of order O(N−1)O(N^{-1}):

We thus obtain (again with the help of a variance estimate)

Then using (39) and since dN(z)d_{N}(z) is bounded we deduce that

We can now give our final master inequality for gN(z)g_{N}(z) following our previous estimates:

where Eσ(z)=∑k=1rγkz−σ2gσ(z)−γk+κ42gσ4(z),{E_{\sigma}(z)=\sum_{k=1}^{r}\frac{\gamma_{k}}{z-\sigma^{2}g_{\sigma}(z)-\gamma_{k}}+\frac{\kappa_{4}}{2}g_{\sigma}^{4}(z)}, κ4\kappa_{4} is the fourth cumulant of the distribution μ\mu.

Note that Eσ(z)E_{\sigma}(z) can be written in terms of the distinct eigenvalues θj\theta_{j} of ANA_{N} as

where ss is a centered semicircular random variable with variance σ2\sigma^{2}.

2 Estimation of |gσ​(z)−gN​(z)+1N​Lσ​(z)||g_{\sigma}(z)-g_{N}(z)+\frac{1}{N}L_{\sigma}(z)|

is roughly the same as the one described in Section 3.6 in CD. Nevertheless we choose to develop it here for the reader’s convenience. We have for any zz in On′\mathcal{O}^{\prime}_{n}, by using (34) and (38),

We get from Theorem 4.2, (35), (39), (4.2), (23),

Finally, using also (36), we get for any zz in On′\mathcal{O}^{\prime}_{n},

Now, for z∉On′z\notin\mathcal{O}^{\prime}_{n}, such that ℑm(z)>0\Im m(z)>0,

Thus, for any zz such that ℑm(z)>0\Im m(z)>0,

Let us denote for a while gN=gNANg_{N}=g_{N}^{A_{N}} and Lσ=LσANL_{\sigma}=L_{\sigma}^{A_{N}}. Note that we get exactly the same estimation (50) dealing with −AN-A_{N} instead of ANA_{N}. Hence since gσ(z)=−gσ(−z)g_{\sigma}(z)=-g_{\sigma}(-z), gN−AN(z)=−gNAN(−z)g_{N}^{-A_{N}}(z)=-g_{N}^{A_{N}}(-z) (using the symmetry assumption on μ\mu) and Lσ−AN(z)=LσAN(−z)L_{\sigma}^{-A_{N}}(z)=L_{\sigma}^{A_{N}}(-z), it readily follows that (50) is also valid for any zz such that ℑm(z)<0\Im m(z)<0. In conclusion:

3 The spectrum of MNM_{N}

The following step now consists of deducing Proposition 4.6 from Proposition 4.4 (from which we will easily deduce the appropriate inclusion of the spectrum of MNM_{N}). Since this transition is based on the inverse Stieltjes transform, we start with establishing the fundamental Proposition 4.5 below concerning the nature of LσL_{\sigma}. To this aim, it will be relevant to rewrite LσL_{\sigma} as

We recall that J+σJ_{+\sigma} (resp., J−σJ_{-\sigma}) denotes the number of jj’s such that θj>σ\theta_{j}>\sigma (resp., θj<−σ\theta_{j}<-\sigma). As in the Introduction, we define

which is >2σ>2\sigma (resp., <−2σ<-2\sigma) when θj>σ\theta_{j}>\sigma (resp., <−σ<-\sigma).

LσL_{\sigma} is the Stieltjes transform of a distribution Λσ\Lambda_{\sigma} with compact support

The proof relies on the following characterization already used in S.

l(z)→0l(z)\rightarrow 0 as ∣z∣→∞,|z|\rightarrow\infty,

The following properties of the Stieltjes transform gσg_{\sigma} will be useful for showing that LσL_{\sigma} fulfills the previous conditions.

The complement of the support of μσ\mu_{\sigma} is characterized as follows:

Now, we are going to show that LσL_{\sigma} satisfies (c1) and (c2) of Theorem 4.3. We have obviously that

Using also (26)–(28), we get readily that for ∣z∣>α|z|>\alpha,

Then, it is clear that ∣Lσ(z)∣→0|L_{\sigma}(z)|\rightarrow 0 when ∣z∣→+∞|z|\rightarrow+\infty and (c1) is satisfied.

Now we follow the approach of S (Lemma 5.5) to prove (c2). Denote by E\mathcal{E} the convex envelope of Kσ(θ1,…,θJ)K_{\sigma}(\theta_{1},\ldots,\theta_{J}) and define the interval

Hence (c2) is satisfied with C=max⁡(C0,C1,C2)C=\max(C_{0},C_{1},C_{2}) and n=7n=7 and Proposition 4.5 follows from Theorem 4.3.

We are now in position to deduce the following proposition from the estimate (51).

For any smooth function φ\varphi with compact support,

Consequently, for φ\varphi smooth, constant outside a compact set and such that supp⁡(φ)∩Kσ(θ1,…,θJ)=∅\operatorname{supp}(\varphi)\cap K_{\sigma}(\theta_{1},\ldots,\theta_{J})=\varnothing,

We refer the reader to the Appendix of CD where it is proved using the ideas of HT that

Dealing with h(z)=N2rN(z)h(z)=N^{2}r_{N}(z), we deduce that

Following the proof of Lemma 5.6 in S, one can show that Λσ(1)=0\Lambda_{\sigma}(1)=0. Then, the rest of the proof of (54) sticks to the proof of Lemma 6.3 in HT (using Lemma 3.1).

Such a method can be carried out in the case of Wigner real symmetric matrices; then the approximate master equation is the following [compare with (4.1)]:

where ss is a centered semicircular variable with variance σ2\sigma^{2}. Hence by similar arguments as in the complex case, one gets the master equation

Let (MN)N(M_{N})_{N} be any real or complex Deformed model satisfying (i) and (ii) in Section 2. Let J+σJ_{+\sigma} (resp., J−σJ_{-\sigma}) be the number of j’s such that θj>σ\theta_{j}>\sigma (resp., θj<−σ\theta_{j}<-\sigma). Then for any ε>0\varepsilon>0, almost surely, there is no eigenvalue of MNM_{N} in

As soon as ϵ>0\epsilon>0 is small enough, the union (4.4) is made of nonempty disjoint intervals.

4 The almost sure convergence result

As announced in the Introduction, Theorem 2.1 is the analogue of the main statement of BS3 established for general spiked population models (1). The previous Theorem 4.4 is the main step of the proof since now, we can adapt the arguments needed for the conclusion of BS3 viewing the Deformed Wigner model (2) as the additive analogue of the spiked population model (1).

Let us consider one of the positive eigenvalues θj\theta_{j} of the ANA_{N}’s. We recall that this implies that λk1+⋯+kj−1+i(AN)=θj\lambda_{k_{1}+\cdots+k_{j-1}+i}(A_{N})=\theta_{j} for all 1≤i≤kj1\leq i\leq k_{j}. We want to show that if θj>σ\theta_{j}>\sigma (i.e., with our notation, if j∈{1,…,J+σ}j\in\{1,\ldots,J_{+\sigma}\}), the corresponding eigenvalues of MNM_{N} almost surely jump above the right endpoint 2σ2\sigma of the semicircle support as

whereas the rest of the asymptotic spectrum of MNM_{N} lies below 2σ2\sigma with

Analogous results hold for the negative eigenvalues θj\theta_{j} [see points (c) and (d) of Theorem 2.1]. To describe the phenomenon, one can say that, when NN is large enough, the (first extremal) eigenvalues of MNM_{N} can be viewed as a “smoothed” deformation of the (first extremal) eigenvalues of ANA_{N}. According to the analysis made in the previous section [Lemma 4.4(b)], we already know that the limits ρθj\rho_{\theta_{j}} are related to the θj\theta_{j}’s through the Stieltjes transform gσg_{\sigma}. More precisely, one has

Our main purpose now is to establish the asymptotic link between the spectra of the matrices MN=XN+ANM_{N}=X_{N}+A_{N} and ANA_{N}.

Intuitively, this link seems rather natural when σ\sigma is close to zero. Indeed, when NN goes to infinity, since the spectrum of XNX_{N} is concentrated in [−2σ,2σ][-2\sigma,2\sigma] [recall (11)], the spectrum of MNM_{N} should be close to the one of ANA_{N} as soon as σ\sigma will be close to zero (in other words, the spectrum of MNM_{N} is, viewed as a deformation of the one of ANA_{N}, continuous in σ\sigma in a neighborhood of zero). Thus given an interval [a,b]⊂c ⁣Kσ(θ1,…,θJ)[a,b]\subset{}{{}^{c}\!K}_{\sigma}(\theta_{1},\ldots,\theta_{J}), the result of Theorem 4.4 saying that [a,b][a,b] does not contain eigenvalues of MNM_{N} should be improved: it should correspond to [a,b][a,b] some interval II close to [a,b][a,b], lying outside the spectrum of ANA_{N} and such that the number of eigenvalues of MNM_{N} in one side of [a,b][a,b] is equal to the one of ANA_{N} in the corresponding side of II. Following BS2, we will say that there is exact separation of eigenvalues of the matrices ANA_{N} and MNM_{N}.

In the following section, we justify that the exact separation phenomenon occurs regardless of the size of σ\sigma. The proof of Theorem 2.1 will then follow from some suitable choices of [a,b][a,b] (see Section 4.4.2).

According to the previous discussion, we need to refine the analysis made on gσg_{\sigma} in order to identify and understand the link between intervals in cKσ(θ1,…,θJ){}^{c}K_{\sigma}(\theta_{1},\ldots,\theta_{J}) and the complement of the spectrum of the ANA_{N}’s. We also need to understand the dependence on σ\sigma. This is the aim of the following important Lemma 4.5.

As before, we denote (recall Lemma 4.4) by zσz_{\sigma} the inverse function of gσg_{\sigma} which is given by

Using Lemma 4.4, one readily sees that the set cKσ(θ1,…,θJ){}^{c}K_{\sigma}(\theta_{1},\ldots,\theta_{J}) can be characterized as follows:

Obviously, one has g=gσ(x)g=g_{\sigma}(x) if x∈c ⁣Kσ(θ1,…,θJ)x\in{}{{}^{c}\!K}_{\sigma}(\theta_{1},\ldots,\theta_{J}).

Let [a,b][a,b] be a compact set contained in cKσ(θ1,…,θJ){}^{c}K_{\sigma}(\theta_{1},\ldots,\theta_{J}). Then:

[1gσ(a),1gσ(b)]⊂(Spect⁡(AN))c{[\frac{1}{g_{\sigma}(a)},\frac{1}{g_{\sigma}(b)}]}\subset(\operatorname{Spect}(A_{N}))^{c}.

For all 0<σ^<σ0<\hat{\sigma}<\sigma, the interval [zσ^(gσ(a)),zσ^(gσ(b))][z_{\hat{\sigma}}(g_{\sigma}(a)),z_{\hat{\sigma}}(g_{\sigma}(b))] is contained in cKσ^(θ1,…,θJ){}^{c}K_{\hat{\sigma}}(\theta_{1},\ldots,\theta_{J}) and zσ^(gσ(b))−zσ^(gσ(a))≥b−az_{\hat{\sigma}}(g_{\sigma}(b))-z_{\hat{\sigma}}(g_{\sigma}(a))\geq b-a.

The function 1/gσ1/{g_{\sigma}} being increasing, (i) readily follows from (56).

Noticing that Gσ⊂Gσ^\mathcal{G}_{{\sigma}}\subset\mathcal{G}_{\hat{\sigma}} for all σ^<σ\hat{\sigma}<\sigma implies (recall also that gσg_{\sigma} decreases on [a,b][a,b]) that [gσ(b),gσ(a)]⊂Gσ^[g_{\sigma}(b),g_{\sigma}(a)]\subset\mathcal{G}_{\hat{\sigma}}. Relation (56) combined with the fact that the function zσ^z_{\hat{\sigma}} is decreasing on [gσ(b),gσ(a)][g_{\sigma}(b),g_{\sigma}(a)] leads to

and the first part of (ii) is stated. Now, we have

The exact separation result can now be stated. Let [a,b][a,b] be an interval contained in cKσ(θ1,…,θJ){}^{c}K_{\sigma}(\theta_{1},\ldots,\theta_{J}). By Theorem 4.4, [a,b][a,b] is outside the spectrum of MNM_{N}. Moreover, from Lemma 4.5(i), there corresponds an interval I=[a′,b′]I=[a^{\prime},b^{\prime}] outside the spectrum of ANA_{N}, that is, there is iN∈{0,…,N}i_{N}\in\{0,\ldots,N\} such that

aa and a′a^{\prime} (resp., bb and b′b^{\prime}) are linked as follows:

We claim that [a,b][a,b] splits the eigenvalues of MNM_{N} exactly as II splits the spectrum of ANA_{N}. In other words:

With iNi_{N} satisfying (\refsep1)(\ref{sep1}), one has

This result is the analogue of the main statement of BS2 (cf. Theorem 1.2 of BS2) established in the spiked population setting (and in fact for quite general sample covariance matrices). Its proof is quite technical and is inspired by the work BS2. It mainly relies on results on eigenvalues of the rescaled Wigner matrix XNX_{N} combined with the following classical result (due to Weyl).

Let B and C be two N×NN\times N Hermitian matrices. For any pair of integers j,kj,k such that 1≤j,k≤N1\leq j,k\leq N and j+k≤N+1j+k\leq N+1, we have

For any pair of integers j,kj,k such that 1≤j,k≤N1\leq j,k\leq N and j+k≥N+1j+k\geq N+1, we have

Note that this lemma is the additive analogue of Lemma 1.1 of BS2 needed for the investigations of the spiked population model.

In particular, Lemma 4.6 gives that λiN+1(MN)≤λiN+1(AN)+λ1(XN)\lambda_{i_{N}+1}(M_{N})\leq\lambda_{i_{N}+1}(A_{N})+\lambda_{1}(X_{N}) and λiN(MN)≥λiN(AN)+λN(XN).\lambda_{i_{N}}(M_{N})\geq\lambda_{i_{N}}(A_{N})+\lambda_{N}(X_{N}). Besides, as both λ1(XN)\lambda_{1}(X_{N}) and −λN(XN)-\lambda_{N}(X_{N}) tend toward 2σ2\sigma as N→∞N\to\infty [this is (11)], the statement of Theorem 4.5 can be quite easily derived when σ\sigma is close enough to zero. To handle the general case, the key idea is that one can reduce to the previous situation by introducing some parameters. More precisely, given L>0L>0 and k≥0k\geq 0, we will introduce the Wigner matrix

be the Deformed Wigner matrix of parameter

The proof will be organized as follows. On the one hand, as σk,L→0\sigma_{k,L}\to 0 when k→∞k\to\infty (for any fixed L>0L>0), we will readily prove that exact separation occurs for the matrices ANA_{N} and MNK,LM_{N}^{K,L} as soon as KK is large enough. On the other hand, we will show that exact separation also occurs for the eigenvalues of MN=MN0,LM_{N}=M_{N}^{0,L} and MNK,LM_{N}^{K,L} choosing LL large enough. This latter point will be established by induction on kk; the underlying idea is that when the parameter LL is large, the matrices MNk,LM_{N}^{k,L} and MNk+1,LM_{N}^{k+1,L} are close to each other and hence split their spectrum in a similar way.

[Proof of Theorem 4.5] With our choice of [a,b][a,b] and the very definition of the spectrum of the ANA_{N}’s, one can consider ϵ′>0\epsilon^{\prime}>0 small enough such that, for all large NN,

Given L>0L>0 and k≥0k\geq 0 (their size will be determined later), we define

where we recall that zσk,L(g)=1/g+σk,L2g.z_{\sigma_{k,L}}(g)={1}/{g}+\sigma_{k,L}^{2}g. Note that for all L>0L>0, one has a0,L=aa_{0,L}=a and b0,L=bb_{0,L}=b.

We first choose the size of LL as follows. We take L0L_{0} large enough such that for all L≥L0L\geq L_{0},

From the very definition of the ak,La_{k,L}’s and bk,Lb_{k,L}’s, one can easily see that bk,L−ak,L≥b−ab_{k,L}-a_{k,L}\geq b-a [using the last point of (ii) in Lemma 4.5] and that this choice of L0L_{0} ensures that, for all L≥L0L\geq L_{0} and for all k≥0k\geq 0,

Now, we fix LL such that L≥L0L\geq L_{0} and we write ak=ak,La_{k}=a_{k,L}, bk=bk,Lb_{k}=b_{k,L} and σk=σk,L\sigma_{k}=\sigma_{k,L}.

We first show that there exists KK large enough such that, for all k≥Kk\geq K, there is exact separation of the eigenvalues of the matrices ANA_{N} and MNk,LM_{N}^{k,L}, that is,

Furthermore, according to (11), the two first extremal eigenvalues of XNX_{N} are such that almost surely and for all NN large enough,

Thus for all kk, almost surely, at least for NN large enough (NN does not depend on kk),

As σk→0\sigma_{k}\to 0 when k→+∞k\to+\infty, there is KK large enough such that for all k≥Kk\geq K,

and then, almost surely, for all NN large enough

Since λN+1(MNk,L)=−λ0(MNk,L)=−∞\lambda_{N+1}(M_{N}^{k,L})=-\lambda_{0}(M_{N}^{k,L})=-\infty, (63) [resp., (64)] is obviously satisfied if iN=Ni_{N}=N (resp., iN=0i_{N}=0). Thus, we have established that for any iN∈{0,…,N}i_{N}\in\{0,\ldots,N\} satisfying (58), (62) holds for all k≥Kk\geq K. In particular,

Now, we shall show that with probability 11: for NN large, [aK,bK][a_{K},b_{K}] and [a,b][a,b] split the eigenvalues of, respectively, MNK,LM_{N}^{K,L} and MNM_{N} having equal amount of eigenvalues to the left sides of the intervals. To this aim, we will proceed by induction on kk and establish that, for all k≥0k\geq 0, [ak,bk][a_{k},b_{k}] and [a,b][a,b] split the eigenvalues of MNk,LM_{N}^{k,L} and MNM_{N} (recall that MN=MN0,LM_{N}=M_{N}^{0,L}) in exactly the same way. To begin, let us consider for all k≥0k\geq 0, the set

This can be done by induction calling, one more time, on Lemma 4.6. By (66), this is true for k=0k=0. Now, let us assume that (67) holds true. One has

But, for NN large enough, 0<−λN(XN)≤3σ0<-\lambda_{N}(X_{N})\leq 3\sigma a.s., so by the condition (60) on LL,

By (61), one readily observes that a^k−ak+1=ak−ak+1+(b−a)/4>0\hat{a}_{k}-a_{k+1}=a_{k}-a_{k+1}+({b-a})/{4}>0 and similarly that b^k−bk+1<0\hat{b}_{k}-b_{k+1}<0. This implies that

As a consequence, (67) holds for all k≥0k\geq 0 and in particular for k=Kk=K. Comparing this with (65), we deduce that jN=iNj_{N}=i_{N} a.s. and

Now, we are in position to prove the main Theorem 2.1.

4.2 Proof of Theorem 2.1

Our reasoning is close to the last Section 4 of BS3. It is enough to establish parts (a) and (b) since the assertions (c) and (d) can then be deduced by taking −MN-M_{N} instead of MNM_{N}.

The proof of (a) is mainly based on successive applications of Theorem 4.5. Fix an integer 1≤j≤J+σ1\leq j\leq J_{+\sigma}, and let us consider for ϵ>0\epsilon>0, the interval [a,b]=[ρθj+ϵ,ρθj−1−ϵ][a,b]=[\rho_{\theta_{j}}+\epsilon,\rho_{\theta_{j-1}}-\epsilon] which is included in the union (4.4) (at least for ϵ\epsilon small enough). We define Kj(−1)=k1+⋯+kj(−1)K_{j(-1)}=k_{1}+\cdots+k_{j(-1)}. We also take θ0:=+∞\theta_{0}:=+\infty and recall the conventions that λ0(MN)=λ0(AN)=+∞\lambda_{0}(M_{N})=\lambda_{0}(A_{N})=+\infty and K0=0K_{0}=0. Since 1/gσ(ρθk)=θk1/g_{\sigma}(\rho_{\theta_{k}})=\theta_{k} for k=j−1k=j-1 and jj and since the function 1/gσ1/g_{\sigma} is continuous and increasing on [a,b][a,b], the compact interval [a,b][a,b] satisfies (58) with iN=Kj−1i_{N}=K_{j-1}. Hence by Theorem 4.5, one has

Similar arguments imply that for all j∈{1,…,J+σ−1}j\in\{1,\ldots,J_{+\sigma}-1\},

As a result, we deduce that for all 1≤j≤J+σ−11\leq j\leq J_{+\sigma}-1,

So, letting ϵ\epsilon go to zero, we obtain (a) for each integer jj of {1,…,J+σ−1}\{1,\ldots,J_{+\sigma}-1\}.

Let us now quickly consider the case where j=J+σj=J_{+\sigma}. Note first that, from the preceding discussion, we still have (for ϵ\epsilon small enough)

Then, using the fact that 1/gσ1/g_{\sigma} increases continuously on ]2σ,+∞[]2\sigma,+\infty[ with 1/\penaltygσ(]2σ,+∞[)= ]σ,+∞[1/\penalty g_{\sigma}(]2\sigma,+\infty[)=\,]\sigma,+\infty[, one can show that once ϵ>0\epsilon>0 is small enough, the compact set [a,b]=[2σ+ϵ,ρθJ+σ−ϵ][a,b]=[2\sigma+\epsilon,\rho_{\theta_{J_{+\sigma}}}-\epsilon] satisfies the assumptions of Theorem 4.5 with iN=KJ+σi_{N}=K_{J_{+\sigma}}. This leads to

Letting ϵ→0\epsilon\to 0, we deduce that (4.4.2) holds for j=J+σj=J_{+\sigma} and the assertion (a) is established. For point (b), the preceding analysis gives that lim sup⁡NλKJ+σ+1(MN)≤2σ\mboxa.s.\limsup_{N}\lambda_{K_{J_{+\sigma}}+1}(M_{N})\leq 2\sigma\mbox{ a.s.} and it remains to prove that

This inequality follows from the fact that the spectral measure of MNM_{N} converges a.s. toward the semicircle law μsc\mu_{sc} which is compactly supported in [−2σ,2σ][-2\sigma,2\sigma]. This completes the proof of Theorem 2.1.

Fluctuations

The (complex or real) Wigner matricial models under consideration are the same as previously [i.e., defined by (i) in Section 2] but now we assume that the perturbation ANA_{N} is diagonal: AN=diag⁡(θ,0,…,0)A_{N}=\operatorname{diag}(\theta,0,\ldots,0) with unique nonnull eigenvalue θ>σ\theta>\sigma. According to the previous section, the a.s. convergence of λ1(MN)\lambda_{1}(M_{N}) toward ρθ\rho_{\theta} is universal in the sense that it does not depend on μ\mu.

In the first part of this section, we will show that the fluctuations of λ1(MN)\lambda_{1}(M_{N}) around this universal limit are not universal any more. Indeed, we are going to prove that N(1−σ2/θ2)−1(λ1(MN)−ρθ)\sqrt{N}(1-{\sigma^{2}}/{\theta^{2}})^{-1}(\lambda_{1}(M_{N})-\rho_{\theta}) converges in distribution toward the convolution of μ\mu and a Gaussian distribution. Hence, the limiting distribution clearly varies with μ\mu and in particular cannot be Gaussian unless μ\mu is Gaussian.

In the second part of this section, we will sharpen the analysis of the particular Deformed GOE model and explain how this gives Theorem 2.4.

Let B=(bij)B=(b_{ij}) be an N×NN\times N Hermitian matrix and YNY_{N} be a vector of size NN which contains i.i.d. standardized entries with bounded fourth moment. Then there is a constant K>0K>0 such that

there exists a constant a>0a>0 (not depending on NN) such that ∥B∥≤a\|B\|\leq a,

1NTr⁡B2\frac{1}{N}\operatorname{Tr}B^{2} converges in probability to a number a2a_{2},

1N∑i=1Nbii2\frac{1}{N}\sum_{i=1}^{N}b_{ii}^{2} converges in probability to a number a12a_{1}^{2}.

Then the random variable (1/N)(YN∗BYN−Tr⁡B)({1}/{\sqrt{N}})(Y_{N}^{*}BY_{N}-\operatorname{Tr}B) converges in distribution to a Gaussian variable with mean zero and variance

where t=4t=4 when y1y_{1} is real and is 22 when y1y_{1} is complex.

This result is in fact a particular case of a more general result of BY2 (Theorems 7.1 and 7.2) which follows from the method of moments. We give an alternative elegant proof by J. Baik and J. Silverstein in the Appendix of the present paper.

Let ff be an analytic function on an open set of the complex plane including [−2σ,2σ][-2\sigma,2\sigma]. If the entries of a general Wigner matrix WN=((WN)ij)1≤i≤j≤NW_{N}=((W_{N})_{ij})_{1\leq i\leq j\leq N} satisfy the conditions:

2 Proof of Theorem 2.2

The approach is the same for the complex and real settings and is close to the one of Pa and the ideas of BBPbis. Let M^N−1\widehat{M}_{N-1} be the N−1×N−1N-1\times N-1 matrix obtained from MNM_{N} removing the first row and the first column. Thus, N/(N−1)M^N−1\sqrt{N/(N-1)}\widehat{M}_{N-1} is a non-Deformed Wigner matrix associated with the measure μ\mu. We denote by λ1(M^N−1)\lambda_{1}(\widehat{M}_{N-1}) [resp., λN−1(M^N−1)\lambda_{N-1}(\widehat{M}_{N-1})] the largest (resp., lowest) eigenvalue of M^N−1\widehat{M}_{N-1}.

Let 0<δ<(ρθ−2σ)/40<\delta<{(\rho_{\theta}-2\sigma)}/{4}. Let us define the event

On ΩN\Omega_{N}, λ1(MN)\lambda_{1}(M_{N}) is not an eigenvalue of M^N−1\widehat{M}_{N-1} and one can write the eigen-equations using the resolvent G^(λ1(MN)):=(λ1(MN)IN−1−M^N−1)−1\widehat{G}(\lambda_{1}(M_{N})):=(\lambda_{1}(M_{N})I_{N-1}-\widehat{M}_{N-1})^{-1} as follows:

Since v1v_{1} is obviously nonequal to zero, one gets from (70)

Moreover, on ΩN\Omega_{N}, ρθ\rho_{\theta} is not an eigenvalue of M^N−1\widehat{M}_{N-1} (recall that ρθ>2σ\rho_{\theta}>2\sigma) and the resolvent G^(ρθ):=(ρθIN−1−M^N−1)−1\widehat{G}(\rho_{\theta}):=(\rho_{\theta}I_{N-1}-\widehat{M}_{N-1})^{-1} is well defined, too. Thus, (71) is equivalent to

Using G^(λ1(MN))−G^(ρθ)=−(λ1(MN)−ρθ)G^(ρθ)G^(λ1(MN))\widehat{G}(\lambda_{1}(M_{N}))-\widehat{G}(\rho_{\theta})=-(\lambda_{1}(M_{N})-\rho_{\theta})\widehat{G}(\rho_{\theta})\widehat{G}(\lambda_{1}(M_{N})) and gσ(ρθ)=1/θ,g_{\sigma}(\rho_{\theta})={1}/{\theta}, one gets (on ΩN\Omega_{N})

defining fθ(z):=1ρθ−z1∣z∣≤2σ+δ{f_{\theta}(z):=\frac{1}{\rho_{\theta}-z}1_{|z|\leq 2\sigma+\delta}}, we can easily deduce from the previous equality the following identity on ΩN\Omega_{N}:

[using Lemma 3.2(v)]. By the law of large numbers 1N∑j=2N∣(WN)j1∣2\frac{1}{N}\sum_{j=2}^{N}|(W_{N})_{j1}|^{2} converges a.s. toward σ2\sigma^{2} and according to Theorem 2.1, ∣λ1(MN)−ρθ∣|\lambda_{1}(M_{N})-\rho_{\theta}| converges a.s. to zero. Hence δ1(N)\delta_{1}(N) converges obviously in probability toward zero.

Now, since fθf_{\theta} is analytic on an open set including [−2σ,2σ][-2\sigma,2\sigma], we deduce from Theorem 5.3 the convergence in probability of δ3(N)\delta_{3}(N) toward zero and of cNc_{N} toward σ2∫fθ2 dμsc=σ2θ2−σ2{\sigma^{2}\int f_{\theta}^{2}\,d\mu_{sc}=\frac{\sigma^{2}}{\theta^{2}-\sigma^{2}}}.

According to Theorem 5.1 and using Lemma 3.2(v),

The convergence in probability of δ2(N)\delta_{2}(N) toward zero readily follows by Chebyshev inequality.

Let us check that G^(ρθ)1∥M^N−1∥≤2σ+δ\widehat{G}(\rho_{\theta})1_{\|\widehat{M}_{N-1}\|\leq 2\sigma+\delta} satisfies the conditions of Theorem 5.2.

∥G^(ρθ)1∥M^N−1∥≤2σ+δ∥≤1ρθ−2σ−δ\|\widehat{G}(\rho_{\theta})1_{\|\widehat{M}_{N-1}\|\leq 2\sigma+\delta}\|\leq\frac{1}{\rho_{\theta}-2\sigma-\delta} by Lemma 3.2(v).

Thus, choosing α\alpha such that 2α(ρθ−2σ−δ)3<ϵ3\frac{2\alpha}{(\rho_{\theta}-2\sigma-\delta)^{3}}<\frac{\epsilon}{3}, we readily deduce the convergence in probability of

Since G^(ρθ)1∥M^N−1∥≤2σ+δ\widehat{G}(\rho_{\theta})1_{\|\widehat{M}_{N-1}\|\leq 2\sigma+\delta} and Mˇ⋅1\check{M}_{\bm{\cdot}1} are independent, we can deduce from Theorem 5.2 that dNd_{N} converges in distribution toward a Gaussian law with mean zero and variance

where t=4t=4 in the real setting and t=2t=2 in the complex one. Note that one readily verifies that vθv_{\theta} satisfies (14) in Section 2.

Let 0<ϵ<10<\epsilon<1. Since δ1(N)+δ2(N)\delta_{1}(N)+\delta_{2}(N) converges in probability toward zero, the probability of the event

tends to 1. Now, since cN≥0c_{N}\geq 0 we have the following identity on Ω~N\widetilde{\Omega}_{N}:

with uN:=1+cN+δ1(N)+δ2(N)u_{N}:=1+c_{N}+\delta_{1}(N)+\delta_{2}(N) converging in distribution toward (1−σ2/θ2)−1(1-{\sigma^{2}}/{\theta^{2}})^{-1}. Moreover, since (WN)11(W_{N})_{11} and dNd_{N} are independent, (WN)11+\penaltyN/(N−1)dN+N/(N−1)δ3(N)(W_{N})_{11}+\penalty\sqrt{N/(N-1)}d_{N}+\sqrt{N/(N-1)}\delta_{3}(N) converges in distribution toward the convolution of μ\mu and a Gaussian distribution N(0,vθ)\mathcal{N}(0,v_{\theta}).

Finally, we can conclude that N(1−σ2/θ2)−1(λ1(MN)−ρθ)\sqrt{N}(1-{\sigma^{2}}/{\theta^{2}})^{-1}(\lambda_{1}(M_{N})-\rho_{\theta}) converges in distribution toward μ∗N(0,vθ)\mu\ast\mathcal{N}(0,v_{\theta}).

3 Proof of Theorem 2.4

As before, θ\theta is assumed to be >σ>\sigma. In Theorem 2.4, we consider the real Deformed models and claim that the full deformation ANA_{N} defined by (AN)ij=θ/N(A_{N})_{ij}={\theta}/{N} exhibits universality of the Gaussian fluctuations of the largest eigenvalue around ρθ\rho_{\theta}. As already stated, the analogue of this result holds in the complex setting. This is one of the conclusions of the work FePe which also partly solves the real case (we recall to the reader that all the results of FePe readily extend to the framework of Theorem 2.4 calling on Ru). In order to explain this more precisely, let us summarize the main arguments developed by FePe in the complex setting. First, it is shown that the universality of the fluctuations follows from the universality of limits of expectations of traces of suitable high powers of any Deformed Wigner matrices (the powers are of the order of N\sqrt{N}). Second (this is the main part of the work FePe), to handle such expectations, the authors perform a combinatorial method inspired by So and then deduce that in the large limit N→∞N\to\infty, the previous expectations behave as in the Gaussian case. The last step of the analysis calls on the investigations of Pe on the Deformed GUE which allow to identify the value of these limits.

Actually, the combinatorial arguments also work in the real setting (see in particular Section 2 in FePe) and reduce the universality problem to the knowledge of the Deformed GOE. Thus, to get the result of Theorem 2.4, it suffices to prove (using the orthogonal invariance of the GOE) the following limit.

Call LθL_{\theta} the Laplace transform of the law N(0,2σθ2)\mathcal{N}(0,2\sigma_{\theta}^{2}). Let MNGM_{N}^{G} be the Deformed GOE with AN=diag⁡(θ,0,…,0)A_{N}=\operatorname{diag}(\theta,0,\ldots,0) and θ>σ\theta>\sigma.

The starting point of our computations is the following result which states that the previous expectation only involves (as N→∞N\to\infty) the rescaled largest eigenvalue of the Deformed GOE

This formula does not appear explicitly in FePe but all the arguments needed for its justification can be found in it (actually one can show that the formula holds for any Deformed Wigner model MNM_{N} satisfying the assumptions of Theorem 2.4). We will not give the proof and refer the reader to Section 2 in FePe.

Hence, to derive Proposition 5.1, it remains to show the next lemma on ξ1G\xi_{1}^{G}.

Observe first that it is enough to show that

where the event Ω~N{\widetilde{\Omega}_{N}} was defined above by (73) choosing δ>0\delta>0 smaller than min⁡{ρθ−2σ4;13∫1ρθ−x dμsc(x)}\min\{\frac{\rho_{\theta}-2\sigma}{4};\frac{1}{3}\int\frac{1}{\rho_{\theta}-x}\,d\mu_{sc}(x)\}. Indeed, by the Cauchy–Schwarz inequality,

The previous right-hand side is negligible as N→∞N\to\infty since the probability vanishes and the expectation is bounded since FePe proved that the left-hand side of (76) is bounded, too.

The occurrence of the event Ω~N{\widetilde{\Omega}_{N}} allows to make use of the relevant representation (74) of ξ1G\xi_{1}^{G} obtained in the previous Section 5.2:

Second, by Fubini’s theorem one can check that

where N\mathcal{N} is a centered Gaussian variable with variance 2σθ22\sigma_{\theta}^{2}. We want to deduce (78) from (5.3) and (81) by the dominated convergence theorem. Thus, we are going to prove that there exists a function hh such that for NN large enough and for any xx,

where E=E′E′′\mathcal{E}=\mathcal{E}^{\prime}\mathcal{E}^{\prime\prime} with

Using the Gaussian assumptions (see Sa, pages 90–91), one has

for large enough NN, where the βi\beta_{i}’s are the eigenvalues of G^(ρθ)1∥M^N−1G∥≤2σ+δ\widehat{G}(\rho_{\theta})1_{\|\widehat{M}^{G}_{N-1}\|\leq 2\sigma+\delta}. Note that 0≤βi<13δ0\leq\beta_{i}<\frac{1}{3\delta} so that the last identities make sense, for instance, for N>16σ4δ2N>\frac{16\sigma^{4}}{\delta^{2}}. Hence,

Let α>12\alpha>\frac{1}{2}; using that for any yy in [0,1−12α][0,1-\frac{1}{2\alpha}], we have −ln⁡(1−y)−y≤αy2-\ln(1-y)-y\leq\alpha y^{2}. So, as βi<13δ\beta_{i}<\frac{1}{3\delta}, we get that for N>16σ4δ−2(1−12α)−2N>{16\sigma^{4}}{\delta^{-2}(1-\frac{1}{2\alpha})^{-2}},

Thus, there is some constant Cα,σ,δC_{\alpha,\sigma,\delta} such that E′E′′≤Cα,σ,δ\mathcal{E}^{\prime}\mathcal{E}^{\prime\prime}\leq C_{\alpha,\sigma,\delta} and

Now, for 0<t<2(1−ε)0<t<2(1-\varepsilon), ∫0+∞exp⁡(x−2x(1−ε)/t) dx<∞\int_{0}^{+\infty}\exp(x-2x{(1-\varepsilon)}/{t})\,dx<\infty. The proof is complete.

Appendix: By J. Baik and J. Silverstein

This Appendix presents the proof by J. Baik and J. Silverstein of the CLT (given by Theorem 5.2) needed in the previous section for the proof of Theorem 2.2. Their proof is based on a writing of the expression

as a sum of martingale differences, and uses the following CLT.

For each NN, let ZN1,…,ZNrNZ_{N1},\ldots,Z_{Nr_{N}} be a real martingale difference sequence with respect to the increasing σ\sigma-field {FN,j}\{\mathcal{F}_{N,j}\} having second moments. If, as N→∞N\to\infty,

where v2v^{2} is a positive constant, and for each ϵ>0\epsilon>0,

[Proof of Theorem 5.2] First, one can write (1) as a sum of martingale differences:

We will show the conditions of Theorem .1 are met.

To verify the Lindeberg condition (3), we need to show this property is closed under addition. This will follow from the following fact. For random variables X1X_{1}, X2X_{2}, and positive ϵ\epsilon,

The same bound starting with X2X_{2} leads to (4).

Write Zi=X1i+X2iZ_{i}=X_{1}^{i}+X_{2}^{i}, with X1i=(1/N)(∣yi∣2−1)biiX_{1}^{i}=(1/\sqrt{N})(|y_{i}|^{2}-1)b_{ii}. Then for ϵ>0\epsilon>0,

as N→∞N\to\infty, by the dominated convergence theorem.

Thus, by (5), (6) and (4), {Zi}\{Z_{i}\} satisfies (3).

Now, we shall verify condition (2). We have

Let BLB_{L} denote the strictly lower triangular part of BB. We have

We apply the following bound (due to Mathias; see Mt): ∥BL∥≤γN∥B∥\|B_{L}\|\leq\gamma_{N}\|B\| where γN=O(ln⁡N)\gamma_{N}=O(\ln N), and the bound ∥B∥≤a\|B\|\leq a to conclude that

where t=4t=4 when y1y_{1} is real, and is 2 when y1y_{1} is complex.

Besides, from Lemma 2.7 in BS1 (recalled in Theorem 5.1) we have

(8) implies that condition (2) holds with

Thus, by Theorem .1, we deduce that (1/N)(YN∗BYN−Tr⁡B)({1}/{\sqrt{N}})(Y_{N}^{*}BY_{N}-\operatorname{Tr}B) converges in distribution to a Gaussian variable with mean zero and variance v2v^{2}.

Acknowledgments

The authors are very grateful to Jack Silverstein and Jinho Baik for providing them their proof of Theorem 5.2 (which is a fundamental argument in the proof of Theorem 2.2) presented in the Appendix of the present article. The authors also wish to thank an anonymous referee for useful comments which led to an improvement of this paper.

References