The singular values and vectors of low rank perturbations of large rectangular random matrices

Florent Benaych-Georges, Raj Rao Nadakuditi

Introduction

In many applications, the n×mn\times m signal-plus-noise data or measurement matrix formed by stacking the mm samples or measurements of n×1n\times 1 observation vectors alongside each other can be modeled as:

where uiu_{i} and viv_{i} are left and right “signal” column vectors, σi\sigma_{i} are the associated “signal” values and XX is the noise-only matrix of random noises. This model is ubiquitous in signal processing , statistics and machine learning and is known under various guises as a signal subspace model , a latent variable statistical model , or a probabilistic PCA model .

Relative to this model, a common application-driven objective is to estimate the signal subspaces Span{u1,…,ur}\{u_{1},\ldots,u_{r}\} and Span{v1,…,vr}\{v_{1},\ldots,v_{r}\} that contain signal energy. This is accomplished by computing the singular value decomposition (in brief SVD) of X~\widetilde{X} and extracting the rr largest singular values and the associated singular vectors of X~\widetilde{X} - these are referred to as the rr principal components and the Eckart-Young-Mirsky theorem states that they provide the best rank-rr approximation of the matrix X~\widetilde{X} for any unitarily invariant norm . This theoretical justification combined with the fact that these vectors can be efficiently computed using now-standard numerical algorithms for the SVD has led to the ubiquity of the SVD in applications such as array processing , genomics , wireless communications , information retrieval to list a few .

In this paper, motivated by emerging high-dimensional statistical applications , we place ourselves in the setting where nn and mm are large and the SVD of X~\widetilde{X} is used to form estimates of {σi}\{\sigma_{i}\}, {ui}i=1r\{u_{i}\}_{i=1}^{r} and {vi}i=1r\{v_{i}\}_{i=1}^{r}. We provide a characterization of the relationship between the estimated extreme singular values of X~\widetilde{X} and the true “signal” singular values σi\sigma_{i} (and also the angle between the estimated and true singular vectors).

In the limit of large matrices, the extreme singular values only depend on integral transforms of the distribution of the singular values of the noise-only matrix XX in (1) and exhibit a phase transition about a critical value: this is a new occurrence of the so-called BBP phase transition, named after the authors of the seminal paper . The critical value also depends on the aforementioned integral transforms which arise from rectangular free probability theory . We also characterize the fluctuations of the singular values about these asymptotic limit. The results obtained are precise in the large matrix limit and, akin to our results in , go beyond answers that might be obtained using matrix perturbation theory .

Our results are in a certain sense very general (in terms of possible distributions for the noise model XX) and recover as a special case results found in the literature for the eigenvalues and eigenvectors of X~X~∗\widetilde{X}\widetilde{X}^{*} in the setting where XX in (1) is Gaussian. For the Gaussian setting we provide new results for the right singular vectors. Such results had already been proved in the particular case where XX is a Gaussian matrix, but our approach brings to light a general principle, which can be applied beyond the Gaussian case. Roughly speaking, this principle says that for XX a n×pn\times p matrix (with n,p≫1n,p\gg 1), if one adds an independent small rank perturbation ∑i=1rσiuivi∗\sum_{i=1}^{r}\sigma_{i}u_{i}v_{i}^{*} to XX, then the extreme singular values will move to positions which are approximately the solutions zz of the equations

In the case where these equations have no solutions (which means that the θi\theta_{i}’s are below a certain threshold), then the extreme singular values of XX will not move significantly. We also provide similar results for the associated left and right singular vectors and give limit theorems for the fluctuations. These expressions provide the basis for the parameter estimation algorithm developed by Hachem et al in .

The papers were devoted to the analogue problem for the eigenvalues of finite rank perturbations of Hermitian matrices. We follow the strategy developed in these papers for our proofs: we derive master equation representations that implicitly encode the relationship between the singular values and singular vectors of XX and X~\widetilde{X} and use concentration results to obtain the stated analytical expressions. Of course, because of these similarities in the proofs, we chose to focus, in the present paper, in what differs from .

At a certain level, our proof also present analogies with the ones of other papers devoted to other occurrences of the BBP phase transition, such as . We mention that the approach of the paper could also be used to consider large deviations of the extreme singular values of X~\widetilde{X}.

This paper is organized as follows. We state our main results in Section 2 and provide some examples in Section 3. The proofs are provided in Sections 4-7 with some technical details relegated to the appendix in Section 8.

Main results

Let XnX_{n} be a n×mn\times m real or complex random matrix. Throughout this paper we assume that n≤mn\leq m so that we may simplify the exposition of the proofs. We may do so without loss of generality because in the setting where n>mn>m, the expressions derived will hold for Xn∗X_{n}^{*}. Let the n≤mn\leq m singular values of XnX_{n} be σ1≥σ2≥…≥σn\sigma_{1}\geq\sigma_{2}\geq\ldots\geq\sigma_{n}. Let μXn\mu_{X_{n}} be the empirical singular value distribution, i.e., the probability measure defined as

Let mm depend on nn – we denote this dependence explicitly by mnm_{n} which we will sometimes omit for brevity by substituting mm for mnm_{n}. Assume that as n⟶∞n\longrightarrow\infty, n/mn⟶c∈n/m_{n}\longrightarrow c\in. In the following, we shall need some of the following hypotheses.

The probability measure μXn\mu_{X_{n}} converges almost surely weakly to a non-random compactly supported probability measure μX\mu_{X}.

Examples of random matrices satisfying this hypothesis can be found in e.g. . Note however that the question of isolated extreme singular values is not addressed in papers like (where moreover the perturbation has a non bounded rank).

Let aa be infimum of the support of μX\mu_{X}. The smallest singular value of XnX_{n} converges almost surely to aa.

Let bb be supremum of the support of μX\mu_{X}. The largest singular value of XnX_{n} converges almost surely to bb.

Examples of random matrices satisfying the above hypotheses can be found in e.g. .

In this problem, we shall consider the extreme singular values and the associated singular vectors of X~n\widetilde{X}_{n}, which is the random n×mn\times m matrix:

where PnP_{n} is defined as described below.

Setting uiu_{i} and viv_{i} to equal the ii-th column of 1nGu(n)\frac{1}{\sqrt{n}}G_{u}^{(n)} and 1mGv(n)\frac{1}{\sqrt{m}}G^{(n)}_{v} respectively or,

Setting uiu_{i} and viv_{i} to equal to the vectors obtained from a Gram-Schmidt (or QR factorization) of Gu(n)G_{u}^{(n)} and Gv(n)G^{(n)}_{v} respectively.

In the orthonormalized model, the θi\theta_{i}’s are the non zero singular values of PnP_{n} and the uiu_{i}’s and the viv_{i}’s are the left and right associated singular vectors.

We make the following hypothesis on the law ν\nu of the entries of Gu(n)G_{u}^{(n)} and Gv(n)G_{v}^{(n)} (see [3, Sect. 2.3.2] for the definition of log-Sobolev inequalities).

The probability measure ν\nu has mean zero, variance one and that satisfies a log-Sobolev inequality.

We also note if ν\nu is the standard real or complex Gaussian distribution, then the singular vectors produced using the orthonormalized model will have uniform distribution on the set of rr orthogonal random vectors.

If XnX_{n} is random but has a bi-unitarily invariant distribution and PnP_{n} is non-random with rank rr, then we are in same setting as the orthonormalized model for the results that follow. More generally, our idea in defining both of our models (the i.i.d. one and the orthonormalized one) was to show that if PnP_{n} is chosen independently from XnX_{n} in a somehow “isotropic way” (i.e. via a distribution which is not faraway from being invariant by the action of the orthogonal group by conjugation), then a BBP phase transition occurs, which is governed by a certain integral transform of the limit empirical singular values distribution of XnX_{n}, namely μX\mu_{X}.

We note that there is small albeit non-zero probability that rr i.i.d. copies of a random vector are not linearly independent. Consequently, there is a small albeit non-zero probability that the rr vectors obtained as in (2) via the Gram-Schmidt orthogonalization may not be well defined. However, in the limit of large matrices, this process produces well-defined vectors with overwhelming probability (indeed, by Proposition 8.2, the determinant of the associated r×rr\times r Gram matrix tends to one). This is implicitly assumed in what follows.

2. Notation

we also let ⟶a.s.\overset{\textrm{a.s.}}{\longrightarrow} denote almost sure convergence. The (ordered) singular values of an n×mn\times m Hermitian matrix MM will be denoted by σ1(M)≥⋯≥σn(M)\sigma_{1}(M)\geq\cdots\geq\sigma_{n}(M). Lastly, for a subspace FF of a Euclidian space EE and a unit vector x∈Ex\in E, we denote the norm of the orthogonal projection of xx onto FF by ⟨x,F⟩\langle x,F\rangle.

3. Largest singular values and singular vectors phase transition

In Theorems 2.8, 2.9 and 2.10, we suppose Assumptions 2.1, 2.3 and 2.4 to hold.

We define θ‾\overline{\theta}, the threshold of the phase transition, by the formula

with the convention that (+∞)−1/2=0(+\infty)^{-1/2}=0, and where DμXD_{\mu_{X}}, the DD-transform of μX\mu_{X} is the function, depending on cc, defined by

In the theorems below, DμX−1(⋅)D_{\mu_{X}}^{-1}(\cdot) will denote its functional inverse on [b,+∞)[b,+\infty).

The rr largest singular values of the n×mn\times m perturbed matrix X~n\widetilde{X}_{n} exhibit the following behavior as n,mn→∞n,m_{n}\to\infty and n/mn→cn/m_{n}\to c. We have that for each fixed 1≤i≤r1\leq i\leq r,

Moreover, for each fixed i>ri>r, we have that σi(X~n)⟶a.s.b\sigma_{i}(\widetilde{X}_{n})\overset{\textrm{a.s.}}{\longrightarrow}b.

Consider indices i0∈{1,…,r}{i_{0}}\in\{1,\ldots,r\} such that θi0>θ‾\theta_{i_{0}}>\overline{\theta} . For each nn, define σ~i0=σi0(X~n)\widetilde{\sigma}_{i_{0}}=\sigma_{i_{0}}(\widetilde{X}_{n}) and let u~\widetilde{u} and v~\widetilde{v} be left and right unit singular vectors of X~n\widetilde{X}_{n} associated with the singular value σ~i0\widetilde{\sigma}_{i_{0}}. Then we have, as n⟶∞n\longrightarrow\infty,

When r=1r=1, let the sole singular value of PnP_{n} be denoted by θ\theta. Suppose that

For each nn, let u~\widetilde{u} and v~\widetilde{v} denote, respectively, left and right unit singular vectors of X~n\widetilde{X}_{n} associated with its largest singular value. Then

The following proposition allows to assert that in many classical matrix models, the threshold θ‾\overline{\theta} of the above phase transitions is positive. The proof relies on a straightforward computation which we omit.

Assume that the limiting singular distribution μX\mu_{X} has a density fμXf_{\mu_{X}} with a power decay at bb, i.e., that, as t→bt\to b with t<bt<b, fμX(t)∼M(b−t)αf_{\mu_{X}}(t)\sim M(b-t)^{\alpha} for some exponent α>−1\alpha>-1 and some constant MM. Then:

so that the phase transitions in Theorems 2.8 and 2.10 manifest for α=1/2\alpha=1/2.

Under additional hypotheses on the manner in which the empirical singular distribution of Xn⟶a.s.μXX_{n}\overset{\textrm{a.s.}}{\longrightarrow}\mu_{X} as n⟶∞n\longrightarrow\infty, Theorem 2.10 can be generalized to any singular value with limit bb such that DμX′(ρ)D_{\mu_{X}}^{\prime}(\rho) is infinite. The specific hypothesis has to do with requiring the spacings between the singular values of XnX_{n} to be more “random matrix like” and exhibit repulsion instead of being “independent sample like” with possible clumping. We plan to develop this line of inquiry in a separate paper.

4. Smallest singular values and vectors for square matrices

We now consider the phase transition exhibited by the smallest singular values and vectors. We restrict ourselves to the setting where X~n\widetilde{X}_{n} is a square matrix; this restriction is necessary because the non-monotonicity of the function DμXD_{\mu_{X}} on [0,a)[0,a) when c=lim⁡n/m<1c=\lim n/m<1, poses some technical difficulties that do not arise in the square setting. Moreover, in Theorems 2.13, 2.14 and 2.15, we suppose Assumptions 2.1, 2.2 and 2.4 to hold.

We define θ‾\underline{\theta}, the threshold of the phase transition, by the formula

When a>0a>0 and m=nm=n, the rr smallest singular values of X~n\widetilde{X}_{n} exhibit the following behavior. We have that for each fixed 1≤i≤r1\leq i\leq r,

Moreover, for each fixed i>ri>r, we have that σn+1−i(X~n)⟶a.s.a\sigma_{n+1-i}(\widetilde{X}_{n})\overset{\textrm{a.s.}}{\longrightarrow}a.

Consider indices i0∈{1,…,r}{i_{0}}\in\{1,\ldots,r\} such that θi0>θ‾\theta_{i_{0}}>\underline{\theta}. For each nn, define σ~i0=σn+1−i0(X~n)\widetilde{\sigma}_{i_{0}}=\sigma_{n+1-i_{0}}(\widetilde{X}_{n}) and let u~\widetilde{u} and v~\widetilde{v} be left and right unit singular vectors of X~n\widetilde{X}_{n} associated with the singular value σ~i0\widetilde{\sigma}_{i_{0}}. Then we have, as n⟶∞n\longrightarrow\infty,

When r=1r=1 and m=nm=n, let the smallest singular value of X~n\widetilde{X}_{n} be denoted by σ~n\widetilde{\sigma}_{n} with u~\widetilde{u} and v~\widetilde{v} representing associated left and right unit singular vectors respectively. Suppose that

The analogue of Remark 2.12 also applies here.

5. The D𝐷D-transform in free probability theory

is the analogue of the logarithm of the Fourier transform for the rectangular free convolution with ratio cc (see for an introduction to the theory of rectangular free convolution) in the sense described next.

Let AnA_{n} and BnB_{n} be independent n×mn\times m rectangular random matrices that are invariant, in law, by conjugation by any orthogonal (or unitary) matrix. Suppose that, as n,m→∞n,m\to\infty with n/m→cn/m\to c, the empirical singular values distributions μAn\mu_{A_{n}} and μBn\mu_{B_{n}} of AnA_{n} and BnB_{n} satisfy μAn⟶μA\mu_{A_{n}}\longrightarrow\mu_{A} and μBn⟶μB\mu_{B_{n}}\longrightarrow\mu_{B}. Then by , the empirical singular values distribution μAn+Bn\mu_{A_{n}+B_{n}} of An+BnA_{n}+B_{n} satisfies μAn+Bn⟶μA⊞cμB\mu_{A_{n}+B_{n}}\longrightarrow\mu_{A}\boxplus_{c}\mu_{B}, where μA⊞cμB\mu_{A}\boxplus_{c}\mu_{B} is a probability measure which can be characterized in terms of the CC-transform as

The coefficients of the series expansion of U(z)U(z) are the rectangular free cumulants with ratio cc of μ\mu (see for an introduction to the rectangular free cumulants). The connection between free rectangular additive convolution and Dμ−1D_{\mu}^{-1} (via the CC-transform) and the appearance of Dμ−1D_{\mu}^{-1} in Theorem 2.8 could be of independent interest to free probabilists: the emergence of this transform in the study of isolated singular values completes the picture of , where the transforms linearizing additive and multiplicative free convolutions already appeared in similar contexts.

6. Fluctuations of the largest singular value

Assume that the empirical singular value distribution of XnX_{n} converges to μX\mu_{X} faster than 1/n.1/\sqrt{n}. More precisely,

r=1r=1, θ:=θ1>θ‾\theta:=\theta_{1}>\overline{\theta} and

for ρ=DμX−1(1/θ2)\rho=D_{\mu_{X}}^{-1}(1/\theta^{2}) the limit of σ1(X~n)\sigma_{1}(\widetilde{X}_{n}).

We also make the following hypothesis on the law ν\nu (note that it doesn’t contains the fact that ν\nu is symmetric). In fact, wouldn’t it hold, we would still have a limit theorem on the fluctuations of the largest singular value, like in Theorem 3.4 of , but we chose not to develop this case.

Note that we do not ask ν\nu to be symmetric and make no hypothesis about its third moment. The reason is that the main ingredient of the following theorem is Theorem 6.4 of (or Theorem 7.1 of ), where no hypothesis of symmetry or about the third moment is done.

Suppose Assumptions 2.1, 2.3, 2.4, 2.16 and 2.17 to hold. Let σ~1\widetilde{\sigma}_{1} denote the largest singular value of X~n\widetilde{X}_{n}. Then as n⟶∞n\longrightarrow\infty,

where ρ=DμX−1(c,1/θ2)\rho=D_{\mu_{X}}^{-1}(c,1/\theta^{2}) and

with β=1\beta=1 (or 22) when XX is real (or complex) and

with μ~X=cμX+(1−c)δ0\widetilde{\mu}_{X}=c\mu_{X}+(1-c)\delta_{0}.

7. Fluctuations of the smallest singular value of square matrices

When mn=nm_{n}=n so that c=1c=1, assume that:

For all nn, mn=nm_{n}=n, r=1r=1, θ:=θ1>θ‾\theta:=\theta_{1}>\underline{\theta} and

for ρ:=φμX−1(1/θ)\rho:=\varphi_{\mu_{X}}^{-1}(1/\theta) the limit of the smallest singular value of X~n\widetilde{X}_{n}.

Suppose Assumptions 2.1, 2.2, 2.4, 2.19 and 2.17 to hold. Let σ~n\widetilde{\sigma}_{n} denote the smallest singular value of X~n\widetilde{X}_{n}. Then as n⟶∞n\longrightarrow\infty

Examples

Let XnX_{n} be an n×mn\times m real (or complex) matrix with independent, zero mean, normally distributed entries with variance 1/m1/m. It is known that, as n,m⟶∞n,m\longrightarrow\infty with n/m→c∈(0,1]n/m\to c\in(0,1], the spectral measure of the singular values of XnX_{n} converges to the distribution with density

where a=1−ca=1-\sqrt{c} and b=1+cb=1+\sqrt{c} are the end points of the support of μX\mu_{X}. It is known that the extreme eigenvalues converge to the bounds of this support.

Associated with this singular measure, we have, by an application of the result in [10, Sect. 4.1] and Equation (8),

Thus for any n×mn\times m deterministic matrix PnP_{n} with rr non-zero singular values θ1≥⋯≥θr\theta_{1}\geq\cdots\geq\theta_{r} (rr independent of n,mn,m), for any fixed i≥1i\geq 1, by Theorem 2.8, we have

as n⟶∞n\longrightarrow\infty. As far as the i.i.d. model is concerned, this formula allows us to recover some of the results of .

Now, let us turn our attention to the singular vectors. In the setting where r=1r=1, let Pn=θuv∗P_{n}=\theta uv^{*}. Then, by Theorems 2.9 and 2.10, we have

The phase transitions for the eigenvectors of X~n∗X~n\widetilde{X}_{n}^{*}\widetilde{X}_{n} or for the pairs of singular vectors of X~n\widetilde{X}_{n} can be similarly computed to yield the expression:

2. Square Haar unitary matrices

Let XnX_{n} be Haar distributed unitary (or orthogonal) random matrix. All of its singular values are equal to one, so that it has limiting spectral measure

with a=b=1a=b=1 being the end points of the support of μX\mu_{X}.

Associated with this spectral measure, we have (of course, c=1c=1)

Thus for any n×nn\times n, rank rr perturbing matrix PnP_{n} with rr non-zero singular values θ1≥⋯≥θr\theta_{1}\geq\cdots\geq\theta_{r} where neither rr, nor the θi\theta_{i}’s depend on nn, for any fixed i=1,…,ri=1,\ldots,r, by Theorem 2.8 we have

while for any fixed i≥r+1i\geq r+1, both σi(Xn+Pn)\sigma_{i}(X_{n}+P_{n}) and σn+1−i(Xn+Pn)\sigma_{n+1-i}(X_{n}+P_{n}) ⟶a.s.1\overset{\textrm{a.s.}}{\longrightarrow}1.

Proof of Theorems 2.8 and 2.13

The proofs of both theorems are quite similar. As a consequence, we only prove Theorem 2.8.

The sequence of steps described below yields the desired proof (which is very close to the one of Theorem 2.1 of ):

The first, rather trivial, step in the proof of Theorem 2.8 is to use Weyl’s interlacing inequalities to prove that any fixed-rank singular value of X~n\widetilde{X}_{n} which does not tend to a limit >b>b tends to bb.

Then, we utilize Lemma 4.1 below to express the extreme singular values of X~n\widetilde{X}_{n} as the zz’s such that a certain random 2r×2r2r\times 2r matrix Mn(z)M_{n}(z) is singular.

We then exploit convergence properties of certain analytical functions (derived in the appendix) to prove that almost surely, Mn(z)M_{n}(z) converges to a certain deterministic matrix M(z)M(z), uniformly in zz.

We then invoke a continuity lemma (see Lemma 8.1 in the appendix) to claim that almost surely, the zz’s such that Mn(z)M_{n}(z) is singular (i.e. the extreme singular values of X~n\widetilde{X}_{n}) converge to the zz’s such that M(z)M(z) is singular.

We conclude the proof by noting that, for our setting, the zz’s such that M(z)M(z) is singular are precisely the zz’s such that for some i∈{1,…,r}i\in\{1,\ldots,r\}, DμX(z)=1θi2D_{\mu_{X}}(z)=\frac{1}{\theta^{2}_{i}}. Part (ii) of Lemma 8.1 , about the rank of Mn(z)M_{n}(z), will be useful to assert that when the θi\theta_{i}’s are pairwise distinct, the multiplicities of the isolated singular values are all equal to one.

Firstly, up to a conditioning by the σ\sigma-algebra generated by the XnX_{n}’s, one can suppose them to be deterministic and all the randomness supported by the perturbing matrix PnP_{n}.

Secondly, by [33, Th. 3.1.2], one has, for all i≥1i\geq 1,

with the convention σj(Xn)=+∞\sigma_{j}(X_{n})=+\infty for i≤0i\leq 0 and for i>ni>n. By the same proof as in [17, Sect. 6.2.1], it follows that for all i≥1i\geq 1 fixed,

(we insist here on the fact that ii has to be fixed, i.e. not to depend on nn: of course, for i=n/2i=n/2, (13) is not true anymore in general).

Our approach is based on the following lemma, which reduces the problem to the study of 2r×2r2r\times 2r random matrices. Recall that the constants rr, θ1,…,θr\theta_{1},\ldots,\theta_{r}, and the random column vectors (which depend on nn, even though this dependence does not appear in the notation) u1,…,vru_{1},\ldots,v_{r}, v1,…,vrv_{1},\ldots,v_{r} have been introduced in Section 2.1 and that the perturbing matrix PnP_{n} is given by

Recall also that the singular values of XnX_{n} are denoted by σ1≥⋯≥σn\sigma_{1}\geq\cdots\geq\sigma_{n}. Let us define the matrices

The positive singular values of X~n\widetilde{X}_{n} which are not singular values of XnX_{n} are the z∉{σ1,…,σn}z\notin\{\sigma_{1},\ldots,\sigma_{n}\} such that the 2r×2r2r\times 2r matrix

For the sake of completeness, we provide a proof, even though several related results can be found in the literature (see e.g. ).

Proof. Firstly, [32, Th. 7.3.7] states that the non-zero singular values of X~n\widetilde{X}_{n} are the positive eigenvalues of [0X~nX~n∗0]\begin{bmatrix}0&\widetilde{X}_{n}\\ \widetilde{X}_{n}^{*}&0\end{bmatrix}. Secondly, for any z>0z>0 which is not a singular value of XnX_{n}, by [15, Lem. 6.1],

which allows to conclude, since by hypothesis, det⁡(zIn+m−[0XnXn∗0])−1≠0\det\left(zI_{n+m}-\begin{bmatrix}0&X_{n}\\ X_{n}^{*}&0\end{bmatrix}\right)^{-1}\neq 0. □\square

where φμX\varphi_{\mu_{X}} and φμ~X\varphi_{\widetilde{\mu}_{X}} are the functions defined in the statement of Theorem 2.9.

Now, note that once (12) has been established, our result only concerns the number of singular values of X~n\widetilde{X}_{n} in [b+η,+∞)[b+\eta,+\infty) (for any η>0\eta>0), hence can be proved via Lemma 8.1. Indeed, by Hypothesis 2.3, for nn large enough, XnX_{n} has no singular value >b+η>b+\eta, thus numbers >b+η>b+\eta cannot be in the same time singular values of XnX_{n} and X~n\widetilde{X}_{n}.

In the case where the θi\theta_{i}’s are pairwise distinct, Lemma 8.1 allows to conclude the proof of Theorem 2.8. Indeed, Lemma 8.1 says that exactly as much singular values of X~n\widetilde{X}_{n} as predicted by the theorem have limits >b>b and that their limits are exactly the ones predicted by the Theorem. The part of the theorem devoted to singular values tending to bb can then be deduced from (12) and (13).

In the case where the θi\theta_{i}’s are not pairwise distinct, an approximation approach allows to conclude (proceed for example as in Section 6.2.3 of , using [32, Cor. 7.3.8 (b)] instead of [32, Cor. 6.3.8]).

Proof of Theorems 2.9 and 2.14

The proofs of both theorems are quite similar. As a consequence, we only prove Theorem 2.9.

As above, up to a conditioning by the σ\sigma-algebra generated by the XnX_{n}’s, one can suppose them to be deterministic and all the randomness supported by the perturbing matrix PnP_{n}.

Firstly, by the Law of Large Numbers, even in the i.i.d. model, the uiu_{i}’s and the viv_{i}’s are almost surely asymptotically orthonormalized. More specifically, for all i≠ji\neq j,

(the same being true for the viv_{i}’s). As a consequence, it is enough to prove that

Again, the proof is based on a lemma which reduces the problem to the study of the kernel of a random 2r×2r2r\times 2r matrix. The matrices Θ\Theta, UnU_{n} and VmV_{m} are the ones introduced before Lemma 4.1.

Let zz be a singular value of X~n\widetilde{X}_{n} which is not a singular value of XnX_{n} and let (u,v)(u,v) be a corresponding singular pair of unit vectors. Then the column vector

belongs to the kernel of the 2r×2r2r\times 2r matrix Mn(z)M_{n}(z) introduced in Lemma 4.1. Moreover, we have

Proof. The first part of the lemma is easy to verify with the formula Xn∗f(XnXn∗)=f(Xn∗Xn)Xn∗X_{n}^{*}f(X_{n}X_{n}^{*})=f(X_{n}^{*}X_{n})X_{n}^{*} for any function ff defined on [0,+∞)[0,+\infty). For the second part, use the formulas

to establish u=(z2In−XnXn∗)−1(zPnv+XnPn∗u),u=(z^{2}I_{n}-X_{n}X_{n}^{*})^{-1}(zP_{n}v+X_{n}P_{n}^{*}u), and then use the fact that u∗u=1u^{*}u=1. □\square

Let us consider znz_{n}, (u~,v~)(\widetilde{u},\widetilde{v}) as in the statement of Theorem 2.9. Note firstly that for nn large enough, zn>σ1(Xn)z_{n}>\sigma_{1}(X_{n}), hence Lemma 5.1 can be applied, and the vector

belongs to ker⁡Mn(zn)\ker M_{n}(z_{n}). As explained in the proof of Theorem 2.8, the random matrix-valued function Mn(⋅)M_{n}(\cdot) converges almost surely uniformly to the matrix-valued function M(⋅)M(\cdot) introduced in Equation (14). Hence Mn(zn)M_{n}(z_{n}) converges almost surely to M(ρ)M(\rho), and it follows that the orthogonal projection on (ker⁡M(ρ))⊥(\ker M(\rho))^{\perp} of the vector of (20) tends almost surely to zero.

Note that ρ\rho is precisely defined by the relation θi02φμX(ρ)φμ~X(ρ)=1\theta_{i_{0}}^{2}\varphi_{\mu_{X}}(\rho)\varphi_{\widetilde{\mu}_{X}}(\rho)=1. Hence with β:=−θi0φμ~X(ρ)\beta:=-\theta_{i_{0}}\varphi_{\widetilde{\mu}_{X}}(\rho), we have,

and the orthogonal projection of any vector [xy]\begin{bmatrix}x\\ y\end{bmatrix} on (ker⁡M(ρ))⊥(\ker M(\rho))^{\perp} is the vector [x′y′]\begin{bmatrix}x^{\prime}\\ y^{\prime}\end{bmatrix} such that for all ii,

Then, (17) and (18) are direct consequences of the fact that the projection of the vector of (20) on (ker⁡M(ρ))⊥(\ker M(\rho))^{\perp} tends to zero.

Since the limit of znz_{n} is out of the support of μX\mu_{X}, one can apply Proposition 8.2 to assert that both cnc_{n} and dnd_{n} have almost sure limit zero and that in the sums (22) and (23), any term such that i≠ji\neq j tends almost surely to zero. Moreover, by (17), these sums can also be reduced to the terms with index ii such that θi=θi0\theta_{i}=\theta_{i_{0}}. Tu sum up, we have

Now, note that since znz_{n} tends to ρ\rho,

Moreover, by (18), for all ii such that θi=θi0\theta_{i}=\theta_{i_{0}},

allow to recover the RHS of (16) easily. Via (18), one easily deduces (15).

Proof of Theorems 2.10 and 2.15

Again, we shall only prove Theorem 2.10 and suppose the XnX_{n}’s to be non random.

Let us consider the matrix Mn(z)M_{n}(z) introduced in Lemma 4.1. Here, r=1r=1, so one easily gets, for each nn,

Moreover, for bn:=σ1(Xn)b_{n}:=\sigma_{1}(X_{n}) the largest singular value of XnX_{n}, looking carefully at the term in 1z2−bn2\frac{1}{z^{2}-b_{n}^{2}} in det⁡Mn(z)\det M_{n}(z), it appears that with a probability which tends to one as n⟶∞n\longrightarrow\infty, we have

It follows that with a probability which tends to one as n⟶∞n\longrightarrow\infty, the largest singular value σ~1\widetilde{\sigma}_{1} of X~n\widetilde{X}_{n} is >bn>b_{n}.

Then, one concludes using the second part of Lemma 5.1, as in the proof of Theorem 2.3 of .

Proof of Theorems 2.18 and 2.20

We shall only prove Theorem 2.18, because Theorem 2.20 can be proved similarly.

We have supposed that r=1r=1. Let us denote u=u1u=u_{1} and v=v1v=v_{1}. Then we have

Let us fix an arbitrary b∗b^{*} such that b<b∗<ρb<b^{*}<\rho. Theorem 2.8 implies that almost surely, for nn large enough, det⁡[Mn(⋅)]\det[M_{n}(\cdot)] vanishes exactly once in (b∗,∞)(b^{*},\infty). Since moreover, almost surely, for all nn,

we deduce that almost surely, for nn large enough, det⁡[Mn(z)]>0\det[M_{n}(z)]>0 for b∗<z<σ~1b^{*}<z<\widetilde{\sigma}_{1} and det⁡[Mn(z)]<0\det[M_{n}(z)]<0 for σ~1<z\widetilde{\sigma}_{1}<z.

As a consequence, for any real number xx, for nn large enough,

Therefore, we have to understand the limit distributions of the entries of Mn(ρ+xn)M_{n}\left(\rho+\frac{x}{\sqrt{n}}\right). They are given by the following

For any fixed real number xx, as n⟶∞n\longrightarrow\infty, the distribution of

for X,Y,ZX,Y,Z (resp. X,Y,ℜ(Z),ℑ(Z)X,Y,\Re(Z),\Im(Z)) independent standard real Gaussian variables if β=1\beta=1 (resp. if β=2\beta=2) and for c1,c2,dc_{1},c_{2},d some real constants given by the following formulas:

Proof. Let us define zn:=ρ+xnz_{n}:=\rho+\frac{x}{\sqrt{n}}. We have

Let us for example expand the upper left entry of Γn,1,1\Gamma_{n,1,1} of Γn\Gamma_{n}. We have

The third term of the RHS of (28) tends to xφμX′(ρ)x\varphi_{\mu_{X}}^{\prime}(\rho) as n⟶∞n\longrightarrow\infty. By Taylor-Lagrange Formula, there is ξn∈(0,1)\xi_{n}\in(0,1) such that the second one is equal to

hence tends to zero, by Assumptions 2.1 and 2.19. To sum up, we have

Then the “κ4(ν)=0\kappa_{4}(\nu)=0” case of Theorem 6.4 of allows to conclude. □\square

Let us now complete the proof of Theorem 2.18. By the previous lemma, we have

for some random variables Xn,Yn,ZnX_{n},Y_{n},Z_{n} with converging in distribution to the random variables X,Y,ZX,Y,Z of the previous lemma. Using the relation φμX(ρ)φμ~X(ρ)=θ−2\varphi_{\mu_{X}}(\rho)\varphi_{\widetilde{\mu}_{X}}(\rho)=\theta^{-2}, we get

One can easily recover the formula given in Theorem 2.18 for s2s^{2}, using the relation φμX(ρ)φμ~X(ρ)=θ−2\varphi_{\mu_{X}}(\rho)\varphi_{\widetilde{\mu}_{X}}(\rho)=\theta^{-2}.

Appendix

We now state the continuity lemma that we use in the proof of Theorem 2.8. We note that nothing in its hypotheses is random. As hinted earlier, we will invoke it to localize the extreme eigenvalues of X~n\widetilde{X}_{n}.

φi(z)⟶0\varphi_{i}(z)\longrightarrow 0 as ∣z∣⟶∞|z|\longrightarrow\infty.

Let us define the 2r×2r2r\times 2r-matrix-valued function

and denote by z1>⋯>zpz_{1}>\cdots>z_{p} the zz’s in (b,∞)(b,\infty) such that M(z)M(z) is not invertible, where p∈{0,…,r}p\in\{0,\ldots,r\} is the number of θi\theta_{i}’s such that

Let us also consider a sequence 0<bn0<b_{n} with limit bb and, for each nn, a 2r×2r2r\times 2r-matrix-valued function Mn(⋅)M_{n}(\cdot), defined on

which coefficient are analytic functions, such that

there exists pp real sequences zn,1>⋯>zn,pz_{n,1}>\cdots>z_{n,p} converging respectively to z1,…,zpz_{1},\ldots,z_{p} such that for any ε>0\varepsilon>0 small enough, for nn large enough, the zz’s in (b+ε,∞)(b+\varepsilon,\infty) such that Mn(z)M_{n}(z) is not invertible are exactly zn,1,…,zn,pz_{n,1},\ldots,z_{n,p},

for nn large enough, for each ii, Mn(zn,i)M_{n}(z_{n,i}) has rank 2r−12r-1.

Proof. To prove this lemma, we use the formula

in the appropriate place and proceed as the proof of Lemma 6.1 in . □\square

We also need the following proposition. The uiu_{i}’s and the viv_{i}’s are the random column vectors introduced in Section 2.1.

Let, for each nn, AnA_{n}, BnB_{n} be complex n×nn\times n, n×mn\times m matrices which operator norms, with respect to the canonical Hermitian structure, are bounded independently of nn. Then for any η>0\eta>0, there exists C,α>0C,\alpha>0 such that for all nn, for all i,j,k∈{1,…,r}i,j,k\in\{1,\ldots,r\} such that i≠ji\neq j,

Proof. In the i.i.d. model, this result is an obvious consequence of [15, Prop. 6.2]. In the orthonormalized model, one also has to use [15, Prop. 6.2], which states that the uiu_{i}’s (the same holds for the viv_{i}’s) are obtained from the n×rn\times r matrix Gu(n)G^{(n)}_{u} with i.i.d. entries distributed according to ν\nu by the following formula: for all i=1,…,ri=1,\ldots,r,

where W(n)W^{(n)} is a (random) r×rr\times r matrix such that for certain positive constants D,c,κD,c,\kappa, for all ε>0\varepsilon>0 and all nn,

References