Asymptotic Spectra of Matrix-Valued Functions of Independent Random Matrices and Free Probability

F. Götze, H. Kösters, A. Tikhomirov

Introduction

One of the main questions studied in Random Matrix Theory is the asymptotic universality, meaning the dependence on a few global characteristics of the distribution of the matrix entries, of the distribution of spectra of random matrices when their dimension goes to infinity. This holds for the spectra of Hermitian random matrices with independent entries (up to symmetry), first proved by Wigner in 1955 . Another well studied case is that of sample covariance matrices (i.e. W=XX∗\mathbf{W}=\mathbf{X}\mathbf{X}^{*}, where X\mathbf{X} is a matrix with independent entries), first studied in by Marchenko–Pastur. The spectrum of non Hermitian random matrices with independent identically distributed entries is universal as well. The limiting complex spectrum of this Ginibre–Girko Ensemble is the circular law (i.e. the uniform distribution on the unit circle in the complex plane). The universality here was first proved in by Girko. In the last years different models of random matrices which were derived from Wigner and Ginibre–Girko matrices were studied. For instance, in , the universality of the singular value distribution of powers of Ginibre–Girko matrices was shown. In and the universality of the spectrum of products of independent random matrices from the Ginibre–Girko Ensemble was proved. Moreover, more recently, the local properties of the spectrum have also been investigated in the Gaussian case; see e.g. and .

In this paper we describe a general approach to prove the universality of singular value and eigenvalue distributions of matrix-valued functions of independent random matrices. More precisely, we consider random matrices of the form

Furthermore, we introduce a general approach to identify the limiting eigenvalue distribution of the (square) matrix F\mathbf{F}. Our main results here show how to derive the density of the limiting eigenvalue distribution of the matrix F\mathbf{F} from (the SS-transform of) its limiting singular value distribution. This derivation can be divided into two major steps:

In a first step, we derive equations for the Stieltjes transforms g(z,α)g(z,\alpha) of the (symmetrized) singular value distributions of the shifted matrices F−αI\mathbf{F}-\alpha\mathbf{I} via the SS-transform S(z)S(z) of the (symmetrized) singular value distribution of the unshifted matrix F\mathbf{F}. The key system of equations here reads

where w(z,α)w(z,\alpha) is an unknown auxiliary function and R~α(z)\widetilde{R}_{\alpha}(z) is a known function. To derive this system of equations, we use the asymptotic freeness of the matrices

as well as the calculus for RR-transforms and SS-transforms. Furthermore, we show that it is possible take the limit z→0z\to 0 in (1). Since we are working in a quite general framework, the investigation of the existence of this limit as well as its analytic properties require some work.

In a second step, we identify the density ff of the limiting eigenvalue distribution of the random matrix F\mathbf{F} using logarithmic potential theory. The main observation here is that the function ψ(α):=−w(0,α)g(0,α)\psi(\alpha):=-w(0,\alpha)g(0,\alpha) is closely related to the partial derivatives of the logarithmic potential of the limiting eigenvalue distribution. Thus, under regularity assumptions, we obtain the relation

where uu and vv denote the real and imaginary part of α\alpha, respectively.

Let us emphasize that this identification of the limiting eigenvalue distribution is quite general. In principle, we only need the SS-transform of the limiting singular value distribution and the asymptotic freeness of the matrices in (1.2).

In Section 8 we give several examples for applications of our main universality results (Theorems 3.2 and 4.4). The guiding principle here is (i) to establish universality and (ii) to compute the limits in the Gaussian case, using tools from free probability theory. Here we focus on a special class of matrix-valued functions, namely products of matrices or powers and inverses thereof. Although our framework should, in principle, cover more general functions as well, products of independent matrices represent a convenient class of examples in which the assumptions of our main results can be checked. For instance, the conditions (C0)(C0), (C1)(C1), (C2)(C2) on the large and small singular values can be deduced from existing results by Tao and Vu and Götze and Tikhomirov , here. Moreover, once universality is proved, it suffices to identify the limiting eigenvalue and singular value distributions in the Gaussian case. But if the random matrices X(1),…,X(m)\mathbf{X}^{(1)},\ldots,\mathbf{X}^{(m)} have independent standard Gaussian entries, their distributions are invariant under rotations, and the SS-transforms of the limiting singular value distributions of their products are readily obtained using tools from free probability theory, see e.g. Voiculescu or Hiai and Petz . From here it is possible to obtain the limiting singular value distributions and, as we have seen, the limiting eigenvalue distributions.

Our examples illustrate that our main results provide a unifying framework to derive old and new results for products of independent random matrices. In particular, we determine the limiting singular value and eigenvalue distributions for products of independent random matrices from the so-called spherical ensemble (see e.g. ), i.e. for products of the form X(1)(X(2))−1⋯X(2m−1)(X(2m))−1\mathbf{X}^{(1)}(\mathbf{X}^{(2)})^{-1}\cdots\mathbf{X}^{(2m-1)}(\mathbf{X}^{(2m)})^{-1}, where X(1),…,X(2m)\mathbf{X}^{(1)},\ldots,\mathbf{X}^{(2m)} are independent Girko–Ginibre matrices.

General Framework

We now introduce our main assumptions and notation. Generalizations and specializations will be indicated at the beginnings of later sections.

In order to study the spectral asymptotics of sequences of such matrix tuples we shall make a so-called dimension shape assumption, meaning that nq=nq(n)n_{q}=n_{q}(n), and that for any q=1,…,mq=1,\ldots,m,

Let X=(X(1),…,X(m))\mathbf{X}=(\mathbf{X}^{(1)},\ldots,\mathbf{X}^{(m)}) be an mm-tuple of independent random matrices of dimensions n0×n1,…,nm−1×nmn_{0}\times n_{1},\ldots,n_{m-1}\times n_{m}, respectively, with independent entries. More precisely, we assume that

where the Xjk(q)X^{(q)}_{jk} are independent complex random variables such that for all q=1,…,mq=1,\ldots,m and j=1,…,nq−1; k=1,…,nqj=1,\ldots,n_{q-1};\,k=1,\ldots,n_{q}, we have E Xjk(q)=0\mathbf{E}\,X_{jk}^{(q)}=0 and E ∣Xjk(q)∣2=1\mathbf{E}\,|X_{jk}^{(q)}|^{2}=1.

Furthermore, let Y=(Y(1),…,Y(m))\mathbf{Y}=(\mathbf{Y}^{(1)},\ldots,\mathbf{Y}^{(m)}) be an mm-tuple of independent random matrices of dimensions n0×n1,…,nm−1×nmn_{0}\times n_{1},\ldots,n_{m-1}\times n_{m}, respectively, with independent Gaussian entries. More precisely, we assume that

In particular, E Yjk(q)=0\mathbf{E}\,Y_{jk}^{(q)}=0 and E ∣Yjk(q)∣2=1\mathbf{E}\,|Y_{jk}^{(q)}|^{2}=1.

In Section 8, when we determine the limiting singular value and eigenvalue distributions in the Gaussian case, we will impose the stronger assumption that the Yjk(q)Y^{(q)}_{jk} are standard real or complex Gaussian random variables. By Eq. (2), this entails some restrictions on the second moments of the Xjk(q)X^{(q)}_{jk}.

We shall also assume that the random matrices X(1),…,X(m)\mathbf{X}^{(1)},\ldots,\mathbf{X}^{(m)} and Y(1),…,Y(m)\mathbf{Y}^{(1)},\ldots,\mathbf{Y}^{(m)} are defined on the same probability space and that Y(1),…,Y(m)\mathbf{Y}^{(1)},\ldots,\mathbf{Y}^{(m)} are independent of X(1),…,X(m)\mathbf{X}^{(1)},\ldots,\mathbf{X}^{(m)}. Finally, for any mm-tuple Z=(Z(1),…,Z(m))\mathbf{Z}=(\mathbf{Z}^{(1)},\ldots,\mathbf{Z}^{(m)}) in Mn0×n1×⋯×Mnm−1×nm\mathcal{M}_{n_{0}\times n_{1}}\times\cdots\times\mathcal{M}_{n_{m-1}\times n_{m}}, we set

Note that since we are interested in asymptotic singular value and eigenvalue distributions, we are actually dealing with sequences of matrix tuples of increasing dimension. However, the dependence on nn is usually suppressed in our notation.

Throughout this paper, we use the following notation. For a matrix A=(ajk)\mathbf{A}=(a_{jk}) ∈Mn×p\in\mathcal{M}_{n\times p}, we write ∥A∥\|\mathbf{A}\| for the operator norm of A\mathbf{A} and ∥A∥2:=(∑j=1n∑k=1p∣ajk∣2)1/2\|\mathbf{A}\|_{2}:=(\sum_{j=1}^{n}\sum_{k=1}^{p}|a_{jk}|^{2})^{1/2} for the Frobenius norm of A\mathbf{A}. The singular values of A\mathbf{A} are the square-roots of the eigenvalues of the n×nn\times n matrix AA∗\mathbf{A}\mathbf{A}^{*}. Finally, unless otherwise indicated, CC and cc denote sufficiently large and small positive constants, respectively, which may change from step to step.

Universality of Singular Value Distributions of Functions of Independent Random Matrices

We shall assume that the random variables Xjk(q)X_{jk}^{(q)}, for q=1,…,mq=1,\ldots,m, j=1,…,nq−1; k=1,…,nqj=1,\ldots,n_{q-1};\,k=1,\ldots,n_{q}, satisfy the following Lindeberg condition, i. e.

We now define truncated matrices. Note that by (3.1) there exists a sequence (τn)(\tau_{n}) such that

Clearly, we may additionally require that τn≥n−1/3\tau_{n}\geq n^{-1/3} for all nn. We fix such a sequence and consider the matrix tuple X^=(X^(1),…,X^(m))\widehat{\mathbf{X}}=(\widehat{\mathbf{X}}^{(1)},\ldots,\widehat{\mathbf{X}}^{(m)}) consisting of the matrices X^(q)=(1nqX^jk(q))\widehat{\mathbf{X}}^{(q)}=(\tfrac{1}{\sqrt{n_{q}}}\widehat{X}_{jk}^{(q)}), q=1,…,mq=1,\ldots,m, where

Let B\mathbf{B} be a non-random matrix of order n×pn\times p, let FX\mathbf{F}_{\mathbf{X}} and FX^\mathbf{F}_{\widehat{\mathbf{X}}} be defined as in (2.3), and let s1(X)≥…≥sn(X)s_{1}({\mathbf{X}})\geq\ldots\geq s_{n}({\mathbf{X}}) and s1(X^)≥…≥sn(X^)s_{1}({\widehat{\mathbf{X}}})\geq\ldots\geq s_{n}({\widehat{\mathbf{X}}}) denote the singular values of the matrices FX+B\mathbf{F}_{\mathbf{X}}+\mathbf{B} and FX^+B\mathbf{F}_{\widehat{\mathbf{X}}}+\mathbf{B}, respectively. Let FX(x)\mathcal{F}_{\mathbf{X}}(x) (resp. FX^(x)\mathcal{F}_{\widehat{\mathbf{X}}}(x)) denote the empirical distribution function of the squared singular values of the matrix FX+B\mathbf{F}_{\mathbf{X}}+\mathbf{B} (resp. FX^+B\mathbf{F}_{\widehat{\mathbf{X}}}+\mathbf{B}), i.e.

The corresponding Stieltjes transforms of these empirical distributions are denoted by mX(z)m_{\mathbf{X}}(z) and mX^(z)m_{\widehat{\mathbf{X}}}(z), i.e.

Assume that the conditions (3.1) and (3.2) hold. Then

By the rank inequality of Bai, see , Theorem A.44, we have

Inequalities (3.4)–(3) and assumption (3.3) together complete the proof of the Lemma. ∎

Let Y=(Y(1),…,Y(m))\mathbf{Y}=(\mathbf{Y}^{(1)},\ldots,\mathbf{Y}^{(m)}) be an mm-tuple of independent random matrices with independent Gaussian entries as in Section 1, and let Y^=(Y^(1),…,Y^(m))\widehat{\mathbf{Y}}=(\widehat{\mathbf{Y}}^{(1)},\ldots,\widehat{\mathbf{Y}}^{(m)}) denote the mm-tuple consisting of the matrices Y^(q)=(1nqY^jk(q))\widehat{\mathbf{Y}}^{(q)}=(\tfrac{1}{\sqrt{n_{q}}}\widehat{Y}^{(q)}_{jk}), where

for q=1,…,mq=1,\ldots,m and j=1,…,nq−1; k=1,…,nqj=1,\ldots,n_{q-1};\,k=1,\ldots,n_{q}. Set

Then, using the relation τn≥n−1/3\tau_{n}\geq n^{-1/3} and the special properties of the Gaussian distribution, it is easy to check that we also have

Furthermore, note that for the truncated random variables, the moment identities (2) need not hold anymore. However, we have the relations

as well as the analogous relations for the r.v.’s Y^jk(q)\widehat{Y}_{jk}^{(q)}, which imply that

For the rest of this section, we use the following notation. For any matrix tuple X=(X(1),…,X(m))\mathbf{X}=(\mathbf{X}^{(1)},\ldots,\mathbf{X}^{(m)}) and any n×pn\times p matrix B\mathbf{B}, we introduce the matrix FX\mathbf{F}_{\mathbf{X}} as in (2.3), the Hermitian matrix

as well as the corresponding resolvent matrix

Furthermore, let s1(X)≥…≥sn(X)s_{1}({\mathbf{X}})\geq\ldots\geq s_{n}({\mathbf{X}}) denote the singular values of the matrix FX+B\mathbf{F}_{{\mathbf{X}}}+\mathbf{B}. Note that, apart from a fixed number of zero eigenvalues, the eigenvalues of the matrix VX\mathbf{V}_{\mathbf{X}} are given by ±s1(X),… ±sn(X)\pm s_{1}({\mathbf{X}}),\ldots\,\pm s_{n}({\mathbf{X}}). The corresponding Stieltjes transform will be denoted by

For 0≤φ≤π20\leq\varphi\leq\frac{\pi}{2} and q=1,…,mq=1,\ldots,m, let

and Z(φ):=(Z(1)(φ),…,Z(m)(φ))\mathbf{Z}(\varphi):=(\mathbf{Z}^{(1)}(\varphi),\ldots,\mathbf{Z}^{(m)}(\varphi)). For abbreviation, we shall write F(φ)\mathbf{F}(\varphi), V(φ)\mathbf{V}(\varphi), R(φ)\mathbf{R}(\varphi), and mn(z,φ)m_{n}(z,\varphi) instead of FZ(φ)\mathbf{F}_{\mathbf{Z}(\varphi)}, VZ(φ)\mathbf{V}_{\mathbf{Z}(\varphi)}, RZ(φ)\mathbf{R}_{\mathbf{Z}(\varphi)}, and mn(z,Z(φ))m_{n}(z,\mathbf{Z}(\varphi)). With this notation we have FX^=F(0)\mathbf{F}_{\widehat{\mathbf{X}}}=\mathbf{F}(0), FY^=F(π2)\mathbf{F}_{\widehat{\mathbf{Y}}}=\mathbf{F}(\frac{\pi}{2}), mn(z,X^)=mn(z,0)m_{n}(z,\widehat{\mathbf{X}})=m_{n}(z,0) and mn(z,Y^)=mn(z,π2)m_{n}(z,\widehat{\mathbf{Y}})=m_{n}(z,\frac{\pi}{2}). Also, we may write

The representation of type (3.17) has been used for sums of random variables, for instance, in (second relation on page 367). For random matrices (3.17) has been used, for example, by Pastur and Lytova in (see Equation (60)).

Let gjk(q)(θ)g_{jk}^{(q)}(\theta) denote the function obtained from gjk(q)g_{jk}^{(q)} by replacing the indeterminate Zjk(q)Z_{jk}^{(q)} with θZjk(q)\theta Z_{jk}^{(q)}.

Assume that the Lindeberg condition (3.1) and the rank condition (3.2) hold. Furthermore suppose that there exist constants A0>0A_{0}>0, A1>0A_{1}>0 and A2>0A_{2}>0 such that for any random variable θ\theta which is uniformly distributed on the interval $andindependentofther.v.’sand independent of the r.v.’sX_{jk}^{(q)}andandY_{jk}^{(q)}$, the following conditions hold:

Remark. It follows from the conclusion of the theorem and basic properties of the Stieltjes transform that if the singular value distributions of the matrices FY+B\mathbf{F}_{\mathbf{Y}}+\mathbf{B} are weakly convergent in probability to some limit ν\nu, then so are the singular value distributions of the matrices FX+B\mathbf{F}_{\mathbf{X}}+\mathbf{B}. In this sense Theorem 3.2 proves the universality of singular value distributions.

By Lemma 3.1 and the subsequent remark, it is sufficient to prove the claim with mn(z,X^)m_{n}(z,\widehat{\mathbf{X}}) and mn(z,Y^)m_{n}(z,\widehat{\mathbf{Y}}) instead of mn(z,X)m_{n}(z,\mathbf{X}) and mn(z,Y)m_{n}(z,\mathbf{Y}). Furthermore, according to Lemma A.1 in the Appendix, it is enough to prove that

where θ\theta is a random variable which is uniformly distributed on the unit interval, we get

Here E θ\mathbf{E}\,_{\theta} denotes the expectation with respect to the r.v. θ\theta conditioning on all other r.v.’s. Inserting (3) into (3), we get

and Σ7,…,Σ12\Sigma_{7},\ldots,\Sigma_{12} denote similar terms coming from the second line in (3). Since Σ7,…,Σ12\Sigma_{7},\ldots,\Sigma_{12} can be treated in the same way as Σ1,…,Σ6\Sigma_{1},\ldots,\Sigma_{6}, we provide the details for the latter only.

Since gjk(q)(0,0)g_{jk}^{(q)}(0,0) and Xjk(q)X_{jk}^{(q)}, Yjk(q)Y_{jk}^{(q)} are independent, it follows from (3.13) that

Using again that the random variables ∂gjk(q)∂Zjk(q)(0,0)\frac{\partial g_{jk}^{(q)}}{\partial Z_{jk}^{(q)}}(0,0) and Xjk(q)X_{jk}^{(q)}, Yjk(q)Y_{jk}^{(q)} are independent, we get

By (3.14), the last inequality implies that

Similarly, using (3.9) and (3.16), we get

Combining the preceding estimates, we obtain (3.26). Thus, Theorem 3.2 is proved. ∎

Universality of Eigenvalue Distributions of Functions of Independent Random Matrices

We now turn to the eigenvalue distribution of functions of independent random matrices. We use the assumptions and the notation from Section 1, but throughout this section we assume additionally that n=pn=p, so that FX{\mathbf{F}}_{\mathbf{X}} and FY{\mathbf{F}}_{\mathbf{Y}} are square matrices.

Let μ\mu a probability measure on the complex plane with compact support. Define the logarithmic potential of the measure μ\mu as

Let μX\mu_{\mathbf{X}} (resp. μY\mu_{\mathbf{Y}}) denote the empirical spectral measure of the matrix FX\mathbf{F}_{\mathbf{X}} (resp. FY\mathbf{F}_{\mathbf{Y}}), i.e. μX\mu_{\mathbf{X}} (resp. μY\mu_{\mathbf{Y}}) is the uniform distribution on the eigenvalues {λ1(X),…,λn(X)}\{\lambda_{1}(\mathbf{X}),\ldots,\lambda_{n}(\mathbf{X})\} (resp. {λ1(Y),…λn(Y)}\{\lambda_{1}(\mathbf{Y}),\ldots\lambda_{n}(\mathbf{Y})\}) of the matrix FX\mathbf{F}_{\mathbf{X}} (resp. FY\mathbf{F}_{\mathbf{Y}}). Then

If there exists some p>0p>0 such that the quantity

is bounded in probability as n→∞n\to\infty, we say that the matrices FX\mathbf{F}_{\mathbf{X}} satisfy condition (C0)(C0).

holds, we say that the matrices FX\mathbf{F}_{\mathbf{X}} satisfy condition (C1)(C1).

with n1=[n−nδn]+1n_{1}=[n-n\delta_{n}]+1, n2=[n−nγ]n_{2}=[n-n^{\gamma}], we say that the matrices FX\mathbf{F}_{\mathbf{X}} satisfy condition (C2)(C2).

We now prove the universality of eigenvalue distributions.

for any bounded continuous function ff and any ε>0\varepsilon>0.

The proof of Theorem 4.4 is based on the “replacement principle” by Tao and Vu (see , Theorem 2.1) which builds upon a sort of inversion formula for the logarithmic potential that goes back to Girko and that was also investigated by Bai , .

Note that by condition (C1) the determinants in (4.3) are not zero with probability 1−o(1)1-o(1).

Let L(F,G)L(F,G) denote the Lévy distance between two distribution functions FF and GG. Recall that

Let GX(x,α)\mathcal{G}_{\mathbf{X}}(x,\alpha) and GY(x,α)\mathcal{G}_{\mathbf{Y}}(x,\alpha) denote the distribution functions of the singular values of the matrices FX−αI\mathbf{F}_{\mathbf{X}}-\alpha\mathbf{I} and FY−αI\mathbf{F}_{\mathbf{Y}}-\alpha\mathbf{I}, respectively. Let ϰn=L(GX(⋅,α),GY(⋅,α))\varkappa_{n}=L(\mathcal{G}_{\mathbf{X}}(\cdot,\alpha),\mathcal{G}_{\mathbf{Y}}(\cdot,\alpha)). According to Theorem 3.2 (with B=αI\mathbf{B}=\alpha\mathbf{I}), we have

Note that by (4.4) and Markov’s inequality, we have

Furthermore, put ηn=max⁡{(E ϰn)13,(log⁡n)−1}\eta_{n}=\max\{(\mathbf{E}\,\varkappa_{n})^{\frac{1}{3}},(\log n)^{-1}\}, and introduce the event

Let δn:=1/∣log⁡2ηn∣→0\delta_{n}:=1/|\log 2\eta_{n}|\to 0, let n1=[n−nδn]+1n_{1}=[n-n\delta_{n}]+1 and n2=[n−nγ]n_{2}=[n-n^{\gamma}] be defined as in condition (C2), and introduce the event

Note that on the set {sn1(FX−αI)<2ηn}\{s_{n_{1}}(\mathbf{F}_{\mathbf{X}}-\alpha\mathbf{I})<2\eta_{n}\}, we have, for large enough nn,

by condition (C2). Since a similar estimate holds with FY\mathbf{F}_{\mathbf{Y}} instead of FX\mathbf{F}_{\mathbf{X}}, it follows that

Furthermore, on the set Ω3\Omega_{3} we have, for any a∈[ηn,2ηn]a\in[\eta_{n},2\eta_{n}],

Now fix ε>0\varepsilon>0. The preceding inequality implies that

By condition (C1)(C1) and inequalities (4.5) – (4.7),

By definition of Ω0\Omega_{0} and by condition (C2)(C2), we have, for n→∞n\to\infty,

Moreover, using that for any p>0p>0, the function x−plog⁡xx^{-p}\log x is decreasing in xx for x≥e1/px\geq\text{\rm e}^{1/p}, we get, for large enough nn,

By the inequality sk(F−αI)≤sk(F)+∣α∣s_{k}(\mathbf{F}-\alpha\mathbf{I})\leq s_{k}(\mathbf{F})+|\alpha| and condition (C0)(C0), this quantity converges to zero in probability.

Thus, it remains to bound the last but one summand in (4). Recall that a∈[ηn,2ηn]a\in[\eta_{n},2\eta_{n}]. Integrating by parts, we have

Recall that we need to bound this expression for ω∈Ω2∩Ω1\omega\in\Omega_{2}\cap\Omega_{1} only and that ϰn≤ηn\varkappa_{n}\leq\eta_{n} for such ω\omega. By Chebyshev’s inequality, we have

for any M>0M>0. It therefore follows from condition (C0)(C0) that the second term on the r.h.s. of (4) converges to zero in probability. Furthermore, by the definition of ϰn\varkappa_{n}, we have the following bound

Thus, for ω∈Ω2∩Ω1\omega\in\Omega_{2}\cap\Omega_{1} we may find a∈[ηn,2ηn]a\in[\eta_{n},2\eta_{n}] such that

Because ηn12∣log⁡ηn∣→0\eta_{n}^{\frac{1}{2}}|\log\eta_{n}|\to 0 as n→∞n\to\infty, it follows that the first term on the r.h.s of (4) converges to zero in probability. Finally, using inequality (4.12) again, we obtain

It is straightforward to check that for any 0<ε<a<b0<\varepsilon<a<b and any distribution function F(x)F(x) the following inequality

Therefore, for ω∈Ω2∩Ω1\omega\in\Omega_{2}\cap\Omega_{1} we obtain, for large enough nn,

We may now apply the “replacement principle” by Tao and Vu; see , Theorem 2.1. Note that this theorem is based on two assumptions (i) and (ii). Assumption (ii) is just a reformulation of relation (4.3). Assumption (i) is only needed to show that the probability measures μX\mu_{\mathbf{X}} and μY\mu_{\mathbf{Y}} are tight in probability (see Equations (3.3) and (3.4) in ), and may be replaced with our assumption (C0C0). It therefore follows that μX−μY\mu_{\mathbf{X}}-\mu_{\mathbf{Y}} converges weakly to zero in probability, i.e. for any bounded continuous function ff and any ε>0\varepsilon>0, we have

Asymptotic Freeness of Random Matrices

In this section we consider the asymptotic freeness of random matrices with special structure. Before that, we recall the definition of Voiculescu’s asymptotic freeness as well as some basic notions from free probability theory. See also the survey by Speicher and the lecture notes by Voiculescu .

Let (A,φ)(\mathcal{A},\varphi) be a non-commutative probability space.

1)1) Let (Ai)i∈I(\mathcal{A}_{i})_{i\in I} be a family of unital sub-algebras of A\mathcal{A}. The sub-algebras Ai\mathcal{A}_{i} are called free or freely independent, if, for any positive integer kk, φ(a1⋯ak)=0\varphi(a_{1}\cdots a_{k})=0 whenever the following set of conditions holds: aj∈Ai(j)a_{j}\in\mathcal{A}_{i(j)} (with i(j)∈Ii(j)\in I) for all j=1,…,kj=1,\ldots,k, φ(aj)=0\varphi(a_{j})=0 for all j=1,…,kj=1,\ldots,k, and neighbouring elements are from taken different sub-algebras, i.e. i(1)≠i(2), i(2)≠i(3), …, i(k−1)≠i(k)i(1)\neq i(2),\,i(2)\neq i(3),\,\ldots,\,i(k-1)\neq i(k).

2) Let (Ai′)i∈I(\mathcal{A}_{i}^{\prime})_{i\in I} be a family of subsets of A\mathcal{A}. The subsets Ai′\mathcal{A}_{i}^{\prime} are called free or freely independent, if their generated unital sub-algebras are free, i.e. if (Ai)i∈I(\mathcal{A}_{i})_{i\in I} are free, where, for each i∈Ii\in I, Ai\mathcal{A}_{i} is the smallest unital sub-algebra of A\mathcal{A} which contains Ai′\mathcal{A}_{i}^{\prime}.

3) Let (ai)i∈I(a_{i})_{i\in I} be a family of elements from A\mathcal{A}. The elements aia_{i} are called free or freely independent, if the subsets {ai}\{a_{i}\} are free.

Consider two random variables aa and bb which are free. Then the distributions of a+ba+b and abab (in the sense of linear functionals) depend only on the distributions of aa and bb (see e.g. , Chapter 2), and we can make the following definition:

For free random variables aa and bb, the distributions of a+ba+b and abab are called the free additive convolution and the free multiplicative convolution of μa\mu_{a} and μb\mu_{b} and are denoted by

where Ga−1(z)G_{a}^{-1}(z) denotes the inverse of Ga(z)G_{a}(z) w.r.t. composition of functions. Moreover, when φ(a)≠0\varphi(a)\neq 0, define the SS-transform of aa by

Then, for free random variables aa and bb, we have

and, when ϕ(a)≠0\phi(a)\neq 0 and ϕ(b)≠0\phi(b)\neq 0,

where R~a−1(z){\widetilde{R}}_{a}^{-1}(z) denotes the inverse of R~a(z){\widetilde{R}}_{a}(z) w.r.t. composition of functions. See for instance Nica , Equation (21) in Chapter 13. For clarity, let us emphasize that we call R~a(z)\widetilde{R}_{a}(z) what is called Ra(z)R_{a}(z) in . Furthermore, let us note that the argument in requires that φ(a)≠0\varphi(a)\neq 0. However, when φ(a)=0\varphi(a)=0 and φ(a2)≠0\varphi(a^{2})\neq 0, one can use similar arguments as in Rao and Speicher to show that, similarly as for the SS-transform, there exist two branches of R~a−1(z){\widetilde{R}}_{a}^{-1}(z) and that (5.5) continues to hold with an appropriate choice of these branches. We shall always take the branches such that

The above transforms also have extensions to unbounded probability measures. Let us provide the details for the SS-transform; cf. Section 6 in . For a probability measure ν\nu on (0,∞)(0,\infty), define the function

where ψν−1\psi_{\nu}^{-1} denotes the inverse of ψν\psi_{\nu}.

and the families {ai:i∈Ij}\{a_{i}:i\in I_{j}\}, j=1,…,lj=1,\ldots,l, are free.

Here In\mathbf{I}_{n} denotes the identity matrix of dimension n×nn\times n.

This means that relation (5.7) holds if at least one of the l1,l2,…,lkl_{1},l_{2},\ldots,l_{k} is even. Hence, suppose that l1,…,lkl_{1},\ldots,l_{k} are all odd. In this case we may reduce relation (5.7) to

In order to complete the proof, we proceed similarly as in Hiai and Petz . Using the bi-unitary invariance and the singular value decomposition of the matrix Yn\mathbf{Y}_{n}, we may represent the matrix Yn\mathbf{Y}_{n} as UnΔnVn∗\mathbf{U}_{n}\mathbf{\Delta}_{n}\mathbf{V}_{n}^{*}, where Un\mathbf{U}_{n}, Δn\mathbf{\Delta}_{n}, Vn\mathbf{V}_{n} are independent, Un\mathbf{U}_{n} and Vn\mathbf{V}_{n} are random unitary matrices (with Haar distribution), and Δn\mathbf{\Delta}_{n} is a random diagonal matrix whose diagonal elements are the singular values of Yn\mathbf{Y}_{n}, but with random signs (chosen uniformly at random and independently from everything else). Note that

Thus, the non-zero n×nn\times n blocks in the matrix Anj1~J(α)⋯Anjk~J(α)\widetilde{\mathbf{A}_{n}^{j_{1}}}\mathbf{J}(\alpha)\cdots\widetilde{\mathbf{A}_{n}^{j_{k}}}\mathbf{J}(\alpha) in (5.21) are products of the matrices Un(Δn2p−∫x2pdμI)Un∗\mathbf{U}_{n}(\mathbf{\Delta}_{n}^{2p}-\int x^{2p}d\mu\mathbf{I})\mathbf{U}_{n}^{*}, Vn(Δn2p−∫x2pdμI)Vn∗\mathbf{V}_{n}(\mathbf{\Delta}_{n}^{2p}-\int x^{2p}d\mu\mathbf{I})\mathbf{V}_{n}^{*}, UnΔn2p+1Vn∗\mathbf{U}_{n}\mathbf{\Delta}_{n}^{2p+1}\mathbf{V}_{n}^{*}, VnΔn2p+1Un∗\mathbf{V}_{n}\mathbf{\Delta}_{n}^{2p+1}\mathbf{U}_{n}^{*}, as well as certain powers of α\alpha and α‾\overline{\alpha}, such that each Un∗\mathbf{U}_{n}^{*} is followed by a Vn\mathbf{V}_{n}, and each Vn∗\mathbf{V}_{n}^{*} is followed by a Un\mathbf{U}_{n}.

But this implies that the limit in (5.21) is equal to zero. This completes the proof of the asymptotic freeness of An\mathbf{A}_{n} and Bn\mathbf{B}_{n}. ∎

The notion of bi-unitary invariance is relevant for computing the limiting spectral distributions for random matrices with i.i.d. standard complex Gaussian entries. To compute the limiting distributions for random matrices with i.i.d. standard real Gaussian entries we need the notion of bi-orthogonal invariance. According to a side-remark in Hiai and Petz , the results for this case are analogous.

Moreover, using bi-orthogonal invariance, it is even possible to treat random matrices with i.i.d. entries with a common bivariate real Gaussian distribution. (Thus, we may allow for correlations between the real and imaginary parts, for example.) Indeed, suppose that the matrices Y(1),…,Y(q)\mathbf{Y}^{(1)},\ldots,\mathbf{Y}^{(q)} have i.i.d. Gaussian entries such that E Yjk(q)=0\mathbf{E}\,Y_{jk}^{(q)}=0, E ∣Yjk(q)∣2=1\mathbf{E}\,|Y_{jk}^{(q)}|^{2}=1 and

Since the matrices τ1Yj′+iτ2Yj′′\tau_{1}\mathbf{Y}^{\prime}_{j}+i\tau_{2}\mathbf{Y}^{\prime\prime}_{j} have independent complex entries with mean and variance 11, their limiting singular value distribution is well-known by the Marchenko–Pastur theorem; see e.g. Theorem 3.7 in . Also, from independence and bi-orthogonal invariance, it follows that for each q=m−1,…,1q=m-1,\ldots,1, the matrices

are asymptotically free. Thus, the limiting singular value distribution of the product ∏j=1m(τ1Yj′+iτ2Yj′′)\prod_{j=1}^{m}(\tau_{1}\mathbf{Y}^{\prime}_{j}+i\tau_{2}\mathbf{Y}^{\prime\prime}_{j}) may be found by repeated application of Lemma A.2 in the appendix. Along these lines, many of the results in Section 8 may be extended to random matrices with a more general second moment structure as in (5.22).

Stieltjes Transforms of Spectral Limits of Shifted Matrices

In this section, we assume that n=pn=p, i.e. FY\mathbf{F}_{\mathbf{Y}} is a square matrix, and that the matrices Y(q)\mathbf{Y}^{(q)} have independent standard complex Gaussian entries, up to normalization.

where J(α)\mathbf{J}(\alpha) is defined in Equation (6.1) below. More precisely, we will show that, under appropriate conditions, the mean eigenvalue distributions of the matrices V(α)\mathbf{V}(\alpha) converge in moments to probability measures with compact support. Note that this implies the weak convergence of the mean eigenvalue distributions, and hence the weak convergence in probability of the eigenvalue distributions, by the variance estimate from Section A.1.

from the previous section. This matrix has spectral distribution T(α)=12δ+∣α∣+12δ−∣α∣T(\alpha)=\frac{1}{2}\delta_{+|\alpha|}+\frac{1}{2}\delta_{-|\alpha|}, where δa\delta_{a} denotes the unit atom in the point aa. We now calculate the RR-transform of the distribution T(α)T(\alpha). It is straightforward to check that for the distribution T(α)T(\alpha), we have

(Recall that M(z)M(z) and G(z)G(z) have been introduced above Equation (5.1).) From (6.3) it follows that

Here we consider the principal branch of the square root. In order to obtain a function Rα(z)R_{\alpha}(z) that is analytic at zero, we must take the plus sign. Therefore, (5.1) yields

Remark. Similarly, we find that the SS-transform of the distribution T(α)T(\alpha) is given by

We can now state the main result of this section.

Note that the measure μV\mu_{\mathbf{V}} is symmetric with respect to the origin, and recall that in this case we choose the branch of SS-transform SV(z)S_{\mathbf{V}}(z) as in (5.6).

By asymptotic freeness, the matrices V(α)\mathbf{V}(\alpha) converge in moments to the probability measure μV(α):=μV⊞T(α)\mu_{\mathbf{V}(\alpha)}:=\mu_{\mathbf{V}}\boxplus T(\alpha); see the remark below Definition 5.5. Let RV(α)R_{\mathbf{V}(\alpha)} and R~V(α)\widetilde{R}_{\mathbf{V}(\alpha)} denote the RR-transform and R~\widetilde{R}-transform of μV(α)\mu_{\mathbf{V}(\alpha)}, respectively. Then, by the additivity of the R~\widetilde{R}-transform (see Eq. (5.3)), we have

Using the relation (5.5) again, we finally get

Thus, Theorem 6.1 is proved completely. ∎

In the next section, we will consider the system (6.1) with z=iyz=iy, where y≥0y\geq 0 is any non-negative real number. The next results show that this is possible under appropriate conditions.

Let μV\mu_{\mathbf{V}}, μα\mu_{\alpha} and μV(α)\mu_{\mathbf{V}(\alpha)} be defined as above.

For all y>0y>0, g(iy,α)∈i[0,12∣α∣]g(iy,\alpha)\in i[0,\tfrac{1}{2|\alpha|}].

The function z↦R~α(−g(z,α))z\mapsto\widetilde{R}_{\alpha}(-g(z,\alpha)) has an analytic continuation to an open neighborhood UU of the upper imaginary half-axis.

By abuse of notation, we still write R~α(−g(z,α))\widetilde{R}_{\alpha}(-g(z,\alpha)) for the analytic continuation in part (iv), and we then define w=w(z,α)w=w(z,\alpha) as in (6.1).

For z=iyz=iy with y>0y>0, we have the representation

where the branch of the square root is determined by the analytic continuation in part (iv). (Thus, the square root may be positive or negative!) In particular, R~α(−g(iy,α))∈\widetilde{R}_{\alpha}(-g(iy,\alpha))\in.

(i) follows from a straightforward calculation.

For the proof of (iii), note that (6.12) and (ii) imply

For the proof of (iv), recall that, for z=iyz=iy with y>0y>0 large enough,

This also establishes Equation (6.11). Since h(iy)h(iy) takes values in $$ by part (iii), the rest of part (v) follows immediately.

for all y>0y>0, for it then follows by analytic continuation that the second equation in (6.1) holds for all y>0y>0.

By the definition of w(iy,α)w(iy,\alpha), we have

Since iyg(iy,α)∈(−1,0)iyg(iy,\alpha)\in(-1,0) by part (i) and (6.13) and R~α(−g(iy,α))∈\widetilde{R}_{\alpha}(-g(iy,\alpha))\in by part (v), it follows that

for y>0y>0 large enough, and it remains to show (by continuity) that

for all y>0y>0. Suppose by way of contradiction that 1+w(iy0,α)g(iy0,α)=01+w(iy_{0},\alpha)g(iy_{0},\alpha)=0 for some y0>0y_{0}>0. By (6.15), we may assume without loss of generality that y0>0y_{0}>0 is maximal with this property. Then 1+w(iy,α)g(iy,α)∈(0,1)1+w(iy,\alpha)g(iy,\alpha)\in(0,1) for all y>y0y>y_{0}, and by analytic continuation, the second equation in (6.1) holds for all y>y0y>y_{0}. Letting y↓y0y\downarrow y_{0}, we get

since (−x)SV(−x)→0(-x)S_{\mathbf{V}}(-x)\to 0 as x↓0x\downarrow 0. But this is a contradiction to (i). ∎

where S~V(z):=zSV(z)\widetilde{S}_{\mathbf{V}}(z):=zS_{\mathbf{V}}(z) for z∈(−1,0)z\in(-1,0) and S~V(z)\widetilde{S}_{\mathbf{V}}(z) is defined by continuous extension for z∈{−1,0}.z\in\{-1,0\}.

We proceed by contradiction. Suppose that the limit g(0,α):=lim⁡y↓0g(iy,α)g(0,\alpha):=\lim_{y\downarrow 0}g(iy,\alpha) does not exist. Then, by Lemma 6.3 (iii) and continuity, the set of all accumulation points is a non-degenerate closed interval I⊂i[0,12∣α∣]I\subset i[0,\tfrac{1}{2|\alpha|}]. But, as a consequence of Lemma 6.3 (vi), for each accumulation point g~∈I∖{0}\widetilde{g}\in I\setminus\{0\}, we have

where the square-root can be positive or negative. It is easy to see that this implies that

In view of our remark above Theorem 6.1, this means that μV=T(α)\mu_{\mathbf{V}}=T(\alpha), in contradiction to our assumption that μV\mu_{\mathbf{V}} is not a two-point distribution. Thus, the limit g(0,α)g(0,\alpha) exists, and (6.17) holds.

The existence of the limit (wg)(0,α):=lim⁡y↓0(wg)(iy,α)(wg)(0,\alpha):=\lim_{y\downarrow 0}(wg)(iy,\alpha) as well as the relations (6.4) are now simple consequences. It is worth noting here that the sign of the square root 1+4∣α∣2g(iy,α)2\sqrt{1+4|\alpha|^{2}g(iy,\alpha)^{2}} can only change when g(iy,α)=i2∣α∣g(iy,\alpha)=\frac{i}{2|\alpha|}, and hence must be constant for y≈0y\approx 0 when g(0,α)≠i2∣α∣g(0,\alpha)\neq\frac{i}{2|\alpha|}. ∎

A similar argument shows that under the additional assumption that g(iy,α)g(iy,\alpha) is (jointly) continuous in yy and α\alpha, we have

Note that Equation (6.17) has the “trivial” solution g(0,α)=0g(0,\alpha)=0 when the sign of the square-root is negative. The next result gives a sufficient condition for g(0,α)≠0g(0,\alpha)\neq 0.

On the one hand, it is easy to see that there exists a constant C>0C>0 such that

for all sufficiently small r>0r>0. On the other hand, our assumption implies that

Thus we find that the square-root must be negative for all sufficiently small y>0y>0. Using Taylor expansion, it follows that

Recalling that 1+(wg)(iy,α)1+(wg)(iy,\alpha) takes values in $,wemayconcludethatforallsufficientlysmall, we may conclude that for all sufficiently smally>0$,

By (6.1) and (6.18), it follows that for all sufficiently small y>0y>0,

For ∣α∣<12C|\alpha|<\frac{1}{\sqrt{2}C}, this is a contradiction. Consequently, our assumption that g(0,α)=0g(0,\alpha)=0 is wrong in this case, and Lemma 6.6 is proved. ∎

Density of Limiting Spectral Distribution

In this section, we compute the density of the limit distribution of the empirical spectral distributions of the matrices FY\mathbf{F}_{\mathbf{Y}}. Here we assume that n=pn=p, i.e. FY\mathbf{F}_{\mathbf{Y}} is a square matrix, and that the matrices Y(q)\mathbf{Y}^{(q)} have independent standard complex Gaussian entries, up to normalization.

To study the limiting distribution of the eigenvalue distributions of the matrices FY\mathbf{F}_{\mathbf{Y}}, we use the method of hermitization which goes back to Girko . This method may be summarized as follows:

See e.g. Lemma 4.3 in Bordenave and Chafaï . Let us mention here that Conditions (C0)(C0), (C1)(C1) and (C2)(C2) together with the assumption of weak convergence in probability imply that the function log⁡∣ ⋅ ∣\log|\,\cdot\,| is uniformly integrable in probability for the measures νn( ⋅ ,α)\nu_{n}(\,\cdot\,,\alpha) and that the integrals in (7.1) are finite; see also Lemma A.9 in Appendix A.4.

We now describe the density ff of the limiting spectral distribution μF\mu_{\mathbf{F}} in terms of the SS-transform of the measure μV\mu_{\mathbf{V}}. In doing so, we will not use any special properties of random matrices, but only the probability measures μV\mu_{\mathbf{V}} and μF\mu_{\mathbf{F}} and their properties stated below.

We shall additionally make the following assumptions:

where the square-root is the same as in (6.11). Moreover, the function g(iy,α)g(iy,\alpha) admits a continuous extension g(0,α)g(0,\alpha) as y↓0y\downarrow 0.

The following lemma shows that Assumptions 7.2 and 7.3 are satisfied if the probability measure μV\mu_{\mathbf{V}} has compact support or, more generally, sufficiently small tails. Since the proof is rather technical, it is deferred to Appendix A.4.

Assumptions 7.2 and 7.3 hold for probability measures μV\mu_{\mathbf{V}} such that μV([−x,+x]c)=O(x−η)\mu_{\mathbf{V}}([-x,+x]^{c})=\mathcal{O}(x^{-\eta}) (x→∞)(x\to\infty) for some η>0\eta>0.

The logarithmic transform of the measures ν( ⋅ α)\nu(\,\cdot\,\alpha) is defined by

Note that this is exactly the integral on the right-hand side in (7.3). Similarly as above, we regard the function Φ\Phi as a function of the real parameters uu and vv.

where the function g(0,α)g(0,\alpha) and the sign of the square-root are the same as in (6.17).

Clearly, by Assumption 7.3, the expression in the second line can be made arbitrarily small by choosing CC sufficiently large. Also, note that ϰ(C,α)→0\varkappa(C,\alpha)\to 0 as C→∞C\to\infty. Thus, to complete the proof of the lemma, it remains to show that for any C>0C>0,

Let us mention here that the sign of the first square-root is positive for CC large enough, whereas the sign of the second square-root may be positive or negative, as in Lemma 6.4.

To prove (7.7), note that by Assumption 7.2, we have

for any 0<c<C0<c<C. It therefore follows that

Thus, setting c:=∣h∣c:=|h| and letting h→0h\to 0, we obtain

where we have used the fact that the functions ϰ(y,α)\varkappa(y,\alpha) and 1−4∣α∣2ϰ2(y,α)\sqrt{1-4|\alpha|^{2}\varkappa^{2}(y,\alpha)} are bounded and continuous near the point (0,α)(0,\alpha). This completes the proof of (7.7). ∎

Suppose that μV\mu_{\mathbf{V}} and μF\mu_{\mathbf{F}} are as above and that Assumptions 7.2 and 7.3 hold. For x>0x>0 and α≠0\alpha\neq 0, introduce the functions

where w(z,α)w(z,\alpha) is defined as in Theorem 6.1, and their limits

Then the functions ϰ( ⋅ ,α)\varkappa(\,\cdot\,,\alpha) and ψ( ⋅ ,α)\psi(\,\cdot\,,\alpha) are real-valued with values in [0,12α][0,\tfrac{1}{2\alpha}] and $,respectively.Furthermore,set, respectively. Furthermore, set\xi_{\mathbf{V}}(x):=i(-x)S_{\mathbf{V}}(-x)(x\in[0;1]).Then. Then\xi_{\mathbf{V}}\geq 0$, and we have the relation

Alternatively, and more conveniently for applications, we may rewrite Equation (7.8) in the form of two equations:

One nuisance is that the solution to (7.6) is not unique. Indeed, the pair (ϰ,ψ)≡(0,1)(\varkappa,\psi)\equiv(0,1) is always a solution. However, this trivial solution can be excluded using Lemma 6.6. In typical applications, we will proceed as follows: For given ξV\xi_{\mathbf{V}}, solve the system (7.6). Using Lemma 6.6, argue that the solution is unique. This is possible at least in some cases, and notably in all our applications. Then check that the unique solution is continuously differentiable (except on a finite number of rings) and compute ff using (7.10). Finally, check that ff is indeed a probability density.

Note that the measure ν(⋅,α)\nu(\cdot,\alpha) with corresponding Stieltjes transform g(z,α)g(z,\alpha) is symmetric to the origin. Thus, by Lemma 6.3 (iii) and (vi), we have ϰ(x,α)∈[0,12α]\varkappa(x,\alpha)\in[0,\tfrac{1}{2\alpha}] and ψ(x,α)∈\psi(x,\alpha)\in for all x≥0x\geq 0. Also, the limits ϰ(0,α)\varkappa(0,\alpha) and ψ(0,α)\psi(0,\alpha) exist by Lemma 6.4 and the subsequent remark. Furthermore, by our conventions concerning the SS-transform, we have ξV(x)≥0\xi_{\mathbf{V}}(x)\geq 0 for all x∈(0,1)x\in(0,1).

We now rewrite the Equations (6.1) in terms of the real-valued functions ϰ\varkappa, ψ\psi and ξV\xi_{\mathbf{V}}. Using (6.1) with z=ixz=ix, we have

where the sign of the square-root is determined as in (6.11). Letting x↓0x\downarrow 0, we get

where the sign of the square-root is determined by continuous extension. Taking squares in the first equation and rearranging terms, we deduce that

Eliminating ϰ\varkappa from these equations leads to the equivalent equation (7.8).

Suppose additionally that there exists a finite set AA such that for α∉A\alpha\not\in A, the function ψ(α)\psi(\alpha) is continuously differentiable at α\alpha. Let α∉A\alpha\not\in A. By Lemma 7.5 and Equation (7), we have

Since ψ\psi is continuously differentiable w.r.t. uu, it follows that

Note that all functions depend on ∣α∣|\alpha| only, and are therefore symmetric with respect to uu and vv. Thus we also have

Summing (7.15) and (7.16), we get, for ∣α∣∉A|\alpha|\not\in A,

Moreover, it follows from the preceding discussion that on the open set where ∣α∣∉A|\alpha|\not\in A, Φ(α)\Phi(\alpha) is twice continuously differentiable.

In view of relation (7.3), this means that the log-potential UF(α)U_{\mathbf{F}}(\alpha) is twice continuously differentiable on the open set where ∣α∣∉A|\alpha|\not\in A. It therefore follows by a well known result from potential theory (see e.g. , Theorem II.1.3) that the restriction of μF\mu_{\mathbf{F}} to this set is absolutely continuous with Lebesgue density

This completes the proof of Theorem 7.6. ∎

Let us mention that in Chapter II.1 in , it is assumed that the measure μ\mu under consideration is finite and of compact support. However, a closer inspection of the proof shows that the latter assumption may be relaxed; it is sufficient to assume that the function z↦log⁡+∣z∣z\mapsto\log^{+}|z| is integrable w.r.t. μ\mu.

Applications

We consider applications of Theorems 3.2 and 4.4. These applications show that our main results allow old and new results on products of independent random matrices to be derived in a unified way. In doing so, we shall always assume that either all random variables Xjk(q)X_{jk}^{(q)} are real with

or all random variables Xjk(q)X_{jk}^{(q)} are complex with

Clearly, under these assumptions, the corresponding random variables Yjk(q)Y_{jk}^{(q)} as in (2) have standard real or complex Gaussian distributions, respectively, and we may use the results from the preceding sections. Let us note here that although we have stated these results only for the complex case, there exist analogous results for the real case, as mentioned at the end of Section 5. Furthermore, let us emphasize that the assumptions (8.1) and (8.2) are not needed to establish universality, but only to identify the limiting distributions. Finally, let us mention that the assumptions (8.1) and (8.2) can be relaxed a bit; see the Remark at the end of Section 5.

In this section we consider some applications of Theorem 3.2. We start from the simplest case of Marchenko–Pastur law.

Assume that the random variables XjkX_{jk}, j=1,…,nj=1,\ldots,n, k=1,…,pk=1,\ldots,p satisfy the Lindeberg condition (3.1). Then

where G′(x)=(x−a)(b−x)2πxyG^{\prime}(x)=\frac{\sqrt{(x-a)(b-x)}}{2\pi xy} with a=(1−y)2a=(1-\sqrt{y})^{2}, b=(1+y)2b=(1+\sqrt{y})^{2}.

The probability distribution given by the distribution function G(x)G(x) is called the Marchenko–Pastur distribution with parameter yy.

For simplicity we shall consider only the case that the r.v.’s XjkX_{jk} are real. Let Y\mathbf{Y} be a Gaussian matrix as in Section 1. We prove only the universality of the singular value distribution of the matrix X\mathbf{X}, and then suppose that the limiting distribution of the singular values of the Gaussian matrix Y\mathbf{Y} is known.

To apply Theorem 3.2, we check conditions (3.2), (3.20), (3.22), and (3.24). First we note that in our case

Starting from (8.5), it is straightforward to check that condition (3.20) holds with constant A0=2v−2A_{0}=2v^{-2}, condition (3.22) holds with constant A1=8v−3A_{1}=8v^{-3} and condition (3.24) holds with constant A2=48v−4A_{2}=48v^{-4}.

where a=(1−y)2a=(1-\sqrt{y})^{2} and b=(1+y)2b=(1+\sqrt{y})^{2}. Thus Theorem 8.1 is proved. ∎

The SS-transform of the Marchenko–Pastur distribution with parameter y∈(0,1]y\in(0,1], with density defined in (8.6), is given by

Let gy(z)g_{y}(z) denote the Stieltjes transform of the Marchenko–Pastur distribution. It is well-known that

See for instance , Equations (3.1) and (3.9). Using this equation and the formal identity

Solving this equation with respect to zz, we obtain

Let XX be a r.v. with Marchenko-Pastur distribution with parameter y=1y=1, and let FtF_{t} denote the distribution of (X+t)−1(X+t)^{-1}, t≥0t\geq 0. Then Ft→FF_{t}\to F in Kolmogorov distance as t→0t\to 0, and the SS-transform of the limit X−1X^{-1} is given by

The first part follows from the pointwise convergence of the corresponding densities, which are easily calculated using (8.6). For the second part, we provide a formal proof. Recall that SX(z)=1z+1S_{X}(z)=\frac{1}{z+1}. The corresponding Stieltjes transform is gX(z)=12(−1+z−4z)g_{X}(z)=\frac{1}{2}(-1+\sqrt{\frac{z-4}{z}}). Furthermore, we note that formally MX−1(z)=zgX(z)M_{X^{-1}}(z)=zg_{X}(z), where MX−1(z)M_{X^{-1}}(z) denotes the generic moment generating function of the distribution of X−1X^{-1}. This implies

1.2 Product of Independent Rectangular Matrices

Let m≥1m\geq 1 be fixed. Let n0,…,nmn_{0},\ldots,n_{m} denote integers depending on n≥1n\geq 1 such that n0=nn_{0}=n and

Assume that the random variables Xjk(q)X_{jk}^{(q)}, for q=1,…,mq=1,\ldots,m and j=1,…,nq−1j=1,\ldots,n_{q-1}; k=1,…,nqk=1,\ldots,n_{q}, satisfy the Lindeberg condition (3.1). Then

where the Stieltjes transform s(z)=∫−∞∞1x−zdG(m)(x)s(z)=\int_{-\infty}^{\infty}\frac{1}{x-z}dG^{(m)}(x) of the distribution function G(m)(x)G^{(m)}(x) is determined by the equation

For simplicity we shall assume that all r.v.’s Xjk(q)X^{(q)}_{jk} are real. Let Y(1),…,Y(q)\mathbf{Y}^{(1)},\ldots,\mathbf{Y}^{(q)} be Gaussian matrices as in Section 1.

Let Z(1),…,Z(q)\mathbf{Z}^{(1)},\ldots,\mathbf{Z}^{(q)} be defined as in (3.17). Furthermore, introduce the matrices

for j=1,…,nq−1;k=1,…,nqj=1,\ldots,n_{q-1};k=1,\ldots,n_{q}. From here it follows that

Consider for instance the first term in the right hand side of (8.9). We have

Note that for each q=1,…,mq=1,\ldots,m, the vector V1,q−1ej(q)\mathbf{V}_{1,q-1}\mathbf{e}_{j}^{(q)} is independent of the matrix Z(q)\mathbf{Z}^{(q)} due to the block structure of the matrices H(q)\mathbf{H}^{(q)}. From here it follows that

and by Lemma 5.1 in the Appendix of , we have

q=1,…,mq=1,\ldots,m. Combining these estimates, it follows that

Furthermore, it is straightforward to check that

Using equalities (8.12), (8.8), (8.10) and Lemma 5.1 in the Appendix of , we get

Using equalities (8.13), (8.8), (8.10) and Lemma 5.2 in the Appendix of , it is straightforward to prove that

Inequalities (8.11), (8.12), (8.14) imply that conditions (3.20), (3.22), (3.24) hold. Thus, from Theorem 3.2, it follows that the limit distribution of the singular values of the matrices FX\mathbf{F}_{\mathbf{X}} is the same as the limit distribution of the singular values of the matrices FY\mathbf{F}_{\mathbf{Y}}. For the Gaussian case we may prove that the random matrices (∏q=1l−1Y(q))∗(∏q=1l−1Y(q))(\prod_{q=1}^{l-1}\mathbf{Y}^{(q)})^{*}(\prod_{q=1}^{l-1}\mathbf{Y}^{(q)}) and Y(l)Y(l)∗\mathbf{Y}^{(l)}{\mathbf{Y}^{(l)}}^{*} are asymptotically free for any l=1,…,ml=1,\ldots,m. For details see , Lemma 4.1. From here and Lemma A.2 it follows that the SS-transform of the distribution function G(m)(x)G^{(m)}(x) is given by

For details, see Equations (4.9) and (4.13) in . Thus Theorem 8.5 is proved. ∎

Assume that the conditions of Theorem 8.5 hold and y1=⋯=ym=1y_{1}=\cdots=y_{m}=1. Then

where the Stieltjes transform of G(x)G(x) is determined by the equation

1.3 Powers of Random Matrices

Assume that the random variables XjkX_{jk} satisfy the Lindeberg condition (3.1). Then

where G(m)(x)G^{(m)}(x) is defined by its Stieltjes transform s(z)s(z), which satisfies the equation

The numbers MkM_{k} appearing in (8.17) are called Fuss–Catalan numbers.

Again, for simplicity we shall consider only the case that the r.v.’s XjkX_{jk} are real. Let Y\mathbf{Y} be a Gaussian matrix as in Section 1.

where H=[ZOOZ∗]\mathbf{H}=\begin{bmatrix}&\mathbf{Z}&\mathbf{O}&\\ &\mathbf{O}&\mathbf{Z}^{*}&\end{bmatrix} and J=J(−1)\mathbf{J}=\mathbf{J}(-1). Clearly,

Consider for instance the first sum on the right-hand side. Applying Hölder’s inequality, we get

where H(j,k)\mathbf{H}^{(j,k)} is obtained from H\mathbf{H} by replacing the entry ZjkZ_{jk} with zero. Note that, for q≥2q\geq 2,

By the independence of the matrices H(j,k)\mathbf{H}^{(j,k)} and the random variables ZjkZ_{jk}, we have

for some absolute positive constants C1,C2C_{1},C_{2}. Similarly we have

for some positive constant C>0C>0. (Let us note here that the moment conditions here are a bit different from those in , but using (3.9) – (3.12), it is easy to see that the conclusion still holds.) Furthermore,

Thus, we may rewrite the equality (8.19) in the form

All summands on the r.h.s of (8.1.3) may be bounded similarly. For instance,

where the sum is taken over all t,u,v,wt,u,v,w from the set {j,k,j+n,k+n}\{j,k,j+n,k+n\}. Applying Hölder’s inequality, the representation (8.18) and Lemmas 5.2 and 5.4 in , we get

As follows from Theorem 3.2 the limiting singular value distributions of the matrices FX\mathbf{F}_{\mathbf{X}} and FY\mathbf{F}_{\mathbf{Y}} are the same. In the Gaussian case the limit distribution is computed in Section 4 of . ∎

It follows from equation (8.16) that the SS-transform of the distribution G(m)(x)G^{(m)}(x) is given by the formula

See equality (8.15) for y1=⋯=ym=1y_{1}=\cdots=y_{m}=1 as well.

1.4 Product of Powers of Independent Matrices

Assume that the random variables Xjk(q)X_{jk}^{(q)}, for q=1,…,mq=1,\ldots,m and j,k=1,…,nj,k=1,\ldots,n, satisfy the Lindeberg condition (3.1). Then

where the Stieltjes transform s(z)=∫−∞∞1x−zdG(m)(x)s(z)=\int_{-\infty}^{\infty}\frac{1}{x-z}dG^{(m)}(x) of the distribution function G(m)(x)G^{(m)}(x) is determined by the equation

For simplicity we shall assume that the Xjk(q)X_{jk}^{(q)} are real. Let Y(1),…,Y(m)\mathbf{Y}^{(1)},\ldots,\mathbf{Y}^{(m)} denote the corresponding Gaussian matrices. We shall apply Theorem 3.2. First we note that

This implies condition (3.2). Conditions (3.20), (3.22), (3.24) may be checked similarly as in Subsections 8.1.2 and 8.1.3. For more details see .

Theorem 3.2 now implies that the limit distributions of the singular values of the matrices FX\mathbf{F}_{\mathbf{X}} and FY\mathbf{F}_{\mathbf{Y}} are the same. For the Gaussian case we use the asymptotic freeness of the matrices (∏q=lm(Y(l))ml)(∏q=lm(X(l))ml)∗(\prod_{q=l}^{m}(\mathbf{Y}^{(l)})^{m_{l}})(\prod_{q=l}^{m}(\mathbf{X}^{(l)})^{m_{l}})^{*} and ((Y(l−1))ml−1)∗(Y(l−1))ml−1)((\mathbf{Y}^{(l-1)})^{m_{l-1}})^{*}(\mathbf{Y}^{(l-1)})^{m_{l-1}}) for l=2,…,ml=2,\ldots,m. The proof of this claim repeats the proof of Lemma 4.2 in . From here it follows that the SS-transform of the distribution of Gn(x)\mathcal{G}_{n}(x) is given by

This completes the proof of Theorem 8.10. ∎

1.5 Polynomials of Random Matrices

Let Y(1),…,Y(q)\mathbf{Y}^{(1)},\ldots,\mathbf{Y}^{(q)} be Gaussian random matrices as in Section 1, and let WY\mathbf{W}_{\mathbf{Y}} be defined analogously to WX\mathbf{W}_{\mathbf{X}}. Let GX(x)\mathcal{G}_{\mathbf{X}}(x) and GY(x)\mathcal{G}_{\mathbf{Y}}(x) denote the empirical spectral distribution functions of the matrices WX\mathbf{W}_{\mathbf{X}} and WY\mathbf{W}_{\mathbf{Y}}, respectively.

Let E Xjk(q)=0\mathbf{E}\,X_{jk}^{(q)}=0 and E ∣Xjk(q)∣2=1\mathbf{E}\,|X_{jk}^{(q)}|^{2}=1. Assume that the random variables Xjk(q)X_{jk}^{(q)} for q=1,…,mq=1,\ldots,m; j,k=1,…,mj,k=1,\ldots,m satisfy the Lindeberg condition (3.1). Then

There has recently been considerable progress in computing the limiting spectral distributions for polynomials of random matrices; see the approach by Belinschi, Mai and Speicher for self-adjoint polynomials of self-adjoint random matrices. Possibly this approach can also be used to compute the limiting distributions in Theorem 8.11.

1.6 Spherical Ensemble

In this section we consider the so-called spherical ensemble. Assume that the Xjk(q)X^{(q)}_{jk}, q=1,2q=1,2, j,k=1,…,nj,k=1,\ldots,n, are independent random variables as in (8.1) or (8.2). Moreover, assume that the r.v.’s Xjk(2)X_{jk}^{(2)} satisfy the condition

Let F=X(1)(X(2))−1\mathbf{F}=\mathbf{X}^{(1)}(\mathbf{X}^{(2)})^{-1}, where X(1)\mathbf{X}^{(1)} and X(2)\mathbf{X}^{(2)} denote the n×nn\times n matrices with the entries 1nXjk(1)\frac{1}{\sqrt{n}}X_{jk}^{(1)} and 1nXjk(2)\frac{1}{\sqrt{n}}X_{jk}^{(2)}, respectively.

Remark. It is well-known that under Condition (8.23), the matrix X(2)\mathbf{X}^{(2)} is invertible with probability 1+o(1)1+o(1) as n→∞n\to\infty, see e.g. Lemma A.5 in Appendix A.3. Thus, since we are interested in convergence in probability, we may restrict ourselves to the event where X(2)\mathbf{X}^{(2)} is invertible. This will tacitly be assumed in the subsequent proofs.

Let W=FF∗\mathbf{W}=\mathbf{F}\mathbf{F}^{*}, and let Gn(x)\mathcal{G}_{n}(x) denote the empirical spectral distribution function of the matrix W\mathbf{W}. Then we have the following result, cf. Tikhomirov .

Assume that the random variables Xjk(q)X_{jk}^{(q)}, for q=1,2q=1,2 and j,k=1,…,nj,k=1,\ldots,n satisfy the Lindeberg condition (3.1). Also, assume that Condition (8.23) holds. Then

Remark. Note that if ξ\xi has Cauchy density then η=ξ2\eta=\xi^{2} has density p(x)p(x).

In order to apply Theorem 3.2 we need to regularize the inverse matrix (X(2))−1(\mathbf{X}^{(2)})^{-1}. To begin with, note that (X(2))−1=((X(2))∗X(2))−1(X(2))∗=(X(2))∗(X(2)(X(2))∗)−1.(\mathbf{X}^{(2)})^{-1}=((\mathbf{X}^{(2)})^{*}\mathbf{X}^{(2)})^{-1}(\mathbf{X}^{(2)})^{*}=(\mathbf{X}^{(2)})^{*}(\mathbf{X}^{(2)}(\mathbf{X}^{(2)})^{*})^{-1}. We now introduce the following matrices. For any t>0t>0, let

Because the matrix A~t\widetilde{\mathbf{A}}_{t} is positive definite, we have

Let s1≥⋯≥sns_{1}\geq\cdots\geq s_{n} denote the singular values of the matrix X(2)\mathbf{X}^{(2)}. Then the integral in (8.27) may be represented as

Now, by the Marchenko–Pastur theorem (Theorem 8.1), we have, for any fixed t>0t>0,

By Assumption (8.23) and Lemmas A.4 – A.6 in Appendix A.3, the matrix X(2)\mathbf{X}^{(2)} satisfies Conditions (C0)(C0), (C1)(C1) and (C2)(C2). Thus, by the Marchenko–Pastur theorem and Lemma A.9, we also have

Now fix ε>0\varepsilon>0, and take t>0t>0 sufficiently small so that

It then follows from (8.28) – (8.30) that

In view of (8.27), this implies the statement of the lemma. ∎

where now At=((Z(2))∗Z(2)+tI)−1\mathbf{A}_{t}=((\mathbf{Z}^{(2)})^{*}\mathbf{Z}^{(2)}+t\mathbf{I})^{-1}. We have the representation

Consider the function gjk(2)g_{jk}^{(2)} now. Introduce some auxiliary matrices. Let

By definition of At\mathbf{A}_{t}, we have

Now, writing Z(1)=(Z(1)−E Z(1))+E Z(1)\mathbf{Z}^{(1)}=(\mathbf{Z}^{(1)}-\mathbf{E}\,\mathbf{Z}^{(1)})+\mathbf{E}\,\mathbf{Z}^{(1)} and using Eqs. (3.3) and (3.9), it is straightforward to check that, for any constant vector v\mathbf{v}, we have E ∥Z(1)v∥22≤C∥v∥22\mathbf{E}\,\|\mathbf{Z}^{(1)}\mathbf{v}\|_{2}^{2}\leq C\|\mathbf{v}\|_{2}^{2}. Thus, because the matrix Z(1)\mathbf{Z}^{(1)} and the r.v.’s Xjk(2),Yjk(2)X_{jk}^{(2)},Y_{jk}^{(2)} are independent, we obtain

Inserting the bounds (8.36) and (8.37) into (8.1.6), we get

This inequality implies condition (3.20) for q=2q=2. Now we consider the condition (3.22). We have

Using representation (8.31), it is straightforward to check that

Thus condition (3.22) holds for q=1q=1. Consider q=2q=2 now. We have

Using Hölder’s inequality, we may prove that

Using inequality (8.36) and (8.37), we therefore obtain

Finally, using the block structure of the matrices H(1)\mathbf{H}^{(1)} and H(2)\mathbf{H}^{(2)}, it is easy to see that

Relations (8.41), (8.43), (8.45) and (8.46) together imply

Furthermore, relations (8.32), (8.33), (8.34), together imply

This concludes the proof of (3.22) for q=2q=2. The proof of condition (3.24) is similar and hence omitted.

From now on, let At\mathbf{A}_{t} be defined by At=((Y(2))∗Y(2)+tI)−1\mathbf{A}_{t}=((\mathbf{Y}^{(2)})^{*}\mathbf{Y}^{(2)}+t\mathbf{I})^{-1}. Then the matrices (Y(1))∗Y(1)(\mathbf{Y}^{(1)})^{*}\mathbf{Y}^{(1)} and At(Y(2))∗Y(2)At\mathbf{A}_{t}(\mathbf{Y}^{(2)})^{*}\mathbf{Y}^{(2)}\mathbf{A}_{t} are asymptotically free. By the Marchenko–Pastur theorem (Theorem 8.1), their limiting (mean) empirical spectral distributions are given by ϱ\varrho, the Marchenko–Pastur distribution, and σt\sigma_{t}, the induced measure of ϱ\varrho under the mapping x↦(x+t)−1x(x+t)−1x\mapsto(x+t)^{-1}x(x+t)^{-1}, respectively. Thus, by Lemma A.2, the limiting (mean) empirical spectral distribution of Wt(Y)\mathbf{W}_{t}(\mathbf{Y}) is given by ϱ⊠σt\varrho\boxtimes\sigma_{t}, with SS-transform Sϱ⋅SσtS_{\varrho}\cdot S_{\sigma_{t}}. Clearly, as t→0t\to 0, we have σt→σ\sigma_{t}\to\sigma in Kolmogorov distance, where σ\sigma denotes the induced measure of ϱ\varrho under the mapping x↦x−1x\mapsto x^{-1}. Using (5.10) and (5.18), we obtain ϱ⊠σt→ϱ⊠σ\varrho\boxtimes\sigma_{t}\to\varrho\boxtimes\sigma in Kolmogorov distance and Sϱ⋅Sσt→Sϱ⋅SσS_{\varrho}\cdot S_{\sigma_{t}}\to S_{\varrho}\cdot S_{\sigma}. It therefore follows by Lemma 8.14 that the limiting (mean) empirical spectral distribution of W(Y)\mathbf{W}(\mathbf{Y}) has the SS-transform Sϱ⋅SσS_{\varrho}\cdot S_{\sigma}. Now, by Remarks 8.3 and 8.4, we have

After a simple calculation we get that the density of the limit distribution is given by

This completes the proof of Theorem 8.13. ∎

1.7 Product of Independent Matrices from Spherical Ensemble

In this section we consider products of independent matrices of type X(2q−1)(X(2q))−1\mathbf{X}^{(2q-1)}(\mathbf{X}^{(2q)})^{-1}, for q=1,…,mq=1,\ldots,m, assuming that all matrices and all entries of matrices are independent. Let F=∏q=1mX(2q−1)(X(2q))−1\mathbf{F}=\prod_{q=1}^{m}\mathbf{X}^{(2q-1)}(\mathbf{X}^{(2q)})^{-1} and W=FF∗\mathbf{W}=\mathbf{F}\mathbf{F}^{*}. Let s12≥…≥sn2s_{1}^{2}\geq\ldots\geq s_{n}^{2} denote the eigenvalues of the matrix W\mathbf{W} and let Gn(x)\mathcal{G}_{n}(x) denote the empirical distribution function

We shall assume as usual that X(q)=1n(Xjk(q))\mathbf{X}^{(q)}=\frac{1}{\sqrt{n}}(X_{jk}^{(q)}), and that (8.1) or (8.2) holds. Then we have the following result, which was already announced by the first and third author of this paper in . See also Forrester and Forrester and Liu for the Gaussian case. The density in (8.63) also occurs in Biane in the context of free multiplicative Lévy processes.

Assume that the random variables Xjk(q)X_{jk}^{(q)}, for q=1,…,2mq=1,\ldots,2m and j,k=1,…,nj,k=1,\ldots,n satisfy condition (8.23). Then

where Gm(x)G_{m}(x) denotes the distribution function with density pm(x)=Gm′(x)p_{m}(x)=G_{m}^{\prime}(x) given by

Similarly as the proof of Theorem 8.13, the proof of Theorem 8.15 is rather technical. Before we can apply Theorem 3.2, we must regularize all the inverse matrices, and this requires a slightly more complicated construction than in the previous subsection.

Let us formulate a general result. Let A\mathbf{A} and B\mathbf{B} be random matrices of size n×nn\times n, and let X\mathbf{X} be a Girko-Ginibre matrix of size n×nn\times n satisfying (A.3). Introduce the inverse X0−1:=X−1\mathbf{X}_{0}^{-1}:=\mathbf{X}^{-1} and the regularized inverse

Even more, the convergence is uniform in A\mathbf{A}.

Remark. Note that our assumptions on X\mathbf{X} and B\mathbf{B} imply that

In the following considerations we always work on this event. (This is possible because we are interested in convergence in probability.) In particular, the inverse X0−1\mathbf{X}_{0}^{-1} exists on this event.

The main idea of the proof is as follows. Firstly, we replace the matrix B\mathbf{B} by some regularized version Bs\mathbf{B}_{s} whose singular value distribution is bounded away from zero and infinity. Secondly, we regularize the matrix X−1\mathbf{X}^{-1} as described above. Thirdly, we undo the regularization of the matrix B\mathbf{B}.

Note that as s→0s\to 0, we have f(s,x)→f(0,x):=xf(s,x)\to f(0,x):=x for any x>0x>0. Furthermore, note that for each s>0s>0, the function x↦f(s,x)x\mapsto f(s,x) is increasing in xx, with values in the bounded interval (s1/21+s3/2,s−1)(\frac{s^{1/2}}{1+s^{3/2}},s^{-1}). Also, setting h(s,x):=log⁡f(s,x)h(s,x):=\log f(s,x), we may write

Here, ∂1\partial_{1} denotes the partial derivative w.r.t. the first argument.

Write B=UΛV∗\mathbf{B}=\mathbf{U}\mathbf{\Lambda}\mathbf{V}^{*} (singular value distribution), and set Bs:=Uf(s,Λ)V∗\mathbf{B}_{s}:=\mathbf{U}f(s,\mathbf{\Lambda})\mathbf{V}^{*}, s≥0s\geq 0. Here, f(s,Λ)f(s,\mathbf{\Lambda}) is obtained by applying f(s, ⋅ )f(s,\,\cdot\,) to the diagonal elements of Λ\mathbf{\Lambda}. Note that

Now, for any n×nn\times n matrices M1\mathbf{M}_{1}, M2\mathbf{M}_{2}, M3\mathbf{M}_{3}, with M2\mathbf{M}_{2} self-adjoint, we have

where ∣M2∣|\mathbf{M}_{2}| is defined by spectral calculus. Furthermore, for any t,s≥0t,s\geq 0, we have

It therefore follows from (8.51), (8.54) and (8.55) that

Taking the normalized trace in (8.53) and using the previous estimates, it follows that for any s∈(0,1)s\in(0,1),

Once we have (8.58) and (8.59), it is easy to complete the proof. Indeed, fix ε>0\varepsilon>0. Then, by (8.58), there exists a function s(t)s(t) such that s(t)→0s(t)\to 0 as t→0t\to 0 and still

Thus, it remains to show (8.58) and (8.59). We have already checked that (8.58) follows from our assumption that X\mathbf{X} satisfies Conditions (C0), (C1) and (C2); see the proof of Lemma 8.14. It is straightforward to check that

Since the matrix B\mathbf{B} satisfies (C0), (C1), (C2) and the squared singular value distribution of B\mathbf{B} is weakly convergent to a limit ν\nu with ∫∣log⁡∣ dν<∞\int|\log|\,d\nu<\infty (by Lemma A.9 in the appendix), this implies (8.59). ∎

The proof is by induction on mm. Suppose that m>1m>1 and that the result is true of all smaller values of mm. (The case of a single spherical matrix is Theorem 8.13.) We start with a regularisation of the inverse matrices. We introduce the following matrices

We apply Lemma 8.16. Let Zt(k):=X(2k−1)(X(2k))t−1\mathbf{Z}_{t}^{(k)}:=\mathbf{X}^{(2k-1)}(\mathbf{X}^{(2k)})_{t}^{-1}, k=1,…,mk=1,\ldots,m and

Clearly, writing vk:=(t,…,t,0,…,0)\mathbf{v}_{k}:=(t,\ldots,t,0,\ldots,0) for the vector consisting of kk tt’s and m−km-k ’s, we have

By Lemma 8.16, each of the summands converges to zero in probability. Here we use (i) the inductive hypothesis and (ii) the fact that an arbitrary product of independent spherical matrices satisfies Conditions (C0) – (C2). The latter will be checked in the proof of Theorem 8.24; note that the verification does not rely on the results in this section. This completes the proof of Lemma 8.17. ∎

We continue with the proof of Theorem 8.15. We may now use Theorem 3.2 for the matrix Wt\mathbf{W}_{t}. The Lindeberg condition (3.1) follows from condition (8.23). The check of the remaining conditions of Theorem 3.2 is similar to that in the previous subsection; we omit the details. Thus, by a similar argument as in the proof of Theorem 8.13 (but with Lemma 8.17 instead of Lemma 8.14), it remains to identify the limiting empirical spectral distribution of the matrix W(Y)\mathbf{W}(\mathbf{Y}) in the Gaussian case.

Here we can use the same approach and notation as in the proof of Theorem 8.13. Firstly, by asymptotic freeness, Lemma A.2 and Theorem 8.1, the limiting (mean) empirical spectral distribution of the matrices Wt(Y)\mathbf{W}_{t}(\mathbf{Y}) is given by ϱ⊠m⊠σt⊠m\varrho^{\boxtimes m}\boxtimes\sigma_{t}^{\boxtimes m}, with corresponding SS-transform SWt(z)=Sϱm(z)⋅Sσtm(z)S_{\mathbf{W}_{t}}(z)=S_{\varrho}^{m}(z)\cdot S_{\sigma_{t}}^{m}(z). Secondly, using Lemma 8.17, it follows that the limiting (mean) empirical spectral distribution of the matrices W(Y)\mathbf{W}(\mathbf{Y}) is given by ϱ⊠m⊠σ⊠m\varrho^{\boxtimes m}\boxtimes\sigma^{\boxtimes m}, with corresponding SS-transform SW(z)=Sϱm(z)⋅Sσm(z)S_{\mathbf{W}}(z)=S_{\varrho}^{m}(z)\cdot S_{\sigma}^{m}(z). Thirdly, by Remarks 8.3 and 8.4, we get

Using the representation (8.62) we may now determine the Stieltjes transform of the asymptotic distribution of the eigenvalues of the matrix W\mathbf{W} and from here determine the density of its distribution. Let g(z)g(z) denote the Stieltjes transform of the asymptotic distribution of the eigenvalues of matrix W\mathbf{W}. By definition of the SS-transform, we have

Combining the last two equalities, we get

Thus Theorem 8.13 is proved completely. ∎

2 Applications of Theorem 4.4: Distribution of eigenvalues

Let X\mathbf{X} denote the n×nn\times n random matrix with independent entries 1nXjk\frac{1}{\sqrt{n}}X_{jk} such that (8.1) or (8.2) holds. Assume that the r.v.’s XjkX_{jk} satisfy the condition

Let λ1,…,λn\lambda_{1},\ldots,\lambda_{n} denote the eigenvalues of the matrix X\mathbf{X}. Denote by μn\mu_{n} the empirical spectral distribution of the matrix X\mathbf{X}. Then we have the following result, cf. Girko , Bai , Pan and Zhou , Götze and Tikhomirov as well as Tao and Vu .

Assume that condition (8.64) holds. Then the measures μn\mu_{n} converge weakly in probability to the uniform distribution μ\mu on the unit disc.

To prove Theorem 8.18, we first apply Theorem 4.4. Then we compute the limit distribution for the Gaussian case using the result of Theorem 7.6.

Note that condition (3.1) is implied by condition (8.64) and that conditions (3.2) and (3.20) – (3.25) of Theorem 3.2 have been checked in Subsection 8.1.1. It remains to check conditions (C0), (C1) and (C2) of Theorem 4.4. To this end we can use existing bounds for singular values; see Lemmas A.4, A.5 and A.6 in Appendix A.3. We may therefore apply Theorem 4.4.

Let V=VY\mathbf{V}=\mathbf{V}_{\mathbf{Y}} be defined as in Section 6, with F=FY=Y\mathbf{F}=\mathbf{F}_{\mathbf{Y}}=\mathbf{Y}. By Proposition 5.8, the matrices V\mathbf{V} and J(α)\mathbf{J}(\alpha) are asymptotically free. Furthermore, it follows from Theorem 8.1 that the limiting eigenvalue distribution of the matrices VY\mathbf{V}_{\mathbf{Y}} is given by the semi-circular law. We now compute the limiting eigenvalue distribution of the matrices FY\mathbf{F}_{\mathbf{Y}} using Theorem 7.6. It is well-known, see e.g. Rao and Speicher , Section 3, that

Now recall that (i) ψ(α)∈\psi(\alpha)\in by Theorem 4.4, (ii) ϰ\varkappa is continuous by Remark 6.5, and (iii) ϰ(α)≠0\varkappa(\alpha)\neq 0 for α≈0\alpha\approx 0 by Lemma 6.6. Thus, we obtain a unique solution, namely

Using (7.10), it therefore follows that the density f(u,v)f(u,v) of the limiting empirical spectral distribution of the matrix Y\mathbf{Y} is given by the equality

2.2 Product of Independent Square Matrices

Let m≥1m\geq 1. Consider independent random matrices X(q)\mathbf{X}^{(q)}, q=1,…,mq=1,\ldots,m with independent entries 1nXjk(q)\tfrac{1}{\sqrt{n}}X_{jk}^{(q)}, 1≤j,k≤n1\leq j,k\leq n, q=1,…,mq=1,\ldots,m, and suppose that (8.1) or (8.2) holds. Let F=∏q=1mX(q)\mathbf{F}=\prod_{q=1}^{m}\mathbf{X}^{(q)}. Then we have the following result, cf. Burda, Janik and Waclaw for the Gaussian case and Götze and Tikhomirov as well as O’Rourke and Soshnikov for the general case.

Let the r.v.’s Xjk(q)X_{jk}^{(q)} satisfy the condition

Then the empirical spectral distributions μn\mu_{n} of the matrices F\mathbf{F} converge weakly in probability to the measure μ\mu with Lebesgue density

The limiting measure μ\mu is the induced measure of the uniform distribution on the unit disc under the mapping z↦zmz\mapsto z^{m}. Consequently, the product of mm independent square matrices has the same limiting empirical spectral distribution as the mmth power of a single matrix.

The conditions of Theorem 3.2 were checked in the proof of Theorem 8.1.2. The condition (C0)(C0) of Theorem 4.4 for p=2p=2 follows from Lemma 7.2 in , where it is shown that

Furthermore, in Lemma 5.1 in it is proved that there exist positive constants QQ and AA such that

This implies condition (C1)(C1) of Theorem 4.4. Moreover, inequality (5.16) and Lemma 5.2 in together imply that for some 0<γ<10<\gamma<1, for any sequence δn→0\delta_{n}\to 0

with n1=[n−nδn]+1n_{1}=[n-n\delta_{n}]+1 and n2=[n−nγ]n_{2}=[n-n^{\gamma}]. This implies condition (C2)(C2) of Theorem 4.4. According to Theorem 4.4 we may now consider the Gaussian matrices Y(q)\mathbf{Y}^{(q)}, q=1,…,mq=1,\ldots,m. Let

By Proposition 5.8, the matrices V\mathbf{V} and J(α)\mathbf{J}(\alpha) are asymptotically free. Moreover, it follows from Remark 8.3 and Lemma A.2 that the SS-transform corresponding to the matrices W:=FYFY∗\mathbf{W}:=\mathbf{F}_{\mathbf{Y}}\mathbf{F}_{\mathbf{Y}}^{*} is given by

Thus, the SS-transform of the matrix V\mathbf{V} is given by the formula

We rewrite equations (7.6) for this case:

Solving this system we find by similar arguments as in the previous subsection that

By (7.10), these relations immediately imply that

2.3 Product of Independent Rectangular Matrices

Let m≥1m\geq 1 be fixed. Let for any n≥1n\geq 1 be given integers n0=n,n1≥n,…,nm−1≥nn_{0}=n,n_{1}\geq n,\ldots,n_{m-1}\geq n and nm=nn_{m}=n. Assume that yq=lim⁡n→∞nnq∈(0,1]y_{q}=\lim_{n\to\infty}\frac{n}{n_{q}}\in(0,1], q=1,…,mq=1,\ldots,m. Note that ym=1y_{m}=1. Consider independent random matrices X(q)\mathbf{X}^{(q)} of order nq−1×nqn_{q-1}\times n_{q}, q=1,…,mq=1,\ldots,m, with independent entries 1nqXjk(q)\frac{1}{\sqrt{n_{q}}}X_{jk}^{(q)} as in (8.1) or (8.2). Put F=∏q=1mX(q)\mathbf{F}=\prod_{q=1}^{m}\mathbf{X}^{(q)}. Then we have the following result, see also Burda, Jarosz, Livan, Nowak, Swiech and Tikhomirov for related results in the Gaussian case and in the general case, respectively.

Assume that the r.v.’s Xjk(q)X_{jk}^{(q)} for q=1,…,mq=1,\ldots,m, and j=1,…,pq−1j=1,\ldots,p_{q-1}, k=1,…,pqk=1,\ldots,p_{q} satisfy the condition (8.65). Then the empirical spectral distributions of the matrices F\mathbf{F} weakly converge in probability to the measure μ\mu with Lebesgue density

where α=u+iv\alpha=u+iv and ψ(α)\psi(\alpha) is given by the unique solution in the interval $$ to the equation

The conditions of Theorem 3.2 were checked in Subsubsection 8.1.2. The conditions (C0)(C0), (C1)(C1) and (C2)(C2) of Theorem 4.4 may be checked similarly as in the proof of the previous Theorem; we omit the details. To compute the limit measure μ\mu in the Gaussian case, we may use Theorem 7.6 now. Using Remark 8.3 and Lemma A.2, we may show that

We have used here that ym=1y_{m}=1. Inserting (8.70) into (7.6), we get

Solving this system we find that for ∣α∣≤1|\alpha|\leq 1,

We have used here that the function h(x):=x∏ν=1m−1(1−yν+yνx)h(x):=x\prod_{\nu=1}^{m-1}(1-y_{\nu}+y_{\nu}x) is strictly increasing on [0,∞)[0,\infty) with h(0)=0h(0)=0 and h(1)=1h(1)=1. Now (8.69) follows immediately from (7.10). In order to check that ff is in fact a probability density, regard ψ\psi and ff as functions of r:=u2+v2r:=\sqrt{u^{2}+v^{2}}. Then

and it follows using polar coordinates that

Finally, for m=2m=2 and u2+v2≤1u^{2}+v^{2}\leq 1, we get

Hence, on the set {u2+v2≤1}\{u^{2}+v^{2}\leq 1\}, we obtain

2.4 Spherical Ensemble

Let X(1)\mathbf{X}^{(1)} and X(2)\mathbf{X}^{(2)} be independent n×nn\times n random matrices with independent entries 1nXjk(q)\frac{1}{\sqrt{n}}X_{jk}^{(q)} such that (8.1) or (8.2) holds. Consider the matrix FX=X(1)(X(2))−1\mathbf{F}_{\mathbf{X}}=\mathbf{X}^{(1)}({\mathbf{X}^{(2)}})^{-1}. Let μn\mu_{n} denote the empirical spectral distribution of the matrix FX\mathbf{F}_{\mathbf{X}}. Then we have the following result, cf. Bordenave .

Let the r.v.’s Xjk(q)X_{jk}^{(q)} for q=1,2q=1,2 and j,k=1,…,nj,k=1,\ldots,n satisfy the condition (8.65). Then the measures μn\mu_{n} weakly converge in probability to the measure μ\mu with Lebesgue density

This density corresponds after stereographic projection of the complex plane to the uniform distribution on the sphere.

This implies that the first factor on the r.h.s. of (8.71) is bounded in probability. For p<29p<\frac{2}{9} we have β=2p2−p<14\beta=\frac{2p}{2-p}<\frac{1}{4}. Denote by Gn(x)\mathcal{G}_{n}(x) the empirical distribution function of the squared singular values of the matrix X(2)\mathbf{X}^{(2)}. By the Marchenko–Pastur theorem

This implies that for p<2(1−γ)2Q+1−γp<\frac{2(1-\gamma)}{2Q+1-\gamma}

From n−n2n≤ϰn\frac{n-n_{2}}{n}\leq\varkappa_{n} it follows that

It remains to prove that the last summand on the r.h.s. of (8.72) is bounded in probability. Again by Lemma A.8, with probability 1−o(1)1-o(1), we have sn2(X(2))≥cϰns_{n_{2}}(\mathbf{X}^{(2)})\geq c\varkappa_{n} and therefore

The last integral on the r.h.s. of (8.74) is bounded for β<1\beta<1. Integrating by parts in the first integral on the r.h.s. of (8.74), we get

The last inequalities imply that the last summand on the r.h.s. of (8.72) is bounded in probability. This concludes the proof of the condition (C0)(C0). The condition (C1)(C1) follows from the bound

and Lemmas A.4 and A.7. To prove the condition (C2)(C2), fix a sequence (δn)(\delta_{n}) with δn≥n−γ\delta_{n}\geq n^{-\gamma} for all nn and δn→0\delta_{n}\to 0, and let n2:=n[1−δn]n_{2}:=n[1-\delta_{n}]. Then, by the arguments for condition (C1)(C1) as well as Theorem 3.3.4 in , we have, with probability 1−o(1)1-o(1),

Note that both sums on the r.h.s. converge to zero in probability. For the first sum, this follows from Lemma A.8 applied to the matrix X(1)−αX(2)\mathbf{X}^{(1)}-\alpha\mathbf{X}^{(2)}, while for the second sum, this follows from the observation that we have, with probability 1−o(1)1-o(1),

Combining these estimates, we come to the conclusion that 1n∑k=n2+1n∣log⁡sk(F−αI)∣\frac{1}{n}\sum_{k=n_{2}+1}^{n}|\log s_{k}(\mathbf{F}-\alpha\mathbf{I})| converges to zero in probability, i.e. condition (C2) is proved.

Thus, the conditions (C0), (C1), (C2) have been checked for the matrices FX\mathbf{F}_{\mathbf{X}}. For the Gaussian matrices FY\mathbf{F}_{\mathbf{Y}}, the proof is the same. We may now apply Theorem 4.4 to conclude that the limiting eigenvalue distributions of the matrices FX\mathbf{F}_{\mathbf{X}} and FY\mathbf{F}_{\mathbf{Y}} are the same (if existent). Thus, it remains to compute the limit of the empirical distribution of the eigenvalues of the matrix FY\mathbf{F}_{\mathbf{Y}}.

From now on, let the matrices Ft:=Ft(Y)\mathbf{F}_{t}:=\mathbf{F}_{t}(\mathbf{Y}) be defined as in the proof of Theorem 8.13. We shall use the asymptotic freeness of matrices

As we have seen in the proof of Theorem 8.13, the limiting (mean) empirical spectral distribution of the matrices FtFt∗\mathbf{F}_{t}\mathbf{F}_{t}^{*} is given by σ⊠ϱt\sigma\boxtimes\varrho_{t}, with SS-transform Sσ⋅SϱtS_{\sigma}\cdot S_{\varrho_{t}}. Here, σ\sigma and ϱt\varrho_{t} denote the Marchenko-Pastur distribution and its induced measure under the mapping x↦(x+t)−1x(x+t)−1x\mapsto(x+t)^{-1}x(x+t)^{-1}, respectively. Thus, the limiting (mean) empirical spectral distribution of the matrices Vt\mathbf{V}_{t} is given by Q−1(σ⊠ϱt)\mathcal{Q}^{-1}(\sigma\boxtimes\varrho_{t}), where Q\mathcal{Q} is as in (5.15), and the limiting spectral distribution of the matrices Vt(α):=Vt+J(α)\mathbf{V}_{t}(\alpha):=\mathbf{V}_{t}+\mathbf{J}(\alpha) is given by Q−1(σ⊠ϱt)⊞T(α)\mathcal{Q}^{-1}(\sigma\boxtimes\varrho_{t})\boxplus T(\alpha), where T(α)T(\alpha) is as in Section 6. By a variant of Lemma 8.14 for shifted matrices as well as relations (5.9) and (5.10), it then follows that the limiting (mean) empirical spectral distributions of the matrices V\mathbf{V} and V(α):=V+J(α)\mathbf{V}(\alpha):=\mathbf{V}+\mathbf{J}(\alpha) are given by Q−1(σ⊠ϱ)\mathcal{Q}^{-1}(\sigma\boxtimes\varrho) and Q−1(σ⊠ϱ)⊞T(α)\mathcal{Q}^{-1}(\sigma\boxtimes\varrho)\boxplus T(\alpha), respectively.

Moreover, the SS-transform of the limiting eigenvalue distribution of FF∗\mathbf{F}\mathbf{F}^{*} is given by

as we have seen in the proof of Theorem 8.13. Thus, by (5.16), the SS-transform of the limiting eigenvalue distribution of V\mathbf{V} is given by

Finally, by Theorem 6.1, the Stieltjes transform gt(z,α)g_{t}(z,\alpha) associated with the matrices Vt(α)\mathbf{V}_{t}(\alpha) satisfies the system of equations (6.4) with SVS_{\mathbf{V}} replaced by SVtS_{\mathbf{V}_{t}}. It therefore follows by continuity that the Stieltjes transform g(z,α)g(z,\alpha) associated with the matrices V(α)\mathbf{V}(\alpha) satisfies the system of equations (6.4).

Thus, the assumptions stated above Assumption 7.3 are satisfied, and we may apply Theorem 7.6. Solving now the system

The last equality and equality (7.10) together imply

2.5 Product of Independent Matrices from Spherical Ensemble

For fixed m≥1m\geq 1, let X(q)\mathbf{X}^{(q)}, q=1,…,2mq=1,\ldots,2m, be independent n×nn\times n random matrices with independent entries 1nXjk(q)\frac{1}{\sqrt{n}}X_{jk}^{(q)}. Suppose that (8.1) or (8.2) holds. Consider the matrix Fm=Fm(X)=∏q=1mX(2q−1)(X(2q))−1\mathbf{F}_{m}=\mathbf{F}_{m}(\mathbf{X})=\prod_{q=1}^{m}\mathbf{X}^{(2q-1)}(\mathbf{X}^{(2q)})^{-1}, and denote by μn\mu_{n} its empirical spectral distribution.

Let the r.v.’s Xjk(q)X_{jk}^{(q)} for q=1,…,2mq=1,\ldots,2m and j,k=1,…,nj,k=1,\ldots,n satisfy the condition (8.65). Then the measures μn\mu_{n} weakly converge in probability to the measure μ\mu with Lebesgue density

Write s1(A)≥⋯≥sn(A)s_{1}(\mathbf{A})\geq\cdots\geq s_{n}(\mathbf{A}) for the singular values of the matrix A\mathbf{A}. Then, similarly as in , Theorem 3.3.14 (c), we have, for any k=1,…,nk=1,\ldots,n and for any function ff such that φ(t)=f(et)\varphi(t)=f({\rm e}^{t}) is increasing and convex,

Now use similar arguments as in the proof of Theorem 8.22. Thus, Theorem 4.4 and Remark 4.5 are applicable, and it remains to determine the limiting empirical spectral distribution in the Gaussian case.

Write F=F(Y)\mathbf{F}=\mathbf{F}(\mathbf{Y}) for the products of independent “Gaussian” spherical matrices. To find their limiting empirical spectral distribution, we use the results from Section 7. For brevity, we give only a formal proof; it is straightforward (although a bit cumbersome) to make this proof rigorous by using similar arguments as in the proof of Theorem 8.22 (using a variant of the regularization lemma 8.17 this time). First, we find the SS-transform SFF∗(z)S_{\mathbf{F}\mathbf{F}^{*}}(z) associated with the matrices W:=FF∗\mathbf{W}:=\mathbf{F}\mathbf{F}^{*}. Remember that, for any q=1,…,mq=1,\ldots,m, the SS-transform associated with the matrices Y(q)Y(q+1)−1(Y(q+1)−1)∗(Y(q))∗\mathbf{Y}^{(q)}{\mathbf{Y}^{(q+1)}}^{-1}({\mathbf{Y}^{(q+1)}}^{-1})^{*}(\mathbf{Y}^{(q)})^{*} is given by

by (8.75). Thus, by the multiplicative property of SS-transform, we formally have

The last equality and equality (7.10) together imply

Appendix A Appendix

In this subsection, F=FX\mathbf{F}=\mathbf{F}_{\mathbf{X}} is defined as in Section 1, and V\mathbf{V} and R\mathbf{R} are the Hermitian matrices defined by

where B\mathbf{B} is a non-random matrix and z=u+ivz=u+iv with v>0v>0.

Suppose that the rank condition (3.2) holds. Then

We introduce the σ\sigma-algebras \textfrakMq,j=σ{Xlk(q), j<l≤nq−1,k=1,…,nq;Xpk(r)\textfrak{M}_{q,j}=\sigma\{X^{(q)}_{lk},\,j<l\leq n_{q-1},k=1,\ldots,n_{q};X^{(r)}_{pk}, r=q+1,…m, p=1,…,nr−1, k=1,…,nr}r=q+1,\ldots m,\,p=1,\ldots,n_{r-1},\,k=1,\ldots,n_{r}\} and use the representation

where E q,j\mathbf{E}\,_{q,j} denotes conditional expectation given the σ\sigma-algebra \textfrakMq,j\textfrak{M}_{q,j}. Note that \textfrakMq,nq−1=\textfrakMq+1,0\textfrak{M}_{q,n_{q-1}}=\textfrak{M}_{q+1,0}. Furthermore, we introduce the matrices X(q,j)\mathbf{X}^{(q,j)} obtained from X(q)\mathbf{X}^{(q)} by replacing the entries Xjk(q)X_{jk}^{(q)} (k=1,…,nqk=1,\ldots,n_{q}) by zero’s. Define the matrices

A.2 S-Transform for Rectangular Matrices

In the case of rectangular matrices this relation is not true anymore. The next lemma gives the correct relation for products of rectangular matrices.

where δ0\delta_{0} denotes the unit atom at zero. From here it follows that

By asymptotic freeness and the multiplicative property of the SS-transform, we have

(In particular, the limiting eigenvalue distribution μYY∗X∗X\mu_{\mathbf{Y}\mathbf{Y}^{*}\mathbf{X}^{*}\mathbf{X}} exists.) By the same argument as for (A.2), we get

The three last equalities together imply the result of Lemma. ∎

Using the preceding results, it is easy to see why the mmth power of a random square matrix Yn\mathbf{Y}_{n} and the product Yn(1)⋯Yn(m)\mathbf{Y}^{(1)}_{n}\cdots\mathbf{Y}^{(m)}_{n} of mm independent copies of this matrix should have the same limiting singular value and eigenvalue distributions. Indeed, let Yn\mathbf{Y}_{n} be bi-unitary invariant random square matrices such that the empirical spectral distribution of the matrices YnYn∗\mathbf{Y}_{n}\mathbf{Y}^{*}_{n} converges weakly in probability as well as in moments to a compactly supported probability measure μYY∗\mu_{\mathbf{Y}\mathbf{Y}^{*}}.

Then, similarly as in Hiai and Petz , using the singular value decomposition of the matrix Yn\mathbf{Y}_{n}, one can show that YnYn∗\mathbf{Y}_{n}\mathbf{Y}^{*}_{n} and Yn∗Yn\mathbf{Y}^{*}_{n}\mathbf{Y}_{n} are asymptotically free, and it follows from Equation (A.1) (and induction) that Ynm(Ynm)∗\mathbf{Y}^{m}_{n}(\mathbf{Y}^{m}_{n})^{*} converges in moments to μYY∗⊠m\mu_{\mathbf{Y}\mathbf{Y}^{*}}^{\boxtimes m}. A similar argument, also based on Equation (A.1), shows that the same is true for (Yn(1)⋯Yn(m))(Yn(1)⋯Yn(m))∗(\mathbf{Y}^{(1)}_{n}\cdots\mathbf{Y}^{(m)}_{n})(\mathbf{Y}^{(1)}_{n}\cdots\mathbf{Y}^{(m)}_{n})^{*}, where Yn(1),…,Yn(m)\mathbf{Y}^{(1)}_{n},\ldots,\mathbf{Y}^{(m)}_{n} are mm independent copies of mm. Thus, the matrices Ynm\mathbf{Y}^{m}_{n} and Yn(1)⋯Yn(m)\mathbf{Y}^{(1)}_{n}\cdots\mathbf{Y}^{(m)}_{n} will have the same limiting singular value distributions.

Now the SS-transform of the limiting singular value distribution of the shifted matrices F−αI\mathbf{F}-\alpha\mathbf{I} is well defined by α\alpha and the limiting singular value distribution of the matrix F\mathbf{F}. To prove this we must use the additive property of the RR-transform and the correspondence between RR- and SS-transforms. Furthermore, note that the limit measure for the eigenvalue distribution is well defined by its logarithmic potential and that we may reconstruct the logarithmic potential from the family of the singular value distribution of the shifted matrices. It therefore follows that the limiting eigenvalue distribution of the matrices Ynm\mathbf{Y}^{m}_{n} and Yn(1)⋯Yn(m)\mathbf{Y}^{(1)}_{n}\cdots\mathbf{Y}^{(m)}_{n} will also be the same.

But the eigenvalues of the mmth power of a matrix are the mmth powers of the eigenvalues of that matrix. For example, if the limiting eigenvalue distribution of the random matrix Yn\mathbf{Y}_{n} is the circular law, then the limiting eigenvalue distribution of the product Yn(1)⋯Yn(m)\mathbf{Y}^{(1)}_{n}\cdots\mathbf{Y}^{(m)}_{n} is the mmth power of the uniform distribution in the unit disc.

A.3 Bounds on Singular Values

Throughout this subsection, let X\mathbf{X} denote an n×nn\times n random matrix with independent entries 1nXjk\frac{1}{\sqrt{n}}X_{jk} such that E Xjk=0\mathbf{E}\,X_{jk}=0 and E ∣Xjk∣2=1\mathbf{E}\,|X_{jk}|^{2}=1. Let s1(X)≥⋯≥sn(X)s_{1}(\mathbf{X})\geq\cdots\geq s_{n}(\mathbf{X}) denote the singular values of the matrix X\mathbf{X}. Then we have the following result:

We have lim⁡t→∞lim sup⁡n→∞Pr⁡{1n∑k=1nsk2(X)≥t}=0\lim_{t\to\infty}\limsup_{n\to\infty}\Pr\{\tfrac{1}{n}\sum_{k=1}^{n}s_{k}^{2}(\mathbf{X})\geq t\}=0.

Henceforward, assume additionally that the r.v.’s XjkX_{jk} satisfy the condition

Under these assumptions, we have the following bounds on the small singular values, see Götze and Tikhomirov, , Theorem 4.1 and , Lemma 5.2 and Proposition 5.1. (For the i.i.d. case similar results were obtained by Tao and Vu, , Lemma 4.1 and 4.2.) Let s1(X−αI)≥⋯≥sn(X−αI)s_{1}(\mathbf{X}-\alpha\mathbf{I})\geq\cdots\geq s_{n}(\mathbf{X}-\alpha\mathbf{I}) denote the singular values of the matrix X−αI\mathbf{X}-\alpha\mathbf{I}.

For a proof of this lemma see the proof of Theorem 4.1 in .

with n1=[n−nδn]+1n_{1}=[n-n\delta_{n}]+1 and n2=[n−nγ]n_{2}=[n-n^{\gamma}].

For a proof of this lemma see the proof of inequality (5.17) and Lemma 5.2 in .

For the investigation of the spherical ensembles, we need the following extensions of these results; see Equations (5.9) and (5.17) in .

Suppose that condition (A.3) holds. Then, for any K>0K>0 and L>0L>0, there exist positive constants QQ and BB such that for any non-random matrix M\mathbf{M} with ∥M∥2≤KnL\|\mathbf{M}\|_{2}\leq Kn^{L}, we have

Suppose that condition (A.3) holds. Then, for any fixed K>0K>0 and L>0L>0, there exist constants 0<γ<10<\gamma<1 and c>0c>0 such that for any non-random matrix M\mathbf{M} with ∥M∥2≤KnL\|\mathbf{M}\|_{2}\leq Kn^{L}, we have

A.4 Technical Details for Section 7

In this subsection, we state some technical lemmas which have been used in Section 7.

Using similar arguments as in the proof of Theorem 4.4 / Remark 4.5, we may prove the following.

Assume that the matrices FY\mathbf{F}_{\mathbf{Y}} satisfy the conditions (C0)(C0), (C1)(C1) and (C2)(C2). Moreover, assume that the singular value distributions of the matrices FY\mathbf{F}_{\mathbf{Y}} converge weakly in probability to a non-random probability measure ν\nu. Then the logarithm is integrable w.r.t. ν\nu, and we have

Clearly, the main problem is to show that the logarithm is integrable w.r.t. ν\nu. Once this is shown, it is straightforward to adapt the proof of Theorem 4.4, replacing the singular value distributions of the matrices FY\mathbf{F}_{\mathbf{Y}} with the fixed distribution ν\nu. We will show separately that log⁡+:=max⁡{+log⁡,0}\log^{+}:=\max\{+\log,0\} and log⁡−:=max⁡{−log⁡,0}\log^{-}:=\max\{-\log,0\} are integrable w.r.t. ν\nu. In doing so, we write νn\nu_{n} for the singular value distribution of FY\mathbf{F}_{\mathbf{Y}}.

Let us begin with the positive part. First of all, passing to a suitable subsequence, we may assume w.l.o.g. that νn⇒ν\nu_{n}\Rightarrow\nu almost surely. Now, by assumption (C0), there exists a constant K>0K>0 such that

Let us now consider the negative part. Again, we may select a subsequence (νnk)(\nu_{n_{k}}) such that νnk⇒ν\nu_{n_{k}}\Rightarrow\nu almost surely. Moreover, using monotone convergence and weak convergence, we have

We may assume w.l.o.g. that the sequence (k(l))(k(l)) is increasing. Thus, we obtain a subsequence (which we again denote by νnk\nu_{n_{k}}, by abuse of notation) such that

Thus, we may select a sequence (k(l))(k(l)) such that with probability 11, we have

We now prove Lemma 7.4. For convenience, we repeat the statement of the lemma.

Assumptions 7.2 and 7.3 hold for probability measures μV\mu_{\mathbf{V}} such that μV([−x,+x]c)=O(x−η)\mu_{\mathbf{V}}([-x,+x]^{c})=\mathcal{O}(x^{-\eta}) (x→∞)(x\to\infty) for some η>0\eta>0.

The proof consists of several parts. We will use the fact that the free additive convolution is monotone with respect to stochastic order ≤st\leq_{\mathop{st}} (see e.g. Proposition 4.16 in Bercovici and Voiculescu ), i.e. we have

It follows from (A.5) that μV([−x,+x]c)=O(x−η)\mu_{\mathbf{V}}([-x,+x]^{c})=\mathcal{O}(x^{-\eta}) (x→∞x\to\infty) implies μV(α)([−x,+x]c)=O(x−η)\mu_{\mathbf{V}(\alpha)}([-x,+x]^{c})=\mathcal{O}(x^{-\eta}) (x→∞x\to\infty), where the O\mathcal{O}-bound is locally uniform in α\alpha. Thus, using integration by parts, we find that for any continuously differentiable (possibly complex-valued) function ff such that f′(x)=O(∣x∣−1)f^{\prime}(x)=\mathcal{O}(|x|^{-1}) as ∣x∣→∞|x|\to\infty, we have

Suppose w.l.o.g. that ∣α∣≤∣β∣|\alpha|\leq|\beta|, and set m:=∣α∣+∣β∣2m:=\frac{|\alpha|+|\beta|}{2}, ε:=∣β∣−∣α∣\varepsilon:=|\beta|-|\alpha| and ξ:=μV⊞T(m)\xi:=\mu_{\mathbf{V}}\boxplus T(m). Then, by (A.5), we have

In particular, the integral ∫f(x) dμV(α)(x)\int f(x)\,d\mu_{\mathbf{V}(\alpha)}(x) is continuous in α\alpha.

By general properties of the Stieltjes transform, the function g(iy,α)g(iy,\alpha) is locally uniformly continuous in yy, uniformly in α\alpha. Thus, it remains to show that the function g(iy,α)g(iy,\alpha) is continuous in α\alpha. This follows by taking f(x):=1x−iyf(x):=\frac{1}{x-iy} in (A.6), with y>0y>0 fixed.

By general properties of the Stieltjes transform, the function g(iy,α)g(iy,\alpha) is differentiable with respect to yy, with derivative

It therefore follows by the same arguments as in the preceding paragraph that ∂g∂y(iy,α)\frac{\partial g}{\partial y}(iy,\alpha) is continuous.

Suppose that the bivariate power series converge on the set of all (z,α)(z,\alpha) with ∣z−z0∣<ε|z-z_{0}|<\varepsilon and ∣α−α0∣<ε|\alpha-\alpha_{0}|<\varepsilon, where ε=ε(z0,α0)>0\varepsilon=\varepsilon(z_{0},\alpha_{0})>0.

Let us investigate the growth of the coefficients, and hence the radius of convergence, of the univariate power series in zz. Since

Since for fixed α\alpha, g(z,α)g(z,\alpha) is a non-constant analytic function in a certain open set containing the upper imaginary half-axis, there exists an at most countable set Yα={y1,y2,y3,…}Y_{\alpha}=\{y_{1},y_{2},y_{3},\ldots\} such that for all y∉Yαy\not\in Y_{\alpha},

For y∉Yαy\not\in Y_{\alpha}, differentiating the second equation in (6.1) with respect to yy, we get

where S~V(z):=zSV(z)\widetilde{S}_{\mathbf{V}}(z):=zS_{\mathbf{V}}(z).

Now fix (iy0,α0)(iy_{0},\alpha_{0}) with α0≠0\alpha_{0}\neq 0 and y0∉Yα0y_{0}\not\in Y_{\alpha_{0}}. Then there exists a small neighborhood NN such that for (iy,α)∈N(iy,\alpha)\in N, we have (A.7), (A.8), and

and the sign of the square-root is constant in NN. Note that FF is an analytic function with

Moreover, comparing (A.8) and (A.9), we see that

It therefore follows from the implicit function theorem for real-analytic functions that there exists a small neighborhood N~⊂N\widetilde{N}\subset N of the point (iy0,α0)(iy_{0},\alpha_{0}) such that g(iy,α)g(iy,\alpha), the solution to the equation F(ζ,iy,α)=0F(\zeta,iy,\alpha)=0, is analytic on N~\widetilde{N}, with gradient

Equation (7.4) now follows by a straightforward calculation.

This follows from Lemma 6.4 and the subsequent Remark 6.5.

Let KK be a compact set as in Assumption 7.3, and let α,β∈K\alpha,\beta\in K. For fixed C>0C>0, consider the function f(y):=log⁡(1+y2/C2)f(y):=\log(1+y^{2}/C^{2}). Since f′(y)=2yC2+y2f^{\prime}(y)=\frac{2y}{C^{2}+y^{2}}, this function satisfies the conditions of Equation (A.6), and we obtain

from which Assumption 7.3 follows immediately. ∎

We thank Peter J. Forrester for pointing out some relevant references.

References