The single ring theorem

Alice Guionnet, Manjunath Krishnapur, Ofer Zeitouni

The problem

Horn asked the question of describing the eigenvalues of a square matrix with prescribed singular values. If AA is a n×nn\times n matrix with singular values s1≥…≥sn≥0s_{1}\geq\ldots\geq s_{n}\geq 0 and eigenvalues λ1,…,λn\lambda_{1},\ldots,\lambda_{n} in decreasing order of absolute values, then the inequalities

were shown by Weyl to hold. Horn established that these were all the relationships between singular values and eigenvalues.

In this paper we study the natural probabilistic version of this problem and show that for “typical matrices”, the singular values almost determine the eigenvalues. To frame the problem precisely, fix s1≥…≥sn≥0s_{1}\geq\ldots\geq s_{n}\geq 0 and consider n×nn\times n matrices with these singular values. They are of the form A=PTQA=PTQ, where TT is diagonal with entries sjs_{j} on the diagonal, and P,QP,Q are arbitrary unitary matrices.

We make AA into a random matrix by choosing PP and QQ independently from Haar measure on U(n)\mathcal{U}(n), the unitary group of n×nn\times n matrices, and independent from TT. Let λ1,…,λn\lambda_{1},\ldots,\lambda_{n} be the (random) eigenvalues of AA. The following natural questions arise.

Are there deterministic or random sets {sj}\{s_{j}\}, for which one can find the exact distribution of {λj}\{\lambda_{j}\}?

For finite nn, for fixed SS, is LΛL_{\Lambda} concentrated in the space of probability measures on the plane?

In this paper, we concentrate on the second question and answer it in the affirmative, albeit with some restrictions. In this context, we note that Fyodorov and Wei [8, Theorem 2.1] gave a formula for the mean eigenvalues density of AA, yet in terms of a large sum which does not offer an easy handle on asymptotic properties (see also for the case where TT is a projection). The authors of explicitely state the second question as an open problem.

Of course, questions 1–3. above are not new, and have been studied in various formulations. We now describe a partial and necessarily brief history of what is known concerning questions 1. and 2.; partial results concerning question 3. will be discussed elsewhere.

The most famous case of a positive answer to question 1. is the Ginibre ensemble, see , and its asymmetric variant, see . (There are some pitfalls in the standard derivation of Ginibre’s result. We refer to for a discussion.) Another situation is the truncation of random unitary matrices, described in .

Concerning question 2., the convergence of the empirical measure of eigenvalues in the Ginibre ensemble (and other ensembles related to question 1.) is easy to deduce from the explicit formula for the joint distribution of eigenvalues. Generalizations of this convergence in the absence of such explicit formula, for matrices with iid entries, is covered under Girko’s circular law, which is described in ; the circular law was proved under some conditions in and finally, in full generality, in and . Such matrices, however, do not possess the invariance properties discussed in connection of question 2. The single ring theorem of Feinberg and Zee is, to our knowledge, the first example where a partial answer to this question is offered. (Various issues of convergence are glossed over in and, as it turns out, require a significant effort to overcome.) As we will see in Section 3, the asymptotics of the spectral measure appearing in question 2. are described by the Brown measure of RR-diagonal operators. (The Brown measure is a continuous analogue of the spectral distribution of non-normal operators, introduced in .) RR-diagonal operators were introduced by Nica and Speicher in the context of free probability; they represent the weak*-limit (or more precisely, the limit in ∗*-moments) of operators of the form UTUT with UU unitary with size going to infinity and TT diagonal, and were intensively studied in the last decade within the theory of free probability, in particular in connection with the problem of classifying invariant subspaces .

Limiting spectral density of a non-normal matrix

GμG_{\mu} is analytic off the support of μ\mu. We let Hn{\mathcal{H}}_{n} denote the Haar measure on the nn-dimensional unitary group U(n)\mathcal{U}(n). Let {Pn,Qn}n≥1\{P_{n},Q_{n}\}_{n\geq 1} denote a sequence of independent, Hn{\mathcal{H}}_{n}-distributed matrices. Let TnT_{n} denote a sequence of diagonal matrices, independent of (Pn,Qn)(P_{n},Q_{n}), with real positive entries Sn={si(n)}S_{n}=\{s^{(n)}_{i}\} on the diagonal, and introduce the empirical measure of the symmetrized version of TnT_{n} as

Let An=PnTnQnA_{n}=P_{n}T_{n}Q_{n}, let Λn={λi(n)}\Lambda_{n}=\{\lambda_{i}^{(n)}\} denote the set of eigenvalues of AnA_{n}, and set

There exist constants κ,κ1>0\kappa,\kappa_{1}>0 such that

LAnL_{A_{n}} converges in probability to a limiting probability measure μA\mu_{A}.

The support of μA\mu_{A} is a single ring: there exist constants 0≤a<b<∞0\leq a<b<\infty so that

Further, a=0a=0 if and only if ∫x−2dΘ(x)=∞\int x^{-2}d\Theta(x)=\infty.

See Remark 7 for an explicit characterization of the free convolution appearing in Theorem 1, and [1, Ch. 5] for general background. A different characterization of ρA\rho_{A}, borrowed from and instrumental in the proof of part (c) of Theorem 1, is provided in Remark 8 in Section 3.1.

As a corollary of Theorem 1, we prove the Feinberg-Zee “single ring theorem”.

Let VV denote a polynomial with positive leading coefficient. Let the nn-by-nn complex matrix XnX_{n} be distributed according to the law

where ZnZ_{n} is a normalization constant and dXdX the Lebesgue measure on nn-by-nn complex matrices. Let LXnL_{X_{n}} be the ESD of XnX_{n}. Then {LXn}n\{L_{X_{n}}\}_{n} satisfies the conclusions of Theorem 1 with Θ\Theta the unique minimizer of the functional

Theorem 3 will follow by checking that the assumptions of Theorem 1 are satisfied for the spectral decomposition Xn=UnTnVnX_{n}=U_{n}T_{n}V_{n}, see Section 6.

The second hypothesis in Theorem 1 may seem difficult to verify in general; we show in the next proposition that adding a small Gaussian matrix guarantees it.

Let (Tn)n≥0(T_{n})_{n\geq 0} be a sequence of matrices satisfying the assumptions of Theorem 1 except for (3) and assume that ∥Tn−1∥\|T_{n}^{-1}\| is uniformly bounded. Let NnN_{n} be a n×nn\times n matrix with independent (complex) Gaussian entries of zero mean and covariance equal identity. Let Un,VnU_{n},V_{n} follow the Haar measure on unitary n×nn\times n matrices, independently of Tn,NnT_{n},N_{n}. Then, the empirical measure of the eigenvalues of Yn:=UnTnVn+n−γNnY_{n}:=U_{n}T_{n}V_{n}+n^{-\gamma}N_{n} converges weakly in probability to μA\mu_{A} as in Theorem 1 for any γ∈(12,∞)\gamma\in(\frac{1}{2},\infty).

satisfies the hypotheses of Proposition 4.

A rather straightforward generalization of Theorem 1 concerns the limiting spectral measure of Pn+BnP_{n}+B_{n}, where PnP_{n} is Hn{\mathcal{H}}_{n} distributed and the sequence of n×nn\times n matrices BnB_{n} converges in ∗*-moments to an operator bb in a non-commutative probability space (A,τ)({\mathcal{A}},\tau). (The latter means that for all polynomial PP in two non-commutative variables,

An example of matrices BnB_{n} which satisfy the hypotheses of Proposition 6 is given by the diagonal matrices Bn=\mboxdiag(s1n,…,snn)B_{n}=\mbox{diag}(s_{1}^{n},\ldots,s_{n}^{n}) with entries sins_{i}^{n} satisfying the hypotheses of Example 5. This is easily verified from the fact that the eigenvalues of Dn(w)D_{n}(w) are given by (∣w−s1n∣,…,∣w−snn∣)(|w-s_{1}^{n}|,\ldots,|w-s_{n}^{n}|).

The main difficulty in studying the ESD LAnL_{A_{n}} is that AnA_{n} is not a normal matrix, that is AnAn∗≠An∗AnA_{n}A^{*}_{n}\not=A^{*}_{n}A_{n}, almost surely. For normal matrices, the limit of ESDs can be found by the method of moments or by the method of Stieltjes’ transforms. For non-normal matrices, the only known method of proof is more indirect and follows an idea of Girko that we describe now (the details are a little different from what is presented in Girko or Bai ).

From Green’s formula, for any polynomial P(z)=∏j=1n(z−λj)P(z)=\prod_{j=1}^{n}(z-\lambda_{j}), we have

It will be convenient for us to introduce the 2n×2n2n\times 2n matrix

It may be checked easily that eigenvalues of HnzH_{n}^{z} are the positive and negative of the singular values of zI−AnzI-A_{n}. Therefore, if we let νnz\nu_{n}^{z} denote the ESD of HnzH_{n}^{z},

This is Girko’s formula in a different form and its utility lies in the following attack on finding the limit of LAnL_{A_{n}}.

Justify that ∫log⁡∣x∣dνnz(x)→∫log⁡∣x∣dνz(x)\int\log|x|d\nu_{n}^{z}(x)\rightarrow\int\log|x|d\nu^{z}(x) for (almost every) zz. But for the fact that “log⁡\log” is not a bounded function, this would have followed from the weak convergence of νnz\nu_{n}^{z} to νz\nu^{z}. As it stands, this is the hardest technical part of the proof.

A standard weak convergence argument is then used in order to convert the convergence for (almost every) zz of νnz\nu_{n}^{z} to a convergence of integrals over zz. Indeed, setting h(z):=∫log⁡∣x∣dνz(x)h(z):=\int\log|x|d\nu^{z}(x), we will get from (6) that

Show that hh is smooth enough so that one can integrate the previous equation by parts to get

which identifies Δh(z)\Delta h(z) as the density (with respect to Lebesgue measure) of the limit of LAnL_{A_{n}}.

Identify the function hh sufficiently precisely to be able to deduce properties of Δh(z)\Delta h(z). In particular, show the single ring phenomenon, which states that the support of the limiting spectral measure is a single annulus (the surprising part being that it cannot consist of several disjoint annuli).

Girko’s equation (6) and these five steps give a general recipe for finding limiting spectral measures of non-normal random matrices. Whether one can overcome the technical difficulties depends on the model of random matrix one chooses. For the model of random matrices with i.i.d. entries having zero mean and finite variance, this has been achieved in stages by Bai , Götze and Tikhomirov , Pan and Zhou and Tao and Vu . While we heavily borrow from that sequence, a major difficulty in the problem considered here is that there is no independence between entries of the matrix AnA_{n}. Instead, we will rely on properties of the Haar measure, and in particular on considerations borrowed from free probability and the so called Schwinger–Dyson (or master-loop) equations. Such equations were already the key to obtaining fine estimates on the Stieltjes transform of Gaussian generalized band matrices in . In , they were used to study the asymptotics of matrix models on the unitary group. Our approach combines ideas of to estimate Stieltjes transforms and the necessary adaptations to unitary matrices as developped in . The main observation is that one can reduce attention to the study of the ESD of matrices of the form (T+U)(T+U)∗(T+U)(T+U)^{*} where TT is real diagonal and UU is Haar distributed. In the limit (i.e., when TT and UU are replaced by operators in a C∗C^{*}-algebra that are freely independent, with TT bounded and self adjoint and UU unitary), the limit ESD has been identified by Haagerup and Larsen . The Schwinger–Dyson equations give both a characterization of the limit and, more important to us, a discrete approximation that can be used to estimate the discrepancy between the pre-limit ESD and its limit. These estimates play a crucial role in integrating the singularity of the log in Step two above, but only once an a-priori (polynomial) estimate on the minimal singular value has been obtained. The latter is deduced from assumption 3. In the context of the Feinberg–Zee single ring theorem, the latter assumption holds due to an adaptation of the analysis of .

Notation

We describe our convention concerning constants. Throughout, by the word constant we mean quantities that are independent of nn (or of the complex variables zz, z1z_{1}). Generic constants denoted by the letters CC,cc or RR, have values that may change from line to line, and they may depend on other parameters. Constants denoted by CiC_{i}, KK, κ\kappa and κ′\kappa^{\prime} are fixed and do not change from line to line.

where Wnz=z‾QnPn/∣z∣W_{n}^{z}=\overline{z}Q_{n}P_{n}/|z| is unitary and Hn\mathcal{H}_{n} distributed. Throughout, we will write ρ=∣z∣\rho=|z|. We also will assume in this section that the sequence TnT_{n} is deterministic. We are thus led to the study of the ESD for a sequence of matrices of the form

with Bn=ρUn+TnB_{n}=\rho U_{n}+T_{n}, TnT_{n} being a real, diagonal matrix of uniformly bounded norm, and UnU_{n} a Hn{\mathcal{H}}_{n} unitary matrix. Because ∥Tn∥\|T_{n}\| is uniformly bounded, it will be enough to consider throughout ρ\rho uniformly bounded.

where we used the notation A⊗B♯C=ACBA\otimes B\sharp C=ACB.

By the invariance of μ\mu under unitary conjugation, see [27, Proposition 5.17] or [1, (5.4.31)], we have the Schwinger–Dyson equation

We continue to use the notation Y{\bf Y}, U,U∗{\bf U},{\bf U^{*}} and T{\bf T} in a way similar to (17) and (18). So, we let Y=ρ(U+U∗)+T{\bf Y}=\rho({\bf U}+{\bf U}^{*})+{\bf T} with

We extend μ\mu to the algebra generated by U,U∗{\bf U},{\bf U}^{*} and T{\bf T} by putting for any A,B,C,D∈AA,B,C,D\in{\cal A},

Observe that this extension is still tracial.

The non-commutative derivative ∂\partial in (20) extends naturally to the algebra generated by the matrix-valued U,U∗,T{\bf U},{\bf U}^{*},{\bf T}, using the Leibniz rule (19) together with the relations

where A(B⊗C)D=(AB)⊗(CD)A(B\otimes C)D=(AB)\otimes(CD). Further, (21) extends also in this context.

We apply the derivative ∂\partial to the analytic function P=(z1−Y)−1(z2−T)−1U ,P=(z_{1}-{\bf Y})^{-1}(z_{2}-{\bf T})^{-1}{\bf U}\,, while noticing that, by (19) and (23),

Applying (21), with μ(P)=GU(z1,z2)\mu(P)=G_{U}(z_{1},z_{2}) and μ(p)=1/2\mu(p)=1/2, we find

Note that Pp=PPp=P and thus μ(pP)=μ(P)\mu(pP)=\mu(P). Further, for any smooth function QQ, μ(U∗QU)\mu({\bf U}^{*}Q{\bf U}) equals μ((1−p)Q)\mu((1-p)Q) due to the traciality of μ\mu and UU∗=1−p{\bf U}{\bf U}^{*}=1-p. By symmetry (note that (1−p)(z1−Y)−1(z2−T)−1(1-p)(z_{1}-{\bf Y})^{-1}(z_{2}-{\bf T})^{-1} and p(z1−Y)−1(z2−T)−1p(z_{1}-{\bf Y})^{-1}(z_{2}-{\bf T})^{-1} are given by the same formula up to replacing (U,U∗)({\bf U},{\bf U}^{*}) by (U∗,U)({\bf U}^{*},{\bf U}), which has the same law) we get that μ(U∗P)\mu({\bf U}^{*}P) equals

The first equality holds without the last factor (z2−T)−1(z_{2}-{\bf T})^{-1}, thus implying that μ((z1−Y)−1p)=μ((z1−Y)−1)/2=G(z1)/2\mu((z_{1}-{\bf Y})^{-1}p)=\mu((z_{1}-{\bf Y})^{-1})/2=G(z_{1})/2 and so we get from (26) that

Noticing that GU(z1)G_{U}(z_{1}) is the limit of z2GU(z1,z2)z_{2}G_{U}(z_{1},z_{2}) as z2→∞z_{2}\to\infty, we find by (28) that

and therefore, as GU(z1)G_{U}(z_{1}) goes to zero as z1→∞z_{1}\to\infty,

(Again, here and in the rest of this subsection, the proper branch of the square root is determined by analyticity.) Let RρR_{\rho} denote the RR-transform of the Bernoulli law λρ:=(δ−ρ+δ+ρ)/2\lambda_{\rho}:=(\delta_{-\rho}+\delta_{+\rho})/2, that is,

see [1, Definition 5.3.22 and Exercise 5.3.27], so that we have

Repeating the computation with GU∗G_{U^{*}}, we have GU∗=GUG_{U^{*}}=G_{U}. Algebraic manipulations yield

Therefore, we get by substituting (31) and (32) into (33) that

and has an analytic continuation to a neighborhood of D{\cal D}, and F′>0F^{\prime}>0 on D{\cal D}. Further, with μA\mu_{A} as above, ρA(reiθ)=ρA(r)\rho_{A}(re^{i\theta})=\rho_{A}(r) and it holds that

Finally, ρA\rho_{A} has an analytic continuation to a neighborhood of (a,b](a,b], and μA\mu_{A} is a probability measure, see [13, Pg 333].

In the next section, we will need the following estimate.

If ∣ℑGT(⋅)∣≤κ1|\Im G_{T}(\cdot)|\leq\kappa_{1} on {z:ℑ(z)≥ϵ}\{z:\Im(z)\geq\epsilon\} then ∣ℑG(⋅)∣≤κ1|\Im G(\cdot)|\leq\kappa_{1} on {z:ℑ(z)≥ϵ}\{z:\Im(z)\geq\epsilon\}.

2 Finite n𝑛n equations and convergence

We next turn to the evaluation of the law of Yn{\bf Y}_{n}. We assume throughout that the sequence TnT_{n} is uniformly bounded by some constant MM, that LTn→μTL_{T_{n}}\to\mu_{T} weakly in probability, and further that (4) is satisfied. All constants in this section are independent of ρ\rho, but depend implicitly on MM, the uniform bound on ∥Tn∥\|T_{n}\| and on ρ\rho.

We get by taking P=(z1−Yn)−1(z2−Tn)−1UnP=(z_{1}-{\bf Y}_{n})^{-1}(z_{2}-{\bf T}_{n})^{-1}{\bf U}_{n} that

with ∥P∥L\|P\|_{L} the Lipschitz constant of PP given by

if DD is the cyclic derivative given by D=m∘∂D=m\circ\partial with m(A⊗B)=BAm(A\otimes B)=BA and ∥DP∥∞\|DP\|_{\infty} denotes the operator norm. (The appearance of the cyclic derivative in the evaluation of the Lipshitz constant can be seen by approximating PP by polynomials.) Applying (41) to each term of ∂P\partial P (recall formula (25)), we get that for ℑ(z1),ℑ(z2)>0\Im(z_{1}),\Im(z_{2})>0, and with a∧b=min⁡(a,b)a\wedge b=\min(a,b),

(The inequality uses that for any Hermitian matrix, ∥(z−H)−1∥∞≤1/∣ℑ(z)∣\|(z-H)^{-1}\|_{\infty}\leq 1/|\Im(z)|.) Multiplying by z2z_{2} and taking the limit as z2→∞z_{2}\to\infty we deduce from (40) that

with again the choice of the square root determined by analyticity and behavior at infinity.

and therefore, as (ℜ(z)−x)2+ℑ(z)2≤4R2(\Re(z)-x)^{2}+\Im(z)^{2}\leq 4R^{2} for all z,x∈B(0,R)z,x\in B(0,R)

Moreover, since ∣GUn(z)∣≤1/∣ℑ(z)∣|G_{U}^{n}(z)|\leq 1/|\Im(z)|, we deduce from (42) that for some constant cc independent of nn and all nn large,

Combining this estimate and (48), we get that

as soon as ℑ(z)>C1n−1/3\Im(z)>C_{1}n^{-1/3} for an appropriate C1C_{1}, and ∣z∣<R|z|<R. The conclusion follows.

There exists a constant C3C_{3} such that if ℑ(z)>C3n−1/4\Im(z)>C_{3}n^{-1/4}, then ℑ(ψn(z))≥ℑ(z)/2.\Im(\psi_{n}(z))\geq\Im(z)/2.

Then, for any z∈Δnz\in\Delta_{n}, and whatever choice of branch of the square root made in (43), if en−1/2O1(n,z){\bf e}_{n}^{-1/2}O_{1}(n,z) is small enough (smaller than en/2{\bf e}_{n}/2 is fine), then that choice can be extended to include a neighborhood of the point w=Gn(z)w=G^{n}(z) such that with this choice, the function rρ(w)=14ρ(−1+1+4ρ2w2)r_{\rho}(w)=\frac{1}{4\rho}(-1+\sqrt{1+4\rho^{2}w^{2}}) is Lipschitz in the sense that

Combining the last display with the relation Rρ(θ)=2rρ(θ)/θR_{\rho}(\theta)=2r_{\rho}(\theta)/\theta, (50) and (48), one obtains that for z∈Δnz\in\Delta_{n},

Since the above right hand side is smaller than ℑ(z)/2\Im(z)/2 for ℑ(z)>n−1/4\Im(z)>n^{-1/4}, we conclude that for z∈Δn∩{ℑ(z)>n−1/4}z\in\Delta_{n}\cap\{\Im(z)>n^{-1/4}\}

as, regardless of the branch taken in the definition of Rρ(⋅)R_{\rho}(\cdot), ℑRρ(Gn(z))≤0\Im R_{\rho}(G^{n}(z))\leq 0.

We thus conclude from the last display and (51) the existence of a constant C3C_{3} such that if ℑ(z)>C3n−1/4\Im(z)>C_{3}n^{-1/4} then

We have made all preparatory steps in order to state the main result of this subsection.

There exist positive finite constants C6,C7,C8C_{6},C_{7},C_{8} such that, for n>C6n>C_{6} and all z∈En:={z:ℑ(z)>n−C7}z\in{\cal E}_{n}:=\{z:\Im(z)>n^{-C_{7}}\},

Proof This is immediate from Lemma 11, Lemma 12, the definition of ψn\psi_{n}, the assumption (4) on GTnG_{T_{n}}, and the equality (46). ∎

in probability. (ii) Fix R>0R>0. For any smooth compactly supported deterministic function φ\varphi on BRB_{R},

Before bringing the proof of Proposition 14, we recall the following elementary lemma.

We can now provide the Proof of Proposition 14

(i) Assume z∈BRz\in B_{R} for some R>0R>0. By (3), we can replace the lower limit of integration in (53) with n−δn^{-\delta}. Let GnzG_{n}^{z} denote the Stieltjes transform of E[νnz]E[\nu_{n}^{z}]. By Lemma 13 and Lemma 9, there exist positive constants c1=c1(R),c2=c2(R)c_{1}=c_{1}(R),c_{2}=c_{2}(R) such that whenever ℑ(u)>n−c1\Im(u)>n^{-c_{1}}, it holds that ∣ℑGnz(u)∣<c2|\Im G_{n}^{z}(u)|<c_{2}. We may and will assume that c1<δc_{1}<\delta.

Since GnzG_{n}^{z} is the Stieltjes transform of E[νnz]E[\nu_{n}^{z}], by Lemma 15, we have for any y>0y>0 that

Thus, we get that for any z∈BRz\in B_{R} and with α∈\alpha\in,

where 2J−1n−c1<ϵ≤2Jn−c12^{J-1}n^{-c_{1}}<\epsilon\leq 2^{J}n^{-c_{1}}. Note that by Lemma 15 and the estimate on GnzG_{n}^{z}, for j≥0j\geq 0,

where the constant C=C(R)C=C(R). To obtain the estimate (53), we will consider α=1\alpha=1 and argue as follows. Due to (3), for α<2\alpha<2 we have

by Hölder’s inequality. The first factor goes to zero because

By (3), the second factor is bounded by (δ′)α/2(\delta^{\prime})^{\alpha/2}. We thus get (53) from (57). By Chebycheff’s inequality, the convergence in expectation implies the convergence in probability and therefore for any δ,δ′>0\delta,\delta^{\prime}>0 there exists ϵ>0\epsilon>0 small enough so that

On the other hand, ∫ϵ∞log⁡∣x∣dνnz(x)\int_{\epsilon}^{\infty}\log|x|d\nu_{n}^{z}(x) converges to ∫ϵ∞log⁡∣x∣dνz(x)\int_{\epsilon}^{\infty}\log|x|d\nu^{z}(x) by the weak convergence of νnz\nu_{n}^{z} to νz\nu^{z} in probability for any ϵ>0\epsilon>0, and ∫0ϵlog⁡∣x∣dνz(x)\int_{0}^{\epsilon}\log|x|d\nu^{z}(x) converges to as ϵ→0\epsilon\to 0 since νz\nu^{z} has a bounded density by Lemma 9. Hence, we get (54).

and set fn(z)=fn1(z)+fn2(z)f_{n}(z)=f_{n}^{1}(z)+f_{n}^{2}(z). Because νnz\nu_{n}^{z} is supported in BR+MB_{R+M} on ∥Tn∥≤M\|T_{n}\|\leq M for all z∈BRz\in B_{R}, fnf_{n} is bounded above by log⁡(R+M)\log(R+M). By (57), E[∣fn2(⋅)∣2E[|f_{n}^{2}(\cdot)|^{2} is bounded, uniformly in z∈BRz\in B_{R}. On the other hand, by (3), again uniformly in z∈BRz\in B_{R}, E(fn1(z)2)<δ′E(f_{n}^{1}(z)^{2})<\delta^{\prime}, and therefore

is bounded in probability. This uniform integrability and the weak convergence (54) are enough to conclude, using dominated convergence (see [25, Lemma 3.1] for a similar argument). ∎

Proof of Theorem 1

in probability. Since the sequence LAnL_{A_{n}} is tight, it thus follows that it converges, in the sense of distribution, to the measure

From Remark 8 (based on [13, Corollary 4.5]), we have that μA\mu_{A} is a probability measure that possesses a radially symmetric density ρA\rho_{A} satisfying the properties stated in parts b and c of the theorem. ∎

Proof of Theorem 3

which goes to zero for κ∈(0,(1−2κ′)/4)\kappa\in(0,(1-2\kappa^{\prime})/4). Together with [22, Equation (2.32)], this proves point 3 of the assumptions. Thus, it remains only to check point 2 of the assumptions. Toward this end, define Gn={σ1n<M+1}{\cal G}_{n}=\{\sigma_{1}^{n}<M+1\} and note that we may and will restrict attention to ∣z∣<M+2|z|<M+2 when checking (3). We begin with the following proposition, due to .

Let A‾\overline{A} be an arbitrary nn-by-nn matrix, and let A=A‾+σNA=\overline{A}+\sigma N where NN is a matrix with independent (complex) Gaussian entries of zero mean and unit variances. Let σn(A)\sigma_{n}(A) denote the minimal singular value of AA. Then, there exists a constant C12C_{12} independent of A‾\overline{A}, σ\sigma or nn such that

The proof of Proposition 16 is identical to [23, Theorem 3.3], with the required adaptation in moving from real to complex entries. (Specifically, in the right side of the display in [23, Lemma A.2], ϵ2/π/σ\epsilon\sqrt{2/\pi}/\sigma is replaced by its square.) We omit further details.

On the event Gn{\cal G}_{n}, all entries of the matrix XnX_{n} are bounded by a constant multiple of n\sqrt{n}. Let NnN_{n} be a Gaussian matrix as in Proposition 16. With α>2\alpha>2 a constant to be determined below, set

Turning to the construction, observe first that from (59),

Let Xn(α)=Xn+n−αNn1Gn′X_{n}^{(\alpha)}=X_{n}+n^{-\alpha}N_{n}{\bf 1}_{{\cal G}_{n}^{\prime}}. Let {θi}\{\theta_{i}\} and {μi}\{\mu_{i}\} denote the eigenvalues of Wn=XnXn∗W_{n}=X_{n}X_{n}^{*} and of Wn(α)=(Xn(α))(Xn(α))∗W_{n}^{(\alpha)}=(X_{n}^{(\alpha)})(X_{n}^{(\alpha)})^{*}, respectively, arranged in decreasing order. Note that the density of XnX_{n} is of the form

where the variable x={xi,j}1≤i,j≤n{\bf x}=\{x_{i,j}\}_{1\leq i,j\leq n} is matrix valued and dx=∏1≤i,j≤ndxi,jd{\bf x}=\prod_{1\leq i,j\leq n}dx_{i,j}, while that of Xn(α)X_{n}^{(\alpha)} is of the form

where ENE_{N} denotes expectation with respect to the law of NnN_{n}, and ZnZ_{n} is the same in both expressions. Note that σ1(Xn(α))∈[σ1(Xn)−1,σ1(Xn)+1]\sigma_{1}(X_{n}^{(\alpha)})\in[\sigma_{1}(X_{n})-1,\sigma_{1}(X_{n})+1]. Because V(⋅)V(\cdot) is locally Lipschitz, we have that if either σ1(Xn)≤M+1\sigma_{1}(X_{n})\leq M+1 or σ1(Xn(α))≤M+1\sigma_{1}(X_{n}^{(\alpha)})\leq M+1, then there exists a constant C13C_{13} independent of α\alpha so that

where the Cauchy–Schwarz inequality was used in the third inequality and the Hoffman–Wielandt inequality in the next (see e.g. [1, Lemma 2.1.19]). On the event Gn{\cal G}_{n}, all entries of Wn−WnαW_{n}-W_{n}^{\alpha} are bounded by n(3−α)/2n^{(3-\alpha)/2}. Therefore,

where the constant C14C_{14} does not depend on α\alpha. In particular, if α>(C14+1)∨2\alpha>(C_{14}+1)\vee 2 we obtain that on Gn{\cal G}_{n}, the ratio of the functions fn=e−n\mboxtr(V(Wn))f_{n}=e^{-n{\mbox{\rm tr}}(V(W_{n}))} and gn=e−n\mboxtr(V(Wn(α)))g_{n}=e^{-n{\mbox{\rm tr}}(V(W_{n}^{(\alpha)}))} is bounded e.g. by 1+n(C14+1−α)/21+n^{(C_{14}+1-\alpha)/2}; in particular, it holds that

Therefore, the variational distance between the law of XnX_{n} conditioned on σ1(Xn)<M\sigma_{1}(X_{n})<M and that of Xn(α)X_{n}^{(\alpha)} conditioned on σ1(Xn(α))<M\sigma_{1}(X_{n}^{(\alpha)})<M, is bounded by

It follows that one can construct a matrix YnY_{n} of law identical to the law of Xn(α)X_{n}^{(\alpha)} conditioned on σ1(Xnα)<M\sigma_{1}(X_{n}^{\alpha})<M, together with XnX_{n}, on the same probability space so that

where α\alpha was chosen as function of xx. This yields immediately point 2 of the assumptions of Theorem 1, if δ>5C17/2\delta>5C_{17}/2.

We have checked now that in the setup of Theorem 3, all the assumptions of Theorem 1 hold. Applying now the latter theorem completes the proof of Theorem 3. ∎

The proof of Theorem 3 carries over to more general situations; indeed, VV does not need to be a polynomial, it is enough that its growth at infinity is polynomial and that it is locally Lipschitz, so that the results of still apply. We omit further details.

Proof of Proposition 4

with C(∥Tn−1∥,∥Tn∥)C(\|T_{n}^{-1}\|,\|T_{n}\|) a finite constant depending only on ∥Tn−1∥,∥Tn∥\|T_{n}^{-1}\|,\|T_{n}\| which we assumed bounded. (In deriving the last estimate, we used that ∥(I+B)1/2−I∥≤∥B∥\|(I+B)^{1/2}-I\|\leq\|B\| when ∥B∥<1/2\|B\|<1/2.) As a consequence, the third condition is satisfied since

with γ′=min⁡{κ,12(γ−12)}\gamma^{\prime}=\min\{\kappa,\frac{1}{2}(\gamma-\frac{1}{2})\} and ℑ(z)≥n−max⁡{12(γ−12),κ′}\Im(z)\geq n^{-\max\{\frac{1}{2}(\gamma-\frac{1}{2}),\kappa^{\prime}\}}. Hence, the results of Lemma 13 hold and we need only check, as in Proposition 14, that with νnz\nu_{n}^{z} the empirical measure of the singular values of zI−YnzI-Y_{n},

where we finally used that F−1F^{-1} is Hölder continuous with index α\alpha. ∎

Extension to orthogonal conjugation

In this section, we generalize Theorem 1 to the case where we conjugate TnT_{n} by orthogonal matrices instead of unitary matrices.

Proof of Proposition 6

whereas our hypotheses allow us to bound uniformly the Stieltjes transform of νzn\nu^{n}_{z} on {z1:ℑ(z1)≥n−C7}\{z_{1}:\Im(z_{1})\geq n^{-C_{7}}\} as in Lemma 13, hence providing a control of the integral on the interval [n−C7,ϵ][n^{-C_{7}},\epsilon]. The control of the integral for x<n−C7x<n^{-C_{7}} uses a regularization by the Gaussian matrix n−γNnn^{-\gamma}N_{n} as in Proposition 4 .∎

Acknowledgments: We thank Greg Anderson for many fruitful and encouraging discussions. We thank Yan Fyodorov for pointing out the paper and Philippe Biane for suggesting that our technique could be applied to the examples in . We thank the referee for a careful reading of the manuscript.

References