Circular law, Extreme Singular values and Potential theory

Guangming Pan, Wang Zhou

Introduction

Let {XkjX_{kj}}, k,j=⋯ ,k,j=\cdots, be a double array of independent and identically distributed (i.i.d.) complex random variables (r.v.’s) with EX11=0EX_{11}=0 and E∣X11∣2=1E|X_{11}|^{2}=1. The complex eigenvalues of the matrix n−1/2X=n−1/2(Xkj)n^{-1/2}{\bf X}=n^{-1/2}(X_{kj}) are denoted by λ1,⋯ ,λn\lambda_{1},\cdots,\lambda_{n}. The two-dimensional empirical spectral distribution μn(x,y)\mu_{n}(x,y) is defined as

The study of μn(x,y)\mu_{n}(x,y) is related to understanding the random behavior of slow neutron resonances in nuclear physics. See . Since 1950’s it has been conjectured that, under the unit varaince condition, μn(x,y)\mu_{n}(x,y) converges to the so-called circular law, i.e. the uniform distribution over the unit disk in the complex plane. Up to now, this conjecture is only proved in some partial cases.

The first answer for complex normal matrices was given in based on the joint density function of the eigenvalues of n−1/2Xn^{-1/2}{\bf X}. Huang in reported that this result was obtained in an unpublished paper of Silverstein (1984). After more than one decade, Edelman also showed that the expected empirical spectral distribution converges to the circular law for real normal matrices. It is Girko who investigated the circular law for general matrix with independent entries for the first time in . But Girko imposed, not only moment conditions, but also strong smooth conditions on matrix entries. Later on, he further published a series of papers (for example, ) about this problem. However, as pointed out in and , Girko’s argument includes serious mathematical gaps. The rigorous argument of the conjecture was given by Bai in his 1997 celebrated paper for general random matrices. In addition to the finite (4+ε)th(4+\varepsilon)th moment condition Bai still assumed that the joint density of the real and imaginary part of the entries is bounded. Again, the result was further improved by Bai and Silverstein under the assumption E∣X11∣2+η<∞E|X_{11}|^{2+\eta}<\infty in their comprehensive book , but the finiteness condition of the density of matrix entries is still there. Recently, Götze and Tikhomirov gave a proof of the convergence of Eμn(x,y)E\mu_{n}(x,y) to the circular law under the strong moment assumption that the entries have sub-Gaussian tails or are sparsely non-zero instead of the condition about the density of the entries in .

Generally speaking, there are five approaches to studying the spectral distribution of random matrices. The difficulty of the circular conjecture is that the methodologies used in Hermitian matrices do not work well in non-Hermitian ones. There was no powerful tool to attack this conjecture.

1. Moment method. Moments are very important characteristics of r.v.’s. They have many applications in probability and statistics. For example, we have moment estimators in statistics. As far as we know, it is Wigner who introduced moment method into random matrices. Since then, the moment method has been very successful in establishing the convergence of the empirical spectral distribution of Hermitian matrices. Bai did a lot of important work. One can refer to . But moment method fails to work in non-Hermitian ones, because for any complex r.v. ZZ uniformly distributed over any disk centered at , one can verify that for any m≥1m\geq 1

2. Stietjes transform. Another powerful tool in random matrices theory is the Stieltjes transform, which is defined by

for any distribution function G(x)G(x). The basic property of Stieltjes transform is that it is a representing class of probability measures. This property offers one a strong analytic machine. Still see and the references therein. However, the Stieltjes transform of n−1/2Xn^{-1/2}{\bf X} is unbounded if zz coincides with one eigenvalue. So this leads to serious difficulties when dealing with the Stieltjes transform of n−1/2Xn^{-1/2}{\bf X}.

3. Orthogonal polynomials. The study of orthogonal polynomials goes back as far as Hermite. For the deep connections between orthogonal polynomials and random matrices, one can refer to . Orthogonal polynomials are usually limited to Guassian random matrices. Moreover, orthogonal polynomials are only suitable to deriving the spacing between consecutive eigenvalues for large classes of random matrices (see ).

4. Characteristic functions. There is a long history of characteristic functions. In 1810, Laplace used Fourier transform, i.e. characteristic functions to prove central limit theorem for bounded r.v.’s. Then in 1934 P. Lévy reproved Linderberg central limit theorem by characteristic functions. From that time on, characteristic functions are well known to almost every mathematician. Surprisingly, one can not see any application of characteristic functions in random matrices until 1984. Girko combined together the characteristic function of μn(x,y)\mu_{n}(x,y) and the Stieltjes transform, trying to prove the conjecture in . Developing ideas proposed by Girko , Bai reduced the conjecture to estimating the smallest singular value of n−1/2X−zIn^{-1/2}{\bf X}-z{\bf I} in . However, one should note that some uniform estimate of the smallest singular values of n−1/2X−zIn^{-1/2}{\bf X}-z{\bf I} with respect to zz will be required if the method in is employed.

5. Potential theory. Potential theory is the terminology given to the wide area of analysis encompassing such topics as harmonic and subharmonic functions, the boundary problem, harmonic measure, Green’s function, potentials and capacity. Since Doob’s famous book appeared, it is widely accepted that potential theory and probability theory are closely related. For example, superharmonic functions correspond to supermartingales.

The logarithmic potential of a measure μ\mu (see ) is defined by

where μ(t)\mu(t) is any positive finite Borel measure with support in a compact subset of the complex plane. There is also an inversion formula, i.e. μ\mu can be defined through UμU^{\mu} as dμ=−(2π)−1ΔUμd\mu=-(2\pi)^{-1}\Delta U^{\mu}, where Δ\Delta is the two dimensional Laplacian operator. This relation makes Khoruzhenko in suggest to use potential theory to derive the circular law. Then Götze and Tikhomirov in used the logarithmic potential of EμnE\mu_{n} convoluted by a smooth distribution to provide a proof for the convergence of EμnE\mu_{n} to the circular law with entries being sub-Gaussian or sparsely non-zero.

In this paper, the conjecture, the convergence of μn(x,y)\mu_{n}(x,y) to the circular law with probability one, is established under the assumption that the underlying r.v.’s have finite fourth moment. Compared with , we work on the logarithmic potential of μn(x,y)\mu_{n}(x,y) directly, while depends on the logarithmic potential of a convolution of Eμn(x,y)E\mu_{n}(x,y) and the uniform distribution on the disk of radius rr.

The main result of this paper is formulated as follows.

Suppose that {Xjk}\{X_{jk}\} are i.i.d. complex r.v.’s with EX11=0,EX_{11}=0, E∣X11∣2=1E|X_{11}|^{2}=1 and E∣X11∣4<∞E|X_{11}|^{4}<\infty. Then, with probability one, the empirical spectral distribution function μn(x,y)\mu_{n}(x,y) converges to the uniform distribution over the unit disk in two dimensional space.

The bounded density condition in and the sub-Gaussian assumption in are not needed any more.

Theorem 1 will be handled by potential theory in conjunction with estimates for the smallest singular value of n−1/2X−zIn^{-1/2}{\bf X}-z{\bf I}.

The research of the smallest singular values originates from von Neumann and his colleagues. They guessed that

with sn(X)s_{n}({\bf X}) being the smallest singular value of X{\bf X}. Edelman in proved it for random Gaussian matrices, i.e., for each ε≥0\varepsilon\geq 0

Rudelson and Vershynin in solved it for real random matrices, i.e., for every δ>0\delta>0 there exist ε>0\varepsilon>0 and n0n_{0} depending only on δ\delta and the fourth moment of XjkX_{jk} so that

Moreover, since (1.5) fails to hold for the random sign matrices (XjkX_{jk} being symmetric ±1\pm 1 r.v.’s), Spielman and Teng speculated that for random sign matrices for any ε≥0\varepsilon\geq 0

Again, (1.7) has been proved for real random matrices with i.i.d. subgaussian entries in .

We will adapt Rudelson and Vershynin’s method to obtain the order of the smallest singular value for complex matrices perturbed by a constant matrix.

Formally, let W=X+An{\bf W}={\bf X}+{\bf A}_{n}, where An{\bf A}_{n} is a fixed complex matrix and X=(Xjk){\bf X}=(X_{jk}), a random matrix. Denote the singular values of W{\bf W} by s1,⋯ ,sns_{1},\cdots,s_{n} arranged in the non-increasing order. Particularly, the smallest singular value is

where ∥⋅∥2\|\cdot\|_{2} means Euclidean norm, and we denote the spectral norm of a matrix by ∥⋅∥\|\cdot\|.

Let {Xjk}\{X_{jk}\} be i.i.d. complex r.v.’s with EX11=0, E∣X11∣2=1EX_{11}=0,\ E|X_{11}|^{2}=1 and E∣X11∣3<BE|X_{11}|^{3}<B. Let K≥1K\geq 1. Then for every ε≥0\varepsilon\geq 0,

where C>0C>0 and c∈(0,1)c\in(0,1) depend only on KK, BB, E\big{(}Re(X_{11})\big{)}^{2}, E\big{(}Im(X_{11})\big{)}^{2}, and ERe(X11)Im(X11)ERe(X_{11})Im(X_{11}).

In Theorem 2, ε\varepsilon is arbitrary. It can depend on nn. KK is a constant not smaller than 11. In Section 3 when we apply (1.8) in the proof of Theorem 1, we will select ε=n−1−δ, K>4\varepsilon=n^{-1-\delta},\ K>4.

Theorem 2 includes Theorem 5.1 in as a special case, where An=0{\bf A}_{n}=0, the r.v.’s are real and have finite fourth moment. Therefore, (1.6) is true with X{\bf X} replaced by W{\bf W} when X11X_{11} has finite fourth moment and ∥An∥≤Cn\|{\bf A}_{n}\|\leq C\sqrt{n} (C≥0C\geq 0), i.e.,

Moreover, if X{\bf X} is a subgaussian matrix and ∥An∥≤Cn\|{\bf A}_{n}\|\leq C\sqrt{n}, by Lemma 2.4 of or Fact 2.4 of , (1.7) holds with X{\bf X} replaced by W{\bf W}, i.e.,

This exponential rate is better than the polynomial rate in Tao and Vu .

Furthermore, for general random matrices, similar to steps (3.3)-(3.4) in Section 3 one can conclude that

In addition to the assumptions of Theorem 2, suppose that ∣Xij∣≤nεn|X_{ij}|\leq\sqrt{n}\varepsilon_{n} and ∥An∥≤Cn\|{\bf A}_{n}\|\leq C\sqrt{n} with 0≤C<∞0\leq C<\infty, then for any ε≥0\varepsilon\geq 0

where ll is any positive number and εn→0\varepsilon_{n}\rightarrow 0 with the convergence rate slower than any preassigned one as n→∞n\to\infty.

Taking ε=0\varepsilon=0, Corollary 1 then leads to a polynomial bound for the singularity probability:

For random sign matrices Tao and Vu showed that for every A>0A>0 there exists B>0B>0 so that

Recently, Tao and Vu reported a result concerning the smallest singular value of a perturbed matrix too. Under some mild conditions, they proved that

Compared with their results, (1.9) gives an explicit dependence between the bound on sn(W)s_{n}({\bf W}) and probability, while the relationship between AA and BB in and is implicit. In addition, (1.9) holds for general random matrices, while Tao and Vu’s theorem basically applies to discrete random matrices.

In this paper, we will use the letters B,K1,K2B,K_{1},K_{2} to denote some finite absolute constants.

The argument of Theorem 2 is presented in the next section and the proof of the circular law is given in the last section.

Smallest singular value

In this section the smallest singular value of the matrix X{\bf X} perturbed by a constant matrix will be characterized. We begin first with the estimation of the so-called small ball probability.

We first establish a small ball probability for big ε\varepsilon via central limit theorem for complex r.v.’s η1,⋯ ,ηn\eta_{1},\cdots,\eta_{n}. Before we state the next result, let us introduce some more notation and terminology. Re(z)Re(z) and Im(z)Im(z) will denote the real and imaginary part of a complex number zz. Write η1k=Re(ηk),\eta_{1k}=Re(\eta_{k}), η2k=Im(ηk)\eta_{2k}=Im(\eta_{k}), σ12=σ1k2=E(η1k−Eη1k)2, σ22=σ2k2=E(η2k−Eη2k)2, σ12=σ12k=E(η1k−Eη1k)(η2k−Eη2k)\sigma_{1}^{2}=\sigma_{1k}^{2}=E(\eta_{1k}-E\eta_{1k})^{2},\ \sigma_{2}^{2}=\sigma_{2k}^{2}=E(\eta_{2k}-E\eta_{2k})^{2},\ \sigma_{12}=\sigma_{12k}=E(\eta_{1k}-E\eta_{1k})(\eta_{2k}-E\eta_{2k}) for k=1,2,⋯ ,nk=1,2,\cdots,n. For real r.v.’s ξ\xi and η\eta, if \big{(}E(\xi-E\xi)(\eta-E\eta)\big{)}^{2}=E(\xi-E\xi)^{2}E(\eta-E\eta)^{2}>0, then we will say that ξ\xi and η\eta are linearly correlated.

Let η1,⋯ ,ηn\eta_{1},\cdots,\eta_{n} be i.i.d. complex r.v.’s with variances at least 11, E∣η1∣3<BE|\eta_{1}|^{3}<B and let b1,⋯ ,bnb_{1},\cdots,b_{n} be complex numbers such that 0<K1≤∣bk∣≤K20<K_{1}\leq|b_{k}|\leq K_{2} for all kk. Then for every ε>0\varepsilon>0,

where CC is a finite constant depending only on BB, σ1, σ2\sigma_{1},\ \sigma_{2} and σ12\sigma_{12}.

The case where Re(ηk)=0Re(\eta_{k})=0 or Im(ηk)=0Im(\eta_{k})=0 a.s.a.s. follows from Berry-Esseen inequality directly.

Now suppose Re(ηk)Re(\eta_{k}) and Im(ηk)Im(\eta_{k}) are not linearly correlated, and P\big{(}Re(\eta_{k})=0\big{)}<1, P\big{(}Im(\eta_{k})=0\big{)}<1. Let bk=b1k+ib2kb_{k}=b_{1k}+ib_{2k} and v=v1+iv2v=v_{1}+iv_{2}. Define η^1k=b1kη1k−b2kη2k\hat{\eta}_{1k}=b_{1k}\eta_{1k}-b_{2k}\eta_{2k} and η^2k=b1kη2k+b2kη1k\hat{\eta}_{2k}=b_{1k}\eta_{2k}+b_{2k}\eta_{1k}. Obviously, ∑k=1nE∣η^jk−Eη^jk∣3≤∑k=1nE∣bk(ηk−Eηk)∣3≤8B∥b∥33,j=1,2\sum\limits_{k=1}^{n}E|\hat{\eta}_{jk}-E\hat{\eta}_{jk}|^{3}\leq\sum\limits_{k=1}^{n}E|b_{k}(\eta_{k}-E\eta_{k})|^{3}\leq 8B\|b\|_{3}^{3},j=1,2, where ∥b∥33=∑k=1n∣bk∣3\|b\|_{3}^{3}=\sum\limits_{k=1}^{n}|b_{k}|^{3}. In order to apply Berry-Esseen inequality, we need to get a lower bound for E∣η^jk−Eη^jk∣2E|\hat{\eta}_{jk}-E\hat{\eta}_{jk}|^{2}. For j=1j=1, we have

For t∈t\in, let f(t)=(tσ1−1−t2σ2)2+2t1−t2(σ1σ2±σ12)f(t)=(t\sigma_{1}-\sqrt{1-t^{2}}\sigma_{2})^{2}+2t\sqrt{1-t^{2}}(\sigma_{1}\sigma_{2}\pm\sigma_{12}). So the smallest value a=min⁡t∈f(t)a=\min_{t\in}f(t) of f(t)f(t) in $isattaintedatoris attainted at or1orsomeor somet_{0}\in(0,1).Therefore,. Therefore,aisapositiveconstantdependingonlyonis a positive constant depending only on\sigma_{1},\ \sigma_{2}andand\sigma_{12}.Hence. HenceE|\hat{\eta}_{1k}-E\hat{\eta}_{1k}|^{2}\geq a|b_{k}|^{2}.Similarly,. Similarly,E|\hat{\eta}_{2k}-E\hat{\eta}_{2k}|^{2}\geq a|b_{k}|^{2}$. By Berry-Esseen inequality, one can then conclude that

where CC is a constant depending only on BB, σ1, σ2\sigma_{1},\ \sigma_{2} and σ12\sigma_{12}.

Thus (2.4) follows from (2.5), (2.6) and the following inequality

Theorem 3 only yields a polynomial rate n−1/2n^{-1/2}. Next, an improved small ball probability is needed for our future use. To this end, some concepts will be presented which are parallel to those of .

where the subset σ(b)⊆{1,⋯ ,n}\sigma({\bf b})\subseteq\{1,\cdots,n\} is given by {k:K1≤n∣bk∣≤K2}\{k:\quad K_{1}\leq\sqrt{n}|b_{k}|\leq K_{2}\}. Similarly, for j=1,2j=1,2, define

where b1kb_{1k} and b2kb_{2k} denote, respectively, the real part and imaginary part of bkb_{k}.

Similar to the real case, the complex incompressible vector are also evenly spread, i.e. many coordinates are of the order n−1/2n^{-1/2}.

Let b∈Incomp(γ,ρ){\bf b}\in Incomp(\gamma,\rho). Then there is a set σ1(b)⊂{1,⋯ ,n}\sigma_{1}({\bf b})\subset\{1,\cdots,n\} of cardinality ∣σ1(b)∣≥cn|\sigma_{1}({\bf b})|\geq cn with c≥ρ2γ/4c\geq\rho^{2}\gamma/4 so that for j=1j=1 or 22,

By Lemma 3.4 in , for b∈Incomp(γ,ρ){\bf b}\in Incomp(\gamma,\rho), there is a set σ(b)\sigma({\bf b}) of cardinality ∣σ(b)∣≥12ρ2γn|\sigma({\bf b})|\geq\frac{1}{2}\rho^{2}\gamma n so that

Hence ∣b1k∣≤1/γn|b_{1k}|\leq 1/\sqrt{\gamma n} and ∣b2k∣≤1/γn|b_{2k}|\leq 1/\sqrt{\gamma n} if k∈σ(b)k\in\sigma({\bf b}). On the other hand, either b1kb_{1k} or b2kb_{2k} must be bigger than ρ(22n)−1\rho(2\sqrt{2n})^{-1}. The assertion follows. ∎

(1) Suppose that η1,⋯ ,ηn\eta_{1},\cdots,\eta_{n} are i.i.d. real r.v.’s, or imaginary r.v.’s, or complex ones with linearly correlated Re(ηk)Re(\eta_{k}) and Im(ηk),k=1,2,⋯ ,nIm(\eta_{k}),k=1,2,\cdots,n. If E∣ηk−Eηk∣2=1E|\eta_{k}-E\eta_{k}|^{2}=1 and E∣ηk∣3<BE|\eta_{k}|^{3}<B, for any ε≥0\varepsilon\geq 0, then

where C,c>0C,c>0 depend only on B,K1,K2B,K_{1},K_{2}.

(2) Let η1,⋯ ,ηn\eta_{1},\cdots,\eta_{n} be i.i.d. complex r.v.’s with E∣ηk−Eηk∣2=1E|\eta_{k}-E\eta_{k}|^{2}=1 and E∣ηk∣3<BE|\eta_{k}|^{3}<B, then (2.8) holds or

where C,c>0C,c>0 depend only on B, K1, K2B,\ K_{1},\ K_{2}, σ1, σ2\sigma_{1},\ \sigma_{2} and σ12\sigma_{12}.

(1). We only consider the case where the r.v.’s {ηk}\{\eta_{k}\} are real. The other two cases follow from the real case. Let bk=b1k+ib2kb_{k}=b_{1k}+ib_{2k} and v=v1+iv2v=v_{1}+iv_{2}. Noting that

Let bk=b1k+ib2kb_{k}=b_{1k}+ib_{2k}, ηk=η1k+iη2k\eta_{k}=\eta_{1k}+i\eta_{2k} and v=v1+iv2v=v_{1}+iv_{2}. It is observed that Theorem 3 implies Theorem 4 for big values of ε\varepsilon (constant order or even larger). Therefore we can suppose in what follows that

where l1l_{1} is a constant which will be specified later.

If the real part of η1\eta_{1} is linearly correlated to the imaginary part of η1\eta_{1}, then we have (2.8). Therefore we assume in the sequel that η11\eta_{11} is not linearly correlated to η21\eta_{21}.

Set ζk=1∣bk∣∣ξk−ξk′∣\zeta_{k}=\frac{1}{|b_{k}|}|\xi_{k}-\xi_{k}^{\prime}| where ξk=b1kη1k−b2kη2k\xi_{k}=b_{1k}\eta_{1k}-b_{2k}\eta_{2k} and ξk′\xi_{k}^{\prime} is an independent copy of ξk\xi_{k}. Then

where aa is some positive constant depending only on σ1, σ2\sigma_{1},\ \sigma_{2} and σ12\sigma_{12}.

On the other hand, Eζk3≤64BE\zeta_{k}^{3}\leq 64B. The Paley-Zygmund inequality () gives that

which is a positive constant depending only on BB, σ1\sigma_{1}, σ2\sigma_{2} and σ12\sigma_{12}. Following we introduce a new r.v. ζ^k\hat{\zeta}_{k} conditioned on ζk>a\zeta_{k}>\sqrt{a}, that is, for any measurable function gg

With the notation ϕk(t)=Eexp⁡(iξkt)\phi_{k}(t)=E\exp(i\xi_{k}t), it is observed that

by taking ε<(πa)/4=l1\varepsilon<(\pi\sqrt{a})/4=l_{1}. All the remaining arguments including the analysis for the level sets T(m,r)T(m,r) are similar to those of and so we here omit the details. Thus, one can conclude that for every ε≥0\varepsilon\geq 0

where 0<τ<n0<\tau<n, ∣b∣=(∣b1∣,⋯ ,∣bn∣)|{\bf b}|=(|b_{1}|,\cdots,|b_{n}|) and C,c>0C,c>0 are positive constants depending only on BB, σ1\sigma_{1}, σ2\sigma_{2} and σ12\sigma_{12}..

Finally, combining (2.14) and Lemma 2.1 in one can obtain the small ball probability for complex case (when applying (2.14) to the spread part of the vector b{\bf b} one can suppose that K1=1K_{1}=1 by re-scaling bkb_{k} and α\alpha). Thus we complete the proof. ∎

To treat the compressible vector, the following lemma is needed.

Suppose that η1,⋯ ,ηn\eta_{1},\cdots,\eta_{n} are i.i.d. centered complex r.v.’s with E∣ηk∣2=1E|\eta_{k}|^{2}=1 and E∣ηk∣3≤BE|\eta_{k}|^{3}\leq B. Let {ajk,j,k=1,⋯ ,n}\{a_{jk},j,k=1,\cdots,n\} be complex numbers. Then for 0<λ<10<\lambda<1 and any vector b=(b1,⋯ ,bn)∈Sn−1{\bf b}=(b_{1},\cdots,b_{n})\in S^{n-1} there is μ∈(0,1)\mu\in(0,1) such that the sum Snj=∑k=1nbk(ηk−ajk)S_{nj}=\sum\limits_{k=1}^{n}b_{k}(\eta_{k}-a_{jk}) satisfy

where μ\mu depends only on λ\lambda and BB.

On the other hand by Burkholder inequality we have

Hence Paley-Zygmund inequality gives that

where μ\mu depends only on λ\lambda and BB. ∎

2. Proof of Theorem 2

The whole argument is similar to that of and we only sketch the proof. For more details one can refer to .

Since Sn−1S^{n-1} can be decomposed as the union of CompComp and IncompIncomp, we then consider the smallest singular value on each set separately.

By Lemma 2 there are c1>0c_{1}>0 and v∈(0,1)v\in(0,1) depending on μ\mu only so that

Actually, the proof is similar to that of Proposition 3.4 in . The only difference is that we should use our Lemma 2 instead of Lemma 3.6 in . Therefore similar to Lemma 3.3 in , we have, there exist γ,ρ,c2,c3>0\gamma,\rho,c_{2},c_{3}>0 so that

Let X1,⋯ ,Xn{\bf X}_{1},\cdots,{\bf X}_{n} denote the column vectors of W{\bf W} and HkH_{k} the span of all columns except the kk-th column. One can check that Lemma 3.5 in is still true in complex case and hence

When all {Xjk}\{X_{jk}\} are real r.v.’s, or when Re(Xjk)Re(X_{jk}) and Im(Xjk)Im(X_{jk}) are linearly correlated or when Re(Xjk)=0Re(X_{jk})=0 we have

where UKU_{K} denotes the event that ∥W∥≤Kn1/2\|{\bf W}\|\leq Kn^{1/2}. One can check that Lemma 3.6 in applies to complex case and hence

where c4c_{4} is a constant depending only on B, KB,\ K, σ1, σ2\sigma_{1},\ \sigma_{2} and σ12\sigma_{12}. Further,

where V1kV_{1k} and V2kV_{2k} denote, respectively, the events that the real part and imaginary part of the vector Yk∈Incomp{\bf Y}_{k}\in Incomp satisfy (2.7) in Lemma 1, Y^1k\hat{{\bf Y}}_{1k} and Y^2k\hat{{\bf Y}}_{2k} denote, respectively, the spread part of the real part and imaginary part of the vector Yk{\bf Y}_{k}. By (2.8) in Theorem 4 and (2.3) we have

where c5,c6,c7c_{5},c_{6},c_{7} are positive constants depending only on BB, σ1, σ2\sigma_{1},\ \sigma_{2} and σ12\sigma_{12}.

Here the level set SD⊆Sn−1S_{D}\subseteq S^{n-1} is defined as

where α\alpha and D0D_{0} are some constants. For more details about α\alpha and D0D_{0}, see . Further, one can similarly prove that Lemma 5.8 in holds in our case and therefore we obtain

which, combined with the fact that the cardinal number ∣D∣|\mathcal{D}| is of order nn, then implies that

where c8>0c_{8}>0. Similarly, one may also show that

Picking up the above argument one can conclude that

where C>0C>0 and c∈(0,1)c\in(0,1) depend only on KK, BB, σ1, σ2\sigma_{1},\ \sigma_{2} and σ12\sigma_{12}.

For all the remaining case, i.e. Re(Xjk)Im(Xjk)≢0Re(X_{jk})Im(X_{jk})\not\equiv 0, and Re(Xjk)Re(X_{jk}), Im(Xjk)Im(X_{jk}) are not linearly correlated, one has

and one can similarly obtain (2.18) for complex case. Theorem 2 follows from (2.15)-(2.19) immediately.

The convergence of logarithmic potential and circular law

In this part the logarithmic potential will be used to show that the circular law is true. According to Lower Envelop Theorem and Unicity Theorem (see Theorem 6.9, p.73, and Corollary 2.2, p.98, in ), it suffices to show that the corresponding potential converges to the potential of the circular law.

To make use of Theorem 2 one needs to bound the maximum singular value of W{\bf W}. To this end, we would like to present an important fact which was proved in , that is, if (1) EXjk=0EX_{jk}=0, (2) ∣Xjk∣≤nεn,|X_{jk}|\leq\sqrt{n}\varepsilon_{n}, (3) E∣Xjk∣2≤1 and 1≥E∣Xjk∣2→1E|X_{jk}|^{2}\leq 1\ \text{and}\ 1\geq E|X_{jk}|^{2}\rightarrow 1 and (4) E∣Xjk∣l≤c(nεn)l−3 for l≥3E|X_{jk}|^{l}\leq c(\sqrt{n}\varepsilon_{n})^{l-3}\ \text{for}\ l\geq 3, where εn→0\varepsilon_{n}\rightarrow 0 with the convergence rate slower than any preassigned one as n→∞n\to\infty. Then for any K>4K>4

where ll is any positive number (proved for real case in , for complex case see Chapter 5 of ).

Let the random matrix X^=(X^jk)\hat{{\bf X}}=(\hat{X}_{jk}) with X^jk=XjkI(∣Xjk∣≤nεn)\hat{X}_{jk}=X_{jk}I(|X_{jk}|\leq\sqrt{n}\varepsilon_{n}). Then one can show that

see Lemma 2.2 of (the argument of the complex case is similar to that of the real one). Here the notation i.o.i.o. means infinitely often. Thus it is sufficient to consider the random matrix X^\hat{{\bf X}} in order to prove the conjecture.

Taking An=EX^−znI{\bf A}_{n}=E\hat{{\bf X}}-z\sqrt{n}{\bf I} in Theorem 2 one can obtain that

where EX^=(EX^kj)E\hat{{\bf X}}=(E\hat{X}_{kj}). Here one should note that from (3.3) re-scaling the underlying r.v.’s is trivial. Moreover

Therefore, applying (3.1) and choosing an appropriate KK in (3.3), we have

where both C>0C>0 and c∈(0,1)c\in(0,1) depend only on KK, E∣X11∣3E|X_{11}|^{3}, E\big{(}Re(X_{11})\big{)}^{2}, E\big{(}Im(X_{11})\big{)}^{2}, and ERe(X11)Im(X11)ERe(X_{11})Im(X_{11}).

In the sequel, to simplify the notation, we still use the notation X{\bf X} instead of X^\hat{{\bf X}} and μn(x,y)\mu_{n}(x,y) instead of the empirical spectral distribution corresponding to X^\hat{{\bf X}}. But one should keep in mind that {Xkj}\{X_{kj}\} are non-centered and ∣Xkj∣≤nεn|X_{kj}|\leq\sqrt{n}\varepsilon_{n}.

Before we prove the convergence of the logarithmic potential of μn(x,y)\mu_{n}(x,y), we will characterize the relation between the potential of the circular law μ(x,y)\mu(x,y) and the integral of logarithmic function with respect to v(x,z)v(x,z), the limiting distribution of vn(x,z)v_{n}(x,z) as below.

Let x+iy=reiθ,r>0x+iy=re^{i\theta},r>0. One can then verify that

On the other hand by Lemma 4.4 in one has

Therefore for any z=s+it,z1=s1+itz=s+it,z_{1}=s_{1}+it with ∣z1∣>1|z_{1}|>1, we have

Let s1→∞s_{1}\rightarrow\infty and then ∣z1∣→∞|z_{1}|\rightarrow\infty. Therefore, from Lemma 4.2 of the left and right end point, x1x_{1} and x2x_{2}, of the support of v(⋅,z1)v(\cdot,z_{1}) satisfy

as s1→∞s_{1}\rightarrow\infty. In addition,

We now proceed to prove the convergence of the potential of μn(x,y)\mu_{n}(x,y). The potential of μn(x,y)\mu_{n}(x,y) is

where I{\bf I} is the identity matrix. We will prove

as n→∞n\to\infty. Observe that by the fourth moment condition

where λmax⁡(Hn)\lambda_{\max}({\bf H}_{n}) denotes the maximum eigenvalue of Hn{\bf H}_{n}. It follows that for any δ>0\delta>0 and sufficiently large nn

Here we do not present the proof of the convergence of vn(x,z)v_{n}(x,z) to v(x,z)v(x,z) with the desired convergence rate for each zz. Indeed, the rank inequality (see Theorem 11.43 in ) can be used to re-centralize XjkX_{jk} and then Lemma 10.15 in provides the convergence rate under the assumption E∣X11∣2+δ<∞E|X_{11}|^{2+\delta}<\infty.

On the other hand, by (3.4) and Borel-Cantelli lemma,

Here we take ε=n−1−δ,δ>0\varepsilon=n^{-1-\delta},\delta>0 in (3.4). One should observe that ε\varepsilon in Theorem 5.1 in can be dependent on nn, so does ε\varepsilon in Theorem 2. Moreover, from Lemma 4.2 in one can conclude that

So for all large nn, almost surely μn\mu_{n} is compactly supported on the disk {z:∣z∣≤2+δ}\{z:|z|\leq 2+\delta\}. Here we have used the fact that all the eigenvalues of an n×nn\times n matrix are dominated by the largest singular value of the same matrix. Consequently Theorem 1 follows from Lemma 3 combined with Lower Envelop Theorem and Unicity Theorem for logarithmic potential of measures (see Theorem 6.9, p.73, and Corollary 2.2, p.98, in ).

Acknowledgments

The authors would like to thank Prof. Z. D. Bai for his helpful discussions when we read Chapter 10 of Bai and Silverstein’s book.

References