Large information plus noise random matrix models and consistent subspace estimation in large sensor networks

Walid Hachem, Philippe Loubaton, Xavier Mestre, Jamal Najim, Pascal Vallet

Introduction

The classical source localization problem consists in estimating vector θ=(θ1,…,θK)T\boldsymbol{\theta}=(\theta_{1},\ldots,\theta_{K})^{T} from NN samples collected in the M×NM\times N matrix YN=(y1,…,yN)\mathbf{Y}_{N}=({\bf y}_{1},\ldots,{\bf y}_{N}). This problem was extensively studied in the past (see e.g. and the references therein). The so-called subspace estimator of θ=(θ1,…,θK)T\boldsymbol{\theta}=(\theta_{1},\ldots,\theta_{K})^{T} is based on the observation that if matrices A(θ)\mathbf{A}(\boldsymbol{\theta}) and SN=(s1,…,sN)\mathbf{S}_{N}=({\bf s}_{1},\ldots,{\bf s}_{N}) have both full rank KK, then the angles (θk)k=1,…,K(\theta_{k})_{k=1,\ldots,K} are solutions The KK angles are the unique solutions under certain assumptions on function θ→a(θ)\theta\rightarrow\mathbf{a}(\theta) of the equation a(θ)∗ΠNa(θ)=0\mathbf{a}(\theta)^{*}\boldsymbol{\Pi}_{N}\mathbf{a}(\theta)=0, where ΠN\boldsymbol{\Pi}_{N} represents the orthogonal projection matrix on the kernel of matrix A(θ)SNSN∗A(θ)∗\mathbf{A}(\boldsymbol{\theta})\mathbf{S}_{N}\mathbf{S}_{N}^{*}\mathbf{A}(\boldsymbol{\theta})^{*}. The existing subspace methods consist in estimating for each θ\theta the quadratic form ηN(θ)=a(θ)∗ΠNa(θ)\eta_{N}(\theta)=\mathbf{a}(\theta)^{*}\boldsymbol{\Pi}_{N}\mathbf{a}(\theta) of ΠN\boldsymbol{\Pi}_{N} by a certain term η^N(θ)\hat{\eta}_{N}(\theta), and then to estimate the KK angles as the argument of the KK most significant local minima of function θ→η^N(θ)\theta\rightarrow\hat{\eta}_{N}(\theta). This approach has been extensively developed when N→+∞N\rightarrow+\infty and MM fixed. In this context, ηN(θ)\eta_{N}(\theta) can be estimated consistently for each θ\theta by η^N(θ)=a(θ)∗Π^Na(θ)\hat{\eta}_{N}(\theta)=\mathbf{a}(\theta)^{*}\hat{\boldsymbol{\Pi}}_{N}\mathbf{a}(\theta) with Π^N\hat{\boldsymbol{\Pi}}_{N} the orthogonal projection matrix on the eigenspace associated to the M−KM-K smallest eigenvalues of the empirical covariance matrix 1NYNYN∗\frac{1}{N}\mathbf{Y}_{N}\mathbf{Y}_{N}^{*}. It clearly holds that sup⁡θ∈[−π,π]∣η^N(θ)−ηN(θ)∣\sup_{\theta\in[-\pi,\pi]}\left|\hat{\eta}_{N}(\theta)-\eta_{N}(\theta)\right| converges torwards 0 almost surely, and this allows to prove that the corresponding estimators (θ^k)k=1,…,K(\hat{\theta}_{k})_{k=1,\ldots,K} of the direction of arrivals are consistent.

If however MM and NN are of the same order of magnitude, a quite common situation if the number of sensors MM is large, then the above estimators show poor performances because Π^N\hat{\boldsymbol{\Pi}}_{N} is no longer an accurate estimator of ΠN\boldsymbol{\Pi}_{N}. In order to study this context, Mestre & Lagunas were the first to propose consistent estimators of ηN(θ)\eta_{N}(\theta) when M,N→+∞M,N\rightarrow+\infty in such a way that cN=MN→cc_{N}=\frac{M}{N}\rightarrow c, with c>0c>0. In Mestre & Lagunas , it is assumed that the source signals (sk,n)k=1,…,K(s_{k,n})_{k=1,\ldots,K} are mutually independent complex Gaussian i.i.d. time series with unit variance elements. Under this assumption, yn\mathbf{y}_{n} can be written as

2 General notations and useful results

We now introduce various notations and results used throughout the paper.

The quantity CC will represent a generic positive constant whose main feature is to be deterministic and independent of MM and NN. The value of CC may change from one line to another.

We also notice that if m(z)m(z) is the Stieltjes transform of positive measure μ\mu, then it holds that

Background on the Information plus Noise model and on the estimator of [23]

Assumption A-1: 0<cN<1  \mboxand  0<c<10<c_{N}<1\;\mbox{and}\;0<c<1.

In this section, ΣN\boldsymbol{\Sigma}_{N} represents the complex valued M×NM\times N random matrix given by

where BN=A(θ)SNN\mathbf{B}_{N}=\frac{\mathbf{A}(\boldsymbol{\theta})\mathbf{S}_{N}}{\sqrt{N}} and WN=VNN\mathbf{W}_{N}=\frac{\mathbf{V}_{N}}{\sqrt{N}}. Matrices BN\mathbf{B}_{N} and WN\mathbf{W}_{N} are assumed to satisfy the following assumptions

Assumption A-2: Matrix BN\mathbf{B}_{N} is deterministic and satisfies sup⁡N∥BN∥<+∞\sup_{N}\|\mathbf{B}_{N}\|<+\infty

Assumption A-4: The entries of matrix WN\mathbf{W}_{N} are i.i.d and follow the complex normal distribution CN(0,σ2N)\mathcal{CN}(0,\frac{\sigma^{2}}{N}).

We assume moreover that the non zero eigenvalues of BNBN∗\mathbf{B}_{N}\mathbf{B}_{N}^{*} have multiplicities 1 in order to simplify the notations. In the following, we denote by 0=λ1,N=…=λM−K,N<λM−K+1,N<…<λM,N0=\lambda_{1,N}=\ldots=\lambda_{M-K,N}<\lambda_{M-K+1,N}<\ldots<\lambda_{M,N} and (uk,N)k=1,…,M({\bf u}_{k,N})_{k=1,\ldots,M} the ordered eigenvalues and associated eigenvectors of BNBN∗\mathbf{B}_{N}\mathbf{B}_{N}^{*}. The eigenvalues and the eigenvectors of matrix ΣNΣN∗\boldsymbol{\Sigma}_{N}\boldsymbol{\Sigma}_{N}^{*} are denoted (λ^k,N)k=1,…,M(\hat{\lambda}_{k,N})_{k=1,\ldots,M} and (u^k,N)k=1,…,M(\hat{{\bf u}}_{k,N})_{k=1,\ldots,M}, and μ^N\hat{\mu}_{N} represents the empirical eigenvalue distribution of ΣNΣN∗\boldsymbol{\Sigma}_{N}\boldsymbol{\Sigma}_{N}^{*} defined by

As we assume cN<1c_{N}<1, the joint probability distribution of (λ^k,N)k=1,…,M(\hat{\lambda}_{k,N})_{k=1,\ldots,M} is absolutely continuous (see e.g. James ) and it holds that the (λ^k,N)k=1,…,M(\hat{\lambda}_{k,N})_{k=1,\ldots,M} have multiplicity 1 almost surely. We finally denote by QN(z)\mathbf{Q}_{N}(z) the resolvent of matrix ΣNΣN∗\boldsymbol{\Sigma}_{N}\boldsymbol{\Sigma}_{N}^{*}, i.e. QN(z)=(ΣNΣN∗−zIM)−1\mathbf{Q}_{N}(z)=\left(\boldsymbol{\Sigma}_{N}\boldsymbol{\Sigma}_{N}^{*}-z\mathbf{I}_{M}\right)^{-1}.

It is well-known ([11, Th. 7.4], [9, Th. 1.1]) that it exists a sequence of deterministic probability measures (μN)(\mu_{N}) such that μ^N−μN→N0\hat{\mu}_{N}-\mu_{N}\to_{N}0 weakly almost surely. Measure μN\mu_{N} is characterized by its Stieltjes transform mN(z)m_{N}(z) which is known to satisfy the equation

with P1P_{1} and P2P_{2} two polynomials with positive coefficients, independent of NN. Then,

Taking into account the previous result, it is shown in that

If we denote by TN(z)\mathbf{T}_{N}(z) the matrix-valued function defined by

then TN{\bf T}_{N} coincides with the Stieltjes transform of a positive matrix valued measure μN\boldsymbol{\mu}_{N} with support SN\mathcal{S}_{N} such that μN(SN)=IM\boldsymbol{\mu}_{N}(\mathcal{S}_{N})=\mathbf{I}_{M} (see Hachem et al [13, Th. 2.4 & Prop. 2.2]), i.e.

In the remainder of the paper, we will make use of the following result proved in if WN{\bf W}_{N} is complex Gaussian and in Hachem et al in the non Gaussian case.

Consider two sequences of deterministic vectors (bN)(\mathbf{b}_{N}), (dN)(\mathbf{d}_{N}) such that sup⁡N∥bN∥<+∞\sup_{N}\|\mathbf{b}_{N}\|<+\infty and sup⁡N∥dN∥<+∞\sup_{N}\|\mathbf{d}_{N}\|<+\infty. Then, it holds that

In order to present the chacterization of SN\mathcal{S}_{N}, we first introduce the following notations. We denote by fN,ϕNf_{N},\phi_{N} and wNw_{N} the functions defined by

We are now in position to characterize SN\mathcal{S}_{N}.

The function ϕN\phi_{N} admits 2Q2Q non-negative local extrema counting multiplicities (with 1≤Q≤K+11\leq Q\leq K+1) whose preimages are denoted w1,N−<0<w1,N+≤w2,N−…≤wQ,N−<wQ,N+w_{1,N}^{-}<0<w_{1,N}^{+}\leq w_{2,N}^{-}\ldots\leq w_{Q,N}^{-}<w_{Q,N}^{+}. Define xq,N−=ϕN(wq,N−)x_{q,N}^{-}=\phi_{N}(w_{q,N}^{-}) and xq,N+=ϕN(wq,N+)x_{q,N}^{+}=\phi_{N}(w_{q,N}^{+}) for q=1…Qq=1\ldots Q. Then,

and the support SN\mathcal{S}_{N} of μN\mu_{N} is given by

Moreover, for q=1,…,Qq=1,\ldots,Q, each interval ]wq,N−,wq,N+[]w_{q,N}^{-},w_{q,N}^{+}[ contains at least an element of the set {0,λM−K+1,N,…,λM,N}\{0,\lambda_{M-K+1,N},\ldots,\lambda_{M,N}\} and each eigenvalue of BNBN∗\mathbf{B}_{N}\mathbf{B}_{N}^{*} belongs to one of these intervals.

The second statement of the theorem shows that each eigenvalue of BNBN∗\mathbf{B}_{N}\mathbf{B}_{N}^{*} corresponds to a certain interval of SN\mathcal{S}_{N}. More precisely, an eigenvalue of BNBN∗\mathbf{B}_{N}\mathbf{B}_{N}^{*} will be said to be associated to cluster [xq,N−,xq,N+][x_{q,N}^{-},x_{q,N}^{+}] if it belongs to the interval (wq,N−,wq,N+)(w_{q,N}^{-},w_{q,N}^{+}). We note that the eigenvalue is necessarily associated to the first cluster [x1,N−,x1,N+][x_{1,N}^{-},x_{1,N}^{+}].

wN(xq,N−)=wq,N−w_{N}(x_{q,N}^{-})=w_{q,N}^{-} and wN(xq,N+)=wq,N+w_{N}(x_{q,N}^{+})=w_{q,N}^{+} for each 1≤q≤Q1\leq q\leq Q,

3 Some useful evaluations

In this paragraph, we gather some useful bounds related to certain Stieltjes transforms. We first recall that the inequality

By using the identity, TN(z)−TN(z)∗=TN(z)(TN(z)−∗−TN(z)−1)TN(z)∗\mathbf{T}_{N}(z)-\mathbf{T}_{N}(z)^{*}=\mathbf{T}_{N}(z)\left(\mathbf{T}_{N}(z)^{-*}-\mathbf{T}_{N}(z)^{-1}\right)\mathbf{T}_{N}(z)^{*}, we get after some algebra

Indeed, TN(z)\mathbf{T}_{N}(z) can be written as TN(z)=(1+σ2cNmN(z))(BNBN∗−wN(z)IM)−1\mathbf{T}_{N}(z)=(1+\sigma^{2}c_{N}m_{N}(z))\left(\mathbf{B}_{N}\mathbf{B}_{N}^{*}-w_{N}(z)\mathbf{I}_{M}\right)^{-1}. Therefore, ∥TN(z)∥\|\mathbf{T}_{N}(z)\| is equal to

Since m^N(z)\hat{m}_{N}(z) is the Stieltjes transform of the distribution 1M∑k=1Mδλ^k,N\frac{1}{M}\sum_{k=1}^{M}\delta_{\hat{\lambda}_{k,N}}, it holds that

where 1\mathbf{1} denotes vector 1=(1,1,…,1)T\mathbf{1}=(1,1,\ldots,1)^{T}. We denote ω^1,N≤…≤ω^M,N\hat{\omega}_{1,N}\leq\ldots\leq\hat{\omega}_{M,N} its eigenvalues. Then we have the following straighforward properties.

The zeros of z↦1+σ2cNm^N(z)z\mapsto 1+\sigma^{2}c_{N}\hat{m}_{N}(z) are included in the set {ω^1,N,…,ω^M,N}\{\hat{\omega}_{1,N},\ldots,\hat{\omega}_{M,N}\}.

If the eigenvalues λ^1,N,…,λ^M,N\hat{\lambda}_{1,N},\ldots,\hat{\lambda}_{M,N} of ΣNΣN∗\boldsymbol{\Sigma}_{N}\boldsymbol{\Sigma}_{N}^{*} have multiplicity one, the equation 1+σ2cNm^N(z)=01+\sigma^{2}c_{N}\hat{m}_{N}(z)=0 has MM multiplicity one solutions which coincide with the (ω^k,N)k=1,…,M(\hat{\omega}_{k,N})_{k=1,\ldots,M}. Moreover, λ^1,N<ω^1,N<…<λ^M,N<ω^M,N\hat{\lambda}_{1,N}<\hat{\omega}_{1,N}<\ldots<\hat{\lambda}_{M,N}<\hat{\omega}_{M,N}.

If the eigenvalue λ^k,N\hat{\lambda}_{k,N} has multiplicity p>1p>1, i.e. λ^k−1,N<λ^k,N=λ^k+p−1,N<λ^k+p,N\hat{\lambda}_{k-1,N}<\hat{\lambda}_{k,N}=\hat{\lambda}_{k+p-1,N}<\hat{\lambda}_{k+p,N}, then,

and the ω^k,N\hat{\omega}_{k,N} that do not coincide with some eigenvalues of ΣNΣN∗\boldsymbol{\Sigma}_{N}\boldsymbol{\Sigma}_{N}^{*} are zeros of 1+σ2cNm^N(z)1+\sigma^{2}c_{N}\hat{m}_{N}(z).

Since cN<1c_{N}<1, we recall that the eigenvalues (λ^k,N)k=1,…,M(\hat{\lambda}_{k,N})_{k=1,\ldots,M} have multiplicity 1 almost surely. However, in subsection 3.2, it will be necessary to define properly the solutions of 1+σ2cNm^N(z)=01+\sigma^{2}c_{N}\hat{m}_{N}(z)=0 everywhere. This explains why the case where some of the (λ^k,N)k=1,…,M(\hat{\lambda}_{k,N})_{k=1,\ldots,M} are multiple has to be considered.

Function z↦−1z(1+σ2cNm^N(z))z\mapsto\frac{-1}{z(1+\sigma^{2}c_{N}\hat{m}_{N}(z))} is the Stieltjes transform of a probability measure whose support coincides with the set of all roots of the equation z(1+σ2cNm^N(z))=0z(1+\sigma^{2}c_{N}\hat{m}_{N}(z))=0, which is included into the set {0,ω^1,N,…,ω^M,N}\{0,\hat{\omega}_{1,N},\ldots,\hat{\omega}_{M,N}\}. Therefore, it holds that

We recall the two following useful results of and .

for each N>N0N>N_{0}. Then, with probability one, no eigenvalue of ΣNΣN∗\boldsymbol{\Sigma}_{N}\boldsymbol{\Sigma}_{N}^{*} belongs to [a,b][a,b] for NN large enough.

It is useful to mention that sup⁡NxQN,N+<+∞\sup_{N}x_{Q_{N},N}^{+}<+\infty and that these two theorems are still valid if b=+∞b=+\infty (see ).

The approach of is valid under the following assumptions.

Assumption A-5: For NN large enough, none of the strictly positive eigenvalues of BNBN∗\mathbf{B}_{N}\mathbf{B}_{N}^{*} is associated to the first cluster [x1,N−,x1,N+][x_{1,N}^{-},x_{1,N}^{+}], i.e. λM−K+1,N>wN(x1,N+)\lambda_{M-K+1,N}>w_{N}(x_{1,N}^{+}) for NN large enough.

Using theorems 3 and 4, we deduce that if t1−,t1+,t2−,t2+t_{1}^{-},t_{1}^{+},t_{2}^{-},t_{2}^{+} are real numbers independent of NN satisfying

then, almost surely, for NN large enough, it holds that

Assumptions 2.5 and 2.5 thus imply that, almost surely, the smallest M−KM-K eigenvalues of ΣNΣN∗\boldsymbol{\Sigma}_{N}\boldsymbol{\Sigma}_{N}^{*} are separated from the KK greatest ones for NN large enough in the sense that the 2 sets of eigenvalues are included into 2 disjoint intervals that do not depend on NN. It is interesting to remark that Assumptions 2.5 and 2.5 are "deterministic conditions" depending only on σ2,cN=MN\sigma^{2},c_{N}=\frac{M}{N},and on the eigenvalues of BNBN∗\mathbf{B}_{N}\mathbf{B}_{N}^{*}. If KK remains fixed, recent results of Benaych-Rao (see also ) imply that Assumptions A-5 and A-6 hold if and only if lim inf⁡N→+∞λM−K+1,N>σ2c\liminf_{N\rightarrow+\infty}\lambda_{M-K+1,N}>\sigma^{2}\sqrt{c}. If however KK scales with NN, the derivation of more explicit conditions equivalent to Assumptions 2.5 and 2.5 is still an open problem.

We are now in position to present the consistent estimator of ηN\eta_{N} proposed in . It is based on the observation that

where C{\cal C} represents a contour enclosing and not the strictly positive eigenvalues of BNBN∗\mathbf{B}_{N}\mathbf{B}_{N}^{*}, and the symbol C−{\cal C}^{-} means that the contour is oriented clockwise. The estimator of is based on the observation that under Assumptions 2.5 and 2.5, function wN(z)w_{N}(z) provides such a contour for NN large enough. In the following, for y>0y>0 and ϵ>0\epsilon>0, ϵ<y3\epsilon<\frac{y}{3} small enough, we consider the rectangle Ry\mathcal{R}_{y} defined by

and its boundary ∂Ry\partial\mathcal{R}_{y}. Then, the properties of function wN(z)w_{N}(z) (see Proposition 1) imply that for NN large enough, the set wN(∂Ry)w_{N}(\partial\mathcal{R}_{y}) is a contour enclosing the origin, but not the other eigenvalues of BNBN∗\mathbf{B}_{N}\mathbf{B}_{N}^{*}. Therefore, ΠN\boldsymbol{\Pi}_{N} can also be written as

because (BNBN∗−wN(z)IM)−1=TN(z)1+σ2cNmN(z)(\mathbf{B}_{N}\mathbf{B}_{N}^{*}-w_{N}(z)\mathbf{I}_{M})^{-1}=\frac{{\bf T}_{N}(z)}{1+\sigma^{2}c_{N}m_{N}(z)}. Using (5) and (8) as well as the following lemma

Almost surely, for NN large enough, the MM solutions (ω^k,N)k=1,…,M(\hat{\omega}_{k,N})_{k=1,\ldots,M} of the equation 1+σ2cNmN(z)=01+\sigma^{2}c_{N}m_{N}(z)=0 satisfy

almost surely for NN large enough. In pratice, the above estimator is quite easy to implement because, as the localization of the poles of the integrand in (22) w.r.t. the contour ∂Ry\partial\mathcal{R}_{y} is known (see lemma 2), the contour integral in (22) can be solved, and expressed in closed form in terms of the (u^k,N,λ^k,N,ω^k,N)k=1,…,M(\hat{{\bf u}}_{k,N},\hat{\lambda}_{k,N},\hat{\omega}_{k,N})_{k=1,\ldots,M}.

From now on, we assume that vector a(θ)\mathbf{a}(\theta) is given by (2) and that assumptions 2.5 and 2.5 hold. We consider t1−t_{1}^{-}, t1+t_{1}^{+}, t2−t_{2}^{-} and t2+t_{2}^{+} satisfying (18) as well a rectangle Ry\mathcal{R}_{y} defined by (20). We prove here the following result.

Assume assumptions A-1 to A-6 hold. Then, we have

In the following, we denote by Tϵ\mathcal{T}_{\epsilon} the set

We first establish in Sections 3.1 and 3.2 that the events E1,N\mathcal{E}_{1,N} and E2,N\mathcal{E}_{2,N} defined by

and ϕ(λ)∈(0,1)\phi(\lambda)\in(0,1) elsewhere, and define the random variable

The purpose of this section is to prove the following technical result.

and ψ0(λ)∈(0,1)\psi_{0}(\lambda)\in(0,1) elsewhere. From this definition, we clearly have

we finally obtain that (29) holds for l=1l=1.

The first term of the r.h.s of (31) can be upperbounded as follows

Using Hölder’s inequality, we get immediately that

Plugging the previous estimates into (32), we get

In this section, we will prove the following result.

We follow the same approach than in Section 3.1 and first prove that the (ω^k,N)k=1,…,M(\hat{\omega}_{k,N})_{k=1,\ldots,M} satisfy a property similar to (7). For this, we study the behaviour of the Stieltjes transform n^N(z)\hat{n}_{N}(z) of the distribution 1M∑k=1Mδω^k,N\frac{1}{M}\sum_{k=1}^{M}\delta_{\hat{\omega}_{k,N}} defined by

and use Lemma 1 as well as the inverse Stieltjes transform formula (3). Our starting point is the following result showing that the empirical eigenvalue distribution of Ω^N\hat{\boldsymbol{\Omega}}_{N} is very similar to the distribution of the eigenvalues of ΣNΣN∗\boldsymbol{\Sigma}_{N}\boldsymbol{\Sigma}_{N}^{*}. The following auxiliary result will be useful.

Assume assumptions A-1 to A-6 hold. It holds that

Proof: The proof is given in Appendix 5.1. □\square

We now prove the fundamental following result.

Proof: Using that Ω^N\hat{\boldsymbol{\Omega}}_{N} is a rank 1 perturbation of Λ^N\hat{\boldsymbol{\Lambda}}_{N}, we obtain immediately that

In order to study the expectation of this expression, we use (11) and (15). Moreover, (6) and a straightforward application of the Poincaré inequality to m^N(z)\hat{m}_{N}(z) considered for zz fixed as a function of the entries of WN\mathbf{W}_{N} leads immediately to

This immediately implies (35). Now define the function hN(z)h_{N}(z) by

When y→0y\rightarrow 0, this converges towards log⁡(1+σ2cNmN(xq,N+))−log⁡(1+σ2cNmN(xq,N−))\log(1+\sigma^{2}c_{N}m_{N}(x_{q,N}^{+}))-\log(1+\sigma^{2}c_{N}m_{N}(x_{q,N}^{-})), a real quantity because xq,N−x_{q,N}^{-} and xq,N+x_{q,N}^{+} belong to ∂SN\partial\mathcal{S}_{N}. This shows that κN([xq,N−,xq,N+])=0\kappa_{N}([x_{q,N}^{-},x_{q,N}^{+}])=0. Consequently,

where Π^l,N\hat{\boldsymbol{\Pi}}_{l,N} represents the orthogonal projection matrix on the 1–dimensionnal eigenspace associated to the eigenvalue λ^l,N\hat{\lambda}_{l,N} of ΣNΣN∗\boldsymbol{\Sigma}_{N}\boldsymbol{\Sigma}_{N}^{*}.

a property also established in the appendix. We now prove the following result.

Assume assumptions A-1 to A-6 hold. It holds that

Indeed, if (v^k,N)k=1,…,M(\hat{{\bf v}}_{k,N})_{k=1,\ldots,M} represent the eigenvectors of Ω^\hat{\boldsymbol{\Omega}}, then

As ∑k=1M∣elTv^k,N∣2=1\sum_{k=1}^{M}|{\bf e}_{l}^{T}\hat{{\bf v}}_{k,N}|^{2}=1, Jensen’s inequality yields to (41). Therefore, it holds that

As sup⁡N∥BNBN∗∥<+∞\sup_{N}\|\mathbf{B}_{N}\mathbf{B}_{N}^{*}\|<+\infty, we get using lemma 7 that

We remark that ∥ΣNΣN∗∥<t2++ϵ\|\boldsymbol{\Sigma}_{N}\boldsymbol{\Sigma}_{N}^{*}\|<t_{2}^{+}+\epsilon on the set E1,Nc\mathcal{E}_{1,N}^{c},and write the righthandside of (42) as

for each integer pp. This completes the proof of lemma 9. □\square

Assume that (38) holds until integer l−1l-1. We write as previously that

The Cauchy-Schwarz inequality leads immediately to

As for the second term of the r.h.s. of (43), we use Poincaré inequality and Hölder’s inequality to obtain

3 End of the proof of theorem 5

We now complete the proof of Theorem 5 when function θ→a(θ)\theta\rightarrow\mathbf{a}(\theta) is given by

for θ∈[−π,π]\theta\in[-\pi,\pi]. We recall that EN\mathcal{E}_{N} is defined by

Assume assumptions A-1 to A-6 hold. For each NN, it holds that

and remark that for each θ∈[−π,π]\theta\in[-\pi,\pi] and for each NN, there exists θN∈ϑN\theta_{N}\in\vartheta_{N} such that ∣θ−θN∣≤2πN2|\theta-\theta_{N}|\leq\frac{2\pi}{N^{2}}. For each θ∈[−π,π]\theta\in[-\pi,\pi], it holds that

It is easy to check that the third term of the r.h.s. of (45) satisfies

In order to evaluate the behaviour of the supremum over θ\theta of the first term of the r.h.s. of (45), we prove that for each α>0\alpha>0,

for some constant term CC. Inequality (46) thus implies that

for each integer ll. Borel-Cantelli’s lemma eventually implies that

We finally study the supremum of the second term of (45). We denote by νk,N\nu_{k,N} the elements of ϑN\vartheta_{N}. Let α>0\alpha>0, then

In order to complete the proof of Theorem 5, we establish the following proposition.

where the constant CC does not depend on the sequence (aN)(\mathbf{a}_{N}).

Proof: In order to shorten the notations, we denote by g^N(z)\hat{g}_{N}(z) and gN(z)g_{N}(z) the functions defined by

If we denote by A1,N\mathcal{A}_{1,N} and A2,N\mathcal{A}_{2,N} the events defined by

then [D1]i,j,N=0[\mathbf{D}_{1}]_{i,j,N}=0 on A1,Nc\mathcal{A}_{1,N}^{c} and [D2]i,j,N=0[\mathbf{D}_{2}]_{i,j,N}=0 on A2,Nc\mathcal{A}_{2,N}^{c}.

We now establish (49) by induction on ll, and first consider the case l=1l=1. We write the second moment of (g^N(z)−gN(z))χN2(\hat{g}_{N}(z)-g_{N}(z))\chi_{N}^{2} as

As χN≠0\chi_{N}\neq 0 implies that λ^M,N=∥ΣNΣN∗∥≤t2++2ϵ\hat{\lambda}_{M,N}=\|\boldsymbol{\Sigma}_{N}\boldsymbol{\Sigma}_{N}^{*}\|\leq t_{2}^{+}+2\epsilon, Lemma 10 implies that

The same conclusions hold when the derivatives w.r.t. variables W‾i,j,N\overline{W}_{i,j,N} are considered. This shows that the first term of the r.h.s. of (52) is a O(1N)\mathcal{O}\left(\frac{1}{N}\right) term. We now evaluate the behaviour of the second term of the r.h.s. of (52), and establish that

for each integer pp. We express ∂χN2∂Wi,j,N\frac{\partial\chi_{N}^{2}}{\partial W_{i,j,N}} as 2χN∂χN∂Wi,j,N2\chi_{N}\frac{\partial\chi_{N}}{\partial W_{i,j,N}}. Lemma 10 implies that sup⁡z∈∂RyχN2∣g^N(z)−gN(z)∣2<C\sup_{z\in\partial\mathcal{R}_{y}}\chi_{N}^{2}|\hat{g}_{N}(z)-g_{N}(z)|^{2}<C. Therefore, it is sufficient to check that

for each integer pp. ∂χN∂Wi,j,N\frac{\partial\chi_{N}}{\partial W_{i,j,N}} can be written as

for each integer pp. Using similar calculations and Proposition 3, we obtain that

for each integer pp. This completes the proof of (55) and establishes that

Assume assumptions A-1 to A-6 hold. It holds that

We express (g^N(z)−gN(z))χN2(\hat{g}_{N}(z)-g_{N}(z))\chi_{N}^{2} as β1,N(z)+β2,N(z)\beta_{1,N}(z)+\beta_{2,N}(z) where

Using Lemma 10, (59) for β1,N\beta_{1,N} will be established if we show that

This completes the proof of (59) for β1,N\beta_{1,N}. In order to show (59) for β2,N\beta_{2,N}, we first remark that by Lemma 10, ∣aN∗TN(z)aN∣|\mathbf{a}_{N}^{*}\mathbf{T}_{N}(z)\mathbf{a}_{N}| is uniformly bounded on ∂Ry\partial\mathcal{R}_{y}, and write that

The Poincaré inequality and Lemma 12 imply that

for some deterministic constant CC (see Lemma 10). This completes the proof of (49) for l=1l=1.

We now assume that (49) holds until integer l−1l-1 and write that

The Cauchy-Schwarz inequality implies that

Consistency of the angular estimates

For k=1,…,Kk=1,\ldots,K, with probability one,

In order to establish the proposition, we follow a classical approach initiated by Hannan to study sinusoid frequency estimates. For this, we first recall the following useful lemma.

Appendix

We first give the following useful technical result. Its proof, based on Poincaré’s inequality, is elementary and therefore omitted.

Moreover, the same results still hold when QN(z)\mathbf{Q}_{N}(z) is replaced by QN(z)2\mathbf{Q}_{N}(z)^{2}.

We are now in position to establish Lemma 4. We have to establish that

We first notice that (62) is equivalent to

In order to prove (63), we first show that

where ΔN(z)\boldsymbol{\Delta}_{N}(z) is given by ΔN(z)=Δ1,N(z)+Δ2,N(z)+Δ3,N(z)\boldsymbol{\Delta}_{N}(z)=\boldsymbol{\Delta}_{1,N}(z)+\boldsymbol{\Delta}_{2,N}(z)+\boldsymbol{\Delta}_{3,N}(z) with

In order to complete the proof of the lemma, we establish that

Proof: We first observe that (66) and (67) imply that

We first need to establish the following useful Lemma.

Therefore, it holds that max⁡(∣f(xn)−f(yn)∣,∣f(zn)−f(x)∣)<dn2\max(|f(x_{n})-f(y_{n})|,|f(z_{n})-f(x)|)<d_{n}^{2}. Writing f(xn)−f(x)=f(xn)−f(yn)+f(yn)−f(zn)+f(zn)−f(x)f(x_{n})-f(x)=f(x_{n})-f(y_{n})+f(y_{n})-f(z_{n})+f(z_{n})-f(x), we obtain that f(xn)−f(x)=f(yn)−f(zn)+o(dn)f(x_{n})-f(x)=f(y_{n})-f(z_{n})+o(d_{n}). By differentiability of ff on O{\mathcal{O}} and continuity of f′f^{\prime} at xx,

Using Lemma 4.6 in Haagerup-Thorbjornsen , we obtain

Proof: We start by observing that for any integers m1,m2,…,mtm_{1},m_{2},\ldots,m_{t}, matrix A=Λ^Nm111TΛ^Nm2⋯11TΛ^Nmt\mathbf{A}=\hat{\boldsymbol{\Lambda}}_{N}^{m_{1}}{\bf 11}^{T}\hat{\boldsymbol{\Lambda}}_{N}^{m_{2}}\cdots{\bf 11}^{T}\hat{\boldsymbol{\Lambda}}_{N}^{m_{t}} writes

Using Theorem 6 and for p≥2p\geq 2 the inequality,

4 Proof of Lemma 11: differentiability of the regularization factor

and it remains to apply the composition formula for differentials to obtain (50).

hence the derivative (50) is zero on A1,Nc\mathcal{A}_{1,N}^{c}.

if λ^k,N=λ^l,N\hat{\lambda}_{k,N}=\hat{\lambda}_{l,N}. Indeed, given ε>0\varepsilon>0, let ϕε(x)=ϕ(x)+ε\phi_{\varepsilon}(x)=\phi(x)+\varepsilon. Since ϕε(Ω^N)>0\phi_{\varepsilon}(\hat{\boldsymbol{\Omega}}_{N})>0,

5 Proof of lemma 12: various estimates

Thus, (80) follows immediately from (13). □\square

Lemma 18 immediately implies that the following uniform bounds hold.

Proof: We first recall that inequality (10) holds. Therefore, the uniform convergence result (79) implies that

for NN large enough. This establishes (82) that holds for NN large enough. In order to prove (83), we express Rr,N(z)\mathbf{R}_{r,N}(z) as

and use (78) and (80). The proof of (84) is similar, and is based on the identity

and that ∣wr,N(z)∣|w_{r,N}(z)| and ∣wN(z)∣|w_{N}(z)| are uniformly bounded from below by (13) and (80) (recall that is one of the eigenvalues of BNBN∗\mathbf{B}_{N}\mathbf{B}_{N}^{*}). □\square

with sup⁡z∈∂Ry∣ϵi,N(z)∣≤C<∞\sup_{z\in\partial\mathcal{R}_{y}}|\epsilon_{i,N}(z)|\leq C<\infty.

As for the use of the Poincaré inequality, we have:

Moreover, the same kind of uniform bounds still hold when QN(z)\mathbf{Q}_{N}(z) is replaced by QN(z)2\mathbf{Q}_{N}(z)^{2}.

The proofs of these results are based on elementary arguments, and are thus omitted. Following the calculations of and , we obtain that

After some calculations using Lemmas 19, 20, 21, we eventually obtain that

for all large NN. In order to prove (56) and (57), it remains to handle the terms involving the difference Rr,N(z)−TN(z)\mathbf{R}_{r,N}(z)-\mathbf{T}_{N}(z). We show in the following that

for all large NN. We start as usual with the identity Rr,N(z)−TN(z)=Rr,N(z)(TN(z)−1−Rr,N(z)−1)TN(z)\mathbf{R}_{r,N}(z)-\mathbf{T}_{N}(z)=\mathbf{R}_{r,N}(z)\left(\mathbf{T}_{N}(z)^{-1}-\mathbf{R}_{r,N}(z)^{-1}\right)\mathbf{T}_{N}(z), to get

There exists a constant C>0C>0 independent of NN such that

Proof: It is shown in and that Δ‾N(z)\underline{\Delta}_{N}(z) is the determinant of the following 2×22\times 2 linear system

We deduce from this that inf⁡z∈∂Ry∣ΔN(z)∣≥C>0\inf_{z\in\partial\mathcal{R}_{y}}\left|\Delta_{N}(z)\right|\geq C>0 for all large NN. Therefore, we can invert the system (95) and obtain

for all large NN. This establishes (57) and completes the proof of (56).

References