Stein's method and the multivariate CLT for traces of powers on the classical compact groups

Christian Döbler, Michael Stolz

Exchangeable pairs and multivariate normal approximation

Both [CM08] and [Mec09] contain an “infinitesimal version of Stein’s method of exchangeable pairs”, i.e., they provide error bounds for multivariate normal approximations in the case that for each t>0t>0 there is an exchangeable pair (W,Wt)(W,W_{t}) and that some further limiting properties hold as t→0t\to 0. It is such an infinitesimal version that will be applied in what follows.

The abstract multivariate normal approximation theorem provided by Meckes will be sufficient for the special orthogonal and symplectic cases of the present note, and the unitary case needs only a slight extension via basic linear algebra. So it is not really necessary to explain how these approximation theorems relate to the fundamentals of Stein’s method (see [CS05] for a readable account). But it is crucial to understand that there are fundamental differences between the univariate and multivariate cases. So a few remarks are in order. The starting point of any version of Stein’s method of normal approximation in the univariate case is the observation that a random variable WW is standard normal if, and only if, for all ff from a suitable class of piecewise C⁡1\operatorname{C}^{1} functions one has

holds (where HS stands for Hilbert-Schmidt, see below). Among the consequences of this is that the multivariate approximation theorems are phrased in Wasserstein distance (see below) rather than Kolmogorov’s distance

for probability measures μ,ν\mu,\nu on the real line, i.e., the distance concept in which Fulman’s univariate theorems are cast.

If Σ\Sigma is actually positive definite, then

In applications it is often easier to verify the following stronger condition in the place of (iii) of Proposition 1.1:

Now we turn to a version of Proposition 1.1 for complex random vectors, which will be needed for the case of the unitary group. ∥⋅∥op\|\cdot\|_{\rm op} and ∥⋅∥HS\|\cdot\|_{\rm HS} extending in the obvious way, we now denote by Z=(Z1,…,Zd)Z=(Z_{1},\ldots,Z_{d}) a dd-dimensional complex standard normal random vector, i.e., there are iid N⁡(0,1/2)\operatorname{N}(0,1/2) distributed real random variables X1,Y1,…,Xd,YdX_{1},Y_{1},\ldots,X_{d},Y_{d} such that Zj=Xj+iYjZ_{j}=X_{j}+iY_{j} for all j=1,…,dj=1,\ldots,d.

If Σ\Sigma is actually positive definite, then

Remark 1.2 above also applies to the present condition (iv).

and that for d=1d=1 one has to specify whether a scalar is to be interpreted as a vector or as a matrix. Writing JJ for the Kronecker product \left(\begin{array}[]{rr}1&0\\ 0&-1\end{array}\right)\otimes\operatorname{I}_{d}, one easily verifies the following identity of real 2d×2d2d\times 2d matrices:

Now we construct a family of exchangeable pairs to fit into Meckes’ set-up. We do this along the lines of Fulman’s univariate approach in [Ful10]. In a nutshell, let (Mt)t≥0(M_{t})_{t\geq 0} be Brownian motion on KnK_{n} with Haar measure as initial distribution. Set M:=M0M:=M_{0}. Brownian motion being reversible w.r.t. Haar measure, (M,Mt)(M,M_{t}) is an exchangeable pair for any t>0t>0. Suppose that the centered version WW of the statistic we are interested in is given by W=(f1(M),…,fd(M))TW=(f_{1}(M),\ldots,f_{d}(M))^{T} for suitable measurable functions f1,…,fdf_{1},\ldots,f_{d}. Defining Wt=(f1(Mt),…,fd(Mt))TW_{t}=(f_{1}(M_{t}),\ldots,f_{d}(M_{t}))^{T} clearly yields an exchangeable pair (W,Wt)(W,W_{t}).

To be more specific, let kn{\mathfrak{k}}_{n} be the Lie algebra of KnK_{n}, endowed with the scalar product ⟨X,Y⟩=Tr⁡(X∗Y).\langle X,Y\rangle=\operatorname{Tr}(X^{*}Y). Denote by Δ=ΔKn\Delta=\Delta_{K_{n}} the Laplace–Beltrami operator of KnK_{n}, i.e., the diffential operator corresponding to the Casimir element of the enveloping algebra of kn{\mathfrak{k}}_{n}. Then (Mt)t≥0(M_{t})_{t\geq 0} will be the diffusion on KnK_{n} with infinitesimal generator Δ\Delta and Haar measure as initial distribution. Reversibility then follows from general theory (see [Hel00, Sec. II.2.4], [IW89, Sec. V.4], for the relevant facts). Let (Tt)t≥0(T_{t})_{t\geq 0}, often symbolically written as (etΔ)t≥0\left(e^{t\Delta}\right)_{t\geq 0}, be the corresponding semigroup of transition operators on C⁡2(Kn)\operatorname{C}^{2}(K_{n}). From the Markov property of Brownian motion we obtain that

Note that (t,g)↦(Ttf)(g)(t,g)\mapsto(T_{t}f)(g) satisfies the heat equation, so a Taylor expansion in tt yields the following lemma, which is one of the cornerstones in Fulman’s approach in that it shows that in first order in tt a crucial quantity for the regression conditions (i) of Propositions 1.1 and 1.3 is given by the action of the Laplacian.

An elementary construction of the semigroup (Tt)(T_{t}) via an eigenfunction expansion in irreducible characters of KnK_{n} can be found in [Ste70]. This construction immediately implies Lemma 1.5 for power sums, see Sec. 2, since they are characters of (reducible) tensor power representations.

Power sums and Laplacians

For indeterminates X1,…,XnX_{1},\ldots,X_{n} and a finite family λ=(λ1,…,λr)\lambda=(\lambda_{1},\ldots,\lambda_{r}) of positive integers, the power sum symmetric polynomial with index λ\lambda is given by

Using this formula, we will extend the definition of pλp_{\lambda} to integer indices by p0(A)=Tr⁡(I)p_{0}(A)=\operatorname{Tr}(I), p−k(A)=Tr⁡((A−1)k)p_{-k}(A)=\operatorname{Tr}((A^{-1})^{k}) and pλ(A)=∏j=1rpj(A).p_{\lambda}(A)=\prod_{j=1}^{r}p_{j}(A).

In particular, the pλp_{\lambda} may be viewed as functions on KnK_{n}, and the action of the Laplacian ΔKn\Delta_{K_{n}} on them, useful in view of Lemma 1.5, is available from [Rai97, Lév08]. We specialize their formulae to the cases that will be needed in what follows:

For the Laplacian ΔU⁡n\Delta_{\operatorname{U}_{n}} on U⁡n\operatorname{U}_{n}, one has

For the Laplacian ΔSO⁡n\Delta_{\operatorname{SO}_{n}} on SO⁡n\operatorname{SO}_{n},

For the Laplacian ΔUSp⁡2n\Delta_{\operatorname{USp}_{2n}} on USp⁡2n\operatorname{USp}_{2n},

In what follows, we will need to integrate certain pλp_{\lambda} over the group KnK_{n}. Thus we will need special cases of the Diaconis-Shahshahani moment formulae that we now recall (see [DS94, HR03, PV04, Sto05] for proofs). Let a=(a1,…,ar)a=(a_{1},\ldots,a_{r}), b=(b1,…,bq)b=(b_{1},\ldots,b_{q}) be families of nonnegative integers and define

Here we have used the notation (2m−1)!!=(2m−1)(2m−3)⋅…⋅3⋅1.(2m-1)!!=(2m-1)(2m-3)\cdot\ldots\cdot 3\cdot 1. Further, we will write

Let M=MnM=M_{n} be a Haar-distributed element of U⁡n\operatorname{U}_{n}, Z1,…,ZrZ_{1},\ldots,Z_{r} iid complex standard normals. Then, if ka≠kbk_{a}\neq k_{b},

If M=MnM=M_{n} is a Haar-distributed element of SO⁡n\operatorname{SO}_{n}, n−1≥kan-1\geq k_{a}, Z1,…,ZrZ_{1},\ldots,Z_{r} iid real standard normals, then

If M=M2nM=M_{2n} is a Haar-distributed element of USp⁡2n\operatorname{USp}_{2n}, 2n≥ka2n\geq k_{a}, Z1,…,ZrZ_{1},\ldots,Z_{r} iid real standard normals, then

The unitary group

Let Z:=(Zd−r+1,…,Zd)TZ:=(Z_{d-r+1},\ldots,Z_{d})^{T} be an rr-dimensional standard complex normal random vector, i.e., there are iid real random variables Xd−r+1,…,Xd,X_{d-r+1},\ldots,X_{d}, Yd−r+1,…,YdY_{d-r+1},\ldots,Y_{d} with distribution N⁡(0,1/2)\operatorname{N}(0,1/2) such that Zj=Xj+iYjZ_{j}=X_{j}+iY_{j} for j=d−r+1,…,dj=d-r+1,\ldots,d. Furthermore, we take Σ\Sigma to denote the diagonal matrix diag⁡(d−r+1,d−r+2,…,d)\operatorname{diag}(d-r+1,d-r+2,\ldots,d) and write ZΣ:=Σ1/2Z.Z_{\Sigma}:=\Sigma^{1/2}Z. The present section is devoted to the proof of the following

If n≥2dn\geq 2d, the Wasserstein distance between WW and ZΣZ_{\Sigma} is

If 1≤r=⌊cd⌋1\leq r=\lfloor cd\rfloor for 0<c<10<c<1, then

The case r≡1r\equiv 1 means that one considers a single power Tr⁡(Md)\operatorname{Tr}(M^{d}). In his study [Ful10] of the univariate case, Fulman considers the random variable Tr⁡(Md)d\frac{\operatorname{Tr}(M^{d})}{\sqrt{d}} instead. By the scaling properties of the Wasserstein metric, the present result implies

So in this special case we recover the rate of convergence that was obtained by Fulman, albeit in Wasserstein rather than Kolmogorov distance.

In order to prove Theorem 3.1, we invoke the construction that was explained in Section 1 above, yielding a family (W,Wt)t>0(W,W_{t})_{t>0} of exchangeable pairs such that Wt=Wt(d,r,n)=(Tr⁡(Mtd−r+1),Tr⁡(Mtd−r+2),…,Tr⁡(Mtd))TW_{t}=W_{t}(d,r,n)=(\operatorname{Tr}(M_{t}^{d-r+1}),\operatorname{Tr}(M_{t}^{d-r+2}),\ldots,\operatorname{Tr}(M_{t}^{d}))^{T}. We have to check the conditions of Prop. 1.3. As to (i), from Lemma 1.5 and Lemma 2.1(i) we obtain

tends, as t→0t\to 0, a.s. and in L⁡1(∥⋅∥HS)\operatorname{L}^{1}(\|\cdot\|_{\rm HS}) to

where Λ=diag⁡(nj: j=d−r+1,…,d)\Lambda=\operatorname{diag}(nj:\ j=d-r+1,\ldots,d) and R=(Rd−r+1,…,Rd)TR=(R_{d-r+1},\ldots,R_{d})^{T} with Rj=−j∑l=1j−1pl,j−l(M).R_{j}=-j\sum_{l=1}^{j-1}p_{l,j-l}(M). (See Remark 1.7 above for the type of convergence.)

To verify conditions (ii) and (iii) of Prop. 1.3, we first prove the following lemma.

By well known properties of conditional expectation,

Applying Lemma 1.5 and 2.1(ii) to the first term yields

proving the first assertion. For the second, we compute analogously

establishing (i). Turning to (ii), one calculates that

By exchangeability, we have S1=S9S_{1}=S_{9}, S3=S2‾S_{3}=\overline{S_{2}}, S4=S5S_{4}=S_{5}, S7=S2S_{7}=S_{2}, S8=S7‾=S3S_{8}=\overline{S_{7}}=S_{3}.

Now, for n≥2dn\geq 2d, i.e., large enough for the moment formulae of Lemma 2.4 to apply for all j=d−r+1,…,dj=d-r+1,\ldots,d,

Hence, S3=S2=S7‾=S8=2j2−2tnj3+O⁡(t2).S_{3}=S_{2}=\overline{S_{7}}=S_{8}=2j^{2}-2tnj^{3}+\operatorname{O}(t^{2}). On the other hand,

With the conditions of Theorem 1.3 in place, we have

To bound the quantities on the right hand side, we first observe that ∥Λ−1∥op=1n(d−r+1)\|\Lambda^{-1}\|_{\rm op}=\frac{1}{n(d-r+1)} and ∥Σ−1/2∥op=1d−r+1\|\Sigma^{-1/2}\|_{\rm op}=\frac{1}{\sqrt{d-r+1}}. Now

As to ∥T∥HS\|T\|_{\rm HS}, for n≥2dn\geq 2d we have

Plugging these bounds into (7), we obtain

whereas in the case that d−r<rd-r<r one obtains

This yields the first claim, and the others follow easily.

The special orthogonal group

Let Z=(Zd−r+1,…,Zd)TZ=(Z_{d-r+1},\ldots,Z_{d})^{T} denote an rr-dimensional real standard normal random vector and write ZΣ:=Σ1/2ZZ_{\Sigma}:=\Sigma^{1/2}Z, where Σ:=diag⁡(d−r+1,d−r+2,…,d)\Sigma:=\operatorname{diag}(d-r+1,d-r+2,\dots,d). The objective of this section is to prove the following

If n≥4d+1n\geq 4d+1, the Wasserstein distance between WW and ZΣZ_{\Sigma} is of the same order as in the unitary case, namely

To prove Theorem 4.1, we invoke a family (W,Wt)t>0(W,W_{t})_{t>0} of exchangeable pairs such that Wt:=Wt(d,r,n):=(fd−r+1(Mt),fd−r+2(Mt),…,fd(Mt))TW_{t}:=W_{t}(d,r,n):=(f_{d-r+1}(M_{t}),f_{d-r+2}(M_{t}),\ldots,f_{d}(M_{t}))^{T} for all t>0t>0. We will apply Prop. 1.1, so the first step is to verify its conditions. For condition (i) we will need the following

First observe that always fj(Mt)−fj(M)=pj(Mt)−pj(M)f_{j}(M_{t})-f_{j}(M)=p_{j}(M_{t})-p_{j}(M), no matter what the parity of jj is. By Lemmas 1.5 and 2.2

From Lemma 4.2 (see also Remark 1.7) we conclude

where Λ=diag⁡((n−1)j2 , j=d−r+1,…,d)\Lambda=\operatorname{diag}\left(\frac{(n-1)j}{2}\,,\,j=d-r+1,\ldots,d\right) and R=(Rd−r+1,…,Rd)TR=(R_{d-r+1},\ldots,R_{d})^{T}. Thus, condition (i) of Prop. 1.1 is satisfied. In order to verify condition (ii) we will first prove the following

By well-known properties of conditional expectation

By Lemmas 1.5 and part (ii) of Lemma 2.2, we see that

Also, by Lemma 1.5 and part (i) of Lemma 2.2

we see that many terms cancel, and finally obtain

Observing as above that, regardless of the parity of jj, we have fj(Mt)−fj(M)=pj(Mt)−pj(M)f_{j}(M_{t})-f_{j}(M)=p_{j}(M_{t})-p_{j}(M), we can now easily compute

for all j,k=1,…,dj,k=1,\ldots,d. Noting that for j=kj=k the last expression is j2n−j2p2j(M)j^{2}n-j^{2}p_{2j}(M) and that 2ΛΣ=diag⁡((n−1)j2 , j=1,…,d)2\Lambda\Sigma=\operatorname{diag}((n-1)j^{2}\,,\,j=1,\ldots,d) we see that condition (ii) of Prop. 1.1 is satisfied with the matrix S=(Sj,k)j,k=1,…,dS=(S_{j,k})_{j,k=1,\ldots,d} given by

In order to show that condition (iii) of Prop. 1.1 holds, we will need the following facts:

where the last equality follows from exchangeability. By Lemma 1.5 and part (ii) of Lemma 2.2 for the case k=jk=j

Again by Lemma 1.5 and part (i) of Lemma 2.2,

Case 2: jj is even. Then, again by Lemma 2.5

where the last equality follows again by Lemma 2.5. If j/2j/2 is odd, then, clearly, k≠j/4k\not=j/4 is not really a restriction, and by Lemma 2.5

Now we are in a position to check condition (iii)′(iii)^{\prime} of Prop. 1.1. By Hölder’s inequality

and it suffices to show that for all j=d−r+1,…,dj=d-r+1,\ldots,d

But this follows from Lemma 4.4, since by the Cauchy-Schwarz inequality

The case that jj is even can be treated similarly as can be seen by observing that in this case

Proceeding as in the unitary case and treating the cases that r≤d−rr\leq d-r and d−r<rd-r<r separately, one obtains that the last expression is of order

The symplectic group

Let Z=(Zd−r+1,…,Zd)TZ=(Z_{d-r+1},\ldots,Z_{d})^{T} denote an rr-dimensional real standard normal random vector, Σ:=diag⁡(d−r+1,d−r+2,…,d)\Sigma:=\operatorname{diag}(d-r+1,d-r+2,\dots,d), and write ZΣ:=Σ1/2ZZ_{\Sigma}:=\Sigma^{1/2}Z.

The objective of this section is to prove the following

If n≥2dn\geq 2d, the Wasserstein distance between WW and ZΣZ_{\Sigma} is again of the same order as in the unitary case, namely, as given in (6).

For the proof, we use again exchangeable pairs (W,Wt)t>0(W,W_{t})_{t>0} such that

and apply Proposition 1.1. The verification of condition (i) involves the following analog of Lemma 4.2 above.

First observe that always fj(Mt)−fj(M)=pj(Mt)−pj(M)f_{j}(M_{t})-f_{j}(M)=p_{j}(M_{t})-p_{j}(M), no matter what the parity of jj is. By Lemmas 1.5 and 2.3,

where Λ=diag⁡((2n+1)j2 , j=d−r+1,…,d)\Lambda=\operatorname{diag}\left(\frac{(2n+1)j}{2}\,,\,j=d-r+1,\ldots,d\right) and R=(Rd−r+1,…,Rd)TR=(R_{d-r+1},\ldots,R_{d})^{T}. Thus, condition (i) of Prop. 1.1 is satisfied. The validity of condition (ii) follows from the next lemma, whose proof only differs from that of Lemma 4.3 in that it makes use of Lemma 2.3 in the place of Lemma 2.2.

Observing that for all j=d−r+1,…,dj=d-r+1,\ldots,d we have fj(Mt)−fj(M)=pj(Mt)−pj(M)f_{j}(M_{t})-f_{j}(M)=p_{j}(M_{t})-p_{j}(M) and using Lemma 5.3 we can now easily compute

for all d−r+1≤j,k≤dd-r+1\leq j,k\leq d. Noting that for j=kj=k the last expression is 2j2n−j2p2j(M)2j^{2}n-j^{2}p_{2j}(M) and that 2ΛΣ=diag⁡((2n+1)j2 , j=d−r+1,…,d)2\Lambda\Sigma=\operatorname{diag}((2n+1)j^{2}\,,\,j=d-r+1,\ldots,d), we see that condition (ii) of Prop. 1.1 is satisfied with the matrix S=(Sj,k)j,k=d−r+1,…,dS=(S_{j,k})_{j,k=d-r+1,\ldots,d} given by

The validity of condition (iii′)(iii^{\prime}) of Prop. 1.1 is based on the following lemma, which can be proven in the same way as its analog in Section 4.

If n≥2dn\geq 2d, for all j=1,…,dj=1,\ldots,d there holds

References