Stein's method and the multivariate CLT for traces of powers on the classical compact groups
Christian Döbler, Michael Stolz
Exchangeable pairs and multivariate normal approximation
Both [CM08] and [Mec09] contain an “infinitesimal version of Stein’s method of exchangeable pairs”, i.e., they provide error bounds for multivariate normal approximations in the case that for each there is an exchangeable pair and that some further limiting properties hold as . It is such an infinitesimal version that will be applied in what follows.
The abstract multivariate normal approximation theorem provided by Meckes will be sufficient for the special orthogonal and symplectic cases of the present note, and the unitary case needs only a slight extension via basic linear algebra. So it is not really necessary to explain how these approximation theorems relate to the fundamentals of Stein’s method (see [CS05] for a readable account). But it is crucial to understand that there are fundamental differences between the univariate and multivariate cases. So a few remarks are in order. The starting point of any version of Stein’s method of normal approximation in the univariate case is the observation that a random variable is standard normal if, and only if, for all from a suitable class of piecewise functions one has
holds (where HS stands for Hilbert-Schmidt, see below). Among the consequences of this is that the multivariate approximation theorems are phrased in Wasserstein distance (see below) rather than Kolmogorov’s distance
for probability measures on the real line, i.e., the distance concept in which Fulman’s univariate theorems are cast.
If is actually positive definite, then
In applications it is often easier to verify the following stronger condition in the place of (iii) of Proposition 1.1:
Now we turn to a version of Proposition 1.1 for complex random vectors, which will be needed for the case of the unitary group. and extending in the obvious way, we now denote by a -dimensional complex standard normal random vector, i.e., there are iid distributed real random variables such that for all .
If is actually positive definite, then
Remark 1.2 above also applies to the present condition (iv).
and that for one has to specify whether a scalar is to be interpreted as a vector or as a matrix. Writing for the Kronecker product \left(\begin{array}[]{rr}1&0\\ 0&-1\end{array}\right)\otimes\operatorname{I}_{d}, one easily verifies the following identity of real matrices:
Now we construct a family of exchangeable pairs to fit into Meckes’ set-up. We do this along the lines of Fulman’s univariate approach in [Ful10]. In a nutshell, let be Brownian motion on with Haar measure as initial distribution. Set . Brownian motion being reversible w.r.t. Haar measure, is an exchangeable pair for any . Suppose that the centered version of the statistic we are interested in is given by for suitable measurable functions . Defining clearly yields an exchangeable pair .
To be more specific, let be the Lie algebra of , endowed with the scalar product Denote by the Laplace–Beltrami operator of , i.e., the diffential operator corresponding to the Casimir element of the enveloping algebra of . Then will be the diffusion on with infinitesimal generator and Haar measure as initial distribution. Reversibility then follows from general theory (see [Hel00, Sec. II.2.4], [IW89, Sec. V.4], for the relevant facts). Let , often symbolically written as , be the corresponding semigroup of transition operators on . From the Markov property of Brownian motion we obtain that
Note that satisfies the heat equation, so a Taylor expansion in yields the following lemma, which is one of the cornerstones in Fulman’s approach in that it shows that in first order in a crucial quantity for the regression conditions (i) of Propositions 1.1 and 1.3 is given by the action of the Laplacian.
An elementary construction of the semigroup via an eigenfunction expansion in irreducible characters of can be found in [Ste70]. This construction immediately implies Lemma 1.5 for power sums, see Sec. 2, since they are characters of (reducible) tensor power representations.
Power sums and Laplacians
For indeterminates and a finite family of positive integers, the power sum symmetric polynomial with index is given by
Using this formula, we will extend the definition of to integer indices by , and
In particular, the may be viewed as functions on , and the action of the Laplacian on them, useful in view of Lemma 1.5, is available from [Rai97, Lév08]. We specialize their formulae to the cases that will be needed in what follows:
For the Laplacian on , one has
For the Laplacian on ,
For the Laplacian on ,
In what follows, we will need to integrate certain over the group . Thus we will need special cases of the Diaconis-Shahshahani moment formulae that we now recall (see [DS94, HR03, PV04, Sto05] for proofs). Let , be families of nonnegative integers and define
Here we have used the notation Further, we will write
Let be a Haar-distributed element of , iid complex standard normals. Then, if ,
If is a Haar-distributed element of , , iid real standard normals, then
If is a Haar-distributed element of , , iid real standard normals, then
The unitary group
Let be an -dimensional standard complex normal random vector, i.e., there are iid real random variables with distribution such that for . Furthermore, we take to denote the diagonal matrix and write The present section is devoted to the proof of the following
If , the Wasserstein distance between and is
If for , then
The case means that one considers a single power . In his study [Ful10] of the univariate case, Fulman considers the random variable instead. By the scaling properties of the Wasserstein metric, the present result implies
So in this special case we recover the rate of convergence that was obtained by Fulman, albeit in Wasserstein rather than Kolmogorov distance.
In order to prove Theorem 3.1, we invoke the construction that was explained in Section 1 above, yielding a family of exchangeable pairs such that . We have to check the conditions of Prop. 1.3. As to (i), from Lemma 1.5 and Lemma 2.1(i) we obtain
tends, as , a.s. and in to
where and with (See Remark 1.7 above for the type of convergence.)
To verify conditions (ii) and (iii) of Prop. 1.3, we first prove the following lemma.
By well known properties of conditional expectation,
Applying Lemma 1.5 and 2.1(ii) to the first term yields
proving the first assertion. For the second, we compute analogously
establishing (i). Turning to (ii), one calculates that
By exchangeability, we have , , , , .
Now, for , i.e., large enough for the moment formulae of Lemma 2.4 to apply for all ,
Hence, On the other hand,
With the conditions of Theorem 1.3 in place, we have
To bound the quantities on the right hand side, we first observe that and . Now
As to , for we have
Plugging these bounds into (7), we obtain
whereas in the case that one obtains
This yields the first claim, and the others follow easily.
The special orthogonal group
Let denote an -dimensional real standard normal random vector and write , where . The objective of this section is to prove the following
If , the Wasserstein distance between and is of the same order as in the unitary case, namely
To prove Theorem 4.1, we invoke a family of exchangeable pairs such that for all . We will apply Prop. 1.1, so the first step is to verify its conditions. For condition (i) we will need the following
First observe that always , no matter what the parity of is. By Lemmas 1.5 and 2.2
From Lemma 4.2 (see also Remark 1.7) we conclude
where and . Thus, condition (i) of Prop. 1.1 is satisfied. In order to verify condition (ii) we will first prove the following
By well-known properties of conditional expectation
By Lemmas 1.5 and part (ii) of Lemma 2.2, we see that
Also, by Lemma 1.5 and part (i) of Lemma 2.2
we see that many terms cancel, and finally obtain
Observing as above that, regardless of the parity of , we have , we can now easily compute
for all . Noting that for the last expression is and that we see that condition (ii) of Prop. 1.1 is satisfied with the matrix given by
In order to show that condition (iii) of Prop. 1.1 holds, we will need the following facts:
where the last equality follows from exchangeability. By Lemma 1.5 and part (ii) of Lemma 2.2 for the case
Again by Lemma 1.5 and part (i) of Lemma 2.2,
Case 2: is even. Then, again by Lemma 2.5
where the last equality follows again by Lemma 2.5. If is odd, then, clearly, is not really a restriction, and by Lemma 2.5
Now we are in a position to check condition of Prop. 1.1. By Hölder’s inequality
and it suffices to show that for all
But this follows from Lemma 4.4, since by the Cauchy-Schwarz inequality
The case that is even can be treated similarly as can be seen by observing that in this case
Proceeding as in the unitary case and treating the cases that and separately, one obtains that the last expression is of order
The symplectic group
Let denote an -dimensional real standard normal random vector, , and write .
The objective of this section is to prove the following
If , the Wasserstein distance between and is again of the same order as in the unitary case, namely, as given in (6).
For the proof, we use again exchangeable pairs such that
and apply Proposition 1.1. The verification of condition (i) involves the following analog of Lemma 4.2 above.
First observe that always , no matter what the parity of is. By Lemmas 1.5 and 2.3,
where and . Thus, condition (i) of Prop. 1.1 is satisfied. The validity of condition (ii) follows from the next lemma, whose proof only differs from that of Lemma 4.3 in that it makes use of Lemma 2.3 in the place of Lemma 2.2.
Observing that for all we have and using Lemma 5.3 we can now easily compute
for all . Noting that for the last expression is and that , we see that condition (ii) of Prop. 1.1 is satisfied with the matrix given by
The validity of condition of Prop. 1.1 is based on the following lemma, which can be proven in the same way as its analog in Section 4.
If , for all there holds