Bounds on the constant in the mean central limit theorem
Larry Goldstein
Introduction
The classical central limit theorem allows the approximation of the distribution of sums of “comparable” independent real-valued random variables by the normal. As this theorem is an asymptotic, it provides no information as to whether the resulting approximation is useful. For that purpose, one may turn to the Berry–Esseen theorem, the most classical version giving supremum norm bounds between the distribution function of the normalized sum and that of the standard normal. Various authors have also considered Berry–Esseen-type bounds using other metrics and, in particular, bounds in . The case , where the value
is used to measure the distance between distribution functions and , is of some particular interest and results using this metric are known as mean central limit theorems (see, e.g., dedecker , Erickson , HoChen and Ibragimov ; the latter three of these works consider nonindependent summand variables). One motivation for studying -bounds is that, combined with one of type , bounds on -distance for all may be obtained by the inequality
For , let be the collection of distributions with mean zero, variance and finite absolute third moment. We prove the following Berry–Esseen-type result for the mean central limit theorem.
In particular, when are identically distributed with distribution ,
For the case where all variables are identically distributed as having distribution , letting
the second part of Theorem 1.1 yields the upper bound . Regarding lower bounds, we also prove
Clearly, the elements of the sequence are nonnegative and decreasing in , and so have a limit, say . Regarding limiting behavior, Esseen esseen showed that
for an explicit constant depending only on . Zolotarev Zolotarev provides the representation
where and is the span of the distribution in the case where is lattice, and is zero otherwise. Zolotarev obtains
showing that , hence giving the asymptotic -Berry–Esseen constant value.
for all absolutely continuous functions for which these expectations exist. The zero bias transformation, mapping the distribution of to that of , was motivated by the Stein characterization of the normal distribution Stein81 , which states that is normal with mean zero and variance if and only if
for all absolutely continuous functions for which these expectations exist. Hence, the mean zero normal with variance is the unique fixed point of the zero bias transformation. How closeness to normality may be measured by the closeness of a distribution to its zero bias transform, and related applications, are the topics of res , zsm and L1bounds .
As shown in L1bounds and Goldstein-Reinert , for a random variable with and , the distribution of is absolutely continuous with density and distribution functions given, respectively, by
Theorem 1.1 results by showing that the functional
is bounded by 1 for all with distribution . As in (3), one may write out a more “explicit” form for using (6) and expressions for the moments on which depends, but such expressions appear to be of little value for the purposes of proving Theorem 1.1. In turn, the proof here employs convexity properties of which depend on the behavior of the zero bias transformation on mixtures. We also note that the functional has a different character than ; for instance, is zero for all nonlattice distributions with vanishing third moment, whereas is zero only for mean zero normal distributions.
by replacing by and by in equation (16) of Theorem 2.1 of L1bounds , we obtain the following.
Under the hypotheses of Theorem 1.1, we have
For a collection of nontrivial mean zero distributions with finite absolute third moments, we let
Clearly, Theorem 1.1 follows immediately from Proposition 1.1 and the following result.
The equality to 1 in Lemma 1.1 improves the upper bound of 3 shown in L1bounds . Although our interest here is in best universal constants, we note that Proposition 1.1 shows that is a distribution-specific -Berry–Esseen constant, in that
when are identically distributed according to . For instance, when is a mean zero uniform distribution and when is a mean zero two-point distribution; see Corollary 2.1 of L1bounds , and Lemmas 1.2 and 1.3 below.
We close this section with two preliminaries. The first collects some facts shown in L1bounds and the second demonstrates that to prove Lemma 1.1, it suffices to consider the class of random variables . Then, following Hoeffding Hoeffding (see also Utev ), in Section 2, we use a continuity property of to show that its supremum over is attained on finitely supported distributions. Exploiting a convexity-type property of the zero bias transformation on mixtures over distributions having equal variances, we reduce the calculation further to the calculation of the supremum over , the collection of all mean zero distributions with variance 1 and supported on at most three points. As three-point distributions are, in general, a mixture of two two-point distributions with unequal variances, an additional argument is given in Section 3, where a coupling of an with distribution to a variable having the zero bias distribution is constructed, using the optimal -couplings on the component two-point distributions of which is the mixture, in order to obtain for all . The lower bound (2) on is calculated in Section 4.
The following simple formula will be of some use. For and , we have
Let be the distribution of a nontrivial mean zero random variable supported on the two points . Then is uniformly distributed on ,
Being nontrivial, has positive variance and, from (6), we see that the density of at , which is proportional to , is zero outside and constant within it, so for . That has mean zero implies that the support points and satisfy and that gives positive probabilities and to and , respectively. The moment identities are immediate.
Making the change of variable and applying (9) with and yields
Let for some , let have distribution and, for , let denote the distribution of . Then and, in particular,
That has the same distribution as follows from (4). The identities , and (8) now imply the first claim. The second claim now follows from
Reduction to three-point distributions
is measurable. When is a probability measure on , the set function given by
is a probability measure, called the mixture of . With some slight abuse of notation, we let and denote expectations with respect to and , and let and be random variables with distributions and , respectively. For instance, for all functions which are integrable with respect to , we have
In particular, if is a collection of mean zero distributions with variances and absolute third moments , the mixture distribution has variance and third absolute moment given by
where both may be infinite. Note that implies that -almost surely and, therefore, that , the zero bias distribution, exists -almost surely.
Theorem 2.1 shows that the zero bias distribution of a mixture is a mixture of zero bias distributions, the mixing measure of which is the original measure weighted by the variance and rescaled. Define (arbitrarily, see Remark 2.1) the zero bias distribution of , a point mass at zero, to be . Write when and have the same distribution.
In particular, if and only if is a constant a.s.
The distribution exists as has mean zero and finite nonzero variance. Let have the zero bias distribution and let have distribution . For any absolutely continuous function for which the expectations below exist, we have
Since for all such , we conclude that .
If for any , then and, therefore,
Hence, the mixture gives zero weight to all corresponding , showing that may be defined arbitrarily.
and having distributions and , respectively. With a slight abuse of notation, we may write in place of when has distribution .
If is the mixture of a collection of mean zero, variance 1 random variables satisfying , then
If is a collection of mean zero, variance 1 random variables with finite absolute third moments and such that every distribution in can be represented as a mixture of distributions in , then
Since the variances of are constant, the distribution is the mixture of , by Theorem 2.1. Hence, applying (2), we have
Noting that and applying (2), we find that
Regarding (13), clearly, and the reverse inequality follows from (12).
Note that no bound of the type provided by Theorem 2.2 holds, in general, when taking mixtures of variables that have unequal variances. In particular, if and is not constant in , then is a mixture of normals with unequal variances, which is not normal. Hence, in this case, , whereas for all .
To apply Theorem 2.2 to reduce the computation of to finitely supported distributions, we apply the following continuity property of the zero bias transformation; see Lemma 5.2 in lin . We write for the convergence of to in distribution.
Let and be mean zero random variables with finite, nonzero variances. If
If is uniform on $F^{-1}(U)FX_{n}XF_{n}FX_{n}\Rightarrow XF_{n}^{-1}(U)\rightarrow F^{-1}(U)FG$, we have
where the infimum is over all joint distributions on which have marginals and , respectively, and the variables and achieve the minimal -coupling, that is,
With the use of Lemma 2.1, we are able to prove the following continuity property of the functional .
By Lemma 2.1, we have . Let be a uniformly distributed variable and set
By (4) with , we find, for , for example, that
Hence, as , we have and
Combining (2) with the convergence of the variances and the absolute third moments, as provided by (18), the proof is complete.
Lemmas 2.3 and 2.4 borrow much from Theorem 2.1 of Hoeffding , the latter lemma indeed being implicit. However, the results of Hoeffding cannot be applied directly as is not expressed as the expectation of for some when . For , let denote the collection of all mean zero, variance 1 distributions which are supported on at most points.
Letting be the collection of distributions in which have compact support, we first show that
we have , by Slutsky’s theorem, so, in view of (2), the hypotheses of Lemma 2.2 are satisfied, yielding
Since a.s., each is supported on finitely many points and uniformly bounded. Clearly, a.s. and (2) holds by the bounded convergence theorem. Now, defining by (22), the hypotheses of Lemma 2.2 are satisfied, yielding
showing . Combining this inequality with (20) yields and therefore the lemma, the reverse inequality being obvious.
Every distribution in can be expressed as a finite mixture of distributions.
The lemma is trivially true for , so consider and assume that the lemma holds for all integers from to .
The distribution of any is determined by the supporting values and a vector of probabilities . If any of the components of are zero, then for and the induction would be finished, so we assume that all components of are strictly positive. As , the vector must satisfy
Since and the equation specified by the first row of is , the vector contains both positive and negative numbers. Since the vector has strictly positive components, the numbers and given by
by (23), so that and are probability vectors since their components are nonnegative and sum to one. Additionally, the corresponding distributions have mean zero and variance 1, and in each of these two vectors, at least one component has been set to zero. Hence, we may express the -point probability vector as the mixture
of probability vectors on at most support points, thus showing to be the mixture of two distributions in , completing the induction.
The following theorem is an immediate consequence of Theorem 2.2 and Lemmas 2.3 and 2.4.
Hence, we now restrict our attention to .
Lemma 1.3 and Theorem 2.3 imply that . Hence, Lemma 1.1 follows from Theorem 3.1 below, which shows that . We prove Theorem 3.1 with the help of the following result.
Let , and let and be the unique mean zero distributions with support and , respectively, that is,
Let and denote the distribution functions of and , respectively. By Lemma 1.2, is uniform over . There are two cases, depending on the relative magnitudes of and .
Letting and , we have
Since for all ,
Recalling that , applying (9) with , and , after the change of variable , yields
Now, subtracting from (3) and simplifying by noting that the terms inside the parentheses in the numerators of these two expressions are equal, we find that
The denominator in (3) is positive, as is and . For the remaining term, (25) yields
Hence, (3) is positive, thus proving (24) when .
When , we have for all as is zero in and equals in , and is increasing over . Hence,
Now, since and , using Lemma 1.2, we obtain
thus proving inequality (24) when and, therefore, proving the lemma.
Lemma 1.2 shows that if is supported on two points, so and it only remains to consider positively supported on three points. We first prove that
implies that . After proving (3), we treat the remaining case, where , by a continuity argument.
Let be supported on with . Lemma 1.3 with implies that , so we may assume, without loss of generality, that . Let and be the unique mean zero distributions supported on and , respectively, and let and . As, in general, every mean zero distribution having no atom at zero can be represented as a mixture of mean zero two-point distributions (as in the Skorokhod representation, see durrett ), letting
we have for some ; in fact, for the given , one may verify that and that (32) holds when assumes this value. Therefore, to prove (3), it suffices to show that
and, by (32), the variance of is given by
Applying Theorem 2.1 with and being the probability measure putting mass and on the points and , respectively, in view of (34) and (3), , the zero bias distribution, is given by the mixture
Let and denote the distribution functions of and , respectively. Let be a standard uniform variable and, with the inverse functions below given by (15), set
Then , for and, by (17), all pairs of the variables achieve the -distance between their respective distributions. Now, let be defined on the same space with joint distribution given by the mixture
Then has marginals and , hence, by (16),
Lemma 1.2 shows that , that is, for , so (32) yields
and, by (34), (3) and (36), we now find that
Lemma 3.1 shows that the right-hand side and, therefore, the left-hand side of (3) are bounded by (3), that is, that , completing the proof of (33) and hence of (3).
Lower bound
By (1), with and ,
Motivated by Theorem 2.3, that two-point distributions achieve the suprema of , for , let
where is a Bernoulli variable with . The distribution function of is given by
and, therefore, the -distance between and the standard normal is given by
As for all and , letting
inequality (39) gives for all and yields (2).
Remarks
This article was submitted on November 18th, 2008. In the article Tyu , submitted on June 8th, 2009, Ilya Tyurin independently proved Theorem 1.1, also by applying the zero bias method. The current article was posted on arXiv on June 28, 2009; article Tyu was posted on December 3rd, 2009. In Tyu , Theorem 1.1 is used to prove the upper bound on the -Berry–Esseen constant.
Acknowledgment
The author would like to sincerely thank Sergey Utev for helpful suggestions.