Bounds on the constant in the mean central limit theorem

Larry Goldstein

Introduction

The classical central limit theorem allows the approximation of the distribution of sums of “comparable” independent real-valued random variables by the normal. As this theorem is an asymptotic, it provides no information as to whether the resulting approximation is useful. For that purpose, one may turn to the Berry–Esseen theorem, the most classical version giving supremum norm bounds between the distribution function of the normalized sum and that of the standard normal. Various authors have also considered Berry–Esseen-type bounds using other metrics and, in particular, bounds in LpL^{p}. The case p=1p=1, where the value

is used to measure the distance between distribution functions FF and GG, is of some particular interest and results using this metric are known as mean central limit theorems (see, e.g., dedecker , Erickson , HoChen and Ibragimov ; the latter three of these works consider nonindependent summand variables). One motivation for studying L1L^{1}-bounds is that, combined with one of type L∞L^{\infty}, bounds on LpL^{p}-distance for all p∈(1,∞)p\in(1,\infty) may be obtained by the inequality

For σ∈(0,∞)\sigma\in(0,\infty), let Fσ\mathcal{F}_{\sigma} be the collection of distributions with mean zero, variance σ2\sigma^{2} and finite absolute third moment. We prove the following Berry–Esseen-type result for the mean central limit theorem.

In particular, when X1,…,XnX_{1},\ldots,X_{n} are identically distributed with distribution G∈FσG\in\mathcal{F}_{\sigma},

For the case where all variables are identically distributed as XX having distribution GG, letting

the second part of Theorem 1.1 yields the upper bound c1≤1c_{1}\leq 1. Regarding lower bounds, we also prove

Clearly, the elements of the sequence {cm}m≥1\{c_{m}\}_{m\geq 1} are nonnegative and decreasing in mm, and so have a limit, say c∞c_{\infty}. Regarding limiting behavior, Esseen esseen showed that

for an explicit constant A(G)A(G) depending only on GG. Zolotarev Zolotarev provides the representation

where ω=∣EX3∣/(3σ2)\omega=|EX^{3}|/(3\sigma^{2}) and hh is the span of the distribution GG in the case where GG is lattice, and is zero otherwise. Zolotarev obtains

showing that c∞=1/2c_{\infty}=1/2, hence giving the asymptotic L1L^{1}-Berry–Esseen constant value.

for all absolutely continuous functions ff for which these expectations exist. The zero bias transformation, mapping the distribution of XX to that of X∗X^{*}, was motivated by the Stein characterization of the normal distribution Stein81 , which states that ZZ is normal with mean zero and variance σ2\sigma^{2} if and only if

for all absolutely continuous functions ff for which these expectations exist. Hence, the mean zero normal with variance σ2\sigma^{2} is the unique fixed point of the zero bias transformation. How closeness to normality may be measured by the closeness of a distribution to its zero bias transform, and related applications, are the topics of res , zsm and L1bounds .

As shown in L1bounds and Goldstein-Reinert , for a random variable XX with EX=0EX=0 and Var⁡(X)=σ2\operatorname{Var}(X)=\sigma^{2}, the distribution of X∗X^{*} is absolutely continuous with density and distribution functions given, respectively, by

Theorem 1.1 results by showing that the functional

is bounded by 1 for all XX with distribution G∈FσG\in\mathcal{F}_{\sigma}. As in (3), one may write out a more “explicit” form for B(G)B(G) using (6) and expressions for the moments on which B(G)B(G) depends, but such expressions appear to be of little value for the purposes of proving Theorem 1.1. In turn, the proof here employs convexity properties of B(G)B(G) which depend on the behavior of the zero bias transformation on mixtures. We also note that the functional B(G)B(G) has a different character than A(G)A(G); for instance, A(G)A(G) is zero for all nonlattice distributions with vanishing third moment, whereas B(G)B(G) is zero only for mean zero normal distributions.

by replacing σi2\sigma_{i}^{2} by σi2/σ2\sigma_{i}^{2}/\sigma^{2} and ∥Gi∗−Gi∥1\|G_{i}^{*}-G_{i}\|_{1} by ∥Gi∗−Gi∥1/σ\|G_{i}^{*}-G_{i}\|_{1}/\sigma in equation (16) of Theorem 2.1 of L1bounds , we obtain the following.

Under the hypotheses of Theorem 1.1, we have

For F\mathcal{F} a collection of nontrivial mean zero distributions with finite absolute third moments, we let

Clearly, Theorem 1.1 follows immediately from Proposition 1.1 and the following result.

The equality to 1 in Lemma 1.1 improves the upper bound of 3 shown in L1bounds . Although our interest here is in best universal constants, we note that Proposition 1.1 shows that B(G)B(G) is a distribution-specific L1L^{1}-Berry–Esseen constant, in that

when X1,…,XnX_{1},\ldots,X_{n} are identically distributed according to G∈FσG\in\mathcal{F}_{\sigma}. For instance, B(G)=1/3B(G)=1/3 when GG is a mean zero uniform distribution and B(G)=1B(G)=1 when GG is a mean zero two-point distribution; see Corollary 2.1 of L1bounds , and Lemmas 1.2 and 1.3 below.

We close this section with two preliminaries. The first collects some facts shown in L1bounds and the second demonstrates that to prove Lemma 1.1, it suffices to consider the class of random variables F1\mathcal{F}_{1}. Then, following Hoeffding Hoeffding (see also Utev ), in Section 2, we use a continuity property of B(G)B(G) to show that its supremum over F1\mathcal{F}_{1} is attained on finitely supported distributions. Exploiting a convexity-type property of the zero bias transformation on mixtures over distributions having equal variances, we reduce the calculation further to the calculation of the supremum over D3D_{3}, the collection of all mean zero distributions with variance 1 and supported on at most three points. As three-point distributions are, in general, a mixture of two two-point distributions with unequal variances, an additional argument is given in Section 3, where a coupling of an XX with distribution G∈D3G\in D_{3} to a variable X∗X^{*} having the XX zero bias distribution is constructed, using the optimal L1L^{1}-couplings on the component two-point distributions of which GG is the mixture, in order to obtain B(G)≤1B(G)\leq 1 for all G∈D3G\in D_{3}. The lower bound (2) on c1c_{1} is calculated in Section 4.

The following simple formula will be of some use. For a≥0,b>0a\geq 0,b>0 and l>0l>0, we have

Let GG be the distribution of a nontrivial mean zero random variable XX supported on the two points x<yx<y. Then X∗X^{*} is uniformly distributed on [x,y][x,y],

Being nontrivial, GG has positive variance and, from (6), we see that the density g∗g^{*} of G∗G^{*} at uu, which is proportional to E[X1(X>u)]E[X\mathbf{1}(X>u)], is zero outside [x,y][x,y] and constant within it, so G∗(w)=(w−x)/(y−x)G^{*}(w)=(w-x)/(y-x) for w∈[x,y]w\in[x,y]. That GG has mean zero implies that the support points xx and yy satisfy x<0<yx<0<y and that GG gives positive probabilities y/(y−x)y/(y-x) and −x/(y−x)-x/(y-x) to xx and yy, respectively. The moment identities are immediate.

Making the change of variable u=w−xu=w-x and applying (9) with a=y/(y−x),b=−x/(y−x)a=y/(y-x),b=-x/(y-x) and l=y−xl=y-x yields

Let G∈FσG\in\mathcal{F}_{\sigma} for some σ∈(0,∞)\sigma\in(0,\infty), let XX have distribution GG and, for a≠0a\not=0, let GaG_{a} denote the distribution of aXaX. Then B(Ga)=B(G)B(G_{a})=B(G) and, in particular,

That aX∗aX^{*} has the same distribution as (aX)∗(aX)^{*} follows from (4). The identities σaX2=a2σX2\sigma_{aX}^{2}=a^{2}\sigma_{X}^{2}, E∣aX∣3=∣a∣3E∣X3∣E|aX|^{3}=|a|^{3}E|X^{3}| and (8) now imply the first claim. The second claim now follows from

Reduction to three-point distributions

is measurable. When μ\mu is a probability measure on (S,Σ)(S,\Sigma), the set function given by

is a probability measure, called the μ\mu mixture of {ms}s∈S\{m_{s}\}_{s\in S}. With some slight abuse of notation, we let EμE_{\mu} and EsE_{s} denote expectations with respect to mμm_{\mu} and msm_{s}, and let XμX_{\mu} and XsX_{s} be random variables with distributions mμm_{\mu} and msm_{s}, respectively. For instance, for all functions ff which are integrable with respect to μ\mu, we have

In particular, if {ms}s∈S\{m_{s}\}_{s\in S} is a collection of mean zero distributions with variances σs2=EXs2\sigma_{s}^{2}=EX_{s}^{2} and absolute third moments γs=E∣Xs3∣\gamma_{s}=E|X_{s}^{3}|, the mixture distribution mμm_{\mu} has variance σμ2\sigma_{\mu}^{2} and third absolute moment γμ\gamma_{\mu} given by

where both may be infinite. Note that σμ2<∞\sigma_{\mu}^{2}<\infty implies that σs2<∞\sigma_{s}^{2}<\infty μ\mu-almost surely and, therefore, that ms∗m_{s}^{*}, the msm_{s} zero bias distribution, exists μ\mu-almost surely.

Theorem 2.1 shows that the zero bias distribution of a mixture is a mixture of zero bias distributions, the mixing measure of which is the original measure weighted by the variance and rescaled. Define (arbitrarily, see Remark 2.1) the zero bias distribution of δ0\delta_{0}, a point mass at zero, to be δ0\delta_{0}. Write X=dYX=_{d}Y when XX and YY have the same distribution.

In particular, ν=μ\nu=\mu if and only if σs2\sigma_{s}^{2} is a constant μ\mu a.s.

The distribution mμ∗m_{\mu}^{*} exists as mμm_{\mu} has mean zero and finite nonzero variance. Let Xμ∗X_{\mu}^{*} have the mμm_{\mu} zero bias distribution and let YY have distribution mμ∗m_{\mu}^{*}. For any absolutely continuous function ff for which the expectations below exist, we have

Since Ef′(Xμ∗)=Ef′(Y)Ef^{\prime}(X_{\mu}^{*})=Ef^{\prime}(Y) for all such ff, we conclude that Xμ∗=dYX_{\mu}^{*}=_{d}Y.

If ms=δ0m_{s}=\delta_{0} for any s∈Ss\in S, then σs2=0\sigma_{s}^{2}=0 and, therefore,

Hence, the mixture Xμ∗X_{\mu}^{*} gives zero weight to all corresponding ms∗m_{s}^{*}, showing that (δ0)∗(\delta_{0})^{*} may be defined arbitrarily.

XX and YY having distributions FF and GG, respectively. With a slight abuse of notation, we may write B(X)B(X) in place of B(G)B(G) when XX has distribution GG.

If XμX_{\mu} is the μ\mu mixture of a collection {Xs,s∈S}\{X_{s},s\in S\} of mean zero, variance 1 random variables satisfying E∣Xμ3∣<∞E|X_{\mu}^{3}|<\infty, then

If C\mathcal{C} is a collection of mean zero, variance 1 random variables with finite absolute third moments and D⊂C\mathcal{D}\subset\mathcal{C} such that every distribution in C\mathcal{C} can be represented as a mixture of distributions in D\mathcal{D}, then

Since the variances σs2\sigma_{s}^{2} of XsX_{s} are constant, the distribution Xμ∗X_{\mu}^{*} is the μ\mu mixture of {Xs∗,s∈S}\{X_{s}^{*},s\in S\}, by Theorem 2.1. Hence, applying (2), we have

Noting that Var⁡(Xμ)=∫SEXs2 dμ=1\operatorname{Var}(X_{\mu})=\int_{S}EX_{s}^{2}\,d\mu=1 and applying (2), we find that

Regarding (13), clearly, B(D)≤B(C)B(\mathcal{D})\leq B(\mathcal{C}) and the reverse inequality follows from (12).

Note that no bound of the type provided by Theorem 2.2 holds, in general, when taking mixtures of variables that have unequal variances. In particular, if Xs∼N(0,σs2)X_{s}\sim\mathcal{N}(0,\sigma_{s}^{2}) and σs2\sigma_{s}^{2} is not constant in ss, then XμX_{\mu} is a mixture of normals with unequal variances, which is not normal. Hence, in this case, B(Xμ)>0B(X_{\mu})>0, whereas B(Xs)=0B(X_{s})=0 for all ss.

To apply Theorem 2.2 to reduce the computation of B(F1)B(\mathcal{F}_{1}) to finitely supported distributions, we apply the following continuity property of the zero bias transformation; see Lemma 5.2 in lin . We write Xn⇒XX_{n}\Rightarrow X for the convergence of XnX_{n} to XX in distribution.

Let XX and Xn,n=1,2,…,X_{n},n=1,2,\ldots, be mean zero random variables with finite, nonzero variances. If

If UU is uniform on $,then, thenF^{-1}(U)hasdistributionfunctionhas distribution functionF.If. IfX_{n}andandXhavedistributionfunctionshave distribution functionsF_{n}andandF,respectively,and, respectively, andX_{n}\Rightarrow X,then, thenF_{n}^{-1}(U)\rightarrow F^{-1}(U)a.s.(see,e.g.,Theorem2.1ofdurrett,Chapter2).Fordistributionfunctionsa.s. (see, e.g., Theorem 2.1 of durrett , Chapter 2). For distribution functionsFandandG$, we have

where the infimum is over all joint distributions on X,YX,Y which have marginals FF and GG, respectively, and the variables F−1(U)F^{-1}(U) and G−1(U)G^{-1}(U) achieve the minimal L1L^{1}-coupling, that is,

With the use of Lemma 2.1, we are able to prove the following continuity property of the functional B(X)B(X).

By Lemma 2.1, we have Xn∗⇒X∗X_{n}^{*}\Rightarrow X^{*}. Let UU be a uniformly distributed variable and set

By (4) with f(x)=x2\mboxsgn(x)f(x)=x^{2}\mbox{sgn}(x), we find, for YY, for example, that

Hence, as n→∞n\rightarrow\infty, we have EYn2=EXn2→EX2=EY2EY_{n}^{2}=EX_{n}^{2}\rightarrow EX^{2}=EY^{2} and

Combining (2) with the convergence of the variances and the absolute third moments, as provided by (18), the proof is complete.

Lemmas 2.3 and 2.4 borrow much from Theorem 2.1 of Hoeffding , the latter lemma indeed being implicit. However, the results of Hoeffding cannot be applied directly as B(G)B(G) is not expressed as the expectation of K(X)K(X) for some KK when L(X)=G\mathcal{L}(X)=G. For m≥2m\geq 2, let DmD_{m} denote the collection of all mean zero, variance 1 distributions which are supported on at most mm points.

Letting M\mathcal{M} be the collection of distributions in F1\mathcal{F}_{1} which have compact support, we first show that

we have Xn⇒XX_{n}\Rightarrow X, by Slutsky’s theorem, so, in view of (2), the hypotheses of Lemma 2.2 are satisfied, yielding

Since ∣X∣≤M|X|\leq M a.s., each YnY_{n} is supported on finitely many points and uniformly bounded. Clearly, Yn→XY_{n}\rightarrow X a.s. and (2) holds by the bounded convergence theorem. Now, defining XnX_{n} by (22), the hypotheses of Lemma 2.2 are satisfied, yielding

showing B(M)≤B(⋃m≥3Dm)B(\mathcal{M})\leq B(\bigcup_{m\geq 3}D_{m}). Combining this inequality with (20) yields B(F1)≤B(⋃m≥3Dm)B(\mathcal{F}_{1})\leq B(\bigcup_{m\geq 3}D_{m}) and therefore the lemma, the reverse inequality being obvious.

Every distribution in ⋃m≥3Dm\bigcup_{m\geq 3}D_{m} can be expressed as a finite mixture of D3D_{3} distributions.

The lemma is trivially true for m=3m=3, so consider m>3m>3 and assume that the lemma holds for all integers from 33 to m−1m-1.

The distribution of any X∈DmX\in D_{m} is determined by the supporting values a1<⋯<ama_{1}<\cdots<a_{m} and a vector of probabilities p=(p1,…,pm)′\mathbf{p}=(p_{1},\ldots,p_{m})^{\prime}. If any of the components of p\mathbf{p} are zero, then X∈DkX\in D_{k} for k<mk<m and the induction would be finished, so we assume that all components of p\mathbf{p} are strictly positive. As X∈DmX\in D_{m}, the vector p\mathbf{p} must satisfy

Since v≠0\mathbf{v}\not=0 and the equation specified by the first row of AA is ∑ivi=0\sum_{i}v_{i}=0, the vector v\mathbf{v} contains both positive and negative numbers. Since the vector p\mathbf{p} has strictly positive components, the numbers t1t_{1} and t2t_{2} given by

by (23), so that p1\mathbf{p}_{1} and p2\mathbf{p}_{2} are probability vectors since their components are nonnegative and sum to one. Additionally, the corresponding distributions have mean zero and variance 1, and in each of these two vectors, at least one component has been set to zero. Hence, we may express the mm-point probability vector p\mathbf{p} as the mixture

of probability vectors on at most m−1m-1 support points, thus showing XX to be the mixture of two distributions in Dm−1D_{m-1}, completing the induction.

The following theorem is an immediate consequence of Theorem 2.2 and Lemmas 2.3 and 2.4.

Hence, we now restrict our attention to D3D_{3}.

Lemma 1.3 and Theorem 2.3 imply that B(Fσ)=B(F1)=B(D3)B(\mathcal{F}_{\sigma})=B(\mathcal{F}_{1})=B(D_{3}). Hence, Lemma 1.1 follows from Theorem 3.1 below, which shows that B(D3)=1B(D_{3})=1. We prove Theorem 3.1 with the help of the following result.

Let x<y<0<zx<y<0<z, and let m1m_{1} and m0m_{0} be the unique mean zero distributions with support {x,z}\{x,z\} and {y,z}\{y,z\}, respectively, that is,

Let F1,F0F_{1},F_{0} and F1∗F_{1}^{*} denote the distribution functions of m1,m0m_{1},m_{0} and m1∗m_{1}^{*}, respectively. By Lemma 1.2, m1∗m_{1}^{*} is uniform over [x,z][x,z]. There are two cases, depending on the relative magnitudes of F1∗(y)=(y−x)/(z−x)F_{1}^{*}(y)=(y-x)/(z-x) and F0(y)=z/(z−y)F_{0}(y)=z/(z-y).

Letting J1=[x,y)J_{1}=[x,y) and J2=[y,z]J_{2}=[y,z], we have

Since F1∗(w)≥0=F0(w)F_{1}^{*}(w)\geq 0=F_{0}(w) for all w∈J1w\in J_{1},

Recalling that F1∗(y)≤F0(y)F_{1}^{*}(y)\leq F_{0}(y), applying (9) with a=zz−y−y−xz−xa=\frac{z}{z-y}-\frac{y-x}{z-x}, b=−yz−yb=-\frac{y}{z-y} and l=z−yl=z-y, after the change of variable u=w−yu=w-y, yields

Now, subtracting from (3) and simplifying by noting that the terms inside the parentheses in the numerators of these two expressions are equal, we find that

The denominator in (3) is positive, as is −y-y and y−xy-x. For the remaining term, (25) yields

Hence, (3) is positive, thus proving (24) when F1∗(y)≤F0(y)F_{1}^{*}(y)\leq F_{0}(y).

When F1∗(y)>F0(y)F_{1}^{*}(y)>F_{0}(y), we have F1∗(w)≥F0(w)F_{1}^{*}(w)\geq F_{0}(w) for all w∈[x,z)w\in[x,z) as F0(w)F_{0}(w) is zero in [x,y)[x,y) and equals F0(y)F_{0}(y) in [y,z)[y,z), and F1∗(w)F_{1}^{*}(w) is increasing over [y,z)[y,z). Hence,

Now, since (x+z)(x−z)=x2−z2≤z2+x2(x+z)(x-z)=x^{2}-z^{2}\leq z^{2}+x^{2} and z−x>0z-x>0, using Lemma 1.2, we obtain

thus proving inequality (24) when F1∗(y)>F0(y)F_{1}^{*}(y)>F_{0}(y) and, therefore, proving the lemma.

Lemma 1.2 shows that B(X)=1B(X)=1 if XX is supported on two points, so B(D3)≥1B(D_{3})\geq 1 and it only remains to consider XX positively supported on three points. We first prove that

EX=0EX=0 implies that x<0<zx<0<z. After proving (3), we treat the remaining case, where y=0y=0, by a continuity argument.

Let XX be supported on x<y<zx<y<z with y≠0y\not=0. Lemma 1.3 with a=−1a=-1 implies that B(−X)=B(X)B(-X)=B(X), so we may assume, without loss of generality, that x<y<0<zx<y<0<z. Let m1m_{1} and m0m_{0} be the unique mean zero distributions supported on {x,z}\{x,z\} and {y,z}\{y,z\}, respectively, and let L(X1)=m1\mathcal{L}(X_{1})=m_{1} and L(X0)=m0\mathcal{L}(X_{0})=m_{0}. As, in general, every mean zero distribution having no atom at zero can be represented as a mixture of mean zero two-point distributions (as in the Skorokhod representation, see durrett ), letting

we have L(X)=L(Xα)\mathcal{L}(X)=\mathcal{L}(X_{\alpha}) for some α∈\alpha\in; in fact, for the given XX, one may verify that P(X=x)/P(X1=x)∈(0,1)P(X=x)/P(X_{1}=x)\in(0,1) and that (32) holds when α\alpha assumes this value. Therefore, to prove (3), it suffices to show that

and, by (32), the variance of XαX_{\alpha} is given by

Applying Theorem 2.1 with S={0,1}S=\{0,1\} and μ\mu being the probability measure putting mass α\alpha and 1−α1-\alpha on the points 11 and , respectively, in view of (34) and (3), mα∗m_{\alpha}^{*}, the XαX_{\alpha} zero bias distribution, is given by the mixture

Let F1,F0,F1∗F_{1},F_{0},F_{1}^{*} and F0∗F_{0}^{*} denote the distribution functions of m1,m0,m1∗m_{1},m_{0},m_{1}^{*} and m0∗m_{0}^{*}, respectively. Let UU be a standard uniform variable and, with the inverse functions below given by (15), set

Then Yi=dXiY_{i}=_{d}X_{i}, Yi∗=dXi∗Y_{i}^{*}=_{d}X_{i}^{*} for i∈{1,2}i\in\{1,2\} and, by (17), all pairs of the variables Y1,Y0,Y1∗,Y0∗Y_{1},Y_{0},Y_{1}^{*},Y_{0}^{*} achieve the L1L^{1}-distance between their respective distributions. Now, let (Yα,Yα∗)(Y_{\alpha},Y_{\alpha}^{*}) be defined on the same space with joint distribution given by the mixture

Then (Yα,Yα∗)(Y_{\alpha},Y_{\alpha}^{*}) has marginals Yα=dXαY_{\alpha}=_{d}X_{\alpha} and Yα∗=dYα∗Y_{\alpha}^{*}=_{d}Y_{\alpha}^{*}, hence, by (16),

Lemma 1.2 shows that G(Xi)=1G(X_{i})=1, that is, E∣Xi3∣=2EXi2∥mi∗−mi∥1E|X_{i}^{3}|=2EX_{i}^{2}\|m_{i}^{*}-m_{i}\|_{1} for i=1,2i=1,2, so (32) yields

and, by (34), (3) and (36), we now find that

Lemma 3.1 shows that the right-hand side and, therefore, the left-hand side of (3) are bounded by (3), that is, that B(Xα)=2EXα2∥mα∗−mα∥1/E∣Xα3∣≤1B(X_{\alpha})=2EX_{\alpha}^{2}\|m_{\alpha}^{*}-m_{\alpha}\|_{1}/E|X_{\alpha}^{3}|\leq 1, completing the proof of (33) and hence of (3).

Lower bound

By (1), with m=1m=1 and L(X)=G∈F1\mathcal{L}(X)=G\in\mathcal{F}_{1},

Motivated by Theorem 2.3, that two-point distributions achieve the suprema of B(G)B(G), for p∈(0,1)p\in(0,1), let

where ξ\xi is a Bernoulli variable with P(ξ=1)=p=1−P(ξ=0)P(\xi=1)=p=1-P(\xi=0). The distribution function GpG_{p} of XX is given by

and, therefore, the L1L^{1}-distance between GpG_{p} and the standard normal is given by

As Gp∈F1G_{p}\in\mathcal{F}_{1} for all p∈(0,1)p\in(0,1) and E∣X3∣=(p2+q2)/pqE|X^{3}|=(p^{2}+q^{2})/\sqrt{pq}, letting

inequality (39) gives c1≥ψ(p)c_{1}\geq\psi(p) for all p∈(0,1)p\in(0,1) and ψ(1/2)\psi(1/2) yields (2).

Remarks

This article was submitted on November 18th, 2008. In the article Tyu , submitted on June 8th, 2009, Ilya Tyurin independently proved Theorem 1.1, also by applying the zero bias method. The current article was posted on arXiv on June 28, 2009; article Tyu was posted on December 3rd, 2009. In Tyu , Theorem 1.1 is used to prove the upper bound 0.47850.4785 on the L∞L^{\infty}-Berry–Esseen constant.

Acknowledgment

The author would like to sincerely thank Sergey Utev for helpful suggestions.

References