New estimates of the convergence rate in the Lyapunov theorem

Ilya Tyurin

Introduction

Consider centered independent (real-valued) random variables (r.v.) X1,…,XnX_{1},\ldots,X_{n} with variances σ12,…,σn2\sigma_{1}^{2},\ldots,\sigma_{n}^{2} and finite absolute moments β1,…,βn\beta_{1},\ldots,\beta_{n}. We denote

According to the Lyapunov theorem, Sn:=(X1+…+Xn)/σ(n)S_{n}:=(X_{1}+\ldots+X_{n})/\sigma(n) converges in distribution to the standard normal r.v. when εn→0\varepsilon_{n}\to 0. From both theoretical and practical points of view, it is very important to estimate the convergence rate in this theorem. It is known that there exists a minimal numerical constant CC such that for the Kolmogorov distance between SnS_{n} and the standard normal variable NN holds the inequality

There are plenty of works devoted to estimation of this constant. Esseen showed that C⩽7.5C\leqslant 7{.}5. Bergström obtained the bound C⩽4.8C\leqslant 4{.}8. Takano established that in the case of independent identically distributed (i.i.d.) summands C⩽2.031C\leqslant 2{.}031. Zolotarev obtained a new inequality allowing to estimate the proximity of two sums of independent r.v. With the help of this inequality he showed successively that C⩽1.322C\leqslant 1{.}322 and C⩽0.9051C\leqslant 0{.}9051, while in the case of i.i.d. variables C⩽1.301C\leqslant 1{.}301 and C⩽0.8197C\leqslant 0{.}8197. The proposed method was further developed in the works of van Beek and Shiganov , who proved the estimates C⩽0.7975C\leqslant 0{.}7975 and C⩽0.7915C\leqslant 0{.}7915, respectively. For the sums of identically distributed r.v. Shiganov obtained the bound C⩽0.7655C\leqslant 0{.}7655, which was sharpened in 2006 by Shevtsova . She showed that in this case C⩽0.7056C\leqslant 0{.}7056. In we derived the estimates C⩽0.6379C\leqslant 0{.}6379 in the general case and C⩽0.5894C\leqslant 0.5894 for identically distributed summands. In the present paper we improve them.

From a private communication with Korolev and Shevtsova we know that recently they have established the convergence rate in the central limit theorem in a variety of sences . In these works only the i.i.d. case was considered and the bound for the constant CC is not as sharp as ours. However, interesting estimates of the other kind were obtained.

It is worth mentioning the related problem of determining the asymptotically best constants in Lyapunov’s theorem. As it was shown by Esseen , if all the r.v. Xj,j=1,2,…X_{j},\enskip j=1,2,\ldots have the same distribution, then

and the constant on the right-hand side of this inequality cannot be lowered (hence the lower bound C⩾C1C\geqslant C_{1}). This result was elaborated by Rogozin , who established that under the same assumptions

where ρ(Sn,N):=inf⁡G∈Nρ(Sn,G)\rho(S_{n},\mathcal{N}):=\inf\limits_{G\in\mathcal{N}}\rho(S_{n},G) and N\mathcal{N} is the set of all normal r.v.

Chistyakov generalized (2) and (3) to the case of nonidentically distributed summands. He proved that

where r1(εn),r2(εn)r_{1}(\varepsilon_{n}),r_{2}(\varepsilon_{n}) are o(εn)o(\varepsilon_{n}) when εn→0\varepsilon_{n}\to 0.

There are also the estimates of the convergence rate in the Lyapunov theorem provided that the moments of the order 2+δ2+\delta exist (see ).

Analogues of (1) are known for other probability metrics as well, for example, ζr\zeta_{r} (where r=1,2,3r=1,2,3). The latter will be described in detail in section 2. Estimates in terms of these metrics can be obtained in a natural way using Stein’s method. The proof of the estimates mentioned above uses, in particular, the so-called zero bias transformation of a probability distribution (see ).

For the distance in terms of metrics ζr\zeta_{r} (r=1,2,3)(r=1,2,3) the following estimates (see ) are known:

Hoeffding considered the problem of finding the least upper bound of Ef(X1,…,Xn){\sf E}f(X_{1},\ldots,X_{n}) over the set of all collections of independent simple r.v. satisfying mm restrictions of the form \Egij(Xj)=cij,j=1,…,n\E g_{ij}(X_{j})=c_{ij},j=1,\ldots,n. More precisely, it was established that in this case one has to consider only r.v. taking at most m+1m+1 values. In the present work results of are generalized to the case of arbitrary quasiconvex functional defined on the set of all probability distributions.

The results obtained allowed us to derive an unimprovable estimate for the proximity in the mean metric between a probability distribution and its zero bias transformation. The latter was used to estimate the accuracy of the Gaussian approximation for the sums of independent variates. It was established that the values of constants in (4) can be taken 3 times lower. In addition, our estimate for the metric ζ3\zeta_{3} is optimal. Furthermore, new estimates for the difference between the characteristic functions of the normalized sum and the standard normal r.v. were derived, which allowed us to prove that C⩽0.5606C\leqslant 0{.}5606 and in the case of i.i.d. summands C⩽0.4785C\leqslant 0{.}4785.

Main notions and results

It is easy to see that QQ forms a linear space. And the set D of discrete probability distributions that are concentrated on finite sets of points is a convex subset of QQ. The latter means that αμ1+(1−α)μ2∈D \alpha\mu_{1}+(1-\alpha)\mu_{2}\in D\, for arbitrary μ1,μ2∈D\mu_{1},\mu_{2}\in D and α∈(0,1)\alpha\in(0,1).

Consider the set of all collections consisting of nn independent r.v. X1,…,XnX_{1},\ldots,X_{n}. Then

We assume that on SS some real-valued functions h1,…,hmh_{1},\ldots,h_{m} are defined. Consider the set

In this expression, we assume that the supremum over the empty set is zero.

Theorem 2. Let ff be a nonnegative function on SS, VV – a linear space with the norm ∥⋅∥\|\cdot\|, A:K→VA:K\to V – such a mapping that

for arbitrary μ,ν∈K,α∈(0,1)\mu,\nu\in K,\alpha\in(0,1). Then the least value of γ\gamma such that the inequality

holds for every measure μ∈K\mu\in K, coincides with the least value of γ\gamma such that (\refmincnst)(\ref{mincnst}) is true for every measure μ∈Km+1\mu\in K_{m+1}.

Let WW be a zero-mean r.v. with variance σ2>0\sigma^{2}>0. A r.v. W∗W^{*} is said to have the WW-zero biased distribution if

where Fr\mathcal{F}_{r} is the set of all real bounded functions with Mr(f)⩽1M_{r}(f)\leqslant 1.

Note that ζ1\zeta_{1} has alternative representations. These are the so-called mean metric

Theorem 3. If WW is a centered r.v. with unit variance and finite third absolute moment, then

with equality when WW has a 2-point distribution.

Corollary 1. Consider a r.v. Sn∗S_{n}^{*} having the SnS_{n}-zero biased distribution. Then

Theorem 4. The following inequalities are true:

The latter double inequality is optimal, namely, for every δ>0\delta>0 there exists such a sequence of i.i.d. r.v. X1,X2,…X_{1},X_{2},\ldots, that

and MM is the point where this maximum is attained, M≈3.995896M\approx 3{.}995896.

The quantities δ^1(ε,t)\widehat{\delta}_{1}(\varepsilon,t) and δ^2(ε,t)\widehat{\delta}_{2}(\varepsilon,t) can be expressed in terms of the so-called Dawson integral

which can be computed by the means of several efficient numerical procedures. For example, such a function is available in the GNU Scientific Library (GSL). It is easy to check that

These representations are of great importance, since they allow to reduce significantly the amount of numerical calculations required for the proof of Theorem 7.

In the case of i.i.d. variables the estimates can be slightly improved. Denote τn := 1σ3∑j=1nσj3\tau_{n}~:=~\frac{1}{\sigma^{3}}\sum_{j=1}^{n}\sigma_{j}^{3}, and let X1,X2,…X_{1},X_{2},\ldots be centered i.i.d. r.v. with unit variances and finite third absolute moments β\beta. Then εn=β/n\varepsilon_{n}=\beta/\sqrt{n} and τn=1/n\tau_{n}=1/\sqrt{n}.

Estimates (12)-(18) allowed to establish the following result.

Theorem 7. The constant CC in inequality (\refuniform)(\ref{uniform}) does not exceed 0.56060{.}5606, and in the case of identically distributed summands C⩽0.4785C\leqslant 0{.}4785.

Proofs

Proof of Theorem 1. If K=∅K=\varnothing, then Km+1=∅K_{m+1}=\varnothing, and the statement of our theorem is true. Further we suppose that the set KK is nonempty.

The sequence of the sets K1,K2,…K_{1},K_{2},\ldots increases to the set KK. Therefore,

Let’s take an arbitrary measure μ∈Kj\mu\in K_{j}, where j>m+1j>m+1, and show that there exists μ′∈Kj−1\mu^{\prime}\in K_{j-1} such that g(μ′)⩾g(μ)g(\mu^{\prime})\geqslant g(\mu).

Let μ\mu be concentrated in points s1,…,sj∈Ss_{1},\ldots,s_{j}\in S and

The vector μˉ=(μ1,…,μj)\bar{\mu}=(\mu_{1},\ldots,\mu_{j}) defines a probability distribution, so

Moreover, the conditions  ⟨hi,μ⟩=0,i=1,…,m, \,\langle h_{i},\mu\rangle=0,\enskip i=1,\ldots,m,\, hold and therefore

Vice versa, an arbitrary vector with nonnegative coordinates satisfying the system of linear equations (20) and (21) defines according to (19) an element of the set KjK_{j}, and if one of its coordinates equals zero – an element of Kj−1K_{j-1}. We have m+1m+1 equations and at least m+2m+2 unknowns, so there exists a nonzero solution νˉ=(ν1,…,νj)\bar{\nu}=(\nu_{1},\ldots,\nu_{j}) of the corresponding homogeneous system. Since the sum of the coordinates of this vector is equal to zero, but the vector itself is nonzero, it follows that νˉ\bar{\nu} has both positive and negative coordinates. Therefore, there exist the least α⩾0\alpha\geqslant 0 and the least β⩾0\beta\geqslant 0 such that one of the coordinates of the vector μ∗ˉ=μˉ−ανˉ\bar{\mu_{*}}=\bar{\mu}-\alpha\bar{\nu} equals zero and some coordinate of μ∗ˉ=μˉ+βνˉ\bar{\mu^{*}}=\bar{\mu}+\beta\bar{\nu} is equal to zero. If α=0\alpha=0, then μ∈Kj−1\mu\in K_{j-1}. Otherwise,

and because of the quasiconvexity g(μ)⩽max⁡{g(μ∗),g(μ∗)}g(\mu)\leqslant\max\{g(\mu_{*}),g(\mu^{*})\}, where μ∗,μ∗\mu_{*},\mu^{*} are the distributions defined by μ∗ˉ\bar{\mu_{*}} and μ∗ˉ\bar{\mu^{*}}. Thus, g(μ)⩽g(μ∗)g(\mu)\leqslant g(\mu_{*}) or g(μ)⩽g(μ∗)g(\mu)\leqslant g(\mu^{*}). But μ∗\mu_{*} and μ∗∈Kj−1.□\mu^{*}\in K_{j-1}.\enskip\square

Proof of Theorem 2. According to Theorem 1, it is sufficient to prove that for every fixed value of γ\gamma the function

is quasiconvex. Let α+β=1\alpha+\beta=1. By the properties of the norm

Proof of Theorem 3. We begin by showing that without loss of generality we can consider simple r.v. WW. It is sufficient to establish that for every r.v. WW satisfying the conditions of the theorem there exists a sequence (Wn)n⩾1(W_{n})_{n\geqslant 1} of simple r.v. with zero means and unit variances such that

It is easy to see that WnW_{n} converges to WW in L3L_{3} as well. Therefore, the second condition in (22) is obviously satisfied. It remains to show that the first one also holds.

From the triangle inequality for the metric ζ1\zeta_{1} one can easily derive that

The first summand on the right-hand side of (23) tends to zero, since

Let’s evaluate the second summand. For the function f∈F1f\in\mathcal{F}_{1} we set F(x):=∫0xf(u)duF(x):=\int_{0}^{x}f(u)du. Then

The difference of expectations on the left-hand side of (24) does not change if we replace the function f(x)f(x) by f(x)−f(0)f(x)-f(0). Therefore, we can assume without loss of generality that f(0)=0f(0)=0. Then ∣f(x)−f(y)∣⩽∣x−y∣|f(x)-f(y)|\leqslant|x-y| yields ∣f(x)∣⩽∣x∣|f(x)|\leqslant|x| and so ∣F(x)∣⩽∣x∣2,∣xf(x)∣⩽∣x∣2|F(x)|\leqslant|x|^{2},|xf(x)|\leqslant|x|^{2}. According to the finite-increment theorem,

where ξ\xi is a number between WW and WnW_{n}. Moreover,

And finally, the Hölder’s inequality yields

Obviously, (E∣W−Wn∣3)13→0\left({\sf E}|W-W_{n}|^{3}\right)^{\frac{1}{3}}\to 0, since WnW_{n} converges to WW in L3L_{3}. Thus, the second summand in (23) tends to zero and (22) is fulfilled.

So, it is sufficient to consider simple r.v. Let A1A_{1} be a function that maps the distribution PXP_{X} of a r.v. XX to its zero-biased distribution PX∗P_{X}^{*}. Moreover, consider a linear operator A2A_{2} that maps a signed measure ν\nu to its cumulative distribution function (c.d.f.) Gν(x):=ν((−∞,x]){G_{\nu}(x):=\nu\left((-\infty,x]\right)}. It is easy to see that

hence the mapping A2−A2A1A_{2}-A_{2}A_{1} satisfies (5). If we set h1(x)=xh_{1}(x)=x, h2(x)=x2−1h_{2}(x)=x^{2}-1 and apply Theorem 2 to f(x)=∣x∣3f(x)=|x|^{3}, A=A2−A2A1A=A_{2}-A_{2}A_{1}, VV – the normed space of integrable functions on the real line with the norm

then the problem reduces to the case of simple r.v. taking at most 3 values. C.d.f. of a simple r.v. WW is a staircase function. Using formula (8) one can easily obtain the c.d.f. of W∗W^{*}. Therefore, it is not difficult to find the explicit expression for ϰ1(W,W∗)\varkappa_{1}(W,W^{*}).

Let WW take exactly two values −x-x and yy with probabilities pp and qq. Then its c.d.f. is piecewise constant and has two steps in points −x-x and yy that are equal to pp and qq, respectively. Since WW is centered, we have px=qypx=qy, which together with (8) yields that W∗W^{*} is uniformly distributed on [−x,y][-x,y]. Therefore, on [−x,y][-x,y] its c.d.f. is linear, and its graph is a segment that connects (−x,0)(-x,0) and (y,1)(y,1). By definition ϰ1(W,W∗)\varkappa_{1}(W,W^{*}) equals the area of the figure bounded by distribution functions of these r.v. (in the case considered it is a union of two triangles, see pic. 1, left).

It follows from the conditions \EW=0\E W=0 and \EW2=1\E W^{2}=1 that x=q/p,y=p/q.x=\sqrt{q/p},\enskip y=\sqrt{p/q}. Hence

Let’s find the area of the figure bounded by c.d.f. of r.v. WW and W∗W^{*}. Density of W∗W^{*} is px=qy=pqpx=qy=\sqrt{pq}. Thus, the slope of the c.d.f. of this r.v. on [-x,y] is pq\sqrt{pq}. The length of the vertical leg of the first triangle is pp, and that of the second one is qq. Hence, the total area of both triangles is

Therefore, if WW takes exactly two values, there is equality in (9).

Consider the case when WW takes three values. We assume without loss of generality that two of them (−a<−b-a<-b) do not exceed zero and one (cc) is positive. As before, the c.d.f. of WW is piecewise constant and the c.d.f. of W∗W^{*} is piecewise linear. However, the form of the figure bounded by them is more complicated (see pic. 1, right). Denote by RR the value of the c.d.f. of W∗W^{*} at −b-b and by SS its value at 00. Let WW take values −a,−b,c-a,-b,c with probabilities p,q,rp,q,r, respectively. Then, because of the moment-type restrictions,

It is a system of linear equations with respect to p,q,rp,q,r. Using Cramer’s rule, we obtain

where Δ=(a+c)(b+c)(a−b)\Delta=(a+c)(b+c)(a-b). Thus, every r.v. with zero mean and variance 1 that takes three values is uniquely determined by these three values. It is easy to see that p,q,rp,q,r are nonnegative iff

In other words, a r.v. WW taking the values −a,−b,c-a,-b,c exists iff (26) is satisfied. Our aim is to prove that the function

does not exceed zero. Its explicit form in terms of variables a,b,ca,b,c depends on how the c.d.f. of r.v. WW and W∗W^{*} are located with respect to each other. There are 5 cases:

I. R⩽pR\leqslant p, S⩽p+qS\leqslant p+q, or, equivalently, a(a−b)⩽1a(a-b)\leqslant 1, c⩾1c\geqslant 1. In this case

II. R⩽pR\leqslant p, S⩾p+q⇔a(a−b)⩽1S\geqslant p+q\Leftrightarrow a(a-b)\leqslant 1, c⩽1c\leqslant 1.

III. R⩾pR\geqslant p, S⩽p+q⇔a(a−b)⩾1S\leqslant p+q\Leftrightarrow a(a-b)\geqslant 1, c⩾1c\geqslant 1 (the latter implies c(b+c)⩾1c(b+c)\geqslant 1).

IV. p⩽R⩽p+qp\leqslant R\leqslant p+q, S⩾p+q⇔a(a−b)⩾1S\geqslant p+q\Leftrightarrow a(a-b)\geqslant 1, c(b+c)⩾1c(b+c)\geqslant 1, c⩽1c\leqslant 1.

V. R⩾p+q⇔c(b+c)⩽1R\geqslant p+q\Leftrightarrow c(b+c)\leqslant 1.

Note that in each of these cases gg is the same function defined by (27). As a result, if the values a,b,ca,b,c satisfy the restrictions of two cases simultaneously, then for the function gg we can use the expression corresponding to any of them.

As one can see, in the cases I, II, and in the cases III, IV the function gg has the same representation. Therefore, further we distinguish three possibilities:

B. a(a−b)⩾1a(a-b)\geqslant 1, c(b+c)⩾1c(b+c)\geqslant 1

We show that in each of the cases A, B and C the function gg does not exceed zero.

because of (26), nonnegativeness of a,b,ca,b,c and the fact that a>ba>b.

Since ac⩾1ac\geqslant 1, it suffices to prove that the expression enclosed by braces does not exceed zero. Consider this expression as a function of the variable cc while holding the others fixed. When c=0c=0, this function equals −b⩽0-b\leqslant 0. Moreover, it decreases with respect to cc, since the coefficients of terms cc and c2c^{2} are negative. Consequently, for all positive values of cc it does not exceed zero.

where k2=1−2a(a−b)(1+ab),k1=b[1−a(a−b)(1+ab)],k0=a(a−b)(1+ab).\enskip k_{2}=1-2a(a-b)(1+ab),\enskip k_{1}=b\left[1-a(a-b)(1+ab)\right],\enskip k_{0}=a(a-b)(1+ab). Assume that g(a,b,c)>0g(a,b,c)>0. Due to the condition a(a−b)⩾1a(a-b)\geqslant 1 one has

Therefore, k2k_{2} and k1k_{1} do not exceed zero and, consequently, k2c2+k1c+k0k_{2}c^{2}+k_{1}c+k_{0} decreases with respect to cc. Consequently, if one reduces the value of the variable cc while holding aa and bb fixed, gg will remain positive. The variable cc is bounded from below by two conditions:

The first of these conditions can be omitted, since it follows from the other two:

Therefore, we can reduce cc to the value c∗c_{*} such that c∗(b+c∗)=1c_{*}(b+c_{*})=1. And gg will remain positive. But the situation, when c(b+c)=1c(b+c)=1, satisfies the restrictions of the case C, for which we established that g⩽0g\leqslant 0 – a contradiction. □\enskip\square

Proof of Corollary 1. Without loss of generality assume σ\sigma = 1. Let II be a random index taking values 1,…,n1,\ldots,n with probabilities σ12,…,σn2\sigma_{1}^{2},\ldots,\sigma_{n}^{2}, independent of X1,…,XnX_{1},\ldots,X_{n}. Construct on an extended probability space

where Xi∗X_{i}^{*} has the XiX_{i}-zero biased distribution and is independent of I,X1,…,XnI,X_{1},\ldots,X_{n}, i=1,…,n{i=1,\ldots,n}. Then Sn∗=SI′S_{n}^{*}=S_{I}^{\prime} (see ). Therefore, for an arbitrary function f∈F1f\in\mathcal{F}_{1} one has

Here we used the statement of Theorem 3 for r.v. 1σkXk\frac{1}{\sigma_{k}}X_{k} as well as the property of homogeneity of the metric ζ1\zeta_{1} (i. e. ζ1(cX,cY)=cζ1(X,Y)\zeta_{1}(cX,cY)=c\zeta_{1}(X,Y)) and (αX)∗=DαX∗(\alpha X)^{*}\stackrel{{\scriptstyle D}}{{=}}\alpha X^{*}. □\enskip\square

Taking into account that Mr(f)⩽1M_{r}(f)\leqslant 1 in the definition of the metric ζr\zeta_{r}, one has

Estimates in terms of εn\varepsilon_{n} are obtained by applying Corollary 1 to (29).

Let’s prove the optimality of (11). We set f(x):=x3/6f(x):=x^{3}/6. Then M3(f)=1M_{3}(f)=1 and \Ef(N)=0{\E f(N)=0}, since the r.v. NN is symmetric and the function ff is odd. Consider X1,X2,…X_{1},X_{2},\ldots – a sequence of i.i.d. variables with zero means and unit variances. Then

It only remains to prove that ∣\EX13∣/\E∣X1∣3|\E X_{1}^{3}|/\E|X_{1}|^{3} can be arbitrarily close to unity. According to (25), the third absolute moment of a centered r.v. X1X_{1} with variance 1 taking two values x=−q/p,  y=p/qx=-\sqrt{q/p},\,\,y=\sqrt{p/q} with probabilities pp and qq, respectively, is equal to

It’s easy to see that the third moment of this r.v. equals

Obviously, ∣\EX1∣/\E∣X1∣3→1|\E X_{1}|/\E|X_{1}|^{3}\to 1 when p→1p\to 1. □\enskip\square

Proof. This can be checked directly by calculating the derivative.□\enskip\square

Lemma 3. If WW is a centered r.v. with variance 1 and f(t)=\EeitWf(t)=\E e^{itW}, then

where f∗(t)f^{*}(t) is the characteristic function of a r.v. W∗W^{*} having the WW-zero biased distribution.

Proof. According to the definition of the WW-zero biased distribution,

Consider the function ψ(t):=f(t)φ(t)\psi(t):=\frac{f(t)}{\varphi(t)}. Note that ψ(0)=1\psi(0)=1. Taking into account (33), we have

Lemma 4. For arbitrary r.v. XX and YY we have

Thus, for arbitrary X~,Y~\widetilde{X},\widetilde{Y} defined on one probability space such that Law(X~)=Law(X)Law(\widetilde{X})=Law(X) and Law(Y~)=Law(Y)Law(\widetilde{Y})=Law(Y), we have

Passing in (34) to the greatest lower bound among every possible X~,Y~\widetilde{X},\widetilde{Y}, we obtain

Proof of Theorem 5. The inequality (12) is a consequence of Lemma 1. Indeed, according to the Lyapunov inequality, we have σj3⩽βj,j=1,…,n\sigma_{j}^{3}\leqslant\beta_{j},j=1,\ldots,n. Hence τn⩽εn\tau_{n}\leqslant\varepsilon_{n}. Now (12) follows from (31) and Lemma 2.

Further we assume without loss of generality that σ=1\sigma=1. Denote fj(t):=\EeitXj, fj∗(t):=\EeitXj∗,  j=1,…,n,f_{j}(t):=\E e^{itX_{j}},\,f_{j}^{*}(t):=\E e^{itX_{j}^{*}},\,\,j=1,\ldots,n, and set W:=SnW:=S_{n} in (32). Using Lemma 4 and Corollary 1 we get

Substituting the latter into (32), we arrive at (13).

It follows from (30) that ∣fj(s)∣⩽exp⁡(−σj2s2/2+2βja∣s∣3)|f_{j}(s)|\leqslant\exp(-{\sigma_{j}^{2}s^{2}/2}+2\beta_{j}a|s|^{3}) for all real ss. As a result,

Since σj3⩽βj⩽εn\sigma_{j}^{3}\leqslant\beta_{j}\leqslant\varepsilon_{n} for j=1,…,nj=1,\ldots,n, we have for such jj

The function exp⁡(s22−2as3)\exp\left(\frac{s^{2}}{2}-2as^{3}\right) increases on the segment [0,1/6a][0,1/6a] and at the point 1/6a1/6a it attains its global maximum equal to 1/l1/l. Therefore,

Combining (37), (38) and (39) gives for sεn1/3⩽1/6as\varepsilon_{n}^{1/3}\leqslant 1/6a

Substituting the expressions obtained into (32), we get the required estimates.

Proof of Theorem 6. At first we prove (15). Denote f1(t):=\EeitX1f_{1}(t):=\E e^{itX_{1}}. According to Lemma 1,

Now (15) follows from the fact that fSn(t)=f1n(t/n)f_{S_{n}}(t)=f_{1}^{n}(t/\sqrt{n}).

To establish (17) we note that 1+x⩽ex1+x\leqslant e^{x} for all real xx. Applying this inequality to (15) gives

It remains to note that the sequence (τm)m⩾1(\tau_{m})_{m\geqslant 1} decreases, which leads to (17).

We set W:=SnW:=S_{n} in Lemma 3. Applying (35) to the r.v. 1nX1,…,1nXn\frac{1}{\sqrt{n}}X_{1},\ldots,\frac{1}{\sqrt{n}}X_{n} yields

The first factor can be estimated with the help of (40) and the second – by means of (36). We have

Substituting the expression obtained into (32), we get (16). To establish (18) we apply the inequality 1+x⩽ex{1+x\leqslant e^{x}} to the first factor on the right-hand side of (42) and note that (τm)m⩾1(\tau_{m})_{m\geqslant 1} is decreasing. □\enskip\square

Proof of Theorem 7. Let D(ε,n)D(\varepsilon,n) denote the least quantity such that for every collection consisting of nn r.v. X1,…,XnX_{1},\ldots,X_{n} with εn=ε\varepsilon_{n}=\varepsilon holds the inequality

Then the constant CC can be determined as

Hence, it suffices to show that for all possible values of ε\varepsilon and nn the quantity D(ε,n)⩽0.5606{D(\varepsilon,n)\leqslant 0{.}5606} (and in the case of i.i.d. r.v. D(ε,n)⩽0.4785D(\varepsilon,n)\leqslant 0{.}4785). For ε⩾1/0.5606\varepsilon\geqslant 1/0{.}5606 (respectively, ε⩾1/0.4785{\varepsilon\geqslant 1/0{.}4785}) the latter is obvious, since ρ(Sn,N)⩽1\rho(S_{n},N)\leqslant 1.

Moreover, denote λn:=σ2(n)/(σ2(n)−max⁡k=1,…,nσk2)\lambda_{n}:=\sigma^{2}(n)/(\sigma^{2}(n)-\max\limits_{k=1,\ldots,n}\sigma_{k}^{2}) and set

Then, according to the inequality (I.52) from , for ε^n+εn′⩽0.2\widehat{\varepsilon}_{n}+\varepsilon^{\prime}_{n}\leqslant 0{.}2 we have

Assume without loss of generality that σ=1\sigma=1. Then the Lyapunov inequality yields σj3⩽βk⩽∑k=1nβk=εn, j=1,…,n\sigma_{j}^{3}\leqslant\beta_{k}\leqslant\sum_{k=1}^{n}\beta_{k}=\varepsilon_{n},\,j=1,\ldots,n. Hence λn⩽(1−εn2/3)−1\lambda_{n}\leqslant(1-\varepsilon_{n}^{2/3})^{-1}. In addition, εn′⩽ε^n\varepsilon^{\prime}_{n}\leqslant\widehat{\varepsilon}_{n} and, as it was shown in , εn′′⩽(εn′)4/3\varepsilon^{\prime\prime}_{n}\leqslant(\varepsilon^{\prime}_{n})^{4/3}. From these inequalities and (43) it follows easily that D(ε)⩽0.5606D(\varepsilon)\leqslant 0{.}5606 when ε⩽0.02\varepsilon\leqslant 0{.}02.

In the case of i.i.d. summands εn=β1/(σ13n)⩾1/n\varepsilon_{n}=\beta_{1}/(\sigma_{1}^{3}\sqrt{n})\geqslant 1/\sqrt{n}. Thus, n⩾⌈1/εn2⌉n\geqslant\lceil 1/\varepsilon_{n}^{2}\rceil and

where n0(ε):=⌈1/ε2⌉.n_{0}(\varepsilon):=\lceil 1/\varepsilon^{2}\rceil. Moreover,

Combining (43), (44) and (45) yields D(ε)⩽0.4785D(\varepsilon)\leqslant 0{.}4785 when ε⩽0.037\varepsilon\leqslant 0{.}037. Therefore, in the general case we have to consider ε\varepsilon from the segment I1=[0.02;1/0.5606]I_{1}=\left[0{.}02;{1}/{0{.}5606}\right] and in the case of i.i.d. r.v. – from I2=[0.037;1/0.4785]I_{2}=\left[0{.}037;{1}/{0{.}4785}\right]. The proof for these values of ε\varepsilon is based on an inequality due to Prawitz

where K(u):=12(1−∣u∣)+i2((1−∣u∣)cot⁡(πu)+sgn(u)π),0<U0⩽U\enskip K(u):=\frac{1}{2}(1-|u|)+\frac{i}{2}\left((1-|u|)\cot(\pi u)+\frac{\textrm{sgn}(u)}{\pi}\right),\quad 0<U_{0}\leqslant U.

It follows from (46) that D(ε)D(\varepsilon) does not exceed the quantity D∗(ε,U0,U)D^{*}(\varepsilon,U_{0},U), which arises on the right-hand side of (46) when we substitute δn(t)\delta_{n}(t) with its estimate min⁡{δ^1(ε,t),δ^2(ε,t)}\min\{\widehat{\delta}_{1}(\varepsilon,t),\widehat{\delta}_{2}(\varepsilon,t)\}, ∣fn(t)∣|f_{n}(t)| – with the estimate f^1(ε,t)\widehat{f}_{1}(\varepsilon,t) and select such parameters U0,UU_{0},U that the resulting expression was as little as possible. This procedure was carried out with the aid of computer for several hundreds values of ε\varepsilon dispersed on the segment I1I_{1}. To obtain the estimates for the intermediate points we used the following property of the quantities D∗D^{*}, which holds due to the monotonicity of the functions f^1,…,f^3\widehat{f}_{1},\ldots,\widehat{f}_{3} and δ^1,…,δ^4\widehat{\delta}_{1},\ldots,\widehat{\delta}_{4} with respect to their first arguments.

The extremal value of the quantity D∗(ε,U0,U)=0.56054D^{*}(\varepsilon,U_{0},U)=0{.}56054 was attained for ε=0.5092,U0=2.4852,U=5.9508.{\varepsilon=0{.}5092},U_{0}=2{.}4852,U=5{.}9508.

In the case of i.i.d. r.v. the estimates were constructed in a different way.

For the fixed value of ε\varepsilon we estimated the quantities D(ε,n), n⩾1,D(\varepsilon,n),\,n\geqslant 1, separately . For n<mn<m, where mm is some natural number, the individual estimates of D(ε,n)D(\varepsilon,n) were given. On the right-hand side of (46) we substituted δn(t)\delta_{n}(t) and ∣fn(t)∣|f_{n}(t)| with their upper estimates δ^3(ε,n,t)\widehat{\delta}_{3}(\varepsilon,n,t) and f^2(ε,n,t)\widehat{f}_{2}(\varepsilon,n,t). After that the computational procedure as described above was carried out to select the optimal parameters UU and U0U_{0}. For n⩾mn\geqslant m the quantities D(ε,n)D(\varepsilon,n) were estimated uniformly. On the right-hand side of (46) the estimates δ^4(ε,m,t)\widehat{\delta}_{4}(\varepsilon,m,t) and f^3(ε,m,t)\widehat{f}_{3}(\varepsilon,m,t) were used. As before, it was sufficient to carry out the calculations only for the finite number of points, since a property similar to (47) holds in this case as well. For the i.i.d. r.v. the extremal value 0.478490{.}47849 was attained for ε=0.3536,n=8\varepsilon=0{.}3536,n=8, U0=2.6157,U=8.9115.U_{0}=2.6157,U=8.9115.

Thus, the constant CC does not exceed 0.56060{.}5606. And if we restrict to the case of i.i.d. r.v., we have C⩽0.4785C\leqslant 0{.}4785.

Acknowledgement

The author would like to thank Professor A. V. Bulinski for useful discussions and valuable advice.

References