New estimates of the convergence rate in the Lyapunov theorem
Ilya Tyurin
Introduction
Consider centered independent (real-valued) random variables (r.v.) with variances and finite absolute moments . We denote
According to the Lyapunov theorem, converges in distribution to the standard normal r.v. when . From both theoretical and practical points of view, it is very important to estimate the convergence rate in this theorem. It is known that there exists a minimal numerical constant such that for the Kolmogorov distance between and the standard normal variable holds the inequality
There are plenty of works devoted to estimation of this constant. Esseen showed that . Bergström obtained the bound . Takano established that in the case of independent identically distributed (i.i.d.) summands . Zolotarev obtained a new inequality allowing to estimate the proximity of two sums of independent r.v. With the help of this inequality he showed successively that and , while in the case of i.i.d. variables and . The proposed method was further developed in the works of van Beek and Shiganov , who proved the estimates and , respectively. For the sums of identically distributed r.v. Shiganov obtained the bound , which was sharpened in 2006 by Shevtsova . She showed that in this case . In we derived the estimates in the general case and for identically distributed summands. In the present paper we improve them.
From a private communication with Korolev and Shevtsova we know that recently they have established the convergence rate in the central limit theorem in a variety of sences . In these works only the i.i.d. case was considered and the bound for the constant is not as sharp as ours. However, interesting estimates of the other kind were obtained.
It is worth mentioning the related problem of determining the asymptotically best constants in Lyapunov’s theorem. As it was shown by Esseen , if all the r.v. have the same distribution, then
and the constant on the right-hand side of this inequality cannot be lowered (hence the lower bound ). This result was elaborated by Rogozin , who established that under the same assumptions
where and is the set of all normal r.v.
Chistyakov generalized (2) and (3) to the case of nonidentically distributed summands. He proved that
where are when .
There are also the estimates of the convergence rate in the Lyapunov theorem provided that the moments of the order exist (see ).
Analogues of (1) are known for other probability metrics as well, for example, (where ). The latter will be described in detail in section 2. Estimates in terms of these metrics can be obtained in a natural way using Stein’s method. The proof of the estimates mentioned above uses, in particular, the so-called zero bias transformation of a probability distribution (see ).
For the distance in terms of metrics the following estimates (see ) are known:
Hoeffding considered the problem of finding the least upper bound of over the set of all collections of independent simple r.v. satisfying restrictions of the form . More precisely, it was established that in this case one has to consider only r.v. taking at most values. In the present work results of are generalized to the case of arbitrary quasiconvex functional defined on the set of all probability distributions.
The results obtained allowed us to derive an unimprovable estimate for the proximity in the mean metric between a probability distribution and its zero bias transformation. The latter was used to estimate the accuracy of the Gaussian approximation for the sums of independent variates. It was established that the values of constants in (4) can be taken 3 times lower. In addition, our estimate for the metric is optimal. Furthermore, new estimates for the difference between the characteristic functions of the normalized sum and the standard normal r.v. were derived, which allowed us to prove that and in the case of i.i.d. summands .
Main notions and results
It is easy to see that forms a linear space. And the set D of discrete probability distributions that are concentrated on finite sets of points is a convex subset of . The latter means that for arbitrary and .
Consider the set of all collections consisting of independent r.v. . Then
We assume that on some real-valued functions are defined. Consider the set
In this expression, we assume that the supremum over the empty set is zero.
Theorem 2. Let be a nonnegative function on , – a linear space with the norm , – such a mapping that
for arbitrary . Then the least value of such that the inequality
holds for every measure , coincides with the least value of such that is true for every measure .
Let be a zero-mean r.v. with variance . A r.v. is said to have the -zero biased distribution if
where is the set of all real bounded functions with .
Note that has alternative representations. These are the so-called mean metric
Theorem 3. If is a centered r.v. with unit variance and finite third absolute moment, then
with equality when has a 2-point distribution.
Corollary 1. Consider a r.v. having the -zero biased distribution. Then
Theorem 4. The following inequalities are true:
The latter double inequality is optimal, namely, for every there exists such a sequence of i.i.d. r.v. , that
and is the point where this maximum is attained, .
The quantities and can be expressed in terms of the so-called Dawson integral
which can be computed by the means of several efficient numerical procedures. For example, such a function is available in the GNU Scientific Library (GSL). It is easy to check that
These representations are of great importance, since they allow to reduce significantly the amount of numerical calculations required for the proof of Theorem 7.
In the case of i.i.d. variables the estimates can be slightly improved. Denote , and let be centered i.i.d. r.v. with unit variances and finite third absolute moments . Then and .
Estimates (12)-(18) allowed to establish the following result.
Theorem 7. The constant in inequality does not exceed , and in the case of identically distributed summands .
Proofs
Proof of Theorem 1. If , then , and the statement of our theorem is true. Further we suppose that the set is nonempty.
The sequence of the sets increases to the set . Therefore,
Let’s take an arbitrary measure , where , and show that there exists such that .
Let be concentrated in points and
The vector defines a probability distribution, so
Moreover, the conditions hold and therefore
Vice versa, an arbitrary vector with nonnegative coordinates satisfying the system of linear equations (20) and (21) defines according to (19) an element of the set , and if one of its coordinates equals zero – an element of . We have equations and at least unknowns, so there exists a nonzero solution of the corresponding homogeneous system. Since the sum of the coordinates of this vector is equal to zero, but the vector itself is nonzero, it follows that has both positive and negative coordinates. Therefore, there exist the least and the least such that one of the coordinates of the vector equals zero and some coordinate of is equal to zero. If , then . Otherwise,
and because of the quasiconvexity , where are the distributions defined by and . Thus, or . But and
Proof of Theorem 2. According to Theorem 1, it is sufficient to prove that for every fixed value of the function
is quasiconvex. Let . By the properties of the norm
Proof of Theorem 3. We begin by showing that without loss of generality we can consider simple r.v. . It is sufficient to establish that for every r.v. satisfying the conditions of the theorem there exists a sequence of simple r.v. with zero means and unit variances such that
It is easy to see that converges to in as well. Therefore, the second condition in (22) is obviously satisfied. It remains to show that the first one also holds.
From the triangle inequality for the metric one can easily derive that
The first summand on the right-hand side of (23) tends to zero, since
Let’s evaluate the second summand. For the function we set . Then
The difference of expectations on the left-hand side of (24) does not change if we replace the function by . Therefore, we can assume without loss of generality that . Then yields and so . According to the finite-increment theorem,
where is a number between and . Moreover,
And finally, the Hölder’s inequality yields
Obviously, , since converges to in . Thus, the second summand in (23) tends to zero and (22) is fulfilled.
So, it is sufficient to consider simple r.v. Let be a function that maps the distribution of a r.v. to its zero-biased distribution . Moreover, consider a linear operator that maps a signed measure to its cumulative distribution function (c.d.f.) . It is easy to see that
hence the mapping satisfies (5). If we set , and apply Theorem 2 to , , – the normed space of integrable functions on the real line with the norm
then the problem reduces to the case of simple r.v. taking at most 3 values. C.d.f. of a simple r.v. is a staircase function. Using formula (8) one can easily obtain the c.d.f. of . Therefore, it is not difficult to find the explicit expression for .
Let take exactly two values and with probabilities and . Then its c.d.f. is piecewise constant and has two steps in points and that are equal to and , respectively. Since is centered, we have , which together with (8) yields that is uniformly distributed on . Therefore, on its c.d.f. is linear, and its graph is a segment that connects and . By definition equals the area of the figure bounded by distribution functions of these r.v. (in the case considered it is a union of two triangles, see pic. 1, left).
It follows from the conditions and that Hence
Let’s find the area of the figure bounded by c.d.f. of r.v. and . Density of is . Thus, the slope of the c.d.f. of this r.v. on [-x,y] is . The length of the vertical leg of the first triangle is , and that of the second one is . Hence, the total area of both triangles is
Therefore, if takes exactly two values, there is equality in (9).
Consider the case when takes three values. We assume without loss of generality that two of them () do not exceed zero and one () is positive. As before, the c.d.f. of is piecewise constant and the c.d.f. of is piecewise linear. However, the form of the figure bounded by them is more complicated (see pic. 1, right). Denote by the value of the c.d.f. of at and by its value at . Let take values with probabilities , respectively. Then, because of the moment-type restrictions,
It is a system of linear equations with respect to . Using Cramer’s rule, we obtain
where . Thus, every r.v. with zero mean and variance 1 that takes three values is uniquely determined by these three values. It is easy to see that are nonnegative iff
In other words, a r.v. taking the values exists iff (26) is satisfied. Our aim is to prove that the function
does not exceed zero. Its explicit form in terms of variables depends on how the c.d.f. of r.v. and are located with respect to each other. There are 5 cases:
I. , , or, equivalently, , . In this case
II. , , .
III. , , (the latter implies ).
IV. , , , .
V. .
Note that in each of these cases is the same function defined by (27). As a result, if the values satisfy the restrictions of two cases simultaneously, then for the function we can use the expression corresponding to any of them.
As one can see, in the cases I, II, and in the cases III, IV the function has the same representation. Therefore, further we distinguish three possibilities:
B. ,
We show that in each of the cases A, B and C the function does not exceed zero.
because of (26), nonnegativeness of and the fact that .
Since , it suffices to prove that the expression enclosed by braces does not exceed zero. Consider this expression as a function of the variable while holding the others fixed. When , this function equals . Moreover, it decreases with respect to , since the coefficients of terms and are negative. Consequently, for all positive values of it does not exceed zero.
where Assume that . Due to the condition one has
Therefore, and do not exceed zero and, consequently, decreases with respect to . Consequently, if one reduces the value of the variable while holding and fixed, will remain positive. The variable is bounded from below by two conditions:
The first of these conditions can be omitted, since it follows from the other two:
Therefore, we can reduce to the value such that . And will remain positive. But the situation, when , satisfies the restrictions of the case C, for which we established that – a contradiction.
Proof of Corollary 1. Without loss of generality assume = 1. Let be a random index taking values with probabilities , independent of . Construct on an extended probability space
where has the -zero biased distribution and is independent of , . Then (see ). Therefore, for an arbitrary function one has
Here we used the statement of Theorem 3 for r.v. as well as the property of homogeneity of the metric (i. e. ) and .
Taking into account that in the definition of the metric , one has
Estimates in terms of are obtained by applying Corollary 1 to (29).
Let’s prove the optimality of (11). We set . Then and , since the r.v. is symmetric and the function is odd. Consider – a sequence of i.i.d. variables with zero means and unit variances. Then
It only remains to prove that can be arbitrarily close to unity. According to (25), the third absolute moment of a centered r.v. with variance 1 taking two values with probabilities and , respectively, is equal to
It’s easy to see that the third moment of this r.v. equals
Obviously, when .
Proof. This can be checked directly by calculating the derivative.
Lemma 3. If is a centered r.v. with variance 1 and , then
where is the characteristic function of a r.v. having the -zero biased distribution.
Proof. According to the definition of the -zero biased distribution,
Consider the function . Note that . Taking into account (33), we have
Lemma 4. For arbitrary r.v. and we have
Thus, for arbitrary defined on one probability space such that and , we have
Passing in (34) to the greatest lower bound among every possible , we obtain
Proof of Theorem 5. The inequality (12) is a consequence of Lemma 1. Indeed, according to the Lyapunov inequality, we have . Hence . Now (12) follows from (31) and Lemma 2.
Further we assume without loss of generality that . Denote and set in (32). Using Lemma 4 and Corollary 1 we get
Substituting the latter into (32), we arrive at (13).
It follows from (30) that for all real . As a result,
Since for , we have for such
The function increases on the segment and at the point it attains its global maximum equal to . Therefore,
Combining (37), (38) and (39) gives for
Substituting the expressions obtained into (32), we get the required estimates.
Proof of Theorem 6. At first we prove (15). Denote . According to Lemma 1,
Now (15) follows from the fact that .
To establish (17) we note that for all real . Applying this inequality to (15) gives
It remains to note that the sequence decreases, which leads to (17).
We set in Lemma 3. Applying (35) to the r.v. yields
The first factor can be estimated with the help of (40) and the second – by means of (36). We have
Substituting the expression obtained into (32), we get (16). To establish (18) we apply the inequality to the first factor on the right-hand side of (42) and note that is decreasing.
Proof of Theorem 7. Let denote the least quantity such that for every collection consisting of r.v. with holds the inequality
Then the constant can be determined as
Hence, it suffices to show that for all possible values of and the quantity (and in the case of i.i.d. r.v. ). For (respectively, ) the latter is obvious, since .
Moreover, denote and set
Then, according to the inequality (I.52) from , for we have
Assume without loss of generality that . Then the Lyapunov inequality yields . Hence . In addition, and, as it was shown in , . From these inequalities and (43) it follows easily that when .
In the case of i.i.d. summands . Thus, and
where Moreover,
Combining (43), (44) and (45) yields when . Therefore, in the general case we have to consider from the segment and in the case of i.i.d. r.v. – from . The proof for these values of is based on an inequality due to Prawitz
where .
It follows from (46) that does not exceed the quantity , which arises on the right-hand side of (46) when we substitute with its estimate , – with the estimate and select such parameters that the resulting expression was as little as possible. This procedure was carried out with the aid of computer for several hundreds values of dispersed on the segment . To obtain the estimates for the intermediate points we used the following property of the quantities , which holds due to the monotonicity of the functions and with respect to their first arguments.
The extremal value of the quantity was attained for
In the case of i.i.d. r.v. the estimates were constructed in a different way.
For the fixed value of we estimated the quantities separately . For , where is some natural number, the individual estimates of were given. On the right-hand side of (46) we substituted and with their upper estimates and . After that the computational procedure as described above was carried out to select the optimal parameters and . For the quantities were estimated uniformly. On the right-hand side of (46) the estimates and were used. As before, it was sufficient to carry out the calculations only for the finite number of points, since a property similar to (47) holds in this case as well. For the i.i.d. r.v. the extremal value was attained for ,
Thus, the constant does not exceed . And if we restrict to the case of i.i.d. r.v., we have .
Acknowledgement
The author would like to thank Professor A. V. Bulinski for useful discussions and valuable advice.