Stein's method and the Laplace distribution

John Pike, Haining Ren

Background and Introduction

If W∼Laplace(0,b)W\sim\textrm{Laplace}(0,b), then its moments are given by

This distribution was introduced by P.S. Laplace in 1774, four years prior to his proposal of the “second law of errors,” now known as the normal distribution. Though nowhere near as ubiquitous as its younger sibling, the Laplace distribution appears in numerous applications, including image and speech compression, options pricing, and modeling sizes of sand particles, diamonds, and beans. For more properties and applications of the Laplace distribution, the reader is referred to the text .

Our interest in the Laplace distribution was sparked by the fact that if X1,X2,...X_{1},X_{2},... is a sequence of random variables (satisfying certain technical assumptions) and Np∼Geometric(p)N_{p}\sim\textrm{Geometric}(p) is independent of the XiX_{i}’s, then the sum p12∑i=1NpXip^{\frac{1}{2}}\sum_{i=1}^{N_{p}}X_{i} converges weakly to the Laplace distribution as p→0p\rightarrow 0 . Such geometric sums arise in a variety of settings , and the general setup (distributional convergence of sums of random variables) is exactly the type of problem for which one expects Stein’s method computations to yield useful results. Indeed, Erol Peköz and Adrian Röllin have applied Stein’s method arguments to generalize a theorem due to Rényi concerning the convergence of sums of a random number of positive random variables to the exponential distribution . By an analogous line of reasoning, we are able to carry out a similar program for convergence of random sums of certain mean zero random variables to the Laplace distribution.

We begin in Section 2 by introducing a Stein operator which we show completely characterizes the mean zero Laplace distribution. Specifically, we prove

Let W∼Laplace(0,b)W\sim\textnormal{Laplace}(0,b) and define the operator A\mathcal{A} by

Finally, in Section 4 we apply these tools to the study of random sums of mean zero random variables. As a special case, we show

Characterizing the Laplace Distribution

Our first order of business is to establish a characterizing operator for the Laplace distribution. As is typical in Stein’s method constructions, we split the proof of Theorem 1.1 into two parts. We begin with

Applying Fubini’s theorem twice shows that

Setting h(y)=g(−y)h(y)=g(-y), it follows from the previous calculation that

Note that since the density of a Laplace(0,b)\textrm{Laplace}(0,b) random variable is given by fW(w)=12be−∣w∣bf_{W}(w)=\frac{1}{2b}e^{-\frac{\left|w\right|}{b}}, the density method suggests the following characterizing equation for the Laplace distribution:

and indeed one can verify that if W∼Laplace(0,b)W\sim\textrm{Laplace}(0,b), then

for all absolutely continuous gg for which these expectations exist. Thus if g′g^{\prime} is such a function as well, setting G(w)=sgn(w)(g(w)−g(0))G(w)=\textrm{sgn}(w)\left(g(w)-g(0)\right) gives

so the general form of the equation in Theorem 1.1 can be ascertained by iterating the density method.

Now, in order to establish the second part of Theorem 1.1, we will show that any XX satisfying the hypotheses has

where dBLd_{BL} denotes the bounded Lipschitz distance given by

In keeping with the general strategy laid out in the introduction, we consider the initial value problem

For h∈HBLh\in\mathcal{H}_{BL}, W∼Laplace(0,b)W\sim\textnormal{Laplace}(0,b), h~(x)=h(x)−Wh,\widetilde{h}(x)=h(x)-Wh, a bounded, twice-differentiable solution to the initial value problem

This solution satisfies ∥gh∥∞≤2\left\|g_{h}\right\|_{\infty}\leq 2, ∥gh′∥∞≤2b\left\|g_{h}^{\prime}\right\|_{\infty}\leq\frac{2}{b}, ∥gh′′∥∞≤4b2\left\|g_{h}^{\prime\prime}\right\|_{\infty}\leq\frac{4}{b^{2}} and ∥gh′′′∥∞≤b+2b3.\left\|g_{h}^{\prime\prime\prime}\right\|_{\infty}\leq\frac{b+2}{b^{3}}.

The general solution to the homogeneous equation g′′(x)−b−2g(x)=0g^{\prime\prime}(x)-b^{-2}g(x)=0 is given by g0(x)=C1exb+C2e−xbg_{0}(x)=C_{1}e^{\frac{x}{b}}+C_{2}e^{-\frac{x}{b}}, so, since the associated Wronskian is nonzero, the variation of parameters method suggests that a solution to the inhomogeneous equation g′′(x)−b−2g(x)=−b−2h~(x)g^{\prime\prime}(x)-b^{-2}g(x)=-b^{-2}\widetilde{h}(x) is given by

To see that the initial condition is satisfied, we observe that

Moreover, since ∥h∥∞≤1\left\|h\right\|_{\infty}\leq 1,

With the preceding result in hand, we can finish of the proof of Theorem 1.1 via

for every twice-differentiable function gg with ∥g∥∞,∥g′∥∞,∥g′′∥∞<∞\left\|g\right\|_{\infty},\left\|g^{\prime}\right\|_{\infty},\left\|g^{\prime\prime}\right\|_{\infty}<\infty, then X∼Laplace(0,b)X\sim\textnormal{Laplace}(0,b).

Let W∼Laplace(0,b)W\sim\textrm{Laplace}(0,b) and, for h∈HBLh\in\mathcal{H}_{BL}, let ghg_{h} be as in Lemma 2.2. Because gh(0)=0g_{h}(0)=0 and gh,gh′,gh′′g_{h},g_{h}^{\prime},g_{h}^{\prime\prime} are bounded, it follows from the above assumptions that

Taking the supremum over h∈HBLh\in\mathcal{H}_{BL} shows that dBL(L(X),L(W))=0d_{BL}\left(\mathscr{L}(X),\mathscr{L}(W)\right)=0. ∎

Before moving on, we observe that the reason we are working with the bounded Lipschitz distance is that the bounds on ghg_{h} and its derivatives depended on both hh and h′h^{\prime} having finite sup norm. As dBLd_{BL} is not especially common (at least explicitly) in the Stein’s method literature, we conclude this section with a proposition relating it to the more familiar Kolmogorov distance

If ZZ is an absolutely continuous random variable whose density, fZf_{Z}, is uniformly bounded by a constant C<∞C<\infty, then for any random variable XX,

We first note that the inequality holds trivially if dBL(L(X),L(Z))=0d_{BL}\left(\mathscr{L}(X),\mathscr{L}(Z)\right)=0 as dBLd_{BL} and dKd_{K} are metrics. Also, since dK(P,Q)≤1d_{K}(P,Q)\leq 1 for all probability measures PP and QQ, C+22≥1\frac{C+2}{2}\geq 1, and dBL(L(X),L(Z))≥1d_{BL}\left(\mathscr{L}(X),\mathscr{L}(Z)\right)\geq 1 implies dBL(L(X),L(Z))≥1\sqrt{d_{BL}\left(\mathscr{L}(X),\mathscr{L}(Z)\right)}\geq 1, we have

whenever dBL(L(X),L(Z))≥1d_{BL}\left(\mathscr{L}(X),\mathscr{L}(Z)\right)\geq 1. Thus it suffices to consider the case where dBL(L(X),L(Z))∈(0,1)d_{BL}\left(\mathscr{L}(X),\mathscr{L}(Z)\right)\in(0,1).

Since dBL(L(X),L(Z))∈(0,1)d_{BL}\left(\mathscr{L}(X),\mathscr{L}(Z)\right)\in(0,1), if we take ε=dBL(L(X),L(Z))∈(0,1)\varepsilon=\sqrt{d_{BL}\left(\mathscr{L}(X),\mathscr{L}(Z)\right)}\in(0,1), then εhx,ε∈HBL\varepsilon h_{x,\varepsilon}\in\mathcal{H}_{BL} and thus

When C≥1C\geq 1, we can take ε=1CdBL(L(X),L(Z))\varepsilon=\sqrt{\frac{1}{C}d_{BL}\left(\mathscr{L}(X),\mathscr{L}(Z)\right)} in the above argument to obtain an improved bound of

To the best of the authors’ knowledge, the above proposition is original, though the proof follows the same basic line of reasoning as the well-known bound on the Kolmogorov distance by the Wasserstein distance (see Proposition 1.2 in ). It seems that the primary reason for using the Wasserstein metric, dWd_{W}, is that it enables one to work with smoother test functions while still implying convergence in the more natural Kolmogorov distance. Proposition 2.4 shows that dBLd_{BL} also upper-bounds dKd_{K} while enjoying all of the resulting smoothness of Wasserstein test functions and with additional boundedness properties to boot. Moreover, the Wasserstein distance is not always well-defined (e.g. if one of the distributions does not have a first moment), whereas dBLd_{BL} always exists. Finally, dBLd_{BL} is a fairly natural measure of distance since it metrizes weak convergence . However, we always have dW(L(X),L(Z))≥dBL(L(X),L(Z))d_{W}\left(\mathscr{L}(X),\mathscr{L}(Z)\right)\geq d_{BL}\left(\mathscr{L}(X),\mathscr{L}(Z)\right), and it is possible for a sequence to converge in dBLd_{BL} but not in dWd_{W} or dKd_{K}. Furthermore, the bounded Lipschitz metric does not scale as nicely as the Kolmogorov or Wasserstein distances when the associated random variables are multiplied by a positive constant. For the remainder of this paper, we will state our results in terms of dBLd_{BL} with the corresponding Kolmogorov bound being implicit therein, though one should note that, as with the Wasserstein bound on dKd_{K}, Kolmogorov bounds obtained in this fashion are not necessarily optimal, often giving the root of the true rate.

The Centered Equilibrium Transformation

Our next task is to use the characterization in Theorem 1.1 to obtain bounds on the error terms resulting from approximation by the Laplace distribution. To this end, we introduce the following definition.

For any nondegenerate random variable XX with mean zero and finite variance, we say that the random variable XLX^{L} has the centered equilibrium distribution with respect to XX if

for all twice-differentiable functions ff such that ff, f′f^{\prime}, and f′′f^{\prime\prime} are bounded. We call the map X↦XLX\mapsto X^{L} the centered equilibrium transformation.

Suppose that XX has mean zero and variance σ2∈(0,∞)\sigma^{2}\in(0,\infty). Let XzX^{z} have the zero bias distribution with respect to XX and let B∼Beta(2,1)B\sim\textnormal{Beta}(2,1) be independent of XzX^{z}. Then XL:=BXzX^{L}:=BX^{z} satisfies

for all twice-differentiable ff with ∥f∥∞,∥f′∥∞,∥f′′∥∞<∞\left\|f\right\|_{\infty},\left\|f^{\prime}\right\|_{\infty},\left\|f^{\prime\prime}\right\|_{\infty}<\infty.

Applying the fundamental theorem of calculus, Fubini’s theorem, the definition of XzX^{z}, and the fact that BB has density p(x)=2x1(x)p(x)=2x1_{}(x) gives

The assumptions ensure that all of the functions are integrable. ∎

In an earlier version of this paper, we established the existence of the centered equilibrium distribution by showing that for certain random variables XX, XLX^{L} can be obtained by iterating the X−PX-P bias transformation from with P(x)=sgn(x)P(x)=\textrm{sgn}(x). Though there may be some merit to such a strategy and it provides another example of how results for higher order Stein operators may be obtained by iterating more traditional techniques, in our case it required the rather artificial assumption that the variates in the domain of the transformation have median zero. Those interested in the iterated X−PX-P bias approach are referred to the article by Christian Döbler, which contains the essential technical details of our original argument.

Lemma 2.3 shows that, up to scaling, the mean zero Laplace distribution is the unique fixed point of the centered equilibrium transformation. Thus one expects that if a random variable is close to its centered equilibrium transform, then its distribution is close to the Laplace. Theorem 1.2 formalizes this intuition.

For the example in Section 4, we will also need the following complementary result.

Of course, neither ff, f′f^{\prime}, nor f′′f^{\prime\prime} is bounded, so we must proceed by approximation. To this end, define

Fatou’s lemma shows that f′′(YL)f^{\prime\prime}(Y^{L}) is integrable since

so another application of dominated convergence gives

Random Sums

Suppose that X1,X2,...X_{1},X_{2},... are i.i.d., symmetric, and nondegenerate random variables with finite variance σ2\sigma^{2}, and let Np∼Geometric(p)N_{p}\sim\textrm{Geometric}(p) be independent of the Xi′sX_{i}^{\prime}s. If

then there exists γ>0\gamma>0 such that ap=p12γ+o(p12)a_{p}=p^{\frac{1}{2}}\gamma+o(p^{\frac{1}{2}}) and XX has the Laplace distribution with mean and variance σ2γ2\sigma^{2}\gamma^{2}.

A recent theorem due to Alexis Toda gives the following Lindeberg-type conditions for the existence of the distributional limit in Theorem 4.1.

Then as p→0p\rightarrow 0, the sum p12∑i=1NpXip^{\frac{1}{2}}\sum_{i=1}^{N_{p}}X_{i} converges weakly to the Laplace distribution with mean and variance σ2\sigma^{2}.

The original statement of Toda’s theorem is slightly more general, allowing for convergence to a possibly asymmetric Laplace distribution.

Now, taking V=μ−12∑i=1NXiV=\mu^{-\frac{1}{2}}\sum_{i=1}^{N}X_{i}, we claim that VL=μ−12(∑i=1M−1Xi+XML)V^{L}=\mu^{-\frac{1}{2}}\left(\sum_{i=1}^{M-1}X_{i}+X_{M}^{L}\right) has the centered equilibrium distribution with respect to VV. (Throughout, XmLX_{m}^{L} is taken to be independent of MM, NN, and XkX_{k} for k≠mk\neq m.) To see that this is so, let ff be any function satisfying the assumptions in Definition 3.1. Then, using the notation

letting νm\nu_{m} denote the distribution of

Finally, since the Xi′sX_{i}^{\prime}s are independent with mean zero, the Cauchy-Schwarz inequality gives

We conclude our discussion with a proof of Theorem 1.3, which gives sufficient conditions for weak convergence in the setting of Theorem 4.1. Though it requires that the Xi′sX_{i}^{\prime}s have uniformly bounded third absolute moments, the condition of symmetry is dropped and the identical distribution assumption is reduced to the requirement that the Xi′sX_{i}^{\prime}s have common variance. This result is not quite as general as Theorem 4.2, but it does provide bounds on the error terms.

In the language of Theorem 4.3, we have μ=1p\mu=\frac{1}{p}, σ=b2μ\sigma=b\sqrt{2\mu}, and

Acknowledgements

The authors would like to thank Larry Goldstein for suggesting the use of Stein’s method to get convergence rates in the general setting of Theorem 4.2 and for many helpful comments throughout the preparation of this paper. Thanks also to Alex Rozinov for several illuminating conversations concerning the material in Section 4 and to the anonymous referees whose careful notes were of great help to us.

References