Stein's method, logarithmic Sobolev and transport inequalities

Michel Ledoux, Ivan Nourdin, Giovanni Peccati

Introduction

is the relative entropy of dν=hdγd\nu=hd\gamma with respect to γ\gamma and

where Hess(φ){\rm Hess}(\varphi) stands for the Hessian of φ\varphi, whereas ⟨⋅,⋅⟩HS{\langle\cdot,\cdot\rangle}_{\rm HS} and ∥⋅∥HS{\|\cdot\|}_{\rm HS} denote the usual Hilbert-Schmidt scalar product and norm, respectively. Note that while Stein kernels appear implicitly in the literature about Stein’s method (see the original monograph [Ste, Lecture VI] of C. Stein, as well as [C2, C3, G-S1, G-S2]…), they gained momentum in recent years, specially in connection with probabilistic approximations involving random variables living on a Gaussian (Wiener) space (see the recent monograph [N-P2] for an overview of this emerging area). The terminology ‘kernel’ with respect to ‘factor’ seems the most appropriate to avoid confusion with related but different existing notions.

becomes relevant as a measure of the proximity of ν\nu and γ\gamma. This quantity is actually at the root of the Stein method [C-G-S, N-P2]. For example, in dimension one, the classical Stein bound expresses that the total variation distance TV(ν,γ){\rm TV}(\nu,\gamma) between a probability measure ν\nu and the standard Gaussian distribution γ\gamma is bounded from above as

justifying therefore the interest in the Stein discrepancy (see also [C-P-U]). It is actually a main challenge addressed in [N-P-S1] and this work to investigate the multidimensional setting in which inequalities such as (1.2) are no more available.

With the Stein discrepancy S(ν ∣ γ){\rm S}(\nu\,|\,\gamma), we emphasize here the inequality, for every probability dν=hdγd\nu=hd\gamma,

as a new improved form of the logarithmic Sobolev inequality (1.1). In addition, this inequality (1.3) transforms bounds on the Stein discrepancy into entropic bounds, hence allowing for entropic approximations (under finiteness of the Fisher information). Indeed as is classical, the relative entropy H(ν ∣ γ){\rm H}(\nu\,|\,\gamma) is another measure of the proximity between two probabilities ν\nu and γ\gamma (note that H(ν ∣ γ)≥0{\rm H}(\nu\,|\,\gamma)\geq 0 and H(ν ∣ γ)=0{\rm H}(\nu\,|\,\gamma)=0 if and only if ν=γ\nu=\gamma), which is moreover stronger than the total variation distance by the Pinsker-Csizsár-Kullback inequality

The proof of (1.3) is achieved by the classical interpolation scheme along the Ornstein-Uhlenbeck semigroup (Pt)t≥0{(P_{t})}_{t\geq 0} towards the logarithmic Sobolev inequality, but modified for time tt away from by a further integration by parts involving the Stein kernel. Indeed, while the exponential decay Iγ(Pth)≤e−2t Iγ(h){\rm I}_{\gamma}(P_{t}h)\leq e^{-2t}\,{\rm I}_{\gamma}(h) of the Fisher information classically produces the logarithmic Sobolev inequality (1.1), the argument is supplemented by a different control of Iγ(Pth){\rm I}_{\gamma}(P_{t}h) by the Stein discrepancy for t>0t>0.

We call the inequality (1.3) HSI, connecting entropy H, Stein discrepancy S and Fisher information I, by analogy with the celebrated Otto-Villani HWI inequality [O-V] relating entropy H, (quadratic) Wasserstein distance W (W2{\rm W}_{2}) and Fisher information I. We actually provide in Section 3 a comparison between the HWI and HSI inequalities (suggesting even an HWSI inequality). Moreover, based on the approach developed in [O-V], we prove that

an inequality that improves upon the celebrated Talagrand quadratic transportation cost inequality [T]

(since arccos⁡(e−r)≤2r\arccos(e^{-r})\leq\sqrt{2r} for every r≥0r\geq 0). We shall refer to (1.4) as the ‘WSH inequality’. Note also that W2(ν,γ) ≤ S(ν ∣ γ){\rm W}_{2}(\nu,\gamma)\,\leq\,{\rm S}(\nu\,|\,\gamma) so that, as entropy, the Stein discrepancy is a stronger measurement than the Wasserstein metric W2{\rm W}_{2}.

The new HSI inequality put forward in this work has a number of significant applications to exponential convergence to equilibrium and concentration inequalities. For example, the standard exponential decay of entropy H(νt ∣ γ)≤e−2t H(ν0 ∣ γ){\rm H}(\nu^{t}\,|\,\gamma)\leq e^{-2t}\,{\rm H}(\nu^{0}\,|\,\gamma) along the flow dνt=Pthdγd\nu^{t}=P_{t}hd\gamma, t≥0t\geq 0 (ν0=ν\nu^{0}=\nu, ν∞=γ\nu^{\infty}=\gamma), which characterizes the logarithmic Sobolev inequality (1.1) may be strengthened under finiteness of the Stein discrepancy S=S(ν ∣ γ)=S(ν0 ∣ γ){\rm S}={\rm S}(\nu\,|\,\gamma)={\rm S}(\nu^{0}\,|\,\gamma) into

for all 0≤r≤rn0\leq r\leq r_{n} where rn→∞r_{n}\to\infty according to the growth of Sp(ν ∣ γ){\rm S}_{p}(\nu\,|\,\gamma) as p→∞p\to\infty.

As alluded to above, the HSI inequality (1.3) is designed to yield entropic central limit theorems for sequences of probability measures of the form dνn=hndγd\nu_{n}=h_{n}d\gamma, n≥1n\geq 1, such that sn=S(νn ∣ γ)→0s_{n}={\rm S}(\nu_{n}\,|\,\gamma)\to 0 and

of the relative entropy of the distribution νF\nu_{F} of a vector F=(F1,…,Fd)F=(F_{1},\ldots,F_{d}) on (E,μ,Γ)(E,\mu,\Gamma) with respect to γ\gamma by the Stein discrepancy S(νF ∣ γ){\rm S}(\nu_{F}\,|\,\gamma), where CF,C~F>0C_{F},\widetilde{C}_{F}>0 depend on integrability properties of FF, the carré du champ operators Γ(Fi,Fj)\Gamma(F_{i},F_{j}), i,j=1,…,di,j=1,\ldots,d, and the inverse of the determinant of the matrix (Γ(Fi,Fj))1≤i,j≤d{(\Gamma(F_{i},F_{j}))}_{1\leq i,j\leq d}. In particular, H(νF ∣ γ)→ 0{\rm H}(\nu_{F}\,|\,\gamma)\to~{}0 as S(νF ∣ γ)→0{\rm S}(\nu_{F}\,|\,\gamma)\to 0 providing therefore entropic convergence under the Stein discrepancy. The general results obtained here cover not only normal approximation but also gamma approximation.

The inequality (1.7) thus transfers bounds on the Stein discrepancy to entropic bounds. The issue of controlling the Stein discrepancy S(νF ∣ γ){\rm S}(\nu_{F}\,|\,\gamma) itself (in terms of moment conditions for example) is not addressed here, and has been the subject of numerous recent studies around the so-called Nualart-Peccati fourth moment theorem (cf. [N-P2]). This investigation is in particular well adapted to functionals F=(F1,…,Fd)F=(F_{1},\ldots,F_{d}) whose coordinates are eigenfunctions of the underlying Markov generator. See [A-C-P, A-M-P, L3] for several results in this direction and [N-P2, Chapters 5-6] for a detailed discussion of estimates on S(νF ∣ γ){\rm S}(\nu_{F}\,|\,\gamma) that are available for random vectors FF living on the Wiener space.

The structure of the paper thus consists of two main parts, the first one devoted to the new HSI and WSH inequalities, the second one to an investigation of entropic bounds via the Stein discrepancy. Section 2 is devoted to the proof and discussions of the HSI inequality in the Gaussian case, with a first sample of illustrations and applications to convergence to equilibrium and measure concentration. In Section 3, we investigate connections between the Stein discrepancy, Wasserstein distances and transportation cost inequalities, in particular the HWI inequality, and establish the WSH inequality. Extensions of the HSI inequality to more general distributions arising as invariant probability measures of second order differential operators are addressed in Section 4. The second part consists of Section 5 which develops a general methodology (in the context of Markov Triples) to reach entropic bounds on densities of families of functionals under conditions which do not necessarily involve the Fisher information.

Logarithmic Sobolev inequality and Stein discrepancy

Observe from (2.1) that, without loss of generality, one may and will assume in the sequel that τνij(x)=τνji(x)\tau_{\nu}^{ij}(x)=\tau_{\nu}^{ji}(x) ν\nu-a.e., i,j=1,…,di,j=1,\ldots,d. Also, by choosing φ=xi\varphi=x_{i}, i=1,…,di=1,\ldots,d, in (2.1) one sees that, if ν\nu admits a Stein kernel, then ν\nu is necessarily centered. Moreover, by selecting φ=xixj\varphi=x_{i}x_{j}, i,j=1,…,di,j=1,\ldots,d, and since τνij=τνji\tau^{ij}_{\nu}=\tau^{ji}_{\nu},

(and in particular ν\nu has finite second moments).

In dimension d≥2d\geq 2, a Stein kernel τν\tau_{\nu} may not be unique – see [N-P-S2, Appendix A].

It is important to notice that, in dimension d≥2d\geq 2, the definition (2.1) of Stein kernel is actually weaker than the one used in [N-P-S1, N-P-S2]. Indeed, in those references a Stein kernel τν\tau_{\nu} is required to satisfy the stronger ‘vector’ (as opposed to the trace identity (2.1)) relation

Definition (2.1) is directly inspired by the Gaussian integration by parts formula according to which

so that the proximity of τν\tau_{\nu} with the identity matrix Id{\rm Id} indicates that ν\nu should be close to γ\gamma. In particular, it should be clear that the notion of Stein kernel in the sense of (2.1) is motivated by normal approximation. Section 4 will introduce analogous definitions adapted to the target measure in the context of the generator approach to Stein’s method. Whenever a Stein kernel exists, we consider to this task the quantity, called Stein discrepancy of ν\nu with respect to γ\gamma in the introduction,

(Note that S(ν ∣ γ){\rm S}(\nu\,|\,\gamma) may be infinite if one of the τνij\tau_{\nu}^{ij}’s is not in L2(ν){\rm L}^{2}(\nu).) Whenever S(ν ∣ γ)=0{{\rm S}(\nu\,|\,\gamma)=0}, then ν=γ\nu=\gamma since τν\tau_{\nu} is the identity matrix (see e.g. [N-P2, Lemma 4.1.3]). Observe also that if CC denotes the covariance matrix of ν\nu, then

where Varν{\rm Var}_{\nu} indicates the variance under the probability measure ν\nu.

2 The Gaussian HSI inequality

The following result emphasizes the Gaussian HSI inequality connecting entropy H{\rm H}, Stein discrepancy S{\rm S} and Fisher information I{\rm I}. In the statement, we use the conventions 0log⁡(1+s0)=00\log(1+\frac{s}{0})=0 and ∞log⁡(1+s∞)=s\infty\log(1+\frac{s}{\infty})=s for every s∈[0,∞]s\in[0,\infty], and rlog⁡(1+∞r)=∞r\log(1+\frac{\infty}{r})=\infty for every r∈(0,∞)r\in(0,\infty).

from which I(ν ∣ γ)=0{\rm I}(\nu\,|\,\gamma)=0, and therefore ν=γ\nu=\gamma, which is in contrast with the assumption S(ν ∣ γ)>0{\rm S}(\nu\,|\,\gamma)>0. As a consequence, we infer that S(ν ∣ γ)=0{\rm S}(\nu\,|\,\gamma)=0, from which it follows that ν=γ\nu=\gamma.

Let γC\gamma_{C} be as above (with CC non-singular), and let dν=hdγCd\nu=hd\gamma_{C} be centered with smooth probability density hh with respect to γC\gamma_{C}. Assume that ν\nu admits a Stein kernel τν\tau_{\nu} in the sense of (2.1). Then,

where C−12C^{-\frac{1}{2}} denotes the unique symmetric non-singular matrix such that (C−12)2=C−1{(C^{-\frac{1}{2}})}^{2}=C^{-1}.

Corollary 2.3 is easily deduced from Theorem 2.2 and details are left to the reader. The argument simply uses that if MM is the unique non-singular symmetric matrix such that C=M2C=M^{2}, then H(ν ∣ γC)=H(ν0 ∣ γ){\rm H}(\nu\,|\,\gamma_{C})={\rm H}(\nu^{0}\,|\,\gamma) where dν0(x)=h(Mx)dγ(x)d\nu^{0}(x)=h(Mx)d\gamma(x).

3 Proof of the Gaussian HSI inequality

According to our conventions, if either S(ν ∣ γ){\rm S}(\nu\,|\,\gamma) or I(ν ∣ γ){\rm I}(\nu\,|\,\gamma) is infinite, then (2.6) coincides with the logarithmic Sobolev inequality (1.1). On the other hand, if S(ν ∣ γ){\rm S}(\nu\,|\,\gamma) or I(ν ∣ γ){\rm I}(\nu\,|\,\gamma) equals zero, then ν=γ\nu=\gamma, and therefore H(ν ∣ γ)=0{\rm H}(\nu\,|\,\gamma)=0. It follows that, in order to prove (2.6), we can assume without loss of generality that S(ν ∣ γ){\rm S}(\nu\,|\,\gamma) and I(ν ∣ γ){\rm I}(\nu\,|\,\gamma) are both non-zero and finite.

owing to a standard integration by parts of the Gaussian density.

The generator L\cal L is a diffusion and satisfies the integration by parts formula

(These expressions should actually be considered for h+εh+\varepsilon as ε→0\varepsilon\to 0.) Using PthP_{t}h instead of hh in the previous relations and writing vt=log⁡Pthv_{t}=\log P_{t}h, one deduces from the symmetry of PtP_{t} that

Recall finally that if dνt=Pthdγd\nu^{t}=P_{t}hd\gamma, t≥0t\geq 0 (with ν0=ν\nu^{0}=\nu and νt→γ\nu^{t}\to\gamma), the classical de Bruijn’s formula (see e.g. [B-G-L, Proposition 5.2.2]) indicates that

Theorem 2.2 will follow from the next Proposition 2.4. In this proposition, (i) corresponds to the integral version of (2.13) whereas (ii) describes the well-known exponential decay of the Fisher information along the Ornstein-Uhlenbeck semigroup. This decay actually yields the logarithmic Sobolev inequality (1.1), see [B-G-L, Section 5.7]. The new third point (iii) is a reformulation of [N-P-S1, Theorem 2.1] for which we provide a self-contained proof. It describes an alternate bound on the Fisher information along the semigroup in terms of the Stein discrepancy for values of t>0t>0 away from . It is the combination of (ii) and (iii) which will produce the HSI inequality. Point (iv)(iv) will be needed in the forthcoming proof of the WSH inequality (1.4), as well as in the proof of Proposition 3.1 providing a direct bound of the Wasserstein distance W2{\rm W}_{2} by the Stein discrepancy.

Under the above notation and assumptions, denote by τν\tau_{\nu} a Stein kernel of dν=hdγd\nu=hd\gamma. For every t>0t>0, recall dνt=Pth dγd\nu^{t}=P_{t}h\,d\gamma, and write vt=log⁡Pthv_{t}=\log P_{t}h. Then,

(Exponential decay of Fisher information) For every t≥0t\geq 0,

(Exponential decay of Stein discrepancy) For every t≥0t\geq 0,

In view of the preceding discussion, only the proofs of (iii) and (iv)(iv) need to be detailed. Throughout the various analytical arguments below, it may be assumed that the density hh is regular enough, the final conclusions being then reached by approximation arguments as e.g. in [O-V, B-G-L]. Starting with (iii), use (2.12) and the definition (2.1) of τν\tau_{\nu} to write, for any t>0t>0,

Now, for all i,j=1,…,di,j=1,\ldots,d, by (2.8) and (2.9),

which is (2.16). To deduce the estimate (2.17), it suffices to apply (twice) the Cauchy-Schwarz inequality to the right-hand side of (2.16) in such a way that, by integrating out the yy variable,

by symmetry of PtP_{t}, the proof of (2.17) is complete.

By the integral representation of PtP_{t},

Use now the definition of τν\tau_{\nu} in the xx variable and integration by parts in the yy variable to get that

As a consequence, a Stein kernel for νt\nu^{t} is

By the Cauchy-Schwarz inequality along PtP_{t},

that is the announced result (iv). Proposition 2.4 is established.

For every t>0t>0, it is easily checked that the mapping x↦τνt(x)x\mapsto\tau_{\nu^{t}}(x) appearing in (2.20) admits the probabilistic representation

We are now in a position to prove Theorem 2.2.

Proof of Theorem 2.2. As announced, on the basis of the interpolation (2.14), we apply (2.15) and (2.17) respectively to bound the Fisher information Iγ(Pth){\rm I}_{\gamma}(P_{t}h) for tt around and away from . We thus get, for every u>0u>0,

Optimizing in uu (set 1−e−2u=r∈(0,1)1-e^{-2u}=r\in(0,1)) concludes the proof. ∎

It is worth mentioning that a slight modification of the proof of (iii) in Proposition 2.4 leads to the improved form of the exponential decay (2.15) of the Fisher information

As for the classical logarithmic Sobolev inequality, the inequality (2.22) may be integrated along de Bruijin’s formula (2.13) towards the better, although less tractable, HSI inequality

(understood in the limit as S2=I{\rm S}^{2}={\rm I}), where H=H(ν ∣ γ){\rm H}={\rm H}(\nu\,|\,\gamma), S=S(ν ∣ γ){\rm S}={\rm S}(\nu\,|\,\gamma) and I=I(ν ∣ γ){\rm I}={\rm I}(\nu\,|\,\gamma).

Together with the de Bruijn identity (2.13), the classical logarithmic Sobolev inequality (1.1) ensures the exponential decay in t≥0t\geq 0 of the relative entropy

along the Ornstein-Uhlenbeck semigroup (cf. e.g. [B-G-L, Theorem 5.2.1]). The new HSI produces a reinforcement of this exponential convergence to equilibrium under finiteness of the Stein discrepancy.

Let ν\nu with Stein discrepancy S(ν ∣ γ)=S{\rm S}(\nu\,|\,\gamma)={\rm S}. For any t≥0t\geq 0,

Together with (2.18) and since r\mapsto r\log\big{(}1+\frac{s}{r}\big{)} is increasing for any fixed ss, the HSI inequality applied to νt\nu^{t} implies that

Set U(t)=e4tS2 H(νt ∣ γ)U(t)=\frac{e^{4t}}{{\rm S}^{2}}\,{\rm H}(\nu^{t}\,|\,\gamma), t≥0t\geq 0, so that by (2.13), U′=4U−e4tS2 I(νt ∣ γ)U^{\prime}=4U-\frac{e^{4t}}{{\rm S}^{2}}\,{\rm I}(\nu^{t}\,|\,\gamma). The latter inequality therefore rewrites as

Since er−1−r≥r22e^{r}-1-r\geq\frac{r^{2}}{2} for r≥0r\geq 0, this inequality may be relaxed into −2U+2U2 ≤ −U′{-2U+2U^{2}\,\leq\,-U^{\prime}}. Setting V(t)=e−2tU(t)V(t)=e^{-2t}U(t), t≥0t\geq 0, it follows that 2e2tV2(t)≤−V′(t)2e^{2t}V^{2}(t)\leq-V^{\prime}(t) so that, after integration,

By definition of VV, this inequality amounts to the conclusion of Corollary 2.7 and the proof is complete. ∎

4 Stein discrepancy and concentration inequalities

Equivalently (up to numerical constants) in terms of moment growth,

Hence S2(ν ∣ γ)=S(ν ∣ γ){\rm S}_{2}(\nu\,|\,\gamma)={\rm S}(\nu\,|\,\gamma) is the Stein discrepancy as defined earlier. Recall ∥⋅∥op{\|\cdot\|}_{\rm op} the operator norm on the d×dd\times d matrices.

Before turning to the proof of this result, let us comment on its measure concentration content. One first important aspect is that the constant CC is dimension free. When ν=γ\nu=\gamma, (2.28) exactly fits the Gaussian case (2.27). In general, the moment growth in pp describes various concentration regimes of ν\nu (cf. [L2, Section 1.3], [B-L-M, Chapter 14]) according to the growth of the pp-Stein discrepancy Sp(ν ∣ γ){\rm S}_{p}(\nu\,|\,\gamma).

In view of the elementary estimate ∥τν∥op≤1+∥τν−Id∥HS{\|\tau_{\nu}\|}_{\rm op}\leq 1+{\|\tau_{\nu}-{\rm Id}\|}_{\rm HS}, the conclusion (2.28) immediately yields the moment growth (1.6) emphasized in the introduction

Note that there is already an interest to write this bound for p=2p=2,

Together with E. Milman’s Lipschitz characterization of Poincaré inequalities for log-concave measures [M], it shows that the Stein discrepancy S(ν ∣ γ){\rm S}(\nu\,|\,\gamma) with respect to the standard Gaussian measure is another control of the spectral properties in this class of measures.

Similar inequalities hold for arbitrary covariances by suitably adapting the Stein kernel as in Corollary 2.3.

By the triangle inequality, the latter is bounded from above by

which produces a first bound of interest. If it is assumed in addition that the covariance matrix of XX is the identity, we may use classical inequalities for sums of independent centered random vectors (in Euclidean space) to the family τν(Xk)−Id\tau_{\nu}(X_{k})-{\rm Id}, k=1,…,nk=1,\ldots,n. Hence, by for example Rosenthal’s inequality (see e.g. [B-L-M, M-J-C-F-T]), for p≥2p\geq 2,

for some numerical C>0C>0. Note that the bound is optimal both for ν=γ\nu=\gamma and as n→∞n\to\infty describing the standard Gaussian concentration (2.27). By Markov’s inequality, optimizing in p≥2p\geq 2, one deduces that for some numerical C′>0C^{\prime}>0,

for all 0≤r≤rn0\leq r\leq r_{n} where rn→∞r_{n}\to\infty according to the growth of Sp{\rm S}_{p} as p→∞p\to\infty. For example, if Sp=O(pα){\rm S}_{p}=O(p^{\alpha}) for some α>0\alpha>0 (see below for such illustrations), then

(for some possibly different numerical C>0C>0) for every p≤n12α+2p\leq n^{\frac{1}{2\alpha+2}}. By Markov’s inequality in this range of pp,

and with p∼r24C2p\sim\frac{r^{2}}{4C^{2}}, the claims follows with rnr_{n} of the order of n14α+4n^{\frac{1}{4\alpha+4}}.

For concreteness, we describe two classes of one-dimensional distributions such that the associated Stein kernel has finite moments of all orders. Denote by XX a centered real-valued random variable with law ν\nu and Stein kernel τν\tau_{\nu}. Recall from (2.2) of Remark 2.1, that if ν\nu has density ρ\rho with respect to the Lebesgue measure, a version of τν\tau_{\nu} is given by τν(x)=ρ(x)−1∫x∞yρ(y)dy\tau_{\nu}(x)=\rho(x)^{-1}\int_{x}^{\infty}y\rho(y)dy for xx inside the support of ρ\rho.

Assume that the support of ρ\rho coincides with an open interval of the type (a,b)(a,b), with −∞≤a<b≤+∞-\infty\leq a<b\leq+\infty. Say then that the law ν\nu of XX is a (centered) member of the Pearson family of continuous distributions if the density ρ\rho satisfies the differential equation

for some real numbers a0,a1,b0,b1,b2a_{0},a_{1},b_{0},b_{1},b_{2}. We refer the reader e.g. to [DZ, Sec. 5.1] for an introduction to the Pearson family. It is a well-known fact that there are basically five families of distributions satisfying (2.30): the centered normal distributions, centered gamma and beta distributions, and distributions that are obtained by centering densities of the type ρ(x)=Cx−αe−β/x\rho(x)=Cx^{-\alpha}e^{-\beta/x} or ρ(x)=\rho(x)= C(1+x)−αexp⁡(βarctan⁡(x)){C(1+x)^{-\alpha}\exp(\beta\arctan(x))} (CC being a suitable normalizing constant). According to [Ste, Theorem 1, p. 65], if τν\tau_{\nu} satisfies

then τν(x)=αx2+βx+γ\tau_{\nu}(x)=\alpha x^{2}+\beta x+\gamma, x∈(a,b)x\in(a,b) (with α, β, γ\alpha,\,\beta,\,\gamma real constants) if and only if ν\nu is a member of the Pearson family in the sense that ρ\rho satisfies (2.30) for every x∈(a,b)x\in(a,b) with a0=βa_{0}=\beta, a1=2α+1a_{1}=2\alpha+1, b0=γb_{0}=\gamma, b1=βb_{1}=\beta and b2=αb_{2}=\alpha. It follows that if ν\nu is centered member of the Pearson family such that (2.31) is satisfied and XX has finite moments of all orders, so has τν(X)\tau_{\nu}(X). This includes the case of Gaussian, gamma and beta distribution for example.

In concrete instances, such as Wiener chaos for example, the latter expression may be easily controled so to yield concentration properties of the underlying distribution of FF. For example, in the setting of the recent [A-C-P], it may be shown by hypercontractive means that for the Hermite, Laguerre or Jacobi (or mixed ones) chaos structures, for any p≥2p\geq 2,

According to the respective growth in pp of Cp,λC_{p,\lambda}, concentration properties on FF may be achieved.

Using that ∣∇u∣≤1|\nabla u|\leq 1 (since uu is 11-Lipschitz) and furthermore

it easily follows as in the previous section that for every tt,

where α=2q\alpha=2q and (2q−1)β=2q(2q-1)\beta=2q, and

where α′=q\alpha^{\prime}=q and (2q−2)β′=2q(2q-2)\beta^{\prime}=2q. Therefore (2.32) implies that, for every tt,

Integrating this differential inequality yields that

where C~(t)=∫t∞C(s)ds{\widetilde{C}}(t)=\int_{t}^{\infty}C(s)ds, t≥0t\geq 0. It follows that ϕ(0)≤eC~(0)∫0∞D(s)ds\phi(0)\leq e^{{\widetilde{C}}(0)}\int_{0}^{\infty}D(s)ds and therefore

Since C~(0){\widetilde{C}}(0) is bounded above by CqCq for some numerical C>0C>0, the announced claim follows. The proof of Theorem 2.8 is therefore complete. ∎

5 On the rate of convergence in the entropic central limit theorem

On the other hand, as in the previous paragraph,

where α(a)=∑i=1nai4\alpha(a)=\sum_{i=1}^{n}a_{i}^{4}.

As a consequence therefore of the HSI inequality of Theorem 2.2,

This result has to be compared with the works [A-B-B-N] and [B-J] (cf. [J]) which produce the bound

under the hypothesis that ν\nu satisfies a Poincaré inequality with constant c>0c>0.

For the classical average Tn=1n∑k=1nXkT_{n}=\frac{1}{\sqrt{n}}\sum_{k=1}^{n}X_{k}, (2.34) yields a rate O(1n)O(\frac{1}{n}) in the entropic central limit theorem while (2.33) only produces O(log⁡nn)O(\frac{\log n}{n}), however at a cheap expense and under potentially different conditions as described in Remark 2.9. For this classical average, the recent works [B-C-G1, B-C-G2] actually provide a complete picture with rate O(1n)O(\frac{1}{n}) under a fourth-moment condition on XX based on local central limit theorems and Edgeworth expansions. General sums T=∑k=1nakXkT=\sum_{k=1}^{n}a_{k}X_{k} are studied in [B-C-G3] as a particular case of sums of independent non-identically distributed random variables. Vector-valued random variables may be considered similarly.

Transport distances and Stein discrepancy

We shall subdivide the analysis into two parts. In Section 3.1, we deal with the special case of the quadratic Wasserstein distance W2{\rm W}_{2}, for which we use the definition (2.1) of a Stein kernel. In Section 3.2, we deal with general Wasserstein distances Wp{\rm W}_{p} possibly of order p≠2p\neq 2, for which it seems necessary to use the stronger definition (2.3) adopted in [N-P-S1, N-P-S2].

Note that (3.2) is actually the central argument in the Otto-Villani theorem [O-V] asserting that a logarithmic Sobolev inequality implies a Talagrand transport inequality. Here, by making use of (3.2) and then (2.17) we get that

The general case is obtained by a simple regularization procedure which is best presented in probabilistic terms. Fix ε>0\varepsilon>0 and introduce the auxiliary random variable Fε=e−εF+1−e−2εZF_{\varepsilon}=e^{-\varepsilon}F+\sqrt{1-e^{-2\varepsilon}}Z where FF and ZZ are independent with respective laws ν\nu and γ\gamma. It is immediately checked that: (a) the distribution of FεF_{\varepsilon}, denoted by νε\nu^{\varepsilon}, admits a smooth density hεh_{\varepsilon} with respect to γ\gamma (of course, this density coincides with PεhP_{\varepsilon}h whenever the distribution of FF admits a density hh with respect to γ\gamma as in the first part of the proof); (b) a Stein kernel for νε\nu^{\varepsilon} is given by

(consistent with (2.21)); (c) S(νε ∣ γ)≤e−2ε S(ν ∣ γ){\rm S}(\nu^{\varepsilon}\,|\,\gamma)\leq e^{-2\varepsilon}\,{\rm S}(\nu\,|\,\gamma); (d) as ε→0\varepsilon\to 0, FεF_{\varepsilon} converges to FF in L2{\rm L}^{2}, so that, in particular, W2(νε,γ)→W2(ν,γ){\rm W}_{2}(\nu^{\varepsilon},\gamma)\to{\rm W}_{2}(\nu,\gamma). One therefore infers that

The inequality (3.1) may of course be compared to the Talagrand quadratic transportation cost inequality [T, V, B-G-L]

As announced in the introduction, one can actually further refine (3.1) in order to deduce an improvement of (3.3) in the form of a WSH inequality. The refinement relies on the HSI inequality itself.

Proof. For any t≥0t\geq 0, recall dνt=Pthdγd\nu^{t}=P_{t}hd\gamma (in particular, ν0=ν\nu^{0}=\nu and νt→γ\nu^{t}\to\gamma as t→∞t\to\infty). The HSI inequality (2.6) applied to νt\nu^{t} yields that

Now, S2(νt ∣ γ)≤S2(ν ∣ γ){\rm S}^{2}(\nu^{t}\,|\,\gamma)\leq{\rm S}^{2}(\nu\,|\,\gamma) by (2.18) and r\mapsto r\log\big{(}1+\frac{s}{r}\big{)} is increasing for any fixed ss from which it follows that

By exponentiating both sides, this inequality is equivalent to

Combining with (3.2) and recalling (2.13) leads to

The desired conclusion is achieved by integrating between t=0t=0 and t=∞t=\infty. The proof of Theorem 3.2 is complete. ∎

Proposition 3.1 and Theorem 3.2 raise a number of observations.

Since arccos⁡(e−r)≤2r\arccos(e^{-r})\leq\sqrt{2r} for every r≥0r\geq 0, the WSH inequality thus represents an improvement upon the Talagrand inequality (3.3). Moreover, as for the HSI inequality, the WSH inequality produces the case of equality in (3.3) since arccos⁡(e−r)≤2r\arccos(e^{-r})\leq\sqrt{2r} is an equality only at r=0r=0.

The Talagrand inequality may combined with the HSI inequality of Theorem 2.2 to yield the bound

(HWI inequality). As described in the introduction, a fundamental estimate connecting entropy H{\rm H}, Wassertein distance W2{\rm W}_{2} and Fisher information I{\rm I} is the so-called HWI inequality of Otto and Villani [O-V] stating that, for all dν=hdγd\nu=hd\gamma with density hh with respect to γ\gamma,

(see, e.g. [V, pp. 529-542] or [B-G-L, Section 9.3.1] for a general discussion). Recall that the HWI inequality (3.5) improves upon both the logarithmic Sobolev inequality (1.1) and the Talagrand inequality (3.3). It is natural to look for a more general inequality, involving all four quantities H{\rm H}, W2{\rm W}_{2}, I{\rm I} and the Stein discrepancy S{\rm S}, and improving both the HSI and HWI inequalities. One strategy towards this task would be to follow again the heat flow approach of the proof of Theorem 2.2 and write, for 0<u≤t0<u\leq t,

Here, we used (2.15) and (2.17), as well as the known reverse Talagrand inequality along the semigroup given by

(cf. e.g. [B-G-L, p. 446]). Setting α=1−e−2u≤1−e−2t=β\alpha=1-e^{-2u}\leq 1-e^{-2t}=\beta, the preceding estimate yields

However, elementary computations show that, unless the rather unnatural inequality 2W2(ν,γ)≤S(ν ∣ γ)2{\rm W}_{2}(\nu,\gamma)\leq{\rm S}(\nu\,|\,\gamma) is verified, the minimum in the above expression is attained at a point (α,β)(\alpha,\beta) such that either α=β\alpha=\beta (and in this case one recovers HWI) or β=1\beta=1 (yielding HSI). Hence, at this stage, it seems difficult to outperform both HWI and HSI estimates with a single ‘HWSI’ inequality. In the subsequent point (d), we provide an elementary explicit example in which the HSI estimate perform better than the HWI inequality.

In this item, we thus compare the HWI and HSI inequalities on a specific example in dimension d=1d=1. For every n≥1n\geq 1, consider the probability measure dνn(x)=ρn(x)dxd\nu_{n}(x)=\rho_{n}(x)dx with density

where (an)n≥1{(a_{n})}_{n\geq 1} is such that an∈a_{n}\in for every n≥1n\geq 1, a_{n}=o\big{(}\frac{1}{\log n}\big{)} and n2/3an→∞{n^{2/3}a_{n}\to\infty}. A direct computation easily shows that H(νn ∣ γ)→0{\rm H}(\nu_{n}\,|\,\gamma)\to 0. Also, since

one may show after simple (but a bit lengthy) computations that

We next examine the Stein discrepancy S(νn ∣ γ){\rm S}(\nu_{n}\,|\,\gamma) and Wassertein distance W2(νn,γ){\rm W}_{2}(\nu_{n},\gamma). Since a Stein kernel τn\tau_{n} of νn\nu_{n} is given by

Concerning the Wasserstein distance, from the inequality (3.1), we deduce that W2(νn,γ)≤an{\rm W}_{2}(\nu_{n},\gamma)\leq\sqrt{a_{n}}. On the other hand, by the Lipschitz characterization of W1{\rm W_{1}} (specializing to the Lipschitz function x↦∣cos⁡(x)∣x\mapsto|\cos(x)|), cf. e.g. [V, Remark 6.5]),

Now, the right-hand side of this inequality multiplied by 1an\frac{1}{a_{n}} is equal to

which, by dominated convergence, converges to a non-zero limit. As a consequence, there exists c>0c>0 such that, for nn large enough, W2(νn,γ)≥c an{\rm W}_{2}(\nu_{n},\gamma)\geq c\,a_{n}.

Summarizing the conclusions, the quantity

is bigger than a sequence of the order of nan3/2=(n2/3an)3/2na_{n}^{3/2}=(n^{2/3}a_{n})^{3/2}, which (by construction) diverges to infinity as n→∞n\to\infty. This fact implies that, in this specific case, the bound in the HWI inequality diverges to infinity, whereas H(νn ∣ γ)→0{\rm H}(\nu_{n}\,|\,\gamma)\to 0. On the other hand, the HSI bound converges to zero, since

2 General Wasserstein distances under a stronger notion of Stein kernel

(where δij=1\delta_{ij}=1 if i=ji=j and if not), possibly infinite if τνij∉Lp(ν)\tau_{\nu}^{ij}\notin{\rm L}^{p}(\nu). In particular, ∥τν−Id∥2,ν=S(ν ∣ γ){\|\tau_{\nu}-{\rm Id}\|}_{2,\nu}={\rm S}(\nu\,|\,\gamma).

Let p∈[2,∞)p\in[2,\infty). If ν\nu has finite moments of order pp, then (with the same CpC_{p} as in (i))

In particular, for p=2p=2 we recover (3.1).

Owing to an approximation argument analogous to the one rehearsed at end of the proof of Proposition 3.1, it is sufficient to consider the case dν=h dγd\nu=h\,d\gamma where hh is a smooth density. Write as before vt=log⁡Pthv_{t}=\log P_{t}h and dνt=Pthdγd\nu^{t}=P_{t}hd\gamma. By virtue of [N-P-S1, Lemma 2.9], under thus the strengthened assumption (2.3), a version of ∇vt\nabla v_{t}, t>0t>0, is given by

where, as in Remark 2.5, FF and ZZ are independent with respective law ν\nu and γ\gamma, and Ft=e−tF+1−e−2tZF_{t}=e^{-t}F+\sqrt{1-e^{-2t}}Z. Moreover, one can straightforwardly modify the proof of [O-V, Lemma 2] (cf. also [V, Theorem 24.2(iv)]) in order to obtain the general estimate

yielding (i). On the other hand, if p≥2p\geq 2, then

which immediately yields (ii). The proof of Proposition 3.4 is complete. ∎

Specializing (3.6) to the case p=1p=1 yields the estimate

which improves previous dimensional bounds obtained by an application of the multidimensional Stein method (cf. the proof of [N-P2, Theorem 6.1.1]). It is important to note that, apart from the results obtained in the present paper, there is no other version of Stein’s method allowing one to deal with Wasserstein distances of order p>1p>1. Observe that coupling results from [C3] (that are based on completely different methods) may be used to deduce analogous estimates in the case when d=1d=1 and the Stein kernel τν\tau_{\nu} is bounded.

HSI inequalities for further distributions

The operator L\mathcal{L} satisfies the chain rule formula and defines a diffusion operator. We assume that L\mathcal{L} is the generator of a symmetric Markov semigroup (Pt)t≥0{(P_{t})}_{t\geq 0}, where the symmetry is with respect to an invariant probability measure μ\mu.

A central object of interest in this context is the carré du champ operator Γ\Gamma defined from the generator L\mathcal{L} by

for all (f,g)∈A×A(f,g)\in\mathcal{A}\times\mathcal{A}. Note that Γ\Gamma is bilinear and symmetric and Γ(f,f)≥0\Gamma(f,f)\geq 0. Moreover, the integration by parts property for L\mathcal{L} with respect to the invariant measure μ\mu is expressed by the fact that, for functions f,g∈Af,g\in\mathcal{A},

The structure (E,μ,Γ)(E,\mu,\Gamma) then defines a Markov Triple in the sense of [B-G-L] to which we refer for the necessary background.

The requested semigroup analysis toward HSI inequalities will actually involve in addition the iterated gradient operators Γn\Gamma_{n}, n≥1n\geq 1, defined inductively for (f,g)∈A×A(f,g)\in\mathcal{A}\times\mathcal{A} via the relations Γ0(f,g)=fg\Gamma_{0}(f,g)=fg and

In particular Γ1=Γ\Gamma_{1}=\Gamma and the operators Γn\Gamma_{n}, n≥1n\geq 1, are similarly symmetric and bilinear. In what follows, we shall often adopt the shorthand notation Γn(f)\Gamma_{n}(f) instead of Γn(f,f)\Gamma_{n}(f,f). The Γ2\Gamma_{2} operator is part of the famous Bakry-Émery criterion for logarithmic Sobolev inequalities [B-E], [B-G-L, Section 5.7]. As a new feature of the analysis here, the iterated gradient Γ3\Gamma_{3} will turn essential towards a suitable analogue of (iii) in Proposition 2.4.

Given thus the preceding Markov Triple (E,μ,Γ)(E,\mu,\Gamma) associated to the second order differential operator L\mathcal{L} of (4.1), let dν=hdμd\nu=hd\mu where hh is a smooth probability density with respect to μ\mu. As in the Gaussian case, the relative entropy of ν\nu with respect to μ\mu is the quantity

Similarly, the Fisher information of ν\nu (or hh) with respect to μ\mu is defined as

The (integrated) de Bruijn’s identity (cf. Proposition 5.2.2 in [B-G-L]) reads as in (i) of Proposition 2.4,

is a Stein kernel for the probability ν\nu on EE with respect to the generator of L\cal L of (4.1), where b=(bi(x))1≤i≤db={(b_{i}(x))}_{1\leq i\leq d} is part of the definition of L\cal L. For the Ornstein-Uhlenbeck operator L=Δ−x⋅∇{\cal L}=\Delta-x\cdot\nabla, the definition corresponds to (2.1). Since ∫ELf dμ=0\int_{E}\mathcal{L}f\,d\mu=0, observe that aa is a Stein kernel for μ\mu. The main result in this section is an HSI inequality that relates H(ν ∣ μ){\rm H}(\nu\,|\,\mu), I(ν ∣ μ){\rm I}(\nu\,|\,\mu) and the Stein discrepancy of ν\nu with respect to μ\mu

that we regard, as in the Gaussian case of Section 2, as a measure of the distance between ν\nu and μ\mu (since τμ=a\tau_{\mu}=a). Note that choosing a=Ca=C in (4.4), with CC non-singular, yields the quantity arising in Corollary 2.3. It should also be mentioned that the Stein discrepancy (4.4) is somewhat in contrast with the bounds one customarily obtains when applying Stein’s method (see e.g. [N-P1] for the specific example of the one-dimensional Gamma distribution, or [R] for a general reference), which typically involve quantities of the type ∫E∥τν−a∥HS2 dν\int_{E}\|\tau_{\nu}-a\|^{2}_{\rm HS}\,d\nu. The appearance of the inverse matrices a−12a^{-\frac{1}{2}} seems to be inextricably connected with the fact that we deal with information-theoretical functionals.

The following general statement collects the necessary assumptions on the iterated gradients Γ\Gamma, Γ2\Gamma_{2} and Γ3\Gamma_{3} to achieve the expected HSI inequality by the semigroup interpolation scheme. The next paragraphs will provide illustrations in various concrete instances of interest. In Theorem 4.1 below, (i) amounts to the Bakry-Émery Γ2\Gamma_{2} criterion to ensure the logarithmic Sobolev inequality (cf. [B-G-L, Section 5.7]) while condition (ii) linking the Γ2\Gamma_{2} and Γ3\Gamma_{3} operators will provide (together with (iii)) the suitable semigroup bound for the time control of I(Pth){\rm I}(P_{t}h) away from . Recall Ψ(r)=1+log⁡r\Psi(r)=1+\log r if r≥1r\geq 1 and Ψ(r)=r\Psi(r)=r if 0≤r≤10\leq r\leq 1.

In the preceding context, let dν=hdμd\nu=hd\mu where hh is a smooth density with Stein kernel τν\tau_{\nu} with respect to μ\mu. Assume that there exists ρ,κ,σ>0\rho,\kappa,\sigma>0 such that, for any f∈Af\in\mathcal{A},

Γ3(f)≥κ Γ2(f)\Gamma_{3}(f)\geq\kappa\,\Gamma_{2}(f);

Γ2(f)≥σ ∥a12 Hess(f) a12∥HS2\Gamma_{2}(f)\geq\sigma\,\|a^{\frac{1}{2}}\,{\rm Hess}(f)\,a^{\frac{1}{2}}\|_{\rm HS}^{2} (with aa as in (4.1)).

Note that in the Ornstein-Uhlenbeck example, ρ=κ=σ=1\rho=\kappa=\sigma=1 from which we recover the HSI inequality (2.6), however in a slightly weaker formulation.

It is therefore a classical fact (see e.g. [B-G-L, (5.7.4)]) that (i) ensures the exponential decay of the Fisher information along the semigroup

for every t≥0t\geq 0 (and then yields a logarithmic Sobolev inequality for μ\mu.) Now, fix t>0t>0 and let f∈Af\in\mathcal{A}. The Γ\Gamma-calculus as developed in [B-G-L], but at the level of the Γ2\Gamma_{2} and Γ3\Gamma_{3} operators, yields on [0,t][0,t] (by the very definition of Γ3\Gamma_{3} from Γ2\Gamma_{2}),

By (ii), the latter is non-negative so that the map s↦Ps(Γ2(Pt−sf))e−2κss\mapsto P_{s}(\Gamma_{2}(P_{t-s}f))e^{-2\kappa s} is increasing on [0,t][0,t], and thus

Together with (iii), it then follows that

We shall apply (4.6) to vt=log⁡Pthv_{t}=\log P_{t}h (with hh regular enough). First, by symmetry of μ\mu with respect to (Pt)t≥0{(P_{t})}_{t\geq 0},

where the last step follows from (4.6). Since

Finally, using (4.5) for small tt and (4.8) for large tt, one deduces that, for every u>0u>0,

Now, using that 1−rρ≤max⁡(1,ρκ)(1−rκ)1-r^{\rho}\leq\max(1,\frac{\rho}{\kappa})(1-r^{\kappa}) for r∈(0,1)r\in(0,1), a simple (non-optimal) optimization yields the desired conclusion. The proof of Theorem 4.1 is complete. ∎

It should be pointed out that, on the basis of (4.8), transport inequalities as studied in Section 3 may be investigated similarly in the preceding general context, and with similar illustrations as developed below. For example, as an analogue of (3.1),

In order not to expand too much the exposition, we leave the details to the reader.

The next paragraphs present various illustrations of Theorem 4.1.

2 Multivariate gamma distribution

After some easy but cumbersome calculations, it may be checked that, along suitable smooth functions ff,

Note that (recall xi,xj,xk≥0x_{i},x_{j},x_{k}\geq 0)

Since pi≥32p_{i}\geq\frac{3}{2}, it follows at once that Γ3(f)≥12 Γ2(f)\Gamma_{3}(f)\geq\frac{1}{2}\,\Gamma_{2}(f). Analogous computations lead to

As a consequence, Theorem 4.1 applies with ρ=κ=σ=12\rho=\kappa=\sigma=\frac{1}{2} to yield the following result (the numerical constants there are not sharp). The restrictions pi≥32p_{i}\geq\frac{3}{2}, i=1,…,d{i=1,\ldots,d}, are probably not optimal. For example, it is not difficult to see from the preceding computations that in the one-dimensional case d=1d=1, it is actually enough to assume that p≥12p\geq\frac{1}{2}.

3 One-dimensional uniform distribution on [−1,+1]11[-1,+1]

1[-1,+1] In this section, we examine the case of the one-dimensional Jacobi operator of parameters α=β=1\alpha=\beta=1, that is,

whose associated invariant measure μ\mu is uniform distribution on [−1,+1][-1,+1]. The general family of parameters with the beta distributions as invariant measures (cf. [B-G-L, Section 2.7.4]) may be considered similarly, at the expense however of tedious computations, as well as multivariate (product) versions. For simplicity, we only detail this case to better illustrate the conclusion.

Easy calculations lead to, for a smooth function ff on [−1,+1][-1,+1],

so that Γ3(f)≥Γ2(f)\Gamma_{3}(f)\geq\Gamma_{2}(f). Also,

Hence, Theorem 4.1 applies with ρ=κ=1\rho=\kappa=1 and σ=12\sigma=\frac{1}{2} (note that a(x)=1−x2a(x)=1-x^{2}) to yield the following conclusion. Again, the numerical constants are not sharp.

Let μ\mu be uniform probability measure on [−1,+1][-1,+1]. Then, for any dν=hdμd\nu=hd\mu where hh is a smooth probability density,

4 Families of log-concave distributions

We consider here a diffusion operator on the line of the type

Assume that there exists c>0c>0 such that, uniformly, u′′≥cu^{\prime\prime}\geq c,

Then Γ2(f)≥c Γ(f)\Gamma_{2}(f)\geq c\,\Gamma(f), Γ2(f)≥f′′2\Gamma_{2}(f)\geq{f^{\prime\prime}}^{2} and Γ3(f)≥3c Γ2(f)\Gamma_{3}(f)\geq 3c\,\Gamma_{2}(f) for every ff. Hence, Theorem 4.1 applies with ρ=c\rho=c, κ=3c\kappa=3c and σ=1\sigma=1.

Recall that in this context, the only condition u′′≥c>0u^{\prime\prime}\geq c>0 ensures the logarithmic Sobolev inequality for μ\mu [B-G-L, Corollary 5.7.2]. It is not difficult to find (simple) examples outside the Gaussian model (corresponding to c=13c=\frac{1}{3}) such that conditions (4.9) and (4.10) are fulfilled. For example, if u(x)=x22+εx4u(x)=\frac{x^{2}}{2}+\varepsilon x^{4}, it is easily seen that these hold for c=14c=\frac{1}{4} and ε=112\varepsilon=\frac{1}{12} (for instance). In the Gaussian case, the estimate obtained in this proposition is somewhat worse than the HSI inequality of Theorem 2.2. At the expenses of more involved conditions (4.9) and (4.10), multidimensional versions may be considered similarly.

Entropy bounds on laws of functionals

Referring as before to [B-G-L] for a complete account, we thus deal with a Markov Triple (E,μ,Γ)(E,\mu,\Gamma) on a probability space (E,E,μ)(E,{\cal E},\mu), with Markov semigroup (Pt)t≥0{(P_{t})}_{t\geq 0} with symmetric and invariant probability measure μ\mu, infinitesimal generator L\rm L, associated carré du champ operator Γ\Gamma and underlying algebra of (smooth) functions A\cal A. Integration by parts expresses that

The second order differential operators of Section 4 provide instances of this general framework. Gaussian and Wiener spaces with associated Ornstein-Uhlenbeck semigroup and generator are a prototypical example for the illustrations. Note in particular that Wiener chaoses as investigated in [N-P-S1] are eigenfunctions of the Ornstein-Uhlenbeck generator. Eigenfunctions of the underlying operator L\rm L are actually of special interest in the context of the Stein method as illustrated in Section 5.1.

The first statement shows that, whenever the vector FF is composed of eigenfunctions of L{\rm L}, a Stein kernel τνF\tau_{\nu_{F}} of νF\nu_{F} with respect to γ\gamma as defined in (2.1) can be expressed in terms of the carré du champ operator Γ\Gamma.

Let F=(F1,…,Fd)F=(F_{1},\ldots,F_{d}) on (E,E,μ)(E,{\cal E},\mu) such that, for every i=1,…,di=1,\ldots,d, the random variable FiF_{i} is an eigenfunction of −L-{\rm L}, with eigenvalue λi>0\lambda_{i}>0. Assume moreover that Γ(Fi,Fj)∈L1(μ)\Gamma(F_{i},F_{j})\in{\rm L}^{1}(\mu) for every i,j=1,…,di,j=1,\ldots,d. Then, the matrix-valued map τνF\tau_{\nu_{F}} defined as

is a Stein kernel for νF\nu_{F}, that is, it satisfies (2.1). (The right-hand side of (5.2) indicates a version of the conditional expectation of Γ(Fi,Fj)\Gamma(F_{i},F_{j}) with respect to FF under the probability measure μ\mu.)

The proof is concluded by taking conditional expectations. ∎

As a consequence, together with (2.5) and Jensen’s inequality,

where CC denotes the covariance matrix of νF\nu_{F}, providing therefore a tractable way to control the Stein discrepancy in this case. In addition, combining with the HSI inequality of Theorem 2.2 immediately yields the following statement.

Under the assumptions and notation of Proposition 5.1,

In particular, if d=1d=1 and C=1C=1, H(νF ∣ γ)→0{\rm H}(\nu_{F}\,|\,\gamma)\to 0 whenever Var(Γ(F))→0{\rm Var}(\Gamma(F))\to 0 (cf. [N-P2, L3]).

A typical example of a Markov Triple for which the quantity V2{\rm V}^{2} appearing in the above bound can be estimated explicitly corresponds to the case where (E,E,μ)(E,\mathcal{E},\mu) is a probability space supporting an isonormal Gaussian process X={X(h):h∈H}X=\{X(h):h\in\mathfrak{H}\} over some real separable Hilbert space H\mathfrak{H}, and L{\rm L} is the generator of the associated Ornstein-Uhlenbeck semigroup. In this case, Γ(F,G)=⟨DF,DG⟩H\Gamma(F,G)={\langle DF,DG\rangle}_{\mathfrak{H}} for smooth functionals FF and GG, where DD stands for the Malliavin derivative operator, and the eigenspaces of −L-{\rm L} are the so-called Wiener chaoses {Ck:k≥0}\{C_{k}:k\geq 0\} of XX. For k=0,1,2,…k=0,1,2,\ldots, the eigenvalue of CkC_{k} is given by kk. A detailed discussion about how to bound a quantity such as V2{\rm V}^{2} in the case of random vectors with components inside a Wiener chaos can be found in [N-P2, Chapter 6]. In particular, if d=1d=1 and FF belongs to CkC_{k}, then V2{\rm V}^{2} can be controlled by the second and fourth moments of FF as

In particular, such an estimate provides a proof of the famous ‘fourth moment theorem’ for chaotic random variables, cf. [N-P2, Theorem 5.2.7].

While eigenfunctions appear as functionals of particular interest for the control of the Stein discrepancy itself, the Γ\Gamma-calculus actually provides a formal description of Stein kernels of a given functional FF on (E,μ,Γ)(E,\mu,\Gamma) (in dimension one for simplicity) as the conditional expectation with respect to FF of Γ(F,L−1F)\Gamma(F,{\rm L}^{-1}F) (where L−1F=∫0∞PtFdt{\rm L}^{-1}F=\int_{0}^{\infty}P_{t}Fdt). This observation further expands on the preceding example, allowing for a rather general analysis.

2 Bounds on the Fisher information

When dealing with the upper-bound (5.4), the Fisher information I(νF ∣ γ)=Iγ(h){\rm I}(\nu_{F}\,|\,\gamma)={\rm I}_{\gamma}(h) of the density hh of the law νF\nu_{F} of FF cannot always be explicitly deduced from the data concerning the random vector FF. The task of this paragraph is therefore to deduce some useful bounds on I(νF ∣ γ){\rm I}(\nu_{F}\,|\,\gamma) in terms of FF and its gradients.

Let Γ~{\widetilde{\Gamma}} be the symmetric matrix with entries Γ(Fi,Fj)\Gamma(F_{i},F_{j}), i,j=1,…,di,j=1,\ldots,d. Applying the latter to w=wijw=w_{ij}, symmetric in i,ji,j, yields

Applied to ϕ=v=log⁡h\phi=v=\log h, by the Cauchy-Schwarz inequality and (4.2),

The consequences of the previous computations are gathered together in the next statement, where we point out a set of sufficient conditions on FF and its gradients Γ(Fi,Fj)\Gamma(F_{i},F_{j}) ensuring that the random variable UU is indeed square-integrable.

Let F=(F1,…,Fd)F=(F_{1},\ldots,F_{d}) be a vector of elements of A\cal A on (E,μ,Γ)(E,\mu,\Gamma). Assume that all the FiF_{i}, LFi{\rm L}F_{i}, Γ(Fi,Fj)\Gamma(F_{i},F_{j}), i,j=1,…,di,j=1,\ldots,d, and 1det(Γ~)\frac{1}{{\rm det}({\widetilde{\Gamma}})} are in Lp(μ){\rm L}^{p}(\mu) for every p≥1p\geq 1. Then, ∫E∣U∣2dμ<∞\int_{E}|U|^{2}d\mu<\infty and

The condition on 1det(Γ~)\frac{1}{{\rm det}({\widetilde{\Gamma}})} in Proposition 5.5 has some similarity with basic assumptions in Malliavin calculus (cf. [N, N-P2]).

3 Fisher information growth and normal approximation

One evident drawback of Proposition 5.5 of the previous paragraph is that, since the quantity ∣U∣|U| is singular as the determinant of Γ~{\widetilde{\Gamma}} is close to , one is forced to assume that 1det(Γ~)\frac{1}{{\rm det}({\widetilde{\Gamma}})} is in all Lp(μ){\rm L}^{p}(\mu) spaces (or at least for some pp large enough depending on dd). This assumption is in general too strong, and very difficult to check in concrete situations. The idea developed in this section (which generalizes the approach initiated in [N-P-S1]) is that, under weaker moment assumptions, while the Fisher information Iγ(h){\rm I}_{\gamma}(h) might be infinite, it is nevertheless possible to control the growth as t→0t\to 0 of Iγ(Pth){\rm I}_{\gamma}(P_{t}h). Together with the control in terms of the Stein discrepancy for large time achieved in Section 2, one may then reach entropic bounds which can be handled in concrete examples (such as those of random vectors whose components belong to some Wiener chaos).

As before, let F=(F1,…,Fd)F=(F_{1},\ldots,F_{d}) be a general vector of centered and square-integrable random variables (in the algebra A\cal A or some natural extension), with distribution dνF=hdγd\nu_{F}=hd\gamma. As a crucial assumption, νF\nu_{F} has a Stein kernel τνF\tau_{\nu_{F}} with respect to γ\gamma as defined in (2.1) (see also Proposition 5.1 and Remark (5.4)). Recall the matrix Γ~{\widetilde{\Gamma}} with entries Γ(Fi,Fj)\Gamma(F_{i},F_{j}), i,j=1,…,di,j=1,\ldots,d. Also, in what follows we use the convention that, if Γ~\widetilde{\Gamma} is singular, then the matrix det(Γ~) Γ~−1{\rm det}({\widetilde{\Gamma}})\,{\widetilde{\Gamma}}^{-1} must be understood as the transpose of usual adjugate matrix operator of Γ~\widetilde{\Gamma} (both quantities being of course equal for non-singular matrices).

Choose W=det(Γ~) Γ~−1det(Γ~)+εW=\frac{{\rm det}({\widetilde{\Gamma}})\,{\widetilde{\Gamma}}^{-1}}{{\rm det}({\widetilde{\Gamma}})+\varepsilon} in (5.5), so that

Apply now the preceding to ϕ=Ptvt\phi=P_{t}v_{t}, vt=log⁡Pthv_{t}=\log P_{t}h, t>0t>0. Since ∇Ptvt(F)=e−tPt(∇vt)\nabla P_{t}v_{t}(F)=e^{-t}P_{t}(\nabla v_{t}) and

by the Cauchy-Schwarz inequality, assuming for simplicity that 0<ε≤10<\varepsilon\leq 1,

On the other hand, using the same semigroup computations as in Section 2,

Collecting the preceding bounds and recalling from (4.7) that

yields that, for t>0t>0 and 0<ε≤10<\varepsilon\leq 1,

Let F=(F1,…,Fd)F=(F_{1},\ldots,F_{d}) be a vector of centered elements of A\cal A on (E,μ,Γ)(E,\mu,\Gamma). Assume that all the FiF_{i}, LFi{\rm L}F_{i}, Γ(Fi,Fj)\Gamma(F_{i},F_{j}), i,j=1,…,di,j=1,\ldots,d, are in Lp(μ){\rm L}^{p}(\mu) for every p≥1p\geq 1, and that

for some α>0\alpha>0. Then, AF<∞A_{F}<\infty (as defined in (5.7)) and

where κ=2+α2(4+3α) (<14 )\kappa=\frac{2+\alpha}{2(4+3\alpha)}\,(<\frac{1}{4}\,). In particular, under the assumptions on FF, H(νF ∣ γ)→0{\rm H}(\nu_{F}\,|\,\gamma)\to 0 as S(νF ∣ γ)→0{\rm S}(\nu_{F}\,|\,\gamma)\to 0.

First of all, we have that the parameter AFA_{F} is finite, since the expressions det(Γ~)Γ~−1LF{\rm det}({\widetilde{\Gamma}}){\widetilde{\Gamma}}^{-1}{\rm L}F, V1V_{1} and V2V_{2} only involve products of FiF_{i}, LFi{\rm L}F_{i} and Γ(Fi,Fj)\Gamma(F_{i},F_{j}), i,j=1,…,di,j=1,\ldots,d.

Now, for every ε>0\varepsilon>0 and r>0r>0,

The choice of r=ε2α+2r=\varepsilon^{\frac{2}{\alpha+2}} yields (5.8) with δ(ε)=(BF+1)ε−42+α\delta(\varepsilon)=(B_{F}+1)\varepsilon^{-\frac{4}{2+\alpha}}. Let then ε=ε(t)=(1−e−2t)κ\varepsilon=\varepsilon(t)=(1-e^{-2t})^{\kappa}, t≥0t\geq 0, for κ=2+α2(4+3α)\kappa=\frac{2+\alpha}{2(4+3\alpha)} (<14<\frac{1}{4}). Then

from which, as a consequence of (5.9), for every t>0t>0,

To conclude, recall, as in the proof of Theorem 2.2, the decomposition for every u>0u>0,

and the bound (5.11) in the statement follows by optimizing in u>0u>0 (set (1−e−2u)1−4κ=r∈(0,1)(1-e^{-2u})^{1-4\kappa}=r\in(0,1).) Theorem 5.7 is established. ∎

so that, under the assumptions of Theorem 5.7, one also has that H(νF ∣ γ)<∞{\rm H}(\nu_{F}\,|\,\gamma)<\infty, a conclusion of independent interest.

The quantity AFA_{F} of (5.7) involves integrability conditions on FF and its gradients (they may actually be weakened according to the precise expression of AFA_{F}). On the other hand, BFB_{F} of (5.10) is rather concerned with a small ball behavior. For a vector F=(F1,…,Fd)F=(F_{1},\ldots,F_{d}) of eigenvectors of the underlying Markov generator L{\rm L}, Theorem 5.7 may be combined with (5.3) to fully control the relative entropy in terms of FF and its gradients as now illustrated in some instances.

We describe, in part following [N-P-S1], how the preceding developments may be applied to concrete examples of interest.

for every r>0r>0, where N≥1N\geq 1 is an integer related to the degrees of the FiF_{i}’s. Under (5.14), the second hypothesis of Theorem 5.7 clearly holds for any α<1N\alpha<\frac{1}{N} (cf. (5.12)). The latter then applies to basically recover the main conclusion of [N-P-S].

4 Fisher information growth and gamma approximation

This final section develops the analogous investigation towards gamma approximation, for simplicity one-dimensional. Denote by γp\gamma_{p} the gamma distribution (on the positive real line) with parameter p>0p>0, invariant measure of the Laguerre operator

for every smooth test function φ\varphi. In particular, ∫EFdμ=p\int_{E}Fd\mu=p. Note that, in this case,

From the study of Gaussian chaoses for example, and as already mentioned earlier, it appears that the latter S(νF ∣ γp){\rm S}(\nu_{F}\,|\,\gamma_{p}) might not always be the relevant quantity of interest (cf. [N-P1, R]). Indeed, for an eigenfunction FF with eigenvalue −λ-\lambda, λ>0\lambda>0, the Stein kernel τν(F)\tau_{\nu}(F) may be identified with the conditional expectation of λ−1Γ(F)\lambda^{-1}\Gamma(F) knowing FF. Now, for such a functional, moment conditions on FF may be used to rather control the variance of λ−1Γ(F)−F\lambda^{-1}\Gamma(F)-F, and similarly higher moments (cf. [A-C-P, A-M-P, L3]). Of course, by Hölder’s inequality,

for r>1r>1, 1r+1s=1\frac{1}{r}+\frac{1}{s}=1. Provided it may be ensured that ∫EF−2rdμ<∞\int_{E}F^{-2r}d\mu<\infty for some r>1r>1, the results here are nevertheless still of interest.

We assume below that p≥12p\geq\frac{1}{2} so that the estimates (4.6) and (4.8) are verified, with the choice of parameters d=1d=1 and ρ=κ=σ=12\rho=\kappa=\sigma=\frac{1}{2} (see the comment preceding Proposition 4.3). The proof of the following statement will follow the one developed for Theorem 5.7.

On (E,μ,Γ)(E,\mu,\Gamma), let F≥0F\geq 0 in A\cal A. Assume that FF, LF{\rm L}F, Γ(F)\Gamma(F) and Γ(F,Γ(F))\Gamma(F,\Gamma(F)) are in Lq(μ){\rm L}^{q}(\mu) for every q≥1q\geq 1 and that

where κ=2+α2(4+3α) (<14 )\kappa=\frac{2+\alpha}{2(4+3\alpha)}\,(<\frac{1}{4}\,). In particular, under the assumptions on FF, H(νF ∣ γp)→0{\rm H}(\nu_{F}\,|\,\gamma_{p})\to 0 as S(νF ∣ γp)→0{\rm S}(\nu_{F}\,|\,\gamma_{p})\to 0.

Denoting by (Pt)t≥0{(P_{t})}_{t\geq 0} the semigroup with infinitesimal generator Lp\mathcal{L}_{p}, we have as in (4.7),

where vt=log⁡Pthv_{t}=\log P_{t}h. Now, for every ε>0\varepsilon>0,

As a consequence, with the notation introduced in the statement,

Since Γ(Ptvt)≤e−tPt(Γ(vt))\Gamma(P_{t}v_{t})\leq e^{-t}P_{t}(\Gamma(v_{t})) (Theorem 3.2.4 in [B-G-L]),

On the other hand, the estimate (4.6) yields the bound

Gathering together all the previous estimates, we deduce that, for every 0<ε≤10<\varepsilon\leq 1 and t>0t>0,

On the basis of this estimate, we then conclude exactly as in the proof of Theorem 5.7. ∎

References