Semi Log-Concave Markov Diffusions

Patrick Cattiaux, Arnaud Guillin

Introduction and main results.

In this paper we shall investigate some properties of time marginals (at time TT finite or infinite) of Markov diffusion processes satisfying some logarithmic semi-convexity like property. The properties we are interested in are functional inequalities (Poincaré, log-Sobolev) or transportation inequalities. We shall also give some consequences for the long time behavior of such processes.

Our main tools are on one hand coupling techniques and on the other hand stochastic calculus. We shall mainly use the so called “synchronous” coupling, i.e. using the same Brownian motion, but we also give some new results by using the “mirror” coupling (or coupling by reflection) introduced by Lindvall and Rogers in . The main stochastic tool is (a very simple form of) Girsanov theory and hh-processes.

The use of coupling techniques for obtaining analytic estimates is far to be new. It is impossible (and dangerous) to give here, even an account of the existing literature (see however and references therein). The use of Girsanov theory for this goal is not new too. We shall recall later some references. The conjunction of both techniques is not usual.

We deliberately decided to present in details the simplest situation of a Brownian motion with a gradient drift, for which almost everything is well known, and then to extend our method to new situations. Some specialists would certainly find that these parts of the present paper are lengthy, but we think that the understanding of how the method works in this simple case is an useful guide for generalizations.

The meaning of logarithmic semi-convexity will generalize the “usual” one we recall now.

This property is called KK-semi-convexity of UU. It is clearly equivalent to the convexity of U(x)−K∣x∣2U(x)-K|x|^{2}. We denote Υ(dx)=e−U(x)dx\Upsilon(dx)=e^{-U(x)}dx the Boltzmann measure associated to the potential UU. If e−Ue^{-U} is dxdx integrable, we also introduce the normalized μ(dx)=1ZU e−U(x) dx\mu(dx)=\frac{1}{Z_{U}}\,e^{-U(x)}\,dx which is a probability measure. If UU is semi-convex, μ\mu is said to be semi log-concave.

Consider first the diffusion process, given by the solution of the Ito stochastic differential system

and Γ\Gamma denotes the carré du champ, namely here

It is known that Υ\Upsilon is a symmetric (reversible) measure for the diffusion process, and is actually the unique invariant (stationary) measure for the process. If Υ\Upsilon is bounded, μ\mu is ergodic.

In the latter case, PtP_{t} is thus a symmetric semi-group on L2(μ)L^{2}(\mu). The domain D(L)D(L) of its generator contains the algebra A\mathcal{A} generated by the constant functions and Cc∞C_{c}^{\infty}. In particular, if f∈Af\in\mathcal{A}, ∂tPtf=LPtf=PtLf\partial_{t}P_{t}f=LP_{t}f=P_{t}Lf in L2(μ)L^{2}(\mu), so that since ∂t−L\partial_{t}-L is hypo-elliptic (t,x)↦Ptf(x)∈C∞(t,x)\mapsto P_{t}f(x)\in C^{\infty}.

LL is the basic example of generator satisfying the celebrated C(K/2,+∞)C(K/2,+\infty) Bakry-Emery curvature condition (see ). Indeed if we define

(H.C.K) is equivalent to Γ2(f)≥(K/2) Γ(f)\Gamma_{2}(f)\geq(K/2)\,\Gamma(f).

This curvature condition is known to imply (and is in fact equivalent to) a lot of nice inequalities for the semi-group, in particular for all T>0T>0 and all xx, a commutation between Γ\Gamma and the semi group PtP_{t} holds, namely

which in turn implies powerful functional inequalities such as

Recall that ν\nu satisfies a log-Sobolev inequality with constant CLSC_{LS} if

(1.3) is exactly what is (a little bit improperly) called a “local” log-Sobolev inequality in (theorem 5.4.7). For further informations and more, see the forthcoming book . The reader has to be careful with the constants, since we are writing the log-Sobolev inequality (as well as all other inequalities) in the usual form, while the “Bakry-Emery” version uses Γ\Gamma, hence with an extra factor 2.

It is well known that a log-Sobolev inequality implies a Poincaré inequality

with CP=12 CLSC_{P}=\frac{1}{2}\,C_{LS}, as well as a T2T_{2} transportation inequality

with CW=12 CLSC_{W}=\frac{1}{2}\,C_{LS}. Here W2W_{2} denotes the Wasserstein distance between the probability measures η\eta and ν\nu, i.e.

where π\pi is a coupling of η\eta and ν\nu (i.e. has respective marginals equal to η\eta and ν\nu) and

denotes the Kullback-Leibler information or relative entropy of η\eta w.r.t. ν\nu. The latter property is due to Otto-Villani . Another approach and related properties were developed by Bobkov, Gentil and Ledoux . For a nice survey on transportation inequalities we refer to . One can find in all these references another remarkable consequence of semi log-concavity, namely that a log-Sobolev inequality derives from a transportation inequality. This is a consequence of the following (H.W.I) inequality that holds for any nice μ\mu density of probability hh,

As a consequence, if (H.C.K) holds for some K≤0K\leq 0, a T2T_{2} transportation inequality for μ\mu implies a log-Sobolev inequality with constant CLS≤(4/CW) (1+(K/CW))−2C_{LS}\leq(4/C_{W})\,(1+(K/C_{W}))^{-2} provided 1+(K/CW)>01+(K/C_{W})>0, in particular if K=0K=0. Let us finally remark that the starting point of this approach is the Γ2\Gamma_{2} commutation property (1.2) which fails however to give a direct proof of the T2T_{2} inequality.

Our first goal is to show that functional and transportation inequalities can be derived, in the previous situation, by using synchronous coupling and simple tools of stochastic calculus. This is done in section 2. The methods are then extended to a more general framework which is as natural for studying properties of time marginals as the Γ2\Gamma_{2} framework.

Indeed, consider a classical diffusion process, given by the solution of an Ito stochastic differential system

B.B_{.} being a standard brownian motion. For simplicity we assume that σ\sigma is a squared matrix. We extend (H.C.K) to this new situation

Notice that if σ\sigma and bb are CC-Lipschitz, i.e. each component is CC Lipschitz, (H.C.K) is satisfied for K=−(C2n2+C)K=-(C^{2}n^{2}+C), but if σ\sigma is CC-Lipschitz, (H.C.K) can be satisfied for a non-negative KK provided bb is sufficiently repealing. Contrary to the case of a constant diffusion coefficient, (H.C.K) is not related to the Bakry-Emery curvature condition which involves in this situation controls on derivatives of higher order of the coefficients.

For simplicity in the sequel we shall assume that σ∈Cb2\sigma\in C_{b}^{2} hence is CC-Lipschitz and that bb is C2C^{2}, but not necessarily bounded nor with bounded derivatives. With these assumptions, once again if we assume that (H.C.K) is in force, (1.7) admits a unique non explosive strong solution using x↦x2x\mapsto x^{2} as a Lyapunov function for non explosion. We shall show this and other properties of the process in subsection 3.1.

We still use the notations introduced before, but now

Our first results can be gathered in Theorem 1.8 below, after introducing some additional assumptions.

Hypothesis (R). One of the following assumptions is satisfied (in addition to the fact that σ∈Cb2\sigma\in C_{b}^{2} and b∈C2b\in C^{2}).

σ∈Cb∞\sigma\in C_{b}^{\infty}, b∈C∞b\in C^{\infty} and has at most polynomial growth and ∂t−L\partial_{t}-L is hypo-elliptic,

σ=a1/2\sigma=a^{1/2}, i.e. σ\sigma is symmetric.

Actually, assumptions (R1)-(R4) ensure that for f∈Af\in\mathcal{A},

which is what we really need. (R5) is a limiting situation for (R3) as we will see in (sub)subsection 3.1.2. Of course the time marginal distributions only depend on LL and not on σ\sigma, but the constant KK is related to a1/2a^{1/2} in (R5) (note that the Hilbert Schmidt norm of σ(x)−σ(y)\sigma(x)-\sigma(y) can change when we change σ\sigma without modifying aa).

Assume that (R) and (H.C.K) are satisfied. Let M=sup⁡∣u∣=1 sup⁡x ∣σ(x)u∣2M=\sup_{|u|=1}\,\sup_{x}\,|\sigma(x)u|^{2}.

If μ0\mu_{0} satisfies a Poincaré inequality with constant CP(0)C_{P}(0) then μT\mu_{T} satisfies a Poincaré inequality with constant

When K=0K=0 one has to replace (1−e−KT)K\frac{(1-e^{-KT})}{K} by TT. This applies in particular to P(T,x,.)P(T,x,.) with CP(0)=0C_{P}(0)=0.

P(T,x,.)P(T,x,.) satisfies a T2T_{2} transportation inequality with constant CT=CT(M,K)C_{T}=C_{T}(M,K), bounded in time if K>0K>0, linear in time if K=0K=0 and exploding exponentially in time if K<0K<0.

If μ0\mu_{0} satisfies T2T_{2} with constant CW(0)C_{W}(0), μT\mu_{T} satisfies T2T_{2} with a constant CW(T)=CW(T,M,K,CW(0))C_{W}(T)=C_{W}(T,M,K,C_{W}(0)), bounded in time if K>0K>0, linear in time if K=0K=0 and exploding exponentially in time if K<0K<0.

When σ=Id\sigma=Id, if μ0\mu_{0} satisfies a log-Sobolev inequality with constant CLS(0)C_{LS}(0), μT\mu_{T} satisfies a log-Sobolev inequality with constant

This applies in particular to P(T,x,.)P(T,x,.) with CLS(0)=0C_{LS}(0)=0.

Some other consequences, as for example convergence to equilibrium (when it exists) are also discussed in particular in subsection 3.7.

Of course, (1) is a weaker version of the commutation relation (1.2) and (5) is nothing else than (1.3) when 2b=−∇U2b=-\nabla U. When a diffusion coefficient is present, (1) is however very different from the usual commutation property. For example we will show that it holds even in the negative infinite curvature case, but that it still enables us to provide interesting local functional inequalities. (3) as well as the general version of (H.C.K) we have introduced appeared (for this kind of application and to our knowledge) for the first time in the paper by Djellout, Guillin and Wu , Theorem 5.6 and condition 4.5 therein, for K>0K>0. Our scheme of proof for the transportation inequality, based on Girsanov theory, is actually a simplified version of the one in , but instead of looking at the full law of the process on a time interval we shall use hh-processes in order to look at time marginals. What we shall show is that the same scheme of proof also furnishes functional inequalities. This unified treatment of functional inequalities and transportation inequalities using an ad-hoc coupling is the novelty here. It easily extends to time dependent coefficients as shown in section 4.

In addition in this section we show how to directly obtain convergence to equilibrium and properties of the invariant measure for non linear diffusions of Mc Kean-Vlasov type, simplifying arguments in .

The use of stochastic calculus in deriving such inequalities is not new but only a small number of papers dealt with. One can trace back to the paper of Borell , who used Girsanov theory to study the propagation of log-concavity along the Schrödinger dynamics (not the Fokker-Planck one we are looking at here). In addition to for transportation inequalities, one can also mention where similar ideas are used to study hyper-boundedness. More recently, using similar arguments, Lehec has studied gaussian functional inequalities and Fontbona and Jourdain obtained a pathwise version of the Γ2\Gamma_{2} theory.

Let us come back to (1.1). (H.C.K) (or the Γ2\Gamma_{2} theory) for K>0K>0 applies to potentials UU which are the sum of (K/2)∣x∣2(K/2)|x|^{2} and of a concave (hence sub-linear) potential VV. In particular, for “super-convex” potentials like ∣x∣β|x|^{\beta} with β>2\beta>2, or more generally for (smooth) potentials UU which are uniformly convex at “infinity”, (H.C.K) holds but with a negative KK due to the behavior of UU near the origin, so that, according to theorem 1.8, μT\mu_{T} satisfies functional inequalities but with exploding constants in TT.

It is however well known, since U=V+WU=V+W with VV KK-uniformly convex and WW bounded, that μ(dx)=e−U(x)dx\mu(dx)=e^{-U(x)}dx satisfies a log-Sobolev (and a Poincaré) inequality with a constant CLS=(2/K) exp⁡(OscW)C_{LS}=(2/K)\,\exp(\textrm{Osc}W) where Osc denotes the oscillation of WW. One can thus expect that CLS(T)C_{LS}(T) is bounded in TT. In section 5, we introduce the following extension of (H.C.K).

When α(a)=1\alpha(a)=1, we may take ε=0\varepsilon=0 and we recognize (H.C.K). In Proposition 5.4 we show that U(x)=∣x∣2βU(x)=|x|^{2\beta} (with β≥1\beta\geq 1) satisfies (H.α\alpha.KβK_{\beta}) for α(a)=aβ−1\alpha(a)=a^{\beta-1} and an explicit Kβ>0K_{\beta}>0.

The main result of this section is then that, for suitable functions α\alpha,

if (H.α\alpha.K) holds (for K>0K>0), then μ\mu satisfies a log-Sobolev inequality.

See theorems 5.20 and 5.24. These theorems thus (partly) extend the Bakry-Emery criterion (1.3) to some non uniformly convex potentials. However, they are dealing with the invariant measure only and not with the law at time TT (only incomplete results are proved in this section for these distributions).

The next section 6 is devoted to the use of mirror coupling. In a recent work , Eberle has adapted the mirror coupling to get estimates of W1W_{1} convergence for drifted brownian motions when the drift satisfies some “convexity at infinity” property. We recall Eberle’s method and obtain some new consequences of his result. In addition, up to an extra condition, we show that his result (and all the consequences we derived) can be extended to general elliptic diffusion processes. We will also use this mirror coupling to show that we may get a weak version of the commutation property in the log concave case with the “convexity at infinity” property at least in dimension one, which is the first result we know of in this direction. Still in dimension one, we will also consider using mirror coupling for non linear diffusions.

Section 7 is peculiar. Using the results we have described for the Ornstein-Uhlenbeck process we show how to recover known results on the stability of functional inequalities under convolution (provided one of the terms is gaussian).

Semi log-concave drifted brownian motion.

In this first warming up section we shall look at the usual situation given by (1.1)

For functional inequalities the key is a commutation property of the gradient and the semi group. This commutation property is almost immediate using an appropriate coupling as explained below :

In the situation of (1.1), assume (H.C.K). Then for all f∈Af\in\mathcal{A},

Applying Ito formula yields (almost surely)

for some ztz_{t} sandwiched by XtxX_{t}^{x} and XtyX_{t}^{y}. It remains to use the continuity (and boundedness) of ∇f\nabla f and the fact that XtyX_{t}^{y} goes almost surely to XtxX_{t}^{x} as y→xy\to x to conclude. ∎

As is seen from the proof, in fact, the sole convergence of the Wasserstein distance is not sufficient to get the commutation property exposed here. It will however be our starting point for the result when a diffusion coefficient is present. The synchronous coupling here enables us however to get an almost sure “deterministic” control of Xtx−XtyX_{t}^{x}-X_{t}^{y} which is far more powerful. ♢\diamondsuit

We recall previously that (H.C.K) is exactly the Γ2\Gamma_{2} condition of Bakry-Emery in this context, which is in fact equivalent to (2.3). However the proof is very different from ours: it relies on a tricky calculus on ψ(s)=e−Ks/2 PsΓ(Pt−sf)\psi(s)=e^{-Ks/2}\,P_{s}\sqrt{\Gamma(P_{t-s}f)} to show that ψ′(s)≥0\psi^{\prime}(s)\geq 0. ♢\diamondsuit

If instead of (x,y)(x,y) the processes start with initial distribution π0\pi_{0} the “optimal coupling” between μ0\mu_{0} and ν0\nu_{0} for the W2W_{2} distance, the previous shows that W22(μT,νT)≤e−KT W22(μ0,ν0)W_{2}^{2}(\mu_{T},\nu_{T})\leq e^{-KT}\,W_{2}^{2}(\mu_{0},\nu_{0}). As discussed in the Appendix, this result can be used to show the existence and uniqueness of the invariant measure. ♢\diamondsuit

2. hℎh-processes and functional inequalities.

We now introduce the standard notion of hh-process. Let T>0T>0 and hh be a non-negative function such that ∫ h dμT=1\int\,h\,d\mu_{T}=1. For simplicity, we assume for the moment that there exist cc and CC such that C≥h≥c>0C\geq h\geq c>0. We thus may define on the path-space up to time TT a new probability measure

In this situation, it is well known (Girsanov transform theory) that one can find a progressively measurable process usu_{s} such that

Actually, if h∈Ah\in\mathcal{A}, it is immediate to check (applying Ito formula) that

If hh is smooth we may apply Proposition 2.2 in order to get

where we have used Cauchy-Schwarz inequality for the second inequality and the Markov property for the third one. The previous inequality then extends to any hh in C1C^{1} for which the right hand side makes sense, by density. We have thus obtained the following

In the situation of (1.1), assume (H.C.K). If μ0\mu_{0} satisfies a log-Sobolev inequality with constant CLS(0)C_{LS}(0), μT\mu_{T} satisfies a log-Sobolev inequality with constant

When K=0K=0 one has to replace (1−e−KT)K\frac{(1-e^{-KT})}{K} by TT. This applies in particular to μT=P(T,x,.)\mu_{T}=P(T,x,.) since δx\delta_{x} satisfies a log-Sobolev inequality with constant equal to .

Apply the log-Sobolev inequality to μ0\mu_{0}. It furnishes (since ∫PThdμ0=1\int P_{T}hd\mu_{0}=1),

similarly as what we did in (2.9). Hence the result applying (2.9). ∎

As we recalled in the introduction a log-Sobolev inequality implies a T2T_{2} transportation inequality. It is interesting to see that one can directly obtain such an inequality for semi log-concave measures, by using the previous construction. But before to do this, just remark that the above proof using h=1+εgh=1+\varepsilon g with ∫gdμT=0\int gd\mu_{T}=0 allows us to obtain a similar result replacing the log-Sobolev inequality by a Poincaré inequality i.e.

In the situation of (1.1), assume (H.C.K). If μ0\mu_{0} satisfies a Poincaré inequality with constant CP(0)C_{P}(0), μT\mu_{T} satisfies a Poincaré inequality with constant

When K=0K=0 one has to replace (1−e−KT)K\frac{(1-e^{-KT})}{K} by TT. This applies in particular to μT=P(T,x,.)\mu_{T}=P(T,x,.) since δx\delta_{x} satisfies a Poincaré inequality with constant equal to .

Once again, the proof presented here is very different from the one based on the Γ2\Gamma_{2} calculus of Bakry-Emery which relies on the commutation property and the control of the derivative of ψ(s)=Ps(Pt−sflog⁡(Pt−sf))\psi(s)=P_{s}(P_{t-s}f\log(P_{t-s}f)) to get a local logarithmic Sobolev inequality. Note that considering rather ψ(s)=Ps((Pt−sf)2)\psi(s)=P_{s}((P_{t-s}f)^{2}) leads to a local Poincaré inequality. ♢\diamondsuit

3. Transportation inequalities.

The existence of usu_{s} and (2.7) are ensured as soon as H(hμT∣μT)<+∞H(h\mu_{T}|\mu_{T})<+\infty (see ). For our goal we do not need the explicit expression of usu_{s}.

We have then different alternatives depending on the sign of KK.

1) First in the case where K>0K>0, one has using that 2ab≤Ka2+b2/K2ab\leq Ka^{2}+b^{2}/K

so that we recover an uniform transportation inequality when η0=0\eta_{0}=0, which is moreover optimal for the invariant measure, considering logarithmic Sobolev inequality and Poincaré inequality. If μ0\mu_{0} satisfies some transportation inequality then one obtains that μT\mu_{T} satisfies a transportation inequality with constant the sum of the initial constant plus 2K\frac{2}{K}.

2) The previous simple argument has however a serious drawback in the sense that in positive curvature, μT\mu_{T} does not forget the “initial measure”. Let us see how to deal with this problem. Start once again from the first estimation, but using Itô’s formula between tt and t+εt+\varepsilon and (H.C.K)

so that we may differentiate in time to get for all positive λ\lambda

Note that this is once again optimal for the limiting measure, and captures the fact that it forgets the initial condition. When K<0K<0, we then have

Note however the presence of the additional parameter λ\lambda.

3) Let us see how a direct approach may get rid of this additional parameter, which is particularly important in negative curvature. Define

Since eKtK≤eKTK\frac{e^{Kt}}{K}\leq\frac{e^{KT}}{K} we obtain

If η0=0\eta_{0}=0, (2.14) can be improved in

Again if η0=0\eta_{0}=0, (2.16) can be improved in

The previous inequalities then extend to any non-negative hh (not necessarily bounded below nor above).

If we choose μ0=δx\mu_{0}=\delta_{x}, we have μT=P(T,x,.)\mu_{T}=P(T,x,.), 1=∫hdμT=PTh(x)1=\int hd\mu_{T}=P_{T}h(x) and so ν0=δx\nu_{0}=\delta_{x} and νT=μT\nu_{T}=\mu_{T}. Hence

In the situation of (1.1), assume (H.C.K). Then P(T,x,.)P(T,x,.) satisfies a T2T_{2} transportation inequality

If we choose ν0=μ0\nu_{0}=\mu_{0}, we may use the convexity of t↦tlog⁡tt\mapsto t\log t, i.e

In the situation of (1.1), assume (H.C.K). If μ0\mu_{0} satisfies T2T_{2} with constant C(0)C(0), then μT\mu_{T} satisfies T2T_{2} with a constant C(T)C(T) given,

and when K≤0K\leq 0, C(T)=CT+22+BT C(0)C(T)=C_{T}+\frac{\sqrt{2}}{2}+B_{T}\,C(0) with

All what precedes holds even if Υ\Upsilon is not bounded (i.e. if the process is not positive recurrent), in which case of course, K<0K<0. ♢\diamondsuit

If we choose μ0=μ\mu_{0}=\mu (assuming that Υ\Upsilon is bounded), we have to choose ν0=PThμ\nu_{0}=P_{T}h\mu hence νT=P2Th μ0\nu_{T}=P_{2T}h\,\mu_{0}. After noticing that we can slightly refine the previous bound replacing H(hμT∣μT)H(h\mu_{T}|\mu_{T}) by H(hμT∣μT)−H(PThμ0∣μ0)H(h\mu_{T}|\mu_{T})-H(P_{T}h\mu_{0}|\mu_{0}) according to (2.7), we obtain

The latter has to be compared with remark 4.9 in which shows that the inequality

If K>0K>0 we may let TT go to +∞+\infty in Proposition 2.10 and recover that μ\mu satisfies a log-Sobolev inequality with constant 2/K2/K, hence a T2T_{2} transportation inequality with constant 1/K1/K (in particular we are loosing a factor 44 in Proposition 2.18). Similarly, when T→+∞T\to+\infty, (2.22) shows that if K>0K>0, μ\mu satisfies a T2T_{2} inequality , and since μ\mu is log-concave, satisfies a log-Sobolev inequality. This scheme of proof does not require Proposition 2.2, but the (H.W.I) inequality. Unfortunately it does not furnish the optimal constant. ♢\diamondsuit

4. Transportation-Fisher Inequalities.

Let us see look now at another type of Transportation Information inequality recently introduced in , which is weaker but quite close to logarithmic Sobolev inequality (in fact equivalent under bounded curvature). We are obliged to come back to the initial inequality in (2.3) which becomes in our new situation

Replacing the pair (0,t)(0,t) by (t,t+ε)(t,t+\varepsilon) we thus have

It follows that t↦ηtt\mapsto\eta_{t} is differentiable and satisfies,

(for the second inequality, recall that (H.C.) is satisfied so that, for short, ∣∇Ps∣≤Ps∣∇∣|\nabla P_{s}|\leq P_{s}|\nabla|.) To explore (2.25) we shall use the usual trick ab≤λa2+1λ b2ab\leq\lambda a^{2}+\frac{1}{\lambda}\,b^{2} for a,b,λa,b,\lambda positive. Hence

This inequality is close to what is called a W2IW_{2}I inequality (see definition 10.4 or for examples and details on properties of WI inequality). Here we obtain a defective W2IW_{2}I inequality. However, as T→+∞T\to+\infty, we recover the true W2IW_{2}I inequality for the invariant distribution, which together with the (H.W.I) inequality allows us to recover the log-Sobolev inequality. Nevertheless, we get

Assume (H.C.K), then P(T,x,⋅)P(T,x,\cdot) satisfies a WI inequality of constant 2(1−e−AT)Aλ\frac{2(1-e^{-AT})}{A\lambda}. If we suppose moreover that μ0\mu_{0} satisfies a WI inequality with constant D(0)D(0) then μT\mu_{T} satisfies a WI inequality with constant D(T)=e−ATD(0)+2(1−e−AT)AλD(T)=e^{-AT}D(0)+\frac{2(1-e^{-AT})}{A\lambda}.

As remarked, under (H.C.K), the inequalities verified by the law μT\mu_{T} depend on the inequalities verified by the initial measure, in the range between Poincaré and logarithmic Sobolev inequality. Indeed, a logarithmic Sobolev inequality implies a WI inequality, but to get the the WI inequality for PTP_{T} we need only a WI inequality for the initial measure. As seen by the example of the Gaussian measure, which satisfies (H.C.K), no stronger inequalities can be obtained. ♢\diamondsuit

Instead of the hh-process one should consider Schrödinger bridges allowing to choose both the initial and the final time marginals. Indeed if hh is bounded it is known that one can find non-negative functions such that the measure

and the drift usu_{s} is given by us=∇ log⁡PT−sgu_{s}=\nabla\,\log P_{T-s}g. As before it is immediately checked that

For all this we refer to p.162 and section 6. Even if hh is bounded from below, we do not know whether gg inherits this property. Nevertheless, at least formally we have the relation

Unfortunately, these inequalities do not seem to give new results. ♢\diamondsuit

In almost all what we did we may replace the drift −12 ∇U-\frac{1}{2}\,\nabla U by a general (non-gradient) smooth drift bb satisfying

All the results of this section remain true in this more general situation, as far as we do not use reversibility. The only result where we used reversibility actually is (2.22). Indeed, if the initial law is ν0=PThμ\nu_{0}=P_{T}h\mu, νT=PT∗PTh μ\nu_{T}=P_{T}^{*}P_{T}h\,\mu where PT∗P_{T}^{*} denotes the μ\mu adjoint semi-group.

The only delicate point is the smoothness of PthP_{t}h and the fact that ∂t(Pth)=LPth\partial_{t}(P_{t}h)=LP_{t}h in the usual sense. This will be discussed in an even more general setting in the next section, where we shall look at more general cases with non constant diffusion coefficient.

General diffusion processes.

We shall now extend the results of the previous section to the general situation of (1.7),

First of all we have to discuss some properties of the process and the associated quantities. As we said in the introduction, we need some regularity for PtfP_{t}f at least if f∈Af\in\mathcal{A}. So there is a technical price to pay. We decided to pay this price at the level of the study of the process, rather than in deriving inequalities.

Since we assume that σ∈Cb2\sigma\in C_{b}^{2}, when (H.C.K) is fulfilled, bb satisfies

Thus, if SkS_{k} denotes the exit time from the ball B(x,k)B(x,k), and tk=t∧Skt_{k}=t\wedge S_{k} it holds

where, since σ\sigma is bounded, we have defined

It is then easily seen that one can perform similar calculations with g(t,x)=exp⁡(e−Ct∣x∣2)g(t,x)=\exp(e^{-Ct}|x|^{2}) for a large enough CC in order to kill the integrated term, i.e

It is interesting to notice that one can similarly obtain some “deviation” bound from the starting point. Indeed

so that arguing as we did in order to get (2.14) and (2.16) we obtain the existence of constants α(T,D)\alpha(T,D) and β(T,D)\beta(T,D) such that for 0≤t≤T0\leq t\leq T,

1.2. Properties of the semi-group.

Let us mimic what we did to get Proposition 2.2 i.e. apply Ito formula to get

Notice that with our assumptions, the right hand side of (3.4) is a (true) martingale, so that

Moreover, if σσ∗\sigma\sigma^{*} is positive then (3.7) implies back (H.C.K.). If we suppose moreover for some m≥2m\geq 2

The contraction in W2W_{2} distance inherited from (H.C.K.) has already been proved. The contraction in WmW_{m} distance is done exactly in the same way using once again synchronous coupling. The necessary part comes from (or more precisely section 4. in the Arxiv version 1110.3606). Let us explain the ideas of the proof. In fact, one may compute the time derivative of the Wasserstein distance: note M:N=∑i,jMijNijM:N=\sum_{i,j}M_{ij}N_{ij} when MM and NN are two matrices, then denoting νt\nu_{t} and μt\mu_{t} two solutions starting respectively from  u0\,u_{0} and μ0\mu_{0}

where if νt=∇ϕt#μt\nu_{t}=\nabla\phi_{t}\#\mu_{t},

Then the contraction property implies that at time 0 for ν0=δy\nu_{0}=\delta_{y} and μ0=δx\mu_{0}=\delta_{x}

A clever choice of ϕ\phi then enables to prove the result. ∎

Let ff be CC-Lipschitz continuous. It holds ∣f(Xtx)−f(Xty)∣≤C ∣Xtx−Xty∣|f(X_{t}^{x})-f(X_{t}^{y})|\leq C\,|X_{t}^{x}-X_{t}^{y}| so that, using (3.5), PtfP_{t}f is Lipschitz continuous with Lipschitz constant less than C e−Kt/2C\,e^{-Kt/2}.

As we said at the end of the previous section, when K>0K>0 one deduces the existence and uniqueness of an invariant probability measure μ∞\mu_{\infty}, to which μT\mu_{T} converges weakly.

The rest of this (sub)subsection is devoted to give a proof of the following: if f∈Af\in\mathcal{A} (see the introduction), then (t,x)↦Ptf(x)(t,x)\mapsto P_{t}f(x) is regular and satisfies (for t>0t>0)

The reader who takes this result as granted can skip what follows.

First if f∈Af\in\mathcal{A}, LfLf is Cc0C_{c}^{0} and we have Ptf(x)−f(x)=∫0t Ps(Lf)(x) dsP_{t}f(x)-f(x)=\int_{0}^{t}\,P_{s}(Lf)(x)\,ds. It follows that lim⁡s→01s (Pt+sf(x)−Ptf(x))=Pt(Lf)(x)\lim_{s\to 0}\frac{1}{s}\,(P_{t+s}f(x)-P_{t}f(x))=P_{t}(Lf)(x) for all xx, since v↦PvLf(x)v\mapsto P_{v}Lf(x) is continuous. So ∂tPtf=PtLf\partial_{t}P_{t}f=P_{t}Lf. The first delicate point is of course the commutation of LL and PtP_{t}. The second delicate point is the smoothness of (t,x)↦Ptf(x)(t,x)\mapsto P_{t}f(x).

This commutation property is known if σ\sigma and bb are in Cb∞C_{b}^{\infty} (see p.254-258, boundedness of derivatives is important) in which case (t,x)↦Ptf(x)(t,x)\mapsto P_{t}f(x) is actually C∞C^{\infty}. But assuming boundedness of bb and its derivatives will exclude the cases of positive KK.

We shall first show that PtfP_{t}f is a mild solution, provided bb does not grow too fast.

It follows from lemma 3.2 that if f∈Ef\in E, PtfP_{t}f is bounded by c(t)∥f∥c(t)\parallel f\parallel, hence PtP_{t} is a continuous semi-group on EE (whose range is included into CbC_{b}). To see that f∈Af\in\mathcal{A} belongs to the domain of the generator of PtP_{t}, we have to show that the convergence of lim⁡s→01s (Psf(x)−f(x))\lim_{s\to 0}\frac{1}{s}\,(P_{s}f(x)-f(x)) holds for the norm defined on EE. But

according to (3.3). Since ∣b(x)∣≤C (1+∣x∣k)|b(x)|\leq C\,(1+|x|^{k}), convergence holds for the norm on EE. The proof is completed. ∎

If the coefficients are C∞C^{\infty}, and ∂t−L\partial_{t}-L is hypo-elliptic (for instance if LL is uniformly elliptic) it follows that x↦Ptf∈C∞x\mapsto P_{t}f\in C^{\infty} and that the last equalities hold in the usual sense.

If we do not want to assume too much regularity on the coefficients, we have first to assume that LL is uniformly elliptic and call upon P.D.E. theory. If what follows is certainly well known by specialists, we include the argument.

Let f∈Af\in\mathcal{A}. For kk large enough, Bk=B(0,k)B_{k}=B(0,k) contains the support of ff. Consider the parabolic equation

Since f=0f=0 on ∂Bk\partial B_{k} this makes sense. If LL is uniformly elliptic, it is known that there exists an unique solution uku_{k} in C1,2((0,T)×Bk)C^{1,2}((0,T)\times B_{k}) of (3.10), and that this solution is represented as

where SkxS^{x}_{k} denotes the exit time from BkB_{k} of XtxX_{t}^{x}. For all this see , in particular Theorem 5.2. p.147. It follows in particular that for all kk, ∥uk∥∞≤∥f∥∞\parallel u_{k}\parallel_{\infty}\leq\parallel f\parallel_{\infty} and that for all (t,x)∈(0,T)×Bj(t,x)\in(0,T)\times B_{j}, uk(t,x)→PT−tf(x)u_{k}(t,x)\to P_{T-t}f(x) as k→+∞k\to+\infty since Skx→+∞S_{k}^{x}\to+\infty. Now let jj be fixed, and look at k>jk>j. The parabolic Schauder estimate tells us that there exists a constant CjC_{j} depending on jj, the ellipticity constant and the C2(Bj+1)C^{2}(B_{j+1}) norms of σ\sigma and bb such that

where Ck,12C^{k,\frac{1}{2}} is the set of CkC^{k} functions with 12\frac{1}{2}-Hölder kthk^{th} derivatives. Arzela-Ascoli theorem tells us that a subsequence of uku_{k} converges in C2((0,T)×Bj)C^{2}((0,T)\times B_{j}), and since the limit is PT−tfP_{T-t}f, that the latter is C2C^{2}.

If LL is not uniformly elliptic, we may approximate it by Lε=L+12 εΔL_{\varepsilon}=L+\frac{1}{2}\,\varepsilon\Delta for ε→0\varepsilon\to 0. If σ∈Cb2\sigma\in C_{b}^{2}, the diffusion matrix (field) aε=σε σε∗+ε Ida_{\varepsilon}=\sigma_{\varepsilon}\,\sigma^{*}_{\varepsilon}+\varepsilon\,Id is Cb2C_{b}^{2} and uniformly elliptic). It is known that its square root (in the sense of symmetric matrices) aε1/2a^{1/2}_{\varepsilon} is bounded and Cb2C_{b}^{2} too. In addition aε1/2→a1/2a^{1/2}_{\varepsilon}\to a^{1/2} as ε→0\varepsilon\to 0 the convergence taking place in Cb1C_{b}^{1}. In particular, (H.C.K(ε\varepsilon)) holds for the new diffusion process with a constant K(ε)K(\varepsilon) going to KK as ε→0\varepsilon\to 0, provided (H.C.K) is satisfied for σ=a12\sigma=a^{\frac{1}{2}}.

It follows that the convergence of time marginals holds in W2W_{2} Wasserstein distance, hence in the weak topology.

The conclusion of the previous discussion is the following: if a1/2=σ∈Cb2a^{1/2}=\sigma\in C_{b}^{2}, we may assume that LL is uniformly elliptic as far as the bounds we get do not depend on the ellipticity constant, and then go to the limit. We shall use this trick in the sequel.

2. Commutation property with the gradient

Once again, we will see how synchronous coupling or a contraction in Wasserstein distance provides the commutation property.

where the last step is done by using Theorem 3.6. Provided we know that ∇Ptf\nabla P_{t}f exists, we have thus obtained a weaker form of Proposition 2.2,

Assume (R) and (H.C.K) or the weaker contraction property (3.7). Let f∈Af\in\mathcal{A}. If ∇Ptf\nabla P_{t}f exists (which is true except possibly for (R5)), it holds

Notice that, contrary to the Bakry-Emery bounded curvature case, the previous commutation property holds with the usual gradient and not with the natural one i.e. Γ12\Gamma^{\frac{1}{2}}.

If Proposition 2.2 allowed us to obtain logarithmic Sobolev inequalities, the weaker Proposition 3.11 will allow us to obtain a weaker inequality, namely a Poincaré inequality.

It is worth mentioning here the following alternate proof of the commutation property, starting from Wasserstein contraction, as derived in the recent paper following our suggestion, i.e. using Kantorovitch-Rubinstein duality we have for all bounded Lipschitz ϕ\phi denoting the inf convolution operator Qtϕ(x)=inf⁡y{ϕ(y)+∣x−y∣22t}Q_{t}\phi(x)=\inf_{y}\{\phi(y)+\frac{|x-y|^{2}}{2t}\} and initial measure μ0\mu_{0} and ν0\nu_{0}

Choose now μ0=δx\mu_{0}=\delta_{x}, ν0=δy\nu_{0}=\delta_{y} to get for all yy

which by homogeneity of the inf-convolution operator gives

This assertion is in fact stronger than the gradient commutation property which can be deduced by using the fact that the inf-convolution operator is the Hopf-Lax solution of the Hamilton-Jacobi equation. ♢\diamondsuit

3. h-processes and functional inequalities.

We now introduce the corresponding hh-process. Let T>0T>0 and h>0h>0 be such that

We thus may define on the path-space up to time TT a new probability measure

For simplicity, we assume in what follows that there exist cc and CC such that C≥h≥c>0C\geq h\geq c>0. In this situation, using again Girsanov transform theory, we know that we can find a progressively measurable process usu_{s} such that

If h∈Ah\in\mathcal{A} we may apply Proposition 3.11 in order to get (recall that h≥ch\geq c)

where we have used the Markov property for the second inequality.

Now let g∈Cc∞g\in C_{c}^{\infty} be such that ∫gdμT=∫PTg dμ0=0\int gd\mu_{T}=\int P_{T}g\,d\mu_{0}=0 and choose h=1+ηg∈Ah=1+\eta g\in\mathcal{A} so that ∫PThdμ0=1\int P_{T}hd\mu_{0}=1 and h>c>0h>c>0 for η\eta small enough. Actually we will let η\eta go to so that in the limit c=1c=1. Standard manipulations thus yield

We can eventually use first the density of Cc∞C_{c}^{\infty} and then the trick we formerly described in order to relax the uniform ellipticity assumption (recall that M=sup⁡y sup⁡∣u∣=1∣⟨u,a(y)u⟩∣M=\sup_{y}\,\sup_{|u|=1}|\langle u,a(y)u\rangle| hence only depends on aa too).

Arguing as for Proposition 2.10 we have obtained

Assume that (R) and (H.C.K) are satisfied. Let M=∥∣σ∣2∥∞M=\parallel|\sigma|^{2}\parallel_{\infty}. If μ0\mu_{0} satisfies a Poincaré inequality with constant CP(0)C_{P}(0) then μT\mu_{T} satisfies a Poincaré inequality with constant

This applies in particular to P(T,x,.)P(T,x,.) with CP(0)=0C_{P}(0)=0.

Contrary to the log-Sobolev inequality, the Poincaré inequality does not furnish a transportation inequality, so we shall try to adapt what we did in subsection 2.3.

4. Transportation inequalities.

We may thus conclude as in the previous section

Assume that (R) and (H.C.K) are satisfied. Let M=∥∣σ∣2∥∞M=\parallel|\sigma|^{2}\parallel_{\infty}. The conclusions of Proposition 2.18 and Proposition 2.19 are still true, replacing CTC_{T} by MCTMC_{T}

Actually, when (R5) holds, we have proven this result for h∈Ah\in\mathcal{A} and L+ε2 ΔL+\frac{\varepsilon}{2}\,\Delta. But as we have seen, μT(ε)→μT\mu_{T}(\varepsilon)\to\mu_{T} in W2W_{2} distance, so that if hh is bounded the same holds for zε hμT(ε)z_{\varepsilon}\,h\mu_{T}(\varepsilon) to hμTh\mu_{T} (zεz_{\varepsilon} being a normalization constant). Finally if a T2T_{2} inequality holds for all h∈Ah\in\mathcal{A} it extends to all hh using density and the fact that W2(ν,μ)≤lim inf⁡W2(νn,μ)W_{2}(\nu,\mu)\leq\liminf W_{2}(\nu_{n},\mu) if νn\nu_{n} weakly converges to ν\nu.

Of course a T2T_{2} inequality implies a Poincaré inequality, but the constant in Proposition 3.17 is better (in addition we only require that μ0\mu_{0} satisfies a Poincaré inequality).

One of the renowned consequence of such inequalities is the concentration of measure phenomenon for μT\mu_{T}. In particular, under the assumptions of Proposition 3.20, μT\mu_{T} satisfies a gaussian type concentration property. In particular ∣XTx∣2|X_{T}^{x}|^{2} has some exponential moment, fact we have already shown in lemma 3.2. But this integrability does not reflect all the strength of the T2T_{2} inequality whose tensorization property is particularly useful for statistical purposes.

When LL is uniformly elliptic, this concentration property follows from gaussian estimates for the transition kernel. Here we obtain much more explicit constants (even if they are certainly far from optimality) which do not depend on the ellipticity constant. ♢\diamondsuit

Assume that LL is uniformly elliptic, i.e.

According to proposition 5.4.1, this is equivalent to the CD(K/2,∞)CD(K/2,\infty) condition provided M=eM=e hence when σ\sigma is constant times the identity. In the non constant diffusion case, our condition (H.C.K) seems to be really different from the Bakry-Emery curvature condition. ♢\diamondsuit

5. An hypoelliptic example : kinetic Fokker-Planck equation

We present in this section an application of the techniques developed here in an hypoelliptic example where the Bakry-Emery curvature is negative and where (H.C.K.) may not be satisfied also. Let (xt,vt)(x_{t},v_{t}) be the solution of the following SDE

also called stochastic Hamiltonian system. The long time behavior study of such a system has been considered for a long time and have been tackled by different techniques, see for example: hypocoercivity by Villani or Lyapunov function technique by Bakry&\&al . However, due to its hight degeneracy, the Bakry-Emery curvature is −∞-\infty so that we may not apply the Γ2\Gamma_{2} technique. Remark also that the (H.C.K.) condition reads for all (x,v)(x,v) and (y,w)(y,w)

Let us first remark that if ∇V\nabla V is Lipshitz continuous, (H.C.K) is verified for some negative KK and using synchronous coupling, one may remark that we are in the same situation than in Section2 so that we get that for some negative KK the gradient commutation property holds

and thus the logarithmic Sobolev inequality holds for Pt((x,v),⋅)P_{t}((x,v),\cdot). Let us remark once again that those properties are written with the usual gradient and not the Carré-du-Champ operator Γ(f)=∣∇vf∣2\Gamma(f)=|\nabla_{v}f|^{2}.

One may then wonder if it is possible to get the gradient commutation property with K>0K>0. In fact, using synchronous coupling and Itô’s formula applied to the function N((x,v),(y,w))=a∣x−y∣2+b⟨x−y,v−w⟩+∣v−w∣2N((x,v),(y,w))=a|x-y|^{2}+b\langle x-y,v-w\rangle+|v-w|^{2}, following , we get that if V(x)=∣x∣2+W(x)V(x)=|x|^{2}+W(x) where ∇W\nabla W is δ\delta-Lipschitz with δ\delta sufficiently small there exists a,ba,b and K>0K>0 such that HH is equivalent to the euclidean norm and

so that we get as in Section 2 the commutation property for some K>0K>0 and A>1A>1

and thus a Logarithmic Sobolev inequality holds uniformly in time.

It is not hard to extend the result of this simplified setting to the case where the Brownian motion in the velocity has a diffusion coefficient which is bounded and LL-Lipschitz. We may then obtain a weaker gradient commutation property

and local Poincaré type inequality or Transportation information inequality like in Propositions 3.17 or 3.20, and if LL is sufficiently small uniform in time version of these inequalities (using functional NN).

6. Interpolation of the gradient commutation property and local Beckner inequality

We have seen here that we cannot recover a logarithmic Sobolev inequality by our technique when (H.C.K.) is in force. Remember however that we have introduced the stronger (H.C.K.m) condition which implies a contraction in Wasserstein distance WmW_{m}. It is then not hard to deduce some interpolation of the gradient commutation property

Assume (R) and (H.C.K.m) or the weaker contraction property (3.8). Let f∈Af\in{\mathcal{A}}, if ∇Ptf\nabla P_{t}f exists, it holds

Remark once again that this property does hold even if the diffusion coefficient is degenerate, so that variations of the hypoelliptic example of the previous subsection with a diffusion coefficient in the velocity enters into this framework. This contraction property may thus lead to a reinforcement of the Poincaré inequality to a Beckner inequality.

Assume (R) and (H.C.K.m) or the weaker (3.24). Let M=∥∣σ∣2∥∞M=\||\sigma|^{2}\|_{\infty}. Then for all nice ff, we have the following Beckner inequality

The proof unfortunately does not rely on the hh-process introduced previously but on the Γ2\Gamma_{2} type proof. Denote p=2mm+2p=\frac{2m}{m+2}. By (3.24) and Hölder inequality, for all nice non negativeff

7. Convergence to equilibrium in positive curvature.

Still in the uniform elliptic case, assume that K>0K>0. We already mentioned that in this case μT\mu_{T} weakly converges to the unique invariant probability measure μ∞\mu_{\infty} (which exists). In particular , for all smooth gg (say Cb2C_{b}^{2}), VarμT(g)→Varμ∞(g)\textrm{Var}_{\mu_{T}}(g)\to\textrm{Var}_{\mu_{\infty}}(g) as well as ∫∣∇g∣2 dμT→∫∣∇g∣2 dμ∞\int|\nabla g|^{2}\,d\mu_{T}\to\int|\nabla g|^{2}\,d\mu_{\infty}. We deduce that

As we said, this result is not captured by the Γ2\Gamma_{2} theory.

But we can obtain general convergence results, even in the non uniformly elliptic case. Indeed recall that in full generality

Assume that σ\sigma is bounded. Then if (H.C.K) holds for some K>0K>0, defining MM as before, there exists an unique invariant probability measure μ∞\mu_{\infty} and for all nice enough function gg ,

In addition if μ∞\mu_{\infty} is symmetric (i.e. ∫fLgdμ∞=∫gLfdμ∞\int fLgd\mu_{\infty}=\int gLfd\mu_{\infty}), it holds

Remark once again that what is used here is the weak gradient commutation property which is a consequence of (H.C.K.). The last part of the theorem follows from a result in recalled in the Appendix, lemma A.1 (2). Of course, unless we explicitly know the invariant measure, it is not easy to see wether μ∞\mu_{\infty} is symmetric or not.

We may further extends the previous argument to the entropic convergence to equilibrium. Let us suppose that there exists an unique invariant measure μ∞\mu_{\infty}.

Assume that σ\sigma is bounded and that the gradient commutation property

holds for some positive KK. Then for all nice positive function ff (defining MM as before)

The proof is as for the L2L_{2} decay quite standard. Indeed,

One of the important point here is that we do not suppose any non-degeneracy on the diffusion coefficient, so that the result applies to the kinetic Fokker-Planck equation. It then provides an alternative to the approach by Villani , where he obtained such kind of convergence by completely different techniques with assumptions quite similar to the ones described in Section 3.5. One may then complete the approach by regularization of the Fisher Information in small time to obtain an entropic decay controlled by the initial entropy, see or . ♢\diamondsuit

Let us point out that even in the symmetric case, such a control is not sufficient to recover a logarithmic Sobolev inequality as the analog of lemma A.1 is no more valid for the entropy. Remark however that we have shown in section 2 how to recover a logarithmic Sobolev inequality for PtP_{t} using the strong commutation gradient property (3.30). If K>0K>0, we may then let tt goes to infinity to recover a logarithmic Sobolev inequality for the invariant measure. It may be for example be used in the context of kinetic Fokker-Planck equation with non gradient coefficient, for which the invariant measure is unknown. ♢\diamondsuit

Let us consider, as in the Poincaré case via (H.C.K.) condition, a particular class of test function gg such that g≥ε>0g\geq\varepsilon>0, so that Ptg≥εP_{t}g\geq\varepsilon. We then see adapting the preceding proof that a weak commutation of gradient property

obtained for example under (H.C.K.) condition implies that

Non homogeneous diffusion processes.

In the authors extended the Γ2\Gamma_{2} theory to time dependent coefficients (non homogeneous diffusions). Considering the Ito system

If ff only depends on xx, the proof of proposition 2.2 (resp. 3.11) is unchanged using the process starting from (0,x)(0,x) and (0,y)(0,y) and replacing KtKt by K(t)K(t). To obtain the analogue of proposition 2.10 and proposition 3.17, it suffices to remark that σ∇\sigma\nabla is equal to ∇x\nabla_{x}, and use what precedes for hh depending on xx only. ∎

For the transportation inequality we have to slightly modify the method in subsection 2.3. With the notations therein, (2.3) has become,

so that, as in the previous section we have to come back to

Using as usual (ab)12≤λa+1λb(ab)^{\frac{1}{2}}\leq\lambda a+\frac{1}{\lambda}b we obtain (see the details of the derivation in the previous section) that for all increasing function λ(t)\lambda(t)

from which we deduce, provided we choose K(0)=λ(0)=0K(0)=\lambda(0)=0,

If μ0\mu_{0} satisfies a T2T_{2} inequality with constant CT(0)C_{T}(0), then

The best choice of λ\lambda is not clear. If K′(t)K^{\prime}(t) is not positive on the whole [0,T][0,T], it seams that taking λ(T)=λT\lambda(T)=\lambda T for some λ>0\lambda>0 is enough. If K′(t)>0K^{\prime}(t)>0 for all tt (but not necessarily bounded from below by a positive constant), λ(t)=λK(t)\lambda(t)=\lambda K(t) seems to be natural.

Assume that K(t)→+∞K(t)\to+\infty as t→+∞t\to+\infty and that C=∫0+∞e−K(s) ds<+∞C=\int_{0}^{+\infty}e^{-K(s)}\,ds<+\infty. Then, for all tt, P(t,x,.)P(t,x,.) (the distribution of the process starting from xx and t0=0t_{0}=0) satisfies a Poincaré inequality (and a log-Sobolev inequality when σ=Id\sigma=Id) with a constant bounded by MCMC (or 2C2C). The family (P(t,x,.))t>0(P(t,x,.))_{t>0} is then tight, but we do not know whether it is weakly convergent or not. Nevertheless any weak limit satisfies the same functional inequality. When σ=Id\sigma=Id we know that ∣Xtx−Xt∣≤e−K(t)∣x−X0∣|X_{t}^{x}-X_{t}|\leq e^{-K(t)}|x-X_{0}| for any initial random variable X0X_{0}. It follows that if a sequence P(tk,x,.)P(t_{k},x,.) is weakly convergent to some μ\mu, the sequence μtk\mu_{t_{k}} weakly converges to the same limit. In particular if we consider σ=Id\sigma=Id, b(t,x)=−12 (∇U(x)+K′(t) x)b(t,x)=-\frac{1}{2}\,(\nabla U(x)+K^{\prime}(t)\,x), for some convex potential UU, (H.C.K(t)) is satisfied, so that any weak limit satisfies a log-Sobolev inequality. If dμ=e−Udxd\mu=e^{-U}dx does not satisfy a log-Sobolev inequality, it cannot be a weak limit, even if K′(t)→0K^{\prime}(t)\to 0. In this situation one should expect that the “perturbation” of ∇U\nabla U being smaller and smaller when tt growths, the convergence to μ\mu will still hold. This is not the case. ♢\diamondsuit

2. Application to some non-linear diffusions.

We shall now discuss an example that does partly enter the framework of the beginning of this section.

Following consider the following non-linear stochastic differential equation

This is a non-linear diffusion of Mc Kean-Vlasov type modeling, for instance, granular media. We refer to the introduction of for details and motivations. One can approximate the solution of (4.6) by the first coordinate of a linear large particle system with mean field interactions. This is what is done in to study the long time behavior of XtX_{t}.

Let we see how to apply what we have just done. First, under some conditions on VV and WW (we later shall give some of them) existence and weak uniqueness of (4.6) are ensured, provided the initial law admits some big enough polynomial moment. This will imply for all xx, the existence and uniqueness of qtxq_{t}^{x} solution of (4.7) with initial condition δx\delta_{x}. As usual for these non-linear equations, if we consider the linear time inhomogeneous S.D.E.

the pathwise unique solution (up to explosion) Z.x,xZ_{.}^{x,x} is shown to satisfy (4.6) (i.e. L(Ztx,x)=qtx\mathcal{L}(Z^{x,x}_{t})=q_{t}^{x}) so that it coincides with XtxX^{x}_{t}. So, once qtxq_{t}^{x} and qtyq_{t}^{y} are built, we may build our synchronous coupling (Xtx,Xty)(X_{t}^{x},X_{t}^{y}) as before. Now introduce an independent copy (Xˉtx,Xˉty)(\bar{X}_{t}^{x},\bar{X}_{t}^{y}) of (Xtx,Xty)(X_{t}^{x},X_{t}^{y}).

If we assume in addition (as usual) that W(−x)=W(x)W(-x)=W(x), and remember that Xˉ\bar{X} is a copy of XX, it is still equal to

VV, WW and their first two derivatives have at most polynomial growth of order mm and W(−x)=W(x)W(-x)=W(x),

VV satisfies (H.C.KVK_{V}) and WW satisfies (H.C.KWK_{W}).

Let a=max⁡(m(m+3),2m2)a=\max(m(m+3),2m^{2}). If μ0\mu_{0} and ν0\nu_{0} have a polynomial moment of order aa, there exist an unique solution of (4.6) and an unique solution of (4.7) among the set of probability flows having having a polynomial moment of order aa with initial condition μ0\mu_{0} or ν0\nu_{0}. Furthermore

If V=0V=0 and ∫xμ0(dx)=∫xν0(dx)\int x\mu_{0}(dx)=\int x\nu_{0}(dx) then

V=0V=0, ∫xμ0(dx)=∫xν0(dx)\int x\mu_{0}(dx)=\int x\nu_{0}(dx) and KW>0K_{W}>0.

The moment condition ensuring existence and uniqueness is described in section 2.

Of course we may replace the initial δx\delta_{x} and δy\delta_{y} by probability distributions μ0\mu_{0} and ν0\nu_{0} satisfying the required moment conditions. This immediately furnishes the first assertion about the upper bound for the Wasserstein distance.

If V=0V=0 it is easily seen that ∫xqtμ0(x)dx=∫xμ0(dx)\int xq_{t}^{\mu_{0}}(x)dx=\int x\mu_{0}(dx) for all t>0t>0, hence

provided the same holds at time . This furnishes the second assertion for the upper bound.

Finally the convergence under strict positivity of our new “curvature” condition ensures the existence of the limiting measure μ∞\mu_{\infty}. To see that μ∞=q∞(x)dx\mu_{\infty}=q^{\infty}(x)dx is actually invariant, one can for instance use the following trick: first consider the solution qt∞q_{t}^{\infty} of (4.7) with initial condition q∞q^{\infty}. Similar bounds for the Markov non homogeneous process Z.q∞,yZ_{.}^{q^{\infty},y} (when we replace qtxq_{t}^{x} by qt∞q_{t}^{\infty}) are obtained applying the results of the beginning of this section. Hence the law of ZTq∞,q∞Z_{T}^{q^{\infty},q^{\infty}} (which is exactly μT\mu_{T} starting with μ∞\mu_{\infty} as we explained before) converges to some limiting measure μ∞μ∞\mu_{\infty}^{\mu_{\infty}} which in turn is equal to μ∞\mu_{\infty} and is invariant for Z.q∞,yZ_{.}^{q_{\infty},y}. This achieves the proof. ∎

The proof of the above result is new and direct, while the result is mainly contained in using particle approximation. Notice that in the Γ2\Gamma_{2} approach is developed for the non homogeneous Markov diffusion Z.Z_{.} and not for X.X_{.}. Also notice that some direct study of the decay to equilibrium in W2W_{2} distance for granular media is done in . ♢\diamondsuit

for which μ∞(dx)=q∞(x)dx\mu_{\infty}(dx)=q^{\infty}(x)dx the invariant probability measure. Using the results in section 2 we thus have

In the situation of Theorem 4.9, if H’1 or H’2 are satisfied, μ∞\mu_{\infty} satisfies a log-Sobolev inequality with constant CLS=2/KC_{LS}=2/K or CLS=2/KWC_{LS}=2/K_{W}.

All what we have done extends to more general Mc Kean-Vlasov equations, with a diffusion coefficient σ\sigma and a drift bb satisfying hypothesis (R). In particular, positive curvature (in the sense of (H.C.K)) will also imply existence of and convergence to an invariant probability measure. The only difference is that we have to replace log-Sobolev inequality by Poincaré inequality in the latter proposition. Let us explain quickly what kind of model we may consider. We do not aim to be optimal, but will provide a flavor of the results on contraction with some non constant diffusion term. We will not focus also on the existence of solution of such equation. Let XtxX_{t}^{x} be solution of

Let us suppose H1 and H2, that κ\kappa is ll-Lipschitz and that

Suppose moreover that KV−r(1+4l2)+min⁡(KW,0)>0K_{V}-r(1+4l^{2})+\min(K_{W},0)>0, then there exists an unique invariant distribution to (4.12), the convergence to μ∞\mu_{\infty} in W2W_{2} Wasserstein distance being exponential as above.

The proof follows the same line as before except that in the Itô’s formula, there is the diffusion part which comes into play for which we use the Lipschitz condition of the theorem. Note that Bolley&\&al have considered the case of a kinetic McKean-Vlasov equation, but with a constant diffusion coefficient in speed. As before, we may obtain some functional inequality for the invariant distribution as in Prop. 4.11 but we have to replace log-Sobolev inequality by Poincaré inequality .

Extensions to some non uniformly convex potentials.

Let us come back to (1.1), and assume that Υ\Upsilon is bounded. We shall extend (H.C.K) to more general situations. The first natural extension is to replace the squared distance by some other convex functional of the distance. More precisely.

φ\varphi is increasing and convex, with φ(0)=0\varphi(0)=0 and φ(1)=1\varphi(1)=1,

a↦φ(a)/aa\mapsto\varphi(a)/a is non decreasing,

there exist a positive function ψ\psi such that for all a>0a>0 and all λ>0\lambda>0, φ−1(λa)≤ψ(λ) φ−1(a)\varphi^{-1}(\lambda a)\leq\psi(\lambda)\,\varphi^{-1}(a), where φ−1\varphi^{-1} denotes the inverse (reciprocal) function of φ\varphi.

Let φ∈C\varphi\in\mathcal{C}. We shall say that (H.φ\varphi.K) is satisfied for some K>0K>0 if for all (x,y)(x,y),

On one hand, since K>0K>0 and φ≥0\varphi\geq 0, (H.φ\varphi.K) implies that UU is convex. On the other hand, if (H.φ\varphi.K) is satisfied, since UU is smooth, φ(a)/a\varphi(a)/a is necessarily bounded near the origin since lim sup⁡a→0(φ(a)/a)≤inf⁡∣Hess(U)∣\limsup_{a\to 0}(\varphi(a)/a)\leq\inf|Hess(U)|. Here of course if φ∈C\varphi\in\mathcal{C} the latter is automatically satisfied.

If φ(a)=a\varphi(a)=a this is nothing else but (H.C.K). If φ(a)/a→+∞\varphi(a)/a\to+\infty we shall say that UU is super-convex. This terminology is justified by the example below.

Let U(x)=(∣x∣2)βU(x)=(|x|^{2})^{\beta} for some β>1\beta>1. We shall see that (H.φ\varphi.K) is satisfied for φ(a)=aβ\varphi(a)=a^{\beta} and some KK we shall estimate.

We start with the one dimensional case. In this case

If sign(x)=sign(y)sign(x)=sign(y), we may assume that ∣x∣≥∣y∣|x|\geq|y|, write ∣x∣=u+∣y∣|x|=u+|y| for u≥0u\geq 0 and remark that if 2β−1≥12\beta-1\geq 1,

If sign(x)=−sign(y)sign(x)=-sign(y), we have, using the convexity of x↦∣x∣2β−1x\mapsto|x|^{2\beta-1},

Since β>1\beta>1, we may choose Kβ=2β 22−2βK_{\beta}=2\beta\,2^{2-2\beta}.

If α≥0\alpha\geq 0, we write again ∣x∣=∣y∣+a|x|=|y|+a with a≥0a\geq 0. Thus , since 0≤1−α≤10\leq 1-\alpha\leq 1,

It follows, since β≥1\beta\geq 1 and γ2≤1\gamma^{2}\leq 1,

If α<0\alpha<0, since ∣α∣≤1|\alpha|\leq 1, it holds

Let U(x)=(∣x∣2)βU(x)=(|x|^{2})^{\beta} for some β>1\beta>1. Then (H.φ\varphi.K) is satisfied for φ(a)=aβ\varphi(a)=a^{\beta} and Kβ≥2β 23−3βK_{\beta}\geq 2\beta\,2^{3-3\beta}. If n=1n=1 we have the better bound Kβ≥2β 22−2βK_{\beta}\geq 2\beta\,2^{2-2\beta}. ♢\diamondsuit

If φ∈C\varphi\in\mathcal{C}, for all a≥0a\geq 0 and all ε>0\varepsilon>0, it holds

Hence (H.φ\varphi.K) implies the following condition

The latter appears in the study of the granular medium equation in (condition(6)) for power functions φ\varphi. This formulation will be the interesting one. It can be extended in

(H.φ\varphi.K) implies (H.α\alpha.K) with the same KK and α(ε)=φ(ε)/ε\alpha(\varepsilon)=\varphi(\varepsilon)/\varepsilon. In this definition we do not need that a↦aα(a)a\mapsto a\alpha(a) is convex.

Now we shall see how to use (H.φ\varphi.K).

This subsection contains first results which are not really convincing, but have to be tested. If we want to control the gradient ∇Ptf\nabla P_{t}f, we may write for t>ut>u,

Denoting ηt=∣Xtx−Xty∣2\eta_{t}=|X_{t}^{x}-X_{t}^{y}|^{2}, we thus have ηt′≤−K φ(ηt)\eta^{\prime}_{t}\leq-K\,\varphi(\eta_{t}). If φ(a)=aβ\varphi(a)=a^{\beta}, this yields

This result (even after taking expectation) is not really satisfactory. Indeed, first we do not obtain any better control for ∇Ptf\nabla P_{t}f than the one for a general convex potential (in particular we do not obtain a rate of convergence to ). In second place, the decay to of the Wasserstein distance we obtain is desperately slow, while we expected an exponential decay (which we know to hold true for U(x)=∣x∣2βU(x)=|x|^{2\beta} for β≥1\beta\geq 1). Notice however that we recover the exponential decay we obtained previously when β→1\beta\to 1.

If instead of (H.φ\varphi.K) we use (H.α\alpha.K), it is not difficult to show that

If α(ε)=εβ−1\alpha(\varepsilon)=\varepsilon^{\beta-1}, choosing ε=η0 t−θ\varepsilon=\eta_{0}\,t^{-\theta} for some θ<β−1\theta<\beta-1, we get

The method can be extended to the Mc Kean-Vlasov situation studied in subsection 4.2 and allows us to recover (up to the constants) Theorem 4.1 in without the help of a particle approximation. However, better results in this situation are obtained in . ♢\diamondsuit

Mimicking subsection 2.3, in particular (2.3), do we obtain more interesting results ? Using the notation therein we have

so that, if vt=∫0t ηs dsv_{t}=\int_{0}^{t}\,\eta_{s}\,ds,

If φ(a)=aβ\varphi(a)=a^{\beta}, we thus obtain

This result is certainly not fully satisfactory too. On one hand, we get a less explosive bound in time (recall that in the general convex case the bound growths like TT), but on the other hand the relative entropy appears to a power less than 1. In particular such an inequality does not imply a Poincaré inequality (which is obtained for entropies going to ), but furnishes nice concentration properties (obtained for large entropies via Marton’s argument).

2. An improvement of Bakry-Emery criterion.

As we remarked at this end of section 2 we may come back to the initial inequality in (2.3) which becomes in our new situation

(here again (H.C.) is satisfied so that, for short, ∣∇Ps∣≤Ps∣∇∣|\nabla P_{s}|\leq P_{s}|\nabla|.) To explore (5.13) we shall use both the remark 5.5 and the usual trick ab≤λa2+1λ b2ab\leq\lambda a^{2}+\frac{1}{\lambda}\,b^{2} for a,b,λa,b,\lambda positive. Hence

We deduce, denoting A=K φ(ε)ε−2λA=K\,\frac{\varphi(\varepsilon)}{\varepsilon}-2\lambda,

Choose λ=(1/4) K (φ(ε)/ε)\lambda=(1/4)\,K\,(\varphi(\varepsilon)/\varepsilon) so that A=(1/2) K (φ(ε)/ε)>0A=(1/2)\,K\,(\varphi(\varepsilon)/\varepsilon)>0. ηT\eta_{T} is thus bounded in time, but the bound is not tractable except for T=+∞T=+\infty (starting with μ0=μ\mu_{0}=\mu) or if η0=0\eta_{0}=0. In both cases we have obtained

It remains to optimize in ε\varepsilon. In full generality we choose ε\varepsilon such that both terms in the sum of (5.15) are equal (we know that we are loosing a factor less than 22). Remark that we do not use the explicit form of φ\varphi, i.e. we may replace (H.φ\varphi.K) by (H.α\alpha.K) in what we did previously. We have thus obtained

Assume that UU satisfies (H.α\alpha.K) for K>0K>0. Let FF be the inverse (reciprocal) function of ε↦ε α2(ε)\varepsilon\mapsto\varepsilon\,\alpha^{2}(\varepsilon). Denote μT=P(T,x,.)\mu_{T}=P(T,x,.) and μ∞(dx)=μ(dx)=e−U(x) dx\mu_{\infty}(dx)=\mu(dx)=e^{-U(x)}\,dx. Then for all 0<T≤+∞0<T\leq+\infty, μT\mu_{T} satisfies for all nice hh,

where IT(h)=∫∣∇h∣2h dμTI_{T}(h)=\int\frac{|\nabla h|^{2}}{h}\,d\mu_{T} is the Fisher information of hh.

When FF is equal to identity, such an inequality is called a W2IW_{2}I inequality (see definition 10.4). Here we obtained a weak form of W2IW_{2}I inequality (which is clear on (5.15)) in the spirit of the weak Poincaré or the weak log-Sobolev inequalities.

In particular, using (H.W.I), since (H.C.0) is satisfied we obtain

Under the hypotheses of proposition 5.16, μ\mu satisfies the inequality

Weak logarithmic Sobolev inequalities were introduced and studied in . Actually, we are not exactly here in the situation of because we wrote the previous inequality in terms of a density of probability. Let h=f2/∫f2dμh=f^{2}/\int f^{2}d\mu. We deduce from the previous corollary

so that if F(λ a)≤θ(λ) F(a)F(\lambda\,a)\leq\theta(\lambda)\,F(a),

In this sub(sub)section we assume that φ(a)=aβ\varphi(a)=a^{\beta} for some β≥1\beta\geq 1, so that F(a)=a12β−1F(a)=a^{\frac{1}{2\beta-1}}. We thus have

Recall first that if g≥0g\geq 0, then Varμ(g)≤Entμ(g)\textrm{Var}_{\mu}(g)\leq\textrm{Ent}_{\mu}(g) (see e.g. (2.6)). Next recall the following: defining mμ(g)m_{\mu}(g) as a median of gg, we have

We may decompose f−mμ(f)=(f−mμ(f))+−(f−mμ(f))−=g+−g−f-m_{\mu}(f)=(f-m_{\mu}(f))_{+}-(f-m_{\mu}(f))_{-}=g_{+}-g_{-} so that both g+g_{+} and g−g_{-} are non negative with median equal to . In addition, if ff is Lipschitz, so are g+g_{+} and g−g_{-}, ∇f=∇g++∇g−\nabla f=\nabla g_{+}+\nabla g_{-}, and the product of both vanishes. Hence

similarly for g−g_{-}. We have thus obtained

Assume that UU satisfies (H.φ\varphi.K) for K>0K>0 and φ(a)=aβ\varphi(a)=a^{\beta}, for β≥1\beta\geq 1. Then, μ\mu satisfies both a Poincaré inequality with

This theorem applies in particular to U(x)=∣x∣2βU(x)=|x|^{2\beta} for β≥1\beta\geq 1, according to proposition 5.4. The fact that μ\mu satisfies a log-Sobolev inequality in this situation is well known, but here we obtain an explicit (though not really cute) expression for the constant that only depends on β\beta and not on the dimension nn. Unfortunately, in this particular situation, our bounds are not optimal. Indeed, spherically symmetric log-concave probability measures are now well understood.

For the Poincaré constant, it was shown by Bobkov that

It is an (easy) exercise to see that Varμ(x)=Γ((n+2)/2β)/Γ(n/2β)\textrm{Var}_{\mu}(x)=\Gamma((n+2)/2\beta)/\Gamma(n/2\beta), so that CP(μ)≤c(β) n1β−1C_{P}(\mu)\leq c(\beta)\,n^{\frac{1}{\beta}-1} which goes to as n→+∞n\to+\infty. A famous conjecture by Kannan-Lovasz-Simonovitz is that the previous bound for spherically symmetric measures extends (up to a change of the constant 1313) to any log-concave measure. If true, the KLS conjecture will presumably give a better upper bound for the Poincaré constant than ours.

Regarding the log-Sobolev constant, the work by Huet , furnishes a lower bound for the isoperimetric profile of μ\mu (see Theorem 3 and the discussion p.98 therein) which indicates a similar bound for the log-Sobolev constant as above, i.e. depending on the isotropic constant (n1−β2βn^{\frac{1-\beta}{2\beta}}) of μ\mu. ♢\diamondsuit

2.2. Lack of uniform convexity.

Now choose α(a)=aβ\alpha(a)=a^{\beta} for some β≥1\beta\geq 1 and a≤1a\leq 1, and α(a)=1\alpha(a)=1 for a≥1a\geq 1. (H.α\alpha.K) is less restrictive than before since it only implies a linear behavior at infinity for the gradient of the potential.

We now have F(a)=a12β+1F(a)=a^{\frac{1}{2\beta+1}} for a≤1a\leq 1 and F(a)=aF(a)=a for a≥1a\geq 1. It follows

if ∫∣∇f∣2 dμ≤K232 ∫f2dμ\int|\nabla f|^{2}\,d\mu\leq\frac{K^{2}}{32}\,\int f^{2}d\mu and

Assume that UU satisfies (H.α\alpha.K) for K>0K>0 and α(a)=aβ∧1\alpha(a)=a^{\beta}\wedge 1 for β≥1\beta\geq 1. Then, μ\mu satisfies both a Poincaré inequality with

Using the general form of the (H.W.I) inequality it is quite easy to adapt the previous proof in order to show the following result : Let μ\mu satisfying (H.C.K) for some K>−∞K>-\infty. If μ\mu satisfies a weak W2IW_{2}I inequality,

for some 0<β≤10<\beta\leq 1, then μ\mu satisfies a log-Sobolev inequality with a constant depending on C,K,βC,K,\beta only. In particular μ\mu satisfies a T2T_{2} inequality.

In particular if we know that μT\mu_{T} has a bounded below curvature, the previous theorems extend to μT\mu_{T}. ♢\diamondsuit

Using reflection coupling.

As we have seen in the previous section, the simple coupling using the same Brownian motion is not fully well suited to deal with non uniformly convex potentials. In a recent note , Eberle studied the contractivity property in Wasserstein distance W1W_{1}, induced by another well known coupling method: coupling by reflection (or mirror coupling) introduced in . We shall see now how to use this coupling method in the spirit of what we have done before.

In this section we consider XtxX_{t}^{x} the solution starting from xx of the Ito stochastic differential equation

where bb is smooth enough. We introduce another formulation of the semi-convexity property, namely :

We shall say that (H.κ\kappa) is satisfied if

This condition is typically some “uniform convexity at infinity” condition. Indeed if b=−12 ∇Ub=-\frac{1}{2}\,\nabla U where U=U1+U2U=U_{1}+U_{2} with U1U_{1} satisfying (H.C.κ∞\kappa_{\infty}) and U2U_{2} compactly supported, then (H.κ\kappa) is satisfied. We shall come back later to this. Notice that if (6.3) is satisfied, the solution of (6.1) is strongly unique and non explosive, using the same tools as we used before. Now, following we introduce (with some slight change of notations)

If (H.κ\kappa) is satisfied, R0<+∞R_{0}<+\infty so that φmin>0\varphi_{min}>0 and

i.e. D(∣x−y∣)D(|x-y|) which is actually a distance, is equivalent to the euclidean distance. Hence a consequence of Theorem 1 in is the following :

Assume that (H.κ\kappa) is satisfied. Let λ\lambda be defined by

Then for all initial distributions ν\nu and μ\mu, and all tt, the W1W_{1} Wasserstein distance satisfies

In order to prove this result, Eberle adds to (6.1) the following Ito s.d.e.

where et=(Xt−Yt)/∣Xt−Yt∣e_{t}=(X_{t}-Y_{t})/|X_{t}-Y_{t}| and e∗e^{*} is the transposed of ee (remark that if n=1n=1, it just changes BB into −B-B). Of course one has to consider (6.1) and (6.6) together. Existence, strong uniqueness and non explosion are again easy to show. Now introduce the coupling time TcT_{c} defined by

It is easy to see that X.yX_{.}^{y} and Xˉ.y\bar{X}_{.}^{y} have the same law, so that the distribution of (Xtx,Xˉty)(X_{t}^{x},\bar{X}_{t}^{y}) is a coupling of P(t,x,.)P(t,x,.) and P(t,y,.)P(t,y,.). Of course this extends to any initial distributions (μ,ν)(\mu,\nu) and furnishes a coupling of (μt,νt)(\mu_{t},\nu_{t}).

It follows that Zt=Xtx−XˉtyZ_{t}=X_{t}^{x}-\bar{X}_{t}^{y} solves

where Wt=∫0t es∗ dBsW_{t}=\int_{0}^{t}\,e_{s}^{*}\,dB_{s} is a one dimensional brownian motion.

The key of the proof of Theorem 6.5 is then that, if rt=D(∣Xtx−Xˉty∣)r_{t}=D(|X_{t}^{x}-\bar{X}_{t}^{y}|), r.r_{.} is a semi-martingale with decomposition

Taking expectation, it immediately shows that the WDW_{D} Wasserstein distance decays exponentially fast. As remarked by several authors, one can then deduce as we did previously

Assume that b=−12 ∇Ub=-\frac{1}{2}\,\nabla U satisfies(H.κ\kappa) and that μ=(1/ZU)e−U\mu=(1/Z_{U})e^{-U} is a probability measure. Then μ\mu satisfies a Poincaré inequality with constant CP≤(1/2λ)C_{P}\leq(1/2\lambda).

for all Lipschitz function ff. According to Lemma 2.12 in , we deduce that Varμ(Ptf)≤e− 2 λt Varμ(f)\textrm{Var}_{\mu}(P_{t}f)\leq e^{-\,2\,\lambda t}\,\textrm{Var}_{\mu}(f), hence the result. ∎

One can see that the reflection coupling cannot furnish some information on W2W_{2}, the concavity of DD near the origin being crucial. In the same negative direction, theorem 6.11 cannot be extended to the log-Sobolev framework, the key lemma 2.12 in being restricted to the variance control. ♢\diamondsuit

If (H.φ\varphi.K) is satisfied with φ(a)=aβ\varphi(a)=a^{\beta}, we have κ(a)=a2(β−1)\kappa(a)=a^{2(\beta-1)}. We thus have R0=0R_{0}=0, R1=(8/K)12βR_{1}=(8/K)^{\frac{1}{2\beta}}, φmin=1\varphi_{min}=1 and finally

We recover the result in Theorem 5.20, i.e. a bound Cβ K− 1/βC_{\beta}\,K^{-\,1/\beta} but with a better constant CβC_{\beta}.

If (H.α\alpha.K) is satisfied, one can choose R0=2εR_{0}=\sqrt{2\varepsilon}, R1=2ε+(4/Kα(ε))R_{1}=\sqrt{2\varepsilon}+(4/\sqrt{K\alpha(\varepsilon)}), φmin=exp⁡(− 14 ε2 Kα(ε))\varphi_{min}=\exp\left(-\,\frac{1}{4}\,\varepsilon^{2}\,K\alpha(\varepsilon)\right) and finally

Now assume that the potential UU can be written U=U1+U2U=U_{1}+U_{2} where U1U_{1} satisfies (H.C.K) for some K>0K>0 and U2U_{2} satisfies ∥∇U2∥∞=M<+∞\parallel\nabla U_{2}\parallel_{\infty}=M<+\infty. It easily follows that (H.κ\kappa) is satisfied with κ(a)=K−Ma\kappa(a)=K-\frac{M}{a}. We thus have

An old result by Miclo (unpublished but explained in ) indicates that such a result (without the square of the supremum of the gradient but without KK in the exponential) can be obtained by using the usual Holley-Stroock perturbation argument. ♢\diamondsuit

2. The log-concave situation.

Now consider the situation where bb satisfies (H.C.0). In this situation we have λ=0\lambda=0 (R1=+∞R_{1}=+\infty, φ=g=1\varphi=g=1) so that (6.8) becomes

for t≤Tct\leq T_{c} with βt≤0\beta_{t}\leq 0. It follows that rt≤∣x−y∣+2 Wtr_{t}\leq|x-y|+2\,W_{t} up to the first time T∣x−y∣T_{|x-y|} the brownian motion W.W_{.} hits − ∣x−y∣/2-\,|x-y|/2. In particular,

Actually if b=− 12 ∇Ub=-\,\frac{1}{2}\,\nabla U with UU convex (i.e. in the zero curvature situation of the Γ2\Gamma_{2} theory) the inequality ∣∇Ptf∣≤1t  ∥f∥∞|\nabla P_{t}f|\leq\frac{1}{\sqrt{t}}\,\,\parallel f\parallel_{\infty} is well known as a consequence of what is called the reverse (local) Poincaré inequality (see ). The previous proposition extends this result (up to the constant) to a non-gradient drift.

Recall that Xtx=XˉtyX_{t}^{x}=\bar{X}_{t}^{y} for t>Tct>T_{c}. It follows

If one wants to get a contraction bound for the gradient (in the spirit of (6.10) or better of proposition 2.2) we cannot only use a comparison with the brownian motion for which ∇Ptf=Pt∇f\nabla P_{t}f=P_{t}\nabla f.

In the symmetric situation (b=−12 ∇Ub=-\frac{1}{2}\,\nabla U) it is known that t↦∫ ∣∇Ptf∣2 dμt\mapsto\int\,|\nabla P_{t}f|^{2}\,d\mu is non increasing. It easily follows that

If we assume that (H.κ\kappa) is satisfied, we may replace the comparison with a Brownian motion by the comparison with an Ornstein-Uhlenbeck process with parameter λ/2\lambda/2, according to standard comparison theorems for one dimensional Ito processes (see e.g Chapter VI theorem 1.1). For the O-U process, it is known (see ) that

These bounds are interesting as regularization bounds (from bounded to Lipschitz functions), but notice that we have lost a factor 22 in the exponential decay. ♢\diamondsuit

3. Reflection coupling for general diffusions.

The case of a general diffusion process with a non constant diffusion matrix as in section 3 is more delicate to handle, as already remarked in (Theorem 1).

Assume that σ\sigma is a bounded and smooth square matrices field and that it is uniformly elliptic. The quantities we need here are (notations differ from )

Recall the Lindvall-Rogers reflection coupling

Existence and strong uniqueness can be shown as previously. Of course, as in subsection 6.1, we replace Xt′X^{\prime}_{t} by XtX_{t} if t>Tct>T_{c} the coupling time, but not to introduce new notation we still use X.′X^{\prime}_{.}.

Applying Ito formula we thus have for a smooth function DD

We introduce the natural generalization of (H.κ\kappa), namely we assume that

If DD is a non decreasing, concave function we thus get, provided (2/N)−Λ>0(2/N)-\Lambda>0,

Hence looking carefully at the calculations in , we see that, provided (2/N)−Λ>0(2/N)-\Lambda>0, the only thing we have to change in (6.4) is the definition of φ\varphi replacing 1/41/4 by the inverse of (2/N)−Λ>0(2/N)-\Lambda>0, all other definitions being unchanged. We have thus obtained

Assume that (6.17) and (6.18) are satisfied. Assume in addition that (2/N)−Λ>0(2/N)-\Lambda>0. Then defining

the conclusion of Theorem 6.5 is still true with λ=12 (φmin/R12)\lambda=\frac{1}{2}\,(\varphi_{min}/R_{1}^{2}).

All the consequences of Theorem 6.5 still hold (up to the modifications of the constants), in particular one can extend (H.C.K) to the situation of “convexity at infinity” as in Example 6.13 (3). Details are left to the reader.

As we already said, the condition (2/N)−Λ>0(2/N)-\Lambda>0 already appears in and ensures that the coupling by reflection is succesfull. Roughly speaking it means that the fluctuations of σ\sigma are not too big with respect to the uniform ellipticity bound.

4. Gradient commutation property and reflection coupling

It is of course quite disappointing at first glance that the only gradient commutation property that we get using this nice contraction results in W1W_{1} distance, is restricted to Lipschitz function as in (6.10). Let us see however that we may transfer this to stronger gradient commutation properties in some cases. The main tool is the following lemma on Hölder’s type inequality in Wasserstein distance.

The proof is indeed quite simple and relies mainly on Hölder’s inequality. Indeed, in dimension one the optimal transport plan is the same for every convex cost (see for example Villani ), so that there exists a transport plan π\pi such that

The case of product probability measure is deduced using the result in dimension one and the following two direct assertions

and if ν\nu and μ\mu have for ithith marginal νi\nu_{i} and μi\mu_{i}

In fact, as will be seen from our applications, even if cc does depend of ν\nu and μ\mu (in a nice way), it would be sufficient to get new gradient commutation property. ♢\diamondsuit

We are now in position to prove various gradient commutation properties in non standard cases. For simplicity, we suppose here that the diffusion coefficient is constant, i.e.

so that the weak gradient commutation property holds

and thus a local Poincaré inequality holds.

Note that this theorem is the first one to give the commutation gradient property in non strictly convex case with a good behaviour at infinity.

Using synchronous coupling as previously explained and the fact that κ(r)≥−L\kappa(r)\geq-L we have that

In the same time, by Theorem 6.5, we have that

We then use Lemma 6.21 to get the first assertion. The second one is obtained as in Proposition 3.11. ∎

Consider for example the log-concave case b(x)=−4x3b(x)=-4x^{3} for which Bakry-Emery theory enables us to get that we are in 0-curvature and thus

However, using Theorem 6.23 and this last inequality, we easily get that there exists λ>0\lambda>0 such that

which is completely new. It captures both the short time behavior equivalent to the Γ2\Gamma_{2} 0-curvature criterion and the long time behavior for which Ptf→μ(f)P_{t}f\to\mu(f) and thus ∇Ptf∣\nabla P_{t}f| is expected to decay to 0. Note that we may extend this example to a double well potential, in the case when the height of the well is not too large.

Preserving curvature.

A natural question about curvature is the following: is curvature preserved by a diffusion process ? According to a result by Kolesnikov , the Ornstein-Uhlenbeck process is essentially the only one, among diffusion processes, preserving log-concavity (i.e. if ν0\nu_{0} is log-concave, so is νT\nu_{T} for all T>0T>0). One may also wonder if νt\nu_{t} may satisfy other “curvature” like inequality as HWIHWI for example. It would have important applications on local inequalities, indeed transportation inequalities together with a HWI inequality may imply logarithmic Sobolev inequality.

In the spirit of the previous remark, consider a standard Ornstein-Uhlenbeck process X.X_{.}, i.e. the solution of

where GG and ZZ are independent random variables, GG being a standard gaussian variable and ZZ having distribution ν\nu. Hence

or, if we use the notation CLS(Y)=CLS(η)C_{LS}(Y)=C_{LS}(\eta) for a random variable YY with distribution η\eta,

But if AA is a random variable it is clear that CLS(λA)=λ2 CLS(A)C_{LS}(\lambda A)=\lambda^{2}\,C_{LS}(A). It follows

The change of variable Z=Z′+eλT−1λ GZ=Z^{\prime}+\sqrt{\frac{e^{\lambda T}-1}{\lambda}}\,G, yields, using the symmetry of GG

In particular, for λ=−(1/α2)<0\lambda=-(1/\alpha^{2})<0 (α>0\alpha>0), we may let TT go to +∞+\infty and obtain

In particular, the distribution of ZZ satisfies a log-Sobolev inequality if and only if, for some α>0\alpha>0, the distribution of Z+αGZ+\alpha G satisfies a log-Sobolev inequality , and then the considered inequality is satisfied for all α\alpha. Recall that ZZ and GG are independent. This is not surprising since a more general result can be obtained directly (extending a similar result for the Poincaré inequality in ):

Let X and Y be independent random variables and λ∈\lambda\in then,

Conversely, if Y is symmetric (i.e. Y and -Y have the same distribution), we also have

The first result for CPC_{P} is proved in proposition 1. For CLSC_{LS} the proof is very similar. Let ff be smooth. Then

Since ∇f2=2f ∇f\nabla f^{2}=2f\,\nabla f, we may use Cauchy-Schwartz inequality in order to bound the last term in the latter sum. This yields exactly the desired result.

For the second statement, it is enough to use a change of variable, as we did in the gaussian case and the symmetry of YY. ∎

Hence we get a general statement: the distribution of a random variable XX satisfies a Poincaré or a log-Sobolev inequality if and only if, for all or for one, symmetric random variable YY independent of XX whose distribution satisfies a Poincaré or a log-Sobolev inequality, the distribution of X+YX+Y satisfies a Poincaré or a log-Sobolev inequality. ♢\diamondsuit

It is known that any log-concave probability measure satisfies some Poincaré inequality. The result is due to Bobkov (a short proof is contained in ). But if ZZ is a log-concave random variable that do not satisfy a log-Sobolev inequality, Z+αGZ+\alpha G (which is still log-concave according to the Prekopa-Leindler theorem) does not satisfy a log-Sobolev inequality, in particular is not uniformly log-concave.

Here is an amusing proof of the consequence of Prekopa’s result when one variable is gaussian.

Let XX (resp. YY) be a random variable with law e−V(x)dxe^{-V(x)}dx (resp. a standard gaussian variable). We assume that XX and YY are independent. The density of X+λ YX+\sqrt{\lambda}\,Y is thus given by

Let H(x)H(x) be the hessian matrix of log⁡p\log p. Then

(or if one prefers its potential) satisfies (H.C.K+(1/λ)K+(1/\lambda)). If K+(1/λ)>0K+(1/\lambda)>0, it thus satisfies a Poincaré inequality with constant λ/(1+Kλ)\lambda/(1+K\lambda). Applying this Poincaré inequality to the function u↦⟨ξ,(x−u)⟩u\mapsto\langle\xi,(x-u)\rangle, we obtain

Thanks to simple scales we may thus state

Appendix A Some general remarks.

We give in this section some general facts which we used from place to place.

First we recall two facts one can find for instance in remark 2.11 and in lemma 2.12:

Accordingly, since Varμ(Ptf)=12 ∫t+∞ ∣∇Psf∣2 ds\textrm{Var}_{\mu}(P_{t}f)=\frac{1}{2}\,\int_{t}^{+\infty}\,|\nabla P_{s}f|^{2}\,ds, a control ∣∇Psf∣≤cf e−βs|\nabla P_{s}f|\leq c_{f}\,e^{-\beta s} will furnish a Poincaré inequality. Notice that if ∥∇Paf∥∞≤ρ ∥∇f∥∞\parallel\nabla P_{a}f\parallel_{\infty}\leq\rho\,\parallel\nabla f\parallel_{\infty} for some a>0a>0, ρ<1\rho<1 and all Lipschitz functions ff, then using the semi-group property we get ∥∇Ptf∥∞≤C ρt ∥∇f∥∞\parallel\nabla P_{t}f\parallel_{\infty}\leq C\,\rho^{t}\,\parallel\nabla f\parallel_{\infty} for some C>0C>0 and all tt, hence a Poincaré inequality. The latter is a weaker form of the commutation of the gradient and the semi-group up to an exponential rate.

One can use the previous remarks to show that an exponential decay of Wasserstein distances furnishes some Poincaré inequality for μ\mu. In what follows W0W_{0} denotes the total variation distance and W1W_{1} is the usual 11-Wasserstein distance.

Assume that for all bounded (resp. Lipschitz) density of probability hh we have W0(Pthμ,μ)≤ch(t)W_{0}(P_{t}h\mu,\mu)\leq c_{h}(t) (reps. W1W_{1}). Then for all bounded (resp. Lipschitz and bounded) ff, there exist cfc_{f} and hh such that Varμ(Ptf)≤cf ch(2t)\textrm{Var}_{\mu}(P_{t}f)\leq c_{f}\,c_{h}(2t). In particular if ch(t)=ch e−βtc_{h}(t)=c_{h}\,e^{-\beta t}, μ\mu satisfies a Poincaré inequality.

hh is thus a density of probability with ∥h∥∞≤2\parallel h\parallel_{\infty}\leq 2. We have

One can replace W0W_{0} by W1W_{1}, just replacing ∥h∥∞\parallel h\parallel_{\infty} by ∥∇h∥∞\parallel\nabla h\parallel_{\infty} in which case

Even when the decay is not exponential, one gets a weak form of the Poincaré inequality (called a weak Poincaré inequality).

In the situation of proposition A.2, assume that ch(t)=ch c(t)c_{h}(t)=c_{h}\,c(t) with c(t)→0c(t)\to 0 as t→+∞t\to+\infty. Then

for all s>0s>0 where Ψ(f)=ch ∥f−∫fdμ∥∞2\Psi(f)=c_{h}\,\parallel f-\int fd\mu\parallel_{\infty}^{2} for W0W_{0} and Ψ(f)=ch ∥f−∫fdμ∥∞ ∥∇f∥∞\Psi(f)=c_{h}\,\parallel f-\int fd\mu\parallel_{\infty}\,\parallel\nabla f\parallel_{\infty} for W1W_{1} and α(s)=s inf⁡u>0 1u c−1(uexp⁡(1−(u/s)))\alpha(s)=s\,\inf_{u>0}\,\frac{1}{u}\,c^{-1}(u\exp(1-(u/s))) .

Once we notice that the transformation f↦λff\mapsto\lambda f does not change hh the result follows from Theorem 2.3. ∎

Acknowledgements: This research was supported by the French ANR project STAB.

References