Semi Log-Concave Markov Diffusions
Patrick Cattiaux, Arnaud Guillin
Introduction and main results.
In this paper we shall investigate some properties of time marginals (at time finite or infinite) of Markov diffusion processes satisfying some logarithmic semi-convexity like property. The properties we are interested in are functional inequalities (Poincaré, log-Sobolev) or transportation inequalities. We shall also give some consequences for the long time behavior of such processes.
Our main tools are on one hand coupling techniques and on the other hand stochastic calculus. We shall mainly use the so called “synchronous” coupling, i.e. using the same Brownian motion, but we also give some new results by using the “mirror” coupling (or coupling by reflection) introduced by Lindvall and Rogers in . The main stochastic tool is (a very simple form of) Girsanov theory and -processes.
The use of coupling techniques for obtaining analytic estimates is far to be new. It is impossible (and dangerous) to give here, even an account of the existing literature (see however and references therein). The use of Girsanov theory for this goal is not new too. We shall recall later some references. The conjunction of both techniques is not usual.
We deliberately decided to present in details the simplest situation of a Brownian motion with a gradient drift, for which almost everything is well known, and then to extend our method to new situations. Some specialists would certainly find that these parts of the present paper are lengthy, but we think that the understanding of how the method works in this simple case is an useful guide for generalizations.
The meaning of logarithmic semi-convexity will generalize the “usual” one we recall now.
This property is called -semi-convexity of . It is clearly equivalent to the convexity of . We denote the Boltzmann measure associated to the potential . If is integrable, we also introduce the normalized which is a probability measure. If is semi-convex, is said to be semi log-concave.
Consider first the diffusion process, given by the solution of the Ito stochastic differential system
and denotes the carré du champ, namely here
It is known that is a symmetric (reversible) measure for the diffusion process, and is actually the unique invariant (stationary) measure for the process. If is bounded, is ergodic.
In the latter case, is thus a symmetric semi-group on . The domain of its generator contains the algebra generated by the constant functions and . In particular, if , in , so that since is hypo-elliptic .
is the basic example of generator satisfying the celebrated Bakry-Emery curvature condition (see ). Indeed if we define
(H.C.K) is equivalent to .
This curvature condition is known to imply (and is in fact equivalent to) a lot of nice inequalities for the semi-group, in particular for all and all , a commutation between and the semi group holds, namely
which in turn implies powerful functional inequalities such as
Recall that satisfies a log-Sobolev inequality with constant if
(1.3) is exactly what is (a little bit improperly) called a “local” log-Sobolev inequality in (theorem 5.4.7). For further informations and more, see the forthcoming book . The reader has to be careful with the constants, since we are writing the log-Sobolev inequality (as well as all other inequalities) in the usual form, while the “Bakry-Emery” version uses , hence with an extra factor 2.
It is well known that a log-Sobolev inequality implies a Poincaré inequality
with , as well as a transportation inequality
with . Here denotes the Wasserstein distance between the probability measures and , i.e.
where is a coupling of and (i.e. has respective marginals equal to and ) and
denotes the Kullback-Leibler information or relative entropy of w.r.t. . The latter property is due to Otto-Villani . Another approach and related properties were developed by Bobkov, Gentil and Ledoux . For a nice survey on transportation inequalities we refer to . One can find in all these references another remarkable consequence of semi log-concavity, namely that a log-Sobolev inequality derives from a transportation inequality. This is a consequence of the following (H.W.I) inequality that holds for any nice density of probability ,
As a consequence, if (H.C.K) holds for some , a transportation inequality for implies a log-Sobolev inequality with constant provided , in particular if . Let us finally remark that the starting point of this approach is the commutation property (1.2) which fails however to give a direct proof of the inequality.
Our first goal is to show that functional and transportation inequalities can be derived, in the previous situation, by using synchronous coupling and simple tools of stochastic calculus. This is done in section 2. The methods are then extended to a more general framework which is as natural for studying properties of time marginals as the framework.
Indeed, consider a classical diffusion process, given by the solution of an Ito stochastic differential system
being a standard brownian motion. For simplicity we assume that is a squared matrix. We extend (H.C.K) to this new situation
Notice that if and are -Lipschitz, i.e. each component is Lipschitz, (H.C.K) is satisfied for , but if is -Lipschitz, (H.C.K) can be satisfied for a non-negative provided is sufficiently repealing. Contrary to the case of a constant diffusion coefficient, (H.C.K) is not related to the Bakry-Emery curvature condition which involves in this situation controls on derivatives of higher order of the coefficients.
For simplicity in the sequel we shall assume that hence is -Lipschitz and that is , but not necessarily bounded nor with bounded derivatives. With these assumptions, once again if we assume that (H.C.K) is in force, (1.7) admits a unique non explosive strong solution using as a Lyapunov function for non explosion. We shall show this and other properties of the process in subsection 3.1.
We still use the notations introduced before, but now
Our first results can be gathered in Theorem 1.8 below, after introducing some additional assumptions.
Hypothesis (R). One of the following assumptions is satisfied (in addition to the fact that and ).
, and has at most polynomial growth and is hypo-elliptic,
, i.e. is symmetric.
Actually, assumptions (R1)-(R4) ensure that for ,
which is what we really need. (R5) is a limiting situation for (R3) as we will see in (sub)subsection 3.1.2. Of course the time marginal distributions only depend on and not on , but the constant is related to in (R5) (note that the Hilbert Schmidt norm of can change when we change without modifying ).
Assume that (R) and (H.C.K) are satisfied. Let .
If satisfies a Poincaré inequality with constant then satisfies a Poincaré inequality with constant
When one has to replace by . This applies in particular to with .
satisfies a transportation inequality with constant , bounded in time if , linear in time if and exploding exponentially in time if .
If satisfies with constant , satisfies with a constant , bounded in time if , linear in time if and exploding exponentially in time if .
When , if satisfies a log-Sobolev inequality with constant , satisfies a log-Sobolev inequality with constant
This applies in particular to with .
Some other consequences, as for example convergence to equilibrium (when it exists) are also discussed in particular in subsection 3.7.
Of course, (1) is a weaker version of the commutation relation (1.2) and (5) is nothing else than (1.3) when . When a diffusion coefficient is present, (1) is however very different from the usual commutation property. For example we will show that it holds even in the negative infinite curvature case, but that it still enables us to provide interesting local functional inequalities. (3) as well as the general version of (H.C.K) we have introduced appeared (for this kind of application and to our knowledge) for the first time in the paper by Djellout, Guillin and Wu , Theorem 5.6 and condition 4.5 therein, for . Our scheme of proof for the transportation inequality, based on Girsanov theory, is actually a simplified version of the one in , but instead of looking at the full law of the process on a time interval we shall use -processes in order to look at time marginals. What we shall show is that the same scheme of proof also furnishes functional inequalities. This unified treatment of functional inequalities and transportation inequalities using an ad-hoc coupling is the novelty here. It easily extends to time dependent coefficients as shown in section 4.
In addition in this section we show how to directly obtain convergence to equilibrium and properties of the invariant measure for non linear diffusions of Mc Kean-Vlasov type, simplifying arguments in .
The use of stochastic calculus in deriving such inequalities is not new but only a small number of papers dealt with. One can trace back to the paper of Borell , who used Girsanov theory to study the propagation of log-concavity along the Schrödinger dynamics (not the Fokker-Planck one we are looking at here). In addition to for transportation inequalities, one can also mention where similar ideas are used to study hyper-boundedness. More recently, using similar arguments, Lehec has studied gaussian functional inequalities and Fontbona and Jourdain obtained a pathwise version of the theory.
Let us come back to (1.1). (H.C.K) (or the theory) for applies to potentials which are the sum of and of a concave (hence sub-linear) potential . In particular, for “super-convex” potentials like with , or more generally for (smooth) potentials which are uniformly convex at “infinity”, (H.C.K) holds but with a negative due to the behavior of near the origin, so that, according to theorem 1.8, satisfies functional inequalities but with exploding constants in .
It is however well known, since with -uniformly convex and bounded, that satisfies a log-Sobolev (and a Poincaré) inequality with a constant where Osc denotes the oscillation of . One can thus expect that is bounded in . In section 5, we introduce the following extension of (H.C.K).
When , we may take and we recognize (H.C.K). In Proposition 5.4 we show that (with ) satisfies (H..) for and an explicit .
The main result of this section is then that, for suitable functions ,
if (H..K) holds (for ), then satisfies a log-Sobolev inequality.
See theorems 5.20 and 5.24. These theorems thus (partly) extend the Bakry-Emery criterion (1.3) to some non uniformly convex potentials. However, they are dealing with the invariant measure only and not with the law at time (only incomplete results are proved in this section for these distributions).
The next section 6 is devoted to the use of mirror coupling. In a recent work , Eberle has adapted the mirror coupling to get estimates of convergence for drifted brownian motions when the drift satisfies some “convexity at infinity” property. We recall Eberle’s method and obtain some new consequences of his result. In addition, up to an extra condition, we show that his result (and all the consequences we derived) can be extended to general elliptic diffusion processes. We will also use this mirror coupling to show that we may get a weak version of the commutation property in the log concave case with the “convexity at infinity” property at least in dimension one, which is the first result we know of in this direction. Still in dimension one, we will also consider using mirror coupling for non linear diffusions.
Section 7 is peculiar. Using the results we have described for the Ornstein-Uhlenbeck process we show how to recover known results on the stability of functional inequalities under convolution (provided one of the terms is gaussian).
Semi log-concave drifted brownian motion.
In this first warming up section we shall look at the usual situation given by (1.1)
For functional inequalities the key is a commutation property of the gradient and the semi group. This commutation property is almost immediate using an appropriate coupling as explained below :
In the situation of (1.1), assume (H.C.K). Then for all ,
Applying Ito formula yields (almost surely)
for some sandwiched by and . It remains to use the continuity (and boundedness) of and the fact that goes almost surely to as to conclude. ∎
As is seen from the proof, in fact, the sole convergence of the Wasserstein distance is not sufficient to get the commutation property exposed here. It will however be our starting point for the result when a diffusion coefficient is present. The synchronous coupling here enables us however to get an almost sure “deterministic” control of which is far more powerful.
We recall previously that (H.C.K) is exactly the condition of Bakry-Emery in this context, which is in fact equivalent to (2.3). However the proof is very different from ours: it relies on a tricky calculus on to show that .
If instead of the processes start with initial distribution the “optimal coupling” between and for the distance, the previous shows that . As discussed in the Appendix, this result can be used to show the existence and uniqueness of the invariant measure.
2. hℎh-processes and functional inequalities.
We now introduce the standard notion of -process. Let and be a non-negative function such that . For simplicity, we assume for the moment that there exist and such that . We thus may define on the path-space up to time a new probability measure
In this situation, it is well known (Girsanov transform theory) that one can find a progressively measurable process such that
Actually, if , it is immediate to check (applying Ito formula) that
If is smooth we may apply Proposition 2.2 in order to get
where we have used Cauchy-Schwarz inequality for the second inequality and the Markov property for the third one. The previous inequality then extends to any in for which the right hand side makes sense, by density. We have thus obtained the following
In the situation of (1.1), assume (H.C.K). If satisfies a log-Sobolev inequality with constant , satisfies a log-Sobolev inequality with constant
When one has to replace by . This applies in particular to since satisfies a log-Sobolev inequality with constant equal to .
Apply the log-Sobolev inequality to . It furnishes (since ),
similarly as what we did in (2.9). Hence the result applying (2.9). ∎
As we recalled in the introduction a log-Sobolev inequality implies a transportation inequality. It is interesting to see that one can directly obtain such an inequality for semi log-concave measures, by using the previous construction. But before to do this, just remark that the above proof using with allows us to obtain a similar result replacing the log-Sobolev inequality by a Poincaré inequality i.e.
In the situation of (1.1), assume (H.C.K). If satisfies a Poincaré inequality with constant , satisfies a Poincaré inequality with constant
When one has to replace by . This applies in particular to since satisfies a Poincaré inequality with constant equal to .
Once again, the proof presented here is very different from the one based on the calculus of Bakry-Emery which relies on the commutation property and the control of the derivative of to get a local logarithmic Sobolev inequality. Note that considering rather leads to a local Poincaré inequality.
3. Transportation inequalities.
The existence of and (2.7) are ensured as soon as (see ). For our goal we do not need the explicit expression of .
We have then different alternatives depending on the sign of .
1) First in the case where , one has using that
so that we recover an uniform transportation inequality when , which is moreover optimal for the invariant measure, considering logarithmic Sobolev inequality and Poincaré inequality. If satisfies some transportation inequality then one obtains that satisfies a transportation inequality with constant the sum of the initial constant plus .
2) The previous simple argument has however a serious drawback in the sense that in positive curvature, does not forget the “initial measure”. Let us see how to deal with this problem. Start once again from the first estimation, but using Itô’s formula between and and (H.C.K)
so that we may differentiate in time to get for all positive
Note that this is once again optimal for the limiting measure, and captures the fact that it forgets the initial condition. When , we then have
Note however the presence of the additional parameter .
3) Let us see how a direct approach may get rid of this additional parameter, which is particularly important in negative curvature. Define
Since we obtain
If , (2.14) can be improved in
Again if , (2.16) can be improved in
The previous inequalities then extend to any non-negative (not necessarily bounded below nor above).
If we choose , we have , and so and . Hence
In the situation of (1.1), assume (H.C.K). Then satisfies a transportation inequality
If we choose , we may use the convexity of , i.e
In the situation of (1.1), assume (H.C.K). If satisfies with constant , then satisfies with a constant given,
and when , with
All what precedes holds even if is not bounded (i.e. if the process is not positive recurrent), in which case of course, .
If we choose (assuming that is bounded), we have to choose hence . After noticing that we can slightly refine the previous bound replacing by according to (2.7), we obtain
The latter has to be compared with remark 4.9 in which shows that the inequality
If we may let go to in Proposition 2.10 and recover that satisfies a log-Sobolev inequality with constant , hence a transportation inequality with constant (in particular we are loosing a factor in Proposition 2.18). Similarly, when , (2.22) shows that if , satisfies a inequality , and since is log-concave, satisfies a log-Sobolev inequality. This scheme of proof does not require Proposition 2.2, but the (H.W.I) inequality. Unfortunately it does not furnish the optimal constant.
4. Transportation-Fisher Inequalities.
Let us see look now at another type of Transportation Information inequality recently introduced in , which is weaker but quite close to logarithmic Sobolev inequality (in fact equivalent under bounded curvature). We are obliged to come back to the initial inequality in (2.3) which becomes in our new situation
Replacing the pair by we thus have
It follows that is differentiable and satisfies,
(for the second inequality, recall that (H.C.) is satisfied so that, for short, .) To explore (2.25) we shall use the usual trick for positive. Hence
This inequality is close to what is called a inequality (see definition 10.4 or for examples and details on properties of WI inequality). Here we obtain a defective inequality. However, as , we recover the true inequality for the invariant distribution, which together with the (H.W.I) inequality allows us to recover the log-Sobolev inequality. Nevertheless, we get
Assume (H.C.K), then satisfies a WI inequality of constant . If we suppose moreover that satisfies a WI inequality with constant then satisfies a WI inequality with constant .
As remarked, under (H.C.K), the inequalities verified by the law depend on the inequalities verified by the initial measure, in the range between Poincaré and logarithmic Sobolev inequality. Indeed, a logarithmic Sobolev inequality implies a WI inequality, but to get the the WI inequality for we need only a WI inequality for the initial measure. As seen by the example of the Gaussian measure, which satisfies (H.C.K), no stronger inequalities can be obtained.
Instead of the -process one should consider Schrödinger bridges allowing to choose both the initial and the final time marginals. Indeed if is bounded it is known that one can find non-negative functions such that the measure
and the drift is given by . As before it is immediately checked that
For all this we refer to p.162 and section 6. Even if is bounded from below, we do not know whether inherits this property. Nevertheless, at least formally we have the relation
Unfortunately, these inequalities do not seem to give new results.
In almost all what we did we may replace the drift by a general (non-gradient) smooth drift satisfying
All the results of this section remain true in this more general situation, as far as we do not use reversibility. The only result where we used reversibility actually is (2.22). Indeed, if the initial law is , where denotes the adjoint semi-group.
The only delicate point is the smoothness of and the fact that in the usual sense. This will be discussed in an even more general setting in the next section, where we shall look at more general cases with non constant diffusion coefficient.
General diffusion processes.
We shall now extend the results of the previous section to the general situation of (1.7),
First of all we have to discuss some properties of the process and the associated quantities. As we said in the introduction, we need some regularity for at least if . So there is a technical price to pay. We decided to pay this price at the level of the study of the process, rather than in deriving inequalities.
Since we assume that , when (H.C.K) is fulfilled, satisfies
Thus, if denotes the exit time from the ball , and it holds
where, since is bounded, we have defined
It is then easily seen that one can perform similar calculations with for a large enough in order to kill the integrated term, i.e
It is interesting to notice that one can similarly obtain some “deviation” bound from the starting point. Indeed
so that arguing as we did in order to get (2.14) and (2.16) we obtain the existence of constants and such that for ,
1.2. Properties of the semi-group.
Let us mimic what we did to get Proposition 2.2 i.e. apply Ito formula to get
Notice that with our assumptions, the right hand side of (3.4) is a (true) martingale, so that
Moreover, if is positive then (3.7) implies back (H.C.K.). If we suppose moreover for some
The contraction in distance inherited from (H.C.K.) has already been proved. The contraction in distance is done exactly in the same way using once again synchronous coupling. The necessary part comes from (or more precisely section 4. in the Arxiv version 1110.3606). Let us explain the ideas of the proof. In fact, one may compute the time derivative of the Wasserstein distance: note when and are two matrices, then denoting and two solutions starting respectively from and
where if ,
Then the contraction property implies that at time 0 for and
A clever choice of then enables to prove the result. ∎
Let be -Lipschitz continuous. It holds so that, using (3.5), is Lipschitz continuous with Lipschitz constant less than .
As we said at the end of the previous section, when one deduces the existence and uniqueness of an invariant probability measure , to which converges weakly.
The rest of this (sub)subsection is devoted to give a proof of the following: if (see the introduction), then is regular and satisfies (for )
The reader who takes this result as granted can skip what follows.
First if , is and we have . It follows that for all , since is continuous. So . The first delicate point is of course the commutation of and . The second delicate point is the smoothness of .
This commutation property is known if and are in (see p.254-258, boundedness of derivatives is important) in which case is actually . But assuming boundedness of and its derivatives will exclude the cases of positive .
We shall first show that is a mild solution, provided does not grow too fast.
It follows from lemma 3.2 that if , is bounded by , hence is a continuous semi-group on (whose range is included into ). To see that belongs to the domain of the generator of , we have to show that the convergence of holds for the norm defined on . But
according to (3.3). Since , convergence holds for the norm on . The proof is completed. ∎
If the coefficients are , and is hypo-elliptic (for instance if is uniformly elliptic) it follows that and that the last equalities hold in the usual sense.
If we do not want to assume too much regularity on the coefficients, we have first to assume that is uniformly elliptic and call upon P.D.E. theory. If what follows is certainly well known by specialists, we include the argument.
Let . For large enough, contains the support of . Consider the parabolic equation
Since on this makes sense. If is uniformly elliptic, it is known that there exists an unique solution in of (3.10), and that this solution is represented as
where denotes the exit time from of . For all this see , in particular Theorem 5.2. p.147. It follows in particular that for all , and that for all , as since . Now let be fixed, and look at . The parabolic Schauder estimate tells us that there exists a constant depending on , the ellipticity constant and the norms of and such that
where is the set of functions with -Hölder derivatives. Arzela-Ascoli theorem tells us that a subsequence of converges in , and since the limit is , that the latter is .
If is not uniformly elliptic, we may approximate it by for . If , the diffusion matrix (field) is and uniformly elliptic). It is known that its square root (in the sense of symmetric matrices) is bounded and too. In addition as the convergence taking place in . In particular, (H.C.K()) holds for the new diffusion process with a constant going to as , provided (H.C.K) is satisfied for .
It follows that the convergence of time marginals holds in Wasserstein distance, hence in the weak topology.
The conclusion of the previous discussion is the following: if , we may assume that is uniformly elliptic as far as the bounds we get do not depend on the ellipticity constant, and then go to the limit. We shall use this trick in the sequel.
2. Commutation property with the gradient
Once again, we will see how synchronous coupling or a contraction in Wasserstein distance provides the commutation property.
where the last step is done by using Theorem 3.6. Provided we know that exists, we have thus obtained a weaker form of Proposition 2.2,
Assume (R) and (H.C.K) or the weaker contraction property (3.7). Let . If exists (which is true except possibly for (R5)), it holds
Notice that, contrary to the Bakry-Emery bounded curvature case, the previous commutation property holds with the usual gradient and not with the natural one i.e. .
If Proposition 2.2 allowed us to obtain logarithmic Sobolev inequalities, the weaker Proposition 3.11 will allow us to obtain a weaker inequality, namely a Poincaré inequality.
It is worth mentioning here the following alternate proof of the commutation property, starting from Wasserstein contraction, as derived in the recent paper following our suggestion, i.e. using Kantorovitch-Rubinstein duality we have for all bounded Lipschitz denoting the inf convolution operator and initial measure and
Choose now , to get for all
which by homogeneity of the inf-convolution operator gives
This assertion is in fact stronger than the gradient commutation property which can be deduced by using the fact that the inf-convolution operator is the Hopf-Lax solution of the Hamilton-Jacobi equation.
3. h-processes and functional inequalities.
We now introduce the corresponding -process. Let and be such that
We thus may define on the path-space up to time a new probability measure
For simplicity, we assume in what follows that there exist and such that . In this situation, using again Girsanov transform theory, we know that we can find a progressively measurable process such that
If we may apply Proposition 3.11 in order to get (recall that )
where we have used the Markov property for the second inequality.
Now let be such that and choose so that and for small enough. Actually we will let go to so that in the limit . Standard manipulations thus yield
We can eventually use first the density of and then the trick we formerly described in order to relax the uniform ellipticity assumption (recall that hence only depends on too).
Arguing as for Proposition 2.10 we have obtained
Assume that (R) and (H.C.K) are satisfied. Let . If satisfies a Poincaré inequality with constant then satisfies a Poincaré inequality with constant
This applies in particular to with .
Contrary to the log-Sobolev inequality, the Poincaré inequality does not furnish a transportation inequality, so we shall try to adapt what we did in subsection 2.3.
4. Transportation inequalities.
We may thus conclude as in the previous section
Assume that (R) and (H.C.K) are satisfied. Let . The conclusions of Proposition 2.18 and Proposition 2.19 are still true, replacing by
Actually, when (R5) holds, we have proven this result for and . But as we have seen, in distance, so that if is bounded the same holds for to ( being a normalization constant). Finally if a inequality holds for all it extends to all using density and the fact that if weakly converges to .
Of course a inequality implies a Poincaré inequality, but the constant in Proposition 3.17 is better (in addition we only require that satisfies a Poincaré inequality).
One of the renowned consequence of such inequalities is the concentration of measure phenomenon for . In particular, under the assumptions of Proposition 3.20, satisfies a gaussian type concentration property. In particular has some exponential moment, fact we have already shown in lemma 3.2. But this integrability does not reflect all the strength of the inequality whose tensorization property is particularly useful for statistical purposes.
When is uniformly elliptic, this concentration property follows from gaussian estimates for the transition kernel. Here we obtain much more explicit constants (even if they are certainly far from optimality) which do not depend on the ellipticity constant.
Assume that is uniformly elliptic, i.e.
According to proposition 5.4.1, this is equivalent to the condition provided hence when is constant times the identity. In the non constant diffusion case, our condition (H.C.K) seems to be really different from the Bakry-Emery curvature condition.
5. An hypoelliptic example : kinetic Fokker-Planck equation
We present in this section an application of the techniques developed here in an hypoelliptic example where the Bakry-Emery curvature is negative and where (H.C.K.) may not be satisfied also. Let be the solution of the following SDE
also called stochastic Hamiltonian system. The long time behavior study of such a system has been considered for a long time and have been tackled by different techniques, see for example: hypocoercivity by Villani or Lyapunov function technique by Bakryal . However, due to its hight degeneracy, the Bakry-Emery curvature is so that we may not apply the technique. Remark also that the (H.C.K.) condition reads for all and
Let us first remark that if is Lipshitz continuous, (H.C.K) is verified for some negative and using synchronous coupling, one may remark that we are in the same situation than in Section2 so that we get that for some negative the gradient commutation property holds
and thus the logarithmic Sobolev inequality holds for . Let us remark once again that those properties are written with the usual gradient and not the Carré-du-Champ operator .
One may then wonder if it is possible to get the gradient commutation property with . In fact, using synchronous coupling and Itô’s formula applied to the function , following , we get that if where is -Lipschitz with sufficiently small there exists and such that is equivalent to the euclidean norm and
so that we get as in Section 2 the commutation property for some and
and thus a Logarithmic Sobolev inequality holds uniformly in time.
It is not hard to extend the result of this simplified setting to the case where the Brownian motion in the velocity has a diffusion coefficient which is bounded and -Lipschitz. We may then obtain a weaker gradient commutation property
and local Poincaré type inequality or Transportation information inequality like in Propositions 3.17 or 3.20, and if is sufficiently small uniform in time version of these inequalities (using functional ).
6. Interpolation of the gradient commutation property and local Beckner inequality
We have seen here that we cannot recover a logarithmic Sobolev inequality by our technique when (H.C.K.) is in force. Remember however that we have introduced the stronger (H.C.K.m) condition which implies a contraction in Wasserstein distance . It is then not hard to deduce some interpolation of the gradient commutation property
Assume (R) and (H.C.K.m) or the weaker contraction property (3.8). Let , if exists, it holds
Remark once again that this property does hold even if the diffusion coefficient is degenerate, so that variations of the hypoelliptic example of the previous subsection with a diffusion coefficient in the velocity enters into this framework. This contraction property may thus lead to a reinforcement of the Poincaré inequality to a Beckner inequality.
Assume (R) and (H.C.K.m) or the weaker (3.24). Let . Then for all nice , we have the following Beckner inequality
The proof unfortunately does not rely on the -process introduced previously but on the type proof. Denote . By (3.24) and Hölder inequality, for all nice non negative
7. Convergence to equilibrium in positive curvature.
Still in the uniform elliptic case, assume that . We already mentioned that in this case weakly converges to the unique invariant probability measure (which exists). In particular , for all smooth (say ), as well as . We deduce that
As we said, this result is not captured by the theory.
But we can obtain general convergence results, even in the non uniformly elliptic case. Indeed recall that in full generality
Assume that is bounded. Then if (H.C.K) holds for some , defining as before, there exists an unique invariant probability measure and for all nice enough function ,
In addition if is symmetric (i.e. ), it holds
Remark once again that what is used here is the weak gradient commutation property which is a consequence of (H.C.K.). The last part of the theorem follows from a result in recalled in the Appendix, lemma A.1 (2). Of course, unless we explicitly know the invariant measure, it is not easy to see wether is symmetric or not.
We may further extends the previous argument to the entropic convergence to equilibrium. Let us suppose that there exists an unique invariant measure .
Assume that is bounded and that the gradient commutation property
holds for some positive . Then for all nice positive function (defining as before)
The proof is as for the decay quite standard. Indeed,
One of the important point here is that we do not suppose any non-degeneracy on the diffusion coefficient, so that the result applies to the kinetic Fokker-Planck equation. It then provides an alternative to the approach by Villani , where he obtained such kind of convergence by completely different techniques with assumptions quite similar to the ones described in Section 3.5. One may then complete the approach by regularization of the Fisher Information in small time to obtain an entropic decay controlled by the initial entropy, see or .
Let us point out that even in the symmetric case, such a control is not sufficient to recover a logarithmic Sobolev inequality as the analog of lemma A.1 is no more valid for the entropy. Remark however that we have shown in section 2 how to recover a logarithmic Sobolev inequality for using the strong commutation gradient property (3.30). If , we may then let goes to infinity to recover a logarithmic Sobolev inequality for the invariant measure. It may be for example be used in the context of kinetic Fokker-Planck equation with non gradient coefficient, for which the invariant measure is unknown.
Let us consider, as in the Poincaré case via (H.C.K.) condition, a particular class of test function such that , so that . We then see adapting the preceding proof that a weak commutation of gradient property
obtained for example under (H.C.K.) condition implies that
Non homogeneous diffusion processes.
In the authors extended the theory to time dependent coefficients (non homogeneous diffusions). Considering the Ito system
If only depends on , the proof of proposition 2.2 (resp. 3.11) is unchanged using the process starting from and and replacing by . To obtain the analogue of proposition 2.10 and proposition 3.17, it suffices to remark that is equal to , and use what precedes for depending on only. ∎
For the transportation inequality we have to slightly modify the method in subsection 2.3. With the notations therein, (2.3) has become,
so that, as in the previous section we have to come back to
Using as usual we obtain (see the details of the derivation in the previous section) that for all increasing function
from which we deduce, provided we choose ,
If satisfies a inequality with constant , then
The best choice of is not clear. If is not positive on the whole , it seams that taking for some is enough. If for all (but not necessarily bounded from below by a positive constant), seems to be natural.
Assume that as and that . Then, for all , (the distribution of the process starting from and ) satisfies a Poincaré inequality (and a log-Sobolev inequality when ) with a constant bounded by (or ). The family is then tight, but we do not know whether it is weakly convergent or not. Nevertheless any weak limit satisfies the same functional inequality. When we know that for any initial random variable . It follows that if a sequence is weakly convergent to some , the sequence weakly converges to the same limit. In particular if we consider , , for some convex potential , (H.C.K(t)) is satisfied, so that any weak limit satisfies a log-Sobolev inequality. If does not satisfy a log-Sobolev inequality, it cannot be a weak limit, even if . In this situation one should expect that the “perturbation” of being smaller and smaller when growths, the convergence to will still hold. This is not the case.
2. Application to some non-linear diffusions.
We shall now discuss an example that does partly enter the framework of the beginning of this section.
Following consider the following non-linear stochastic differential equation
This is a non-linear diffusion of Mc Kean-Vlasov type modeling, for instance, granular media. We refer to the introduction of for details and motivations. One can approximate the solution of (4.6) by the first coordinate of a linear large particle system with mean field interactions. This is what is done in to study the long time behavior of .
Let we see how to apply what we have just done. First, under some conditions on and (we later shall give some of them) existence and weak uniqueness of (4.6) are ensured, provided the initial law admits some big enough polynomial moment. This will imply for all , the existence and uniqueness of solution of (4.7) with initial condition . As usual for these non-linear equations, if we consider the linear time inhomogeneous S.D.E.
the pathwise unique solution (up to explosion) is shown to satisfy (4.6) (i.e. ) so that it coincides with . So, once and are built, we may build our synchronous coupling as before. Now introduce an independent copy of .
If we assume in addition (as usual) that , and remember that is a copy of , it is still equal to
, and their first two derivatives have at most polynomial growth of order and ,
satisfies (H.C.) and satisfies (H.C.).
Let . If and have a polynomial moment of order , there exist an unique solution of (4.6) and an unique solution of (4.7) among the set of probability flows having having a polynomial moment of order with initial condition or . Furthermore
If and then
, and .
The moment condition ensuring existence and uniqueness is described in section 2.
Of course we may replace the initial and by probability distributions and satisfying the required moment conditions. This immediately furnishes the first assertion about the upper bound for the Wasserstein distance.
If it is easily seen that for all , hence
provided the same holds at time . This furnishes the second assertion for the upper bound.
Finally the convergence under strict positivity of our new “curvature” condition ensures the existence of the limiting measure . To see that is actually invariant, one can for instance use the following trick: first consider the solution of (4.7) with initial condition . Similar bounds for the Markov non homogeneous process (when we replace by ) are obtained applying the results of the beginning of this section. Hence the law of (which is exactly starting with as we explained before) converges to some limiting measure which in turn is equal to and is invariant for . This achieves the proof. ∎
The proof of the above result is new and direct, while the result is mainly contained in using particle approximation. Notice that in the approach is developed for the non homogeneous Markov diffusion and not for . Also notice that some direct study of the decay to equilibrium in distance for granular media is done in .
for which the invariant probability measure. Using the results in section 2 we thus have
In the situation of Theorem 4.9, if H’1 or H’2 are satisfied, satisfies a log-Sobolev inequality with constant or .
All what we have done extends to more general Mc Kean-Vlasov equations, with a diffusion coefficient and a drift satisfying hypothesis (R). In particular, positive curvature (in the sense of (H.C.K)) will also imply existence of and convergence to an invariant probability measure. The only difference is that we have to replace log-Sobolev inequality by Poincaré inequality in the latter proposition. Let us explain quickly what kind of model we may consider. We do not aim to be optimal, but will provide a flavor of the results on contraction with some non constant diffusion term. We will not focus also on the existence of solution of such equation. Let be solution of
Let us suppose H1 and H2, that is -Lipschitz and that
Suppose moreover that , then there exists an unique invariant distribution to (4.12), the convergence to in Wasserstein distance being exponential as above.
The proof follows the same line as before except that in the Itô’s formula, there is the diffusion part which comes into play for which we use the Lipschitz condition of the theorem. Note that Bolleyal have considered the case of a kinetic McKean-Vlasov equation, but with a constant diffusion coefficient in speed. As before, we may obtain some functional inequality for the invariant distribution as in Prop. 4.11 but we have to replace log-Sobolev inequality by Poincaré inequality .
Extensions to some non uniformly convex potentials.
Let us come back to (1.1), and assume that is bounded. We shall extend (H.C.K) to more general situations. The first natural extension is to replace the squared distance by some other convex functional of the distance. More precisely.
is increasing and convex, with and ,
is non decreasing,
there exist a positive function such that for all and all , , where denotes the inverse (reciprocal) function of .
Let . We shall say that (H..K) is satisfied for some if for all ,
On one hand, since and , (H..K) implies that is convex. On the other hand, if (H..K) is satisfied, since is smooth, is necessarily bounded near the origin since . Here of course if the latter is automatically satisfied.
If this is nothing else but (H.C.K). If we shall say that is super-convex. This terminology is justified by the example below.
Let for some . We shall see that (H..K) is satisfied for and some we shall estimate.
We start with the one dimensional case. In this case
If , we may assume that , write for and remark that if ,
If , we have, using the convexity of ,
Since , we may choose .
If , we write again with . Thus , since ,
It follows, since and ,
If , since , it holds
Let for some . Then (H..K) is satisfied for and . If we have the better bound .
If , for all and all , it holds
Hence (H..K) implies the following condition
The latter appears in the study of the granular medium equation in (condition(6)) for power functions . This formulation will be the interesting one. It can be extended in
(H..K) implies (H..K) with the same and . In this definition we do not need that is convex.
Now we shall see how to use (H..K).
This subsection contains first results which are not really convincing, but have to be tested. If we want to control the gradient , we may write for ,
Denoting , we thus have . If , this yields
This result (even after taking expectation) is not really satisfactory. Indeed, first we do not obtain any better control for than the one for a general convex potential (in particular we do not obtain a rate of convergence to ). In second place, the decay to of the Wasserstein distance we obtain is desperately slow, while we expected an exponential decay (which we know to hold true for for ). Notice however that we recover the exponential decay we obtained previously when .
If instead of (H..K) we use (H..K), it is not difficult to show that
If , choosing for some , we get
The method can be extended to the Mc Kean-Vlasov situation studied in subsection 4.2 and allows us to recover (up to the constants) Theorem 4.1 in without the help of a particle approximation. However, better results in this situation are obtained in .
Mimicking subsection 2.3, in particular (2.3), do we obtain more interesting results ? Using the notation therein we have
so that, if ,
If , we thus obtain
This result is certainly not fully satisfactory too. On one hand, we get a less explosive bound in time (recall that in the general convex case the bound growths like ), but on the other hand the relative entropy appears to a power less than 1. In particular such an inequality does not imply a Poincaré inequality (which is obtained for entropies going to ), but furnishes nice concentration properties (obtained for large entropies via Marton’s argument).
2. An improvement of Bakry-Emery criterion.
As we remarked at this end of section 2 we may come back to the initial inequality in (2.3) which becomes in our new situation
(here again (H.C.) is satisfied so that, for short, .) To explore (5.13) we shall use both the remark 5.5 and the usual trick for positive. Hence
We deduce, denoting ,
Choose so that . is thus bounded in time, but the bound is not tractable except for (starting with ) or if . In both cases we have obtained
It remains to optimize in . In full generality we choose such that both terms in the sum of (5.15) are equal (we know that we are loosing a factor less than ). Remark that we do not use the explicit form of , i.e. we may replace (H..K) by (H..K) in what we did previously. We have thus obtained
Assume that satisfies (H..K) for . Let be the inverse (reciprocal) function of . Denote and . Then for all , satisfies for all nice ,
where is the Fisher information of .
When is equal to identity, such an inequality is called a inequality (see definition 10.4). Here we obtained a weak form of inequality (which is clear on (5.15)) in the spirit of the weak Poincaré or the weak log-Sobolev inequalities.
In particular, using (H.W.I), since (H.C.0) is satisfied we obtain
Under the hypotheses of proposition 5.16, satisfies the inequality
Weak logarithmic Sobolev inequalities were introduced and studied in . Actually, we are not exactly here in the situation of because we wrote the previous inequality in terms of a density of probability. Let . We deduce from the previous corollary
so that if ,
In this sub(sub)section we assume that for some , so that . We thus have
Recall first that if , then (see e.g. (2.6)). Next recall the following: defining as a median of , we have
We may decompose so that both and are non negative with median equal to . In addition, if is Lipschitz, so are and , , and the product of both vanishes. Hence
similarly for . We have thus obtained
Assume that satisfies (H..K) for and , for . Then, satisfies both a Poincaré inequality with
This theorem applies in particular to for , according to proposition 5.4. The fact that satisfies a log-Sobolev inequality in this situation is well known, but here we obtain an explicit (though not really cute) expression for the constant that only depends on and not on the dimension . Unfortunately, in this particular situation, our bounds are not optimal. Indeed, spherically symmetric log-concave probability measures are now well understood.
For the Poincaré constant, it was shown by Bobkov that
It is an (easy) exercise to see that , so that which goes to as . A famous conjecture by Kannan-Lovasz-Simonovitz is that the previous bound for spherically symmetric measures extends (up to a change of the constant ) to any log-concave measure. If true, the KLS conjecture will presumably give a better upper bound for the Poincaré constant than ours.
Regarding the log-Sobolev constant, the work by Huet , furnishes a lower bound for the isoperimetric profile of (see Theorem 3 and the discussion p.98 therein) which indicates a similar bound for the log-Sobolev constant as above, i.e. depending on the isotropic constant () of .
2.2. Lack of uniform convexity.
Now choose for some and , and for . (H..K) is less restrictive than before since it only implies a linear behavior at infinity for the gradient of the potential.
We now have for and for . It follows
if and
Assume that satisfies (H..K) for and for . Then, satisfies both a Poincaré inequality with
Using the general form of the (H.W.I) inequality it is quite easy to adapt the previous proof in order to show the following result : Let satisfying (H.C.K) for some . If satisfies a weak inequality,
for some , then satisfies a log-Sobolev inequality with a constant depending on only. In particular satisfies a inequality.
In particular if we know that has a bounded below curvature, the previous theorems extend to .
Using reflection coupling.
As we have seen in the previous section, the simple coupling using the same Brownian motion is not fully well suited to deal with non uniformly convex potentials. In a recent note , Eberle studied the contractivity property in Wasserstein distance , induced by another well known coupling method: coupling by reflection (or mirror coupling) introduced in . We shall see now how to use this coupling method in the spirit of what we have done before.
In this section we consider the solution starting from of the Ito stochastic differential equation
where is smooth enough. We introduce another formulation of the semi-convexity property, namely :
We shall say that (H.) is satisfied if
This condition is typically some “uniform convexity at infinity” condition. Indeed if where with satisfying (H.C.) and compactly supported, then (H.) is satisfied. We shall come back later to this. Notice that if (6.3) is satisfied, the solution of (6.1) is strongly unique and non explosive, using the same tools as we used before. Now, following we introduce (with some slight change of notations)
If (H.) is satisfied, so that and
i.e. which is actually a distance, is equivalent to the euclidean distance. Hence a consequence of Theorem 1 in is the following :
Assume that (H.) is satisfied. Let be defined by
Then for all initial distributions and , and all , the Wasserstein distance satisfies
In order to prove this result, Eberle adds to (6.1) the following Ito s.d.e.
where and is the transposed of (remark that if , it just changes into ). Of course one has to consider (6.1) and (6.6) together. Existence, strong uniqueness and non explosion are again easy to show. Now introduce the coupling time defined by
It is easy to see that and have the same law, so that the distribution of is a coupling of and . Of course this extends to any initial distributions and furnishes a coupling of .
It follows that solves
where is a one dimensional brownian motion.
The key of the proof of Theorem 6.5 is then that, if , is a semi-martingale with decomposition
Taking expectation, it immediately shows that the Wasserstein distance decays exponentially fast. As remarked by several authors, one can then deduce as we did previously
Assume that satisfies(H.) and that is a probability measure. Then satisfies a Poincaré inequality with constant .
for all Lipschitz function . According to Lemma 2.12 in , we deduce that , hence the result. ∎
One can see that the reflection coupling cannot furnish some information on , the concavity of near the origin being crucial. In the same negative direction, theorem 6.11 cannot be extended to the log-Sobolev framework, the key lemma 2.12 in being restricted to the variance control.
If (H..K) is satisfied with , we have . We thus have , , and finally
We recover the result in Theorem 5.20, i.e. a bound but with a better constant .
If (H..K) is satisfied, one can choose , , and finally
Now assume that the potential can be written where satisfies (H.C.K) for some and satisfies . It easily follows that (H.) is satisfied with . We thus have
An old result by Miclo (unpublished but explained in ) indicates that such a result (without the square of the supremum of the gradient but without in the exponential) can be obtained by using the usual Holley-Stroock perturbation argument.
2. The log-concave situation.
Now consider the situation where satisfies (H.C.0). In this situation we have (, ) so that (6.8) becomes
for with . It follows that up to the first time the brownian motion hits . In particular,
Actually if with convex (i.e. in the zero curvature situation of the theory) the inequality is well known as a consequence of what is called the reverse (local) Poincaré inequality (see ). The previous proposition extends this result (up to the constant) to a non-gradient drift.
Recall that for . It follows
If one wants to get a contraction bound for the gradient (in the spirit of (6.10) or better of proposition 2.2) we cannot only use a comparison with the brownian motion for which .
In the symmetric situation () it is known that is non increasing. It easily follows that
If we assume that (H.) is satisfied, we may replace the comparison with a Brownian motion by the comparison with an Ornstein-Uhlenbeck process with parameter , according to standard comparison theorems for one dimensional Ito processes (see e.g Chapter VI theorem 1.1). For the O-U process, it is known (see ) that
These bounds are interesting as regularization bounds (from bounded to Lipschitz functions), but notice that we have lost a factor in the exponential decay.
3. Reflection coupling for general diffusions.
The case of a general diffusion process with a non constant diffusion matrix as in section 3 is more delicate to handle, as already remarked in (Theorem 1).
Assume that is a bounded and smooth square matrices field and that it is uniformly elliptic. The quantities we need here are (notations differ from )
Recall the Lindvall-Rogers reflection coupling
Existence and strong uniqueness can be shown as previously. Of course, as in subsection 6.1, we replace by if the coupling time, but not to introduce new notation we still use .
Applying Ito formula we thus have for a smooth function
We introduce the natural generalization of (H.), namely we assume that
If is a non decreasing, concave function we thus get, provided ,
Hence looking carefully at the calculations in , we see that, provided , the only thing we have to change in (6.4) is the definition of replacing by the inverse of , all other definitions being unchanged. We have thus obtained
Assume that (6.17) and (6.18) are satisfied. Assume in addition that . Then defining
the conclusion of Theorem 6.5 is still true with .
All the consequences of Theorem 6.5 still hold (up to the modifications of the constants), in particular one can extend (H.C.K) to the situation of “convexity at infinity” as in Example 6.13 (3). Details are left to the reader.
As we already said, the condition already appears in and ensures that the coupling by reflection is succesfull. Roughly speaking it means that the fluctuations of are not too big with respect to the uniform ellipticity bound.
4. Gradient commutation property and reflection coupling
It is of course quite disappointing at first glance that the only gradient commutation property that we get using this nice contraction results in distance, is restricted to Lipschitz function as in (6.10). Let us see however that we may transfer this to stronger gradient commutation properties in some cases. The main tool is the following lemma on Hölder’s type inequality in Wasserstein distance.
The proof is indeed quite simple and relies mainly on Hölder’s inequality. Indeed, in dimension one the optimal transport plan is the same for every convex cost (see for example Villani ), so that there exists a transport plan such that
The case of product probability measure is deduced using the result in dimension one and the following two direct assertions
and if and have for marginal and
In fact, as will be seen from our applications, even if does depend of and (in a nice way), it would be sufficient to get new gradient commutation property.
We are now in position to prove various gradient commutation properties in non standard cases. For simplicity, we suppose here that the diffusion coefficient is constant, i.e.
so that the weak gradient commutation property holds
and thus a local Poincaré inequality holds.
Note that this theorem is the first one to give the commutation gradient property in non strictly convex case with a good behaviour at infinity.
Using synchronous coupling as previously explained and the fact that we have that
In the same time, by Theorem 6.5, we have that
We then use Lemma 6.21 to get the first assertion. The second one is obtained as in Proposition 3.11. ∎
Consider for example the log-concave case for which Bakry-Emery theory enables us to get that we are in 0-curvature and thus
However, using Theorem 6.23 and this last inequality, we easily get that there exists such that
which is completely new. It captures both the short time behavior equivalent to the 0-curvature criterion and the long time behavior for which and thus is expected to decay to 0. Note that we may extend this example to a double well potential, in the case when the height of the well is not too large.
Preserving curvature.
A natural question about curvature is the following: is curvature preserved by a diffusion process ? According to a result by Kolesnikov , the Ornstein-Uhlenbeck process is essentially the only one, among diffusion processes, preserving log-concavity (i.e. if is log-concave, so is for all ). One may also wonder if may satisfy other “curvature” like inequality as for example. It would have important applications on local inequalities, indeed transportation inequalities together with a HWI inequality may imply logarithmic Sobolev inequality.
In the spirit of the previous remark, consider a standard Ornstein-Uhlenbeck process , i.e. the solution of
where and are independent random variables, being a standard gaussian variable and having distribution . Hence
or, if we use the notation for a random variable with distribution ,
But if is a random variable it is clear that . It follows
The change of variable , yields, using the symmetry of
In particular, for (), we may let go to and obtain
In particular, the distribution of satisfies a log-Sobolev inequality if and only if, for some , the distribution of satisfies a log-Sobolev inequality , and then the considered inequality is satisfied for all . Recall that and are independent. This is not surprising since a more general result can be obtained directly (extending a similar result for the Poincaré inequality in ):
Let X and Y be independent random variables and then,
Conversely, if Y is symmetric (i.e. Y and -Y have the same distribution), we also have
The first result for is proved in proposition 1. For the proof is very similar. Let be smooth. Then
Since , we may use Cauchy-Schwartz inequality in order to bound the last term in the latter sum. This yields exactly the desired result.
For the second statement, it is enough to use a change of variable, as we did in the gaussian case and the symmetry of . ∎
Hence we get a general statement: the distribution of a random variable satisfies a Poincaré or a log-Sobolev inequality if and only if, for all or for one, symmetric random variable independent of whose distribution satisfies a Poincaré or a log-Sobolev inequality, the distribution of satisfies a Poincaré or a log-Sobolev inequality.
It is known that any log-concave probability measure satisfies some Poincaré inequality. The result is due to Bobkov (a short proof is contained in ). But if is a log-concave random variable that do not satisfy a log-Sobolev inequality, (which is still log-concave according to the Prekopa-Leindler theorem) does not satisfy a log-Sobolev inequality, in particular is not uniformly log-concave.
Here is an amusing proof of the consequence of Prekopa’s result when one variable is gaussian.
Let (resp. ) be a random variable with law (resp. a standard gaussian variable). We assume that and are independent. The density of is thus given by
Let be the hessian matrix of . Then
(or if one prefers its potential) satisfies (H.C.). If , it thus satisfies a Poincaré inequality with constant . Applying this Poincaré inequality to the function , we obtain
Thanks to simple scales we may thus state
Appendix A Some general remarks.
We give in this section some general facts which we used from place to place.
First we recall two facts one can find for instance in remark 2.11 and in lemma 2.12:
Accordingly, since , a control will furnish a Poincaré inequality. Notice that if for some , and all Lipschitz functions , then using the semi-group property we get for some and all , hence a Poincaré inequality. The latter is a weaker form of the commutation of the gradient and the semi-group up to an exponential rate.
One can use the previous remarks to show that an exponential decay of Wasserstein distances furnishes some Poincaré inequality for . In what follows denotes the total variation distance and is the usual -Wasserstein distance.
Assume that for all bounded (resp. Lipschitz) density of probability we have (reps. ). Then for all bounded (resp. Lipschitz and bounded) , there exist and such that . In particular if , satisfies a Poincaré inequality.
is thus a density of probability with . We have
One can replace by , just replacing by in which case
Even when the decay is not exponential, one gets a weak form of the Poincaré inequality (called a weak Poincaré inequality).
In the situation of proposition A.2, assume that with as . Then
for all where for and for and .
Once we notice that the transformation does not change the result follows from Theorem 2.3. ∎
Acknowledgements: This research was supported by the French ANR project STAB.