An entropic interpolation proof of the HWI inequality
Ivan Gentil, Christian Léonard, Luigia Ripani, Luca Tamanini
Introduction
In a seminal paper , Otto and Villani obtained a powerful functional inequality relating the relative entropy with respect to some reference measure , the quadratic transport cost and the Fisher information . This so-called HWI inequality roughly states that: where the real parameter is a curvature lower bound associated to , see (1.2), (1.4) and Theorem 1.6 below for the exact statement and its well-known consequences in terms of Talagrand and logarithmic Sobolev inequalities.
The first part of Otto and Villani’s article is dedicated to a heuristic proof based on Otto calculus, see , where one formally equips the set of probability measures with the Riemannian-like distance and where McCann displacement interpolations are interpreted as geodesics. Since these interpolations suffer from a lack of regularity, the first and second order time derivatives along them are only formal. Consequently, although heuristics led to the right conjecture, the authors presented an alternative rigorous proof based on a significantly different approach.
In the present article, a new proof of the HWI inequality is proposed. The main idea is to replace the ‘irregular’ McCann interpolation between two probability measures and by a family of ‘smooth’ curves of measures , called ‘entropic interpolations’ (see Definition 2.3 below), where is a small fluctuation parameter such that converges to narrowly as . Otto and Villani’s heuristics apply rigorously to , so that it remains to let tend down to zero to obtain the desired result.
The paper is structured as follows. Section 1 is dedicated to the statement of the HWI inequality. Basic material about entropic interpolations which is needed for the proof is gathered at Section 2. Finally the proof of the inequality is done at Section 3; its core is Lemma 3.11 which is the analogue of Otto and Villani’s heuristic approach. In Section 4 some comments about possible extensions and simplifications of our approach are collected.
Statement of the HWI inequality
Before stating the HWI inequality at Proposition 1.5 and Theorem 1.6 below, we need to make clear the framework we shall work within and introduce the quantities , and .
or , where is a smooth Riemannian manifold without boundary and with metric tensor , is the induced distance and is given by
Relative entropy
For any two probability measures and on a measurable space the relative entropy of with respect to is defined by
where it is understood that this quantity is infinite when is not absolutely continuous with respect to . In our case, will be , or .
Quadratic transport cost
Fisher information
The Fisher information of with respect to is defined by
and otherwise. Up to identify with its density, the Fisher information is lower semicontinuous with respect to the weak topology of (see for instance ).
With this premise, the statement of the HWI∗ inequality is
Let be as in Setting 1. Then for any such that ,
In this case it is said that the reference measure satisfies the HWI inequality.
As already shown by Otto and Villani in , different choices of and in Proposition 1.5 entail three important consequences collected here below.
Let be as in Setting 1 with the further assumption that . Then the following inequalities are satisfied:
Logarithmic Sobolev inequality: assume that , then
First of all, since , it follows that is finite for all .
The HWI inequality is obtained by choosing .
The Talagrand inequality is obtained by choosing .
When , the logarithmic Sobolev inequality with follows by taking the supremum with respect to in the right-hand side of the HWI inequality (a). To extend this result to the case where , a standard approximation argument (carried out for instance in Lemma 3.1 below) is sufficient.
When , any such that stands in . This follows from the variational representation of the relative entropy as
whence the claim. In particular, .
Talagrand inequality (b) is irrelevant when . When in view of previous remark it extends to all provided that one sets when does not belong to .
It follows from the logarithmic Sobolev inequality that when (1.2) or (1.4) holds with some , any such that satisfies . Similarly, it follows from the HWI inequality that when (1.2) or (1.4) is only supposed to hold with real, as soon as , then implies that .
Entropic interpolations
In this section we propose a short and self-contained presentation of entropic interpolations. The purpose is twofold: to provide the reader with those notions and results that will be frequently used later on and discuss their physical interpretation via Nelson’s dynamical view of diffusion processes. For sake of simplicity, the latter will be carried out in the more familiar Euclidean setting.
Let , with be the canonical process, defined by
For any path measure and each , we denote by the -th marginal of , that is the law of the position at time of the random path under . Moreover, for any , we shall denote by the joint law of and under , namely .
As reference path measure we consider the law of the Markov diffusion process with generator
with initial law , where the potential appears at (1.1), (1.3), is the Laplace-Beltrami operator on and the Levi-Civita connection associated to the metric (in Setting 1-(a) they are nothing but the standard Laplacian and gradient). It is well-known that is a reversible Markov measure with reversing measure . In particular it is stationary, that is for all . For any we denote by the time-rescaled process defined by , , and by the corresponding path measure. The parameter is meant to tend to zero so that is a slowed down version of , whose generator is
For any , as a time rescaling of , is also -reversible, so that in particular for all .
The 1-parameter semigroup associated to will be denoted by and, in a completely analogous way, the one associated to ; notice that for all . Within Setting 1 it is well-known (see for instance ) that there exists a unique heat kernel associated to which is a smooth function on . Therefore, the semigroup can be represented by
for all . Let us also recall that, in conjunction with (1.2) or (1.4), enjoys the Bakry-Émery contraction estimate
For its proof as well as for all the regularizing properties of that will be used throughout the paper, we address the reader to .
defined for all , are naturally associated to . As it is not difficult to see, the drift in does not affect , since . It is worth mentioning that, with respect to the standard definition provided in , here and are not divided by 2, as the factor 1/2 already appears in , which thus corresponds to an SDE driven by a standard Brownian motion.
The Schrödinger problem
Let be two probability measures: the Schrödinger problem associated with is defined by
and its value is called ‘entropic cost’. As a strictly convex minimization problem, it admits at most one solution.
The solution of (Sε), if it exists, is called the -entropic bridge between and . The -entropic interpolation between and is defined as the time marginal flow of the solution , namely
The name ‘entropic interpolation’ stems from the connection with displacement interpolation. Indeed, it is known from that
This limit is a consequence of a more general result asserting that (Sε) converges to the quadratic Monge-Kantorovich problem as in the sense of -convergence, see . As shown in , if (Sε) admits the solution , then there exist two non-negative measurable functions such that and defining
usually known as ‘Schrödinger system’: indeed, if we interprete them as a nonlinear system where the unknowns are and , then the (unique up to an obvious multiplicative rescaling) solution completely determines (see for instance ). As far as the convergence of entropic interpolations towards displacement ones is investigated, the following functions
are of special interest. We also set in and in . They are called Schrödinger potentials, in connection with Kantorovich ones.
Existence and regularity results
Let us now derive a criterion, in terms of the endpoint marginals and , for the existence of regular functions with well-defined Fisher information solving the Schrödinger system (2.6). As noticed in , the regularity (smoothness and integrability) of (resp. ) is inherited by (resp. ). In the next result we extend this property: if is finite, then so is and analogously with .
Let be as in Setting 1, and consider two probability measures with compact supports. Then the following hold:
The Schrödinger problem (Sε) admits a solution with finite entropy if and only if , . This solution is unique and the -entropic interpolation between and exists.
Suppose in addition that .
Then .
For any the functions as well as and belong to .
In addition to item (b), suppose that is finite (resp. ). Then so is (resp. ).
For (a) and (b-i) see [13, Proposition 2.1]. As regards (b-ii), the fact that , and (2.1) imply for all , while the maximum principle ensures that ; the statements for and follow by the same reason.
(b-iii). Notice that the first equation in the Schrödinger system (2.6) can be rewritten as
Since is positivity improving, is smooth and with compact support, the conclusion follows. A similar argument holds for .
(c). Observe that the Schrödinger system (2.6) and (as is positivity improving) force to have the same support as . Thus has compact support and, as a consequence, in for some . This remark and the chain rule allow us to say that
so that it remains to prove the integrability of the right-hand side. The first term is integrable by assumption, the third one by the regularization properties of , while for the second one notice that and imply ; plugging this information into
For this reason we formulate the following
The endpoint marginals are such that have compact supports, and their densities belong to .
A dynamical viewpoint
As concerns the evolution of entropic interpolations and Schrödinger potentials, let us first notice that under Assumptions 2.8, by the very definition (2.5), item (b-iii) of Proposition 2.7 and the fact that we deduce that and are smooth on and solve
in the classical sense. Moreover, if we look at as curves parametrized by with values in , they belong to the set of all absolutely continuous functions from $W^{1,2}({\rm X},\mathfrak{m})\partial_{t}f_{t}^{\varepsilon},\partial_{t}g_{t}^{\varepsilon}W^{1,2}$-limits.
Relying on that, it follows that the Schrödinger potentials are smooth on and and solve forward and backward Hamilton-Jacobi-Bellman equations respectively, i.e.
while for the continuity equation
is satisfied in , where and denotes the divergence with respect to , i.e. the opposite of the adjoint of the differential in . This last PDE is strongly linked to the dynamical representation of the entropic cost , namely
shown in for Setting 1-(a) and in for rather general metric measure spaces including Setting 1-(b). By the very definition of and since , this implies
The regularity of Schrödinger potentials comes from the one of , the fact that are everywhere positive and the logarithm is smooth on . However, Schrödinger potentials are not integrable in general, as can be arbitrarily close to 0. Thus we cannot study their behaviour as curves with values into some space. This also explains why the regularity of (resp. ) does not extend up to (resp. ).
A physical interpretation
In the Euclidean framework of Setting 1-(a) it is possible to make a bridge between what is presented so far and Nelson’s formalism , thus providing a physical motivation for some results stated above and a further perspective on some objects.
For , it is easily seen that
This allows to rewrite the continuity equation (2.11) as
where now denotes the divergence with respect to if in Setting 1-(a) or if in Setting 1-(b). This is perfectly coherent with (2.11) since for any vector field . Furthermore, (2.12) becomes
Proof of Proposition 1.5
We need to state some preliminary lemmas before completing the proof of Proposition 1.5 at page 3. Throughout the whole section we shall assume to work within as in Setting 1.
Let us start with an approximation result.
Let with . Then:
there exists a sequence with and such that for all and as ;
if in addition , then there exists a sequence with , such that for all , and as .
Let us write and, as a first step, let us prove that both in (a) and (b) it is possible to find a sequence of measures with smooth densities converging to in the desired sense. This can be proved by defining for
which clearly have smooth densities by the regularizing properties of . The convergence of , and to , and respectively as is now a well-known fact in the theory of gradient flows (see for instance Theorem 2.4.15 and Remark 2.4.16 in in conjunction with the fact that the squared slope of the entropy is the Fisher information, as proved in ).
Thus, it is not restrictive to suppose that has smooth density. Under this new assumption, let us prove that we can find a sequence of measures with compact supports and smooth densities converging to in the desired sense. To this aim define
where is the renormalization constant and is a smooth cut-off function with support in , for some , and Lipschitz constant controlled by , where does not depend on (see e.g. for a proof of the existence of such cut-off functions). By dominated convergence it is not difficult to see that and thus for all as ; for the same reason . If we also assume that , then
Since the right-hand side is integrable and as , by dominated convergence we get .
Combining the two steps and using a diagonal argument, the conclusion follows. ∎
The following conservation result was pointed out in and in the case is compact with different approaches; see also . Following , we extend the statement to the present framework.
Under Assumptions 2.8, for any the function
is real-valued and constant. Thus we shall denote it by .
As a first step, for all by algebraic manipulation we have
From integration by parts formula it is straightforward to see that the right-hand side vanishes, whence the conclusion. ∎
Motivated by (2.12), let us investigate separately the convergence of current and osmotic velocities as .
Under Assumptions 2.8, for any we have
Let us first notice that combining (2.12) and (2.4) we get
Since the continuity equation (2.11) is satisfied by , the Benamou-Brenier formula holds for any , that is
and together with the above limit, this leads us to the first identities both in (3.4) and (3.5). From them we immediately deduce that
whence also the second identity in (3.5). Finally observe that
As already noticed at Remark 2.14, although it is smooth on the open interval , the density might be arbitrarily close to zero. Consequently, the Schrödinger potentials might not be integrable enough, and the Fisher information might behave badly around and . Next lemma provides related regularity and growth controls. Its proof is strongly inspired by , , and ; we thus address the reader to these articles for more details.
Under Assumptions 2.8, let and set, for all , as well as
Then all the functions so defined belong to and, as curves parametrized by with values in , to . The time derivatives of are given by
where , , have to be understood both in the classical sense and as strong -limits.
As already explained in Section 2, under Assumptions 2.8 and . Therefore the regularity and integrability properties of are a straightforward consequence of the chain rule and of the fact that the logarithm is smooth with bounded derivatives on . Also the PDEs solved by are easily deduced, when interpreted in the classical sense, as they follow from (2.10) and (2.11). In order to deduce the same identities with , , seen as strong -limits, notice that by the maximum principle
whence . Moreover, the smoothness of the logarithm, the chain and Leibniz rules entail that
whence ; analogous estimates hold for and . These bounds together with the fact that the PDEs in (3.8) hold in the classical sense imply, by a dominated convergence argument, that (3.8) are satisfied also as strong -limits.
Relying on that, (3.9a) and (3.9b) follow by the computations carried out in . Indeed, the validity of (3.8) as strong -limits on $t\mapsto u(\rho_{t}^{\varepsilon,\delta})t\mapsto\langle\nabla\rho_{t}^{\varepsilon,\delta},\nabla\vartheta_{t}^{\varepsilon,\delta}\rangleAC(,L^{2}(\mathfrak{m}))$ and thus we can pass the time derivatives under the integral sign, i.e.
Then the Hamilton-Jacobi-Bellman equations for the Schrödinger potentials and the continuity equation for the entropic interpolation together with , here replaced by , are the only tools needed to deduce (3.9a) and (3.9b). Finally, the fact that:
;
, belong to and ;
imply the continuity on $$ of the right-hand sides of (3.9a) and (3.9b) respectively. ∎
With this results at disposal we can prove our main lemma: a rigorous ‘entropic’ analogue of Otto-Villani’s heuristic argument.
Under Assumptions 2.8, for any it holds
The proof of the lemma is based on the standard calculus identity
consequence of the lower Ricci bounds (1.2) and (1.4) rewritten in the form of the Bochner-Lichnerowicz-Weitzenböck formula, we obtain
where . It is now sufficient to pass to the limit as . By dominated convergence (recall that ) it is easy to see that the left-hand side converges to . In order to apply the dominated convergence theorem also to the first term on the right-hand side, notice that
Thus, to control let us first observe that
and remark that the right-hand side is integrable by Proposition 2.7-(c); secondly
and in this case the right-hand side is integrable as has compact support and is bounded away from 0 therein; finally,
as and the same strategy applies to all the remaining terms that we obtain developing . Therefore, the first term on the right-hand side of (3.14) converges to the first term on the right-hand side of (3.12). As regards the other two summands, by the very definition of and since
the conclusion will follow if we are able to prove that
To this aim, notice that , whence using either or it is easy to infer that
All the right-hand sides above are integrable on (either by (2.13) or by the Bakry-Émery contraction estimate (2.2) together with ), thus by dominated convergence (3.15a) follows. An analogous argument holds for (3.15b), whence the conclusion. ∎
Completion of the proof of Proposition 1.5
We are now ready to complete the proof of the HWI* inequality.
First of all, by Proposition 2.7 and Lemma 3.1, it is sufficient to prove the result for any satisfying Assumptions 2.8. This allows us to invoke Lemma 3.11. Secondly, observe that by the Cauchy-Schwarz inequality
so that if we plug this information into (3.12) we obtain
Let us now pass to the limit as : by Lemma 3.2, at we have
so that (3.6) together with yields
By Lemma 3.3 we can pass to the limit as also in the remaining terms on the right-hand side, thus concluding. ∎
Let us mention that another approach to prove the HWI inequality via the Schrödinger problem is pointed out in a recent work where an entropic counterpart of the HWI inequality is formally obtained by differentiating the convexity estimate of the entropy along the entropic interpolations introduced in [6, Thm. 1.4].
Final remarks and comments
From a heuristic point of view, the expression of the constant quantity can be deduced by standard arguments in Lagrangian and Hamiltonian formalism. Indeed, motivated by (2.12) let us consider the action functional
By means of Legendre’s transform, the corresponding Hamiltonian is given by
and, at least formally, is constant along the critical points of . Since and are prescribed, the Euler-Lagrange equation for (4.1) reads as
and, as it is not difficult to see (e.g. following the computations carried out in which fit to Setting 1), these PDEs are satisfied along . Finally, as in the Hamiltonian represents a momentum density, it is natural to set . From these considerations, the guess on the existence of the conserved quantity of Lemma 3.2 and on its expression follows. This point of view is particularly investigated in the recent paper , to which we refer for more details.
The compact case
As already mentioned in Remark 2.14, in general Schrödinger potentials fail to be smooth in or and are not even integrable for any , because can be arbitrarily close to 0. This problem can be overcome if for some constant : then as well by the maximum principle and from their very definition. However, even if we assume to have densities bounded away from 0 (thus forcing to belong to , a condition which does not follow from our Setting 1, unless , cf. Remark 1.7-(a)), it is not known yet whether this implies an analogous lower bound on and . This is the motivation behind the definitions provided in Lemma 3.7.
On the contrary, if is assumed to be compact, then it is not restrictive to assume for some : indeed any can be approximated by
(where is the renormalization constant) in , entropy and Fisher information in the sense of Lemma 3.1. Secondly, is bounded from above since . These two facts imply that , as shown in . As a consequence the proof of the HWI inequality is easier and the parallelism with the heuristic proof of Otto-Villani is stronger, since no ‘-argument’ in Lemma 3.7 and Lemma 3.11 is needed; in particular, Lemma 3.7 is already known to hold by and thus no proof is required.
𝖱𝖢𝖣𝖱𝖢𝖣{\sf RCD} spaces
All the results presented in this paper, with particular mention of the key Lemma 3.11 and Proposition 1.5, are also true in a more general framework than the one of Setting 1-(b), namely in spaces (introduced in ). If is assumed to be compact, then this has been shown by the last author in his PhD thesis . When is not compact, the most important steps in our entropic approach still hold. Namely:
all the regularity and integrability results concerning Schrödinger potentials and entropic interpolations mentioned in Section 2 as well as the dynamic representation of the entropic cost;
the regularizing and contraction properties of ;
the existence of ‘good’ cut-off functions;
the Benamou-Brenier formula and the Bochner-Lichnerowicz-Weitzenböck inequality.
The reader is addressed to , for the first point and to for all the others.
Acknowledgements. This research was supported by the French ANR-17-CE40-0030 EFI project and the LABEX MILYON ANR-10-LABX-0070.