An entropic interpolation proof of the HWI inequality

Ivan Gentil, Christian Léonard, Luigia Ripani, Luca Tamanini

Introduction

In a seminal paper , Otto and Villani obtained a powerful functional inequality relating the relative entropy H(⋅ ∣ m)H(\cdot\,|\,\mathfrak{m}) with respect to some reference measure m∈P2(X)\mathfrak{m}\in\mathcal{P}_{2}({\rm X}), the quadratic transport cost W22(⋅,m)W_{2}^{2}(\cdot,\mathfrak{m}) and the Fisher information I(⋅ ∣ m)I(\cdot\,|\,\mathfrak{m}). This so-called HWI inequality roughly states that: H≤WI−κW2/2,H\leq W\sqrt{I}-\kappa W^{2}/2, where the real parameter κ\kappa is a curvature lower bound associated to m\mathfrak{m}, see (1.2), (1.4) and Theorem 1.6 below for the exact statement and its well-known consequences in terms of Talagrand and logarithmic Sobolev inequalities.

The first part of Otto and Villani’s article is dedicated to a heuristic proof based on Otto calculus, see , where one formally equips the set of probability measures with the Riemannian-like distance W2W_{2} and where McCann displacement interpolations are interpreted as geodesics. Since these interpolations suffer from a lack of regularity, the first and second order time derivatives along them are only formal. Consequently, although heuristics led to the right conjecture, the authors presented an alternative rigorous proof based on a significantly different approach.

In the present article, a new proof of the HWI inequality is proposed. The main idea is to replace the ‘irregular’ McCann interpolation (μt)(\mu_{t}) between two probability measures μ0\mu_{0} and μ1\mu_{1} by a family of ‘smooth’ curves of measures (μtε)(\mu_{t}^{\varepsilon}), called ‘entropic interpolations’ (see Definition 2.3 below), where ε>0\varepsilon>0 is a small fluctuation parameter such that μtε\mu_{t}^{\varepsilon} converges to μt\mu_{t} narrowly as ε↓0\varepsilon\downarrow 0. Otto and Villani’s heuristics apply rigorously to (μtε)(\mu_{t}^{\varepsilon}), so that it remains to let ε\varepsilon tend down to zero to obtain the desired result.

The paper is structured as follows. Section 1 is dedicated to the statement of the HWI inequality. Basic material about entropic interpolations which is needed for the proof is gathered at Section 2. Finally the proof of the inequality is done at Section 3; its core is Lemma 3.11 which is the analogue of Otto and Villani’s heuristic approach. In Section 4 some comments about possible extensions and simplifications of our approach are collected.

Statement of the HWI inequality

Before stating the HWI inequality at Proposition 1.5 and Theorem 1.6 below, we need to make clear the framework we shall work within and introduce the quantities HH, WW and II.

or (M,dg,m)(M,{\sf d}_{g},\mathfrak{m}), where MM is a smooth Riemannian manifold without boundary and with metric tensor gg, dg{\sf d}_{g} is the induced distance and m\mathfrak{m} is given by

Relative entropy

For any two probability measures pp and rr on a measurable space ZZ the relative entropy of pp with respect to rr is defined by

where it is understood that this quantity is infinite when pp is not absolutely continuous with respect to rr. In our case, ZZ will be X{\rm X}, X×X{\rm X}\times{\rm X} or C(,X)C(,{\rm X}).

Quadratic transport cost

Fisher information

The Fisher information of μ∈P(X)\mu\in\mathcal{P}({\rm X}) with respect to m\mathfrak{m} is defined by

and +∞+\infty otherwise. Up to identify μ\mu with its density, the Fisher information is lower semicontinuous with respect to the weak topology of L1(m)L^{1}(\mathfrak{m}) (see for instance ).

With this premise, the statement of the HWI∗ inequality is

Let (X,d,m)({\rm X},{\sf d},\mathfrak{m}) be as in Setting 1. Then for any μ0,μ1∈P2(X)\mu_{0},\mu_{1}\in\mathcal{P}_{2}({\rm X}) such that H(μ0 ∣ m)<∞H(\mu_{0}\,|\,\mathfrak{m})<\infty,

In this case it is said that the reference measure m\mathfrak{m} satisfies the HWI inequality.

As already shown by Otto and Villani in , different choices of μ0\mu_{0} and μ1\mu_{1} in Proposition 1.5 entail three important consequences collected here below.

Let (X,d,m)({\rm X},{\sf d},\mathfrak{m}) be as in Setting 1 with the further assumption that m∈P2(X)\mathfrak{m}\in\mathcal{P}_{2}({\rm X}). Then the following inequalities are satisfied:

Logarithmic Sobolev inequality: assume that κ>0\kappa>0, then

First of all, since m∈P2(X)\mathfrak{m}\in\mathcal{P}_{2}({\rm X}), it follows that W2(ν,m)W_{2}(\nu,\mathfrak{m}) is finite for all ν∈P2(X)\nu\in\mathcal{P}_{2}({\rm X}).

The HWI inequality is obtained by choosing μ0=m,μ1=ν\mu_{0}=\mathfrak{m},\mu_{1}=\nu.

The Talagrand inequality is obtained by choosing μ0=ν,μ1=m\mu_{0}=\nu,\mu_{1}=\mathfrak{m}.

When κ>0\kappa>0, the logarithmic Sobolev inequality with ν∈P2(X)\nu\in\mathcal{P}_{2}({\rm X}) follows by taking the supremum with respect to W2W_{2} in the right-hand side of the HWI inequality (a). To extend this result to the case where ν∈P(X)\nu\in\mathcal{P}({\rm X}), a standard approximation argument (carried out for instance in Lemma 3.1 below) is sufficient.

When κ>0\kappa>0, any ν∈P(X)\nu\in\mathcal{P}({\rm X}) such that H(ν ∣ m)<∞H(\nu\,|\,\mathfrak{m})<\infty stands in P2(X)\mathcal{P}_{2}({\rm X}). This follows from the variational representation of the relative entropy as

whence the claim. In particular, m∈P2(X)\mathfrak{m}\in\mathcal{P}_{2}({\rm X}).

Talagrand inequality (b) is irrelevant when κ≤0\kappa\leq 0. When κ>0,\kappa>0, in view of previous remark it extends to all ν∈P(X)\nu\in\mathcal{P}({\rm X}) provided that one sets W2(ν,m)=∞W_{2}(\nu,\mathfrak{m})=\infty when ν\nu does not belong to P2(X)\mathcal{P}_{2}({\rm X}).

It follows from the logarithmic Sobolev inequality that when (1.2) or (1.4) holds with some κ>0\kappa>0, any ν∈P(X)\nu\in\mathcal{P}({\rm X}) such that H(ν ∣ m)=∞H(\nu\,|\,\mathfrak{m})=\infty satisfies I(ν ∣ m)=∞I(\nu\,|\,\mathfrak{m})=\infty. Similarly, it follows from the HWI inequality that when (1.2) or (1.4) is only supposed to hold with κ\kappa real, as soon as ν∈P2(X)\nu\in\mathcal{P}_{2}({\rm X}), then H(ν ∣ m)=∞H(\nu\,|\,\mathfrak{m})=\infty implies that I(ν ∣ m)=∞I(\nu\,|\,\mathfrak{m})=\infty.

Entropic interpolations

In this section we propose a short and self-contained presentation of entropic interpolations. The purpose is twofold: to provide the reader with those notions and results that will be frequently used later on and discuss their physical interpretation via Nelson’s dynamical view of diffusion processes. For sake of simplicity, the latter will be carried out in the more familiar Euclidean setting.

Let X=(Xt)0≤t≤1X=(X_{t})_{0\leq t\leq 1}, Xt:Ω→XX_{t}:\Omega\to{\rm X} with Ω:=C(,X)\Omega:=C(,{\rm X}) be the canonical process, defined by

For any path measure Q∈P(Ω){\sf Q}\in\mathcal{P}(\Omega) and each 0≤t≤10\leq t\leq 1, we denote by Qt:=(Xt)#Q∈P(X){\sf Q}_{t}:=(X_{t})_{\#}{\sf Q}\in\mathcal{P}({\rm X}) the tt-th marginal of Q{\sf Q}, that is the law of the position XtX_{t} at time tt of the random path XX under Q{\sf Q}. Moreover, for any 0≤s,t≤10\leq s,t\leq 1, we shall denote by Qst{\sf Q}_{st} the joint law of XsX_{s} and XtX_{t} under Q{\sf Q}, namely Qst:=(Xs,Xt)#Q{\sf Q}_{st}:=(X_{s},X_{t})_{\#}{\sf Q}.

As reference path measure R∈P(Ω){\sf R}\in\mathcal{P}(\Omega) we consider the law of the Markov diffusion process with generator

with initial law R0=m{\sf R}_{0}=\mathfrak{m}, where the potential VV appears at (1.1), (1.3), Δ\Delta is the Laplace-Beltrami operator on X{\rm X} and ∇\nabla the Levi-Civita connection associated to the metric gg (in Setting 1-(a) they are nothing but the standard Laplacian and gradient). It is well-known that R{\sf R} is a reversible Markov measure with reversing measure m\mathfrak{m}. In particular it is stationary, that is Rt=m{\sf R}_{t}=\mathfrak{m} for all 0≤t≤10\leq t\leq 1. For any ε>0\varepsilon>0 we denote by XεX^{\varepsilon} the time-rescaled process defined by Xtε:=XεtX^{\varepsilon}_{t}:=X_{\varepsilon t}, 0≤t≤10\leq t\leq 1, and by Rε:=(Xε)#R{\sf R}^{\varepsilon}:=(X^{\varepsilon})_{\#}{\sf R} the corresponding path measure. The parameter ε\varepsilon is meant to tend to zero so that Rε{\sf R}^{\varepsilon} is a slowed down version of R{\sf R}, whose generator is

For any ε>0\varepsilon>0, as a time rescaling of R{\sf R}, Rε{\sf R}^{\varepsilon} is also m\mathfrak{m}-reversible, so that in particular Rtε=m{\sf R}^{\varepsilon}_{t}=\mathfrak{m} for all tt.

The 1-parameter semigroup associated to L\mathcal{L} will be denoted by (Tt)(\mathcal{T}_{t}) and, in a completely analogous way, (Ttε)(\mathcal{T}^{\varepsilon}_{t}) the one associated to Lε\mathcal{L}^{\varepsilon}; notice that Ttε=Tεt\mathcal{T}^{\varepsilon}_{t}=\mathcal{T}_{\varepsilon t} for all t≥0t\geq 0. Within Setting 1 it is well-known (see for instance ) that there exists a unique heat kernel rt(x,y){\sf r}_{t}(x,y) associated to L\mathcal{L} which is a smooth function on (0,∞)×X×X(0,\infty)\times{\rm X}\times{\rm X}. Therefore, the semigroup (Tt)(\mathcal{T}_{t}) can be represented by

for all f∈L∞(m)f\in L^{\infty}(\mathfrak{m}). Let us also recall that, in conjunction with (1.2) or (1.4), (Tt)(\mathcal{T}_{t}) enjoys the Bakry-Émery contraction estimate

For its proof as well as for all the regularizing properties of T\mathcal{T} that will be used throughout the paper, we address the reader to .

defined for all f,g∈Cc∞(X)f,g\in C^{\infty}_{c}({\rm X}), are naturally associated to L\mathcal{L}. As it is not difficult to see, the drift VV in m\mathfrak{m} does not affect Γ\Gamma, since Γ(f,g)=⟨∇f,∇g⟩\Gamma(f,g)=\langle\nabla f,\nabla g\rangle. It is worth mentioning that, with respect to the standard definition provided in , here Γ\Gamma and Γ2\Gamma_{2} are not divided by 2, as the factor 1/2 already appears in L\mathcal{L}, which thus corresponds to an SDE driven by a standard Brownian motion.

The Schrödinger problem

Let μ0,μ1∈P(X)\mu_{0},\mu_{1}\in\mathcal{P}({\rm X}) be two probability measures: the Schrödinger problem associated with Rε,μ0,μ1{\sf R}^{\varepsilon},\mu_{0},\mu_{1} is defined by

and its value is called ‘entropic cost’. As a strictly convex minimization problem, it admits at most one solution.

The solution Pε{\sf P}^{\varepsilon} of (Sε), if it exists, is called the Rε{\sf R}^{\varepsilon}-entropic bridge between μ0\mu_{0} and μ1\mu_{1}. The Rε{\sf R}^{\varepsilon}-entropic interpolation (μtε)(\mu^{\varepsilon}_{t}) between μ0\mu_{0} and μ1\mu_{1} is defined as the time marginal flow of the solution Pε{\sf P}^{\varepsilon}, namely

The name ‘entropic interpolation’ stems from the connection with displacement interpolation. Indeed, it is known from that

This limit is a consequence of a more general result asserting that (Sε) converges to the quadratic Monge-Kantorovich problem as ε↓0\varepsilon\downarrow 0 in the sense of Γ\Gamma-convergence, see . As shown in , if (Sε) admits the solution Pε{\sf P}^{\varepsilon}, then there exist two non-negative measurable functions fε,gε:X→[0,∞)f^{\varepsilon},g^{\varepsilon}:{\rm X}\to[0,\infty) such that Pε=fε(X0)gε(X1) Rε{\sf P}^{\varepsilon}=f^{\varepsilon}(X_{0})g^{\varepsilon}(X_{1})\,{\sf R}^{\varepsilon} and defining

usually known as ‘Schrödinger system’: indeed, if we interprete them as a nonlinear system where the unknowns are fεf^{\varepsilon} and gεg^{\varepsilon}, then the (unique up to an obvious multiplicative rescaling) solution completely determines Pε{\sf P}^{\varepsilon} (see for instance ). As far as the convergence of entropic interpolations towards displacement ones is investigated, the following functions

are of special interest. We also set φε:=εlog⁡fε\varphi^{\varepsilon}:=\varepsilon\log f^{\varepsilon} in supp(μ0)\mathop{\rm supp}\nolimits(\mu_{0}) and ψε:=εlog⁡gε\psi^{\varepsilon}:=\varepsilon\log g^{\varepsilon} in supp(μ1)\mathop{\rm supp}\nolimits(\mu_{1}). They are called Schrödinger potentials, in connection with Kantorovich ones.

Existence and regularity results

Let us now derive a criterion, in terms of the endpoint marginals μ0\mu_{0} and μ1\mu_{1}, for the existence of regular functions fε,gεf^{\varepsilon},g^{\varepsilon} with well-defined Fisher information solving the Schrödinger system (2.6). As noticed in , the regularity (smoothness and integrability) of μ0\mu_{0} (resp. μ1\mu_{1}) is inherited by fεf^{\varepsilon} (resp. gεg^{\varepsilon}). In the next result we extend this property: if I(μ0 ∣ m)I(\mu_{0}\,|\,\mathfrak{m}) is finite, then so is I(fεm ∣ m)I(f^{\varepsilon}\mathfrak{m}\,|\,\mathfrak{m}) and analogously with μ1,gε\mu_{1},g^{\varepsilon}.

Let (X,d,m)({\rm X},{\sf d},\mathfrak{m}) be as in Setting 1, ε>0\varepsilon>0 and consider two probability measures μ0=ρ0m,μ1=ρ1m∈P(X)\mu_{0}=\rho_{0}\mathfrak{m},\mu_{1}=\rho_{1}\mathfrak{m}\in\mathcal{P}({\rm X}) with compact supports. Then the following hold:

The Schrödinger problem (Sε) admits a solution with finite entropy if and only if H(μ0 ∣ m)H(\mu_{0}\,|\,\mathfrak{m}), H(μ1 ∣ m)<∞H(\mu_{1}\,|\,\mathfrak{m})<\infty. This solution is unique and the Rε{\sf R}^{\varepsilon}-entropic interpolation between μ0\mu_{0} and μ1\mu_{1} exists.

Suppose in addition that ρ0,ρ1∈L∞(m)\rho_{0},\rho_{1}\in L^{\infty}(\mathfrak{m}).

Then fε,gε∈L∞(m)f^{\varepsilon},g^{\varepsilon}\in L^{\infty}(\mathfrak{m}).

For any 0<t<10<t<1 the functions ftε,gtε,ρtεf^{\varepsilon}_{t},g^{\varepsilon}_{t},\rho^{\varepsilon}_{t} as well as f1εf^{\varepsilon}_{1} and g0εg^{\varepsilon}_{0} belong to C∞(X)∩L∞(m)C^{\infty}({\rm X})\cap L^{\infty}(\mathfrak{m}).

In addition to item (b), suppose that I(μ0 ∣ m)I(\mu_{0}\,|\,\mathfrak{m}) is finite (resp. I(μ1 ∣ m)I(\mu_{1}\,|\,\mathfrak{m})). Then so is I(fεm ∣ m)I(f^{\varepsilon}\mathfrak{m}\,|\,\mathfrak{m}) (resp. I(gεm ∣ m)I(g^{\varepsilon}\mathfrak{m}\,|\,\mathfrak{m})).

For (a) and (b-i) see [13, Proposition 2.1]. As regards (b-ii), the fact that fε∈L∞(m)f^{\varepsilon}\in L^{\infty}(\mathfrak{m}), rt∈C∞(X×X){\sf r}_{t}\in C^{\infty}({\rm X}\times{\rm X}) and (2.1) imply ftε∈C∞(X)f_{t}^{\varepsilon}\in C^{\infty}({\rm X}) for all 0<t≤10<t\leq 1, while the maximum principle ensures that ftε∈L∞(m)f_{t}^{\varepsilon}\in L^{\infty}(\mathfrak{m}); the statements for gtεg^{\varepsilon}_{t} and ρtε=ftεgtε\rho_{t}^{\varepsilon}=f_{t}^{\varepsilon}g_{t}^{\varepsilon} follow by the same reason.

(b-iii). Notice that the first equation in the Schrödinger system (2.6) can be rewritten as

Since Tε\mathcal{T}^{\varepsilon} is positivity improving, T1εgε\mathcal{T}^{\varepsilon}_{1}g^{\varepsilon} is smooth and ρ0∈Ck(X)\rho_{0}\in C^{k}({\rm X}) with compact support, the conclusion follows. A similar argument holds for gεg^{\varepsilon}.

(c). Observe that the Schrödinger system (2.6) and g0ε>0g_{0}^{\varepsilon}>0 (as Tε\mathcal{T}^{\varepsilon} is positivity improving) force fεf^{\varepsilon} to have the same support as ρ0\rho_{0}. Thus fεf^{\varepsilon} has compact support and, as a consequence, g0ε≥c>0g_{0}^{\varepsilon}\geq c>0 in supp(fε)\mathop{\rm supp}\nolimits(f^{\varepsilon}) for some cc. This remark and the chain rule allow us to say that

so that it remains to prove the integrability of the right-hand side. The first term is integrable by assumption, the third one by the regularization properties of Tε\mathcal{T}^{\varepsilon}, while for the second one notice that I(μ0 ∣ m)<∞I(\mu_{0}\,|\,\mathfrak{m})<\infty and ρ0∈L∞(m)\rho_{0}\in L^{\infty}(\mathfrak{m}) imply ∣∇ρ0∣∈L2(m)|\nabla\rho_{0}|\in L^{2}(\mathfrak{m}); plugging this information into

For this reason we formulate the following

The endpoint marginals μ0,μ1∈P2(X)\mu_{0},\mu_{1}\in\mathcal{P}_{2}({\rm X}) are such that μ0,μ1\mu_{0},\mu_{1} have compact supports, H(μ0 ∣ m),H(μ1 ∣ m),I(μ1 ∣ m)<∞H(\mu_{0}\,|\,\mathfrak{m}),H(\mu_{1}\,|\,\mathfrak{m}),I(\mu_{1}\,|\,\mathfrak{m})<\infty and their densities ρ0,ρ1\rho_{0},\rho_{1} belong to C∞(X)C^{\infty}({\rm X}).

A dynamical viewpoint

As concerns the evolution of entropic interpolations and Schrödinger potentials, let us first notice that under Assumptions 2.8, by the very definition (2.5), item (b-iii) of Proposition 2.7 and the fact that rt(x,y)∈C∞((0,∞)×X2){\sf r}_{t}(x,y)\in C^{\infty}((0,\infty)\times{\rm X}^{2}) we deduce that ftεf_{t}^{\varepsilon} and gtεg_{t}^{\varepsilon} are smooth on ×X\times{\rm X} and solve

in the classical sense. Moreover, if we look at (ftε),(gtε)(f_{t}^{\varepsilon}),(g_{t}^{\varepsilon}) as curves parametrized by tt with values in W1,2(X,m)W^{1,2}({\rm X},\mathfrak{m}), they belong to the set AC(,W1,2(X,m))AC(,W^{1,2}({\rm X},\mathfrak{m})) of all absolutely continuous functions from $totoW^{1,2}({\rm X},\mathfrak{m}).AsaconsequencethePDEsaboveholdalsowhen. As a consequence the PDEs above hold also when\partial_{t}f_{t}^{\varepsilon},\partial_{t}g_{t}^{\varepsilon}areseenasstrongare seen as strongW^{1,2}$-limits.

Relying on that, it follows that the Schrödinger potentials φtε,ψtε\varphi_{t}^{\varepsilon},\psi_{t}^{\varepsilon} are smooth on (0,1]×X(0,1]\times{\rm X} and [0,1)×X[0,1)\times{\rm X} and solve forward and backward Hamilton-Jacobi-Bellman equations respectively, i.e.

while for (ρtε)(\rho_{t}^{\varepsilon}) the continuity equation

is satisfied in (0,1)×X(0,1)\times{\rm X}, where ϑtε:=(ψtε−φtε)/2\vartheta_{t}^{\varepsilon}:=(\psi_{t}^{\varepsilon}-\varphi_{t}^{\varepsilon})/2 and divm{\rm div}_{\mathfrak{m}} denotes the divergence with respect to m\mathfrak{m}, i.e. the opposite of the adjoint of the differential in W1,2(X,m)W^{1,2}({\rm X},\mathfrak{m}). This last PDE is strongly linked to the dynamical representation of the entropic cost inf⁡\eqrefeq−h06\inf\eqref{eq-h06}, namely

shown in for Setting 1-(a) and in for rather general metric measure spaces including Setting 1-(b). By the very definition of ϑtε\vartheta_{t}^{\varepsilon} and since εlog⁡ρtε=φtε+ψtε\varepsilon\log\rho_{t}^{\varepsilon}=\varphi_{t}^{\varepsilon}+\psi_{t}^{\varepsilon}, this implies

The regularity of Schrödinger potentials comes from the one of ftε,gtεf_{t}^{\varepsilon},g_{t}^{\varepsilon}, the fact that ftε,gtεf_{t}^{\varepsilon},g_{t}^{\varepsilon} are everywhere positive and the logarithm is smooth on (0,∞)(0,\infty). However, Schrödinger potentials are not integrable in general, as ftε,gtεf_{t}^{\varepsilon},g_{t}^{\varepsilon} can be arbitrarily close to 0. Thus we cannot study their behaviour as curves with values into some Lp(m)L^{p}(\mathfrak{m}) space. This also explains why the regularity of φtε\varphi_{t}^{\varepsilon} (resp. ψtε\psi_{t}^{\varepsilon}) does not extend up to t=0t=0 (resp. t=1t=1).

A physical interpretation

In the Euclidean framework of Setting 1-(a) it is possible to make a bridge between what is presented so far and Nelson’s formalism , thus providing a physical motivation for some results stated above and a further perspective on some objects.

For Rε{\sf R}^{\varepsilon}, it is easily seen that

This allows to rewrite the continuity equation (2.11) as

where now div{\rm div} denotes the divergence with respect to Ln\mathcal{L}^{n} if in Setting 1-(a) or Vol{\rm Vol} if in Setting 1-(b). This is perfectly coherent with (2.11) since divm(w)=div(w)−ε2∇V⋅w{\rm div}_{\mathfrak{m}}(w)={\rm div}(w)-\frac{\varepsilon}{2}\nabla V\cdot w for any vector field ww. Furthermore, (2.12) becomes

Proof of Proposition 1.5

We need to state some preliminary lemmas before completing the proof of Proposition 1.5 at page 3. Throughout the whole section we shall assume to work within (X,d,m)({\rm X},{\sf d},\mathfrak{m}) as in Setting 1.

Let us start with an approximation result.

Let μ∈P2(X)\mu\in\mathcal{P}_{2}({\rm X}) with H(μ ∣ m)<∞H(\mu\,|\,\mathfrak{m})<\infty. Then:

there exists a sequence (μn)⊂P2(X)(\mu_{n})\subset\mathcal{P}_{2}({\rm X}) with μn=ρnm\mu_{n}=\rho_{n}\mathfrak{m} and ρn∈Cc∞(X)\rho_{n}\in C^{\infty}_{c}({\rm X}) such that W2(μn,ν)→W2(μ,ν)W_{2}(\mu_{n},\nu)\to W_{2}(\mu,\nu) for all ν∈P2(X)\nu\in\mathcal{P}_{2}({\rm X}) and H(μn ∣ m)→H(μ ∣ m)H(\mu_{n}\,|\,\mathfrak{m})\to H(\mu\,|\,\mathfrak{m}) as n→∞n\to\infty;

if in addition I(μ ∣ m)<∞I(\mu\,|\,\mathfrak{m})<\infty, then there exists a sequence (μn′)⊂P2(X)(\mu^{\prime}_{n})\subset\mathcal{P}_{2}({\rm X}) with μn=ρnm\mu_{n}=\rho_{n}\mathfrak{m}, ρn∈Cc∞(X)\rho_{n}\in C^{\infty}_{c}({\rm X}) such that W2(μn′,ν)→W2(μ,ν)W_{2}(\mu^{\prime}_{n},\nu)\to W_{2}(\mu,\nu) for all ν∈P2(X)\nu\in\mathcal{P}_{2}({\rm X}), H(μn′ ∣ m)→H(μ ∣ m)H(\mu^{\prime}_{n}\,|\,\mathfrak{m})\to H(\mu\,|\,\mathfrak{m}) and I(μn′ ∣ m)→I(μ ∣ m)I(\mu^{\prime}_{n}\,|\,\mathfrak{m})\to I(\mu\,|\,\mathfrak{m}) as n→∞n\to\infty.

Let us write μ=ρ m\mu=\rho\,\mathfrak{m} and, as a first step, let us prove that both in (a) and (b) it is possible to find a sequence of measures with smooth densities converging to μ\mu in the desired sense. This can be proved by defining for ε>0\varepsilon>0

which clearly have smooth densities by the regularizing properties of (Tε)(\mathcal{T}_{\varepsilon}). The convergence of W2(με,ν)W_{2}(\mu_{\varepsilon},\nu), H(με ∣ m)H(\mu_{\varepsilon}\,|\,\mathfrak{m}) and I(με ∣ m)I(\mu_{\varepsilon}\,|\,\mathfrak{m}) to W2(μ,ν)W_{2}(\mu,\nu), H(μ ∣ m)H(\mu\,|\,\mathfrak{m}) and I(μ ∣ m)I(\mu\,|\,\mathfrak{m}) respectively as ε↓0\varepsilon\downarrow 0 is now a well-known fact in the theory of gradient flows (see for instance Theorem 2.4.15 and Remark 2.4.16 in in conjunction with the fact that the squared slope of the entropy is the Fisher information, as proved in ).

Thus, it is not restrictive to suppose that μ\mu has smooth density. Under this new assumption, let us prove that we can find a sequence of measures with compact supports and smooth densities converging to μ\mu in the desired sense. To this aim define

where αn\alpha_{n} is the renormalization constant and χn\chi_{n} is a smooth cut-off function with support in Bn+1(x)B_{n+1}(x), for some x∈Xx\in{\rm X}, χn(x)=1\chi_{n}(x)=1 and Lipschitz constant controlled by C/nC/n, where C>1C>1 does not depend on nn (see e.g. for a proof of the existence of such cut-off functions). By dominated convergence it is not difficult to see that W2(μ,μn)→0W_{2}(\mu,\mu_{n})\to 0 and thus W2(μn,ν)→W2(μ,ν)W_{2}(\mu_{n},\nu)\to W_{2}(\mu,\nu) for all ν∈P2(X)\nu\in\mathcal{P}_{2}({\rm X}) as n→∞n\to\infty; for the same reason H(μn ∣ m)→H(μ ∣ m)H(\mu_{n}\,|\,\mathfrak{m})\to H(\mu\,|\,\mathfrak{m}). If we also assume that I(μ ∣ m)<∞I(\mu\,|\,\mathfrak{m})<\infty, then

Since the right-hand side is integrable and χn→1\chi_{n}\to 1 as n→∞n\to\infty, by dominated convergence we get I(μn ∣ m)→I(μ ∣ m)I(\mu_{n}\,|\,\mathfrak{m})\to I(\mu\,|\,\mathfrak{m}).

Combining the two steps and using a diagonal argument, the conclusion follows. ∎

The following conservation result was pointed out in and in the case X{\rm X} is compact with different approaches; see also . Following , we extend the statement to the present framework.

Under Assumptions 2.8, for any ε>0\varepsilon>0 the function

is real-valued and constant. Thus we shall denote it by QεQ^{\varepsilon}.

As a first step, for all 0<t<10<t<1 by algebraic manipulation we have

From integration by parts formula it is straightforward to see that the right-hand side vanishes, whence the conclusion. ∎

Motivated by (2.12), let us investigate separately the convergence of current and osmotic velocities as ε↓0\varepsilon\downarrow 0.

Under Assumptions 2.8, for any ε>0\varepsilon>0 we have

Let us first notice that combining (2.12) and (2.4) we get

Since the continuity equation (2.11) is satisfied by μtε\mu_{t}^{\varepsilon}, the Benamou-Brenier formula holds for any ε>0\varepsilon>0, that is

and together with the above limit, this leads us to the first identities both in (3.4) and (3.5). From them we immediately deduce that

whence also the second identity in (3.5). Finally observe that

As already noticed at Remark 2.14, although it is smooth on the open interval (0,1)(0,1), the density ρtε\rho^{\varepsilon}_{t} might be arbitrarily close to zero. Consequently, the Schrödinger potentials might not be integrable enough, and the Fisher information I(μtε ∣ m)I(\mu^{\varepsilon}_{t}\,|\,\mathfrak{m}) might behave badly around t=0t=0 and t=1t=1. Next lemma provides related regularity and growth controls. Its proof is strongly inspired by , , and ; we thus address the reader to these articles for more details.

Under Assumptions 2.8, let δ>0\delta>0 and set, for all 0≤t≤10\leq t\leq 1, ρtε,δ:=(ftε+δ)(gtε+δ)\rho_{t}^{\varepsilon,\delta}:=(f_{t}^{\varepsilon}+\delta)(g_{t}^{\varepsilon}+\delta) as well as

Then all the functions so defined belong to C∞(×X)∩W1,2(X,m)C^{\infty}(\times{\rm X})\cap W^{1,2}({\rm X},\mathfrak{m}) and, as curves parametrized by tt with values in W1,2(X,m)W^{1,2}({\rm X},\mathfrak{m}), to AC(,W1,2(X,m))AC(,W^{1,2}({\rm X},\mathfrak{m})). The time derivatives of φtε,δ,ψtε,δ,ρtε,δ\varphi_{t}^{\varepsilon,\delta},\psi_{t}^{\varepsilon,\delta},\rho_{t}^{\varepsilon,\delta} are given by

where ∂tφtε,δ\partial_{t}\varphi_{t}^{\varepsilon,\delta}, ∂tψtε,δ\partial_{t}\psi_{t}^{\varepsilon,\delta}, ∂tρtε,δ\partial_{t}\rho_{t}^{\varepsilon,\delta} have to be understood both in the classical sense and as strong W1,2W^{1,2}-limits.

As already explained in Section 2, under Assumptions 2.8 ftε,gtε∈C∞(,X)∩W1,2(X,m)f_{t}^{\varepsilon},g_{t}^{\varepsilon}\in C^{\infty}(,{\rm X})\cap W^{1,2}({\rm X},\mathfrak{m}) and (ftε),(gtε)∈AC(,W1,2(X,m))(f_{t}^{\varepsilon}),(g_{t}^{\varepsilon})\in AC(,W^{1,2}({\rm X},\mathfrak{m})). Therefore the regularity and integrability properties of φtε,δ,ψtε,δ,ϑtε,δ,ρtε,δ\varphi_{t}^{\varepsilon,\delta},\psi_{t}^{\varepsilon,\delta},\vartheta_{t}^{\varepsilon,\delta},\rho_{t}^{\varepsilon,\delta} are a straightforward consequence of the chain rule and of the fact that the logarithm is smooth with bounded derivatives on [δ,∞)[\delta,\infty). Also the PDEs solved by φtε,δ,ψtε,δ,ρtε,δ\varphi_{t}^{\varepsilon,\delta},\psi_{t}^{\varepsilon,\delta},\rho_{t}^{\varepsilon,\delta} are easily deduced, when interpreted in the classical sense, as they follow from (2.10) and (2.11). In order to deduce the same identities with ∂tφtε,δ\partial_{t}\varphi_{t}^{\varepsilon,\delta}, ∂tψtε,δ\partial_{t}\psi_{t}^{\varepsilon,\delta}, ∂tρtε,δ\partial_{t}\rho_{t}^{\varepsilon,\delta} seen as strong W1,2W^{1,2}-limits, notice that by the maximum principle

whence (φtε,δ)∈L∞((0,∞),L∞(m))(\varphi_{t}^{\varepsilon,\delta})\in L^{\infty}((0,\infty),L^{\infty}(\mathfrak{m})). Moreover, the smoothness of the logarithm, the chain and Leibniz rules entail that

whence (∣∇φtε,δ∣2),(Δφtε,δ)∈L∞((0,1),W1,2(X,m))(|\nabla\varphi_{t}^{\varepsilon,\delta}|^{2}),(\Delta\varphi_{t}^{\varepsilon,\delta})\in L^{\infty}((0,1),W^{1,2}({\rm X},\mathfrak{m})); analogous estimates hold for ∇∣∇φtε,δ∣2\nabla|\nabla\varphi_{t}^{\varepsilon,\delta}|^{2} and ∇Δφtε,δ\nabla\Delta\varphi_{t}^{\varepsilon,\delta}. These bounds together with the fact that the PDEs in (3.8) hold in the classical sense imply, by a dominated convergence argument, that (3.8) are satisfied also as strong W1,2W^{1,2}-limits.

Relying on that, (3.9a) and (3.9b) follow by the computations carried out in . Indeed, the validity of (3.8) as strong W1,2W^{1,2}-limits on $ensuresthatensures thatt\mapsto u(\rho_{t}^{\varepsilon,\delta})andandt\mapsto\langle\nabla\rho_{t}^{\varepsilon,\delta},\nabla\vartheta_{t}^{\varepsilon,\delta}\ranglebelongtobelong toAC(,L^{2}(\mathfrak{m}))$ and thus we can pass the time derivatives under the integral sign, i.e.

Then the Hamilton-Jacobi-Bellman equations for the Schrödinger potentials and the continuity equation for the entropic interpolation together with εlog⁡ρtε=φtε+ψtε\varepsilon\log\rho_{t}^{\varepsilon}=\varphi_{t}^{\varepsilon}+\psi_{t}^{\varepsilon}, here replaced by εlog⁡ρtε,δ=φtε,δ+ψtε,δ\varepsilon\log\rho_{t}^{\varepsilon,\delta}=\varphi_{t}^{\varepsilon,\delta}+\psi_{t}^{\varepsilon,\delta}, are the only tools needed to deduce (3.9a) and (3.9b). Finally, the fact that:

(ρtε,δ),(ϑtε,δ)∈AC(,W1,2(X,m))(\rho_{t}^{\varepsilon,\delta}),(\vartheta_{t}^{\varepsilon,\delta})\in AC(,W^{1,2}({\rm X},\mathfrak{m}));

(ρtε,δ)∈AC(,W1,2(X,m))(\rho_{t}^{\varepsilon,\delta})\in AC(,W^{1,2}({\rm X},\mathfrak{m})), (∣∇ϑtε,δ∣2),(Δϑtε,δ),(∣∇log⁡ρtε,δ∣2),(Δlog⁡ρtε,δ)(|\nabla\vartheta_{t}^{\varepsilon,\delta}|^{2}),(\Delta\vartheta_{t}^{\varepsilon,\delta}),(|\nabla\log\rho_{t}^{\varepsilon,\delta}|^{2}),(\Delta\log\rho_{t}^{\varepsilon,\delta}) belong to L∞((0,1),W1,2(X,m))L^{\infty}((0,1),W^{1,2}({\rm X},\mathfrak{m})) and ϑtε,δ,log⁡ρtε,δ∈C∞(×X)\vartheta_{t}^{\varepsilon,\delta},\log\rho_{t}^{\varepsilon,\delta}\in C^{\infty}(\times{\rm X});

imply the continuity on $$ of the right-hand sides of (3.9a) and (3.9b) respectively. ∎

With this results at disposal we can prove our main lemma: a rigorous ‘entropic’ analogue of Otto-Villani’s heuristic argument.

Under Assumptions 2.8, for any ε>0\varepsilon>0 it holds

The proof of the lemma is based on the standard calculus identity

consequence of the lower Ricci bounds (1.2) and (1.4) rewritten in the form of the Bochner-Lichnerowicz-Weitzenböck formula, we obtain

where μtε,δ:=ρtε,δm\mu_{t}^{\varepsilon,\delta}:=\rho_{t}^{\varepsilon,\delta}\mathfrak{m}. It is now sufficient to pass to the limit as δ↓0\delta\downarrow 0. By dominated convergence (recall that fε,gε∈L∞(m)f^{\varepsilon},g^{\varepsilon}\in L^{\infty}(\mathfrak{m})) it is easy to see that the left-hand side converges to H(μ1 ∣ m)−H(μ0 ∣ m)H(\mu_{1}\,|\,\mathfrak{m})-H(\mu_{0}\,|\,\mathfrak{m}). In order to apply the dominated convergence theorem also to the first term on the right-hand side, notice that

Thus, to control ⟨∇ϑ1ε,δ,∇ρ1ε,δ⟩\langle\nabla\vartheta_{1}^{\varepsilon,\delta},\nabla\rho_{1}^{\varepsilon,\delta}\rangle let us first observe that

and remark that the right-hand side is integrable by Proposition 2.7-(c); secondly

and in this case the right-hand side is integrable as gεg^{\varepsilon} has compact support and f1εf_{1}^{\varepsilon} is bounded away from 0 therein; finally,

as gε+δ≥δg^{\varepsilon}+\delta\geq\delta and the same strategy applies to all the remaining terms that we obtain developing ⟨∇ϑ1ε,δ,∇ρ1ε,δ⟩\langle\nabla\vartheta_{1}^{\varepsilon,\delta},\nabla\rho_{1}^{\varepsilon,\delta}\rangle. Therefore, the first term on the right-hand side of (3.14) converges to the first term on the right-hand side of (3.12). As regards the other two summands, by the very definition of ϑtε,δ\vartheta_{t}^{\varepsilon,\delta} and since

the conclusion will follow if we are able to prove that

To this aim, notice that ρtε,δ=ρtε+δftε+δgtε+δ2\rho_{t}^{\varepsilon,\delta}=\rho_{t}^{\varepsilon}+\delta f_{t}^{\varepsilon}+\delta g_{t}^{\varepsilon}+\delta^{2}, whence using either ftε+δ≥ftεf_{t}^{\varepsilon}+\delta\geq f_{t}^{\varepsilon} or ftε+δ≥δf_{t}^{\varepsilon}+\delta\geq\delta it is easy to infer that

All the right-hand sides above are integrable on ×X\times{\rm X} (either by (2.13) or by the Bakry-Émery contraction estimate (2.2) together with fε∈Cc∞(X)f^{\varepsilon}\in C^{\infty}_{c}({\rm X})), thus by dominated convergence (3.15a) follows. An analogous argument holds for (3.15b), whence the conclusion. ∎

Completion of the proof of Proposition 1.5

We are now ready to complete the proof of the HWI* inequality.

First of all, by Proposition 2.7 and Lemma 3.1, it is sufficient to prove the result for any μ0,μ1\mu_{0},\mu_{1} satisfying Assumptions 2.8. This allows us to invoke Lemma 3.11. Secondly, observe that by the Cauchy-Schwarz inequality

so that if we plug this information into (3.12) we obtain

Let us now pass to the limit as ε↓0\varepsilon\downarrow 0: by Lemma 3.2, at t=1t=1 we have

so that (3.6) together with I(μ1 ∣ m)<∞I(\mu_{1}\,|\,\mathfrak{m})<\infty yields

By Lemma 3.3 we can pass to the limit as ε↓0\varepsilon\downarrow 0 also in the remaining terms on the right-hand side, thus concluding. ∎

Let us mention that another approach to prove the HWI inequality via the Schrödinger problem is pointed out in a recent work where an entropic counterpart of the HWI inequality is formally obtained by differentiating the convexity estimate of the entropy along the entropic interpolations introduced in [6, Thm. 1.4].

Final remarks and comments

From a heuristic point of view, the expression of the constant quantity QεQ^{\varepsilon} can be deduced by standard arguments in Lagrangian and Hamiltonian formalism. Indeed, motivated by (2.12) let us consider the action functional

By means of Legendre’s transform, the corresponding Hamiltonian is given by

and, at least formally, H\mathscr{H} is constant along the critical points of A\mathscr{A}. Since μ0\mu_{0} and μ1\mu_{1} are prescribed, the Euler-Lagrange equation for (4.1) reads as

and, as it is not difficult to see (e.g. following the computations carried out in which fit to Setting 1), these PDEs are satisfied along (ρtε,ϑtε)(\rho_{t}^{\varepsilon},\vartheta_{t}^{\varepsilon}). Finally, as in the Hamiltonian pp represents a momentum density, it is natural to set pt:=νtvtp_{t}:=\nu_{t}v_{t}. From these considerations, the guess on the existence of the conserved quantity of Lemma 3.2 and on its expression follows. This point of view is particularly investigated in the recent paper , to which we refer for more details.

The compact case

As already mentioned in Remark 2.14, in general Schrödinger potentials fail to be smooth in t=0t=0 or t=1t=1 and are not even integrable for any 0<t<10<t<1, because ftε,gtεf_{t}^{\varepsilon},g_{t}^{\varepsilon} can be arbitrarily close to 0. This problem can be overcome if fε,gε≥c>0f^{\varepsilon},g^{\varepsilon}\geq c>0 for some constant cc: then ftε,gtε≥cf_{t}^{\varepsilon},g_{t}^{\varepsilon}\geq c as well by the maximum principle and φtε,ψtε≥εlog⁡c\varphi_{t}^{\varepsilon},\psi_{t}^{\varepsilon}\geq\varepsilon\log c from their very definition. However, even if we assume μ0,μ1\mu_{0},\mu_{1} to have densities bounded away from 0 (thus forcing m\mathfrak{m} to belong to P2(X)\mathcal{P}_{2}({\rm X}), a condition which does not follow from our Setting 1, unless κ>0\kappa>0, cf. Remark 1.7-(a)), it is not known yet whether this implies an analogous lower bound on fεf^{\varepsilon} and gεg^{\varepsilon}. This is the motivation behind the definitions provided in Lemma 3.7.

On the contrary, if X{\rm X} is assumed to be compact, then it is not restrictive to assume μ0,μ1≥cm\mu_{0},\mu_{1}\geq c\mathfrak{m} for some c>0c>0: indeed any μ∈P(X)\mu\in\mathcal{P}({\rm X}) can be approximated by

(where αn\alpha_{n} is the renormalization constant) in W2W_{2}, entropy and Fisher information in the sense of Lemma 3.1. Secondly, rε{\sf r}_{\varepsilon} is bounded from above since rε∈C∞(X){\sf r}_{\varepsilon}\in C^{\infty}({\rm X}). These two facts imply that fε,gε≥c′>0f^{\varepsilon},g^{\varepsilon}\geq c^{\prime}>0, as shown in . As a consequence the proof of the HWI inequality is easier and the parallelism with the heuristic proof of Otto-Villani is stronger, since no ‘δ\delta-argument’ in Lemma 3.7 and Lemma 3.11 is needed; in particular, Lemma 3.7 is already known to hold by and thus no proof is required.

𝖱𝖢𝖣𝖱𝖢𝖣{\sf RCD} spaces

All the results presented in this paper, with particular mention of the key Lemma 3.11 and Proposition 1.5, are also true in a more general framework than the one of Setting 1-(b), namely in RCD∗(K,N){\sf RCD}^{*}(K,N) spaces (introduced in ). If X{\rm X} is assumed to be compact, then this has been shown by the last author in his PhD thesis . When X{\rm X} is not compact, the most important steps in our entropic approach still hold. Namely:

all the regularity and integrability results concerning Schrödinger potentials and entropic interpolations mentioned in Section 2 as well as the dynamic representation of the entropic cost;

the regularizing and contraction properties of (Tt)(\mathcal{T}_{t});

the existence of ‘good’ cut-off functions;

the Benamou-Brenier formula and the Bochner-Lichnerowicz-Weitzenböck inequality.

The reader is addressed to , for the first point and to for all the others.

Acknowledgements. This research was supported by the French ANR-17-CE40-0030 EFI project and the LABEX MILYON ANR-10-LABX-0070.

References