A formula for the time derivative of the entropic cost and applications

Giovanni Conforti, Luca Tamanini

Introduction and statement of the main results

The entropic transportation cost is the optimal value in a probabilistic version of the Monge-Kantorovich optimal transport problem, the Schrödinger problem, whose study has already shown to have far reaching consequences in various fields, ranging from statistical machine learning to functional inequalities. The goal of the present article is to advance in the study of the entropic cost as a function of the time (regularization) parameter TT (see Definition 1.2 below). Following Mikami’s contribution linking the Schrödinger problem to optimal transport, several results have been obtained in the last years concerning the behavior of the entropic cost in the short-time (small noise) limit T→0T\rightarrow 0. In particular, as a byproduct of the research line originated in another step forward was made with the computation of the first derivative of the rescaled entropic cost at T=0T=0, see also , the recent works and references therein. However, very few results beyond the short-time limit have been obtained. In particular, very little is known about the long-time regime T→∞T\rightarrow\infty, where the entropic cost is expected to converge to the sum of the marginal entropies. This lack of knowledge was one of the main motivations for our work and in this respect, our contribution includes:

A formula for the first and second derivative of the (rescaled) entropic cost for a general value of TT in terms of the so-called “energy” (defined at (1.6) below).

A rigorous identification of the large-time limit of the entropic cost as the sum of the marginal entropies.

Sharp exponential convergence rates under a curvature condition in the long-time regime. We obtain this result not only for the classical Schrödinger problem but also for the recently introduced Mean Field Schrödinger problem .

We also establish some results in the short-time limit. Their interest resides in the fact that they allow to obtain a clean expression of the second derivative of the rescaled entropic cost that, to the best of our knowledge, was not known before. However, the regularity assumptions we impose on the marginals are not the weakest ones. For this part, our contribution can be resumed as follows:

A formula for the second derivative of the rescaled cost around T=0T=0 that yields the local convexity of the cost in the time variable.

A non-asymptotic sharp quantitative bound for the convergence to the Wasserstein distance depending only on the integral of the Fisher information functional along the corresponding Wasserstein geodesic.

The measure m\mathfrak{m} is invariant for the SDE

(here BtB_{t} denotes a standard Brownian motion) whose joint law at time and TT will be denoted by R0,TR_{0,T}. Within this framework and given μ,ν∈P(M)\mu,\nu\in\mathcal{P}(M) (as usual, for a measurable space (E,E)(E,\mathcal{E}) we denote P(E)\mathcal{P}(E) the set of probability measures over (E,E)(E,\mathcal{E})), the entropic transportation cost CT(μ,ν)\mathscr{C}_{T}(\mu,\nu) is defined as the optimal value in the corresponding Schrödinger problem, namely

where Π(μ,ν)⊂P(M×M)\Pi(\mu,\nu)\subset\mathcal{P}(M\times M) is the set of couplings of μ\mu and ν\nu and H\mathcal{H} is the relative entropy functional defined for two probability measures p,qp,q on the same (arbitrary) measurable space as

When qq is not a probability, the precise definition is postponed to Section 2.

The physical meaning of the variational problem (1.2) is described in the seminal papers , where E. Schrödinger addressed the problem of finding the most likely evolution of a system of independent random particles driven by (1.1) conditionally on the observation of their initial and final configuration. With this picture in mind, the short- and long-time behavior of the entropic cost sounds perfectly natural.

We state here the additional assumptions we will make either on the reference measure m\mathfrak{m} or on the marginals μ,ν\mu,\nu. Concerning the former, we will often assume that

Notice that when κ>0\kappa>0, (H2) automatically holds and in addition m∈P2(M)\mathfrak{m}\in\mathcal{P}_{2}(M) (see for instance [32, Theorem 4.26]). For the latter, we shall suppose that either

where P2(M)⊂P(M)\mathcal{P}_{2}(M)\subset\mathcal{P}(M) denotes the space of probability measures over MM with finite second moment, or, more frequently, that

The Benamou-Brenier formulation.

The fluid-dynamical formulation of the entropic cost asserts that, at least under (H4), we have

where (ρˉT,vˉT)(\bar{\rho}^{T},\bar{v}^{T}) is the unique optimal curve and the infimum runs over all weak solutions (ρˉ,vˉ)(\bar{\rho},\bar{v}) of the continuity equation

satisfying the marginal constraints ρ0m=μ\rho_{0}\mathfrak{m}=\mu and ρTm=ν\rho_{T}\mathfrak{m}=\nu, namely among all (ρˉt)⊂L∞(m)(\bar{\rho}_{t})\subset L^{\infty}(\mathfrak{m}) with ρˉtm∈P(M)\bar{\rho}_{t}\mathfrak{m}\in\mathcal{P}(M) and all Borel vector fields (vˉt)(\bar{v}_{t}) such that:

t↦ρˉtmt\mapsto\bar{\rho}_{t}\mathfrak{m} is weakly continuous and there exists C>0C>0 such that ρˉt≤C\bar{\rho}_{t}\leq C for all t∈[0,T]t\in[0,T];

From a physical point of view, (1.4) states that the trajectory (ρˉtT)t∈[0,T](\bar{\rho}_{t}^{T})_{t\in[0,T]}, also called entropic interpolation, is the one minimizing a functional consisting of two terms: the former is purely kinetic, while the latter is given by the Fisher information functional. Hence (ρˉtT)t∈[0,T](\bar{\rho}_{t}^{T})_{t\in[0,T]} is given by the balance between deterministic and chaotic behavior.

The energy.

where the role of the function ff in (1.7) is played by the relative entropy H(⋅∣m)\mathcal{H}(\cdot|\mathfrak{m}) in (1.4). For (1.7) it is well known that along any critical curve the total energy (given by the sum of the kinetic and potential components)

is conserved. The analogous of this simple fact for the problem (1.4) is the conservation of (1.6) along optimal flows.

Short-time behavior of the entropic cost.

The fact that the rescaled entropic cost converges to the squared Wasserstein distance of order two in the short-time limit, i.e.

has generated a surge of interest around the Schrödinger problem (henceforth SP), since SP is more regular and numerically easier to solve than the Monge-Kantorovich problem. There exist nowadays several proofs of (1.8), see for instance for Γ\Gamma-convergence results. A proof that is valid under the hypotheses (H1) and (H4) can be found in [18, Remark 5.11]. In , (see also ) a further fundamental step was taken that consists in computing the first order term in the expansion of TCT(μ,ν)T\mathscr{C}_{T}(\mu,\nu) around T=0T=0. Referring to the above mentioned articles for precise statements, we have the following expansionIn the above mentioned references, one often finds H(μ ∣ m)−H(ν ∣ m)\mathcal{H}(\mu\,|\,\mathfrak{m})-\mathcal{H}(\nu\,|\,\mathfrak{m}) instead of H(μ ∣ m)+H(ν ∣ m)\mathcal{H}(\mu\,|\,\mathfrak{m})+\mathcal{H}(\nu\,|\,\mathfrak{m}) in the first order term. This is due to a slightly different choice of reference measure R0,TR_{0,T}. Since we prefer the reference measure to be reversible, we gain an extra H(μ ∣ m)\mathcal{H}(\mu\,|\,\mathfrak{m}) term in the Taylor expansion.

In this work we establish at Theorem 1.6 a non-asymptotic bound for TCT(μ,ν)−W22(μ,ν)/2T\mathscr{C}_{T}(\mu,\nu)-W_{2}^{2}(\mu,\nu)/2 which is sharp in the limit T→0T\to 0 and we compute the second order term in (1.9), thus getting

where (ρt0m)t∈(\rho^{0}_{t}\mathfrak{m})_{t\in} denotes the (unique) Wasserstein geodesic between μ\mu and ν\nu. Let us remark that (1.10) tells that the rescaled cost is convex around T=0T=0 and using functional inequalities such as the HWI inequality one can also estimate from below its second derivative under the Bakry-Émery condition (H1). It is an interesting question to obtain general versions of (1.10) in terms of Γ\Gamma-convergence. In addition, it is worth noticing that the energy, once properly rescaled, also converges to the squared Wasserstein distance, namely

To see, at least formally, why this is expected to be true, one can integrate (1.6) in time, use (1.4) and the energy conservation to get

Multiplying by TT, letting T→0T\rightarrow 0 and using (1.8) we can then handle the first term on the right-hand side as well as H(μ ∣ m)+H(ν ∣ m)\mathcal{H}(\mu\,|\,\mathfrak{m})+\mathcal{H}(\nu\,|\,\mathfrak{m}). The fact that the last integral converges to has been proved in [17, Lemma 3.3] and the same argument can be adapted verbatim to the present setting (see ); yet it is a non-trivial fact, as the entropic interpolation (ρˉtT)t∈[0,T](\bar{\rho}^{T}_{t})_{t\in[0,T]} converges to the Wasserstein geodesic (ρt0)t∈[0,T](\rho^{0}_{t})_{t\in[0,T]} between μ\mu and ν\nu, whose Fisher information needs not to be defined.

Long-time behavior of the entropic cost.

Although its asymptotic regime has been the object of recent studies in connection with the ergodic behavior of Schrödinger bridges and the limiting behavior of Sinkhorn divergence , very few results are available. In this article we prove that under mild assumptions we have

The intuition behind (1.11) is that R0,TR_{0,T} converges towards m⊗m\mathfrak{m}\otimes\mathfrak{m} and therefore the optimal coupling in SP converges towards the independent coupling of μ\mu and ν\nu. We remark that in a different although related context , the convergence of Sinkhorn divergences towards MMD divergences in the limit when the regularization parameter goes to +∞+\infty shares many analogies with (1.11).

Our second main result deals with the approximation error in (1.11) and states that it is exponentially small (see Theorem 1.4 for the rigorous statement)

provided the Bakry-Émery condition (H1) is satisfied. The proof of (1.12) is based on Theorem 1.1 asserting that the time derivative of CT(μ,ν)\mathscr{C}_{T}(\mu,\nu) is precisely −ET(μ,ν)-\mathscr{E}_{T}(\mu,\nu) and on two functional inequalities. The first one is a version of the Talagrand inequality (obtained in and also called entropic Talagrand inequality) and the second one is a functional inequality relating ∣ET(μ,ν)∣|\mathscr{E}_{T}(\mu,\nu)| with CT(μ,ν)\mathscr{C}_{T}(\mu,\nu), that we call “energy-transport” inequality (cf. Lemma 3.8). A similar inequality has been proved very recently in ; there it has been used to obtain the so-called turnpike property for mean field Schrödinger bridges. It is worth noticing that estimates such as (1.12) do not seem to follow from classical functional inequalities such as Talagrand and Log-Sobolev, whereas they can be proven using the new family of functional inequalities involving the entropic cost CT(μ,ν)\mathscr{C}_{T}(\mu,\nu). The final contribution of the article is to derive a bound similar to (1.12) for the Mean Field Schrödinger problem introduced in . Since we could not establish a generalization of the differentiation formula at Theorem 1.1 to the mean field setup, the proof of this estimate follows a different scheme, but is still based on a class of functional inequalities derived in that are the mean field versions of the entropic Talagrand and of the energy-transport inequalities mentioned above.

Organization of the paper.

The document is structured as follows: in the remainder of Section 1 we state and comment the main results, whose proofs are contained in Section 3. In Section 2 we collect, for reader’s sake, all relevant results and bibliographical references on SP. Finally, in Appendix A we prove the sharpness of a functional inequality introduced by the first-named author in , which plays an important role in this paper.

2 First and second derivative of the entropic cost

The regularity of the entropic cost w.r.t. to the time variable TT has never been investigated to the best of our knowledge, so that the following is the first result of such a kind and it plays a pivotal role in the study of both the long- and short-time behavior of CT(μ,ν)\mathscr{C}_{T}(\mu,\nu).

the map T↦CT(μ,ν)T\mapsto\mathscr{C}_{T}(\mu,\nu) is C1((0,∞))C^{1}((0,\infty)), twice differentiable a.e. and the first derivative is given by

the map T↦TCT(μ,ν)T\mapsto T\mathscr{C}_{T}(\mu,\nu) is C1((0,∞))C^{1}((0,\infty)) and twice differentiable a.e. The first derivative is given for all T>0T>0 by

where (ρˉT,vˉT)(\bar{\rho}^{T},\bar{v}^{T}) is the optimal solution in (1.4). The second derivative writes as

As implicitly stated above, by Theorem 1.1 we see that T↦ET(μ,ν)T\mapsto\mathscr{E}_{T}(\mu,\nu) is continuous on (0,∞)(0,\infty) and differentiable a.e.

3 Long-time behavior of entropic cost and energy

Under (H1) with κ≥0\kappa\geq 0 and (H2), for any μ,ν∈P(M)\mu,\nu\in\mathcal{P}(M) satisfying (H3) it holds

If μ,ν∈P(M)\mu,\nu\in\mathcal{P}(M) satisfy (H4), then it also holds

As concerns the long-time behavior of CT(μ,ν)\mathscr{C}_{T}(\mu,\nu), we provide two different proofs:

the former is direct and relies on a Γ\Gamma-convergence argument (see Section 3.1);

the latter is more technical and requires slightly stronger assumptions on μ\mu and ν\nu, but as an advantage it allows us to determine the long-time behavior of the so-called “(f,g)(f,g)-decomposition” of the optimal coupling in SP (see Section 2 for its definition, Section 3.3 for the proof).

As an application of this result we provide a new proof of the logarithmic Sobolev inequality based on entropic interpolations, in the same spirit of the recent paper , where an “entropic” proof of the HWI inequality is established.

Under (H1) with κ>0\kappa>0, for any μ=ρm∈P(M)\mu=\rho\mathfrak{m}\in\mathcal{P}(M) it holds

where the right-hand side is set equal to +∞+\infty if log⁡ρ\log\rho is not locally Sobolev.

As a further step, in the following result we improve Theorem 1.2 by providing sharp rates of convergence for both CT(μ,ν)\mathscr{C}_{T}(\mu,\nu) and ET(μ,ν)\mathscr{E}_{T}(\mu,\nu). The key message is that, under a positive curvature condition, the approximation error is asymptotically smaller than exp⁡(−κT/2)\exp(-\kappa T/2), up to constant factors depending on H(μ ∣ m)\mathcal{H}(\mu\,|\,\mathfrak{m}) and H(ν ∣ m)\mathcal{H}(\nu\,|\,\mathfrak{m}), and the rate exp⁡(−κT/2)\exp(-\kappa T/2) is sharp. As already pointed out, recall that if (H1) holds with κ>0\kappa>0, then (H2) also holds.

Let us assume that (H1) with κ>0\kappa>0 and (H4) are satisfied. Then for all T>0T>0 it holds

Furthermore, the convergence rate exp⁡(−κT/2)\exp(-\kappa T/2) in (1.14) and (1.15) is sharp in the following sense: for all μ,ν\mu,\nu as in (H4) it holds

and there exists a triplet (M′,dg′,m′)(M^{\prime},{\sf d}_{g}^{\prime},\mathfrak{m}^{\prime}) satisfying (H1) with κ>0\kappa>0 such that

holds for all μ,ν\mu,\nu satisfying (H4) if and only if α≤1/2\alpha\leq 1/2.

In other words, it may be possible to improve the constant factor in (1.14), but the convergence rate exp⁡(−κT/2)\exp(-\kappa T/2) is asymptotically sharp.

Actually, in Section 3.4 we shall prove the following (stronger) bounds

However we prefer the more compact, although slightly less precise, formulation given above. ■\blacksquare

4 Short-time behavior of the entropic cost

We provide a non-asymptotic bound for the difference TCT(μ,ν)−W22(μ,ν)/2T\mathscr{C}_{T}(\mu,\nu)-W_{2}^{2}(\mu,\nu)/2 under the assumption that the integral of the Fisher information along the displacement interpolation is finite and we also prove that this bound is sharp, as it coincides with the Taylor expansion of T↦TCT(μ,ν)T\mapsto T\mathscr{C}_{T}(\mu,\nu) around T=0T=0.added in proof: With respect to a previous version of the paper, we have been able to remove quite demanding regularity assumptions. A similar but less general result (valid only in the Euclidean setting) has recently been obtained in by other means, hence our proof is independent. This result is of particular interest as the second order term in the expansion of the rescaled cost has not been computed before (to the best of our knowledge) and its form is general enough to formulate more general version of (1.18) below, for instance in terms of Γ\Gamma-convergence, that deserve to be the object of future work.

Assume that (H1) and (H4) hold. Then for all T>0T>0 we have:

If in addition the Bakry-Émery CD(κ,N){\sf CD}(\kappa,N) condition

Note, for instance, that the CD(κ,N){\sf CD}(\kappa,N) condition is satisfied when m\mathfrak{m} is the volume measure.

5 Long time behavior of the mean field entropic cost

In the above we denoted by ∗\ast the usual convolution operator. The interaction potential WW satisfies the assumptions

In the mean field entropic cost is defined as the optimal value in (MFSP). However, if W=0W=0 this definition does not give back the usual entropic cost CT\mathscr{C}_{T} and the reason is simply the following: in the large deviations formulation of (1.2) particles are sampled with initial distribution the invariant measure m\mathfrak{m}, whereas in MFSP particles are sampled according to μ\mu. For this reason, and to strengthen the analogy with (1.2), we prefer to define the mean field entropic cost CTmf\mathscr{C}^{mf}_{T} as

With this definition we recover CT(μ,ν)\mathscr{C}_{T}(\mu,\nu) defined in (1.2) when W=0W=0.

After this premise, the following plays the same role of assumption (H1)

while regarding the marginal constraints μ\mu and ν\nu, we assume that they have finite second moment, they belong to the domain of F\mathcal{F} and have the same barycenter, i.e.

The proof we gave of Theorem 1.1 is hard to replicate for MFSP essentially because uniqueness of optimizers is currently not known. However, the long-time behavior of the mean field entropic cost can still be studied and exponential rate of convergence can be derived as well.

Assume (H’1)-(H’3) and that F(μ),F(ν)<+∞\mathcal{F}(\mu),\mathcal{F}(\nu)<+\infty. Then:

the rate of convergence is at least exp⁡(−κT)\exp(-\kappa T), i.e. there exists a decreasing function B(⋅)B(\cdot) such that

One could be more precise in the above statement and get that

where δ=F(μ)2+F(ν)2+exp⁡(−κT/2)F(μ)F(ν)\delta=\mathcal{F}(\mu)^{2}+\mathcal{F}(\nu)^{2}+\exp(-\kappa T/2)\mathcal{F}(\mu)\mathcal{F}(\nu). For sake of clarity, we prefer the more compact, although slightly less precise, formulation given above. ■\blacksquare

Preliminaries

In this section we collect some useful results concerning Markov semigroups and Schrödinger problem, either already present in the literature or extended to our framework.

With this remark in mind, let us present all the useful information about Pt{\sf P}_{t}. First of all, it enjoys the following standard a priori estimate

which can be obtained by differentiating t↦∥∣∇Ptf∣∥L2(m)2t\mapsto\||\nabla{\sf P}_{t}f|\|^{2}_{L^{2}(\mathfrak{m})}. The semigroup is also ergodic. This means that:

if m(M)=1\mathfrak{m}(M)=1, then for all f∈L2(m)f\in L^{2}(\mathfrak{m}) it holds

if m(M)=∞\mathfrak{m}(M)=\infty, then Ptf→0{\sf P}_{t}f\to 0 in L2(m)L^{2}(\mathfrak{m}) for all f∈L2(m)f\in L^{2}(\mathfrak{m}).

The curvature assumption (H1) then yields many important consequences, the first of which is the Bakry-Émery commutation estimate

For its proof as well as for all the regularizing properties of Pt{\sf P}_{t} that will be used throughout the paper, we address once more the reader to . Secondly, Pt{\sf P}_{t} enjoys an L∞L^{\infty}-Lipschitz regularization (see ), namely for all u∈L∞(m)u\in L^{\infty}(\mathfrak{m}) it holds

Then let us recall that under (H1) Hamilton’s gradient estimate is satisfied (see ): for any positive function u∈Lp∩L∞(m)u\in L^{p}\cap L^{\infty}(\mathfrak{m}) for some p∈[1,∞)p\in[1,\infty) it holds

pointwise, where κ−:=max⁡{0,−κ}\kappa^{-}:=\max\{0,-\kappa\}. Within our framework it is also well known (see for instance ) that there exists a unique kernel (transition probability) pt(x,y){\sf p}_{t}(x,y) representing Pt{\sf P}_{t} in the following sense:

The function pt{\sf p}_{t} can also be seen as the density of R0,TR_{0,T} (the joint law at time 0 and TT of the solution to (1.1)) w.r.t. m⊗m\mathfrak{m}\otimes\mathfrak{m}; it is a smooth function on (0,∞)×M×M(0,\infty)\times M\times M and the second-named author recently proved in that in great generality (and in particular under the assumption (H1)) the following upper Gaussian estimate for the kernel holds

for all t>0t>0, x,y∈Mx,y\in M and ε>0\varepsilon>0, where CκC_{\kappa} can be chosen equal to 0 if κ≥0\kappa\geq 0 in (H1). As concerns Gaussian lower bounds, if (H2) holds then by [35, Corollary 1.3] the following is satisfied

The relative entropy functional.

Let us first recall the definition of the relative entropy functional in the case of a reference measure with possibly infinite mass (see for more details). Given a σ\sigma-finite measure qq on MM, there exists a measurable function W:M→[0,∞)W:M\to[0,\infty) such that

The (f,g)𝑓𝑔(f,g)-decomposition.

It is important to stress that solving (1.2) is equivalent to finding non-negative Borel functions fT,gTf^{T},g^{T}, also called decomposition, such that

This pair of equations is known as Schrödinger system and its solvability holds under very mild assumptions (see ), in particular under:

(H4), as a consequence of [18, Proposition 2.1], the smoothness and positivity of the heat kernel and the boundedness of supp(μ),supp(ν)\mathop{\rm supp}\nolimits(\mu),\mathop{\rm supp}\nolimits(\nu);

(H1), (H2) and (H3), because of [25, Proposition 2.5] and (2.8).

Still under (H1)-(H4), the couple (fT,gT)(f^{T},g^{T}) solving (2.9) is unique up to the trivial transformation (fT,gT)↦(cfT,gT/c)(f^{T},g^{T})\mapsto(cf^{T},g^{T}/c) with c>0c>0, as proven for instance in [18, Proposition 2.1]: this fact will play an important role in several proofs (e.g. the extension of (2.20) to our setting and Lemma 3.4). A good feature of fT,gTf^{T},g^{T} is that they inherit the regularity (smoothness and integrability) of ρ,σ\rho,\sigma, the densities of μ,ν\mu,\nu respectively. More precisely,

fT,gT∈L∞(m)f^{T},g^{T}\in L^{\infty}(\mathfrak{m}) since so are ρ,σ\rho,\sigma;

A proof of this property can be found in [17, Proposition 2.7].

About point (a) a more quantitative statement is actually possible. In [18, Proposition 2.1] the second-named author proved that for all μ,ν\mu,\nu as in (H4), the following integral bounds hold for the decomposition (fT,gT)(f^{T},g^{T})

where cTc_{T} is a suitable positive constant. If we normalize gTg^{T} in such a way that

and from now on this choice will always be done, then the bounds above become

Aim of the next lemma is to improve this result by showing that the same kind of bounds holds when TT ranges in a compact subset of (0,∞)(0,\infty) or even closed half-line if we also assume (H2).

Given (H1) and μ,ν\mu,\nu as in (H4), the following hold:

for all 0<T0<T1<∞0<T_{0}<T_{1}<\infty there exists cT0,T1>0c_{T_{0},T_{1}}>0 such that

if (H2) is satisfied, then for all T0>0T_{0}>0 there exists cT0>0c_{T_{0}}>0 such that

The first equation in the Schrödinger system (2.9) and the representation formula (2.6) entail

As pT(x,y){\sf p}_{T}(x,y) is smooth in T∈(0,∞)T\in(0,\infty) and x,y∈Mx,y\in M and fT,gTf^{T},g^{T} are uniformly compactly supported (since they are supported in supp(μ)\mathop{\rm supp}\nolimits(\mu) and supp(ν)\mathop{\rm supp}\nolimits(\nu) respectively) we deduce that there exists cT0,T1>0c_{T_{0},T_{1}}>0 such that pT(x,y)≥cT0,T1{\sf p}_{T}(x,y)\geq c_{T_{0},T_{1}} in supp(μ)×supp(ν)\mathop{\rm supp}\nolimits(\mu)\times\mathop{\rm supp}\nolimits(\nu) for all T∈[T0,T1]T\in[T_{0},T_{1}], whence

provided T∈[T0,T1]T\in[T_{0},T_{1}], and this proves the first inequality in (2.11). For the second one it is sufficient to swap the roles of fTf^{T} and gTg^{T}. If we further assume (H2), then the Gaussian lower bound (2.8) holds. Hence for all T0>0T_{0}>0 there exists cT0>0c_{T_{0}}>0 such that pT(x,y)≥cT0{\sf p}_{T}(x,y)\geq c_{T_{0}} in supp(μ)×supp(ν)\mathop{\rm supp}\nolimits(\mu)\times\mathop{\rm supp}\nolimits(\nu) for all T≥T0T\geq T_{0}, whence by the same argument as above the first inequality in (2.12) follows. By interchanging the roles of fTf^{T} and gTg^{T}, the same conclusion follows for gTg^{T}. ∎

The Benamou-Brenier formulation.

The equivalence between (2.9) and (1.2) is extremely fruitful, as it enables to fully describe the optimal pair (ρˉT,vˉT)(\bar{\rho}^{T},\bar{v}^{T}) in the fluid-dynamical description (1.4) of the entropic cost. Indeed,

Actually, (1.4) and the identities above are nothing but a reparametrization of the Benamou-Brenier-like formula for the entropic cost established in [19, Theorem 4.2] and , which reads as

In (2.13) the infimum runs over all weak solutions of the continuity equation (1.5) satisfying the marginal constraints ρ0m=μ\rho_{0}\mathfrak{m}=\mu and ρ1m=ν\rho_{1}\mathfrak{m}=\nu, while (ρT,vT)(\rho^{T},v^{T}) defined above is the unique optimal density-velocity couple for (2.13). It is also useful to see the second term in the right-hand side of (2.13) as an action functional, whose arguments are the curves (ρt)(\rho_{t}) and (vt)(v_{t}), i.e.

When T=0T=0, AT\mathscr{A}_{T} is a purely kinetic energy and by the well-known Benamou-Brenier formulation of the optimal transport problem ,

where the infimum is taken over the same set as in (2.13). In this case optimal couples shall be denoted by (ρ0,v0)(\rho^{0},v^{0}). Let us also recall that

Finally, the expression of the conserved quantity ET(μ,ν)\mathscr{E}_{T}(\mu,\nu) in terms of the rescaled optimal couple (ρT,vT)(\rho^{T},v^{T}) reads as

Dual formulation.

For future reference, it is also worth mentioning the fact that, for all μ,ν\mu,\nu satisfying (H4), the (rescaled) entropic cost admits the following dual representation (see )

where Q1Tϕ:=−Tlog⁡PT(exp⁡(−ϕ/T))Q_{1}^{T}\phi:=-T\log{\sf P}_{T}(\exp(-\phi/T)).

Geometric information.

As regards the relationship between Schrödinger problem and lower Ricci bounds, let us recall that in the first-named author showed that (H1) implies the following distorted κ\kappa-convexity inequality of the entropy along entropic interpolations:

for all t∈t\in, where μtT:=ρtTm\mu_{t}^{T}:=\rho_{t}^{T}\mathfrak{m}. Truth to be told, the result in was stated in the framework of compact Riemannian manifolds satisfying (H1) with κ>0\kappa>0 and endowed with the volume measure, but the proof can be adapted to our setting when (H2) holds, by relying on [17, Lemma 3.7]. Indeed, if we set

where cδc_{\delta} is a normalization constant so that ρtT,δ\rho_{t}^{T,\delta} is still a probability density, then the identities

hold both in the classical sense and as strong W1,2W^{1,2}-limits, where L:=Δ/2−∇U⋅∇{\sf L}:=\Delta/2-\nabla U\cdot\nabla is the generator of (Pt)({\sf P}_{t}). Together with (2.4) and (2.5), this is sufficient to follow the lines of [9, Lemma 3.6, Lemma 3.7] and [17, Lemma 3.7] and deduce the following

Under (H1), given μ,ν\mu,\nu as in (H4) and with the same notations as in (2.21), for all δ>0\delta>0 and t∈t\in define

Then hf∈C()∩C2((0,1])h_{f}\in C()\cap C^{2}((0,1]), hb∈C()∩C2([0,1))h_{b}\in C()\cap C^{2}([0,1)),

By [9, Lemma 4.1] and following the proof of Theorem 1.4 therein, we obtain

and again by dominated convergence the right-hand side above converges to

Proof of the main results

A rather straightforward proof of the long-time behavior of CT(μ,ν)\mathscr{C}_{T}(\mu,\nu) can be obtained by a Γ\Gamma-convergence argument. Its essence is contained in the following

As a first step, we claim that for any lower bounded lower semicontinuous function ϕ\phi on M×MM\times M it holds

Since R0,Tn(M×M)=1R_{0,T_{n}}(M\times M)=1, the claim is trivial for constant functions; hence, without loss of generality we can assume that ϕ≥0\phi\geq 0 and, under this further assumption, the Gaussian lower bound (2.8) together with Fatou’s lemma yields (3.1). By Portmanteau theorem this is equivalent to say that

After this premise, let π∈Π(μ,ν)\pi\in\Pi(\mu,\nu), (πn)⊂Π(μ,ν)(\pi_{n})\subset\Pi(\mu,\nu) and assume that πn⇀π\pi_{n}\rightharpoonup\pi. The lower semicontinuity of the relative entropy w.r.t. to both its arguments together with (3.2) gives that

thus the Γ\Gamma-liminf inequality. To complete the proof, let π∈Π(μ,ν)\pi\in\Pi(\mu,\nu) and find a sequence (πn)⊂Π(μ,ν)(\pi^{n})\subset\Pi(\mu,\nu) such that

To this aim it is not restrictive to assume that H(π ∣ m⊗m)<∞\mathcal{H}(\pi\,|\,\mathfrak{m}\otimes\mathfrak{m})<\infty and it is also easy to see that πn≡π\pi^{n}\equiv\pi satisfies (3.3). Indeed, the Gaussian lower bound (2.8) and the fact that (H1) holds with κ≥0\kappa\geq 0 imply that

for some constant CC independent of nn. Next, observe that by definition of pTn(⋅,⋅){\sf p}_{T_{n}}(\cdot,\cdot) and πn\pi^{n} we have

Let us point out that in the previous lemma and thus in Theorem 1.2, (H2) is required only when κ=0\kappa=0. As regards the long-time behavior of ET(μ,ν)\mathscr{E}_{T}(\mu,\nu), we need to determine in which way ET(μ,ν)\mathscr{E}_{T}(\mu,\nu) is controlled in terms of CT(μ,ν)\mathscr{C}_{T}(\mu,\nu).

whence trivially the desired upper bound for TET(μ,ν)T\mathscr{E}_{T}(\mu,\nu). On the other hand, Young’s inequality and (2.16) yield

so that plugging this inequality into the previous identity gives also the lower bound for TET(μ,ν)T\mathscr{E}_{T}(\mu,\nu). ∎

We are now in the position to prove Theorem 1.2.

It is easily seen that the unique optimal coupling in

Dividing by TT (3.6) and letting T→∞T\to\infty, the long-time behavior of ET(μ,ν)\mathscr{E}_{T}(\mu,\nu) is established as well. ∎

2 Proof of Theorem 1.1 and Theorem 1.6

The proof of Theorem 1.1 requires the preparatory Lemmas 3.3, 3.4 and 3.5. In the first one we show that the Fisher information of the entropic interpolation is non-increasing as a function of TT.

Under the assumptions of Theorem 1.1, let (ρT,vT)(\rho^{T},v^{T}) be optimal for the formulation (2.13) and (ρ0,v0)(\rho^{0},v^{0}) for (2.15). Then the function

Fix 0≤T1<T20\leq T_{1}<T_{2}. Summing the inequalities

and dividing by (T22−T12)/8(T_{2}^{2}-T_{1}^{2})/8 we obtain

In the second lemma we prove the continuity in TT of the functions fTf^{T} and gTg^{T} given by (2.9) with respect to the Lp(m)L^{p}(\mathfrak{m}) norm, for any p∈[1,∞)p\in[1,\infty).

Under the assumptions of Theorem 1.1, the functions T↦fTT\mapsto f^{T} and T↦gTT\mapsto g^{T} are continuous from (0,∞)(0,\infty) to Lp(m)L^{p}(\mathfrak{m}), for any p∈[1,∞)p\in[1,\infty).

As a byproduct, the functions T↦PTfTT\mapsto{\sf P}_{T}f^{T} and T↦PTgTT\mapsto{\sf P}_{T}g^{T} are continuous from (0,∞)(0,\infty) to Lp(m)L^{p}(\mathfrak{m}), p∈[1,∞)p\in[1,\infty), as well.

Let T0>0T_{0}>0, 0<δ<T00<\delta<T_{0} and denote by CδC_{\delta} the positive constant provided by Lemma 2.1-(i) on the interval [T0−δ,T0+δ][T_{0}-\delta,T_{0}+\delta]. From the first bound in (2.11) we immediately deduce that (fT)T∈[T0−δ,T0+δ](f^{T})_{T\in[T_{0}-\delta,T_{0}+\delta]} is bounded in L∞(m)L^{\infty}(\mathfrak{m}), hence in Lp(m)L^{p}(\mathfrak{m}) for all p∈[1,∞]p\in[1,\infty], because all the functions fTf^{T} are supported in supp(μ)\mathop{\rm supp}\nolimits(\mu) and this has finite mass, because bounded. To prove the same property for (gT)T∈[T0−δ,T0+δ](g^{T})_{T\in[T_{0}-\delta,T_{0}+\delta]} requires a more technical argument.

From (2.7) with ε=1\varepsilon=1 we first oberve that

and since supp(fT)⊂supp(μ)\mathop{\rm supp}\nolimits(f^{T})\subset\mathop{\rm supp}\nolimits(\mu), supp(gT)⊂supp(ν)\mathop{\rm supp}\nolimits(g^{T})\subset\mathop{\rm supp}\nolimits(\nu) for all T>0T>0, μ,ν\mu,\nu have bounded support and x↦m(Bt(x))x\mapsto\mathfrak{m}(B_{t}(x)) is a continuous function,

with CδC_{\delta} independent of TT. Thus if we combine this inequality with the previous one and recall the chosen normalization (2.10) we obtain 1≤Cδ′∥fT∥L1(m)1\leq C^{\prime}_{\delta}\|f^{T}\|_{L^{1}(\mathfrak{m})} for all T∈[T0−δ,T0+δ]T\in[T_{0}-\delta,T_{0}+\delta] for some Cδ′>0C^{\prime}_{\delta}>0 (depending on T0T_{0} and κ\kappa too), i.e.

Plugging this inequality into the second bound in (2.11) yields

Hence also (gT)T∈[T0−δ,T0+δ](g^{T})_{T\in[T_{0}-\delta,T_{0}+\delta]} is bounded in Lp(m)L^{p}(\mathfrak{m}) for all p∈[1,∞]p\in[1,\infty], because all the functions gTg^{T} are supported in supp(ν)\mathop{\rm supp}\nolimits(\nu) and this has finite mass.

where both limits have to be understood in Lp(m)L^{p}(\mathfrak{m}) for p∈[1,∞)p\in[1,\infty).

and observe that the second term on the right-hand side trivially vanishes as h→0h\to 0. As regards the first one,

Since, as already remarked, (fT)T∈[T0−δ,T0+δ](f^{T})_{T\in[T_{0}-\delta,T_{0}+\delta]} is bounded in L∞(m)L^{\infty}(\mathfrak{m}) and all the functions fTf^{T} are supported in supp(μ)\mathop{\rm supp}\nolimits(\mu), there exists M>0M>0 sufficiently large such that 0≤fT≤M\mathds1supp(μ)0\leq f^{T}\leq M\mathds{1}_{\mathop{\rm supp}\nolimits(\mu)} for all T∈[T0−δ,T0+δ]T\in[T_{0}-\delta,T_{0}+\delta], whence

by the maximum principle. As the right-hand side above belongs to L1∩L∞(m)L^{1}\cap L^{\infty}(\mathfrak{m}) and does not depend on TT, by (3.10) and the dominated convergence theorem we thus infer that

whence, combining this fact with the previous steps, the validity of the first limit in (3.9). The argument for the second limit is completely analogous.

As a consequence, up to extract a further (not relabeled) subsequence, we have that the limits in (3.9) hold m\mathfrak{m}-a.e. Therefore, if we look at the Schrödinger system (2.9) at time T0+hT_{0}+h, which reads as

which is possible because PT0+hfT0+h,PT0+hgT0+h>0{\sf P}_{T_{0}+h}f^{T_{0}+h},{\sf P}_{T_{0}+h}g^{T_{0}+h}>0, we deduce that

For T↦CT(μ,ν)T\mapsto\mathscr{C}_{T}(\mu,\nu), by (2.9) we observe that the cost can be rewritten as

and by Lemma 3.4, the locally uniform L∞L^{\infty}-bounds provided by Lemma 2.1 and by dominated convergence it is easy to see that the right-hand side above continuously depends on TT.

As regards T↦ET(μ,ν)T\mapsto\mathscr{E}_{T}(\mu,\nu), let us start observing that the second identity in (2.18) can be equivalently rewritten as

so fix t∈(0,1)t\in(0,1) and note that, by integration by parts first and Cauchy-Schwarz inequality then, we obtain

The second summand after the last inequality clearly vanishes as h→0h\to 0, because of Lemma 3.4. As concerns the first one, argue as in the proof of Lemma 3.4 and notice that for hh sufficiently small, say ∣h∣<δ|h|<\delta with 0≤δ≤T0\leq\delta\leq T, all the functions fT+hf^{T+h} are uniformly bounded in L∞(m)L^{\infty}(\mathfrak{m}) and supported in supp(μ)\mathop{\rm supp}\nolimits(\mu), hence also uniformly bounded in L2(m)L^{2}(\mathfrak{m}). By the a priori estimate (2.1), this is sufficient to conclude that there exists M>0M>0 sufficiently large such that ∥LP(T+h)tfT+h∥L2(m)≤M\|{\sf L}{\sf P}_{(T+h)t}f^{T+h}\|_{L^{2}(\mathfrak{m})}\leq M for all ∣h∣<δ|h|<\delta. Hence, by Lemma 3.4, also the first summand after the last inequality converges to 0 in the limit, whence the continuity of T↦ET(μ,ν)T\mapsto\mathscr{E}_{T}(\mu,\nu).

and by the previous discussion the right-hand side is continuous on (0,∞)(0,\infty). ∎

Consider the map T↦AT(ρT,vT)T\mapsto\mathscr{A}_{T}(\rho^{T},v^{T}), let h>0h>0, write

and observe that the first term on the right-hand side is non-positive by optimality of (ρT+h,vT+h)(\rho^{T+h},v^{T+h}) for AT+h\mathscr{A}_{T+h}. For the second one

where ∂TAT\partial_{T}\mathscr{A}_{T} denotes the partial derivative w.r.t. TT of (T,ρ,v)↦AT(ρ,v)(T,\rho,v)\mapsto\mathscr{A}_{T}(\rho,v). Combining these two facts we deduce

and remark that now the second term on the right-hand side is non-negative by optimality of (ρT,vT)(\rho^{T},v^{T}) for AT\mathscr{A}_{T}. As regards the first one,

which together with (3.13) implies that T↦AT(ρT,vT)T\mapsto\mathscr{A}_{T}(\rho^{T},v^{T}) is everywhere right differentiable on (0,∞)(0,\infty). Left differentiability follows by an analogous argument. Indeed if h<0h<0, then the first term on the right-hand side in (3.11) is non-negative and (3.12) still holds true with h↑0h\uparrow 0 instead of h↓0h\downarrow 0, whence

Applying the same considerations to (3.14) and (3.15) gives

and combining this inequality with the previous one entails the desired left differentiability. Therefore T↦AT(ρT,vT)T\mapsto\mathscr{A}_{T}(\rho^{T},v^{T}) is everywhere differentiable on (0,∞)(0,\infty), a fortiori continuous therein, and

Therefore T↦TCT(μ,ν)T\mapsto T\mathscr{C}_{T}(\mu,\nu) is everywhere differentiable (hence continuous) on (0,∞)(0,\infty) as well and, by (2.13) and the very definition of CT(μ,ν)\mathscr{C}_{T}(\mu,\nu) and ET(μ,ν)\mathscr{E}_{T}(\mu,\nu),

which proves both formulas for the first derivative of T↦TCT(μ,ν)T\mapsto T\mathscr{C}_{T}(\mu,\nu), thanks to (2.14). Now notice that the right-hand side above is continuous on (0,∞)(0,\infty): T↦CT(μ,ν)T\mapsto\mathscr{C}_{T}(\mu,\nu) is continuous on (0,∞)(0,\infty) since so is T↦TCT(μ,ν)T\mapsto T\mathscr{C}_{T}(\mu,\nu), while T↦TET(μ,ν)T\mapsto T\mathscr{E}_{T}(\mu,\nu) is continuous by Lemma 3.5. Hence T↦TCT(μ,ν)T\mapsto T\mathscr{C}_{T}(\mu,\nu) belongs to C1((0,∞))C^{1}((0,\infty)). The fact that T↦TCT(μ,ν)T\mapsto T\mathscr{C}_{T}(\mu,\nu) is also twice differentiable a.e. follows from the fact that

and the last term on the right-hand side is the product of a linear function and a monotone (non-increasing) one, thanks to Lemma 3.3.

This concludes the proof of (ii). Claim (i) is a straightforward consequence as well as the formula for the second derivative of T↦TCT(μ,ν)T\mapsto T\mathscr{C}_{T}(\mu,\nu). ∎

From Theorem 1.1, the proof of Theorem 1.6 follows.

As regards (1.17), we only prove the upper bound, the lower bound being an immediate consequence of the Benamou-Brenier formulation (2.13). Since T↦TCT(μ,ν)T\mapsto T\mathscr{C}_{T}(\mu,\nu) belongs to C1((0,∞))C^{1}((0,\infty)) we can use the fundamental theorem of calculus that gives for all ε>0\varepsilon>0

where the inequality is motivated by Lemma 3.3. The conclusion follows letting ε→0\varepsilon\rightarrow 0 and using the convergence of the rescaled entropic cost towards W22(μ,ν)/2W^{2}_{2}(\mu,\nu)/2.

The coefficients can be identified thanks to Theorem 1.1. Indeed

thanks to (a). Therefore it only remains to show (a), (b), and (c); we shall start with the first one. On the one hand, by Lemma 3.3

On the other hand, as CD(κ,N){\sf CD}(\kappa,N) holds, we can and shall rely on powerful facts established in : by Propositions 5.1 and 5.4 therein we know that

and by [18, Lemma 4.2] we also know that for all xˉ∈M\bar{x}\in M there exist C′,C′′,r>0C^{\prime},C^{\prime\prime},r>0 such that

With this said, let B⊂MB\subset M be any Borel set and let us prove that

To this end, since μ,ν\mu,\nu are compactly supported, there exists a bounded set containing the supports of ρt0\rho_{t}^{0} for all t∈t\in, so that up to choose a larger rr we can assume that supp(ρt0)⊂Br(xˉ)\mathop{\rm supp}\nolimits(\rho_{t}^{0})\subset B_{r}(\bar{x}) for all t∈t\in. By (3.18) and the trivial fact that \mathds1B∩Br(xˉ)∈L1(m)\mathds{1}_{B\cap B_{r}(\bar{x})}\in L^{1}(\mathfrak{m}) this implies that (ρtTm)(B∩Br(xˉ))→(ρt0m)(B∩Br(xˉ))(\rho_{t}^{T}\mathfrak{m})(B\cap B_{r}(\bar{x}))\to(\rho_{t}^{0}\mathfrak{m})(B\cap B_{r}(\bar{x})) as T↓0T\downarrow 0 for all t∈t\in. By (3.19) and the fact that under the CD(κ,N){\sf CD}(\kappa,N) condition the function in the right-hand side of (3.19) is integrable and converges monotonically to 0 as T↓0T\downarrow 0, it also follows

whence (3.20). By [2, Theorem 7.6] this implies

Combining this inequality with (3.17) yields the desired continuity at T=0T=0 and thus (a).

On the other hand, by following the argument of Theorem 1.1 and recalling (2.15) we see that

where (vt0)(v_{t}^{0}) is any choice of gradient of intermediate Kantorovich potentials associated with the geodesic (ρt0m)(\rho_{t}^{0}\mathfrak{m}), and by (a) this implies

Combining this inequality with (3.21) yields the differentiability of T↦TCT(μ,ν)T\mapsto T\mathscr{C}_{T}(\mu,\nu) at T=0T=0 and by (3.16) we see that the derivative is continuous there. Finally, by (2.13), the computations above for determining the coefficients AA and BB, and the properties (a) and (b) it is not difficult to see that (c) holds too. ∎

3 Proof of Theorem 1.2 and Corollary 1.3

In this section we provide the reader with a second proof of Theorem 1.2, whose crucial ingredient is a long-time analogue of Lemma 3.4. This is precisely the content of the next result. With respect to the proof given in Section 3.1, here the marginals μ,ν\mu,\nu satisfy (H4), hence a condition stronger than (H3); as a consequence we gain information on the long-time behavior of fTf^{T} and gTg^{T} separately, determining their limits as T→∞T\to\infty. Let us point out again that (H2) below is required only when κ=0\kappa=0.

Under the assumptions (H1) with κ≥0\kappa\geq 0, (H2) and (H4), for the functions fT,gTf^{T},g^{T} given by (2.9) it holds

where both limits are in Lp(m)L^{p}(\mathfrak{m}) for any p∈[1,∞)p\in[1,\infty).

where both limits are in Lp(m)L^{p}(\mathfrak{m}) for any p∈[1,∞)p\in[1,\infty).

Let T0>0T_{0}>0 and denote by CC the positive constant provided by Lemma 2.1-(ii) on the interval [T0,∞)[T_{0},\infty). From the first bound in (2.12) we immediately deduce that (fT)T≥T0(f^{T})_{T\geq T_{0}} is bounded in L∞(m)L^{\infty}(\mathfrak{m}), hence in Lp(m)L^{p}(\mathfrak{m}) for all p∈[1,∞]p\in[1,\infty], as already argued in the proof of Lemma 3.4. From (2.7) with Cκ=0C_{\kappa}=0 (since κ≥0\kappa\geq 0), (2.10), the fact that μ,ν\mu,\nu have bounded support and arguing as in Lemma 3.4 we also have

Plugging this inequality into the second bound in (2.12) yields that also (gT)T≥T0(g^{T})_{T\geq T_{0}} is bounded in L∞(m)L^{\infty}(\mathfrak{m}), thus in Lp(m)L^{p}(\mathfrak{m}) for all p∈[1,∞]p\in[1,\infty]. As a direct consequence of this and of the maximum, we deduce that also (PTfT)T≥T0({\sf P}_{T}f^{T})_{T\geq T_{0}} and (PTgT)T≥T0({\sf P}_{T}g^{T})_{T\geq T_{0}} are bounded in L∞(m)L^{\infty}(\mathfrak{m}).

where both limits have to be understood in Lp(m)L^{p}(\mathfrak{m}) for p∈[1,∞)p\in[1,\infty). To this aim observe that

where the second inequality is motivated by the fact that PT=PT−T0∘PT0{\sf P}_{T}={\sf P}_{T-T_{0}}\circ{\sf P}_{T_{0}} and PT−T0{\sf P}_{T-T_{0}} is a contraction in Lp(m)L^{p}(\mathfrak{m}). Therefore, arguing as in the proof of Lemma 3.4 we can see that the first term on the right-hand side above vanishes as T→∞T\to\infty. On the other hand, the ergodicity of PT{\sf P}_{T} and (H2) entail that

hence in Lp(m)L^{p}(\mathfrak{m}) for all p∈p\in, while for p>2p>2 it is sufficient to observe that

and use (3.27). Therefore, also the second term on the second line on the right-hand side in (3.26) converges to 0 as T→∞T\to\infty and this proves the first identity in (3.25). An analogous argument leads to the second one.

As a consequence, up to extract a further (not relabeled) subsequence, the limits in (3.25) hold m\mathfrak{m}-a.e. and if we pass to the limit as T→∞T\to\infty in the Schrödinger system (2.9) we get

It is now sufficient to combine this information with (3.25) to get (3.23). ∎

We are now in the position to determine the long-time behavior of the entropic cost.

Let us first observe that for all T0>0T_{0}>0 there exists C>0C>0 such that for all T≥T0T\geq T_{0} it holds

Indeed, both upper bounds have been proven in Lemma 3.6, while the lower ones can be deduced relying on (2.8) and the boundedness of supp(μ)\mathop{\rm supp}\nolimits(\mu), supp(ν)\mathop{\rm supp}\nolimits(\nu). More precisely, for (3.28a) there exists C>0C>0 such that

the last identity being motivated by (2.10). For the same reason and taking (3.24) into account

which proves also (3.28b). With this said, by the very definition of the entropic cost and its equivalence with (2.9) we have

and by (3.28a), (3.28b), (3.23) and the dominated convergence theorem we can pass to the limit as T→∞T\to\infty in the identity above, thus getting

As concerns Corollary 1.3, let us recall an approximation result, whose proof can be found in [17, Lemma 3.1].

Let μ∈P(M)\mu\in\mathcal{P}(M) with H(μ ∣ m)<∞\mathcal{H}(\mu\,|\,\mathfrak{m})<\infty and

Then there exists (μn)⊂P(M)(\mu_{n})\subset\mathcal{P}(M) with μn=ρnm\mu_{n}=\rho_{n}\mathfrak{m}, ρn∈Cc∞(M)\rho_{n}\in C^{\infty}_{c}(M) such that H(μn ∣ m)→H(μ ∣ m)\mathcal{H}(\mu_{n}\,|\,\mathfrak{m})\to\mathcal{H}(\mu\,|\,\mathfrak{m}) and

We can now provide an “entropic” proof of the logarithmic Sobolev inequality.

First of all, we can assume that the right-hand side in (1.13) is finite, otherwise the statement is trivially true; by Lemma 3.7 we can also assume that μ\mu has compact support, is absolutely continuous w.r.t. m\mathfrak{m} and its density ρ\rho is smooth. Hence, if we take any probability measure ν=σm\nu=\sigma\mathfrak{m} with compact support and smooth density, (H4) is satisfied. In addition, fTf^{T} and gTg^{T} are also compactly supported and smooth, as they inherit the regularity of ρ\rho and σ\sigma respectively.

After this premise, let us show that t↦H(μtT ∣ m)t\mapsto\mathcal{H}(\mu_{t}^{T}\,|\,\mathfrak{m}) is differentiable at t=0t=0 and

To this aim, let us recall that from [17, Lemma 3.7] for every δ>0\delta>0 the map t↦H(μtT,δ ∣ m)t\mapsto\mathcal{H}(\mu_{t}^{T,\delta}\,|\,\mathfrak{m}) belongs to C()∩C2((0,1))C()\cap C^{2}((0,1)) and it holds

where ρtT,δ\rho_{t}^{T,\delta}, μtT,δ\mu_{t}^{T,\delta} and vtT,δv_{t}^{T,\delta} are defined as in (2.21). If we integrate (3.30) on [0,t][0,t] with t≤1t\leq 1 we obtain

and by the dominated convergence theorem it is easy to see that the left-hand side converges to H(μtT ∣ m)−H(μ ∣ m)\mathcal{H}(\mu_{t}^{T}\,|\,\mathfrak{m})-\mathcal{H}(\mu\,|\,\mathfrak{m}) as δ↓0\delta\downarrow 0. As regards the right-hand one, to prove that

we borrow an argument used in [17, Lemma 3.11] and we report it here for reader’s sake. By the very definition of vtT,δv_{t}^{T,\delta} and since

the desired conclusion is achieved if we are able to prove that

To this aim, notice that ρtT,δ=ρtT+δPTtfT+δPT(1−t)gT+δ2\rho_{t}^{T,\delta}=\rho_{t}^{T}+\delta{\sf P}_{Tt}f^{T}+\delta{\sf P}_{T(1-t)}g^{T}+\delta^{2}, whence using either PTtfT+δ≥PTtfT{\sf P}_{Tt}f^{T}+\delta\geq{\sf P}_{Tt}f^{T} or PTtfT+δ≥δ{\sf P}_{Tt}f^{T}+\delta\geq\delta it is easy to infer that

All the right-hand sides above are integrable on ×M\times M (either by (2.17) or by the Bakry-Émery contraction estimate (2.3) together with fT∈Cc∞(M)f^{T}\in C^{\infty}_{c}(M)), thus by dominated convergence (3.32a) follows. An analogous argument holds for (3.32b). Therefore we can pass to the limit as δ↓0\delta\downarrow 0 in (3.31) and get

whence (3.29). This allows us to differentiate (2.20) at t=0t=0 and obtain

from the very definition of v0Tv_{0}^{T}, so that

where L=Δ/2−∇U⋅∇{\sf L}=\Delta/2-\nabla U\cdot\nabla is the generator of Pt{\sf P}_{t}, self-adjoint w.r.t. m\mathfrak{m}. By (3.28b), (3.23) and the boundedness of supp(μ)\mathop{\rm supp}\nolimits(\mu) we infer that the first integral on the right-hand side vanishes as T→∞T\to\infty by dominated convergence, whence (3.34).

From (3.33), (3.34) and Theorem 1.2 the logarithmic Sobolev inequality (1.13) follows. ∎

4 Proof of Theorem 1.4

The proof of Theorem 1.4 relies on Theorem 1.1 and two other ingredients: the entropic Talagrand inequality put forward in and an “energy-transport” inequality relating ET(μ,ν)\mathscr{E}_{T}(\mu,\nu) and CT(μ,ν)\mathscr{C}_{T}(\mu,\nu). The former states that if (H1) holds with κ>0\kappa>0, then for all μ,ν\mu,\nu as in (H3) and for all t∈(0,1)t\in(0,1)

Truth to be told, as for (2.20) also the entropic Talagrand inequality (3.35) was stated in in the framework of compact Riemannian manifolds satisfying (H1) with κ>0\kappa>0 and endowed with the volume measure, but since (3.35) is deduced from (2.20) and the latter has been generalized to the present framework in Section 2, there is no problem in applying it.

The “energy-transport” inequality relating ET(μ,ν)\mathscr{E}_{T}(\mu,\nu) and CT(μ,ν)\mathscr{C}_{T}(\mu,\nu) is expressed in the following

Assume that (H1) and (H4) hold. Then for all T>0T>0 it holds

From the second identity in (2.18) we have

so that by Cauchy-Schwarz inequality and by (2.16), (3.37) follows.

As regards (3.36), using the same notations as in (2.21), let us first point out that, being ρtT,δ\rho_{t}^{T,\delta} and vtT,δv_{t}^{T,\delta} optimal for SP with marginals μ0T,δ\mu_{0}^{T,\delta} and μ1T,δ\mu_{1}^{T,\delta}, we have

for δ>0\delta>0, then by (3.39) evaluated in t=1/2t=1/2 and by Cauchy-Schwarz inequality

In order to estimate the right-hand side above, observe that by (2.23)

with hfh_{f}, hbh_{b} defined as in (2.22), so that by Lemma 2.2

while on the other hand, by (3.41a) and a completely analogous argument

Combining these inequalities with (3.40), we obtain

and it is now sufficient to pass to the limit as δ↓0\delta\downarrow 0 to get the conclusion. The fact that H(μ0T,δ ∣ m)→H(μ ∣ m)\mathcal{H}(\mu_{0}^{T,\delta}\,|\,\mathfrak{m})\to\mathcal{H}(\mu\,|\,\mathfrak{m}) and H(μ1T,δ ∣ m)→H(ν ∣ m)\mathcal{H}(\mu_{1}^{T,\delta}\,|\,\mathfrak{m})\to\mathcal{H}(\nu\,|\,\mathfrak{m}) is motivated by dominated convergence; the same is true for ET(μ0T,δ,μ1T,δ)→ET(μ,ν)\mathscr{E}_{T}(\mu_{0}^{T,\delta},\mu_{1}^{T,\delta})\to\mathscr{E}_{T}(\mu,\nu) taking (3.39) into account. ∎

Let us point out that the bound (3.37) is better than the one obtained from (3.36) via a Taylor expansion around κ=0\kappa=0, so that (3.36) is not sharp and the reader might wonder whether (3.36) could be improved by replacing exp⁡(κT/2)\exp(\kappa T/2) with exp⁡(κT)\exp(\kappa T). A posteriori, this is not possible because if (3.36) held with exp⁡(ακT)\exp(\alpha\kappa T) instead of exp⁡(κT/2)\exp(\kappa T/2) for some α>1/2\alpha>1/2, then this would contradict the sharpness of (1.14) as stated in Theorem 1.4. ■\blacksquare

We are now in the position to prove Theorem 1.4.

From Theorem 1.1 and Theorem 1.2 we have that

and by Lemma 3.8 the right-hand side above is controlled as follows

By the entropic Talagrand inequality (3.35) for t=1/2t=1/2 and by (H1) with κ>0\kappa>0, which implies H(⋅ ∣ m)≥0\mathcal{H}(\cdot\,|\,\mathfrak{m})\geq 0, it holds

for all S≥TS\geq T and an analogous inequality holds with μ\mu and ν\nu swapped. Therefore the right-hand side in (3.43) is bounded from above by

The bound (1.16a) follows by calculating explicitly the integral term and this, in turn, trivially implies (1.14).

For the sharpness of (1.14), we shall prove that there exists a triplet (M′,dg′,m′)(M^{\prime},{\sf d}_{g}^{\prime},\mathfrak{m}^{\prime}) satisfying (H1) with κ>0\kappa>0 and μ,ν∈P(M′)\mu,\nu\in\mathcal{P}(M^{\prime}) satisfying (H4) such that

and the stochastic process associated to the SDE (1.1) is the stationary Ornstein-Uhlenbeck process, whence by well-known results (see e.g. ) its transition probabilities w.r.t. m′\mathfrak{m}^{\prime} admit the following explicit representation

and it is now sufficient to remark that the right-hand side is asymptotically positive, so that for TT large enough we are allowed to apply the logarithm to both sides of the inequality above, whence (3.44).

As regards (1.15), it is sufficient to estimate the right-hand side in (3.36) by following the same reasoning as above. To prove that also (1.15) is sharp, it is easy to see that if there exists α>1/2\alpha>1/2 such that

holds for all μ,ν\mu,\nu satisfying (H4), then for any such μ,ν\mu,\nu there exist C,T0>0C,T_{0}>0 sufficiently large such that

and if we plug this inequality into (3.42), instead of Lemma 3.8, then we get ∣CT(μ,ν)−H(μ ∣ m)−H(ν ∣ m)∣≤Cexp⁡(−ακT)|\mathscr{C}_{T}(\mu,\nu)-\mathcal{H}(\mu\,|\,\mathfrak{m})-\mathcal{H}(\nu\,|\,\mathfrak{m})|\leq C\exp(-\alpha\kappa T), whence

and this clearly contradicts the sharpness of (1.14). ∎

5 Proof of Theorem 1.7

for all t∈[0,T]t\in[0,T], see [3, Eq. 92] and also .

After this premise, the first ingredient needed for the proof of Theorem 1.7 is the following

Using the time-reversal relation (3.49) we deduce that

Using the change of variable s=T−ts=T-t and the very definition of time reversal, we rewrite the right-hand side as

so that the conclusion is easily obtained. ∎

As for the proof of Theorem 1.4, also in this case we rely on an entropic Talagrand inequality. In the mean field setting (and taking into account Remark 1.9) this reads as follows

for all t∈(0,1)t\in(0,1) and μ,ν\mu,\nu satisfying (H’3), see [3, Corollary 1.3].

We are now in position to prove Theorem 1.7.

Appendix A On the sharpness of the entropic Talagrand inequality

Under (H1) with κ>0\kappa>0 it holds m∈P2(M)\mathfrak{m}\in\mathcal{P}_{2}(M), so that m\mathfrak{m} satisfies (H3) and it is then licit to choose ν=m\nu=\mathfrak{m} in (3.35), which gives rise to the following version of the entropic Talagrand inequality, closer to the classical Talagrand inequality known in optimal transport:

for all μ\mu as in (H3). In neither the sharpness of (A.1) nor the one of (3.35) was investigated. Aim of this appendix is to remedy this lack.

Assume that (H1) holds with κ>0\kappa>0. Then the entropic Talagrand inequality (A.1) is sharp, since there exists a triplet (M′,dg′,m′)(M^{\prime},{\sf d}_{g}^{\prime},\mathfrak{m}^{\prime}) satisfying (H1) with κ>0\kappa>0 such that

As a byproduct, also the entropic Talagrand inequality (3.35) is sharp in the following sense: there exists a triplet (M′,dg′,m′)(M^{\prime},{\sf d}_{g}^{\prime},\mathfrak{m}^{\prime}) satisfying (H1) with κ>0\kappa>0 such that

As a first step, we shall prove that (A.2) is equivalent to

where QtTϕ:=−Tlog⁡PTt(exp⁡(−ϕ/T)){\sf Q}_{t}^{T}\phi:=-T\log{\sf P}_{Tt}(\exp(-\phi/T)). To this aim, notice that by (2.19) and the symmetry of the entropic cost, i.e. CT(μ,m)=CT(m,μ)\mathscr{C}_{T}(\mu,\mathfrak{m})=\mathscr{C}_{T}(\mathfrak{m},\mu), (A.2) is equivalent to

for all μ∈P(M)\mu\in\mathcal{P}(M) satisfying (H3) and for all ϕ∈Cb(M)\phi\in C_{b}(M). Then observe that

where ψ+:=max⁡{ψ,0}\psi^{+}:=\max\{\psi,0\}, the identity being a byproduct of the well-known variational representation of the entropy (see ). Hence by taking the supremum over μ\mu in (A.4), we get the equivalent formulation

which is clearly equivalent to (A.3) by algebraic manipulations.

holds for ϕ(x):=αx\phi(x):=\alpha x. On the one hand

Combining these identities with (A.5) we obtain

and this inequality is satisfied if and only if C≥(1−exp⁡(−κT))−1C\geq(1-\exp(-\kappa T))^{-1}, as claimed.

The sharpness of (3.35) immediately follows from the one of (A.1). Indeed, if there exists α>κ\alpha>\kappa such that

holds for all μ,ν∈P(M)\mu,\nu\in\mathcal{P}(M) satisfying (H3), then by choosing ν=m\nu=\mathfrak{m} and letting t→1t\to 1 we get a contradiction. ∎

The second author gratefully acknowledges support by the European Union through the ERC-AdG “RicciBounds” for Prof. K. T. Sturm.

References