A formula for the time derivative of the entropic cost and applications
Giovanni Conforti, Luca Tamanini
Introduction and statement of the main results
The entropic transportation cost is the optimal value in a probabilistic version of the Monge-Kantorovich optimal transport problem, the Schrödinger problem, whose study has already shown to have far reaching consequences in various fields, ranging from statistical machine learning to functional inequalities. The goal of the present article is to advance in the study of the entropic cost as a function of the time (regularization) parameter (see Definition 1.2 below). Following Mikami’s contribution linking the Schrödinger problem to optimal transport, several results have been obtained in the last years concerning the behavior of the entropic cost in the short-time (small noise) limit . In particular, as a byproduct of the research line originated in another step forward was made with the computation of the first derivative of the rescaled entropic cost at , see also , the recent works and references therein. However, very few results beyond the short-time limit have been obtained. In particular, very little is known about the long-time regime , where the entropic cost is expected to converge to the sum of the marginal entropies. This lack of knowledge was one of the main motivations for our work and in this respect, our contribution includes:
A formula for the first and second derivative of the (rescaled) entropic cost for a general value of in terms of the so-called “energy” (defined at (1.6) below).
A rigorous identification of the large-time limit of the entropic cost as the sum of the marginal entropies.
Sharp exponential convergence rates under a curvature condition in the long-time regime. We obtain this result not only for the classical Schrödinger problem but also for the recently introduced Mean Field Schrödinger problem .
We also establish some results in the short-time limit. Their interest resides in the fact that they allow to obtain a clean expression of the second derivative of the rescaled entropic cost that, to the best of our knowledge, was not known before. However, the regularity assumptions we impose on the marginals are not the weakest ones. For this part, our contribution can be resumed as follows:
A formula for the second derivative of the rescaled cost around that yields the local convexity of the cost in the time variable.
A non-asymptotic sharp quantitative bound for the convergence to the Wasserstein distance depending only on the integral of the Fisher information functional along the corresponding Wasserstein geodesic.
The measure is invariant for the SDE
(here denotes a standard Brownian motion) whose joint law at time and will be denoted by . Within this framework and given (as usual, for a measurable space we denote the set of probability measures over ), the entropic transportation cost is defined as the optimal value in the corresponding Schrödinger problem, namely
where is the set of couplings of and and is the relative entropy functional defined for two probability measures on the same (arbitrary) measurable space as
When is not a probability, the precise definition is postponed to Section 2.
The physical meaning of the variational problem (1.2) is described in the seminal papers , where E. Schrödinger addressed the problem of finding the most likely evolution of a system of independent random particles driven by (1.1) conditionally on the observation of their initial and final configuration. With this picture in mind, the short- and long-time behavior of the entropic cost sounds perfectly natural.
We state here the additional assumptions we will make either on the reference measure or on the marginals . Concerning the former, we will often assume that
Notice that when , (H2) automatically holds and in addition (see for instance [32, Theorem 4.26]). For the latter, we shall suppose that either
where denotes the space of probability measures over with finite second moment, or, more frequently, that
The Benamou-Brenier formulation.
The fluid-dynamical formulation of the entropic cost asserts that, at least under (H4), we have
where is the unique optimal curve and the infimum runs over all weak solutions of the continuity equation
satisfying the marginal constraints and , namely among all with and all Borel vector fields such that:
is weakly continuous and there exists such that for all ;
From a physical point of view, (1.4) states that the trajectory , also called entropic interpolation, is the one minimizing a functional consisting of two terms: the former is purely kinetic, while the latter is given by the Fisher information functional. Hence is given by the balance between deterministic and chaotic behavior.
The energy.
where the role of the function in (1.7) is played by the relative entropy in (1.4). For (1.7) it is well known that along any critical curve the total energy (given by the sum of the kinetic and potential components)
is conserved. The analogous of this simple fact for the problem (1.4) is the conservation of (1.6) along optimal flows.
Short-time behavior of the entropic cost.
The fact that the rescaled entropic cost converges to the squared Wasserstein distance of order two in the short-time limit, i.e.
has generated a surge of interest around the Schrödinger problem (henceforth SP), since SP is more regular and numerically easier to solve than the Monge-Kantorovich problem. There exist nowadays several proofs of (1.8), see for instance for -convergence results. A proof that is valid under the hypotheses (H1) and (H4) can be found in [18, Remark 5.11]. In , (see also ) a further fundamental step was taken that consists in computing the first order term in the expansion of around . Referring to the above mentioned articles for precise statements, we have the following expansionIn the above mentioned references, one often finds instead of in the first order term. This is due to a slightly different choice of reference measure . Since we prefer the reference measure to be reversible, we gain an extra term in the Taylor expansion.
In this work we establish at Theorem 1.6 a non-asymptotic bound for which is sharp in the limit and we compute the second order term in (1.9), thus getting
where denotes the (unique) Wasserstein geodesic between and . Let us remark that (1.10) tells that the rescaled cost is convex around and using functional inequalities such as the HWI inequality one can also estimate from below its second derivative under the Bakry-Émery condition (H1). It is an interesting question to obtain general versions of (1.10) in terms of -convergence. In addition, it is worth noticing that the energy, once properly rescaled, also converges to the squared Wasserstein distance, namely
To see, at least formally, why this is expected to be true, one can integrate (1.6) in time, use (1.4) and the energy conservation to get
Multiplying by , letting and using (1.8) we can then handle the first term on the right-hand side as well as . The fact that the last integral converges to has been proved in [17, Lemma 3.3] and the same argument can be adapted verbatim to the present setting (see ); yet it is a non-trivial fact, as the entropic interpolation converges to the Wasserstein geodesic between and , whose Fisher information needs not to be defined.
Long-time behavior of the entropic cost.
Although its asymptotic regime has been the object of recent studies in connection with the ergodic behavior of Schrödinger bridges and the limiting behavior of Sinkhorn divergence , very few results are available. In this article we prove that under mild assumptions we have
The intuition behind (1.11) is that converges towards and therefore the optimal coupling in SP converges towards the independent coupling of and . We remark that in a different although related context , the convergence of Sinkhorn divergences towards MMD divergences in the limit when the regularization parameter goes to shares many analogies with (1.11).
Our second main result deals with the approximation error in (1.11) and states that it is exponentially small (see Theorem 1.4 for the rigorous statement)
provided the Bakry-Émery condition (H1) is satisfied. The proof of (1.12) is based on Theorem 1.1 asserting that the time derivative of is precisely and on two functional inequalities. The first one is a version of the Talagrand inequality (obtained in and also called entropic Talagrand inequality) and the second one is a functional inequality relating with , that we call “energy-transport” inequality (cf. Lemma 3.8). A similar inequality has been proved very recently in ; there it has been used to obtain the so-called turnpike property for mean field Schrödinger bridges. It is worth noticing that estimates such as (1.12) do not seem to follow from classical functional inequalities such as Talagrand and Log-Sobolev, whereas they can be proven using the new family of functional inequalities involving the entropic cost . The final contribution of the article is to derive a bound similar to (1.12) for the Mean Field Schrödinger problem introduced in . Since we could not establish a generalization of the differentiation formula at Theorem 1.1 to the mean field setup, the proof of this estimate follows a different scheme, but is still based on a class of functional inequalities derived in that are the mean field versions of the entropic Talagrand and of the energy-transport inequalities mentioned above.
Organization of the paper.
The document is structured as follows: in the remainder of Section 1 we state and comment the main results, whose proofs are contained in Section 3. In Section 2 we collect, for reader’s sake, all relevant results and bibliographical references on SP. Finally, in Appendix A we prove the sharpness of a functional inequality introduced by the first-named author in , which plays an important role in this paper.
2 First and second derivative of the entropic cost
The regularity of the entropic cost w.r.t. to the time variable has never been investigated to the best of our knowledge, so that the following is the first result of such a kind and it plays a pivotal role in the study of both the long- and short-time behavior of .
the map is , twice differentiable a.e. and the first derivative is given by
the map is and twice differentiable a.e. The first derivative is given for all by
where is the optimal solution in (1.4). The second derivative writes as
As implicitly stated above, by Theorem 1.1 we see that is continuous on and differentiable a.e.
3 Long-time behavior of entropic cost and energy
Under (H1) with and (H2), for any satisfying (H3) it holds
If satisfy (H4), then it also holds
As concerns the long-time behavior of , we provide two different proofs:
the former is direct and relies on a -convergence argument (see Section 3.1);
the latter is more technical and requires slightly stronger assumptions on and , but as an advantage it allows us to determine the long-time behavior of the so-called “-decomposition” of the optimal coupling in SP (see Section 2 for its definition, Section 3.3 for the proof).
As an application of this result we provide a new proof of the logarithmic Sobolev inequality based on entropic interpolations, in the same spirit of the recent paper , where an “entropic” proof of the HWI inequality is established.
Under (H1) with , for any it holds
where the right-hand side is set equal to if is not locally Sobolev.
As a further step, in the following result we improve Theorem 1.2 by providing sharp rates of convergence for both and . The key message is that, under a positive curvature condition, the approximation error is asymptotically smaller than , up to constant factors depending on and , and the rate is sharp. As already pointed out, recall that if (H1) holds with , then (H2) also holds.
Let us assume that (H1) with and (H4) are satisfied. Then for all it holds
Furthermore, the convergence rate in (1.14) and (1.15) is sharp in the following sense: for all as in (H4) it holds
and there exists a triplet satisfying (H1) with such that
holds for all satisfying (H4) if and only if .
In other words, it may be possible to improve the constant factor in (1.14), but the convergence rate is asymptotically sharp.
Actually, in Section 3.4 we shall prove the following (stronger) bounds
However we prefer the more compact, although slightly less precise, formulation given above.
4 Short-time behavior of the entropic cost
We provide a non-asymptotic bound for the difference under the assumption that the integral of the Fisher information along the displacement interpolation is finite and we also prove that this bound is sharp, as it coincides with the Taylor expansion of around .added in proof: With respect to a previous version of the paper, we have been able to remove quite demanding regularity assumptions. A similar but less general result (valid only in the Euclidean setting) has recently been obtained in by other means, hence our proof is independent. This result is of particular interest as the second order term in the expansion of the rescaled cost has not been computed before (to the best of our knowledge) and its form is general enough to formulate more general version of (1.18) below, for instance in terms of -convergence, that deserve to be the object of future work.
Assume that (H1) and (H4) hold. Then for all we have:
If in addition the Bakry-Émery condition
Note, for instance, that the condition is satisfied when is the volume measure.
5 Long time behavior of the mean field entropic cost
In the above we denoted by the usual convolution operator. The interaction potential satisfies the assumptions
In the mean field entropic cost is defined as the optimal value in (MFSP). However, if this definition does not give back the usual entropic cost and the reason is simply the following: in the large deviations formulation of (1.2) particles are sampled with initial distribution the invariant measure , whereas in MFSP particles are sampled according to . For this reason, and to strengthen the analogy with (1.2), we prefer to define the mean field entropic cost as
With this definition we recover defined in (1.2) when .
After this premise, the following plays the same role of assumption (H1)
while regarding the marginal constraints and , we assume that they have finite second moment, they belong to the domain of and have the same barycenter, i.e.
The proof we gave of Theorem 1.1 is hard to replicate for MFSP essentially because uniqueness of optimizers is currently not known. However, the long-time behavior of the mean field entropic cost can still be studied and exponential rate of convergence can be derived as well.
Assume (H’1)-(H’3) and that . Then:
the rate of convergence is at least , i.e. there exists a decreasing function such that
One could be more precise in the above statement and get that
where . For sake of clarity, we prefer the more compact, although slightly less precise, formulation given above.
Preliminaries
In this section we collect some useful results concerning Markov semigroups and Schrödinger problem, either already present in the literature or extended to our framework.
With this remark in mind, let us present all the useful information about . First of all, it enjoys the following standard a priori estimate
which can be obtained by differentiating . The semigroup is also ergodic. This means that:
if , then for all it holds
if , then in for all .
The curvature assumption (H1) then yields many important consequences, the first of which is the Bakry-Émery commutation estimate
For its proof as well as for all the regularizing properties of that will be used throughout the paper, we address once more the reader to . Secondly, enjoys an -Lipschitz regularization (see ), namely for all it holds
Then let us recall that under (H1) Hamilton’s gradient estimate is satisfied (see ): for any positive function for some it holds
pointwise, where . Within our framework it is also well known (see for instance ) that there exists a unique kernel (transition probability) representing in the following sense:
The function can also be seen as the density of (the joint law at time 0 and of the solution to (1.1)) w.r.t. ; it is a smooth function on and the second-named author recently proved in that in great generality (and in particular under the assumption (H1)) the following upper Gaussian estimate for the kernel holds
for all , and , where can be chosen equal to 0 if in (H1). As concerns Gaussian lower bounds, if (H2) holds then by [35, Corollary 1.3] the following is satisfied
The relative entropy functional.
Let us first recall the definition of the relative entropy functional in the case of a reference measure with possibly infinite mass (see for more details). Given a -finite measure on , there exists a measurable function such that
The (f,g)𝑓𝑔(f,g)-decomposition.
It is important to stress that solving (1.2) is equivalent to finding non-negative Borel functions , also called decomposition, such that
This pair of equations is known as Schrödinger system and its solvability holds under very mild assumptions (see ), in particular under:
(H4), as a consequence of [18, Proposition 2.1], the smoothness and positivity of the heat kernel and the boundedness of ;
(H1), (H2) and (H3), because of [25, Proposition 2.5] and (2.8).
Still under (H1)-(H4), the couple solving (2.9) is unique up to the trivial transformation with , as proven for instance in [18, Proposition 2.1]: this fact will play an important role in several proofs (e.g. the extension of (2.20) to our setting and Lemma 3.4). A good feature of is that they inherit the regularity (smoothness and integrability) of , the densities of respectively. More precisely,
since so are ;
A proof of this property can be found in [17, Proposition 2.7].
About point (a) a more quantitative statement is actually possible. In [18, Proposition 2.1] the second-named author proved that for all as in (H4), the following integral bounds hold for the decomposition
where is a suitable positive constant. If we normalize in such a way that
and from now on this choice will always be done, then the bounds above become
Aim of the next lemma is to improve this result by showing that the same kind of bounds holds when ranges in a compact subset of or even closed half-line if we also assume (H2).
Given (H1) and as in (H4), the following hold:
for all there exists such that
if (H2) is satisfied, then for all there exists such that
The first equation in the Schrödinger system (2.9) and the representation formula (2.6) entail
As is smooth in and and are uniformly compactly supported (since they are supported in and respectively) we deduce that there exists such that in for all , whence
provided , and this proves the first inequality in (2.11). For the second one it is sufficient to swap the roles of and . If we further assume (H2), then the Gaussian lower bound (2.8) holds. Hence for all there exists such that in for all , whence by the same argument as above the first inequality in (2.12) follows. By interchanging the roles of and , the same conclusion follows for . ∎
The Benamou-Brenier formulation.
The equivalence between (2.9) and (1.2) is extremely fruitful, as it enables to fully describe the optimal pair in the fluid-dynamical description (1.4) of the entropic cost. Indeed,
Actually, (1.4) and the identities above are nothing but a reparametrization of the Benamou-Brenier-like formula for the entropic cost established in [19, Theorem 4.2] and , which reads as
In (2.13) the infimum runs over all weak solutions of the continuity equation (1.5) satisfying the marginal constraints and , while defined above is the unique optimal density-velocity couple for (2.13). It is also useful to see the second term in the right-hand side of (2.13) as an action functional, whose arguments are the curves and , i.e.
When , is a purely kinetic energy and by the well-known Benamou-Brenier formulation of the optimal transport problem ,
where the infimum is taken over the same set as in (2.13). In this case optimal couples shall be denoted by . Let us also recall that
Finally, the expression of the conserved quantity in terms of the rescaled optimal couple reads as
Dual formulation.
For future reference, it is also worth mentioning the fact that, for all satisfying (H4), the (rescaled) entropic cost admits the following dual representation (see )
where .
Geometric information.
As regards the relationship between Schrödinger problem and lower Ricci bounds, let us recall that in the first-named author showed that (H1) implies the following distorted -convexity inequality of the entropy along entropic interpolations:
for all , where . Truth to be told, the result in was stated in the framework of compact Riemannian manifolds satisfying (H1) with and endowed with the volume measure, but the proof can be adapted to our setting when (H2) holds, by relying on [17, Lemma 3.7]. Indeed, if we set
where is a normalization constant so that is still a probability density, then the identities
hold both in the classical sense and as strong -limits, where is the generator of . Together with (2.4) and (2.5), this is sufficient to follow the lines of [9, Lemma 3.6, Lemma 3.7] and [17, Lemma 3.7] and deduce the following
Under (H1), given as in (H4) and with the same notations as in (2.21), for all and define
Then , ,
By [9, Lemma 4.1] and following the proof of Theorem 1.4 therein, we obtain
and again by dominated convergence the right-hand side above converges to
Proof of the main results
A rather straightforward proof of the long-time behavior of can be obtained by a -convergence argument. Its essence is contained in the following
As a first step, we claim that for any lower bounded lower semicontinuous function on it holds
Since , the claim is trivial for constant functions; hence, without loss of generality we can assume that and, under this further assumption, the Gaussian lower bound (2.8) together with Fatou’s lemma yields (3.1). By Portmanteau theorem this is equivalent to say that
After this premise, let , and assume that . The lower semicontinuity of the relative entropy w.r.t. to both its arguments together with (3.2) gives that
thus the -liminf inequality. To complete the proof, let and find a sequence such that
To this aim it is not restrictive to assume that and it is also easy to see that satisfies (3.3). Indeed, the Gaussian lower bound (2.8) and the fact that (H1) holds with imply that
for some constant independent of . Next, observe that by definition of and we have
Let us point out that in the previous lemma and thus in Theorem 1.2, (H2) is required only when . As regards the long-time behavior of , we need to determine in which way is controlled in terms of .
whence trivially the desired upper bound for . On the other hand, Young’s inequality and (2.16) yield
so that plugging this inequality into the previous identity gives also the lower bound for . ∎
We are now in the position to prove Theorem 1.2.
It is easily seen that the unique optimal coupling in
Dividing by (3.6) and letting , the long-time behavior of is established as well. ∎
2 Proof of Theorem 1.1 and Theorem 1.6
The proof of Theorem 1.1 requires the preparatory Lemmas 3.3, 3.4 and 3.5. In the first one we show that the Fisher information of the entropic interpolation is non-increasing as a function of .
Under the assumptions of Theorem 1.1, let be optimal for the formulation (2.13) and for (2.15). Then the function
Fix . Summing the inequalities
and dividing by we obtain
In the second lemma we prove the continuity in of the functions and given by (2.9) with respect to the norm, for any .
Under the assumptions of Theorem 1.1, the functions and are continuous from to , for any .
As a byproduct, the functions and are continuous from to , , as well.
Let , and denote by the positive constant provided by Lemma 2.1-(i) on the interval . From the first bound in (2.11) we immediately deduce that is bounded in , hence in for all , because all the functions are supported in and this has finite mass, because bounded. To prove the same property for requires a more technical argument.
From (2.7) with we first oberve that
and since , for all , have bounded support and is a continuous function,
with independent of . Thus if we combine this inequality with the previous one and recall the chosen normalization (2.10) we obtain for all for some (depending on and too), i.e.
Plugging this inequality into the second bound in (2.11) yields
Hence also is bounded in for all , because all the functions are supported in and this has finite mass.
where both limits have to be understood in for .
and observe that the second term on the right-hand side trivially vanishes as . As regards the first one,
Since, as already remarked, is bounded in and all the functions are supported in , there exists sufficiently large such that for all , whence
by the maximum principle. As the right-hand side above belongs to and does not depend on , by (3.10) and the dominated convergence theorem we thus infer that
whence, combining this fact with the previous steps, the validity of the first limit in (3.9). The argument for the second limit is completely analogous.
As a consequence, up to extract a further (not relabeled) subsequence, we have that the limits in (3.9) hold -a.e. Therefore, if we look at the Schrödinger system (2.9) at time , which reads as
which is possible because , we deduce that
For , by (2.9) we observe that the cost can be rewritten as
and by Lemma 3.4, the locally uniform -bounds provided by Lemma 2.1 and by dominated convergence it is easy to see that the right-hand side above continuously depends on .
As regards , let us start observing that the second identity in (2.18) can be equivalently rewritten as
so fix and note that, by integration by parts first and Cauchy-Schwarz inequality then, we obtain
The second summand after the last inequality clearly vanishes as , because of Lemma 3.4. As concerns the first one, argue as in the proof of Lemma 3.4 and notice that for sufficiently small, say with , all the functions are uniformly bounded in and supported in , hence also uniformly bounded in . By the a priori estimate (2.1), this is sufficient to conclude that there exists sufficiently large such that for all . Hence, by Lemma 3.4, also the first summand after the last inequality converges to 0 in the limit, whence the continuity of .
and by the previous discussion the right-hand side is continuous on . ∎
Consider the map , let , write
and observe that the first term on the right-hand side is non-positive by optimality of for . For the second one
where denotes the partial derivative w.r.t. of . Combining these two facts we deduce
and remark that now the second term on the right-hand side is non-negative by optimality of for . As regards the first one,
which together with (3.13) implies that is everywhere right differentiable on . Left differentiability follows by an analogous argument. Indeed if , then the first term on the right-hand side in (3.11) is non-negative and (3.12) still holds true with instead of , whence
Applying the same considerations to (3.14) and (3.15) gives
and combining this inequality with the previous one entails the desired left differentiability. Therefore is everywhere differentiable on , a fortiori continuous therein, and
Therefore is everywhere differentiable (hence continuous) on as well and, by (2.13) and the very definition of and ,
which proves both formulas for the first derivative of , thanks to (2.14). Now notice that the right-hand side above is continuous on : is continuous on since so is , while is continuous by Lemma 3.5. Hence belongs to . The fact that is also twice differentiable a.e. follows from the fact that
and the last term on the right-hand side is the product of a linear function and a monotone (non-increasing) one, thanks to Lemma 3.3.
This concludes the proof of (ii). Claim (i) is a straightforward consequence as well as the formula for the second derivative of . ∎
From Theorem 1.1, the proof of Theorem 1.6 follows.
As regards (1.17), we only prove the upper bound, the lower bound being an immediate consequence of the Benamou-Brenier formulation (2.13). Since belongs to we can use the fundamental theorem of calculus that gives for all
where the inequality is motivated by Lemma 3.3. The conclusion follows letting and using the convergence of the rescaled entropic cost towards .
The coefficients can be identified thanks to Theorem 1.1. Indeed
thanks to (a). Therefore it only remains to show (a), (b), and (c); we shall start with the first one. On the one hand, by Lemma 3.3
On the other hand, as holds, we can and shall rely on powerful facts established in : by Propositions 5.1 and 5.4 therein we know that
and by [18, Lemma 4.2] we also know that for all there exist such that
With this said, let be any Borel set and let us prove that
To this end, since are compactly supported, there exists a bounded set containing the supports of for all , so that up to choose a larger we can assume that for all . By (3.18) and the trivial fact that this implies that as for all . By (3.19) and the fact that under the condition the function in the right-hand side of (3.19) is integrable and converges monotonically to 0 as , it also follows
whence (3.20). By [2, Theorem 7.6] this implies
Combining this inequality with (3.17) yields the desired continuity at and thus (a).
On the other hand, by following the argument of Theorem 1.1 and recalling (2.15) we see that
where is any choice of gradient of intermediate Kantorovich potentials associated with the geodesic , and by (a) this implies
Combining this inequality with (3.21) yields the differentiability of at and by (3.16) we see that the derivative is continuous there. Finally, by (2.13), the computations above for determining the coefficients and , and the properties (a) and (b) it is not difficult to see that (c) holds too. ∎
3 Proof of Theorem 1.2 and Corollary 1.3
In this section we provide the reader with a second proof of Theorem 1.2, whose crucial ingredient is a long-time analogue of Lemma 3.4. This is precisely the content of the next result. With respect to the proof given in Section 3.1, here the marginals satisfy (H4), hence a condition stronger than (H3); as a consequence we gain information on the long-time behavior of and separately, determining their limits as . Let us point out again that (H2) below is required only when .
Under the assumptions (H1) with , (H2) and (H4), for the functions given by (2.9) it holds
where both limits are in for any .
where both limits are in for any .
Let and denote by the positive constant provided by Lemma 2.1-(ii) on the interval . From the first bound in (2.12) we immediately deduce that is bounded in , hence in for all , as already argued in the proof of Lemma 3.4. From (2.7) with (since ), (2.10), the fact that have bounded support and arguing as in Lemma 3.4 we also have
Plugging this inequality into the second bound in (2.12) yields that also is bounded in , thus in for all . As a direct consequence of this and of the maximum, we deduce that also and are bounded in .
where both limits have to be understood in for . To this aim observe that
where the second inequality is motivated by the fact that and is a contraction in . Therefore, arguing as in the proof of Lemma 3.4 we can see that the first term on the right-hand side above vanishes as . On the other hand, the ergodicity of and (H2) entail that
hence in for all , while for it is sufficient to observe that
and use (3.27). Therefore, also the second term on the second line on the right-hand side in (3.26) converges to 0 as and this proves the first identity in (3.25). An analogous argument leads to the second one.
As a consequence, up to extract a further (not relabeled) subsequence, the limits in (3.25) hold -a.e. and if we pass to the limit as in the Schrödinger system (2.9) we get
It is now sufficient to combine this information with (3.25) to get (3.23). ∎
We are now in the position to determine the long-time behavior of the entropic cost.
Let us first observe that for all there exists such that for all it holds
Indeed, both upper bounds have been proven in Lemma 3.6, while the lower ones can be deduced relying on (2.8) and the boundedness of , . More precisely, for (3.28a) there exists such that
the last identity being motivated by (2.10). For the same reason and taking (3.24) into account
which proves also (3.28b). With this said, by the very definition of the entropic cost and its equivalence with (2.9) we have
and by (3.28a), (3.28b), (3.23) and the dominated convergence theorem we can pass to the limit as in the identity above, thus getting
As concerns Corollary 1.3, let us recall an approximation result, whose proof can be found in [17, Lemma 3.1].
Let with and
Then there exists with , such that and
We can now provide an “entropic” proof of the logarithmic Sobolev inequality.
First of all, we can assume that the right-hand side in (1.13) is finite, otherwise the statement is trivially true; by Lemma 3.7 we can also assume that has compact support, is absolutely continuous w.r.t. and its density is smooth. Hence, if we take any probability measure with compact support and smooth density, (H4) is satisfied. In addition, and are also compactly supported and smooth, as they inherit the regularity of and respectively.
After this premise, let us show that is differentiable at and
To this aim, let us recall that from [17, Lemma 3.7] for every the map belongs to and it holds
where , and are defined as in (2.21). If we integrate (3.30) on with we obtain
and by the dominated convergence theorem it is easy to see that the left-hand side converges to as . As regards the right-hand one, to prove that
we borrow an argument used in [17, Lemma 3.11] and we report it here for reader’s sake. By the very definition of and since
the desired conclusion is achieved if we are able to prove that
To this aim, notice that , whence using either or it is easy to infer that
All the right-hand sides above are integrable on (either by (2.17) or by the Bakry-Émery contraction estimate (2.3) together with ), thus by dominated convergence (3.32a) follows. An analogous argument holds for (3.32b). Therefore we can pass to the limit as in (3.31) and get
whence (3.29). This allows us to differentiate (2.20) at and obtain
from the very definition of , so that
where is the generator of , self-adjoint w.r.t. . By (3.28b), (3.23) and the boundedness of we infer that the first integral on the right-hand side vanishes as by dominated convergence, whence (3.34).
From (3.33), (3.34) and Theorem 1.2 the logarithmic Sobolev inequality (1.13) follows. ∎
4 Proof of Theorem 1.4
The proof of Theorem 1.4 relies on Theorem 1.1 and two other ingredients: the entropic Talagrand inequality put forward in and an “energy-transport” inequality relating and . The former states that if (H1) holds with , then for all as in (H3) and for all
Truth to be told, as for (2.20) also the entropic Talagrand inequality (3.35) was stated in in the framework of compact Riemannian manifolds satisfying (H1) with and endowed with the volume measure, but since (3.35) is deduced from (2.20) and the latter has been generalized to the present framework in Section 2, there is no problem in applying it.
The “energy-transport” inequality relating and is expressed in the following
Assume that (H1) and (H4) hold. Then for all it holds
From the second identity in (2.18) we have
so that by Cauchy-Schwarz inequality and by (2.16), (3.37) follows.
As regards (3.36), using the same notations as in (2.21), let us first point out that, being and optimal for SP with marginals and , we have
for , then by (3.39) evaluated in and by Cauchy-Schwarz inequality
In order to estimate the right-hand side above, observe that by (2.23)
with , defined as in (2.22), so that by Lemma 2.2
while on the other hand, by (3.41a) and a completely analogous argument
Combining these inequalities with (3.40), we obtain
and it is now sufficient to pass to the limit as to get the conclusion. The fact that and is motivated by dominated convergence; the same is true for taking (3.39) into account. ∎
Let us point out that the bound (3.37) is better than the one obtained from (3.36) via a Taylor expansion around , so that (3.36) is not sharp and the reader might wonder whether (3.36) could be improved by replacing with . A posteriori, this is not possible because if (3.36) held with instead of for some , then this would contradict the sharpness of (1.14) as stated in Theorem 1.4.
We are now in the position to prove Theorem 1.4.
From Theorem 1.1 and Theorem 1.2 we have that
and by Lemma 3.8 the right-hand side above is controlled as follows
By the entropic Talagrand inequality (3.35) for and by (H1) with , which implies , it holds
for all and an analogous inequality holds with and swapped. Therefore the right-hand side in (3.43) is bounded from above by
The bound (1.16a) follows by calculating explicitly the integral term and this, in turn, trivially implies (1.14).
For the sharpness of (1.14), we shall prove that there exists a triplet satisfying (H1) with and satisfying (H4) such that
and the stochastic process associated to the SDE (1.1) is the stationary Ornstein-Uhlenbeck process, whence by well-known results (see e.g. ) its transition probabilities w.r.t. admit the following explicit representation
and it is now sufficient to remark that the right-hand side is asymptotically positive, so that for large enough we are allowed to apply the logarithm to both sides of the inequality above, whence (3.44).
As regards (1.15), it is sufficient to estimate the right-hand side in (3.36) by following the same reasoning as above. To prove that also (1.15) is sharp, it is easy to see that if there exists such that
holds for all satisfying (H4), then for any such there exist sufficiently large such that
and if we plug this inequality into (3.42), instead of Lemma 3.8, then we get , whence
and this clearly contradicts the sharpness of (1.14). ∎
5 Proof of Theorem 1.7
for all , see [3, Eq. 92] and also .
After this premise, the first ingredient needed for the proof of Theorem 1.7 is the following
Using the time-reversal relation (3.49) we deduce that
Using the change of variable and the very definition of time reversal, we rewrite the right-hand side as
so that the conclusion is easily obtained. ∎
As for the proof of Theorem 1.4, also in this case we rely on an entropic Talagrand inequality. In the mean field setting (and taking into account Remark 1.9) this reads as follows
for all and satisfying (H’3), see [3, Corollary 1.3].
We are now in position to prove Theorem 1.7.
Appendix A On the sharpness of the entropic Talagrand inequality
Under (H1) with it holds , so that satisfies (H3) and it is then licit to choose in (3.35), which gives rise to the following version of the entropic Talagrand inequality, closer to the classical Talagrand inequality known in optimal transport:
for all as in (H3). In neither the sharpness of (A.1) nor the one of (3.35) was investigated. Aim of this appendix is to remedy this lack.
Assume that (H1) holds with . Then the entropic Talagrand inequality (A.1) is sharp, since there exists a triplet satisfying (H1) with such that
As a byproduct, also the entropic Talagrand inequality (3.35) is sharp in the following sense: there exists a triplet satisfying (H1) with such that
As a first step, we shall prove that (A.2) is equivalent to
where . To this aim, notice that by (2.19) and the symmetry of the entropic cost, i.e. , (A.2) is equivalent to
for all satisfying (H3) and for all . Then observe that
where , the identity being a byproduct of the well-known variational representation of the entropy (see ). Hence by taking the supremum over in (A.4), we get the equivalent formulation
which is clearly equivalent to (A.3) by algebraic manipulations.
holds for . On the one hand
Combining these identities with (A.5) we obtain
and this inequality is satisfied if and only if , as claimed.
The sharpness of (3.35) immediately follows from the one of (A.1). Indeed, if there exists such that
holds for all satisfying (H3), then by choosing and letting we get a contradiction. ∎
The second author gratefully acknowledges support by the European Union through the ERC-AdG “RicciBounds” for Prof. K. T. Sturm.