Stability of Entropic Optimal Transport and Schrödinger Bridges

Promit Ghosal, Marcel Nutz, Espen Bernton

Introduction

Computational progress has lead to manifold applications of optimal transport in high-dimensional problems ranging from machine learning and statistics to image and language processing (e.g., ). In this context, entropic regularization is crucial to enable efficient large-scale computation via Sinkhorn’s algorithm, hence has become the focus of dozens of recent studies. We refer to for a survey with numerous references.

Our main contribution is the stability of solutions to the entropically regularized optimal transport problem with respect to the marginals and the cost function. Parallel to the fundamental stability theorem in classical optimal transport, it justifies, for example, that approximations found by solving discretized problems indeed converge to the true solution when the cost function is continuous. Our results are stated in terms of cyclical invariance, a geometric notion inspired by the cc-cyclical monotonicity property in classical optimal transport. When the entropic transport problem has finite value, a coupling is cyclically invariant if and only if it is an optimal transport. Our stability theorem entails a general wellposedness result beyond the realm of optimization: cyclical invariance singles out a unique coupling even if the transport problem has infinite value—i.e., all couplings have infinite cost—and therefore the paradigm of cost minimization does not differentiate couplings from one another.

where Π(μ,ν)\Pi(\mu,\nu) is the set of couplings and H(⋅∣P)H(\cdot|P) denotes relative entropy (or Kullback–Leibler divergence) with respect to the product PP of the marginals, defined as H(π∣P):=∫log⁡(dπdP) dπH(\pi|P):=\int\log(\frac{d\pi}{dP})\,d\pi for π≪P\pi\ll P and H(π∣P):=∞H(\pi|P):=\infty otherwise. If the minimization (1.1) is finite; i.e., if

then it admits a unique minimizer π∈Π(μ,ν)\pi\in\Pi(\mu,\nu) and moreover π∼P\pi\sim P.

A coupling π∈Π(μ,ν)\pi\in\Pi(\mu,\nu) is called (c,ε)(c,\varepsilon)-cyclically invariant if π∼P\pi\sim P and its density admits a version dπ/dP:X×Y→(0,∞)d\pi/dP:\mathsf{X}\times\mathsf{Y}\to(0,\infty) such that

By way of a factorization property that is equivalent to cyclical invariance, known results imply the following relation to the optimization (1.1).

Let (1.1) be finite. Then π∈Π(μ,ν)\pi\in\Pi(\mu,\nu) is the minimizer of (1.1) if and only if π\pi is (c,ε)(c,\varepsilon)-cyclically invariant.

See Section 5 for details and references. We are mainly interested in optimal transport problems on Euclidean spaces X,Y\mathsf{X},\mathsf{Y}. However, the only particular property of such spaces that plays a role for our analysis is that Lebesgue’s theorem on the differentiation of measures holds. Thus, we postulate that property (see Assumption 2.3) and otherwise allow for a general Polish setting. We can now state the aforementioned wellposedness result.

Let c:X×Y→[0,∞)c:\mathsf{X}\times\mathsf{Y}\to[0,\infty) be continuous, ε>0\varepsilon>0 and (μ,ν)∈P(X)×P(Y)(\mu,\nu)\in\mathcal{P}(\mathsf{X})\times\mathcal{P}(\mathsf{Y}). There exists a unique (c,ε)(c,\varepsilon)-cyclically invariant coupling π∈Π(μ,ν)\pi\in\Pi(\mu,\nu). If (1.1) is finite, π\pi is its unique minimizer.

Uniqueness follows from known facts and does not require the continuity of cc. On the other hand, existence beyond the framework of finite cost is a completely novel result. Rather than using convex analysis or variational arguments, it is based on the subsequent stability theorem for cyclical invariance. One example where wellposedness with infinite cost is of interest, is the statistical notion of rank recently proposed in . Multivariate ranks have been defined in nonparametric statistics through Brenier’s optimal transport map to extend the usual scalar notions and tests; see . Leveraging the same idea but computationally less expensive, entropic optimal transport is used in to define “differentiable ranks.” Theorem 1.3 allows one to naturally define such ranks for arbitrary distributions—like in the scalar case—without imposing a second moment condition.

For n≥1n\geq 1, let (μn,νn)∈P(X)×P(Y)(\mu_{n},\nu_{n})\in\mathcal{P}(\mathsf{X})\times\mathcal{P}(\mathsf{Y}), let εn>0\varepsilon_{n}>0 and let cn:X×Y→[0,∞)c_{n}:\mathsf{X}\times\mathsf{Y}\to[0,\infty) be measurable. Let πn∈Π(μn,νn)\pi_{n}\in\Pi(\mu_{n},\nu_{n}) be (cn,εn)(c_{n},\varepsilon_{n})-cyclically invariant. Suppose that μn,νn\mu_{n},\nu_{n} converge weakly to some limits μ,ν\mu,\nu, that εn→ε>0\varepsilon_{n}\to\varepsilon>0 and that cnc_{n} converges uniformly on bounded sets to a continuous function cc. Then πn\pi_{n} converges weakly to a limit π∈Π(μ,ν)\pi\in\Pi(\mu,\nu) and π\pi is (c,ε)(c,\varepsilon)-cyclically invariant.

If the involved optimization problems are finite, the theorem states the stability of the (entropic) optimal transport couplings. A simple yet important application is when the marginals μ,ν\mu,\nu are approximated by discrete measures, as it would be in a computational implementation. Even for this particular case, we are not aware of similar results in the literature. We mention that continuity for the limiting cost cc in Theorem 1.4 is a sharp condition: Proposition 5.3 will demonstrate that any nontrivial discontinuity in cc leads to a failure of Theorem 1.4, for suitable marginals.

A noteworthy application of the stability theorem was presented in the follow-up work where it was observed that convergence of Sinkhorn’s algorithm can be seen as a stability problem. In that context, the marginals μn,νn\mu_{n},\nu_{n} are produced by the algorithm and known to converge to μ,ν\mu,\nu in great generality, while the iterates of the algorithm correspond to πn\pi_{n}. Theorem 1.4 then implies the weak convergence to the correct optimizer π\pi. Its geometric approach completely avoids the difficulty of establishing the integrability properties of μn,νn\mu_{n},\nu_{n} or even the finiteness of the associated entropic optimal transport problems.

The general existence result of Theorem 1.3 is a consequence of Theorem 1.4 applied with cn=cc_{n}=c, εn=ε\varepsilon_{n}=\varepsilon and approximations (μn,νn)→(μ,ν)(\mu_{n},\nu_{n})\to(\mu,\nu) where μn,νn\mu_{n},\nu_{n} are discrete measures with finite support. To solve the problem with marginals (μn,νn)(\mu_{n},\nu_{n}), one could use Proposition 1.2, but this particular case is a finite-dimensional minimization problem that can also be solved by standard calculus arguments. In particular, Theorem 1.4 yields an approach to construct cyclically invariant couplings which is novel even when the optimization problem is finite. This approach does not use the (classical but non-trivial) arguments of convex analysis and density factorization behind Proposition 1.2 (see , among others). It is also quite different from the iterative method of which uses another finiteness condition; see for a modern presentation, analysis and extension of that method. Instead, our approach is close in spirit to the construction of cc-cyclically monotone couplings that is standard in classical optimal transport; cf. [52, pp. 64–65].

The companion paper illustrates the use of cyclical invariance for the limit ε→0\varepsilon\to 0. In this degenerate asymptotic, the limiting object is classical optimal transport as characterized by cc-cyclical monotonicity. The latter property describes the shape of the support of the coupling and, therefore, is readily amenable to weak convergence arguments via Portmanteau’s theorem. The same fact is often exploited in classical optimal transport theory, for instance in the standard proof of its stability theorem [52, p. 77]. In the present study, the limit is entropic optimal transport. Being a property of the density, the relation of cyclical invariance with weak convergence is less direct (especially as the measures in Theorem 1.4 may well be mutually singular). Our general principle is to blow the points (xi,yi)(x_{i},y_{i}) in Definition 1.1 up to small balls, pass to the weak limit, and then recover information about the limiting density by shrinking the balls, via differentiation of measures. This technique appears to be novel in this area.

Starting with , a number of works examine the degenerate asymptotic ε→0\varepsilon\to 0 where the limiting problem is classical optimal transport. The fact that weak limits of entropic optimizers are optimal transports was established by using Gamma-convergence arguments in a more general context of Schrödinger bridges; see also for the case of optimal transport with quadratic cost. As mentioned above, extends this result to transport problems with infinite value by way of cyclical invariance and cc-cyclical monotonicity; moreover, a large deviations principle quantifies the local rate of convergence. Related results can be found in where the expansion of the optimal cost as a function of ε\varepsilon is studied. We remark that Gamma convergence seems difficult to use in the context of Theorem 1.4 due to the reference measures changing along the sequence. The limit ε→0\varepsilon\to 0 can also be analyzed in the associated dual problem, here the solutions are called potentials. Convergence of potentials was shown in for quadratic costs and compactly supported marginals, and recently in for a general Polish setting. Closer to the problem occurring in computational practice as well as the present question of stability, studies the convergence of the discrete Sinkhorn algorithm to an optimal transport in the joint limit when εn→0\varepsilon_{n}\to 0 and the marginals μ,ν\mu,\nu are approximated by discretizations μn,νn\mu_{n},\nu_{n} satisfying a certain density property. Explicit error bounds are derived, for instance for quadratic cost on the torus, to establish near-linear complexity of the resulting algorithm. For more on the computational challenges and remedies in this regime, see for instance and the references therein.

While we are not aware of general stability results for the nondegenerate limit εn→ε>0\varepsilon_{n}\to\varepsilon>0 in the literature, the sampling complexity of entropic optimal transport (with fixed ε\varepsilon) can be seen as a particular form of stability with respect to the marginals. Indeed, study how the the empirical entropic Wasserstein distance, obtained by optimally coupling i.i.d. samples from the marginals, converges to the population version. The results are based on global arguments exploiting the regularity of the Schrödinger potentials, which, in turn, is achieved by imposing compactness and decay conditions on the marginals. See also which studies a related asymptotic regime for a different regularization of optimal transport. Related to the present work at least in spirit, there are several areas where analogues of cc-cyclical monotonicity have recently lead to breakthroughs, including martingale optimal transport , optimal Skorokhod embeddings and weak transport .

The remainder of this paper is organized as follows. Section 2 details the setting and main results in the language of Schrödinger bridges. The first step towards the stability theorem is reported Section 3 where we establish that weak limits of cyclically invariant couplings remain absolutely continuous. This is based on comparing measures of rectangles, an analysis that may be of independent interest. Section 4 continuous the main proof by showing that limits of cyclically invariant couplings are again cyclically invariant. It comprises of two steps; the aforementioned principle of blowing up points and passing to the limit first yields a weakened version of the invariance property, and then measure-theoretic arguments can be used to show that the (proper) invariance property already follows. The concluding Section 5 collects the arguments to prove the main results and their ramifications, including that continuity of the cost function is necessary for stability.

Main Results

Let (X,d)(\mathsf{X},d) be a complete, separable metric space; we write P(X)\mathcal{P}(\mathsf{X}) for the space of probability measures on the Borel σ\sigma-field B(X)\mathcal{B}(\mathsf{X}) endowed with weak convergence (induced by bounded continuous functions). The same is assumed for the second marginal space (Y,d)(\mathsf{Y},d), and we equip X×Y\mathsf{X}\times\mathsf{Y} with the metric d((x,y),(x′,y′))=max⁡{d(x,x′),d(y,y′)}d((x,y),(x^{\prime},y^{\prime}))=\max\{d(x,x^{\prime}),d(y,y^{\prime})\}. Throughout this section, two measures (μ,ν)∈P(X)×P(Y)(\mu,\nu)\in\mathcal{P}(\mathsf{X})\times\mathcal{P}(\mathsf{Y}) play the role of given marginals for the static Schrödinger bridge problem

where R∈P(X×Y)R\in\mathcal{P}(\mathsf{X}\times\mathsf{Y}) is a given reference measure. We refer to for extensive surveys on Schrödinger bridges. The entropic optimal transport problem (1.1) can be recovered (up to constants) as a special case for RR defined by

where a=(∫e−c/ε dP)−1a=(\int e^{-c/\varepsilon}\,dP)^{-1} is the normalizing constant. In particular, R∼PR\sim P, which will also be an important condition in many of our results for (2.1). By way of (2.2), the following generalizes Definition 1.1.

Let π∈Π(μ,ν)\pi\in\Pi(\mu,\nu) and R∈P(X×Y)R\in\mathcal{P}(\mathsf{X}\times\mathsf{Y}). We call (π,R)(\pi,R) cyclically invariant if π∼R∼P\pi\sim R\sim P and there exists a version dπ/dR:X×Y→(0,∞)d\pi/dR:\mathsf{X}\times\mathsf{Y}\to(0,\infty) of the relative density satisfying

The analogue of the finiteness condition (1.2) is that

We summarize the pertinent facts; see Lemma 5.1 for detailed references.

Let R∼PR\sim P and let (2.1) be finite. There exists a unique minimizer π∈Π(μ,ν)\pi\in\Pi(\mu,\nu) for (2.1), it satisfies π∼R\pi\sim R, and it is the unique coupling π\pi such that (π,R)(\pi,R) is cyclically invariant.

In the remainder of the paper we assume that the underlying spaces allow for differentiation of measures in the following sense.

Given ρ,λ∈P(X)\rho,\lambda\in\mathcal{P}(\mathsf{X}) satisfying ρ≪λ\rho\ll\lambda, there exists X′⊂X\mathsf{X}^{\prime}\subset\mathsf{X} of full λ\lambda-measure such that

defines a version of the Radon–Nikodym density dρ/dλd\rho/d\lambda. The analogous property is assumed on the space X×Y\mathsf{X}\times\mathsf{Y}.

For Euclidean spaces X,Y\mathsf{X},\mathsf{Y}, the assumption holds by the standard differentiation theorem [21, Theorem 1.32, p. 53]. More generally, it holds in the context of so-called Vitali covering relations; the classical reference is [23, Theorem 2.9.8, p. 156]. For example, Assumption 2.3 holds when X\mathsf{X} and X×Y\mathsf{X}\times\mathsf{Y} are compact subsets of Riemannian manifolds (due to the “directionally limited” property established in [23, Section 2.8.9, pp. 145-146]), or more generally, countable unions of such sets. See also [31, pp. 4-8, esp. Example 1.15] for an accessible introduction. For our purposes, the main restriction is that differentiation of measures generally fails on infinite-dimensional spaces; see for a counterexample and for a related result on coverings. An alternative to Assumption 2.3, making our results slightly more general, is to impose a doubling condition on the specific marginal measures (μ,ν)(\mu,\nu); cf. Remark 5.2.

We have seen in the Introduction that continuity of cc is essential for the stability of entropic optimal transport. In view of (2.2), it is then clear that the regularity of dR/dPdR/dP is pivotal. The following generalizations of Theorems 1.3 and 1.4 are our main results.

Suppose that R∼P:=μ⊗νR\sim P:=\mu\otimes\nu and that the density dR/dPdR/dP admits a continuous version. Then there exists a unique coupling π∈Π(μ,ν)\pi\in\Pi(\mu,\nu) such that (π,R)(\pi,R) is cyclically invariant. If (2.1) is finite, π\pi is its minimizer.

For n≥1n\geq 1, consider (μn,νn)∈P(X)×P(Y)(\mu_{n},\nu_{n})\in\mathcal{P}(\mathsf{X})\times\mathcal{P}(\mathsf{Y}), Rn∈P(X×Y)R_{n}\in\mathcal{P}(\mathsf{X}\times\mathsf{Y}) and πn∈Π(μn,νn)\pi_{n}\in\Pi(\mu_{n},\nu_{n}). Let (πn,Rn)(\pi_{n},R_{n}) be cyclically invariant and suppose that μn,νn,Rn\mu_{n},\nu_{n},R_{n} converge weakly to some limits μ,ν,R\mu,\nu,R, where R∼P:=μ⊗νR\sim P:=\mu\otimes\nu. Writing Pn:=μn⊗νnP_{n}:=\mu_{n}\otimes\nu_{n}, suppose also that for some versions dRndPn,dRdP:X×Y→(0,∞)\frac{dR_{n}}{dP_{n}},\frac{dR}{dP}:\mathsf{X}\times\mathsf{Y}\to(0,\infty) of the respective densities and some constants αn>0\alpha_{n}>0, it holds that for any fixed z∈spt⁡Rz\in\operatorname{spt}R,

where o(1)o(1) stands for a function ϕz(z′,n)→0\phi_{z}(z^{\prime},n)\to 0 as d(z′,z)+1/n→0d(z^{\prime},z)+1/n\to 0. Then πn\pi_{n} converges weakly to a limit π∼R\pi\sim R and (π,R)(\pi,R) is cyclically invariant.

Schrödinger bridges are closely related to so-called Schrödinger systems (also called Schrödinger equations). For instance, Theorem 2.4 entails the following wellposedness result.

Let (μ,ν)∈P(X)×P(Y)(\mu,\nu)\in\mathcal{P}(\mathsf{X})\times\mathcal{P}(\mathsf{Y}) and let f:X×Y→(0,∞)f:\mathsf{X}\times\mathsf{Y}\to(0,\infty) be continuous with ∫X×Yf dP=1\int_{\mathsf{X}\times\mathsf{Y}}f\,dP=1. There exist Borel functions φ:X→(0,∞)\varphi:\mathsf{X}\to(0,\infty) and ψ:Y→(0,∞)\psi:\mathsf{Y}\to(0,\infty) such that

for μ\mu-a.e. x∈Xx\in\mathsf{X} and ν\nu-a.e. y∈Yy\in\mathsf{Y}. The pair (φ,ψ)(\varphi,\psi) is a.s. unique up to a multiplicative constant.I.e., any solution (φ′,ψ′)(\varphi^{\prime},\psi^{\prime}) satisfies φ′=aφ\varphi^{\prime}=a\varphi μ\mu-a.s., ψ′=a−1ψ\psi^{\prime}=a^{-1}\psi ν\nu-a.s., for some a>0a>0.

As above, the uniqueness follows from known results. Existence for continuous functions ff such that f,f−1f,f^{-1} are uniformly bounded was first shown in . Under the boundedness condition alone, existence is due to . Using the connection with Schrödinger bridges, relaxed the boundedness to a condition of finite entropy, corresponding to the finiteness of (2.1) in our setting. We refer to for a more complete review of the literature which dates back to Schrödinger. In Corollary 2.6, we reintroduce the continuity condition of but avoid any condition of finite entropy or boundedness. Of course, the stability result of Theorem 2.5 also has an analogous corollary for Schrödinger systems.

Absolute Continuity of Limits

In this section we show that if (πn,Rn)(\pi_{n},R_{n}) is cyclically invariant and Rn∼μn⊗νnR_{n}\sim\mu_{n}\otimes\nu_{n} holds in a locally uniform sense (to be made precise), then any weak limit pair (π,R)=(lim⁡nπn,lim⁡nRn)(\pi,R)=(\lim_{n}\pi_{n},\lim_{n}R_{n}) must satisfy π≪R\pi\ll R.

As the method of proof is novel, we first sketch the line of argument. We shall be comparing the measures of rectangles Fi×Gi⊂X×YF_{i}\times G_{i}\subset\mathsf{X}\times\mathsf{Y} for i=1,2i=1,2 with their permutations Fi×Gi+1F_{i}\times G_{i+1}.We use the cyclical convention for i∈{1,2}i\in\{1,2\}; that is, i+1:=1i+1:=1 for i=2i=2. Consider first the trivial case Rn=Pn:=μn⊗νnR_{n}=P_{n}:=\mu_{n}\otimes\nu_{n}, then cyclical invariance of (πn,Rn)(\pi_{n},R_{n}) implies

This will of course no longer hold if Rn≠PnR_{n}\neq P_{n}, but the equivalence Rn∼PnR_{n}\sim P_{n} suggests that the two sides of (3.1) should still be comparable. We will quantify the equivalence Rn∼PnR_{n}\sim P_{n} and assume it to hold uniformly in nn. Then, we prove that the two sides of (3.1) are comparable in the sense that their quotient remains bounded uniformly in nn. For suitable rectangles, the bound propagates to the weak limit (π,R)(\pi,R), accomplishing the first step of the proof. The second step is to argue by contraposition that this bound implies π≪R\pi\ll R. Indeed, we establish that if π∈Π(μ,ν)\pi\in\Pi(\mu,\nu) is singular wrt. μ⊗ν\mu\otimes\nu, then there exist Fi,GiF_{i},G_{i} such that π(F1×G1)π(F2×G2)\pi(F_{1}\times G_{1})\pi(F_{2}\times G_{2}) is above a threshold whereas π(F1×G2)π(F2×G1)\pi(F_{1}\times G_{2})\pi(F_{2}\times G_{1}) is arbitrarily small.

For ease of reference, we first record two measure-theoretic facts.

Let ρ\rho be a σ\sigma-finite measure on a Polish space (Ω,d)(\Omega,d). We say that C⊂ΩC\subset\Omega is ρ\rho-continuous if its boundary ∂C\partial C is a ρ\rho-nullset.

(i) ρ\rho-continuous sets form a field; i.e., unions, intersections, complements, differences of ρ\rho-continuous sets are again ρ\rho-continuous.

(ii) For fixed z∈Ωz\in\Omega, the open ball Br(z)={z′:d(z,z′)<r}B_{r}(z)=\{z^{\prime}:d(z,z^{\prime})<r\} is ρ\rho-continuous for all but countably many values of r>0r>0. In particular, given r>0r>0, there exists 0<r′≤r0<r^{\prime}\leq r such that Br′(z)B_{r^{\prime}}(z) is ρ\rho-continuous.

(iii) If F⊂XF\subset\mathsf{X} is μ\mu-continuous and G⊂YG\subset\mathsf{Y} is ν\nu-continuous, then F×GF\times G is π\pi-continuous for any π∈Π(μ,ν)\pi\in\Pi(\mu,\nu).

(iv) Given A∈B(Ω)A\in\mathcal{B}(\Omega) and ε>0\varepsilon>0, there exists an open set B⊂Aε:={d(⋅,A)<ε}B\subset A_{\varepsilon}:=\{d(\cdot,A)<\varepsilon\} with ρ(∂B)=0\rho(\partial B)=0 and ρ(AΔB)<ε\rho(A\Delta B)<\varepsilon.

Statement (i) is verified directly; (ii) holds because a σ\sigma-finite measure admits at most countably many disjoint sets of positive measure; (iii) follows from ∂(F×G)=(F‾×∂G)∪(∂F×G‾)\partial(F\times G)=(\overline{F}\times\partial G)\cup(\partial F\times\overline{G}). Let A,εA,\varepsilon be as in (iv). By interior and exterior regularity of ρ\rho, there is a compact set K⊂AK\subset A with ρ(A∖K)<ε\rho(A\setminus K)<\varepsilon and an open set A⊂O⊂AεA\subset O\subset A_{\varepsilon} with ρ(O∖A)<ε\rho(O\setminus A)<\varepsilon. We have r(z):=d(z,Oc)>0r(z):=d(z,O^{c})>0 for all z∈Kz\in K by the closedness of OcO^{c}, and KK is covered by the open balls Br(z)(z)B_{r(z)}(z) with z∈Kz\in K. After making r(z)r(z) smaller if necessary, each of these balls is ρ\rho-continuous. Choosing a finite cover {Br(zi)(zi)}i≤N\{B_{r(z_{i})}(z_{i})\}_{i\leq N}, the set B:=∪iBr(zi)(zi)B:=\cup_{i}B_{r(z_{i})}(z_{i}) has the required properties. ∎

The second fact is a conditional version of the differentiation of measures, based on Assumption 2.3 for the marginal space X\mathsf{X}. While not widely known, this concept was already established in , although the author defined differentiation of measures in a slightly different way. For the convenience of the reader, we detail the adaptation to our setting.

Let π∈P(X×Y)\pi\in\mathcal{P}(\mathsf{X}\times\mathsf{Y}) and let μ\mu be its first marginal. Consider for x∈spt⁡μx\in\operatorname{spt}\mu and r>0r>0 the probability measure πx(r)∈P(Y)\pi^{(r)}_{x}\in\mathcal{P}(\mathsf{Y}) defined by

Under Assumption 2.3 on X\mathsf{X}, there exists X0⊂spt⁡μ\mathsf{X}_{0}\subset\operatorname{spt}\mu with μ(X0)=1\mu(\mathsf{X}_{0})=1 such that for all x∈X0x\in\mathsf{X}_{0}, the weak limit

exists. Moreover, πx\pi_{x} defines a regular conditional probability of π\pi given xx.

As π^x\hat{\pi}_{x} is a regular conditional probability, it holds for μ\mu-a.e. x∈Xx\in\mathsf{X} that

We now apply Assumption 2.3 to the pair f dμ≪dμf\,d\mu\ll d\mu and deduce that the right-hand side converges to f(x)f(x) for μ\mu-a.e. x∈Xx\in\mathsf{X}, which is (3.2). ∎

The next result is the main ingredient for the second step as sketched above: the sets {xi}×Ui\{x_{i}\}\times U_{i} constructed in Lemma 3.3 will be “blown up” to rectangles in the proof of Proposition 3.5 and used to show by contraposition that π≪R\pi\ll R.

Let π∈Π(μ,ν)\pi\in\Pi(\mu,\nu) and let π=μ(dx)⊗πx(dy)\pi=\mu(dx)\otimes\pi_{x}(dy) be a disintegration. If π≪̸μ⊗ν\pi\not\ll\mu\otimes\nu, then

In addition, the sets UiU_{i} can be chosen to be disjoint, of arbitrarily small diameter, and such that πxi(∂Uj)=ν(∂Uj)=0\pi_{x_{i}}(\partial U_{j})=\nu(\partial U_{j})=0 for i,j∈{1,2}i,j\in\{1,2\}.

Let π≪̸μ⊗ν\pi\not\ll\mu\otimes\nu; that is, there exists a set A∈B(X×Y)A\in\mathcal{B}(\mathsf{X}\times\mathsf{Y}) with π(A)>0\pi(A)>0 and (μ⊗ν)(A)=0(\mu\otimes\nu)(A)=0. Let Ax={y: (x,y)∈A}A_{x}=\{y:\,(x,y)\in A\} denote the xx-section, then ν(Ax)=0\nu(A_{x})=0 for μ\mu-a.e. x∈Xx\in\mathsf{X}. On the other hand, any B∈B(Y)B\in\mathcal{B}(\mathsf{Y}) satisfies ν(B)=∫Xπx′(B) μ(dx′)\nu(B)=\int_{\mathsf{X}}\pi_{x^{\prime}}(B)\,\mu(dx^{\prime}). In particular, ν(Ax)=0\nu(A_{x})=0 implies that πx′(Ax)=0\pi_{x^{\prime}}(A_{x})=0 for μ\mu-a.e. x′∈Xx^{\prime}\in\mathsf{X}. Therefore,

As π(A)>0\pi(A)>0, the set W={x: πx(Ax)>0}W=\{x:\,\pi_{x}(A_{x})>0\} satisfies μ(W)>0\mu(W)>0 and hence μ2((W×W)∩F)>0\mu^{2}((W\times W)\cap F)>0. For (x1,x2)∈(W×W)∩F(x_{1},x_{2})\in(W\times W)\cap F we have πxi(Axi)>0\pi_{x_{i}}(A_{x_{i}})>0 and πxi(Axi+1)=0\pi_{x_{i}}(A_{x_{i+1}})=0. In particular, the disjoint sets Ui′:=Axi∖Axi+1U^{\prime}_{i}:=A_{x_{i}}\setminus A_{x_{i+1}} satisfy πxi(Ui′)>0\pi_{x_{i}}(U^{\prime}_{i})>0 and πxi(Ui+1′)=0\pi_{x_{i}}(U^{\prime}_{i+1})=0. By intersecting with a suitable ball, the diameter of Ui′U^{\prime}_{i} can be assumed to be arbitrarily small. Finally, let ρ=πx1+πx2+ν\rho=\pi_{x_{1}}+\pi_{x_{2}}+\nu and choose ρ\rho-continuous sets Ui′′U^{\prime\prime}_{i} for Ui′U^{\prime}_{i} as in Lemma 3.1 (iv), with ε>0\varepsilon>0 small enough such that the sets Ui:=Ui′′∖Ui+1′′U_{i}:=U^{\prime\prime}_{i}\setminus U^{\prime\prime}_{i+1} have the required properties; cf. Lemma 3.1 (i). ∎

The next lemma establishes that the two sides of (3.1) are comparable with a bound related to the equivalence Rn∼PnR_{n}\sim P_{n}. The following notation is useful: when k≥1k\geq 1 and a kk-tuple (z1,…,zk)∈(X×Y)k(z_{1},\dots,z_{k})\in(\mathsf{X}\times\mathsf{Y})^{k} are given, and

with the cyclical convention yk+1:=y1y_{k+1}:=y_{1}.

Let π∈Π(μ,ν)\pi\in\Pi(\mu,\nu) and π≪R∼P:=μ⊗ν\pi\ll R\sim P:=\mu\otimes\nu. Consider rectangles Ai∈B(X)×B(Y)A_{i}\in\mathcal{B}(\mathsf{X})\times\mathcal{B}(\mathsf{Y}) for 1≤i≤k1\leq i\leq k and denote Aˉ1×⋯×Aˉk:={(zˉ1,…,zˉk):zi∈Ai}\bar{A}_{1}\times\dots\times\bar{A}_{k}:=\{(\bar{z}_{1},\dots,\bar{z}_{k}):z_{i}\in A_{i}\}. For some α,αˉ>0\alpha,\bar{\alpha}>0, suppose that dR/dP≤αdR/dP\leq\alpha on AiA_{i} and (dR/dP)−1≤αˉ(dR/dP)^{-1}\leq\bar{\alpha} on Aˉi\bar{A}_{i}, for all ii. Then

Using the cyclical invariance of (π,R)(\pi,R) and the rectangular form of AiA_{i},

We can now prove the main result of this section.

Let (μn,νn)∈P(X)×P(Y)(\mu_{n},\nu_{n})\in\mathcal{P}(\mathsf{X})\times\mathcal{P}(\mathsf{Y}), let πn∈Π(μn,νn)\pi_{n}\in\Pi(\mu_{n},\nu_{n}) and Rn∼Pn:=μn⊗νnR_{n}\sim P_{n}:=\mu_{n}\otimes\nu_{n}. Suppose that (πn,Rn)(\pi_{n},R_{n}) is cyclically invariant for each nn and that πn\pi_{n} converges weakly to some limit π\pi. In particular, μn,νn\mu_{n},\nu_{n} converge to some limits μ,ν\mu,\nu, and π∈Π(μ,ν)\pi\in\Pi(\mu,\nu). Set P=μ⊗νP=\mu\otimes\nu and suppose that (Rn,Pn)(R_{n},P_{n}) are uniformly locally equivalent in the following sense: there are versions dRndPn:X×Y→(0,∞)\frac{dR_{n}}{dP_{n}}:\mathsf{X}\times\mathsf{Y}\to(0,\infty) of the relative densities and constants αn>0\alpha_{n}>0 such that given z∈spt⁡Pz\in\operatorname{spt}P, there exist r=r(z)>0r=r(z)>0 and n0=n0(z,r)n_{0}=n_{0}(z,r) with

Note that the weak convergence of πn∈Π(μn,νn)\pi_{n}\in\Pi(\mu_{n},\nu_{n}) implies the weak convergence of its marginals to some limits μ\mu and ν\nu, and then π∈Π(μ,ν)\pi\in\Pi(\mu,\nu). By Lemma 3.2, there is a set X0⊂spt⁡μ\mathsf{X}_{0}\subset\operatorname{spt}\mu of full μ\mu-measure such that for x∈X0x\in\mathsf{X}_{0}, the weak limit

exists, and π=μ⊗xπx\pi=\mu\otimes_{x}\pi_{x} is a disintegration of π\pi. In particular, for any x∈X0x\in\mathsf{X}_{0} and any πx\pi_{x}-continuity set UU,

Suppose for contradiction that π≪̸μ⊗ν\pi\not\ll\mu\otimes\nu. Then Lemma 3.3 yields xi∈X0x_{i}\in\mathsf{X}_{0} (i=1,2i=1,2) and disjoint sets UiU_{i} such that

Moreover, given r0>0r_{0}>0, the sets UiU_{i} can be chosen to be (π1+π2+ν)(\pi_{1}+\pi_{2}+\nu)-continuous and contained in a ball Br0(yi)B_{r_{0}}(y_{i}) around some yi∈spt⁡νy_{i}\in\operatorname{spt}\nu. Using also (3.5), we can choose r0,n0r_{0},n_{0} such that for some α∗>0\alpha_{*}>0,

for all 1≤i≤k1\leq i\leq k and n≥n0n\geq n_{0}. Writing Bi,r:=Br(xi)B_{i,r}:=B_{r}(x_{i}) for brevity, (3.6) implies in particular that

That is, given ε>0\varepsilon>0, choosing r≤r0r\leq r_{0} small enough results in

and we may further choose rr such that μ(∂Bi,r)=0\mu(\partial B_{i,r})=0. Specifically, we choose ε>0\varepsilon>0 such that ε/(q−ε)<α∗−2\varepsilon/(q-\varepsilon)<\alpha_{*}^{-2}, then (3.8) implies

Note that Bi,r×UjB_{i,r}\times U_{j} is a π\pi-continuity set for i,j∈{1,2}i,j\in\{1,2\}; cf. Lemma 3.1 (iii). In view of πn→π\pi_{n}\to\pi, it then follows that

for nn sufficiently large. On the other hand, we apply Lemma 3.4 with k=2k=2 and Ai=Bi,r×UiA_{i}=B_{i,r}\times U_{i}. In view of (3.7), the condition of the lemma holds with α:=αn−1α∗\alpha:=\alpha_{n}^{-1}\alpha_{*} and αˉ:=αnα∗\bar{\alpha}:=\alpha_{n}\alpha_{*}. Noting that the αn\alpha_{n} cancel to yield (ααˉ)k=α∗4(\alpha\bar{\alpha})^{k}=\alpha_{*}^{4}, the lemma yields the inequality opposite to (3.9), a contradiction. ∎

Cyclical Invariance of Limits

In this section we aim to show that limits of cyclically invariant couplings are again cyclically invariant, under suitable conditions. We proceed in two steps. First, we establish that limits are weakly cyclically invariant as defined below. Second, we show that weak cyclical invariance already implies cyclical invariance. This second step has little to do with the passage to the limit; rather, it settles some measure-theoretic aspects to get rid of pesky nullsets.

The “weak” notion is introduced mainly to disentangle the proof of the main result. It weakens in two ways the cyclical invariance of (π,R)(\pi,R) as stated in Definition 2.1: the equivalence of π\pi and RR is reduced to absolute continuity and the cyclical relation only holds for points from specific sets. We recall the notation zˉi\bar{z}_{i} from (3.3).

Let π∈Π(μ,ν)\pi\in\Pi(\mu,\nu) and R∈P(X×Y)R\in\mathcal{P}(\mathsf{X}\times\mathsf{Y}). We call (π,R)(\pi,R) weakly cyclically invariant if π≪R∼P:=μ⊗ν\pi\ll R\sim P:=\mu\otimes\nu and there exist

a version Z:X×Y→[0,∞]Z:\mathsf{X}\times\mathsf{Y}\to[0,\infty] of the density dπ/dRd\pi/dR,

Ω1,Ω0∈B(X×Y)\Omega_{1},\Omega_{0}\in\mathcal{B}(\mathsf{X}\times\mathsf{Y}) with π(Ω1)=R(Ω0)=1\pi(\Omega_{1})=R(\Omega_{0})=1 and 0<Z<∞0<Z<\infty on Ω1\Omega_{1}

Let (μn,νn)∈P(X)×P(Y)(\mu_{n},\nu_{n})\in\mathcal{P}(\mathsf{X})\times\mathcal{P}(\mathsf{Y}), let πn∈Π(μn,νn)\pi_{n}\in\Pi(\mu_{n},\nu_{n}) and Rn∼Pn:=μn⊗νnR_{n}\sim P_{n}:=\mu_{n}\otimes\nu_{n}. Suppose that (πn,Rn)(\pi_{n},R_{n}) is cyclically invariant for each nn and that πn,Rn\pi_{n},R_{n} converge weakly to some limits π,R\pi,R. In particular, μn,νn\mu_{n},\nu_{n} converge to some limits μ,ν\mu,\nu, and π∈Π(μ,ν)\pi\in\Pi(\mu,\nu). Suppose that R∼P:=μ⊗νR\sim P:=\mu\otimes\nu and that there are versions fn,f:Ω→(0,∞)f_{n},f:\Omega\to(0,\infty) of the densities dRndPn,dRdP\frac{dR_{n}}{dP_{n}},\frac{dR}{dP} and constants αn>0\alpha_{n}>0 such that for any fixed z∈spt⁡Rz\in\operatorname{spt}R,

where o(1)o(1) stands for a function ϕz(z′,n)→0\phi_{z}(z^{\prime},n)\to 0 as d(z′,z)+1/n→0d(z^{\prime},z)+1/n\to 0. Then (π,R)(\pi,R) is weakly cyclically invariant. Specifically, the quantities Ω1,Ω0,Z\Omega_{1},\Omega_{0},Z of Definition 4.1 can be chosen as

Assumption 2.3 for the space X×Y\mathsf{X}\times\mathsf{Y} shows that ZZ is a version of dπ/dRd\pi/dR and R(Ω0)=1R(\Omega_{0})=1. As (4.1) implies (3.5), Proposition 3.5 yields that π≪P\pi\ll P, hence π≪R\pi\ll R and the definition of Ω1\Omega_{1} implies π(Ω1)=1\pi(\Omega_{1})=1.

Let z1,…,zk∈Ω1z_{1},\dots,z_{k}\in\Omega_{1} be such that zˉi∈Ω0\bar{z}_{i}\in\Omega_{0} and consider for r>0r>0 the balls Ai=Ai(r):=Br(zi)=Br(xi)×Br(yi)A_{i}=A_{i}^{(r)}:=B_{r}(z_{i})=B_{r}(x_{i})\times B_{r}(y_{i}). To avoid unwieldy formulas, we use the vector notation

together with the convention that functions and measures are evaluated by multiplication over the components, for instance

As dπn/dPn=(dπn/dRn)(dRn/dPn)d\pi_{n}/dP_{n}=(d\pi_{n}/dR_{n})(dR_{n}/dP_{n}), we then have

We also have πn(Ai)>0\pi_{n}(A_{i})>0 for nn large as Z(zi)>0Z(z_{i})>0 and πn→π\pi_{n}\to\pi. On the other hand, Zn(z′)Z_{n}(\bm{z}^{\prime}) and Pn(dz′)P_{n}(d\bm{z}^{\prime}) are invariant under z′↦zˉ′\bm{z}^{\prime}\mapsto\bm{\bar{z}}^{\prime} due to the assumption on πn\pi_{n} and the form of PnP_{n}; cf. (3.4). Thus

Applying the assumption on fnf_{n} to each of the points zi,zˉiz_{i},\bar{z}_{i} then yields

where o(1)o(1) stands for a function of (z,zˉ)(\bm{z},\bm{\bar{z}}) converging to zero as r+1/n→0r+1/n\to 0. For values of r>0r>0 such that Ai,AˉiA_{i},\bar{A}_{i} are continuity sets of PP (and hence also of π\pi and RR), taking n→∞n\to\infty yields

On the other hand, zi,zˉi∈Ω0z_{i},\bar{z}_{i}\in\Omega_{0} also guarantees that R(Ai)P(Ai)=[1+o(1)]f(zi)\frac{R(A_{i})}{P(A_{i})}=[1+o(1)]f({z}_{i}) and similarly for zˉi\bar{z}_{i}. Recalling P(Aˉ)=P(A)P(\bm{\bar{A}})=P(\bm{A}), we deduce

and then letting r→0r\to 0 (along a sequence of rr such that Ai,AˉiA_{i},\bar{A}_{i} are continuity sets of PP) shows Z(z)=Z(zˉ)Z(\bm{z})=Z(\bm{\bar{z}}), as desired. ∎

Assumption (4.1) on fn,ff_{n},f in Proposition 4.2 is a sufficient condition for

it can be replaced by any other condition implying (4.4). We note that (4.4) can be seen as a differentiation of measures intertwined with a weak limit.

As mentioned above, the second step is to upgrade the weak cyclical invariance. Some of these considerations are similar to arguments in the proofs of , where it is shown by variational arguments that minimizers of certain static Schrödinger bridge problems admit a factorization.

Let π∈Π(μ,ν)\pi\in\Pi(\mu,\nu) and let R∈P(X×Y)R\in\mathcal{P}(\mathsf{X}\times\mathsf{Y}) satisfy R∼P:=μ⊗νR\sim P:=\mu\otimes\nu. The following are equivalent:

(π,R)(\pi,R) is weakly cyclically invariant,

there exist Borel functions φ:X→(0,∞)\varphi:\mathsf{X}\to(0,\infty) and ψ:Y→(0,∞)\psi:\mathsf{Y}\to(0,\infty) such that (x,y)↦φ(x)ψ(y)(x,y)\mapsto\varphi(x)\psi(y) is a version of the density dπ/dRd\pi/dR.

The implications (i)⇒(ii)(i)\Rightarrow(ii) and (iii)⇒(i)(iii)\Rightarrow(i) are immediate. The fact that (i)⇒(iii)(i)\Rightarrow(iii) is also well known, cf. , but will not be used directly. For the proof of (ii)⇒(iii)(ii)\Rightarrow(iii), a measure-theoretic fact will be useful. Given a set A⊂X×YA\subset\mathsf{X}\times\mathsf{Y} with P(A)=1P(A)=1, Lemma 4.5 below states that within any set BB of positive measure we can find a point (x∗,y∗)(x_{*},y_{*}) which acts like an airline hub for AA: any two points of AA are connected through (x∗,y∗)(x_{*},y_{*}), modulo marginal nullsets. In particular, any trip can be achieved with at most one stopover, and we may stop within BB. (The bound of one is optimal as the set AA need not be a rectangle; in fact, AA may fail to contain any measurable rectangle of positive measure [22, Exercise 5.4, p. 74].) Lemma 4.5 is refinement of [6, Lemma 4.3] which asserts the connectedness of AA in the sense of —in our analogy, connectedness means that any trip between two points of AA can be achieved with finitely many stopovers at some points in AA.

Let P=μ⊗νP=\mu\otimes\nu and let A,B⊂X×YA,B\subset\mathsf{X}\times\mathsf{Y} be Borel sets with P(B)>0P(B)>0 and P(A)=1P(A)=1. There are Borel sets X0⊂X\mathsf{X}_{0}\subset\mathsf{X} and Y0⊂Y\mathsf{Y}_{0}\subset\mathsf{Y} with μ(X0)=ν(Y0)=1\mu(\mathsf{X}_{0})=\nu(\mathsf{Y}_{0})=1 such that setting A0:=A∩(X0×Y0)A_{0}:=A\cap(\mathsf{X}_{0}\times\mathsf{Y}_{0}) and B0:=B∩(X0×Y0)B_{0}:=B\cap(\mathsf{X}_{0}\times\mathsf{Y}_{0}), there exists a point

(x∗,y∗)∈B0(x_{*},y_{*})\in B_{0} such that (x,y∗),(x∗,y)∈A0(x,y_{*}),(x_{*},y)\in A_{0} for any (x,y)∈X0×Y0(x,y)\in\mathsf{X}_{0}\times\mathsf{Y}_{0}.

Let Cx={y:(x,y)∈C}C_{x}=\{y:(x,y)\in C\} denote the section of a set C⊂X×YC\subset\mathsf{X}\times\mathsf{Y} at x∈Xx\in\mathsf{X}, and analogously for y∈Yy\in\mathsf{Y}. Let X1={x∈X:ν(Ax)=1}\mathsf{X}_{1}=\{x\in\mathsf{X}:\nu(A_{x})=1\}. In view of Fubini’s theorem, P(A)=1P(A)=1 implies μ(X1)=1\mu(\mathsf{X}_{1})=1. Similarly, P(B)>0P(B)>0 implies that {x∈X:ν(Bx)>0}\{x\in\mathsf{X}:\nu(B_{x})>0\} has positive μ\mu-measure. In particular, there exists a point x∗∈X1x_{*}\in\mathsf{X}_{1} with ν(Bx∗)>0\nu(B_{x_{*}})>0.

Next, let Y0={y∈Y:μ(Ay)=1}∩Ax∗\mathsf{Y}_{0}=\{y\in\mathsf{Y}:\mu(A_{y})=1\}\cap A_{x_{*}}. Then again ν(Y0)=1\nu(\mathsf{Y}_{0})=1, and in particular there exists a point y∗∈Y0∩Bx∗y_{*}\in\mathsf{Y}_{0}\cap B_{x_{*}}. Moreover, the set X0:=X1∩Ay∗\mathsf{X}_{0}:=\mathsf{X}_{1}\cap A_{y_{*}} satisfies μ(X0)=1\mu(\mathsf{X}_{0})=1. By passing to Borel subsets of full measure, we may assume that X0,Y0\mathsf{X}_{0},\mathsf{Y}_{0} are themselves Borel.

Writing A0=A∩(X0×Y0)A_{0}=A\cap(\mathsf{X}_{0}\times\mathsf{Y}_{0}) and B0=B∩(X0×Y0)B_{0}=B\cap(\mathsf{X}_{0}\times\mathsf{Y}_{0}), we have by construction that (x∗,y∗)∈B0(x_{*},y_{*})\in B_{0} satisfies (x∗,y)∈A0(x_{*},y)\in A_{0} for all y∈Y0y\in\mathsf{Y}_{0} and (x,y∗)∈A0(x,y_{*})\in A_{0} for all x∈X0x\in\mathsf{X}_{0}. ∎

Let Z,Ω1,Ω0Z,\Omega_{1},\Omega_{0} be as in Definition 4.1. Our aim is to find Borel functions φ:X→(0,∞)\varphi:\mathsf{X}\to(0,\infty) and ψ:Y→(0,∞)\psi:\mathsf{Y}\to(0,\infty) such that Z′(x,y):=φ(x)ψ(y)Z^{\prime}(x,y):=\varphi(x)\psi(y) defines a version of the density dπ/dRd\pi/dR. It is sufficient to construct φ\varphi on a Borel set X0⊂X\mathsf{X}_{0}\subset\mathsf{X} of full marginal measure, as we may then extend φ\varphi by setting φ=1\varphi=1 on X∖X0\mathsf{X}\setminus\mathsf{X}_{0}, and similarly for ψ\psi. In particular, we may assume that proj⁡XΩ1=X\operatorname{proj}_{\mathsf{X}}\Omega_{1}=\mathsf{X} and proj⁡YΩ1=Y\operatorname{proj}_{\mathsf{Y}}\Omega_{1}=\mathsf{Y}.

Noting that P(Ω1)>0P(\Omega_{1})>0 due to π≪P\pi\ll P, we can apply Lemma 4.5 to Ω1\Omega_{1} and Ω0\Omega_{0}, and in view of the above observation, we may assume that X0=X\mathsf{X}_{0}=\mathsf{X} and Y0=Y\mathsf{Y}_{0}=\mathsf{Y} in its assertion. We then obtain a point

(x∗,y∗)∈Ω1(x_{*},y_{*})\in\Omega_{1} such that (x,y∗),(x∗,y)∈Ω0(x,y_{*}),(x_{*},y)\in\Omega_{0} for any (x,y)∈X×Y(x,y)\in\mathsf{X}\times\mathsf{Y}.

Define φ(x∗):=a>0\varphi(x_{*}):=a>0 as an arbitrary number and ψ(y∗):=Z(x∗,y∗)/φ(x∗)\psi(y_{*}):=Z(x_{*},y_{*})/\varphi(x_{*}). Given (x,y)∈X×Y(x,y)\in\mathsf{X}\times\mathsf{Y}, we have (x∗,y)∈Ω0(x_{*},y)\in\Omega_{0} and (x,y∗)∈Ω0(x,y_{*})\in\Omega_{0}, allowing us to define

The fact that ZZ is Borel readily implies that φ,ψ\varphi,\psi are Borel. Define Z′(x,y):=φ(x)ψ(x)Z^{\prime}(x,y):=\varphi(x)\psi(x) for (x,y)∈X×Y(x,y)\in\mathsf{X}\times\mathsf{Y}. Clearly Z′Z^{\prime} is Borel and takes values in (0,∞)(0,\infty). Let (x,y)∈Ω1(x,y)\in\Omega_{1}. Then (x,y∗),(x∗,y)∈Ω0(x,y_{*}),(x_{*},y)\in\Omega_{0} and weak cyclical invariance yields

That is, Z′=ZZ^{\prime}=Z on Ω1\Omega_{1}. To prove the same relation on Ω0\Omega_{0}, consider (x,y)∈Ω0(x,y)\in\Omega_{0}. There are x′∈Xx^{\prime}\in\mathsf{X} and y′∈Yy^{\prime}\in\mathsf{Y} such that z1:=(x,y′)z_{1}:=(x,y^{\prime}) and z2:=(x′,y)z_{2}:=(x^{\prime},y) are in Ω1\Omega_{1}, and of course we also have z3:=(x∗,y∗)∈Ω1z_{3}:=(x_{*},y_{*})\in\Omega_{1}. Thus the already established fact that Z′=ZZ^{\prime}=Z on Ω1\Omega_{1} implies

On the other hand, zˉ1=(x,y)\bar{z}_{1}=(x,y) and zˉ2=(x′,y∗)\bar{z}_{2}=(x^{\prime},y_{*}) and zˉ3=(x∗,y′)\bar{z}_{3}=(x_{*},y^{\prime}) are all in Ω0\Omega_{0}, so that the invariance yields

As a result, Z′(x,y)=φ(x)ψ(y)=Z(x,y)Z^{\prime}(x,y)=\varphi(x)\psi(y)=Z(x,y), showing Z′=ZZ^{\prime}=Z on Ω0\Omega_{0}. Recalling R(Ω0)R(\Omega_{0})=1, it follows that Z′Z^{\prime} is again a version of dπ/dRd\pi/dR. ∎

Proof of Main Results and Ramifications

For ease of reference, we first summarize some known results.

Let (μ,ν)∈P(X)×P(Y)(\mu,\nu)\in\mathcal{P}(\mathsf{X})\times\mathcal{P}(\mathsf{Y}), R∈P(X×Y)R\in\mathcal{P}(\mathsf{X}\times\mathsf{Y}) and R∼P:=μ⊗νR\sim P:=\mu\otimes\nu.

If the static Schrödinger bridge problem (2.1) is finite, it admits a unique minimizer π∈Π(μ,ν)\pi\in\Pi(\mu,\nu). Moreover, (π,R)(\pi,R) is cyclically invariant.

Let π∈Π(μ,ν)\pi\in\Pi(\mu,\nu). If (π,R)(\pi,R) is cyclically invariant and (2.1) is finite, then π\pi is its minimizer.

There exists at most one π∈Π(μ,ν)\pi\in\Pi(\mu,\nu) such that (π,R)(\pi,R) is cyclically invariant.

Recall that cyclical invariance of (π,R)(\pi,R) is equivalent a factorization of the density dπ/dRd\pi/dR into strictly positive functions; cf. Proposition 4.4 or . Taking that into account, (i) and (ii) can be found in [40, Theorem 2.1] in the stated generality. (The original results are due to , among others.) Finally, (iii) follows from (ii), as was also noted in [40, Corollary 2.9]: if π,π′∈Π(μ,ν)\pi,\pi^{\prime}\in\Pi(\mu,\nu) have positive densities dπ/dRd\pi/dR, dπ′/dRd\pi^{\prime}/dR admitting factorizations, then dπ/dπ′d\pi/d\pi^{\prime} also admits a factorization and now (ii), applied with π′\pi^{\prime} as reference measure, implies that π\pi is the unique minimizer of H(⋅∣π′)H(\cdot|\pi^{\prime}). As π′∈Π(μ,ν)\pi^{\prime}\in\Pi(\mu,\nu) is itself a coupling, this minimizer is π′\pi^{\prime}. ∎

Given data as in Theorem 2.5, the sequences (μn)(\mu_{n}) and (νn)(\nu_{n}) are tight, which readily implies the tightness of (πn)(\pi_{n}); cf. [52, Lemma 4.4, p. 44]. In view of Propositions 4.2 and 4.4, any cluster point π\pi is such that (π,R)(\pi,R) is cyclically invariant. The uniqueness of cyclically invariant couplings, see Lemma 5.1 (iii), shows that all cluster points coincide and hence that the original sequence (πn)(\pi_{n}) converges. This proves Theorem 2.5.

To deduce Theorem 1.4 from Theorem 2.5, we choose the reference measure as in (2.2); i.e.,

where an,aa_{n},a are the normalizing constants. Combining the uniform convergence cn/εn→c/εc_{n}/\varepsilon_{n}\to c/\varepsilon on bounded sets with the continuity of cc, we see that (2.4) holds, for instance with αn=an/a\alpha_{n}=a_{n}/a.

The uniqueness part of Theorem 2.4, as well as its last assertion, are stated in Lemma 5.1. To deduce existence from Theorem 2.5, we consider the constant marginals (μn,νn):=(μ,ν)(\mu_{n},\nu_{n}):=(\mu,\nu) and define approximating reference measures RnR_{n} via

where ana_{n} is the (finite) normalizing constant. As dRdP\frac{dR}{dP} is continuous and positive, hence bounded away from zero on small balls, the condition (2.4) is satisfied with Pn=PP_{n}=P and a function o(1)o(1) independent of nn. The static Schrödinger bridge problem (2.1) for RnR_{n} falls into the classical setting of Lemma 5.1 (i) because the product coupling π0:=μ⊗ν\pi_{0}:=\mu\otimes\nu satisfies H(π0∣Rn)<∞H(\pi_{0}|R_{n})<\infty. In particular, the associated cyclically invariant couplings πn∈Π(μ,ν)\pi_{n}\in\Pi(\mu,\nu) exist, and now Theorem 2.5 implies the existence of π∈Π(μ,ν)\pi\in\Pi(\mu,\nu) such that (π,R)(\pi,R) is cyclically invariant. Finally, Theorem 1.3 is a direct consequence of Theorem 2.4 via (2.2). ∎

Define R∈P(X×Y)R\in\mathcal{P}(\mathsf{X}\times\mathsf{Y}) by dR/dP:=fdR/dP:=f and let π\pi be as in Theorem 2.4. As (π,R)(\pi,R) is cyclically invariant, a version of the relative density admits a factorization dπ/dR(x,y)=φ(x)ψ(y)d\pi/dR(x,y)=\varphi(x)\psi(y) into positive Borel functions; cf. Proposition 4.4. The fact that π\pi has marginals (μ,ν)(\mu,\nu) then translates to the fact that (φ,ψ)(\varphi,\psi) solve the Schrödinger system. Uniqueness of φ,ψ\varphi,\psi up to a constant follows from the uniqueness of π\pi (here the fact that R∼PR\sim P is particularly important). ∎

Assumption 2.3 can be replaced by the assumption that the marginals (X,μ)(\mathsf{X},\mu) and (Y,ν)(\mathsf{Y},\nu) satisfy the so-called doubling property. The latter assumption is structurally different as it refers to the specific measures rather than the metric spaces. Indeed, let (X,μ)(\mathsf{X},\mu) be doubling; i.e., there exist C>0C>0 such that

for any ball Br(x)⊂XB_{r}(x)\subset\mathsf{X}. This ensures that differentiation of measures (in the sense of Assumption 2.3) holds for measures ρ≪μ\rho\ll\mu; cf. [31, Theorem 1.8, p. 4]. In particular, Lemma 3.2 holds as stated, and then so does Proposition 3.5. If (Y,ν)(\mathsf{Y},\nu) is also doubling, then so is (X×Y,P)(\mathsf{X}\times\mathsf{Y},P) where P=μ⊗νP=\mu\otimes\nu, showing that differentiation wrt. PP holds for measures ρ≪P\rho\ll P, in particular for ρ:=R∼P\rho:=R\sim P in the context of Proposition 4.2. As dR/dP>0dR/dP>0, it follows that differentiation also holds wrt. RR. This ensures that the proof of Proposition 4.2 remains valid, and hence the main results.

We show that stability of entropic optimal transport fails as soon as the cost function has an “essential” discontinuity. To see why the qualifier is necessary, consider a cost function cc of the form

Fix arbitrary (x0,y0)∈X×Y(x_{0},y_{0})\in\mathsf{X}\times\mathsf{Y}. As (x,y)↦c(x,y)−c(x,y0)−c(x0,y)(x,y)\mapsto c(x,y)-c(x,y_{0})-c(x_{0},y) is discontinuous, there is a sequence (xn,yn)→(x∞,y∞)(x_{n},y_{n})\to(x_{\infty},y_{\infty}) such that

Consider the marginals μn=(δx0+δxn)/2\mu_{n}=(\delta_{x_{0}}+\delta_{x_{n}})/2 and νn=(δy0+δyn)/2\nu_{n}=(\delta_{y_{0}}+\delta_{y_{n}})/2. Let πn∈Π(μn,νn)\pi_{n}\in\Pi(\mu_{n},\nu_{n}) be (c,1)(c,1)-cyclically invariant; that is,

After passing to a subsequence, πn\pi_{n} converge weakly to a coupling π∈Π(μ,ν)\pi\in\Pi(\mu,\nu) of μ:=(δx0+δx∞)/2\mu:=(\delta_{x_{0}}+\delta_{x_{\infty}})/2 and ν:=(δy0+δy∞)/2\nu:=(\delta_{y_{0}}+\delta_{y_{\infty}})/2. Suppose for contradiction that π\pi is cyclically invariant, then

As the left-hand side of (5.2) converges to the left-hand side of (5.3), the convergence of the right-hand sides follows, contradicting (5.1). ∎

References