Subgeometric rates of convergence of Markov processes in the Wasserstein metric

Oleg Butkovsky

Introduction

In this paper, we study rate of convergence of Markov processes to an invariant measure in the Wasserstein metric. We establish subgeometric bounds on the convergence rate, thus generalizing the results of DFG , DFMS , HMS . We apply the obtained estimates to prove subgeometric ergodicity of strong solutions of stochastic differential delay equations (SDDEs) under Veretennikov–Khasminskii-type conditions. This extends the corresponding results Ver97 , VerTVP , Mal , DFG for stochastic differential equations (without delay).

There are quite a few works which deal with convergence of Harris recurrent Markov chains in total variation; see, for example, the monograph MT and the references therein. Less is known about convergence of Markov chains that are not Harris recurrent. Recall HLL that if a Markov chain has a unique invariant measure, then either (a) the chain is positive Harris recurrent in an absorbing set and the invariant measure is nonsingular, or (b) the invariant measure is singular and there are no Harris sets. It is quite clear that in case (b) the marginal distributions of the Markov chain do not converge in total variation, whereas they might converge weakly (and, hence, in the Wasserstein metric). Thus, for non-Harris chains [case (b)] it is natural to study convergence in the Wasserstein metric (rather than in the total variation metric).

Many interesting Markov processes fall into case (b). For instance, following HMS , consider SDDE

where c>0c>0, WW is a one-dimensional Brownian motion and gg is a strictly increasing positive bounded continuous function. One can show that the strong solution of this equation has a unique invariant measure and converges to it weakly, but not in total variation. On the other hand, the Wasserstein distance between X(t)X(t) and the invariant measure decays exponentially to zero as t→∞t\to\infty. Section 3 contains further examples of processes belonging to case (b).

Many methods of estimation of convergence rates in the total variation metric assume that a Markov process is ψ\psi-irreducible and are based on the analysis of small sets. Probably, one of the first results in this area is due to Dobrushin Dobr , who proved that if the whole state space is small, then a Markov chain is exponentially ergodic. Later Popov Pop and Nummelin and Tuominen NT82 replaced the global Dobrushin condition with a combination of a local Dobrushin condition (existence of a “good” small set) and the Lyapunov drift condition (LDC). This result was further extended by Jarner and Roberts JR and Douc and coauthors DFMS , who established polynomial and general subgeometric estimates of convergence rate, correspondingly. Similar results for continuous time Markov processes (under an additional assumption that the state space is locally compact) are due to Fort and Roberts FR and Douc, Fort and Guillin DFG . The latter work provides subgeometric estimates of the convergence rate under condition that a certain functional of a Markov process is a supermartingale. Let us also mention the recent paper of Hairer and Mattingly HM , which contains a new simple proof of the exponential ergodicity of a Markov process under LDC and the local Dobrushin condition.

Thus, many techniques rely on the irreducibility of a Markov process, the existence of a “good” small set, and (for continuous time processes) the local compactness of the state space. However, if the state space is infinite-dimensional, then in most “typical” situations the process is non-Harris and, therefore these assumptions are not fulfilled. For instance, if we go back to the above SDDE, then it is easy to check that this processes is not ψ\psi-irreducible, the state space is not locally compact and, as was pointed in HMS , all small sets of this process are degenerate (i.e., consists of no more than one point).

An alternative to the local Dobrushin condition was suggested by Bakry, Cattiaux and Guillin in BCG . They obtained estimates of convergence rate in the total variation metric, provided that the LDC holds, and a Markov process has a unique invariant measure, which satisfies a local Poincaré inequality on a large enough set.

Let us discuss another alternative to this set of assumptions, which was developed by Hairer, Mattingly, and Scheutzow HMS specifically for establishing exponential convergence rates of SDDEs, stochastic PDEs, and other infinite-dimensional processes in the Wasserstein metric. Exploiting a new notion of a dd-small set (a generalization of the notion of a small set), in conjunction with the LDC, and without any additional assumptions on the irreducibility of the process, the authors proved the existence of a spectral gap in a suitable norm, and, hence, the exponential convergence to stationarity.

We extend this result and consider the more general situation where a spectral gap may not exist. For discrete time Markov processes (Theorem 2.1) we prove that existence of a “good” dd-small set and the LDC implies subgeometrical convergence in the Wasserstein metric. In the continuous time setting (Theorem 2.4) we obtain the same rate of convergence provided that there exists a “good” dd-small set and the Douc–Fort–Guillin supermartingale condition holds. Thus, we also extend the results of DFG , DFMS .

We apply our conditions to study the asymptotic behavior of strong solutions of SDDEs. We prove that Veretennikov–Khasminskii-type conditions are sufficient for subexponential ergodicity (Theorem 3.3). This extends the results of Ver97 , VerTVP , Mal , DFG .

The rest of the paper is organized as follows. Section 2 contains definitions and the main results. Applications to SDDEs and to an autoregressive model are presented in Section 3. The proofs of the main results are placed in Section 4.

Main results

Recall (see, e.g., BK ) that if dd is a semimetric on EE, then the Wasserstein semidistance WdW_{d} between probability measures μ,ν∈P(E)\mu,\nu\in\mathcal{P}(E) is given by

where C(μ,ν)\mathcal{C}(\mu,\nu) is the set of all probability measures on (E×E,B(E×E))(E\times E,\mathcal{B}(E\times E)) with marginals μ\mu and ν\nu. If dd is a proper metric, then WdW_{d} is a distance.

We consider also the total variation metric on the space P(E)\mathcal{P}(E), which is defined by the following formula:

A set A∈B(E)A\in\mathcal{B}(E) is called small for a Markov operator PP if there exists ε>0\varepsilon>0 such that for all x,y∈Ax,y\in A,

For instance, any one-point set is small. However, as discussed above, a Markov process might have no small sets that consist of more than one point. To study such Markov processes Hairer, Mattingly and Scheutzow HMS introduce the following concept.

A set A∈B(E)A\in\mathcal{B}(E) is called dd-small for a Markov operator PP if there exists ε>0\varepsilon>0 such that for all x,y∈Ax,y\in A,

Note that our definition of a dd-small set is a bit different from the definition of HMS . Namely, the multiplier d(x,y)d(x,y) appears on the right-hand side of the above inequality.

Before we present our main result, let us recall that the total variation metric is contracting, that is, for any Markov semigroup (Pt)t≥0(P^{t})_{t\geq 0} one has

whenever 0≤s≤t0\leq s\leq t. In general, the Wasserstein metric WdW_{d} may not be contracting. However, as discussed in detail in HMS , it is natural to focus only on Wasserstein metrics that are contracting for the process XX, since, in the general case, the Lyapunov drift condition is not sufficient even for a weak convergence toward the invariant measure. Note that the contractivity condition itself does not imply any convergence at all, either. It is the combination of the contractivity, the Lyapunov drift condition and the existence of a “good” dd-small set, which yields the existence and uniqueness of the invariant measure and subgeometric convergence in the Wasserstein metric.

Since HfH_{f} is increasing, the inverse function Hf−1H_{f}^{-1} is well defined.

Suppose there exist a measurable function V\dvtxE→[0;∞)V\dvtx E\to[0;\infty) and a metric dd on EE such that the following conditions hold: {longlist}[(3)]

The space (E,d)(E,d) is a complete separable metric space.

The metric dd is contracting and bounded by 11; that is, for any x,y∈Ex,y\in E,

The level set L:={x,y∈E\dvtxV(x)+V(y)≤R}L:=\{x,y\in E\dvtx V(x)+V(y)\leq R\} is dd-small for some R>φ−1(2K)R>\varphi^{-1}(2K); that is, there exists ρ>0\rho>0 such that

Then the process XX has a unique stationary measure π\pi and

Moreover, for any ε>0\varepsilon>0 there exist constants C1C_{1} and C2C_{2} such that for all x∈Ex\in E,

(i) If φ\varphi is a linear function, then the rate of convergence is exponential and this case is covered by HMS , Theorem 4.8.

Conditions (3) and (4) of the theorem are a bit more general than the corresponding conditions from HMS , Theorem 4.8. Namely, we do not assume here that Wd(P(x,⋅),P(y,⋅))≤(1−ρ) d(x,y)W_{d}(P(x,\cdot),P(y,\cdot))\leq(1-\rho)\,d(x,y) for all x,y∈Ex,y\in E such that d(x,y)<1d(x,y)<1. We suppose that this inequality is satisfied only for xx, yy belonging to the sublevel set.

Note that if φ\varphi grows to infinity not very rapidly (as xγx^{\gamma} for some 0<γ<10<\gamma<1 or slower), then the estimate of convergence rate given by (3) can be as close as possible to the estimate of convergence rate in the total variation distance obtained in DFMS , Proposition 2.5. Specific examples of convergence rates (polynomial, logarithmic, etc.) for different functions φ\varphi are given in DFMS , Section 2.3.

While the proof of the theorem is postponed to Section 4, we outline now the main steps.

Sketch of the proof of Theorem 2.1 To prove the theorem we develop the idea of constructing an auxiliary contracting semimetric Hair , HM , HMS . Namely, let ll be a semimetric on the space EE such that d(x,y)≤l(x,y)d(x,y)\leq l(x,y) for all x,y∈Ex,y\in E. It is possible to prove (for some “good” ll) that for any probability measures μ,ν∈Pφ∘V(E)\mu,\nu\in\mathcal{P}_{\varphi\circ V}(E)

where χ\chi is a positive function (this is done in Lemma 4.3). Hence

Of course, since we want to obtain subgeometric estimates of Wd(Pnμ,Pnν)W_{d}(P^{n}\mu,P^{n}\nu), there is no hope that inf⁡μ,ν∈Pφ∘V(E)χ(μ,ν)\inf_{\mu,\nu\in\mathcal{P}_{\varphi\circ V}(E)}\chi(\mu,\nu) is positive (this lower bound was greater than zero in Hair , HM , HMS , where geometric estimates were obtained). Yet, a good (albeit nonuniform) estimate of χ(Pi+1μ,Pi+1ν)\chi(P^{i+1}\mu,P^{i+1}\nu) can be derived. However, this estimate depends not only on Wl(Piμ,Piν)W_{l}(P^{i}\mu,P^{i}\nu) but also on μ(Pi(φ∘V))\mu(P^{i}(\varphi\circ V)) and ν(Pi(φ∘V))\nu(P^{i}(\varphi\circ V)). The latter two expressions are unbounded if μ,ν\mu,\nu are fixed, and ii runs over positive integers. Fortunately, there are sufficiently many integers ii such that these two expressions are “small” (Lemma 4.1). This allows us to overcome this obstacle (Lemma 4.4) and obtain subgeometric bounds on Wd(Pnμ,Pnν)W_{d}(P^{n}\mu,P^{n}\nu). The last step is to prove the existence and uniqueness of the stationary measure (Lemma 4.5).

Now we give a similar result for continuous time Markov processes. Let X=(Xt)t≥0X=(X_{t})_{t\geq 0} be a time-homogeneous strong Markov process, and let (Pt)t≥0(P_{t})_{t\geq 0} be the associated Markov semigroup. Recall DY , Theorem 2, that if a Markov process has càdlàg paths, then the strong Markov property is implied by the Feller property.

Suppose there exist a measurable function V\dvtxE→[0;∞)V\dvtx E\to[0;\infty) and a metric dd on EE such that the following conditions hold: {longlist}[(2)]

The space (E,d)(E,d) is a complete separable metric space.

The metric dd is bounded by 11 and contracting for all t≥t0t\geq t_{0}, for some t0≥0t_{0}\geq 0; that is, for any x,y∈Ex,y\in E

The level set L:={x,y∈E\dvtxV(x)+V(y)≤R}L:=\{x,y\in E\dvtx V(x)+V(y)\leq R\} is dd-small for all R>0R>0 and all t≥t0t\geq t_{0}, that is, there exists ρ=ρ(R,t)>0\rho=\rho(R,t)>0 such that

Then the process XX has a unique stationary measure π\pi and π(φ∘V)≤K\pi(\varphi\circ V)\leq K. Moreover, for any ε>0\varepsilon>0 there exist constants C1C_{1} and C2C_{2} such that for all x∈Ex\in E,

(i) The linear case φ(x)=λx\varphi(x)=\lambda x, λ>0\lambda>0 is HMS , Theorem 4.8.

(i) Condition (1) of Theorem 2.4 is equivalent to the Douc–Fort–Guillin supermartingale condition DFG , equation (3.2); that is, inequality (4) holds if and only if the process Z:=(Zt)t≥0Z:=(Z_{t})_{t\geq 0},

is a supermartingale with respect to the natural filtration of the process XX.

(ii) Let LL be the extended generator (see, e.g., RevuzYor , Definition 7.1.8) of the Markov process XX. If the function VV belongs to the domain of LL and

The proof of this theorem is given in Section 4. Let us describe here the main idea.

Sketch of the proof of Theorem 2.4 Combining the technique from DFG , FR , NT , we find a function W\dvtxE→[0;∞)W\dvtx E\to[0;\infty) such that

Thus Theorems 2.1 and 2.4 suggest a new method for proving results concerning subgeometrical convergence. Namely, one needs to find a suitable contracting metric dd and a suitable Lyapunov function VV with dd-small sublevel sets, such that the conditions of the theorems hold. It extends the ability of the existing methods by allowing to choose the metric dd (which might be different from the discrete metric).

Examples and applications

Let us give some applications of the results of the previous section. The focus here is on stochastic delay equations; however, it is possible to apply the results of this kind to study convergence in the Wasserstein metric for other classes of Markov processes; see, for example, HMS , Section 5.3, for estimates of convergence rates of stochastic partial differential equations.

An invariant measure π\pi is called singular if for any x∈Ex\in E there exists an absorbing set SxS_{x} such that x∈Sxx\in S_{x} and π(Sx)=0\pi(S_{x})=0. In other words, the Markov chain, whatever the starting point is, will remain in the set of π\pi-measure 0.

Consider the following peculiar AR(1) process, which belongs to case (b).

where ε1,ε2,…\varepsilon_{1},\varepsilon_{2},\ldots are i.i.d. random variables uniformly distributed on the set {0,110,…,910}\{0,\frac{1}{10},\ldots,\frac{9}{10}\} and X0∈[0;1)X_{0}\in[0;1). In other words, to get Xn+1X_{n+1} from XnX_{n} one needs to take the decimal notation of XnX_{n} (which starts with 0 followed by the decimal point) and insert a random digit immediately after the decimal point. Other digits in the decimal notation of XnX_{n} are shifted right by one position.

Clearly, XX is a Markov process with state space (E,E)=([0;1),B([0;1)))(E,\mathcal{E})=([0;1),\mathcal{B}([0;1))). Let dd be the Euclidean metric on this space [i.e., d(x,y)=∣x−y∣d(x,y)=|x-y|, x,y∈Ex,y\in E]. One can easily prove that the process XX has a unique invariant measure π\pi, which is uniformly distributed on the interval [0;1)[0;1). Moreover, the sequence {Xn}\{X_{n}\} weakly converges to π\pi as n→∞n\to\infty.

This autoregression has a number of very interesting and unusual features. First, it has a reconstruction property. Namely, if we have just one observation of XnX_{n}, where the integer nn can be arbitrarily large, then it is possible to find an initial value X0X_{0} with probability 11 by the following simple formula: X0={10nXn}X_{0}=\{10^{n}X_{n}\}, where {b}\{b\} denotes the fractional part of a real bb. In other words, one just needs to shift right the decimal point by nn positions and drop all the digits which will be on the left of the decimal point.

Therefore for xx, y∈Ey\in E, x≠yx\neq y, the probability measures P(x,⋅)P(x,\cdot) and P(y,⋅)P(y,\cdot) are singular. Hence the process XX has no nontrivial small sets. On the other hand, the whole state space EE is dd-small. Indeed, it is easily seen that Wd(P(x,⋅),P(y,⋅))≤∣x−y∣/10W_{d}(P(x,\cdot),P(y,\cdot))\leq|x-y|/10, for any x,y∈Ex,y\in E.

2 Stochastic delay equations

In this subsection we present our results on convergence of SDDEs in the Wasserstein metric.

Consider the stochastic differential delay equation

Throughout this section we assume that the drift and the diffusion satisfy the following conditions:

the drift satisfies a one-sided Lipschitz condition, and the diffusion is Lipschitz; that is, there exists K>0K>0 such that for any x,y∈Cx,y\in\mathcal{C}

the diffusion is nondegenerate; that is, for any x∈Cx\in\mathcal{C} the matrix g(x)g(x) admits a right inverse g−1(x)g^{-1}(x) and

(3.4) ff is continuous and bounded on bounded subsets of C\mathcal{C}.

Conditions (7) and (3.4) imply RS the existence and uniqueness of the strong solution of SDDE (6).

Now we give a general theorem, which describes convergence rates in the Wasserstein metric WdβW_{d_{\beta}}. Theorem 3.2(i) is a generalization of HMS , Assumption 5.1.

lim⁡∥x∥→∞V(x)=∞\lim_{\|x\|\to\infty}V(x)=\infty or {longlist}[(ii)]

where the function f1f_{1} is bounded; then SDDE (6) has a unique invariant measure π\pi. Furthermore, for any β>0\beta>0, the rate of convergence of Law⁡(Xt)\operatorname{Law}(X_{t}) to π\pi in the Wasserstein metric WdβW_{d_{\beta}} is given by (5).

Fix β>0\beta>0. Let us check that the process XX and the function VV satisfy the conditions of Theorem 2.4. It follows from HMS , Proposition 5.4, and Shir , Lemma 3.7.2, that the process XX is Feller. Since XX has continuous paths, we see that XX is strongly Markovian. The first condition of the theorem is satisfied by assumption. The second condition also holds. In case (i) it follows directly from HMS , Section 5.2, that there exists a γ∈(0;β)\gamma\in(0;\beta) such that the third and the fourth conditions are met. In case (ii), arguing as in HMS , Proposition 5.3 and Lemma 3.8, one can show that the set {x∈C\dvtx∣x(0)∣≤R}\{x\in\mathcal{C}\dvtx|x(0)|\leq R\}, R≥0R\geq 0 is dγd_{\gamma}-small for some γ∈(0;β)\gamma\in(0;\beta), and the metric dγd_{\gamma} is contracting. Thus, in both cases the conditions of Theorem 2.4 are satisfied.

Apply Theorem 2.4 to the process XX. It follows from this theorem that SDDE (6) has a unique invariant measure π\pi, and the rate of convergence of Law⁡(Xt)\operatorname{Law}(X_{t}) to π\pi in the metric WdγW_{d_{\gamma}} is provided in (5). To complete the proof, it remains to note that for any measures μ1,μ2∈P(E)\mu_{1},\mu_{2}\in\mathcal{P}(E) one has Wdβ(μ1,μ2)≤Wdγ(μ1,μ2)W_{d_{\beta}}(\mu_{1},\mu_{2})\leq W_{d_{\gamma}}(\mu_{1},\mu_{2}).

Ergodic properties of stochastic differential equations (SDE) were studied by Veretennikov Ver97 , VerTVP , Malyshkin Mal , Klokov KV , Douc, Fort and Guillin DFG and many others. It is known that the Veretennikov–Khasminskii condition on the drift combined with a certain nondegeneracy condition on the diffusion is sufficient for the existence and uniqueness of the invariant measure for the strong solution of an SDE. Moreover, these conditions yield exponential, subexponential or polynomial (depending on the value of the constant α\alpha, see below) convergence toward the invariant measure in the total variation metric PV , DFG . The following theorem extends these results to SDDE.

Suppose conditions (7)–(3.4) hold, Λ<∞\Lambda<\infty and the function f1f_{1} in decomposition (5) is bounded. {longlist}[(ii)]

Assume additionally that for some constants α∈(0,1]\alpha\in(0,1], M>0M>0, ϰ>0\varkappa>0, the generalized Veretennikov–Khasminskii condition holds, that is,

Then SDDE (6) has a unique invariant measure π\pi, and Law⁡(Xt)\operatorname{Law}(X_{t}) converges to π\pi in the Wasserstein metric WdβW_{d_{\beta}} subexponentially (if 0<α<10<\alpha<1) or exponentially (if α=1\alpha=1); that is, for any β>0\beta>0 there exists positive constants C1C_{1} and C2C_{2} such that

If (6) holds with α=0\alpha=0 and ϰ>nΛ/2\varkappa>n\Lambda/2, then SDDE (6) has a unique invariant measure π\pi, but Law⁡(Xt)\operatorname{Law}(X_{t}) converges to π\pi in the Wasserstein metric WdβW_{d_{\beta}} only polynomially; that is, for any β>0\beta>0, ε>0\varepsilon>0 there exist C>0C>0 such that

where ϰ0=(ϰ−nΛ/2)λ+−1\varkappa_{0}=(\varkappa-n\Lambda/2)\lambda_{+}^{-1}.

where C1=λ+(α−2)+nΛC_{1}=\lambda_{+}(\alpha-2)+n\Lambda, C2>0C_{2}>0, C3=ϰ−12λ+αk−12C1M0−αC_{3}=\varkappa-\frac{1}{2}\lambda_{+}\alpha k-\frac{1}{2}C_{1}M_{0}^{-\alpha} and in the second inequality we made use of (6).

where C4:=αkϰ/4C_{4}:=\alpha k\varkappa/4, C5:=C4k2/α−2C_{5}:=C_{4}k^{2/\alpha-2} and C6>0C_{6}>0. Thus the function VV satisfies inequality (4). Theorem 3.2(ii) now yields the existence and the uniqueness of the invariant measure π\pi and implies estimate (7).

(ii) Now let U(v)=∣v∣kU(v)=|v|^{k}, where k>2k>2. We take V(x)=U(x(0))V(x)=U(x(0)) and proceed as follows:

where C1=ϰ−k−22λ+−nΛ2C_{1}=\varkappa-\frac{k-2}{2}\lambda_{+}-\frac{n\Lambda}{2}, C2>0C_{2}>0. Set

where ε>0\varepsilon>0. By choosing ε>0\varepsilon>0 small enough we can ensure that k>2k>2. Take φ(u)=u(k−2)/k\varphi(u)=u^{(k-2)/k}. Then

for some C3C_{3}, C4>0C_{4}>0. Thus the function VV satisfies condition (4), and the statement of the theorem follows now from Theorem 3.2(ii).

Proofs of the main results

To prove Theorems 2.1 and 2.4 we introduce some notation. Consider a semimetric l(x,y):=d(x,y)1/p(1+βφ(V(x)+V(y)))1/ql(x,y):=d(x,y)^{1/p}(1+\beta\varphi(V(x)+V(y)))^{1/q}, where β>0\beta>0, p,q>1p,q>1 and 1/p+1/q=11/p+1/q=1. These parameters will be chosen later. We start with two auxiliary lemmas.

Furthermore, if a measure π\pi is invariant for the process XX, then π∈Pφ∘V(E)\pi\in\mathcal{P}_{\varphi\circ V}(E) and π(φ∘V)≤K\pi(\varphi\circ V)\leq K.

To prove the second part of the lemma we combine the first part of the lemma with a cut-off argument; see, for example, Hair06 , Proposition 4.24. Fix L>0L>0. Then, for any nonnegative integer ii, we have

Summing the both sides of the above inequality over all 0≤i<n0\leq i<n, we derive

Lebesgue’s dominated convergence theorem implies that the integral on the right-hand side of the above inequality tends to as n→∞n\to\infty. Thus

and the second part of the lemma follows from Fatou’s lemma.

The following Lemma 4.2 is due to Petrov.

where ψ\dvtx[0;∞)→[0;1]\psi\dvtx[0;\infty)\to[0;1] is a continuous increasing function with ψ(0)=0\psi(0)=0 and ψ(x)>0\psi(x)>0 for x>0x>0. Then

We see that the function g−1g^{-1} is well defined. This follows from the fact that the function gg is nonnegative, unbounded and strictly decreasing. Since ψ\psi is positive, we have an+1≤ana_{n+1}\leq a_{n}. By the mean value theorem, there exists s∈[an+1;an]s\in[a_{n+1};a_{n}] such that

Hence g(an)≥ng(a_{n})\geq n and an≤g−1(n)a_{n}\leq g^{-1}(n).

The next key lemma gives the estimate of the contraction rate in one step.

Assume that the conditions of Theorem 2.1 hold. Then there exist β=β(p,q)\beta=\beta(p,q) and positive c1(p,q),c2(p,q),c3(p,q)c_{1}(p,q),c_{2}(p,q),c_{3}(p,q) such that for any μ,ν∈Pφ∘V(E)\mu,\nu\in\mathcal{P}_{\varphi\circ V}(E) one has

where c:=c3(μ(φ∘V)+ν(φ∘V))pc:=c_{3}(\mu(\varphi\circ V)+\nu(\varphi\circ V))^{p} and the semimetric ll was introduced at the beginning of this section.

Here, as usual, a∧b=min⁡(a,b)a\wedge b=\min(a,b) and a∨b=max⁡(a,b)a\vee b=\max(a,b) for real aa, bb. To simplify the formulas, we will drop a pair of parentheses and write 1−a∧b1-a\wedge b for 1−(a∧b)1-(a\wedge b).

Proof of Lemma 4.3 We start as in the proof of HMS , Theorem 4.8, by observing that since WlW_{l} is convex, the Jensen inequality implies

for any μ,ν∈Pφ∘V(E)\mu,\nu\in\mathcal{P}_{\varphi\circ V}(E) and any α∈C(μ,ν)\alpha\in\mathcal{C}(\mu,\nu). Applying the Cauchy–Schwarz inequality and the Jensen inequality for concave functions, we find that

where the infimum is taken over all measures λ∈C(P(x,⋅),P(y,⋅))\lambda\in\mathcal{C}(P(x,\cdot),P(y,\cdot)).

To estimate the right-hand side of the last inequality we consider three different cases. Note once again that contrary to the proof of HMS , Theorem 4.8, it is impossible here to obtain a nontrivial upper uniform bound for Wl(P(x,⋅),P(y,⋅))/l(x,y)W_{l}(P(x,\cdot),P(y,\cdot))/l(x,y).

Case 1. V(x)+V(y)≤RV(x)+V(y)\leq R. In this case we proceed similar to Hair , HMS . Using (4) and conditions (1) and (4) of the theorem, we obtain

Case 2. R<V(x)+V(y)≤MR<V(x)+V(y)\leq M. In this case we make use of (1) and the concavity of φ\varphi to derive

Clearly, if u∈(R;M]u\in(R;M], then again by the concavity of φ\varphi we have

where θ:=1−2K/φ(R)\theta:=1-2K/\varphi(R). This inequality, combined with (4), (4) and contraction property (2), yields

Case 3. V(x)+V(y)>MV(x)+V(y)>M. This is the easiest situation because in this case we would like to derive a very weak estimate of Wl(P(x,⋅),P(y,⋅))W_{l}(P(x,\cdot),P(y,\cdot)). Combining (2), (4) and (4), we get

Now we return to the main line of the proof. Introduce

Note that the values of c1c_{1} and c2c_{2} depend neither on the choice of MM nor on measures μ\mu and ν\nu. We see from (10) and the above estimates of Wl(P(x,⋅),P(y,⋅))W_{l}(P(x,\cdot),P(y,\cdot)) that for all M>RM>R one has

The second integral on the right-hand side of (4) is estimated using Chebyshev inequality. Namely,

where C=1/K+β+1C=1/K+\beta+1, and in the second inequality we used the bound φ(M)>K\varphi(M)>K. Note that μ(φ∘V)\mu(\varphi\circ V) as well as ν(φ∘V)\nu(\varphi\circ V) are finite because it was assumed that μ,ν∈Pφ∘V(E)\mu,\nu\in\mathcal{P}_{\varphi\circ V}(E).

Recall that α\alpha is an arbitrary element of C(μ,ν)\mathcal{C}(\mu,\nu). Hence we can take the infimum over all α∈C(μ,ν)\alpha\in\mathcal{C}(\mu,\nu) in (4) and use the above inequality to derive

Now we can choose MM in such a way, that the right-hand side of the above expression is always smaller than Wl(μ,ν)W_{l}(\mu,\nu). Namely, it is sufficient to require that

where c3=c3(p,q,R,K)=2p(1/K+β+1)pc_{3}=c_{3}(p,q,R,K)=2^{p}(1/K+\beta+1)^{p}. The substitution of the last expression into (4) proves the lemma.

We begin by observing that for any measures ζ1,ζ2∈Pφ∘V(E)\zeta_{1},\zeta_{2}\in\mathcal{P}_{\varphi\circ V}(E) one has

where we used the concavity of the function φ\varphi and the bound d≤1d\leq 1. Hence,

and c5,c6,…c_{5},c_{6},\ldots are some positive constants. Note that to obtain the third identity, we made the change of variables u=φ−1(c4t−p)u=\varphi^{-1}(c_{4}t^{-p}). Thus we finally get ank≤c11φ(Hφ−1(c12k))−1/pa_{n_{k}}\leq c_{11}\varphi(H_{\varphi}^{-1}(c_{12}k))^{-1/p} and hence

Under the conditions of Theorem 2.1, the process XX has a unique stationary measure π\pi.

Here the symbol #\# denotes the cardinality of a finite set. It follows from the above definitions that for n<mn<m,

where we used Lemma 4.3 to obtain the first inequality. Recall that the constants C1,C2C_{1},C_{2} are independent of n,mn,m.

It follows from (8) that for any fixed nn there exists an arbitrarily large mm such that A(mn,(m+1)n)≥3n/4A(mn,(m+1)n)\geq 3n/4. Since A(0,n)≥3n/4A(0,n)\geq 3n/4, inequality (17) implies that for any fixed nn there exists an arbitrarily large mm such that B(n,m)≥n/2B(n,m)\geq n/2. It is clear that for all such mm, one has

It is evident that Ψ(n)→0\Psi(n)\to 0, as n→∞n\to\infty.

for all integers k,mk,m. Since the space (P(E),Wd)(\mathcal{P}(E),W_{d}) is complete (see, e.g., BK , Theorem 1.1.3), we see that there exists a measure π∈P(E)\pi\in\mathcal{P}(E) such that Wd(Pnkδx,π)→0W_{d}(P^{n_{k}}\delta_{x},\pi)\to 0.

Let us verify that the measure π\pi is stationary, that is, let us check that Pπ=πP\pi=\pi. Note that the metric WdW_{d} is contractive. Indeed, for any μ,ν∈P(E)\mu,\nu\in\mathcal{P}(E), we have

where we used the Jensen inequality and condition (2).

The first term on the right-hand side of the last expression tends to , as k→∞k\to\infty. To estimate the second term, we observe that if nn is a positive integer, then A(0,n)≥3n/4A(0,n)\geq 3n/4 and A(1,n+1)≥3n/4−1A(1,n+1)\geq 3n/4-1. Therefore, inequality (17) implies B(n,n+1)>n/2−1B(n,n+1)>n/2-1. This, combined with (4), yields

Hence Wl(Pnkδx,Pnk+1δx)→0W_{l}(P^{n_{k}}\delta_{x},P^{n_{k}+1}\delta_{x})\to 0 as k→∞k\to\infty, and we conclude from (4) that Wd(Pπ,π)=0W_{d}(P\pi,\pi)=0, which implies the stationarity of the measure π\pi.

To complete the proof of the lemma it remains to prove the uniqueness of stationary measure. Suppose that, on the contrary, the process XX has two stationary measures π1\pi_{1} and π2\pi_{2} and π1≠π2\pi_{1}\neq\pi_{2}. By Lemma 4.1, π1,π2∈Pφ∘V(E)\pi_{1},\pi_{2}\in\mathcal{P}_{\varphi\circ V}(E) and hence 0<Wl(π1,π2)<∞0<W_{l}(\pi_{1},\pi_{2})<\infty. We make use of stationarity of the measures and Lemma 4.3 to obtain

Proof of Theorem 2.1 It follows from Lemmas 4.1 and 4.5, that the process XX has a unique stationary measure π∈Pφ∘V(E)\pi\in\mathcal{P}_{\varphi\circ V}(E) and π(φ∘V)≤K\pi(\varphi\circ V)\leq K. Fix x∈Ex\in E and consider the following sequence. Let n0=0n_{0}=0 and

We make use of stationarity of π\pi, the bound π(φ∘V)≤K\pi(\varphi\circ V)\leq K and the definition of nkn_{k} to derive

On the other hand, it follows from (8) that nk≤2kn_{k}\leq 2k. To complete the proof, it remains to take 1/p=1−ε1/p=1-\varepsilon and note that

To switch from discrete time to continuous time and prove Theorem 2.4, we combine different methods from DFG , FR , NT . First of all for a set C∈B(E)C\in\mathcal{B}(E), introduce the hitting time delayed by δ>0\delta>0

and the hitting and return times of the skeleton chain

where m>0m>0. Denote for brevity CR:={x∈E\dvtxV(x)≤R}C_{R}:=\{x\in E\dvtx V(x)\leq R\}.

If R>0R>0 and φ(R)>K\varphi(R)>K, then under the conditions of Theorem 2.4

Fix L>δL>\delta. Observe that if δ≤u<τCR(δ)\delta\leq u<\tau_{C_{R}}(\delta), then by definition V(Xu)≥RV(X_{u})\geq R. Combining this with (4) we obtain

The desired inequality follows now from the Fatou lemma.

Let m>0m>0. If R>KmR>Km and φ(R−Km)>K\varphi(R-Km)>K, then under the conditions of Theorem 2.4,

where c1=c1(m,R,K)c_{1}=c_{1}(m,R,K) and c2=c2(m,R,K)c_{2}=c_{2}(m,R,K) are positive functions that do not depend on xx.

The proof of the lemma uses the ideas from the proof of FR , Proposition 22(ii). However, note that we cannot apply this proposition directly because in contrast to Fort and Roberts, we assumed neither that the set {V(x)≤R}\{V(x)\leq R\} is petite nor that the process XX is Harris-recurrent with invariant measure.

Introduce R′<R−KmR^{\prime}<R-Km such that φ(R′)>K\varphi(R^{\prime})>K. The existence of such R′R^{\prime} follows from the conditions of the lemma. Consider the following sequence of stopping times:

It follows from the choice of R′R^{\prime} that γ>0\gamma>0.

We combine this with Lemma 4.6 to finally obtain

for all x∈Ex\in E. This completes the proof of the statement.

Proof of Theorem 2.4 First let us prove that there exist a Lyapunov function W\dvtxE→[0,∞)W\dvtx E\to[0,\infty) and positive constants K1K_{1}, K2K_{2} such that

Choose a sufficiently large RR (such that the conditions of Lemma 4.7 hold with m=t0m=t_{0}), and let

It follows from MT , Theorem 11.3.5(i) that for x∈Ex\in E

Using an argument similar to that in the proof of DFG , Proposition 4.8(i), we obtain for any L>0L>0 and x∈Ex\in E,

Furthermore, using condition (4) and the concavity of the function φ\varphi, we get for any x∈Ex\in E,

Combining this with the previous inequality and using Lemma 4.7 and Fatou’s lemma, we derive for any x∈Ex\in E,

where c1c_{1} and c2c_{2} are defined in Lemma 4.7, c3:=φ(1)+K+φ′(1)Kt0c_{3}:=\varphi(1)+K+\varphi^{\prime}(1)Kt_{0}, c4:=1/t0+c1c3c_{4}:=1/t_{0}+c_{1}c_{3}, c5=c2c3c_{5}=c_{2}c_{3}. Therefore, by the concavity of φ\varphi,

This bound, together with (22) and (4), yields

for some positive c6c_{6}, c7c_{7}. Hence the function WW satisfies (21).

We combine this with condition (4) of the theorem to conclude that for any t>t0t>t_{0},

for some C3>0C_{3}>0. Here ⌊b⌋\lfloor b\rfloor denotes the lower integer part of a real bb. This completes the proof of Theorem 2.4.

Acknowledgments

The author is grateful to Professor A. V. Bulinski and Professor A. Yu. Veretennikov for their help and constant attention to this work. The author also would like to thank Professor M. Hairer and F. V. Petrov for useful discussions and the referee for his valuable comments and suggestions which helped to improve the quality of the paper.

References