Uniform logarithmic Sobolev inequalities for conservative spin systems with super-quadratic single-site potential

Georg Menz, Felix Otto

Introduction and main result

For a real number mm, we consider the N−1N-1 dimensional hyper-plane XN,mX_{N,m} given by

The restriction of μ\mu to XN,mX_{N,m} is called canonical ensemble μN,m\mu_{N,m}, that is,

Here, H⌊XN,mN−1\mathcal{H}_{\lfloor X_{N,m}}^{N-1} denotes the N−1N-1 dimensional Hausdorff measure restricted to the hyperplane XN,mX_{N,m}. For convenience, we introduce the notation:

In 1993, Varadhan (Var , Lemma 5.3 ff.) posed the question for which kind of single-site potential ψ\psi the canonical ensemble μN,m\mu_{N,m} satisfies a spectral gap inequality (SG) uniformly in the system size NN and the mean spin mm. A partial answer was given by Caputo Cap :

Assume that for the single-site potential ψ\psi there exist a splitting ψ=ψc+δψ\psi=\psi_{c}+\delta\psi and constants β−\beta_{-}, β+∈[0,∞)\beta_{+}\in[0,\infty) such that for all x∈[0,∞),x\in[0,\infty),

Then the canonical ensemble μN,m\mu_{N,m} satisfies the SG with constant ϱ>0\varrho>0 uniformly in the system size NN and the mean spin mm. More precisely, for any function ff,

Here, ∇\nabla denotes the gradient determined by the Euclidean structure of XN,mX_{N,m}.

In this article, we give a full answer to the question by Varadhan Var and also show that the last theorem can be strengthened to the logarithmic Sobolev inequality (LSI).

Let XX be a Euclidean space. A Borel probability measure μ\mu on XX satisfies the LSI with constant ϱ>0\varrho>0, if for all functions f≥0f\geq 0

Here, ∇\nabla denotes the gradient determined by the Euclidean structure of XX.

The LSI was originally introduced by Gross Gross . It yields the SG and can be used as a powerful tool for studying spin systems. Like the SG, the LSI implies exponential convergence to equilibrium of the naturally associated conservative diffusion process. The rate of convergence is given by the LSI constant ϱ\varrho; cf. R , Chapter 3.2, and Remark 1.7. Therefore, an appropriate scaling of the LSI constant in the system size indicates the absence of phase transitions. The SG yields convergence in the sense of variances in contrast to the LSI, which yields convergence in the sense of relative entropies. The SG and the LSI are also useful for deducing the hydrodynamic limit; see Var for the SG and GORV for the LSI.

We consider three cases of different potentials: sub-quadratic, quadratic and super-quadratic single-site potentials. In the case of sub-quadratic single-site potentials, Barthe and Wolff BW gave a counterexample where the scaling in the system size of the SG and the LSI constant of the canonical ensemble differs in the system size. More precisely, they showed:

Assume that the single-site potential ψ\psi is given by

Then the SG constant ϱ1\varrho_{1} and the LSI constant ϱ2\varrho_{2} of the canonical ensemble μN,m\mu_{N,m} satisfy

In the case of perturbed quadratic single-site potentials it is known that Theorem 1.1 can be improved to the LSI. More precisely, several authors (cf. LPY , Cha , GORV ) deduced the following statement by different methods:

Assume that the single-site potential ψ\psi is perturbed quadratic in the following sense: There exists a splitting ψ=ψc+δψ\psi=\psi_{c}+\delta\psi such that

Then the canonical ensemble μN,m\mu_{N,m} satisfies the LSI with constant ϱ>0\varrho>0 uniformly in the system size NN and the mean spin mm.

There is only left to consider the super-quadratic case. It is conjectured that the optimal scaling LSI also holds if the single-site potential ψ\psi is a bounded perturbation of a strictly convex function; cf. LPY , page 741, Cha , Theorem 0.3 f., and Cap , page 226. Heuristically, this conjecture seems reasonable: Because the LSI is closely linked to convexity (consider, e.g., the Bakry–Émery criterion), a perturbed strictly convex potential should behave no worse than a perturbed quadratic one. However technically, the methods for the quadratic case are not able to handle the perturbed strictly convex case because they require an upper bound on the second derivative of the Hamiltonian. In the main result of the article we show that the conjecture from above is true:

Assume that the single-site potential ψ\psi is perturbed strictly convex in the sense that there is a splitting ψ=ψc+δψ\psi=\psi_{c}+\delta\psi such that

Then the canonical ensemble μN,m\mu_{N,m} satisfies the LSI with constant ϱ>0\varrho>0 uniformly in the system size NN and the mean spin mm.

Note that the standard criteria for the SG and the LSI (cf. Appendix) fail for the canonical ensemble μN,m\mu_{N,m}:

The Tensorization principle for the SG and the LSI does not apply because of the restriction to the hyper-plane XN,mX_{N,m}; cf. GZ , Theorem 4.4, or Theorem .1.

The Bakry–Émery criterion does not apply because the Hamiltonian HH is not strictly convex; cf. BE , Proposition 3 and Corollary 2, or Theorem .3.

The Holley–Stroock criterion does not help because the LSI constant ϱ\varrho has to be independent of the system size NN; cf. HS , page 1184, or Theorem .2.

Therefore, a more elaborated machinery was needed for the proof of Theorems 1.1 and 1.5. The approach of Caputo to Theorem 1.1 seems to be restricted to the SG because it relies on the spectral nature of the SG. For the proof of Theorem 1.5, Landim, Panizo and Yau LPY and Chafaï Cha used the Lu–Yau martingale method that was originally introduced in LY to deduce an analog version of Theorem 1.5 in the case of discrete spin values. Recently, Grunewald, Villani, Westdickenberg and the second author GORV provided a new technique for deducing Theorem 1.5, called the two-scale approach. We follow this approach in the proof of Theorem 1.6.

The limiting factor for extending Theorem 1.5 to more general single-site potentials is almost the same for the Lu–Yau martingale method and for the two-scale approach: It is the estimation of a covariance term w.r.t. the measure μN,m\mu_{N,m} conditioned on a special event; cf. LPY , (4.6), and GORV , (42). In the two-scale approach one has to estimate for some large but fixed K≫1K\gg 1 and any nonnegative function ff the covariance

In GORV , this term term was estimated by using a standard estimate (cf. Lemma 2.10 and GORV , Lemma 22) that only can be applied for perturbed quadratic single-site potentials ψ\psi. We get around this difficulty by making the following adaptations: Instead of one-time coarse-graining of big blocks, we consider iterative coarse-graining of pairs. As a consequence we only have to estimate the covariance term from above in the case K=2K=2. Because μ2,m\mu_{2,m} is a one-dimensional measure, we are able to apply the more robust asymmetric Brascamp–Lieb inequality (cf. Lemma 2.11) that can also be applied for perturbed strictly convex single-site potentials ψ\psi.

Recently, the optimal scaling LSI was established in Menz by the first author for a weakly interacting Hamiltonian with perturbed quadratic single-site potentials ψ\psi, that is,

Because the original two-scale approach was used, it is an interesting question if one could extend this result to perturbed strictly convex single-site potentials. A direct transfer of the argument of Menz fails because of the iterative structure of the proof of Theorem 1.6.

The remaining part of this article is organized as follows. In Section 2.1 we prove the main result. The auxiliary results of Section 2.1 are proved in Section 2.2. There is one exception: The convexification of the single-site potential by iterated renormalization (see Theorem 2.6) is proved in Section 3. In the short Appendix we state the standard criteria for the SG and the LSI.

Adapted two-scale approach

The proof of Theorem 1.6 is based on an adaptation of the two-scale approach of GORV . We start with introducing the concept of coarse-graining of pairs. We recommend reading GORV , Chapter 2.1, as a guideline.

Due to the coarse-graining operator PP, we can decompose the canonical ensemble μN,m\mu_{N,m} into

where μˉ:=P#μN,m\bar{\mu}:=P_{\#}\mu_{N,m} denotes the push forward of the Gibbs measure μ\mu under PP and μ(dx∣y)\mu(dx|y) is the conditional measure of xx given Px=yPx=y. The last equation has to be understood in a weak sense; that is, for any test function ξ\xi

Now, we are able to state the first ingredient of the proof of Theorem 1.6.

Assume that the single-site potential ψ\psi is perturbed strictly convex in the sense of (6). If the marginal μˉ\bar{\mu} satisfies the LSI with constant ϱ1>0\varrho_{1}>0 uniformly in the system size NN and the mean spin mm, then the canonical ensemble μN,m\mu_{N,m} also satisfies the LSI with constant ϱ2>0\varrho_{2}>0 uniformly in the system size NN and the mean spin mm.

The proof of this statement is given in Section 2.2. Due to the last proposition it suffices to deduce the LSI for the marginal μˉ\bar{\mu}. Hence, let us have a closer look at the structure of μˉ\bar{\mu}. We will characterize the Hamiltonian of the marginal μˉ\bar{\mu} with the help of the renormalization operator R\mathcal{R}, which is introduced as follows.

The renormalized single-site potential Rψ\mathcal{R}\psi can be interpreted in the following way: A change of variables (cf. Evans , Section 3.3.3) and the invariance of the Hausdorff measure under translation yield the identity

Therefore, the renormalized single-site potential Rψ\mathcal{R}\psi describes the free energy of two independent spins X1X_{1} and X2X_{2} [identically distributed asZ−1exp⁡(−ψ)Z^{-1}\exp(-\psi)] conditioned on a fixed mean value 12(X1+X2)=y\frac{1}{2}(X_{1}+X_{2})=y.

Assume that the single-site potential ψ\psi is perturbed strictly convex in the sense of (6). Then the renormalized Hamiltonian Rψ\mathcal{R}\psi is also perturbed strictly convex in the sense of (6).

Direct calculation using the coarea formula (cf. Evans , Section 3.4.2) reveals the following structure of the marginal μˉ\bar{\mu}.

Let ψ\psi be a perturbed strictly convex single-site potential in the sense of (6). Then there is an integer M0M_{0} such that for all M≥M0M\geq M_{0} the MM-times renormalized single-site potential RMψ\mathcal{R}^{M}\psi is uniformly strictly convex independently of the system size NN and the mean spin mm.

We conclude this section by giving some remarks and pointing out the central tools needed for the proof of the auxiliary results. The next remark shows how Theorem 1.6 is verified in the case of an arbitrary number NN of sites.

Note that an arbitrary number of sites NN can be written as

which is independent on NN and mm. Therefore, an iterated application of the hierarchic criterion of the LSI (cf. Proposition 2.1) yields Theorem 1.6.

It is a natural question whether this approach can be applied to the case of inhomogeneous single-site potentials. In this case, the single-site potentials are allowed to depend on the sites; that is, the Hamiltonian has the form H=∑i=1NψiH=\sum_{i=1}^{N}\psi_{i} where each ψi\psi_{i} is a perturbed strictly-convex potential. In principle, we believe that our approach can be adapted to this situation even if not in a straightforward way. The reason is that only one step of the proof of Theorem 1.6 has to be adapted: It is the convexification of the single-site potentials by iterated renormalization (see Theorem 2.6).

Let us make a comment on the proof of Theorem 2.6, which is stated in Section 3. Starting point for the proof is the observation that the MM-times renormalized single-site potential RMψ\mathcal{R}^{M}\psi corresponds to the coarse-grained Hamiltonian related to coarse-graining with block size 2M2^{M}; cf. GORV .

Because the last statement is verified by a straightforward application of the area and coarea formula, we omit the proof. In Lemma 2.9 one could easily determine the exact value of the constant C(2M)C(2^{M}). However, the exact value is not important because we are only interested in the convexity of RMψ\mathcal{R}^{M}\psi. In GORV , the convexification of HˉK\bar{H}_{K} was deduced from a local Cramér theorem; cf. GORV , Proposition 31. For the proof of Theorem 2.6 we follow the same strategy generalizing the argument to perturbed strictly convex single-site potentials ψ\psi.

Now, we make some comments on the proof of Proposition 2.1 and Lemma 2.4, which are stated in Section 2.2. One of the limiting factors in the proof of Theorem 1.5 is the application of a classical covariance estimate; cf. GORV , Lemma 22. In our framework this estimate can be formulated as:

In GORV , the last estimate was applied to the function g=ψ′g=\psi^{\prime}. Note that the function ∣g′(x)∣=∣ψ′′(x)∣|g^{\prime}(x)|=|\psi^{\prime\prime}(x)| is only bounded in the case of a perturbed quadratic single-site potential ψ\psi. The main new ingredient for the proof of the hierarchic criterion for the LSI (cf. Proposition 2.1) and the invariance principle (cf. Lemma 2.4) is an asymmetric Brascamp–Lieb inequality, which does not exhibit this restriction.

where osc⁡δψ:=sup⁡xδψ(x)−inf⁡xδψ(x)\operatorname{osc}\delta\psi:=\sup_{x}\delta\psi(x)-\inf_{x}\delta\psi(x).

We start by deriving the following integral representation of the covariance of μ\mu:

where the nonnegative kernel Kμ(x,y)K_{\mu}(x,y) is given by

and Mμ(x):=μ((−∞,x))M_{\mu}(x):=\mu((-\infty,x)) so that (1−Mμ)(x)=μ((x,∞))(1-M_{\mu})(x)=\mu((x,\infty)). Indeed, we start by noting that

where we do not distinguish between the measure μ(dx)\mu(dx) and its Lebesgue density μ(x)\mu(x) in our notation. Using Mμ′(x)=μ(x)M_{\mu}^{\prime}(x)=\mu(x), we can use integration by parts to rewrite each factor in terms of the derivative

where I(x<z)I(x<z) assumes the value 11 if x<zx<z and zero otherwise. Inserting this and the corresponding identity for g(y)g(y) into (12), we obtain

with kernel Kμ(x,y)K_{\mu}(x,y) as desired, given by

We now establish the following identity for the above kernel:

Let us now consider the Gibbs measures ν(dx)\nu(dx) and νc(dx)\nu_{c}(dx), given by

By the integral representation (11) of the covariance we have the estimate

By a straight-forward calculation, we can estimate

Together with a similar estimate for (1−Mν(y)),(1-M_{\nu}(y)), this yields the kernel estimate

Applying this to the covariance estimate from above yields

Using the identity (14) for μ=νc\mu=\nu_{c}, we may easily conclude

For the entertainment of the reader, let us argue how the identity (14) also yields the traditional Brascamp–Lieb inequality in the case H′′>0H^{\prime\prime}>0. Indeed, by the symmetry of the kernel Kμ(x,y)K_{\mu}(x,y), identity (14) yields, for all xx and yy,

The integral representation of the covariance (11) yields

Then a combination of Hölder’s inequality and the identity (15) for the kernel Kμ(x,y)K_{\mu}(x,y) yields the Brascamp–Lieb inequality,

2 Proof of auxiliary results

In this section we outline the proof of Proposition 2.1 and Lemma 2.4. We start with Proposition 2.1, which is the hierarchic criterion for the LSI. Unfortunately, we cannot directly apply the two-scale criterion of GORV , Theorem 3. The reason is that the number

which measures the interaction between the microscopic and macroscopic scales, can be infinite for a perturbed strictly convex single-site potential ψ\psi. However, we follow the proof of GORV , Theorem 3, with only one major difference: Instead of applying the classical covariance estimate (cf. Lemma 2.10), we apply the asymmetric Brascamp–Lieb inequality; cf. Lemma 2.11. Let us assume for the rest of this section that the single-site potential ψ\psi is perturbed strictly convex in the sense of (6).

For convenience, we set X:=XN,mX:=X_{N,m} and Y:=XN/2,mY:=X_{{N}/{2},m}. We choose on XX and YY the standard Euclidean structure given by

The coarse-graining operator P\dvtxX→YP\dvtx X\to Y given by (7) satisfies the identity

where Pt\dvtxY→XP^{t}\dvtx Y\to X is the adjoint operator of PP. Note that our PtP^{t} differs from the PtP^{t} of GORV , because the Euclidean structure on Y differs from the Euclidean structure used in GORV by a factor. The last identity yields that 2PtP2P^{t}P is the orthogonal projection of XX to im⁡Pt\operatorname{im}P^{t}. Hence, one can decompose XX into the orthogonal sum of microscopic fluctuations and macroscopic variables according to

We apply this decomposition to the gradient ∇f\nabla f of a smooth function ff on XX. The gradient ∇f\nabla f is decomposed into a macroscopic gradient and a fluctuation gradient satisfying

The conditional measure μ(dx∣y)\mu(dx|y) given by (8) satisfies the LSI with constant ϱ>0\varrho>0 uniformly in the system size NN, the macroscopic profile yy and the mean spin mm. More precisely, for any nonnegative function ff

Proof of Lemma 2.12 Observe that the conditional measure μ(dx∣y)\mu(dx|y) has a product structure: We decompose {Px=y}\{Px=y\} into a product of Euclidean spaces. Namely for

It follows from the coarea formula (cf. Evans , Section 3.4.2) that

where we make use of the notation introduced in (2). Because the single-site potential ψ\psi is perturbed strictly convex in the sense of (6), a combination of the Bakry–Émery criterion (cf. Theorem .3) and the Holley–Stroock criterion (cf. Theorem .2) yield that the measure μ2,m(dx1,dx2)\mu_{2,m}(dx_{1},dx_{2}) satisfies the LSI with constant ϱ>0\varrho>0 uniformly in mm. Then the tensorization principle (cf. Theorem .1) implies the desired statement.

For convenience, let us introduce the following notation: Let ff be an arbitrary function. Then its conditional expectation fˉ\bar{f} is defined by

The second main ingredient of the proof of Proposition 2.1 is the following proposition, which is the analog statement of GORV , Proposition 20.

Assume that the marginal μˉ(dy)\bar{\mu}(dy) given by (8) satisfies the LSI uniformly in the system size NN and the mean spin mm. Then for any nonnegative function ff,

uniformly in the macroscopic profile yy and the system size NN.

Before we verify Proposition 2.13, let us show how it can be used in the proof of Proposition 2.1. {pf*}Proof of Proposition 2.1 Using Lemma 2.12 and Proposition 2.13 from above, the argument is exactly the same as in the proof of GORV , Theorem 3:

Let ϕ\phi denote the function ϕ(x):=xlog⁡x\phi(x):=x\log x. The additive property of the entropy implies

An application of Lemma 2.12 yields the estimate

By assumption the marginal μˉ\bar{\mu} satisfies the LSI with constant λ>0\lambda>0. Together with Proposition 2.13 this yields the estimate

A combination of the last three formulas and the observations (8) and (2.2) yield

uniformly in the system size NN and the mean spin mm.

Because the hierarchic criterion for the LSI is an important ingredient in the proof of the main result, we outline the proof of Proposition 2.13 in full detail. We follow the proof of GORV , Proposition 20, which is based on two lemmas. We directly take over the first lemma (cf. GORV , Lemma 21), which in our notation becomes:

For any function ff on XX and any y∈Yy\in Y, it holds

The notational difference compared to GORV , Lemma 21, is based on our choice of the Euclidean structure on Y=XN/2,mY=X_{{N}/{2},m}. Compared to the notation in Lemma 21 of GORV , we have

Hence we omit the proof, which is a straightforward calculation.

The more interesting ingredient of the proof of GORV , Proposition 20, is the estimate (see GORV , (42), (43))

In GORV , the last estimate is deduced by direct calculation from the standard covariance estimate given by Lemma 2.10. In contrast to GORV we cannot use this estimate because the constant κ\kappa given by (17) may be infinite for a perturbed strictly convex single-site potential ψ\psi. We avoid this problem by applying the more robust asymmetric Brascamp–Lieb inequality given by Lemma 2.11. Our substitute for the last estimate is:

uniformly in the system size NN, the macroscopic profile yy and the mean spin mm.

it follows form the definition (7) of PP that for any xx,

By successively using Lemma 2.14 and Jensen’s inequality (with the convex function (a,b)↦∣b∣2/a(a,b)\mapsto|b|^{2}/a), we have

On the first term on the r.h.s. we apply the estimate (20). On the second term we apply Lemma 2.16, which yields the desired estimate.

Now, we prove Lemma 2.16, which also represents one of the main differences compared to the two-scale approach of GORV . The main ingredients are the product structure (19) of μ(dx∣y)\mu(dx|y) and the asymmetric Brascamp–Lieb inequality; cf. Lemma 2.11. {pf*}Proof of Lemma 2.16 We have to estimate the covariance

Therefore, let us consider for j∈{1,…,N2}j\in\{1,\ldots,\frac{N}{2}\} the term cov⁡μ(dx∣y)(f,(2P∇H)j)\operatorname{cov}_{\mu(dx|y)}(f,(2P\nabla H)_{j}). Note that the function

only depends of the variables x2j−1x_{2j-1} and x2jx_{2j}. Hence, the product structure (19) of μ(dx∣y)\mu(dx|y) yields the identity

As we will show below, we obtain, by using the asymmetric Brascamp–Lieb inequality of Lemma 2.11 and the Csiszár–Kullback–Pinsker inequality, the estimate

uniformly in jj and yjy_{j}. Therefore, a combination of identity (2.2), the last estimate and Hölder’s inequality yield

which implies the desired estimate by the identity (21).

It is only left to deduce estimate (2.2). We assume w.l.o.g. j=1j=1. Recall the splitting ψ=ψc+δψ\psi=\psi_{c}+\delta\psi given by (6). We use the bound on ∣δψ′∣|\delta\psi^{\prime}| to estimate

A reparametrization of the one-dimensional Hausdorff measure implies

for any measurable function ξ\xi. We may assume w.l.o.g. that f(x)=f(x1,x2)f(x)=f(x_{1},x_{2}) just depends on the variables x1x_{1} and x2x_{2}. Hence for

an application of the asymmetric Brascamp–Lieb inequality (cf. Lemma 2.11) yields

From the last inequality and from (26) follows the estimate

We turn to the second term on the r.h.s. of (2.2). For convenience, let us write fˉ(y1):=∫fμ2,y1(dx1,dx2)\bar{f}(y_{1}):=\int f\mu_{2,y_{1}}(dx_{1},dx_{2}). An application of the well-known Csiszár–Kullback–Pinsker inequality (cf. Csi , Kul ) yields

An application of the LSI for the measure μ2,y1(dx1,dx2)\mu_{2,y_{1}}(dx_{1},dx_{2}) implies (cf. proof of Lemma 2.12)

A combination of (2.2), (2.2), and the last inequality yield the estimate (2.2).

We turn to the proof of Lemma 2.4. Again, the main ingredient of the proof is the asymmetric Brascamp–Lieb inequality. {pf*}Proof of Lemma 2.4 We define

Now, we show that the splitting Rψ=ψ‾c+δψ‾\mathcal{R}\psi=\overline{\psi}_{c}+\overline{\delta\psi} satisfies the conditions given by (6). Using the strict convexity of ψc\psi_{c} it follows by a standard argument based on the Brascamp–Lieb inequality (cf. BL and (2.1)) that the first condition is preserved, that is,

We turn to the perturbation δψ‾\overline{\delta\psi}. Analogously to the measure ν(dz∣m)\nu(dz|m) given by (25), we introduce the measure νc(dz∣m)\nu_{c}(dz|m) via the density

Direct calculation using the bound ∣δψ∣≲1|\delta\psi|\lesssim 1 yields

We turn to the first derivative of δψ‾\overline{\delta\psi}. A direct calculation based on the definition of δψ‾\overline{\delta\psi} yields

For s∈s\in we define the measure νs(dz)\nu^{s}(dz) by the probability density

Note that νs\nu^{s} interpolates between ν0=νc\nu^{0}=\nu_{c} and ν1=ν\nu^{1}=\nu. By the mean-value theorem there is s∈s\in such that

The first term on the r.h.s. is controlled by the assumption ∣δψ′∣≲1|\delta\psi^{\prime}|\lesssim 1. We turn to the estimation of the first covariance term. An application of the asymmetric Brascamp–Lieb inequality of Lemma 2.11 and ∣δψ∣+∣δψ′∣≲1|\delta\psi|+|\delta\psi^{\prime}|\lesssim 1 yields the estimate

The second covariance term can be estimated by using ∣δψ∣+∣δψ′∣≲1|\delta\psi|+|\delta\psi^{\prime}|\lesssim 1. Summing up, we have deduced the desired estimate ∣δψ‾′∣≲1|\overline{\delta\psi}^{\prime}|\lesssim 1.

Convexification by iterated renormalization

In this section we prove Theorem 2.6 that states the convexification of a perturbed strictly convex single-site potential ψ\psi by iterated renormalization. The proof relies on a local Cramér theorem and some auxiliary results. The proof of Theorem 2.6 is given in Section 3.1. The proofs of the auxiliary results are given in Section 3.2.

Let us consider the coarse-grained Hamiltonian HˉK\bar{H}_{K} given by (10). In view of Lemma 2.9, it suffices to show the strict convexity of HˉK\bar{H}_{K} for large K≫1K\gg 1. The strategy is the same as in GORV , Proposition 31. Let φ\varphi denote the Cramér transform of ψ\psi, namely

Because φ\varphi is the Legendre transform of the strictly convex function

From basic properties of the Legendre transform, it follows that σ\sigma is determined by the equation

The starting point of the proof of the convexification of the coarse-grained Hamiltonian HˉK(m)\bar{H}_{K}(m) is the explicit representation

where XiX_{i} are KK real-valued independent random variables identically distributed according to

Let ψ(x)\psi(x) be a smooth function that is increasing sufficiently fast as ∣x∣↑∞|x|\uparrow\infty for all subsequent integrals to exist. Note that the probability measure μσ\mu^{\sigma} defined by (32) depends on the field strength σ\sigma. We introduce its mean mm and variance s2s^{2}

We assume that uniformly in the field strength σ\sigma, the probability measure μσ\mu^{\sigma} has its standard deviation ss as unique length scale in the sense that

Consider KK independent random variables X1,…,XKX_{1},\ldots,X_{K} identically distributed according to μσ\mu^{\sigma}. Let gK,σg_{K,\sigma} denote the Lebesgue density of the distribution of the normalized sum 1K∑i=1KXi−ms\frac{1}{\sqrt{K}}\sum_{i=1}^{K}\frac{X_{i}-m}{s}.

Then gK,σ(0)g_{K,\sigma}(0) converges for K↑∞K\uparrow\infty to the corresponding value for the normalized Gaussian. This convergence is uniform in mm, of order 1K\frac{1}{\sqrt{K}}, and C2C^{2} in σ\sigma:

Let us comment a bit on this result: Quantitative versions of the central limit theorem like (36) are abundant in the literature; see, for instance, Feller , Chapter XVI, KipLand , Appendix 2, GPV , Section 3, and LPY , page 752 and Section 5. In his work on the spectral gap, Caputo appeals even to a finer estimate that makes the first terms in an error expansion in K−1/2K^{-{{1}/{2}}} explicit Cap , Theorem 2.1. The coefficients of the higher order terms are expressed in terms of moments of μσ\mu^{\sigma}. However, following GORV , Proposition 31, for our two-scale argument we need pointwise control of the Lebesgue density gK,σg_{K,\sigma} [in form of gK,σ(0)g_{K,\sigma}(0)] and, in addition, control of derivatives of gK,σg_{K,\sigma} w.r.t. the field parameter σ\sigma; cf. (37), (38). Note that the derivative ddσ\frac{d}{d\sigma} has units of length (because σ\sigma, which multiplies xx in the Hamiltonian [cf. (32)] has units of inverse length) so that 1sddσ\frac{1}{s}\frac{d}{d\sigma} is the properly nondimensionalized derivative. Pointwise control means that control of the moments [cf. (34)] is not sufficient. One also needs to know that μσ\mu^{\sigma} has no fine structure on scales much smaller than ss. This property is ensured the upper bound (35).

As opposed to GORV , Proposition 31, the Hamiltonian ψ\psi we want Proposition 3.1 to apply is not a perturbation of the quadratic 12x2\frac{1}{2}x^{2}, but of a general, strictly convex potential ψ\psi. As a consequence, the variance s2s^{2} can be a strongly varying function of the field strength σ\sigma. Nevertheless, Lemma 3.2 from below shows that every element μσ\mu^{\sigma} in the family of measures is characterized by the single length scale ss, uniformly in σ\sigma in the sense of (34) and (35). For the verification of (34) in Lemma 3.2, one could take over the argument of Cap , Lemma 2.2, that relies on a result by Bobkov Bob stating that the SG constant ϱ\varrho of the measure μσ\mu^{\sigma} can be estimated by its variance, that is, ϱ≳1s2\varrho\gtrsim\frac{1}{s^{2}}. However, we provide a self-contained argument for the verification of (34) and (35) in Lemma 3.2 just using basic calculus of one variable. The merit of Proposition 3.1 consists in providing a version of the central limit theorem that is C2C^{2} in the field strength σ\sigma even if the variance s2s^{2} varies strongly with σ\sigma.

Assume that the single-site potential ψ\psi is perturbed strictly convex in the sense of (6). Then s≲1s\lesssim 1 uniformly in mm, and conditions (34) and (35) of Proposition 3.1 are satisfied.

Using Proposition 3.1, Lemma 3.2, and the Cramér representation (31) we could easily deduce a local Cramér theorem (cf. GORV , Proposition 31) for general perturbed strictly convex potentials ψ\psi. However, because we are just interested in the convexification of HˉK\bar{H}_{K}, we just consider the convergence of the second derivatives of φ\varphi and HˉK\bar{H}_{K}.

where s2s^{2} is defined as in Proposition 3.1.

We start with some formulas on the derivatives of φ\varphi. Differentiation of identity (29) yields

A direct calculation reveals that [see (61) below]

where s2s^{2} is defined as in Proposition 3.1. Hence, a second differentiation of φ\varphi yields the identity

if K≥K0K\geq K_{0} for some large K0K_{0}. The statement follows from the uniform bound s≲1s\lesssim 1 provided by Lemma 3.2.

2 Proof of the local Cramér theorem and of the auxiliary results

In this section we prove the auxiliary statements of the last subsection. Before turning to the proof of Proposition 3.1 we sketch the strategy. For convenience we introduce the notation

The definition of gK,σg_{K,\sigma} (cf. Proposition 3.1) suggests to introduce the shifted and rescaled variable

We note that by (33) the first and second moment in x^\hat{x} are normalized

Proposition 3.1 is a version of the central limit theorem that, like most others, is best proved with help of the Fourier transform. Indeed, since the random variables X^1:=X1−ms,…,X^K:=XK−ms\hat{X}_{1}:=\frac{X_{1}-m}{s},\ldots,\hat{X}_{K}:=\frac{X_{K}-m}{s} in the statement of Proposition 3.1 are independent and identically distributed, the distribution of their sum is the KK-fold convolution of the distribution of X^1\hat{X}_{1}. Therefore, the Fourier transform of the distribution of the ∑n=1KX^n\sum_{n=1}^{K}\hat{X}_{n} is the KKth power of the Fourier transform of the distribution of X^\hat{X}. The latter is given by

where ξ^\hat{\xi} denotes the variable dual to x^\hat{x}. Hence, the Fourier transform of the distribution of the normalized sum 1K∑n=1KX^K\frac{1}{\sqrt{K}}\sum_{n=1}^{K}\hat{X}_{K} is given by ⟨exp⁡(ix^1Kξ^)⟩K\langle\exp(i\hat{x}\frac{1}{\sqrt{K}}\hat{\xi})\rangle^{K}. Applying the inverse Fourier transform, we obtain the representation

In order to make use of formula (44), we need estimates on ⟨exp⁡(ix^ξ^)⟩\langle\exp(i\hat{x}\hat{\xi})\rangle. Because of

the moment bounds (43) translate into control of ⟨exp⁡(ix^ξ^)⟩\langle\exp(i\hat{x}\hat{\xi})\rangle for ∣ξ^∣≪1|\hat{\xi}|\ll 1. Together with the normalization (42), we obtain, in particular,

We will use the latter in the following form: There exists a complex-valued function h(ξ^)h(\hat{\xi}) such that for ∣ξ^∣≪1|\hat{\xi}|\ll 1,

This estimate, showing that the Fourier transform of the normalized probability ⟨⋅⟩\langle\cdot\rangle is close for ∣ξ^∣≪1|\hat{\xi}|\ll 1 to the Fourier transform of the normalized Gaussian, is at the core of most proofs of the central limit theorem.

Estimate (46) provides good control over ⟨exp⁡(ix^ξ^)⟩\langle\exp(i\hat{x}\hat{\xi})\rangle for ∣ξ^∣≪1|\hat{\xi}|\ll 1. Another key ingredient is uniform decay for ∣ξ^∣≫1|\hat{\xi}|\gg 1. In our new variables, (35) takes on the form

As usual in central limit theorems, we also need control of the characteristic function for intermediate values of ∣ξ^∣|\hat{\xi}|. This can be inferred from (43) and (47) by a soft argument (in particular, it does not require the more intricate argument for Cap , (2.10), from Cap , Lemma 2.5):

Under the assumptions of Proposition 3.1 and for any δ>0\delta>0, there exists λ<1\lambda<1 such that for all σ\sigma,

So far, the strategy is standard; now comes the new ingredient: In view of formula (44), in order to control σ\sigma-derivatives of gK,σ(0)g_{K,\sigma}(0), we need to control 1sddσ⟨exp⁡(ix^ξ^)⟩\frac{1}{s}\frac{d}{d\sigma}\langle\exp(i\hat{x}\hat{\xi})\rangle. Relying on the identities

that will be established in the proof of Lemma 3.5 below, we see that the estimate again follows from the moment control (43). Lemma 3.5 is the only new element of our analysis.

Under the assumptions of Proposition 3.1 we have

Before we deduce Proposition 3.1, we prove Lemma 3.4 and Lemma 3.5.

Proof of Lemma 3.4 In view of (43) and (47), it suffices to show: For any C<∞C<\infty and δ>0\delta>0 there exists λ<1\lambda<1 with the following property: Suppose ⟨⋅⟩\langle\cdot\rangle is a probability measure (in x^\hat{x}) such that

We give an indirect argument for this statement and thus assume that there is a sequence {⟨⋅⟩ν}\{\langle\cdot\rangle_{\nu}\} of probability measures satisfying (52) and (53) and a sequence {ξ^ν}\{\hat{\xi}_{\nu}\} of numbers in [δ,1δ][\delta,\frac{1}{\delta}] such that

In view of (52), after passage to a subsequence, we may assume that there exists a probability measure ⟨⋅⟩∞\langle\cdot\rangle_{\infty} and a number ξ^∞>0\hat{\xi}_{\infty}>0 such that

Since ∣exp⁡(ix^ξ^ν)−exp⁡(ix^ξ^∞)∣≤∣x^∣∣ξ^ν−ξ^∞∣|\exp(i\hat{x}\hat{\xi}_{\nu})-\exp(i\hat{x}\hat{\xi}_{\infty})|\leq|\hat{x}||\hat{\xi}_{\nu}-\hat{\xi}_{\infty}|, we obtain the following from (52), (55) and (56):

On the other hand, (53) is preserved under (55) so that we have, in particular,

We claim that (57) and (58) contradict each other. Indeed, since x^↦\breakexp⁡(ix^ξ^∞)\hat{x}\mapsto\break\exp(i\hat{x}\hat{\xi}_{\infty}) is S1S^{1}-valued, it follows from (57) that there is a fixed ζ∈S1\zeta\in S^{1} such that

which, in view of ξ^∞≠0\hat{\xi}_{\infty}\not=0 and thus ∣nξ^∞∣↑∞|n\hat{\xi}_{\infty}|\uparrow\infty as n↑∞n\uparrow\infty, contradicts (58).

Proof of Lemma 3.5 We restrict our attention to estimate (51); estimate (50) is easier and can be derived by the same arguments. We start with the identities (48) and (49). Deriving (40) w.r.t. σ\sigma yields

In view of definition (41), the latter turns into (48).

We now turn to identity (49) and note that, in view of definitions (33) and (41), the identity (60) yields, in particular,

We now combine formulas (48) and (49) to express derivatives of ⟨f(x^)⟩\langle f(\hat{x})\rangle. We start with the first derivative,

[As a consistency check we note that 1sddσ⟨f(x^)⟩=(\ref18)−⟨(ddx^−x^)f⟩−12⟨x^3⟩⟨x^dfdx^⟩\frac{1}{s}\frac{d}{d\sigma}\langle f(\hat{x})\rangle\stackrel{{\scriptstyle\scriptsize{(\ref{18})}}}{{=}}-\langle(\frac{d}{d\hat{x}}-\hat{x})f\rangle-\frac{1}{2}\langle\hat{x}^{3}\rangle\langle\hat{x}\frac{df}{d\hat{x}}\rangle vanishes if ψ\psi is quadratic since then the distribution of x^\hat{x} under ⟨⋅⟩\langle\cdot\rangle is the normalized Gaussian so that both ⟨(ddx^−x^)f⟩=0\langle(\frac{d}{d\hat{x}}-\hat{x})f\rangle=0 and ⟨x^3⟩=0\langle\hat{x}^{3}\rangle=0.]

Iterating this formula, we obtain for the second derivative,

This formula and the normalization (42) yield that (1sddσ)2⟨exp⁡(iξ^x^)⟩(\frac{1}{s}\frac{d}{d\sigma})^{2}\langle\exp(i\hat{\xi}\hat{x})\rangle vanishes to second order in ξ^\hat{\xi}. More precisely, for k∈{0,1,2}k\in\{0,1,2\}

Therefore, we consider the third derivative w.r.t. ξ^\hat{\xi} given by (3.2). For this purpose we apply the formula for (1sddσ)2⟨f(x^)⟩(\frac{1}{s}\frac{d}{d\sigma})^{2}\langle f(\hat{x})\rangle from above to the function

Using the abbreviation e:=exp⁡(iξ^x^)e:=\exp(i\hat{\xi}\hat{x}), we obtain

From this formula and the moment estimates (43), we obtain the estimate

In combination with (65), this estimate yields (51).

Proof of Proposition 3.1 We focus on (36) and (38). The intermediate estimate (37) can be established as (38).

We start with (36). Fix a δ>0\delta>0 so small such that the expansion (46) of ⟨exp⁡(ix^ξ^)⟩\langle\exp(i\hat{x}\hat{\xi})\rangle holds for ∣ξ^∣≤δ|\hat{\xi}|\leq\delta. We split the integral representation (44) accordingly:

We consider the first term II on the r.h.s. of (3.2), which will turn out to be of leading order. Since δ\delta is so small that (46) holds, we may rewrite it as

We note that for ∣1Kξ^∣≤δ|\frac{1}{\sqrt{K}}\hat{\xi}|\leq\delta we have by (46),

Inserting this estimate into (3.2) we obtain

since ∫{∣(1/K)ξ^∣>δ}exp⁡(−12ξ^2) dξ^\int_{\{|({1}/{\sqrt{K}})\hat{\xi}|>\delta\}}\exp(-\frac{1}{2}\hat{\xi}^{2})\,d\hat{\xi} is exponentially small in KK.

We now address the second term II\mathit{II} on the r.h.s. of (3.2); on the integrand we use Lemma 3.4 (on K−2K-2 of the KK factors) and (47) (on the remaining 22 factors).

It follows that the second term II\mathit{II} on the r.h.s. of (3.2) is exponentially small and thus of higher order:

We now turn to (38). We take the second σ\sigma-derivative of the integral representation (44),

As for (36), we split the integral representation (3.2) according to δ\delta:

On the integrand of the second r.h.s. term in (3.2) we use Lemma 3.4 (on K−12K-12 of the K−2K-2 factors) and (47) (on the remaining 1010 factors):

Hence, we see that this second term in (3.2) is exponentially small and thus of higher order:

For the proof of Lemma 3.2 we need the following auxiliary statement, based on elementary calculus.

Let MM denote the maximum of the density of ν\nu, that is,

Proof of Lemma 3.6 We may assume w.l.o.g. that

and M:=sup⁡xexp⁡(−ψ(x))M:=\sup_{x}\exp(-\psi(x)) is attained at x=0x=0, which means

We start with an analysis of the convex single-site potential ψ\psi. We first argue that

Indeed in view of the monotonicity (74), we have

We now argue that for ∣x∣≥eM|x|\geq\frac{e}{M},

W.l.o.g. we may restrict ourselves to x≥eMx\geq\frac{e}{M}. By convexity of ψ\psi, we have

The convexity of ψ\psi, the last estimate and (75) yield for x≥eMx\geq\frac{e}{M}, as desired,

We finished the analysis on ψ\psi and turn to the verification of the estimate of Lemma 3.6. We split the integral according to

A similar estimate for the integral ∫−∞0∣x∣kexp⁡(−ψ(x)) dx\int_{-\infty}^{0}|x|^{k}\exp(-\psi(x))\,dx follows from the same argument by symmetry. We split the integral

The first integral on the r.h.s. can be estimated as

For the estimation of the second integral, we apply (76), which yields, by the change of variables Me(x−eM)=x^\frac{M}{e}(x-\frac{e}{M})=\hat{x},

Equipped with Lemma 3.6, we are able to give an elementary proof of Lemma 3.2:

Proof of Lemma 3.2 We argue that s≲1s\lesssim 1. Because ψ\psi is a bounded perturbation of a uniformly strictly convex function, the measure μσ\mu^{\sigma} given by (32) satisfies the SG uniformly in σ\sigma. This implies, in particular,

for some convex function ψ^\hat{\psi}, which is normalized in the sense that

An application of Lemma 3.6 yields the estimate

where MM is given by M:=max⁡x^exp⁡(−ψ^(x^))M:=\max_{\hat{x}}\exp(-\hat{\psi}(\hat{x})). Now, we argue that due to the normalization of ψ^\hat{\psi}, we have

for some universal constant C>0C>0. The latter verifies the desired estimate (34). Indeed normalization (78) implies

Hence, there exists an x^0∈(−2,2)\hat{x}_{0}\in(-2,2) such that exp⁡(−ψ(x^0))≥38\exp(-\psi(\hat{x}_{0}))\geq\frac{3}{8}, which yields

Let us turn to the statement (35) of Proposition 3.1. Writing

For convenience, we introduce the Hamiltonian ψ^(x)=−σx+ψc(x)\hat{\psi}(x)=-\sigma x+\psi_{c}(x) and assume w.l.o.g. that ∫exp⁡(−ψ^(x)) dx=1\int\exp(-\hat{\psi}(x))\,dx=1. The splitting ψ=ψc+δψ\psi=\psi_{c}+\delta\psi with ∣δψ∣|\delta\psi|, ∣δψ′∣≲1|\delta\psi^{\prime}|\lesssim 1 and definition (28) of φ∗\varphi^{*} yield the estimate

where ss is defined as in Proposition 3.1. Because s≲1s\lesssim 1 by (77), we only have to consider the first term of the r.h.s. of the last inequality. We argue that for

For the proof of the last statement, we only need the fact that ψ^(x)=−σx+ψc(x)\hat{\psi}(x)=-\sigma x+\psi_{c}(x) is convex. W.l.o.g. we may assume that MM is attained at x=0x=0, which means M=exp⁡(−ψ^(0))M=\exp(-\hat{\psi}(0)). It follows from convexity of ψ^\hat{\psi} that

Therefore, Lemma 3.6 applied to k=2k=2 and ψ\psi replaced by ψ^\hat{\psi} yields

Before we turn to the proof of Lemma 3.3, we will deduce the following auxiliary result.

Assume that (34) of Proposition 3.1 is satisfied. Then, using the notation of Proposition 3.1, it holds that

Proof of Lemma 3.7 We start with restating some basic identities [cf. (61) and (62)]: It holds that

Let us consider (i): It follows from (81) and (82) that

which yields by assumption (34) of Proposition 3.1 the estimate

The statement of (i) is a direct consequence of the last estimate and the identity

We turn to statement (ii): Differentiating the last identity yields

The estimation of the first term on the r.h.s. follows from the estimates

which we have deduced in the first step of the proof. We turn to the estimation of the second term. A direct calculation using (81) yields the identity

Considering the first term on the r.h.s., we get from the identities (81) and (83), and the assumption (34) of Proposition 3.1 that

Before we consider the second term of the r.h.s. of (3.2), we establish the following estimate:

Indeed, direct calculation using (81) and (82) yields

The last identity yields (85) using the assumption (34) of Proposition 3.1. Using (85) and (82), we can estimate the second term of the r.h.s. of (3.2) as

By applying assumption (34) of Proposition 3.1 this yields

Proof of Lemma 3.3 Recall the representation (31), that is,

where XiX_{i} are real-valued independent random variables identically distributed according to μσ\mu^{\sigma}; cf. (32). Let gK,σg_{K,\sigma} denote the density of the normalized random variable

where ss is given by (33). Then the densities are related by

In order to deduce the desired estimate, it thus suffices to show

The first estimate follows directly from the identity

We turn to the second estimate. The identity

and (36) yield for large KK the estimate

The estimation of the first term on the r.h.s. follows from estimate (37) of Proposition 3.1 and the identity

which is a direct consequence of (61). Let us consider the second term. The identity

Now, estimates (37) and (38) of Proposition 3.1 and Lemma 3.7 yield the desired estimate (87).

Appendix: Standard criteria for the SG and the LSI

In this section we quote some standard criteria for the SG and the LSI. For a general introduction to the SG and the LSI we refer to L , R , GZ . Note that even if we only formulate the criteria on the level of the LSI, they also hold on the level of the SG. The first one shows that the LSI is compatible with products; cf., for example, GZ , Theorem 4.4.

Let μ1\mu_{1} and μ2\mu_{2} be probability measures on Euclidean spaces X1X_{1} and X2X_{2}, respectively. If μ1\mu_{1} and μ2\mu_{2} satisfy the LSI with constant ϱ1\varrho_{1} and ϱ2\varrho_{2}, respectively, then the product measure μ1⊗μ2\mu_{1}\otimes\mu_{2} satisfies the LSI with constant min⁡{ϱ1,ϱ2}\min\{\varrho_{1},\varrho_{2}\}.

The next criterion shows how the LSI constant behaves under perturbations; cf. HS , page 1184.

Because of its perturbative nature, the Holley–Stroock criterion is not well adapted for high dimensions. For the proof of the last statement, we refer the reader to L , Lemma 1.2. Now, we state the Bakry–Émery criterion, which connects the convexity of the Hamiltonian to the LSI constant; cf. BE , Proposition 3 and Corollary 2, or L , Corollary 1.6.

Let dμ:=Z−1exp⁡(−H(x)) dxd\mu:=Z^{-1}\exp(-H(x))\,dx be a probability measure on a Euclidean spaces XX. If there is a constant ϱ>0\varrho>0 such that in the sense of quadratic forms

uniformly in x∈Xx\in X, then μ\mu satisfies the LSI with constant ϱ\varrho.

A proof using semi-group methods can be found in L , Corollary 1.6. There is also a heuristic interpretation of the Bakry–Émery criterion on a formal Riemannian structure on the space of probability measures; cf. OV , Section 3.

Acknowledgment

G. Menz and F. Otto thank Franck Barthe, Michel Ledoux and Cedric Villani for discussions on this subject.

References