Scaling limit of the Stein variational gradient descent: the mean field regime

Jianfeng Lu, Yulong Lu, James Nolen

Introduction

In this paper we study the following interacting particle system in Rd\mathbf{R}^{d}:

We refer to each of the NN functions xi(⋅)∈Rdx_{i}(\cdot)\in\mathbf{R}^{d} as a particle. The function K:Rd↦RK:\mathbf{R}^{d}\mapsto\mathbf{R} is a smooth, symmetric, and positive definite kernel. The function V:Rd→RV:\mathbf{R}^{d}\to\mathbf{R} is a smooth potential such that e−V(x)e^{-V(x)} is integrable. More specific assumptions about KK and VV are given below.

We are interested in the macroscopic behavior of the particle system (1) as N→∞N\rightarrow\infty in the framework of mean field limit. Formally this mean field limit is described by the following non-local, nonlinear partial differential equation (PDE):

We aim to make a rigorous connection between (1) and (2). Specifically, we prove global existence and uniqueness of a solution to this initial value problem, for ρ0\rho_{0} in the appropriate regularity class, and we show that the empirical measure

converges as N→∞N\to\infty to the solution of (2), assuming μ0N\mu^{N}_{0} converges to ρ0(x)dx\rho_{0}(x)dx in the appropriate sense. We also want to study the long-time behavior of solutions to the mean field PDE (2). It is easy to see that the probability density

with Z=∫e−V(x)dxZ=\int e^{-V(x)}dx is an invariant solution to (2). Under certain assumptions, we prove that ρ(t,⋅)\rho(t,\cdot) converges weakly to ρ∞\rho_{\infty} as t→+∞t\to+\infty.

Our interest in the particle system (1) is mainly motivated by the recent works by Liu and Wang , where a time-discretized form of (1) was introduced as an algorithm called Stein Variational Gradient Descent (SVGD). The idea of the algorithm is to transport a set of NN particles in Rd\mathbf{R}^{d} so that their empirical measure μN\mu^{N} approximates the target probability measure ρ∞(x)dx=Z−1e−V(x)dx\rho_{\infty}(x)dx=Z^{-1}e^{-V(x)}dx, with an unknown normalization factor ZZ. At discrete times, the particles are updated via the map

where SρS_{\rho} is the so-called Stein operator defined by

In view of (4), this leads to the definition of Stein discrepancy

Therefore, interpreting (4) by (5) and using the fact that ρ∞(x)∝e−V(x)\rho_{\infty}(x)\propto e^{-V(x)}, one sees that the optimal solution of (4) is given by

Putting this optimal velocity back into (3) and letting the step size ε↓0\varepsilon\downarrow 0 gives the evolution (1).

The variational picture described above about the particle system (1) suggests that the mean field limit (2) might also admit a variational structure. Indeed, it has been shown heuristically in that equation (2) can be viewed formally as a gradient flow for the KL-divergence functional

with respect to a generalized optimal transport metric whose definition involves the reproducing kernel Hilbert space with kernel K(x)K(x). This in particular implies that the KL-divergence functional is a Lyapunov functional for the PDE (2), namely

Interpreting an evolutionary PDE as a gradient flow in the space of probability measures with respect to certain Wasserstein metric dates back to the seminar work on Fokker-Planck equation by Jordan, Kinderlehrer and Otto . By now, similar gradient flow structures have been identified for a large family of evolution equations, including porous medium equation , McKean-Vlasov equation , etc. In the present paper, we will not pursue further the rigorous definition and analysis of the gradient flow structure of (2). Instead, we take the system (1) as our starting point and prove its connection to the mean field PDE (2).

2 Relevant Literature

Sampling from a density of the form ρ∞(x)=e−V(x)/Z\rho_{\infty}(x)=e^{-V(x)}/Z without knowing the normalization constant ZZ is a fundamental problem in Bayesian statistics and machine learning. One generic approach that has been tremendously successful in recent years is the Markov chain Monte Carlo (MCMC) methodology based on Metropolis-Hastings mechanism. The general principle of Metropolis-Hastings algorithms is to build an ergodic Markov chain whose invariant measure is the target measure ρ∞\rho_{\infty} by first making candidate samples (proposals), which are then tuned to ensure stationarity via acception/rejection. In practice, one common approach to constructing proposals is by discretizing some stochastic dynamics, such as the following (overdamped) Langevin dynamics:

where BB is a standard Brownian motion in Rd\mathbf{R}^{d}. A vanilla Euler-Maruyama discretization scheme associated to (7) together with Metropolis-Hastings step leads to the famous Metropolis Adjusted Langevin Algorithm (MALA) (whose non-Metropolized version known as unadjusted Langevin algorithm (ULA) ).

One advantageous feature of stochastic dynamics-based sampling methods, e.g. MALA or ULA, is that the dynamics tend to explore high probability regions (around the local minima of VV), while the random noise helps the dynamics to escape outside the basin of attraction and thus promotes its exploration of the entire state space. In contrast to this stochastic sampling approach, (1) may be viewed as a deterministic (albeit coupled) particle system for approximating ρ∞\rho_{\infty}. Qualitatively speaking, the terms in (1) which involve ∇V\nabla V tend to drive particles toward local minima of VV (note however the nonlocal interaction due to the presence of KK). On the other hand, the terms involving ∇K\nabla K are repulsive, forcing the particles to disperse; this is seen in the fact that

where E(x)=1N∑i<jK(xi−xj)E({\bf x})=\frac{1}{N}\sum_{i<j}K(x_{i}-x_{j}) is the interaction energy. Here we assumed that ∇K(0)=0\nabla K(0)=0. This interaction term in SVGD plays a role similar to that of the diffusion term in stochastic-dynamics-based sampling methods. Intuitively, one would expect that the empirical measure μtN\mu^{N}_{t} of the particles {xi(t)}\{x_{i}(t)\} tends to be close to ρ∞\rho_{\infty} in the limit of both large sample size and long time. One of the contributions of this paper is to prove this convergence rigorously.

To compare these two sampling approaches at the PDE level, observe that the probability density for X(t)X(t) defined by (7) solves the linear Fokker-Planck equation

It is well known that under some mild assumption on VV, the solution ρ\rho of (8) converges to the equilibrium distribution ρ∞\rho_{\infty} exponentially fast. On the other hand, if we formally set K(x)=δ0(x)K(x)=\delta_{0}(x), the non-local mean-field equation (2) becomes

which is a non-linear porous medium equation with an additional transport due to ∇V\nabla V. So, compared to (8), the mobility term and the transport term in (9) are small where the density is small. This suggests that the convergence of the solution of (9) towards ρ∞\rho_{\infty} may be slower than that of (8). In this paper, we consider only a fixed kernel KK, but if we scale the kernel KK as KN(⋅)=NβK(Nβ⋅)K_{N}(\cdot)=N^{\beta}K(N^{\beta}\cdot), it is natural to expect the large particle limit of (1) to be governed by (9) instead of (2), if β>0\beta>0 is not too large — rigorous justification of such convergence result is still work in process.

One should also compare (1) with the following more standard deterministic interacting particle system:

It is well-known that under suitable assumption on KK and VV, the mean field limit of (10) is the following McKean-Vlasov equation

The particle system (1) differs from (10) in that the external force added to each particle is non-local, and is defined by averaging the individual forces ∇V(xj)\nabla V(x_{j}) with weights defined by the kernel KK. Interestingly, such non-local external force guarantees that ρ∞\rho_{\infty} is a stationary solution of (2) — this explains the rationale for using the deterministic particle system (1) as an approximation algorithm for sampling ρ∞\rho_{\infty}. On the contrary, ρ∞\rho_{\infty} is not a stationary solution of (11). In fact, if VV or KK is non-convex, the equation (11) may have multiple stationary solutions; see e.g. . We also remark that the nonlocal external force makes the analysis of (1) more challenging than that of (10).

Although sampling via a deterministic particle system is less common, the use of deterministic particles is ubiquitous in numerical approximations of partial differential equations arising in physics and biology. For example, the point vortex method have been proved successful for solving equations in fluid mechanics , and similarly the weighted particle method and the diffusion-velocity method for convection-diffusion and nonlinear-wave equations . For a comprehensive discussion on deterministic particle methods we refer the reader to the recent review paper and references therein. Recently, a blob method was proposed in for an aggregation equation, which is the equation (2) with V=0V=0 and with KK being attractive rather than repulsive. One typical aggregation equation is the so-called Keller-Segel equation . The same blob method was generalized by to a more general class of nonlinear diffusion equations, which has a L2L^{2}-Wasserstein gradient flow structure. A key feature of the blob method considered there is that the particle system preserves a similar gradient flow structure as the diffusion equation, which facilitates the proof of large particle limits. On the contrary, the SVGD dynamics (1) is not a gradient flow. This again makes the analysis of the mean field limit non-trivial.

3 Plan of The Paper

The rest of the paper is organized as follows. In Section 2, we first make several technical assumptions on VV and KK and then state our main results under these assumptions. In Section 3, we prove the existence and uniqueness of weak solutions to the mean field equation Eq. 2 as well as the ODE system Eq. 1 of SVGD by use of the mean field characteristic flow. Some useful estimates on the solution Eq. 1 are also derived. Section 4 concerns the regularity of the solution to the mean field equation (2) under additional regularity assumption on KK. Section 5 devotes to the proof of the passage from the particles system Eq. 1 to its mean field PDE Eq. 2. Finally, in Section 6 we prove that the solution ρ\rho of (2) converges to the equilibrium ρ∞\rho_{\infty} as t→∞t\rightarrow\infty.

Throughout the paper we assume that the kernel KK satisfies the following:

K:Rd↦RK:\mathbf{R}^{d}\mapsto\mathbf{R} is at least C4C^{4} with bounded derivatives. In addition, K(x−y)K(x-y) is symmetric and positive definite, meaning that

A canonical choice of KK satisfying Assumption 2.1 is a Gaussian kernel, e.g. K(x)=1(4π)d/2exp(−∣x∣24)K(x)=\frac{1}{(4\pi)^{d/2}}\text{exp}(-\frac{|x|^{2}}{4}). Higher regularity of KK will be needed to obtain higher regularity of the solution of the mean field PDE; see Proposition 2.4. For the long time convergence of the solution, we will need further assumption on KK; see Theorem 2.6.

For the potential function V:Rd↦RV:\mathbf{R}^{d}\mapsto\mathbf{R}, we will assume the following:

V∈C∞(Rd),V≥0V\in C^{\infty}(\mathbf{R}^{d}),V\geq 0 and V(x)→+∞V(x)\rightarrow+\infty if ∣x∣→+∞|x|\rightarrow+\infty.

There exists a constant CV>0C_{V}>0 and some index q>1q>1 such that

For any α,β>0\alpha,\beta>0, there exists a constant Cα,β>0C_{\alpha,\beta}>0 such that if ∣y∣≤α∣x∣+β|y|\leq\alpha|x|+\beta, then

We comment that Section 2.1 (A1)-(A3) will be used in the proofs of the existence, uniqueness and regularity of the solution of mean field equation. Note that by setting α=1,β=0\alpha=1,\beta=0 and y=xy=x in (A3), we have that

for some constant C1>0C_{1}>0. These assumptions are by no means sharp, but proves to be sufficient for the validity of our theorems. Section 2.1 (A2) implies that there is C0C_{0} such that

where q∗=qq−1q^{*}=\frac{q}{q-1}. Indeed, this follows from

where n^=x/∣x∣\hat{n}=x/|x|, and then integrating from t=0t=0 to t=∣x∣t=|x|. It is also easy to check that Section 2.1 is fulfilled by even polynomials up to order q∗q^{*}.

We use PV\mathscr{P}_{V} and Pp\mathscr{P}_{p} denote the set of Borel probability measures μ\mu on Rd\mathbf{R}^{d} satisfying

respectively. Thanks to Eq. 14, we have Pp⊂PV\mathscr{P}_{p}\subset\mathscr{P}_{V} for any p≥q∗=qq−1p\geq q^{\ast}=\frac{q}{q-1}. For μ,ν∈Pp\mu,\nu\in\mathscr{P}_{p}, Wp(μ,ν)\mathcal{W}_{p}(\mu,\nu) denotes the pp-Wasserstein distance . Given a probability measure μ\mu and a Borel-measurable map ff, we denote by f#μf_{\#}\mu the push-forward of the measure μ\mu under the map ff. In places where ρ\rho is time-dependent, we often use notation ρt=ρ(t,⋅)\rho_{t}=\rho(t,\cdot) to emphasize this time dependence in a succinct way; on the other hand, differentiation with respect to the variable tt will always be denoted by ∂tρ\partial_{t}\rho.

For k,p≥1k,p\geq 1, we denote by Wk,p(Rd)W^{k,p}(\mathbf{R}^{d}) the usual Sobolev space of functions whose weak derivatives up to kk-th order belong to Lp(Rd)L^{p}(\mathbf{R}^{d}). When p=2p=2, we write Hk(Rd)=Wk,2(Rd)H^{k}(\mathbf{R}^{d})=W^{k,2}(\mathbf{R}^{d}). For our result on regularity of solutions to the PDE (2), we introduce function spaces

with norms ∥u∥LV1:=∥(1+V)u∥L1(Rd)\|u\|_{L^{1}_{V}}:=\|(1+V)u\|_{L^{1}(\mathbf{R}^{d})} and ∥u∥WV1,1:=∫Rd(1+V(x))(∣u(x)∣+∣∇u(x)∣)dx.\|u\|_{W^{1,1}_{V}}:=\int_{\mathbf{R}^{d}}(1+V(x))(|u(x)|+|\nabla u(x)|)dx. respectively. We set

2 Main Results

Our first result is the global well-posedness of the nonlinear mean field PDE (2). Observe that (2) is a nonlinear transport equation of the form ∂tρ+∇⋅(ρU[ρ])=0\partial_{t}\rho+\nabla\cdot(\rho U[\rho])=0, where U[ρ]U[\rho] is the vector field

Given a measure ρ∈PV\rho\in\mathscr{P}_{V}, U[ρ]U[\rho] is well-defined. In fact, due to Assumption 2.1 (A2), U[ρ](x)U[\rho](x) is Lipschitz continuous and bounded over Rd\mathbf{R}^{d}:

We say that a measure-valued function ρ∈C([0,∞);P)\rho\in C([0,\infty);\mathscr{P}) (where P\mathscr{P} is given the topology of weak convergence) is a weak solution to (2) with initial condition ρ0=ν∈PV\rho_{0}=\nu\in\mathscr{P}_{V} if

holds for all ϕ∈C0∞([0,∞)×Rd)\phi\in C^{\infty}_{0}([0,\infty)\times\mathbf{R}^{d}). Recall that ρt=ρ(t,⋅)\rho_{t}=\rho(t,\cdot).

Let VV satisfy Section 2.1. For any ν∈PV\nu\in\mathscr{P}_{V}, there is a unique ρ∈C([0,∞);PV)\rho\in C([0,\infty);\mathscr{P}_{V}) which is a weak solution to (2) with initial condition ρ0=ν\rho_{0}=\nu. Moreover there is C1>0C_{1}>0 (depending on KK and VV) such that

If ν∈Pp∩PV\nu\in\mathscr{P}_{p}\cap\mathscr{P}_{V}, then ρ∈C([0,∞);Pp)\rho\in C([0,\infty);\mathscr{P}_{p}), as well, with ∥ρt∥Pp≤eC2t∥ν∥Pp\|\rho_{t}\|_{\mathscr{P}_{p}}\leq e^{C_{2}t}\|\nu\|_{\mathscr{P}_{p}}.

The theorem is proved in Section 3.1 Our next result, proved in Section 3.2, establishes that the finite particle system is well-posed, and that the associated empirical measure is a weak solution of the PDE (2):

Let VV satisfy Section 2.1. Then for any initial condition x0={xi0}i=1N∈RdN\mathbf{x}^{0}=\{x_{i}^{0}\}_{i=1}^{N}\in\mathbf{R}^{dN}, the system (1) has a unique global solution x(t)={xi(t)}i=1N∈C1([0,∞);RdN)\mathbf{x}(t)=\{x_{i}(t)\}_{i=1}^{N}\in C^{1}([0,\infty);\mathbf{R}^{dN}), and the measure μtN=1N∑i=1Nδxi(t)\mu^{N}_{t}=\frac{1}{N}\sum_{i=1}^{N}\delta_{x_{i}(t)} is a weak solution to the PDE (2).

In particular, the bound (19) holds for the empirical measure μtN\mu_{t}^{N}. With additional assumptions about the behavior of VV and KK as ∣x∣→∞|x|\to\infty, we are able to improve upon (19) and show that ∥μtN∥PV\|\mu_{t}^{N}\|_{\mathscr{P}_{V}} is bounded in time; see Lemma 3.6 below.

When the initial condition is more regular, then the weak solution inherits higher regularity, as described by the following proposition. We remark that this regularity result will not be used in our proof of the mean field limit, but is of interest on its own account from the PDE perspective.

Let VV satisfy Section 2.1. Suppose that ν∈PV\nu\in\mathscr{P}_{V} has a density ρ0(x)≥0\rho_{0}(x)\geq 0. If ρt\rho_{t} is the unique weak solution to (2) with this initial condition, then ρt\rho_{t} also has a density. Furthermore, if ρ0∈Yk,V1\rho_{0}\in\mathscr{Y}^{1}_{k,V} for some k≥2k\geq 2, and the kernel KK is k+2k+2 times differentiable with bounded derivatives, then ρt\rho_{t} has a density satisfying

where the constants C1,C2C_{1},C_{2} depend only on VV and KK.

In the case that V=0V=0, a similar regularity result to Eq. 20 was proved for aggregation equation by Laurent . The presence of the potential VV makes the problem more difficult since the velocity ∇V\nabla V is unbounded at infinity. This difficulty was circumvented with the help of the mean field characteristic flow (c.f. Definition 3.1), which allows us to express the solution ρt\rho_{t} in terms of the initial condition and the flow map. The regularity of ρt\rho_{t} simply transfers from that of KK provided we can show ρ(t,⋅)∈WV1,1\rho(t,\cdot)\in W^{1,1}_{V}. See the detailed proof in Section 4.

Next, we prove a stability estimate for weak solutions to (2).

Let VV satisfy Section 2.1 with q∈(1,∞)q\in(1,\infty) in (A2). Let pp be the conjugate index of qq, i.e. 1p+1q=1\frac{1}{p}+\frac{1}{q}=1. Let R>0R>0. Assume that ν1,ν2\nu_{1},\nu_{2} are two initial probability measures in Pp\mathscr{P}_{p} satisfying ∥νi∥Pp≤R\|\nu_{i}\|_{\mathscr{P}_{p}}\leq R, i=1,2i=1,2. Let μ1,t\mu_{1,t} and μ2,t\mu_{2,t} be the associated weak solutions to (2). Then given any T>0T>0, there exists a constant C>0C>0 depending on K,V,R,pK,V,R,p and TT such that

Theorem 2.5 addresses the behavior of the particle system as N→∞N\to\infty. Suppose that the initial points {xiN(0)}i=1N\{x_{i}^{N}(0)\}_{i=1}^{N} are such that Wp(μ0N,ν0)→0\mathcal{W}_{p}(\mu^{N}_{0},\nu_{0})\to 0 as N→∞N\to\infty. Then if ρt\rho_{t} is the unique weak solution to (2) with initial condition ν0\nu_{0}, Theorem 2.5 implies that Wp(μtN,ρt)→0\mathcal{W}_{p}(\mu^{N}_{t},\rho_{t})\to 0 uniformly over [0,T][0,T], since μtN\mu^{N}_{t} is a weak solution to (2). This hypothesis of the convergence of the initial empirical measure, i.e., Wp(μ0N,ν0)→0\mathcal{W}_{p}(\mu^{N}_{0},\nu_{0})\rightarrow 0, can be justified rigorously, e.g., when the initial particles {xi0}\{x^{0}_{i}\} are independent samples drawn from ν0\nu_{0}. For a detailed discussion on the convergence of empirical measures in Wp\mathcal{W}_{p}, we refer the interested readers to references .

We prove Theorem 2.5 in Section 5 by following Dobrushin’s coupling argument for the mean field characteristic flow (defined later at (23)). The proof follows closely the proof of Theorem 1.4.1 of , which dealt with the case V=0V=0. The stability estimate there was stated in terms of 11-Wasserstein distance, and mainly resulted from the Lipschitz condition of ∇K\nabla K. However, we are only be able to prove the stability of mean field characteristic flow in pp-Wasserstein distance with pp strictly larger than one. This is again due to the presence of the nonlinear drift term K∗(∇Vρ)K\ast(\nabla V\rho) in the vector field (16).

Our last result pertains to the long time behavior of solutions ρt\rho_{t} of (2) with sufficiently regular initial condition. Since the probability density ρ∞(x)=e−V(x)/Z\rho_{\infty}(x)=e^{-V(x)}/Z is an invariant solution to the PDE (2), it is natural to ask whether ρ∞\rho_{\infty} is the unique invariant measure, and whether ρt→ρ∞\rho_{t}\to\rho_{\infty} as t→∞t\to\infty. Generally speaking, (2) may admit many invariant measures. For example, for any stationary solution to the finite particle system (1), the empirical measure μN\mu^{N} corresponds to a (stationary) weak solution of the PDE (2); there may be many such stationary solutions. However, if we restrict to initial conditions ρ0\rho_{0} which are absolutely continuous with respect to ρ∞\rho_{\infty}, one may expect that solutions to (2) converge to ρ∞\rho_{\infty} as t→∞t\to\infty. The following theorem confirms this intuition. For technical reasons, we need to make further assumptions on the kernel KK.

Let VV satisfy Section 2.1. Assume that KK satisfies Section 2.1 and the following extra assumption:

Theorem 2.6 in particular implies that ρ∞=e−V/Z\rho_{\infty}=e^{-V}/Z is the unique equilibrium of the mean field equation Eq. 2 provided that the initial distribution ρ0\rho_{0} has a density and satisfies KL(ρ0∣∣ρ∞)<∞\text{KL}(\rho_{0}||\rho_{\infty})<\infty. However, if the initial distribution is discrete, such as in the case of the particle system (1), there could be multiple equilibria, in which case the long time behavior of ρt\rho_{t} may depend on the initial distribution.

The proof of Theorem 2.6 is presented in Section 6. A quantitative convergence rate is far from clear to us. The main obstacle is the lack of a generalized logarithmic Sobolev inequality which could lower bound the Stein discrepancy in terms of the relative entropy. This issue is to be investigated in future works. Another important unresolved issue is whether “generic” stationary solutions of the particle system (1) are close in some sense to ρ∞\rho_{\infty}, when NN is large.

In this section we prove Theorem 2.2 and Proposition 2.3. The main ingredient in the proof of Theorem 2.2 is the so-called mean field characteristic flow, introduced in Section 3.1. In Section 3.2, we prove Proposition 2.3 and an additional estimate on the particle system under strong assumptions on VV.

Here we define the mean field characteristic flow for the PDE (2) (c.f ), which will play an essential role in the proof of the large particle limit of (1).

Given a probability measure ν\nu, we say that the map

is a mean field characteristic flow associated to the particle system (1) or to the mean field PDE (2) if XX is C1C^{1} in time and solves the following problem

The expression μt=X(t,⋅,ν)#ν\mu_{t}=X(t,\cdot,\nu)_{\#}\nu means that the measure μt\mu_{t} is the push-forward of ν\nu under the map x↦X(t,⋅,ν)x\mapsto X(t,\cdot,\nu). We think of {X(t,⋅,ν)}t≥0,ν\{X(t,\cdot,\nu)\}_{t\geq 0,\nu} as a family of maps from Rd\mathbf{R}^{d} to Rd\mathbf{R}^{d}, parameterized by tt and ν\nu. We first prove in the theorem below that the mean field characteristic flow (23) is well-defined. To this end, define the set of functions

which is a complete metric space with the uniform metric dY(u,v)=sup⁡x∣u(x)−v(x)∣d_{Y}(u,v)=\sup_{x}|u(x)-v(x)|. Recall the space of measures PV\mathscr{P}_{V} defined in (15).

Assume the conditions of Theorem 2.2 hold, and ν∈PV\nu\in\mathscr{P}_{V}. For any T>0T>0, there exists a unique solution X(⋅,⋅,ν)∈C1([0,T];Y)X(\cdot,\cdot,\nu)\in C^{1}([0,T];Y) to the problem (23). Moreover, the measure μt=X(t,⋅,ν)#ν\mu_{t}=X(t,\cdot,\nu)_{\#}\nu satisfies ∥μt∥PV≤eC∥∇K∥∞t∥ν∥PV\|\mu_{t}\|_{\mathscr{P}_{V}}\leq e^{C\|\nabla K\|_{\infty}t}\|\nu\|_{\mathscr{P}_{V}}, some constant CC that is independent of ν\nu.

We follow the proof of Theorem 1.3.2 in . The proof of the theorem consists of two steps.

Step 1 (local well-posedness): Fix r>0r>0, and define

We prove that there exists T0>0T_{0}>0 such that the problem (23) has a unique solution X(t,x)X(t,x) in the set

which is a complete metric space, with metric

Consider the integral formulation of (23) given by

Let us define the operator F:u(t,⋅)↦F(u)(t,⋅)\mathcal{F}:u(t,\cdot)\mapsto\mathcal{F}(u)(t,\cdot) by

Our goal is to show that F\mathcal{F} is a contraction in SrS_{r}, and thus has a unique fixed point.

We first show that F\mathcal{F} maps SrS_{r} into SrS_{r}. Checking that (t,x)↦F(u)(t,x)(t,x)\mapsto\mathcal{F}(u)(t,x) is continuous is straightforward; we need to establish a bound on ∣F(u)(t,x)−x∣|\mathcal{F}(u)(t,x)-x|. If u∈Sru\in S_{r}, then for any s∈[0,T0]s\in[0,T_{0}] and x′∈Rx^{\prime}\in\mathbf{R},

Then according to Assumptions (2.1) (A3), there exists a positive constant CrC_{r} such that

where we used the assumption that ν∈PV\nu\in\mathscr{P}_{V}. Therefore,

Next, we show that F\mathcal{F} is indeed a contraction on SrS_{r}. If u,v∈Sru,v\in S_{r}, then for any t∈[0,T0]t\in[0,T_{0}] and x∈Rdx\in\mathbf{R}^{d},

The first term on the right side above can be bounded from above by

Thanks to (25) and (13), the second term can be bounded from above by

To bound the last term on the right side of (27), using Section 2.1 (A2) one obtains that

where in the last inequality we have used the fact that θu+(1−θ)v∈Sr\theta u+(1-\theta)v\in S_{r} so that θu+(1−θ)v\theta u+(1-\theta)v also satisfies the inequality (25), which enables us to apply (A3) of Section 2.1. Plugging (29) into the integral of the last term on the right side of (27), we can bound the last term by

which implies that F\mathcal{F} is a contraction on SrS_{r} when T0T_{0} is small enough. By the contraction mapping theorem, F\mathcal{F} has a unique fixed point X(⋅,⋅,ν)∈SrX(\cdot,\cdot,\nu)\in S_{r}, which solves (24). After defining μt=X(t,⋅,ν)#ν\mu_{t}=X(t,\cdot,\nu)_{\#}\nu, one sees that X(t,x,ν)X(t,x,\nu) solves (23) in the small time interval [0,T0][0,T_{0}].

Step 2 (Extension of local solution): Considering the bounds in the previous step, it is clear that the local solution may be extended beyond time T0T_{0} as long as the quantity

remains finite. We now establish an a priori bound on this quantity, showing that the local solution may be extended for all t>0t>0.

The last inequality follows from Section 2.1 (A3) and the fact that KK is positive definite so that the third line above is non-positive. As a consequence,

holds for all t∈[0,T0]t\in[0,T_{0}]. With bound, one can iterate the argument to extend the local solution defined on [0,T0]×Rd[0,T_{0}]\times\mathbf{R}^{d} to all of [0,∞)×Rd[0,\infty)\times\mathbf{R}^{d}, so that ∥μt∥PV≤eCr∥∇K∥∞t∥ν∥PV\|\mu_{t}\|_{\mathscr{P}_{V}}\leq e^{C_{r}\|\nabla K\|_{\infty}t}\|\nu\|_{\mathscr{P}_{V}} holds for all t>0t>0. Similarly, there is C>0C>0 (depending on rr) such that dY(X(t,x,ν),x)≤CeCtd_{Y}(X(t,x,\nu),x)\leq Ce^{Ct} holds for all t≥0t\geq 0. Finally, thanks to the integral formulation (24) ∂tX\partial_{t}X is continuous on [0,∞)×Rd[0,\infty)\times\mathbf{R}^{d}. The proof is complete.

Given ν\nu, let X(t,x,ν)X(t,x,\nu) be the mean field characteristic flow defined in Theorem 3.2, and let ρt=X(t,⋅,ν)#ν\rho_{t}=X(t,\cdot,\nu)_{\#}\nu. Then this ρt\rho_{t} is a weak solution to (2) in the sense described above – this follows immediately from Theorem 5.34 in , for example.

Suppose that ν∈Pp∩PV\nu\in\mathscr{P}_{p}\cap\mathscr{P}_{V}. As shown in the proof of Theorem 3.2, the map X(t,x,ν)X(t,x,\nu) is an element of the space YY with dY(X(t,x,ν),x)≤CeCtd_{Y}(X(t,x,\nu),x)\leq Ce^{Ct}. Therefore, since ∣X(t,x,ν)∣p≤2p∣x∣p+2pdY(X(t,x,ν),x))p|X(t,x,\nu)|^{p}\leq 2^{p}|x|^{p}+2^{p}d_{Y}(X(t,x,\nu),x))^{p}, we have

for all t>0t>0. Hence ρt∈Pp∩PV\rho_{t}\in\mathscr{P}_{p}\cap\mathscr{P}_{V} for all t>0t>0.

2 Estimates on the particle system

In this section, we prove Proposition 2.3, showing that the particle system (1) is well-posed and that the empirical measure is a weak solution to the mean field PDE. It is useful to introduce the function

where x=(x1,x2,⋯ ,xn)\mathbf{x}=(x_{1},x_{2},\cdots,x_{n}).

Since both KK and VV are C2C^{2}, it is well-known that the problem (1) has a unique solution up to some time T0>0T_{0}>0. So, we must show that the solution does not blow up at a finite time. We claim that for some constant CC,

This estimate and Section 2.1 (A1) imply that xi(t)x_{i}(t) remains bounded over [0,T][0,T] for any T>0T>0, whence the solution can be extended up to any finite time. To establish (35), we first differentiate V(xi(t))V(x_{i}(t)) with respect to tt and sum over ii:

Observe that the second term on the right side of above is non-positive since the matrix {K(xi−xj)}i,j=1N\{K(x_{i}-x_{j})\}_{i,j=1}^{N} is positive definite by Assumption (2.1). Then it follows from the inequality in Assumptions (2.1) (A-2) and the fact that ∇K\nabla K is uniformly bounded that there exists a constant C=C(V,K)>0C=C(V,K)>0 such that

Now having established well-posedness of the finite particle system, it now follows from the definition of the mean field characteristic flow X(t,x,μ0N)X(t,x,\mu^{N}_{0}) that

In view of the proof of Theorem 2.2, we conclude that μtN\mu^{N}_{t} is a weak solution to the mean field PDE (2).

The estimate (35) can be regarded as a discrete analogue of the estimate (19) established in Theorem 2.2. We expect that for fixed NN, HN(x)H_{N}(\mathbf{x}) will remain uniformly bounded in time, although we have been able to prove this only with some further restrictions on VV and KK, as the next lemma states.

Fix N≥1N\geq 1. Suppose that for some p≥2p\geq 2 and m,R>0m,R>0, V(x)=m∣x∣pV(x)=m|x|^{p} if ∣x∣>R|x|>R. Suppose also that K(0)>0K(0)>0 and that ∣x∣p−1K(x)|x|^{p-1}K(x) is bounded. Then HN(x(t))H_{N}(x(t)) is uniformly bounded for t∈[0,∞)t\in[0,\infty).

Observe that ∂tHN(x(t))=−(S1+S2)/N2\partial_{t}H_{N}(x(t))=-(S_{1}+S_{2})/N^{2}, where

Because KK is positive definite, we know that S2≥0S_{2}\geq 0. We wish to bound S2S_{2} from below. For y∈Rdy\in\mathbf{R}^{d}, let us define

(As elsewhere in the paper, the constant CC may change from line to line, here). By the assumptions on VV, there is α=(p−2)/(p−1)∈[0,1)\alpha=(p-2)/(p-1)\in[0,1) such that ∣xi∣p−2≤C(1+∣∇V(xi)∣α|x_{i}|^{p-2}\leq C(1+|\nabla V(x_{i})|^{\alpha}. Also, K(xi−xj)∣∣xj−xi∣p−1K(x_{i}-x_{j})||x_{j}-x_{i}|^{p-1} is bounded, by assumption. Consequently,

always holds. Applying Hölder’s inequality with exponents p=2p=2 and with (p,p∗)=(2/(1+α),2/(1−α))(p,p^{*})=(2/(1+\alpha),2/(1-\alpha)) we obtain

In particular, there are constants C1,C2>0C_{1},C_{2}>0 (dependent on K(0)K(0) and α\alpha) so that

Since ∣∇V∣2≥C(V+1)−C′|\nabla V|^{2}\geq C(V+1)-C^{\prime} holds for all xx, for some positive constants C,C′C,C^{\prime}, this implies

which implies that HNH_{N} is uniformly bounded in tt, for NN fixed.

In this section we prove Proposition 2.4 under the assumption that the initial distribution ν∈PV\nu\in\mathscr{P}_{V} has a density ρ0∈Yk,V1\rho_{0}\in\mathscr{Y}^{1}_{k,V} with some fixed k≥2k\geq 2 and that the kernel KK is (k+2)(k+2) times differentiable with bounded derivatives.

By Theorem 2.2, we know that ρt∈C([0,T];PV)\rho_{t}\in C([0,T];\mathscr{P}_{V}) satisfies

Consequently, the vector field (t,x)↦U[ρt](x)(t,x)\mapsto U[\rho_{t}](x) (defined at (16)) satisfies

Thus U(t,x)∈C([0,T];CBk+1(Rd))U(t,x)\in C([0,T];C_{B}^{k+1}(\mathbf{R}^{d})) where we recall that CBk+1(Rd)C_{B}^{k+1}(\mathbf{R}^{d}) is the space of continuous functions with bounded (k+1)(k+1)-th order derivatives. Let Φt(x)=X(t,x,ν)\Phi_{t}(x)=X(t,x,\nu) denote the characteristic flow (Definition 3.1). Since Φt\Phi_{t} satisfies the ODE system ddtΦt(x)=U[ρt](Φt(x))\frac{d}{dt}\Phi_{t}(x)=U[\rho_{t}](\Phi_{t}(x)), it follows from standard theory that the maps x↦Φtx\mapsto\Phi_{t} and its inverse Φt−1\Phi_{t}^{-1} are both CkC^{k} maps (e.g. see Chapter 2 of ). Therefore, if ρ0\rho_{0} has a density, then ρt\rho_{t} also has a density. In fact, ρt(x)\rho_{t}(x) is given by

with the vector field U(t,x)∈C([0,T];CBk+1(Rd))U(t,x)\in C([0,T];C_{B}^{k+1}(\mathbf{R}^{d})), it follows from [24, Lemma 2.8] that ρ∈C([0,T];Hk(Rd))\rho\in C([0,T];H^{k}(\mathbf{R}^{d})) for any T>0T>0 and ∂tρ∈C([0,T];Hk−1(Rd))\partial_{t}\rho\in C([0,T];H^{k-1}(\mathbf{R}^{d})).

It remains to prove that ρ(t,⋅)∈WV1,1\rho(t,\cdot)\in W^{1,1}_{V} for every t∈[0,T]t\in[0,T] and that it satisfies the a priori estimate (20). First, we show that ρ(t,⋅)∈WV1,1\rho(t,\cdot)\in W^{1,1}_{V}. To see this, we differentiate both sides of

with respect to xix_{i} to get the following equation for ∂xiρ\partial_{x_{i}}\rho

Now given δ>0\delta>0, we define the one dimensional function

It is clear that ϕδ(x)→∣x∣\phi_{\delta}(x)\rightarrow|x| as δ→0\delta\rightarrow 0 and that sup⁡x∈R∣ϕδ′(x)∣≤1\sup_{x\in\mathbf{R}}|\phi_{\delta}^{\prime}(x)|\leq 1. Then ϕδ(∂xiρ)\phi_{\delta}(\partial_{x_{i}}\rho) satisfies:

Notice that since ρ∈C([0,T];Hk(Rd))\rho\in C([0,T];H^{k}(\mathbf{R}^{d})) for any T>0T>0 and ∂tρ∈C([0,T];Hk−1(Rd))\partial_{t}\rho\in C([0,T];H^{k-1}(\mathbf{R}^{d})), the above equation holds in the space C([0,T];L2(Rd))C([0,T];L^{2}(\mathbf{R}^{d})). Let ηR\eta_{R} be a smooth cut-off function on Rd\mathbf{R}^{d} such that

Next, we multiply the above equation with (1+V)ηR(1+V)\eta_{R}, and then integrate on the whole space to get

Using the fact that ∣ϕδ′∣≤1|\phi^{\prime}_{\delta}|\leq 1 and that ηR\eta_{R} is uniformly bounded, we have that

For I1(δ,R)I_{1}(\delta,R), using integration by parts and the assumption (13) one obtains that

Consequently, letting δ→0\delta\rightarrow 0 and R→∞R\rightarrow\infty, we obtain from (44) and (38) that

Finally we derive an HkH^{k}-estimate for the solution. For doing so, let α\boldsymbol{\alpha} be a multi-index such that ∣α∣≤k|\boldsymbol{\alpha}|\leq k. Taking ∂α\partial^{\boldsymbol{\alpha}} on the both sides of (40), multiplying the resulting equation with ∂αρ\partial^{\boldsymbol{\alpha}}\rho and then integrating gives

Note that by Leibniz rule and integration by parts,

Plugging the estimates into (47) and using (38), we obtain by summing over α\boldsymbol{\alpha} with ∣α∣≤k|\boldsymbol{\alpha}|\leq k that

The estimate follows from (20) and (48). This finishes the proof of the proposition.

In this section we prove Theorem 2.5 using Dobrushin’s coupling argument, following Theorem 1.4.1 of .

Recall that p=q∗=qq−1p=q^{\ast}=\frac{q}{q-1}. First by the assumption that ∥νi∥Pp≤R<∞\|\nu_{i}\|_{\mathscr{P}_{p}}\leq R<\infty and the fact that Pp⊂PV\mathscr{P}_{p}\subset\mathscr{P}_{V} thanks to Eq. 14, we know that there exists C(R)>0C(R)>0 such that

By the proof of Theorem 2.2 and Definition 3.1 of the mean field characteristic flow, we know that the weak solutions μi,t\mu_{i,t} take the form

So, we must estimate Wpp(μ1,t,μ2,t)\mathcal{W}_{p}^{p}(\mu_{1,t},\mu_{2,t}) in terms of Wpp(ν1,ν2)\mathcal{W}_{p}^{p}(\nu_{1},\nu_{2}). Let π0\pi^{0} be a coupling measure between the probability measures ν1\nu_{1} and ν2\nu_{2}. Define for δ>0\delta>0, ϕδ(x)=1p(∣x∣2+δ)p/2\phi_{\delta}(x)=\frac{1}{p}(|x|^{2}+\delta)^{p/2} to be an approximation to 1p∣x∣p\frac{1}{p}|x|^{p}, Given any two points x1,x2∈Rdx_{1},x_{2}\in\mathbf{R}^{d}, we have from (23) that

Below we bound IiI_{i} individually. First, it is important to notice that

Then thanks to Assumption (2.1) on KK and the fact that the inclusion Lp↪L1L^{p}\xhookrightarrow{}L^{1} is bounded for p>1p>1, we have

For I2I_{2}, it follows from Assumption 2.1 (A2)-(A3) and Hölder’s inequality that

Observe that the integrals involving VV on the right side of above can be bounded in exactly the same way as (31). Hence we can obtain

with the constant CC depending only on VV. Finally, we find an upper bound for I3I_{3}. In fact, an application of the intermediate value theorem to the difference of ∇V\nabla V and the inequality (12) of Assumption 2.1 (A-2) yields that

then by combing the estimates above, we obtain that for any t∈[0,T]t\in[0,T],

Now integrating the above inequality with respect to the coupling π0(dx1dx2)\pi^{0}(dx_{1}dx_{2}), using the fact that

and finally letting δ→0\delta\rightarrow 0 yields

By the Grönwall’s inequality we obtain that

Now since π0∈Γ(ν1,ν2)\pi^{0}\in\Gamma(\nu_{1},\nu_{2}) and μi,t=(X(t,⋅,νi))#νi\mu_{i,t}=(X(t,\cdot,\nu_{i}))_{\#}\nu_{i}, the mapping

satisfies that (Ξt)#π0∈Γ(μ1,t,μ2,t)(\Xi_{t})_{\#}\pi^{0}\in\Gamma(\mu_{1,t},\mu_{2,t}). As a consequence, we have that

This finishes the proof in view of Eq. 49.

In this section we prove Theorem 2.6. For doing so, we recall following extra assumption on the kernel KK:

A canonical kernel satisfying this condition is a Gaussian kernel.

To prove ρt⇀ρ∞\rho_{t}\rightharpoonup\rho_{\infty} as t→∞t\rightarrow\infty, we only need to prove that ρtk⇀ρ∞\rho_{t_{k}}\rightharpoonup\rho_{\infty} for any sequence tk↗∞t_{k}\nearrow\infty. Indeed, suppose that the later is true and that ρt\rho_{t} does not converge weakly to ρ∞\rho_{\infty}. Then there exists a constant ε>0\varepsilon>0 and a bounded continuous function φ\varphi, such that there exists a sequence tk↗∞t_{k}\nearrow\infty such that

where the inequality follows from the fact that K(x−y)K(x-y) is positive definite. Furthermore, noticing that

Step 2: We show that ρˉ\bar{\rho} satisfies

in the sense of distribution. To this end, using Fourier transform and the fact that K^=K^1/22\widehat{K}=\widehat{K}_{1/2}^{2} we can write

Note that we are allowed to take the Fourier transform because ρtkm∈Y2,V1\rho_{t_{k_{m}}}\in\mathscr{Y}^{1}_{2,V} by Theorem 2.6 and the assumption that ρ0∈Y2,V1\rho_{0}\in\mathscr{Y}^{1}_{2,V}. This together with (51) implies that K1/2∗(∇ρtkm+∇Vρtkm)→0K_{1/2}\ast(\nabla\rho_{t_{k_{m}}}+\nabla V\rho_{t_{k_{m}}})\rightarrow 0 in L2(Rd)L^{2}(\mathbf{R}^{d}). On the other hand, using ρtkm⇀ρˉ\rho_{t_{k_{m}}}\rightharpoonup\bar{\rho} with ρˉ∈P(Rd)\bar{\rho}\in\mathscr{P}(\mathbf{R}^{d}) and integration by parts, one sees that

Therefore we have that ∫Rd∇K1/2(x−y)ρˉ(y)+K1/2(x−y)∇V(y)ρˉ(y)dy=0\int_{\mathbf{R}^{d}}\nabla K_{1/2}(x-y)\bar{\rho}(y)+K_{1/2}(x-y)\nabla V(y)\bar{\rho}(y)dy=0 a.e. x∈Rdx\in\mathbf{R}^{d}. This in particular, implies that

Step 3: We show that ρˉ=ρ∞\bar{\rho}=\rho_{\infty}. We first prove that ∇ρˉ+∇Vρˉ=0\nabla\bar{\rho}+\nabla V\bar{\rho}=0 in the sense of tempered distribution. In fact, since ρˉ∈P(Rd)\bar{\rho}\in\mathscr{P}(\mathbf{R}^{d}) and since VV grows at most polynomially (due to Section 2.1 (A2)), we know that (∇ρˉ+∇Vρˉ)∈S′(\nabla\bar{\rho}+\nabla V\bar{\rho})\in\mathcal{S}^{\prime}. Since K1/2∈SK_{1/2}\in\mathcal{S}, it follows from the convolution theorem of Fourier transform (see e.g. [21, Chapter 4.11, Theorem 3 and Proposition 7]) that K1/2∗(∇ρˉ+∇Vρˉ)K_{1/2}\ast(\nabla\bar{\rho}+\nabla V\bar{\rho}) can be understood as a rapidly decreasing distribution whose Fourier transform is given by

By the assumption that K^1/2≠0\hat{K}_{1/2}\neq 0, we have from (52) that ∇ρˉ+∇Vρˉ^=0\widehat{\nabla\bar{\rho}+\nabla V\bar{\rho}}=0 and hence ∇ρˉ+∇Vρˉ=0\nabla\bar{\rho}+\nabla V\bar{\rho}=0. This in addition implies that ∇(eVρˉ)=0\nabla(e^{V}\bar{\rho})=0 in the sense of distribution. Therefore ρˉ=Cρ∞\bar{\rho}=C\rho_{\infty} a.e. for some constant CC. Finally since both ρˉ\bar{\rho} and ρ∞\rho_{\infty} are probability density, C=1C=1 and ρˉ=ρ∞\bar{\rho}=\rho_{\infty} a.e. This finishes the proof.

The authors would like to thank the anonymous referees for their valuable comments and suggestions to improve the structure and quality of the paper.