Approximation by log-concave distributions, with applications to regression

Lutz Duembgen, Richard Samworth, Dominic Schuhmacher

Introduction

Log-concave distributions, that is, distributions with a Lebesgue density the logarithm of which is concave, are an interesting nonparametric model comprising many parametric families of distributions. Bagnoli and Bergstrom (2005) give an overview of many interesting properties and applications in econometrics. Indeed, these distributions have received a lot of attention among statisticians recently as described in the review by Walther (2009). The nonparametric maximum likelihood estimator was studied in the univariate setting by Pal, Woodroofe and Meyer (2007), Rufibach (2006), Dümbgen, Hüsler and Rufibach (2007), Balabdaoui, Rufibach and Wellner (2009) and Dümbgen and Rufibach (2009). These references contain characterizations of the estimators, consistency results and explicit algorithms. Extensions of one or more of these aspects to the multivariate setting are presented by Cule, Samworth and Stewart (2010), Cule and Samworth (2010), Koenker and Mizera (2010), Seregin and Wellner (2010) and Schuhmacher and Dümbgen (2010). Both Cule and Samworth (2010) and Schuhmacher, Hüsler and Dümbgen (2009) show that multivariate log-concave distributions are a very well-behaved nonparametric class. For instance, moments of arbitrary order are continuous statistical functionals with respect to weak convergence.

(provided this exists and is unique). Even if QQ fails to have a density within F\mathcal{F}, one may view f^n\hat{f}_{n} as an estimator of the approximating density

Since the sequence of empirical measures Q^n\hat{Q}_{n} converges weakly to QQ almost surely, this entails strong consistency of the Grenander estimator f(⋅∣Q^n)f(\cdot|\hat{Q}_{n}) in total variation distance.

Some additional properties of f(⋅∣Q)f(\cdot|Q) will be established as well. We show that the mapping Q↦f(⋅∣Q)Q\mapsto f(\cdot|Q) is continuous with respect to Mallows distance [Mallows (1972)] D1(⋅,⋅)D_{1}(\cdot,\cdot), also known as a Wasserstein, Monge–Kantorovich or Earth Mover’s distance. Precisely, let QQ satisfy the properties just mentioned, and let (Qn)n(Q_{n})_{n} be a sequence of probability distributions converging to QQ in D1D_{1}; in other words,

as n→∞n\to\infty. Then f(⋅∣Qn)f(\cdot|Q_{n}) is well defined for sufficiently large nn and

This entails strong consistency of the maximum likelihood estimator f^n\hat{f}_{n}, because (Q^n)n(\hat{Q}_{n})_{n} converges almost surely to QQ with respect to Mallows distance D1(⋅,⋅)D_{1}(\cdot,\cdot). In addition we show that Q↦max⁡f∈F∫log⁡(f) dQQ\mapsto\max_{f\in\mathcal{F}}\int\log(f)\,dQ is convex and upper semicontinuous with respect to weak convergence.

In Section 3 we apply these results to the following type of regression problem: suppose that we observe independent real random variables Y1Y_{1}, Y2,…,YnY_{2},\ldots,Y_{n} such that

Many proofs and technical arguments are deferred to Section 4. A longer and more detailed version of this paper is the technical report by Dümbgen, Samworth and Schuhmacher (2010), referred to as [DSS 2010] hereafter. It contains all proofs, additional examples and plots, a detailed description of our algorithms and extensive simulation studies. There we also indicate potential applications to change-point analyses.

Log-concave approximations

The next theorem provides a complete characterization of all distributions Q∈QQ\in\mathcal{Q} with real profile log-likelihood L(Q)L(Q). To state the result we first define the convex support of a distribution Q∈QQ\in\mathcal{Q} and collect some of its properties.

is itself closed and convex with Q(csupp⁡(Q))=1Q(\operatorname{csupp}(Q))=1. The following three properties of QQ are equivalent:

csupp⁡(Q)\operatorname{csupp}(Q) has nonempty interior;

For any Q∈QQ\in\mathcal{Q}, the value of L(Q)L(Q) is real if and only if

In that case, there exists a unique function

This function ψ\psi satisfies ∫eψ(x) dx=1\int e^{\psi(x)}\,dx=1 and

Let PP be the approximating probability measure with P(dx)=eψ(x) dxP(dx)=e^{\psi(x)}\,dx. It satisfies the following (in)equalities:

The profile log-likelihood LL is convex on Q1\mathcal{Q}^{1}. Precisely, for arbitrary Q0,Q1∈Q1Q_{0},Q_{1}\in\mathcal{Q}^{1} and 0<t<10<t<1,

The two sides are equal and real if and only if Q0,Q1∈Qo∩Q1Q_{0},Q_{1}\in\mathcal{Q}_{o}\cap\mathcal{Q}^{1} with ψ(⋅∣Q0)=ψ(⋅∣Q1)\psi(\cdot|Q_{0})=\psi(\cdot|Q_{1}).

Furthermore, suppose that QQ has a density gg on an open set UU such that ψ>log⁡g\psi>\log g on this set. Then

2 The one-dimensional case

For the case of d=1d=1 one can generalize Theorem 2.4 of Dümbgen and Rufibach (2009) as follows: for a function ϕ∈Φ(1)\phi\in\Phi(1) let

One consequence of this theorem is that the c.d.f. FF of ψ(⋅∣Q)\psi(\cdot|Q) follows the c.d.f. GG of QQ quite closely in that

Let QQ be a rescaled version of Student’s distribution t2t_{2} with density and distribution function

respectively. The best approximating log-concave distribution is the Laplace distribution with density and distribution function

respectively. To verify this claim, note that by symmetry it suffices to show that

Indeed the integral on the left-hand side equals

for all x≤0x\leq 0. Clearly this expression is zero for x=0x=0, and elementary considerations show that it is nonpositive for all x≤0x\leq 0. Numerical calculations reveal that ∣F−G∣|F-G| is smaller than 0.040.04 everywhere.

Suppose that QQ has a continuous but not log-concave density gg. Nevertheless one can say the following about the approximating log-density ψ=ψ(⋅∣Q)\psi=\psi(\cdot|Q):

Suppose that log⁡g\log g is concave on an interval (−∞,a](-\infty,a] with g(a)>0g(a)>0 and ψ(a)≤log⁡g(a)\psi(a)\leq\log g(a). Then there exists a point a′∈[−∞,a]a^{\prime}\in[-\infty,a] such that ψ\psi is linear on (a′,a](a^{\prime},a] and ψ=log⁡g\psi=\log g on (−∞,a′](-\infty,a^{\prime}].

Suppose that log⁡g\log g is differentiable everywhere, convex on a bounded interval [a,b][a,b] and concave on both (−∞,a](-\infty,a] and [b,∞)[b,\infty). Then there exist points a′∈(−∞,a]a^{\prime}\in(-\infty,a] and b′∈[b,∞)b^{\prime}\in[b,\infty) such that ψ\psi is linear on [a′,b′][a^{\prime},b^{\prime}] while ψ=log⁡g\psi=\log g on (−∞,a′]∪[b′,∞)(-\infty,a^{\prime}]\cup[b^{\prime},\infty).

Suppose that log⁡g\log g is convex on an interval (−∞,a](-\infty,a] such that −∞<log⁡g(a)≤ψ(a)-\infty<\log g(a)\leq\psi(a). Then ψ\psi is linear on (−∞,a](-\infty,a].

Let us illustrate part (ii) of Remark 2.11 with a numerical example. Figure 1 shows the bimodal

3 Continuity in Q𝑄Q

For the applications to regression problems to follow we need to understand the properties of both Q↦L(Q)Q\mapsto L(Q) and Q↦ψ(⋅∣Q)Q\mapsto\psi(\cdot|Q) on Q1∩Qo\mathcal{Q}^{1}\cap\mathcal{Q}_{o}. Our first hope was that both mappings would be continuous with respect to the weak topology. It turned out, however, that we need a somewhat stronger notion of convergence, namely, convergence with respect to Mallows distance D1D_{1} which is defined as follows: for two probability distributions Q,Q′∈Q1Q,Q^{\prime}\in\mathcal{Q}^{1},

where the infimum is taken over all pairs (X,X′)(X,X^{\prime}) of random vectors X∼QX\sim Q and X′∼Q′X^{\prime}\sim Q^{\prime} on a common probability space. It is well known that the infimum in D1(Q,Q′)D_{1}(Q,Q^{\prime}) is a minimum. The distance D1D_{1} is also known as Wasserstein, Monge–Kantorovich or Earth Mover’s distance. An alternative representation due to Kantorovič and Rubinšteĭn (1958) is

A good starting point for more detailed information on Mallows distance is Chapter 7 of Villani (2003).

Before presenting the main results of this section we mention two useful facts about the convex support of distributions.

Moreover, if (Qn)n(Q_{n})_{n} is a sequence in Q\mathcal{Q} converging weakly to QQ, then

This lemma implies that the set Qo\mathcal{Q}_{o} is an open subset of Q\mathcal{Q} with respect to the topology of weak convergence. The supremum h(Q,x)h(Q,x) is a maximum over closed halfspaces and is related to Tukey’s halfspace depth [Donoho and Gasko (1992), Section 6]. For a proof of Lemma 2.13 we refer to [DSS 2010]. Now we are ready to state the main results of this section.

Let (Qn)n(Q_{n})_{n} be a sequence of distributions in Qo\mathcal{Q}_{o} converging weakly to some Q∈QoQ\in\mathcal{Q}_{o}. Then

Moreover, lim inf⁡n→∞L(Qn)<L(Q)\liminf_{n\to\infty}L(Q_{n})<L(Q) if and only if

This result already entails continuity of L(⋅)L(\cdot) on Qo∩Q1\mathcal{Q}_{o}\cap\mathcal{Q}^{1} with respect to Mallows distance D1D_{1}. The next theorem extends this result to L\dvtxQ1→(−∞,∞]L\dvtx\mathcal{Q}^{1}\to(-\infty,\infty]:

Let (Qn)n(Q_{n})_{n} be a sequence of distributions in Q1\mathcal{Q}^{1} such that lim⁡n→∞D1(Qn,Q)=0\lim_{n\to\infty}D_{1}(Q_{n},Q)=0 for some Q∈Q1Q\in\mathcal{Q}^{1}. Then

In case of Q∈Qo∩Q1Q\in\mathcal{Q}_{o}\cap\mathcal{Q}^{1}, the probability densities f:=\exp\mbox{{}\circ{}}\psi(\cdot|Q) and f_{n}:=\exp\mbox{{}\circ{}}\psi(\cdot|Q_{n}) are well defined for sufficiently large nn and satisfy

Applications to regression problems

We propose to estimate (ψ,μ)(\psi,\mu) by a maximizer of

maximizes Λ^(ϕ,m)\hat{\Lambda}(\phi,m) over all (ϕ,m)∈Φ×M(\phi,m)\in\Phi\times\mathcal{M} satisfying the additional constraint that \exp\mbox{{}\circ{}}\phi defines a probability density with mean zero.

Define x:=(xi)i=1n\mathbf{x}:=(x_{i})_{i=1}^{n} and m(x):=(m(xi))i=1nm(\mathbf{x}):=(m(x_{i}))_{i=1}^{n}. Then we may write

and this representation is our key to proving the existence of (ψ^,μ^)(\hat{\psi},\hat{\mu}). Before doing so we state a simple inequality of independent interest, which follows from Jensen’s inequality and elementary considerations:

For any distribution Q∈Q1(1)Q\in\mathcal{Q}^{1}(1),

where Med⁡(Q)\operatorname{Med}(Q) is a median of QQ while μ(Q)\mu(Q) denotes its mean ∫xQ(dx)\int xQ(dx).

The constraint Y∉M(x)\mathbf{Y}\notin\mathcal{M}(\mathbf{x}) excludes situations with perfect fit. In that case, the Dirac measure δ0\delta_{0} would be the most plausible error distribution.

The maximum likelihood estimator (ψ^,μ^)(\hat{\psi},\hat{\mu}) need not be unique in general. Nevertheless we will prove it to be consistent under certain regularity conditions. A key point here is Fisher consistency in the following sense: note that the expectation measure of the empirical distribution Q^m(x)\hat{Q}_{m(\mathbf{x})} equals

with equality if and only if μ−m\mu-m is constant on {x1,x2,…,xn}\{x_{1},x_{2},\ldots,x_{n}\}. This follows from a more general inequality which is somewhat reminiscent of Anderson’s lemma [Anderson (1955)]:

Let Q∈Qo(d)∩Q1(d)Q\in\mathcal{Q}_{o}(d)\cap\mathcal{Q}^{1}(d) and R∈Q1(d)R\in\mathcal{Q}^{1}(d). Then Q⋆R∈Qo∩Q1Q\star R\in\mathcal{Q}_{o}\cap\mathcal{Q}^{1} and

2 Consistency

In this subsection we consider a triangular scheme of independent observations (xni,Yni)(x_{ni},Y_{ni}), 1≤i≤n1\leq i\leq n, with fixed design points xni∈Xnx_{ni}\in\mathcal{X}_{n} and

where μn\mu_{n} is an unknown regression function in Mn\mathcal{M}_{n} and εn1,εn2,…,εnn\varepsilon_{n1},\varepsilon_{n2},\ldots,\varepsilon_{nn} are unobserved independent random errors with mean zero and unknown distribution Qn∈Qo(1)∩Q1(1)Q_{n}\in\mathcal{Q}_{o}(1)\cap\mathcal{Q}^{1}(1). Two basic assumptions are:

D1(Qn,Q)→0D_{1}(Q_{n},Q)\to 0 for some distribution Q∈Qo(1)∩Q1(1)Q\in\mathcal{Q}_{o}(1)\cap\mathcal{Q}^{1}(1).

We write (ψ^n,μ^n)(\hat{\psi}_{n},\hat{\mu}_{n}) for a maximizer of L(ϕ,Q^n,m)L(\phi,\hat{Q}_{n,m}) over all pairs (ϕ,m)∈Φ×Mn(\phi,m)\in\Phi\times\mathcal{M}_{n} such that ∫eϕ(x) dx=1\int e^{\phi(x)}\,dx=1 and ∫xeϕ(x) dx=0\int xe^{\phi(x)}\,dx=0, where Q^n,m\hat{Q}_{n,m} stands for the empirical distribution of the residuals Yni−m(xni)Y_{ni}-m(x_{ni}), 1≤i≤n1\leq i\leq n. We also need to consider its expectation measure

It is also convenient to metrize weak convergence. In Theorem 3.6 below we utilize the bounded Lipschitz distance: for probability distributions Q,Q′Q,Q^{\prime} on the real line let

Let assumptions (A.1) and (A.2) be satisfied. Suppose further that: {longlist}[(A.2)]

Then, with f_{n}:=\exp\mbox{{}\circ{}}\psi(\cdot|Q_{n}) and \hat{f}_{n}:=\exp\mbox{{}\circ{}}\hat{\psi}_{n}, the maximum likelihood estimator (f^n,μ^n)(\hat{f}_{n},\hat{\mu}_{n}) of (fn,μn)(f_{n},\mu_{n}) exists with asymptotic probability one and satisfies

We know already that assumption (A.1) is satisfied for multiple linear regression and isotonic regression. Assumption (A.2) is a generalization of assuming a fixed error distribution for all sample sizes. The crucial point, of course, is assumption (A.3). In our two examples it is satisfied under mild conditions:

The proof of Theorem 3.7 is given in Section 4. For the proof of Theorem 3.8, which uses similar ideas and an additional approximation argument, we refer to [DSS 2010].

3 Algorithms and numerical results

Extensive simulation studies in [DSS 2010] suggest that (ψ^,μ^)(\hat{\psi},\hat{\mu}) provides rather accurate estimates even if nn is only moderately large. For various skewed error distributions, μ^\hat{\mu} may be considerably better than the corresponding least squares estimator. As an example consider the simple linear regression model with observations

where X1,…,XnX_{1},\ldots,X_{n} are independent design points from the Unif⁡\operatorname{Unif} distribution and ε1,…,εn\varepsilon_{1},\ldots,\varepsilon_{n} are independent errors from a centered gamma distribution with shape parameter rr and variance 11. Note that the distribution of (ψ^,θ^−θ)(\hat{\psi},\hat{\theta}-\theta) does not depend on cc or θ\theta. Monte Carlo estimation of the root mean squared error based on 1000 simulations of this model gives 0.023 for the estimator θ^\hat{\theta} versus 0.118 for the least squares estimator of θ\theta if r=1r=1, and 0.095 versus 0.113 for the same comparison if r=3r=3.

4 A data example

A familiar task in econometrics is to model expenditure (YY) of households as a function of their income (XX). Not only the mean curve (Engel curve) but also quantile curves play an important role. A related application are growth charts in which, for instance, XX is the age of a newborn or infant and YY is its height or weight.

Interestingly, neither linear nor quadratic nor cubic regression yield convincing fits to these data. Polynomial regression of degree four or cubic splines with knot points at, say, 4.14.1, 4.34.3, 4.54.5, 4.74.7, 4.94.9 seem to fit the data quite well. Moreover, exact Monte Carlo goodness-of-fits test, assuming the regression function to be a cubic spline and based on a Kolmogorov–Smirnov statistic applied to studentized residuals, revealed the regression errors εi\varepsilon_{i} to be definitely non-Gaussian.

Figure 3 shows the data and estimated β\beta-quantile curves for β=0.1\beta=0.1, 0.250.25, 0.50.5, 0.750.75, 0.90.9, based on our additive regression model. Note that the estimated β\beta-quantile curve is simply the estimated mean curve plus the β\beta-quantile of the estimated error distribution. On the left-hand side, we only assumed μ\mu to be nondecreasing, on the right-hand side we fitted the aforementioned spline model. In both cases the fitted quantile curves are similar to the quantile curves in Figure 2 but with fewer irregularities such as big jumps which may be artifacts due to sampling error.

Proofs

For the proof of Theorem 2.2 we need an elementary bound for the Lebesgue measure of level sets of log-concave distributions:

Another key ingredient for the proofs of Theorems 2.2 and 2.15 is a lemma on pointwise limits of sequences in Φ\Phi:

is nonempty. Then there exist a subsequence (ϕn(k))k(\phi_{n(k)})_{k} of (ϕn)n(\phi_{n})_{n} and a function ϕ∈Φ\phi\in\Phi such that C⊆dom⁡(ϕ)={ϕ>−∞}C\subseteq\operatorname{dom}(\phi)=\{\phi>-\infty\} and

Proof of Theorem 2.2 Suppose first that ∫∥x∥Q(dx)=∞\int\|x\|Q(dx)=\infty. Since any ϕ∈Φ\phi\in\Phi is majorized by x↦a−b∥x∥x\mapsto a-b\|x\| for suitable constants aa and b>0b>0, this entails that L(Q)=−∞L(Q)=-\infty.

For the remainder of this proof suppose that ∫∥x∥Q(dx)<∞\int\|x\|Q(dx)<\infty and that csupp⁡(Q)\operatorname{csupp}(Q) has nonempty interior. Since the concave function h(x)=−∥x∥h(x)=-\|x\| satisfies ∫h dQ>−∞\int h\,dQ>-\infty, we have L(Q)>−∞L(Q)>-\infty. When maximizing L(ϕ,Q)L(\phi,Q) over all ϕ∈Φ\phi\in\Phi we may and do restrict our attention to functions ϕ∈Φ\phi\in\Phi such that ∫eϕ(x) dx=1\int e^{\phi(x)}\,dx=1 (see end of Section 1) and dom⁡(ϕ)={ϕ>−∞}⊆csupp⁡(Q)\operatorname{dom}(\phi)=\{\phi>-\infty\}\subseteq\operatorname{csupp}(Q). For if dom⁡(ϕ)⊄csupp⁡(Q)\operatorname{dom}(\phi)\not\subset\operatorname{csupp}(Q), replacing ϕ(x)\phi(x) with −∞-\infty for all x∉csupp⁡(Q)x\notin\operatorname{csupp}(Q) would also increase L(ϕ,Q)L(\phi,Q) strictly. Let Φ(Q)\Phi(Q) be the family of all ϕ∈Φ\phi\in\Phi with these properties.

as M→∞M\to\infty for any fixed c>0c>0. But Lemma 2.1 entails that for sufficiently large cc and sufficiently small δ>0\delta>0,

If ϕn(xo)<Mn\phi_{n}(x_{o})<M_{n}, then xox_{o} is not an interior point of the closed, convex set {ϕn≥ϕn(xo)}\{\phi_{n}\geq\phi_{n}(x_{o})\}. Hence

with h(Q,xo)<1h(Q,x_{o})<1 defined in Lemma 2.13. In the case of ϕn(xo)=Mn\phi_{n}(x_{o})=M_{n} these inequalities are true as well. Thus

which establishes (5). Combining (5) with ϕn≤M∗\phi_{n}\leq M_{*}, we may deduce from Lemma 3.3 of Schuhmacher, Hüsler and Dümbgen (2009) that there exist constants aa and b>0b>0 such that

Since the boundary of csupp⁡(Q)\operatorname{csupp}(Q) has Lebesgue measure zero, it follows from dominated convergence that ∫eψ(x) dx=1\int e^{\psi(x)}\,dx=1. Moreover, applying Fatou’s lemma to the nonnegative functions x↦a−b∥x∥−ϕn(k)(x)x\mapsto a-b\|x\|-\phi_{n(k)}(x) yields

In our proofs of Theorems 2.7 and 2.15 we utilize a special approximation scheme for functions in Φ\Phi:

For any function ϕ∈Φ\phi\in\Phi with nonempty domain and any parameter ε>0\varepsilon>0 set

Proof of Theorem 2.7 Let PP be the distribution corresponding to FF. Suppose first that ϕ=ψ(⋅∣Q)\phi=\psi(\cdot|Q). Then it follows from (4) and Fubini’s theorem that

It remains to be shown that ∫−∞x(F−G)(t) dt≥0\int_{-\infty}^{x}(F-G)(t)\,dt\geq 0 for x∈S(ϕ)x\in\mathcal{S}(\phi). Suppose first that x∈interior⁡(dom⁡(ϕ))x\in\operatorname{interior}(\operatorname{dom}(\phi)). Note that ϕ′:=ϕ′(⋅+)\phi^{\prime}:=\phi^{\prime}(\cdot+) is nonincreasing on the interior of dom⁡(ϕ)\operatorname{dom}(\phi) with

Moreover, x∈S(ϕ)x\in\mathcal{S}(\phi) implies that ϕ′(x−δ)>ϕ′(x+δ)\phi^{\prime}(x-\delta)>\phi^{\prime}(x+\delta) for all δ>0\delta>0 satisfying x±δ∈interior⁡(dom⁡(ϕ))x\pm\delta\in\operatorname{interior}(\operatorname{dom}(\phi)). For such δ>0\delta>0 we define

When x∈S(ϕ)x\in\mathcal{S}(\phi) is the left or right endpoint of dom⁡(ϕ)\operatorname{dom}(\phi), we define Δ(s):=(s−x)+\Delta(s):=(s-x)^{+} and conclude analogously that ∫−∞x(F−G)(t) dt≥0\int_{-\infty}^{x}(F-G)(t)\,dt\geq 0.

Since ∫(F−G)(t) dt=0\int(F-G)(t)\,dt=0, we may continue with

where the first displayed inequality follows from log-concavity of PP with log-density ϕ\phi. Thus ϕ=ψ\phi=\psi.

Theorem 2.14 and the second part of Theorem 2.15 are a consequence of the following result:

Let (Qn)n(Q_{n})_{n} be a sequence of distributions in Qo\mathcal{Q}_{o} such that Qn→wQ∈QoQ_{n}\to_{w}Q\in\mathcal{Q}_{o}, L(Qn)→λ∈[−∞,∞]L(Q_{n})\to\lambda\in[-\infty,\infty] and ∫∥x∥Qn(dx)→γ∈[0,∞]\int\|x\|Q_{n}(dx)\to\gamma\in[0,\infty] as n→∞n\to\infty. Then γ≥∫∥x∥Q(dx)\gamma\geq\int\|x\|Q(dx), and λ>−∞\lambda>-\infty if and only if γ<∞\gamma<\infty. Moreover,

In the latter case, the densities f:=\exp\mbox{{}\circ{}}\psi(\cdot|Q) and f_{n}:=\exp\mbox{{}\circ{}}\psi(\cdot|Q_{n}) are well defined for sufficiently large nn and satisfy

Before presenting the proof of this result, let us recall two elementary facts about weak convergence and unbounded functions:

If the stronger statement lim⁡n→∞∫h dQn=∫h dQ<∞\lim_{n\to\infty}\int h\,dQ_{n}=\int h\,dQ<\infty holds, then

Proof of Theorem 4.4 The asserted inequality γ≥∫∥x∥Q(dx)\gamma\geq\int\|x\|Q(dx) follows from the first part of Lemma 4.5 with h(x):=∥x∥h(x):=\|x\|.

Suppose that γ<∞\gamma<\infty. Then with ϕ(x):=−∥x∥\phi(x):=-\|x\|,

In other words, λ=−∞\lambda=-\infty entails that γ=∞\gamma=\infty.

This can be verified as follows: since L(Qn)=∫ψn dQn≤MnL(Q_{n})=\int\psi_{n}\,dQ_{n}\leq M_{n}, the sequence (Mn)n(M_{n})_{n} satisfies lim inf⁡n→∞Mn≥λ\liminf_{n\to\infty}M_{n}\geq\lambda. With similar arguments as in the proof of Theorem 2.2 one can deduce that (Mn)n(M_{n})_{n} is bounded from above, provided that

Another key property of the functions ψn\psi_{n} is that

by virtue of Lemma 2.13. Combining (5) with (7) we may again deduce that there exist constants aa and b>0b>0 such that

Thus λ<L(Q)\lambda<L(Q) if γ>∫∥x∥Q(dx)\gamma>\int\|x\|Q(dx).

By monotone convergence, applied to the functions ψ(1)−ψ(ε)\psi^{(1)}-\psi^{(\varepsilon)}, and dominated convergence, applied to \exp\mbox{{}\circ{}}\psi^{(\varepsilon)},

Note that the probability densities f=\exp\mbox{{}\circ{}}\psi and f_{n}=\exp\mbox{{}\circ{}}\psi_{n} obviously satisfy

In particular, (fn)n(f_{n})_{n} converges to ff almost everywhere w.r.t. Lebesgue measure, whence ∫∣fn(x)−f(x)∣ dx→0\int|f_{n}(x)-f(x)|\,dx\to 0.

where Bk:=I−uu⊤+kuu⊤B_{k}:=I-uu^{\top}+kuu^{\top} is a real, d×dd\times d matrix and ak:=−krua_{k}:=-kru. Note that det⁡(Bk)=k\det(B_{k})=k and ϕk(x)=log⁡(k)−∥x∥\phi_{k}(x)=\log(k)-\|x\| for x∈Hx\in H. Thus

and the right-hand side tends to infinity as ∥v∥→∞\|\mathbf{v}\|\to\infty. Thus it follows from Lemma 3.1 that

Our goal is to show that (f^n,μ^n)(\hat{f}_{n},\hat{\mu}_{n}), viewed as a function of \boldsεn\bolds{\varepsilon}_{n} and thus fixed, too, is well defined for sufficiently large nn with

Note that we replaced fnf_{n} with f=\exp\mbox{{}\circ{}}\psi(\cdot|Q) because ∫∣fn(x)−f(x)∣ dx\int|f_{n}(x)-f(x)|\,dx tends to .

where M^n:=Q^n,μ^n\hat{M}_{n}:=\hat{Q}_{n,\hat{\mu}_{n}} and R^n:=R(μn−μ^n)(xn)\hat{R}_{n}:=R_{(\mu_{n}-\hat{\mu}_{n})(\mathbf{x}_{n})}.

Note first that μˇn:=μn+∫tQ^n(dt)\check{\mu}_{n}:=\mu_{n}+\int t\hat{Q}_{n}(dt) belongs to M^n\hat{\mathcal{M}}_{n}, whence

Since (R^n)n(\hat{R}_{n})_{n} is tight, to verify (13) we may consider a subsequence (R^n(k))k(\hat{R}_{n(k)})_{k} that converges weakly to some distribution RR as k→∞k\to\infty. Then M^n(k)→wQ⋆R\hat{M}_{n(k)}\to_{w}Q\star R, so

by Theorems 2.14 and 3.5. Because of (14) we even know that L(M^n(k))→L(Q⋆R)=L(Q)L(\hat{M}_{n(k)})\to L(Q\star R)=L(Q) as k→∞k\to\infty. Consequently, we may deduce from Theorems 2.14 and 3.5 that

Hence ∫∣t∣R^n(k)(dt)\int|t|\hat{R}_{n(k)}(dt) is not greater than

as k→∞k\to\infty. As r↑∞r\uparrow\infty, the limit on the right-hand side converges to ∫∣t∣R(dt)=∣a∣\int|t|R(dt)=|a|. Consequently, lim⁡k→∞D1(R^n(k),R)=0\lim_{k\to\infty}D_{1}(\hat{R}_{n(k)},R)=0. But then 0=lim⁡k→∞∫tR^n(k)(dt)0=\lim_{k\to\infty}\int t\hat{R}_{n(k)}(dt) coincides with ∫tR(dt)=a\int tR(dt)=a.

In our proofs of Theorems 3.7 and 3.8 we utilize a simple inequality for the bounded Lipschitz distance in terms of the Kolmogorov–Smirnov distance,

of two distributions Q,Q′∈Q(1)Q,Q^{\prime}\in\mathcal{Q}(1):

Let QQ and Q′Q^{\prime} be distributions on the real line. Then for arbitrary r>0r>0,

Proof of Theorem 3.7 A key insight is that the empirical distributions Q^n,m\hat{Q}_{n,m} are close to their expectations Qn,mQ_{n,m} with respect to Kolmogorov–Smirnov distance, uniformly over all m∈Mnm\in\mathcal{M}_{n}. Namely,

for some universal constant CC [see Pollard (1990), Theorems 2.2 and 3.5, and van der Vaart and Wellner (1996), Theorem 2.6.4 and Lemma 2.6.16].

Acknowledgments

Constructive comments by an Associate Editor and two referees are gratefully acknowledged.

References