Multivariate Log-Concave Distributions as a Nearly Parametric Model

Dominic Schuhmacher, Andre Huesler, Lutz Duembgen

Introduction

It is well-known that certain statistical functionals such as moments fail to be weakly continuous on the set of, say, all probability measures on the real line for which these functionals are well-defined. This is the intrinsic reason why it is impossible to construct nontrivial two-sided confidence intervals for such functionals. For the mean and other moments, this fact was pointed out by Bahadur and Savage (1956). Donoho (1988) extended these considerations by noting that some functionals of interest are at least weakly semi-continuous, so that one-sided confidence bounds are possible.

When looking at the proofs of the results just mentioned, one realizes that they often involve rather strange, e.g. multimodal or heavy-tailed, distributions. Natural questions are whether statistical functionals such as moments become weakly continuous and whether honest confidence intervals exist for these functionals if attention is restricted to a suitable nonparametric class of distributions. For instance, one possibility would be to focus on distributions on a given bounded region. But this may be too restrictive or lead to rather conservative procedures.

Alternatively we propose a qualitative constraint. When asking a statistician to draw a typical probability density, she or he will often sketch a bell-shaped, maybe skewed density. This suggests unimodality as a constraint, but this would not rule out heavy tails. In the present paper we favor the stronger though natural constraint of log-concavity, also called strong unimodality. One should note here that additional assumptions such as given bounded support or log-concavity can never be strictly verified based on empirical data alone; see Donoho (1988, Section 2).

The remainder of this paper is organized as follows. In Section 2 we present our main result and some consequences, including an existence proof of non-trivial confidence sets for moments of log-concave distributions. Section 3 collects some basic inequalities for log-concave distributions which are essential for the main results and of independent interest. Most proofs are deferred to Section 4.

The main results

(i) The sequence (fn)(f_{n}) converges uniformly to ff on any closed set of continuity points of ff.

It is well-known from convex analysis that φ=log⁡f\varphi=\log f is continuous on int⁡({φ>−∞})=int⁡({f>0})\operatorname{int}(\{\varphi>-\infty\})=\operatorname{int}(\{f>0\}). Hence the discontinuity points of ff, if any, are contained in ∂{f>0}\partial\{f>0\}. But {f>0}\{f>0\} is a convex set, so its boundary has Lebesgue measure zero (cf. Lang 1986). Therefore Part (i) of Theorem 2.1 implies that (fn)n(f_{n})_{n} converges to ff pointwise almost everywhere.

Note also that f(x)≤C1exp⁡(−C2∥x∥)f(x)\leq C_{1}\exp(-C_{2}\|x\|) for suitable constants C1=C1(f)>0C_{1}=C_{1}(f)>0 and C2=C2(f)>0C_{2}=C_{2}(f)>0; see Corollary 3.4 in Section 3. Hence one may take A(x)=c∥x∥A(x)=c\|x\| for any c∈[0,C2)c\in[0,C_{2}) in order to satisfy (2.1). Theorem 2.1 is a multivariate version of Hüsler (2008, Theorem 2.1). It is also more general than findings of Cule and Samworth (2010) who treated the special case of A(x)=ϵ∥x∥A(x)=\epsilon\|x\| for some small ϵ>0\epsilon>0 with different techniques.

Note that for any dd-variate polynomial Π\Pi and arbitrary ϵ>0\epsilon>0 there exists an R=R(Π,ϵ)>0R=R(\Pi,\epsilon)>0 such that ∣Π(x)∣≤exp⁡(ϵ∥x∥)|\Pi(x)|\leq\exp(\epsilon\|x\|) for ∥x∥>R\|x\|>R. Hence part (ii) of Theorem 2.1 and Proposition 2.2 entail the first part of the following theorem:

It is well-known from empirical process theory (e.g. van der Vaart and Wellner 1996, Section 2.19) that for any α∈(0,1)\alpha\in(0,1) there exists a universal constant cα,dc_{\alpha,d} such that

Since convergence with respect to ∥⋅∥H\|\cdot\|_{\mathcal{H}} implies weak convergence, Theorem 2.3 implies the consistency of the confidence sets Cα,n(Π)C_{\alpha,n}^{(\Pi)}, in the sense that

Note that this construction proves existence of honest simultaneous confidence sets for arbitrary moments. But their explicit computation requires substantial additional work and is beyond the scope of the present paper.

If the right hand side is less than or equal to one, then

This lemma entails various upper bounds including a subexponential tail bound for log-concave densities.

2 Inequalities for dimension one

In the special case d=1d=1 we denote the cumulative distribution function of PP with FF. The hazard functions f/Ff/F and f/(1−F)f/(1-F) have the following properties:

The monotonicity properties of the hazard functions f/Ff/F and f/(1−F)f/(1-F) have been noted by An (1998) and Bagnoli and Bergstrom (2005) . For the reader’s convenience a complete proof of Lemma 3.5 will be given.

The next lemma provides an inequality for ff in terms of its first and second moments:

Equality holds if, and only if, ff is log-linear on both (−∞,xo](-\infty,x_{o}] and [xo,∞)[x_{o},\infty).

Proofs

Our proof of Lemma 3.1 is based on a particular representation of Lebesgue measure on simplices: Let

Then for any measurable function h:Δo→[0,∞)h:\Delta_{o}\to[0,\infty),

for some u∈Δou\in\Delta_{o}, where u0:=1−∑i=1duiu_{0}:=1-\sum_{i=1}^{d}u_{i}. In particular,

for any u=(ui)i=1d∈Δou=(u_{i})_{i=1}^{d}\in\Delta_{o} and u0=1−∑i=1duiu_{0}=1-\sum_{i=1}^{d}u_{i}. Hence

and by Jensen’s inequality, the latter expected value is not less than

This yields the first assertion of the lemma.

The inequality \prod_{i=0}^{d}f(x_{i})\leq\bigl{(}P(\Delta)/|\Delta|\bigr{)}^{d+1} may be rewritten as

We first prove Lemma 3.3 because this provides a tool for the proof of Lemma 3.2 as well.

Proof of Lemma 3.3.

At first we investigate how the size of Δ\Delta changes if we replace one of its vertices with another point. Note that for any fixed index j∈{0,1,…,d}j\in\{0,1,\ldots,d\},

Hence the set \Delta_{j}(y):=\operatorname{conv}\bigl{(}\{x_{i}:i\neq j\}\cup\{y\}\bigr{)} has Lebesgue measure

where σmax(X)>0\sigma_{\rm max}(X)>0 is the largest singular value of XX.

Now we consider any log-concave probability density ff. Let fminf_{\rm min} and fmaxf_{\rm max} denote the minimum and maximum, respectively, of {f(xi):i=0,…,d}\{f(x_{i}):i=0,\ldots,d\}, where fminf_{\rm min} is assumed to be greater than zero. Applying Lemma 3.1 to Δj(y)\Delta_{j}(y) in place of Δ\Delta with suitably chosen index jj, we may conclude that

where C=C(x0,…,xd):=∣Δ∣(d+1)−1/2σmax(X)−1C=C(x_{0},\ldots,x_{d}):=|\Delta|(d+1)^{-1/2}\sigma_{\rm max}(X)^{-1}. Moreover, in case of Cfmin(∥y∥2+1)1/2≥1Cf_{\rm min}(\|y\|^{2}+1)^{1/2}\geq 1,

Proof of Lemma 3.2.

Let y∈Δy\in\Delta, i.e. y=∑i=0dλixiy=\sum_{i=0}^{d}\lambda_{i}x_{i} with a unique vector λ=(λi)i=0d\lambda=(\lambda_{i})_{i=0}^{d} in d+1^{d+1} whose components sum to one. With Δj(y)\Delta_{j}(y) as in the proof of Lemma 3.3, elementary calculations reveal that

where J:={j:λj>0}J:=\{j:\lambda_{j}>0\}. Moreover, all these simplices Δj(y)\Delta_{j}(y), j∈Jj\in J, have nonvoid interior, and ∣Δj(y)∩Δk(y)∣=0|\Delta_{j}(y)\cap\Delta_{k}(y)|=0 for different j,k∈Jj,k\in J. Consequently it follows from Lemma 3.1 that

This entails the asserted upper bound for f(y)f(y). The lower bound follows from the elementary fact that any concave function on the simplex Δ\Delta attains its minimal value in one of the vertices x0,x1,…,xdx_{0},x_{1},\ldots,x_{d}. □\Box

Proof of Lemma 3.5.

Note that {F<1}=(−∞,tu)\{F<1\}=(-\infty,t_{u}). On {f=0}∩(−∞,tu)\{f=0\}\cap(-\infty,t_{u}), the function f/(1−F)f/(1-F) is equal to zero. For t∈{f>0}∩(−∞,tu)t\in\{f>0\}\cap(-\infty,t_{u}),

is non-decreasing in tt, because t↦φ(t+x)−φ(t)t\mapsto\varphi(t+x)-\varphi(t) is non-increasing in t∈{f>0}t\in\{f>0\} for any fixed x>0x>0, due to concavity of φ\varphi.

Proof of Lemma 3.6.

The asserted upper bound for f(to)f(t_{o}) is strictly positive and continuous in tot_{o}. Hence it suffices to consider a point tot_{o} with 0<F(to)<10<F(t_{o})<1. Since (xo−μ)2+σ2(x_{o}-\mu)^{2}+\sigma^{2} equals ∫(x−xo)2f(x) dx\int(x-x_{o})^{2}f(x)\,dx, we try to bound the latter integral from above. To this end, let gg be a piecewise loglinear probability density, namely,

with a:=f(xo)/F(xo)a:=f(x_{o})/F(x_{o}) and b:=f(xo)/(1−F(xo))b:=f(x_{o})/(1-F(x_{o})), so that

with equality if, and only if, f=gf=g. Now the assertion follows from

2 Proof of the main results

Note first that {f>0}\{f>0\} is a convex set with nonvoid interior. For notational convenience we may and will assume that

In our proof of Theorem 2.1, Part (i), we utilize two simple inequalities for log-concave densities:

Figure 4.1 illustrates the definition of the corner simplices and a key statement in the proof of Lemma 4.1.

This lemma involves three closed balls B(0,δ)B(0,\delta), B(ty,δt)B(ty,\delta_{t}) and B(y,δt)B(y,\delta_{t}); see Figure 4.2 for an illustration of these and the key argument of the proof.

Suppose that all corner simplices satisfy P(Δj)>0P(\Delta_{j})>0. Then for j=0,1,…,dj=0,1,\ldots,d there exists an interior point zjz_{j} of Δj\Delta_{j} with f(zj)>0f(z_{j})>0, that means, zj=2xj−∑i=0dλijxiz_{j}=2x_{j}-\sum_{i=0}^{d}\lambda_{ij}x_{i} with positive numbers λij\lambda_{ij} such that ∑i=0dλij=1\sum_{i=0}^{d}\lambda_{ij}=1. With the matrices

But the matrix 2I−Λ2I-\Lambda is nonsingular with inverse

Since min⁡x∈Δf(x)≤P(Δ)/∣Δ∣≤max⁡x∈Δf(x)\min_{x\in\Delta}f(x)\leq P(\Delta)/|\Delta|\leq\max_{x\in\Delta}f(x), the inequalities

are obvious. By concavity of φ\varphi, its minimum over Δ\Delta equals φ(xjo)\varphi(x_{j_{o}}) for some index jo∈{0,1,…,d}j_{o}\in\{0,1,\ldots,d\}. But then for arbitrary x∈Δx\in\Delta and y:=2xjo−x∈Δjoy:=2x_{j_{o}}-x\in\Delta_{j_{o}}, it follows from xjo=2−1(x+y)x_{j_{o}}=2^{-1}(x+y) and concavity of φ\varphi that

so that φ≤φ(xjo)\varphi\leq\varphi(x_{j_{o}}) on Δjo\Delta_{j_{o}}. Hence

Proof of Lemma 4.2.

The main point is to show that for any point x∈B(y,δt)x\in B(y,\delta_{t}),

i.e. any point w∈B(ty,δt)w\in B(ty,\delta_{t}) may be written as (1−t)v+tx(1-t)v+tx for a suitable v∈B(0,δ)v\in B(0,\delta); see also Figure 4.2. But note that the equation (1−t)v+tx=w(1-t)v+tx=w is equivalent to v=(1−t)−1(w−tx)v=(1-t)^{-1}(w-tx). This vector vv belongs indeed to B(0,δ)B(0,\delta), because

This consideration shows that for any point x∈B(y,δt)x\in B(y,\delta_{t}) and any point w∈B(ty,δt)w\in B(ty,\delta_{t}),

with v=(1−t)−1(w−tx)∈B(0,δ)v=(1-t)^{-1}(w-tx)\in B(0,\delta) and J0:=inf⁡v∈B(0,δ)f(v)J_{0}:=\inf_{v\in B(0,\delta)}f(v). Averaging this inequality with respect to w∈B(ty,δt)w\in B(ty,\delta_{t}) yields

Since x∈B(y,δt)x\in B(y,\delta_{t}) is arbitrary, this entails the assertion of Lemma 4.2. □\Box

Proof of Theorem 2.1, Part (i).

Step 1:

The sequence (fn)n(f_{n})_{n} converges to ff uniformly on any compact subset of int⁡{f>0}\operatorname{int}\{f>0\}.

By compactness, this claim is a consequence of the following statement: For any interior point yy of {f>0}\{f>0\} and any η>0\eta>0 there exists a neighborhood Δ(y,η)\Delta(y,\eta) of yy such that

To prove the latter statement, fix any number ϵ∈(0,1)\epsilon\in(0,1). Since ff is continuous on int⁡{f>0}\operatorname{int}\{f>0\}, there exists a simplex Δ=conv⁡{x0,x1,…,xd}\Delta=\operatorname{conv}\{x_{0},x_{1},\ldots,x_{d}\} such that y∈int⁡Δy\in\operatorname{int}\Delta and

For ϵ\epsilon sufficiently small, both (1−ϵ)/(1+ϵ)≥1−η(1-\epsilon)/(1+\epsilon)\geq 1-\eta and \bigl{(}(1+\epsilon)/(1-\epsilon)\bigr{)}^{d+1}\leq 1+\eta, which proves the assertion of step 1.

Step 2:

For this step we employ Lemma 4.2. Let δ0>0\delta_{0}>0 such that B(0,δ0)B(0,\delta_{0}) is contained in int⁡{f>0}\operatorname{int}\{f>0\}. Furthermore, let J0>0J_{0}>0 be the minimum of ff over B(0,δ0)B(0,\delta_{0}). Then step 1 entails that

Moreover, for any t∈(0,1)t\in(0,1) and δt:=(1−t)δ0/(1+t)\delta_{t}:=(1-t)\delta_{0}/(1+t),

But the latter bound tends to zero as t↑1t\uparrow 1.

Final step:

(fn)n(f_{n})_{n} converges to ff uniformly on any closed set of continuity points of ff.

Let SS be such a closed set. Then Steps 1 and 2 entail that

for any fixed ρ≥0\rho\geq 0, because S∩B(0,ρ)S\cap B(0,\rho) is compact, and any point y∈S∖int⁡{f>0}y\in S\setminus\operatorname{int}\{f>0\} satisfies f(y)=0f(y)=0.

On the other hand, let Δ\Delta be a nondegenerate simplex with corners x0,x1,…,xd∈int⁡{f>0}x_{0},x_{1},\ldots,x_{d}\in\operatorname{int}\{f>0\}. Step 1 also implies that lim⁡n→∞fn(xi)=f(xi)\lim_{n\to\infty}f_{n}(x_{i})=f(x_{i}) for i=0,1,…,di=0,1,\ldots,d, so that Lemma 3.3 entails that

for any ρ≥0\rho\geq 0 with a constant C=C(x0,…,xd)>0C=C(x_{0},\ldots,x_{d})>0. Since this bound tends to zero as ρ→∞\rho\to\infty, the assertion of Theorem 2.1, Part (i) follows. □\Box

Our proof of Theorem 2.1, Part (ii), is based on Part (i) and an elementary result about convex sets:

Proof of Lemma 4.3.

By convexity of C\mathcal{C} and B(0,δ)⊂CB(0,\delta)\subset\mathcal{C}, it follows from y∈Cy\in\mathcal{C} that

for any t∈t\in. In case of y∉Cy\not\in\mathcal{C}, for λ≥1\lambda\geq 1 and arbitrary x∈B(λy,(λ−1)δ)x\in B(\lambda y,(\lambda-1)\delta) we write x=λy+(λ−1)vx=\lambda y+(\lambda-1)v with v∈B(0,δ)v\in B(0,\delta). But then

Hence y∉Cy\not\in\mathcal{C} is a convex combination of a point in B(0,δ)⊂CB(0,\delta)\subset\mathcal{C} and xx, so that x∉Cx\not\in\mathcal{C}, too. □\Box

Proof of Theorem 2.1, Part (ii).

It follows from (4.5) in the proof of Part (i) with ρ=0\rho=0 that

It follows from Assumption (2.1) that for a suitable ρ>0\rho>0,

According to Part (i), (fn)n(f_{n})_{n} converges to ff uniformly on KK. Thus for fixed numbers ϵ′>0\epsilon^{\prime}>0, ϵ′′∈(0,ρ−1)\epsilon^{\prime\prime}\in(0,\rho^{-1}) and sufficiently large nn, the log-densities φn:=log⁡fn\varphi_{n}:=\log f_{n} satisfy the following inequalities:

Proof of Proposition 2.2.

Proof of Theorem 2.3.

and the right hand side tends to infinity as r↑∞r\uparrow\infty. □\Box

References