Nonparametric estimation of multivariate convex-transformed densities

Arseni Seregin, Jon A. Wellner

Introduction and background

Because the class of multivariate log-concave densities contains the class of multivariate normal densities and is preserved under a number of important operations (such as convolution and marginalization), it serves as a valuable nonparametric surrogate or replacement for the class of normal densities. Further study of the class of log-concave densities from this perspective has been undertaken by Schuhmacher, Hüsler and Duembgen (2009).

Such generalizations of log-concave densities and log-concave measures based on means of order rr have been introduced by a series of authors, sometimes with differing terminology, apparently starting with Avriel (1972), and continuing with Borell (1975), Brascamp and Lieb (1976), Prékopa (1973), Rinott (1976) and Uhrin (1984). A nice summary of these connections is given by Dharmadhikari and Joag-Dev (1988). These authors also present results concerning the preservation of rr-concavity under a variety of operations, including products, convolutions and marginalization.

Despite the longstanding and current rapid development of the properties of such classes of densities on the probability side, very little has been done from the standpoint of nonparametric estimation, especially when d≥2d\geq 2.

Here is an outline of the rest of the paper. All of our main results are presented in Section 2. Section 2.1 gives definitions and basic properties of the transformations involved. Section 2.2 establishes existence of the maximum likelihood estimators for both increasing and decreasing transformations hh under suitable conditions on the function hh. In Section 2.3, we give statements concerning consistency of the estimators, both in the Hellinger metric and in uniform metrics under natural conditions. In Section 2.4, we present asymptotic minimax lower bounds for estimation in these classes under natural curvature hypotheses. We conclude the section with a brief discussion of conjectures concerning attainability of the minimax rates by the maximum likelihood estimators. All of the proofs are given in Section 3.

Supplementary material and some proofs omitted here are available in Seregin and Wellner (2010). There, we also summarize a number of definitions and key results from convex analysis in an Appendix, Section A. We use standard notation from convex analysis; see “Notation” for a (partial) list.

2 Convex-transformed density estimation

Main results

the function h(y)h(y) is o(∣y∣−α)o(|y|^{-\alpha}) for some α>d\alpha>d as y→−∞y\to-\infty;

if y∞<+∞y_{\infty}<+\infty, then h(y)≍(y∞−y)−βh(y)\asymp(y_{\infty}-y)^{-\beta} for some β>d\beta>d as y↑y∞y\uparrow y_{\infty};

the function hh is continuously differentiable on the interval (y0,y∞)(y_{0},y_{\infty}).

Note that the assumption (I.1) is satisfied if y0>−∞y_{0}>-\infty.

the function h(y)h(y) is o(y−α)o(y^{-\alpha}) for some α>d\alpha>d as y→+∞y\to+\infty;

if y∞>−∞y_{\infty}>-\infty, then h(y)≍(y−y∞)−βh(y)\asymp(y-y_{\infty})^{-\beta} for some β>d\beta>d as y↓y∞y\downarrow y_{\infty};

if y∞=−∞y_{\infty}=-\infty, then h(y)γh(−Cy)=o(1)h(y)^{\gamma}h(-Cy)=o(1) for some γ,C>0\gamma,C>0 as y→−∞y\to-\infty;

the function hh is continuously differentiable on the interval (y∞,y0)(y_{\infty},y_{0}).

Note that the assumption (D.1) is satisfied if y0<+∞y_{0}<+\infty. We now define the decreasing class of densities P(h)\mathcal{P}(h).

For a monotone transformation hh, we denote by G(h)\mathcal{G}(h) the class of all closed proper convex functions gg such that h∘gh\circ g belongs to a monotone class P(h)\mathcal{P}(h). The following lemma allows us to compare models defined by increasing or decreasing transformations hh.

Consider two decreasing (or increasing) models P(h1)\mathcal{P}(h_{1}) and P(h2)\mathcal{P}(h_{2}). If h1=h2∘fh_{1}=h_{2}\circ f for some convex function ff, then P(h1)⊆P(h2)\mathcal{P}(h_{1})\subseteq\mathcal{P}(h_{2}).

The argument below is for a decreasing model. For an increasing model, the proof is similar. If f(x)>f(y)f(x)>f(y) for some x<yx<y, then ff is decreasing on (−∞,x)(-\infty,x), f(−∞)=+∞f(-\infty)=+\infty and therefore h2h_{2} is constant on (f(x),+∞)(f(x),+\infty), and we can redefine f(y)=f(x)f(y)=f(x) for all y<xy<x. Thus, we can always assume that ff is nondecreasing.

For any convex function gg, the function f∘gf\circ g is also convex. Therefore, if p=h1∘g∈P(h1)p=h_{1}\circ g\in\mathcal{P}(h_{1}), then p=h2∘f∘g∈P(h2)p=h_{2}\circ f\circ g\in\mathcal{P}(h_{2}).

In this section, we discuss several examples of monotone models. The first two families are based on increasing transformations hh.

This increasing model is defined by h(y)=eyh(y)=e^{y}. Limit points are y0=−∞y_{0}=-\infty and y∞=∞y_{\infty}=\infty. Assumption (I.1) holds for any α>d\alpha>d. These classes of densities were considered by An (1998), who established several useful preservation properties. In particular, log-convexity is preserved under mixtures [An (1998), Proposition 3] and under marginalization [An (1998), Remark 8, page 361].

We now consider some models based on decreasing transformations hh.

This decreasing model is defined by the transform h(y)=e−yh(y)=e^{-y}. Limit points are y0=+∞y_{0}=+\infty and y∞=−∞y_{\infty}=-\infty. Assumption (D.1) holds for any α>d\alpha>d. Assumption (D.3) holds for any γ>C>0\gamma>C>0.

Many parametric models are subsets of this model: in particular, uniform, Gaussian, gamma, beta, Gumbel, Fréchet and logistic densities are all log-concave.

This family of decreasing models is defined by the transforms h(y)=y+−sh(y)=y_{+}^{-s} for s>ds>d. Limit points are y0=+∞y_{0}=+\infty and y∞=0y_{\infty}=0. Assumption (D.1) holds for any α∈(d,s)\alpha\in(d,s). Assumption (D.2) holds for β=s\beta=s. As noted in Section 1, the model P(y+1/r)=P(y+−s)\mathcal{P}(y_{+}^{1/r})=\mathcal{P}(y_{+}^{-s}) (with r=−1/s<0r=-1/s<0) corresponds to the class of rr-concave densities. From Lemma 2.5, we have the following inclusion:

The models defined by power transformations include some parametric models with heavier-than-exponential tails. Several examples, including the multivariate generalizations of Pareto, Student-tt, and FF-distributions are discussed in Borell (1975)—none of these families are log-concave; see Johnson and Kotz (1972) and Seregin and Wellner (2010) for explicit computations.

Borell (1975) developed a framework which unifies log-concave and power-convex densities and gives an interesting characterization for these classes. Here, we briefly state the main result.

holds for all ∅≠A,B⊆C\varnothing\neq A,B\subseteq C and all θ∈(0,1)\theta\in(0,1). We define Ms∘(C)\mathcal{M}^{\circ}_{s}(C) as a subfamily of Ms(C)\mathcal{M}_{s}(C) which consists of probability measures such that the affine hull of its support has dimension dd. Here, μ∗\mu_{*} is the inner measure corresponding to μ\mu and the cases s=0,∞s=0,\infty are defined by continuity.

One of the main results of Borell (1975), Prékopa (1973) and Rinott (1976) is as follows.

Theorem 2.11 provides a special case of what has come to be known as the Borell–Brascamp–Lieb inequality; see, for example, Dharmadhikari and Joag-Dev (1988) and Brascamp and Lieb (1976). The current terminology is apparently due to Cordero-Erausquin, McCann and Schmuckenschläger (2001).

2 Existence of the maximum likelihood estimators

is the maximum likelihood estimator of pp over the class P(h)\mathcal{P}(h), assuming it exists and is unique. We also write g^n\hat{g}_{n} for the MLE of gg. We first state our main results concerning existence and uniqueness of the MLEs for the classes P(h)\mathcal{P}(h).

Suppose that hh is an increasing transformation satisfying assumptions (I.1)–(I.3). The MLE p^n\hat{p}_{n} then exists almost surely for the model P(h)\mathcal{P}(h).

Suppose that hh is a decreasing transformation satisfying assumptions (D.1)–(D.4). The MLE p^n\hat{p}_{n} then exists almost surely for the model P(h)\mathcal{P}(h) if

Uniqueness of the MLE is known for the log-concave model P(e−y)\mathcal{P}(e^{-y}); see, for example, Dümbgen and Rufibach (2009) for d=1d=1 and Cule, Samworth and Stewart (2010) for d≥1d\geq 1. For a brief further comment, see Section 2.5.

3 Consistency of the maximum likelihood estimators

Once existence of the MLEs is ensured, our attention shifts to other properties of the estimators: our main concern in this subsection is consistency. While, for a decreasing model, it is possible to prove consistency without any restrictions, for an increasing model, we need the following assumptions about the true density h∘g0h\circ g_{0}:

the function g0g_{0} is bounded by some constant C<y∞C<y_{\infty};

Note that for d=1d=1, the assumption (I.5) follows from assumption (I.4) and integrability of log⁡(1/x)\log(1/x) at zero. This assumption is also true if PP has finite marginal densities.

Our main results about increasing models are as follows.

For an increasing model P(h)\mathcal{P}(h), where hh satisfies assumptions (I.1)–(I.3) and for the true density h∘g0h\circ g_{0} which satisfies assumptions (I.4)–(I.6), the sequence of MLEs {p^n=h∘g^n}\{\hat{p}_{n}=h\circ\hat{g}_{n}\} is Hellinger consistent: H(p^n,p0)=H(h∘g^n,h∘g0)→a.s.0H(\hat{p}_{n},p_{0})=H(h\circ\hat{g}_{n},h\circ g_{0})\to_{a.s.}0.

The results about decreasing models can be formulated in a similar way.

For a decreasing model P(h)\mathcal{P}(h), where hh satisfies assumptions (D.1)–(D.4), the sequence of MLEs {p^n=h∘g^n}\{\hat{p}_{n}=h\circ\hat{g}_{n}\} is Hellinger consistent:

4 Local asymptotic minimax lower bounds

In this section, we establish local asymptotic minimax lower bounds for any estimator of several functionals of interest on the family P(h)\mathcal{P}(h) of convex-transformed densities. We start with several general results following Jongbloed (2000) and then apply them to estimation at a fixed point and to mode estimation.

First, we define minimax risk as in Donoho and Liu (1991).

where tnt_{n} ranges over all possible estimators of TpTp based on X1,…,XnX_{1},\ldots,X_{n}.

The main result (Theorem 1) in Jongbloed (2000) can be formulated as follows.

Let {pn}\{p_{n}\} be a sequence of densities in P\mathcal{P} such that lim sup⁡n→∞nH(pn,p)≤τ\limsup_{n\to\infty}\sqrt{n}H(p_{n},p)\leq\tau for some density pp in P\mathcal{P}. Then,

It will be convenient to reformulate this result in the following form.

Suppose that for any ε>0\varepsilon>0 small enough, there exists pε∈Pp_{\varepsilon}\in\mathcal{P} such that for some r>0r>0, lim⁡ε→0ε−1∣Tpε−Tp∣=1\lim_{\varepsilon\to 0}\varepsilon^{-1}|Tp_{\varepsilon}-Tp|=1 and

There then exists a sequence {pn}\{p_{n}\} such that

where R1R_{1} is the risk which corresponds to l(x)=∣x∣l(x)=|x|.

Corollary 2.21 shows that for a fixed change in the value of the functional TT, a family pεp_{\varepsilon} which is closer to the true density pp with respect to Hellinger distance provides a sharper lower bound. This suggests that for the functional TT which depends only on the local structure of the density, we would like our family {pε}\{p_{\varepsilon}\} to deviate from pp also locally. Below, we formally define such local deviations.

We call a family of measurable functions {pε}\{p_{\varepsilon}\} a deformation of a measurable function pp if pεp_{\varepsilon} is defined for any ε>0\varepsilon>0 small enough, lim⁡ε→0ess⁡sup⁡∣p−pε∣=0\lim_{\varepsilon\to 0}\mathop{\operatorname{ess}\sup}|p-p_{\varepsilon}|=0 and there exists a bounded family of real numbers rεr_{\varepsilon} and a point x0x_{0} such that

If, in addition, we have lim⁡ε→0rε=0\lim_{\varepsilon\to 0}r_{\varepsilon}=0, then we say that {pε}\{p_{\varepsilon}\} is a local deformation at x0x_{0}.

Since, for a deformation pεp_{\varepsilon}, we have μ[supp⁡∣pε(x)−p(x)∣]>0\mu[{\operatorname{supp}}|p_{\varepsilon}(x)-p(x)|]>0 for every ε>0\varepsilon>0, there exists δ>0\delta>0 such that μ{x\dvtx∣pε(x)−p(x)∣>δ}>0\mu\{x\dvtx|p_{\varepsilon}(x)-p(x)|>\delta\}>0 and thus the LrL_{r}-distance from pεp_{\varepsilon} to pp is positive for all ε>0\varepsilon>0. Note that this is always true if pp and pεp_{\varepsilon} are continuous at x0x_{0} and pε(x0)≠p(x0)p_{\varepsilon}(x_{0})\neq p(x_{0}).

We can now state our lower bound for estimation of the convex-transformed density value at a fixed point x0x_{0}. This result relies on the properties of strongly convex functions, as described in Appendix S.A.4, and can be applied to both increasing and decreasing classes of convex-transformed densities.

where the constant C(d)C(d) depends only on the dimension dd.

If, in addition, gg is twice continuously differentiable at x0x_{0} and ∇2g(x0)\nabla^{2}g(x_{0}) is positive definite, then, by Lemma S.A.22, we have curv⁡x0g=det⁡(∇2g(x0))\operatorname{curv}_{x_{0}}g=\det(\nabla^{2}g(x_{0})).

and the analog of Corollary 2.21 then has the following form.

Suppose that for any ε>0\varepsilon>0 small enough, there exists pε∈Pp_{\varepsilon}\in\mathcal{P} such that for some r>0r>0,

There then exists a sequence {pn}\{p_{n}\} such that

The construction of a lower bound for the functional TT is similar to the procedure we presented for estimation of p=h∘gp=h\circ g at a fixed point x0x_{0}. Again, we use two opposite deformations: one is local and changes the functional value, the other is a convex combination with a fixed deformation and negligible-in-Hellinger-distance computation. However, in this case, the minimax rate also depends on the growth rate of gg.

where the constant C(d)C(d) depends only on the dimension dd, and the metric s(x,y)s(x,y) is defined as ∥x−y∥\|x-y\|.

If, in addition, gg is twice continuously differentiable at x0x_{0} and ∇2g(x0)\nabla^{2}g(x_{0}) is positive definite, then, by Lemma S.A.22, we have curv⁡x0g=det⁡(∇2g(x0))\operatorname{curv}_{x_{0}}g=\det(\nabla^{2}g(x_{0})) and gg is locally Hölder continuous at x0x_{0} with exponent γ=2\gamma=2 and any constant L>∥∇2g(x0)∥L>\|\nabla^{2}g(x_{0})\|.

Since curv⁡x0g>0\operatorname{curv}_{x_{0}}g>0, there exists a constant CC such that C∥x−x0∥2≤∣g(x)−g(x0)∣C\|x-x_{0}\|^{2}\leq|g(x)-g(x_{0})| and thus we have γ∈(0,2]\gamma\in(0,2].

5 Conjectures concerning uniqueness of MLEs

There exist counterexamples to uniqueness for nonconvex transformations hh which satisfy assumptions (D.1)–(D.4). They suggest that uniqueness of the MLE does not depend on the tail behavior of the transformation hh, but rather on the local properties of hh in neighborhoods of the optimal values g^n(Xi)\hat{g}_{n}(X_{i}). We conjecture that uniqueness holds for all monotone models if hh is convex and h/∣h′∣h/|h^{\prime}| is nondecreasing convex. Further work on these uniqueness issues is needed.

6 Conjectures about rates of convergence for the MLEs

We conjecture that the (optimal) rate of convergence n2/(d+4)n^{2/(d+4)} appearing in Theorem 2.23 for estimation of f(x0)f(x_{0}) will be achieved by the MLE only for d=2,3d=2,3. For d=4d=4, we conjecture that the MLE will come within a factor (log⁡n)−γ(\log n)^{-\gamma} (for some γ>0\gamma>0) of achieving the rate n1/4n^{1/4}, but for d>4d>4, we conjecture that the rate of convergence will be the suboptimal rate n1/dn^{1/d}. This conjectured rate-suboptimality raises several interesting further issues:

Can we find alternative estimators (perhaps via penalization or sieve methods) which achieve the optimal rates of convergence?

For interesting subclasses, do maximum likelihood estimators remain rate-optimal?

Proofs

for y<+∞y<+\infty, the sublevel sets lev⁡yg\operatorname{lev}_{y}g are bounded and we have

The sublevel set lev⁡yg\operatorname{lev}_{y}g has the same dimension as dom⁡g\mathop{\operatorname{dom}}g [Theorem 7.6 in Rockafellar (1970)], which is dd. By Lemma S.A.1, this set is bounded when y<y0y<y_{0}. Therefore, it is enough to prove that lev⁡y0g\operatorname{lev}_{y_{0}}g is bounded for y0<+∞y_{0}<+\infty.

Since h∘gh\circ g is a density, we have inf⁡g<y0\inf g<y_{0}. If gg is constant on dom⁡g\mathop{\operatorname{dom}}g, then, for all y∈[inf⁡g,+∞)y\in[\inf g,+\infty), we have lev⁡yg=lev⁡inf⁡gh\operatorname{lev}_{y}g=\operatorname{lev}_{\inf g}h and it is therefore bounded. Otherwise, we can choose inf⁡h≤y1<y2<y0\inf h\leq y_{1}<y_{2}<y_{0}. Then, μ[lev⁡y2g]<∞\mu[\operatorname{lev}_{y_{2}}g]<\infty and, by Lemma S.A.3, we have μ[lev⁡y0g]<∞\mu[\operatorname{lev}_{y_{0}}g]<\infty. The argument above shows that lev⁡y0g\operatorname{lev}_{y_{0}}g is also bounded.

2. This follows from the fact that gg is continuous and lev⁡yg\operatorname{lev}_{y}g is bounded and nonempty for y>inf⁡gy>\inf g.

Using the Fubini–Tonelli theorem, we have, with Lac≡(lev⁡ag)cL_{a}^{c}\equiv(\operatorname{lev}_{a}g)^{c},

By assumption (D.1), the function −[hlog⁡h](y)-[h\log h](y) is decreasing to zero as y→+∞y\to+\infty and we have 0<−[hlog⁡h](y)<Cy−d−α′0<-[h\log h](y)<Cy^{-d-\alpha^{\prime}} for CC large enough and α′∈(0,α)\alpha^{\prime}\in(0,\alpha) as y→+∞y\rightarrow+\infty.

By Lemma 3.1, the level sets lev⁡yg\operatorname{lev}_{y}g are bounded and since h∘g∈P(h)h\circ g\in\mathcal{P}(h), we have inf⁡g>y∞\inf g>y_{\infty}. Therefore, the integral exists if and only if the integral

for some a>y∞a>y_{\infty}. Choosing aa large enough and using Lemma 3.2 for the decreasing transformation h1(y)=y−d−α′h_{1}(y)=y^{-d-\alpha^{\prime}}, we obtain

By Lemma S.A.3, we have μ(lev⁡yg)=O(yd)\mu(\operatorname{lev}_{y}g)=O(y^{d}) and therefore the last integral is finite.

Let BB be a ball such that K⊂BK\subset B. Let cc be such that h(c)=1/μ[B]h(c)=1/\mu[B]. The function g≡c+δ(⋅∣B)g\equiv c+\delta(\cdot\mathop{|}B) then belongs to G(h)\mathcal{G}(h).

2 Proofs for existence results

If an MLE for the model P(h)\mathcal{P}(h) exists, then it maximizes the functional

Here are the corresponding results for decreasing transformations hh.

The bounds provided by the following key lemma are the remaining preparatory work for proving existence of the MLE in the case of increasing transformations.

By Lemma S.1.1, we have h∘g(Xi)≤d!ddV(Xi)h\circ g(X_{i})\leq\frac{d!}{d^{d}V({X_{i}})}, which gives the upper bounds C(Xi,X,ε)C(X_{i},X,\varepsilon). By assumption, we have

which gives the uniform lower bound c(Xi,X,ε)c(X_{i},X,\varepsilon) for all Xi∈XX_{i}\in X. Since, by Lemma S.1.1, g(0)≥g(Xi)g(0)\geq g(X_{i}), we also obtain c(0,X,ε)c(0,X,\varepsilon).

If y∞=+∞y_{\infty}=+\infty, then, for a fixed y∈(c,+∞)y\in(c,+\infty), we have

If y∞<+∞y_{\infty}<+\infty, then for a fixed y∈(c,y∞]y\in(c,y_{\infty}], we have

Thus, there exists s0∈(c,y∞)s_{0}\in(c,y_{\infty}) such that T(s0)>1T(s_{0})>1. This implies that g(0)<s0g(0)<s_{0}. Since s0s_{0} depends only on XaX_{a} and min⁡c(Xi,X,ε)\min c(X_{i},X,\varepsilon), this gives an upper bound C(0,X,ε)C(0,X,\varepsilon).

Since we have only a finite number of possible choices for XaX_{a}, we have obtained c(x0,X,ε)c(x_{0},X,\varepsilon), which completes the proof.

Before proving existence of the MLE for a decreasing transformation family, we need two lemmas.

Consider a decreasing model P(h)\mathcal{P}(h). Let {gk}\{g_{k}\} be a sequence of convex functions from G(h)\mathcal{G}(h) and let {nk}\{n_{k}\} be a nondecreasing sequence of positive integers nk≥ndn_{k}\geq n_{d} such that for some ε>−∞\varepsilon>-\infty and ρ>0\rho>0, the following is true:

There then exists m>y∞m>y_{\infty} such that gk≥mg_{k}\geq m for all kk.

For given observations X=(X1,…,Xn)X=(X_{1},\ldots,X_{n}) such that n≥ndn\geq n_{d}, there exist constants m>y∞m>y_{\infty} and MM which depend only on observations XX and ε\varepsilon such that for any g∈N(h,X,ε)g\in\mathcal{N}(h,X,\varepsilon), we have m≤g(x)≤Mm\leq g(x)\leq M on conv⁡(X)\operatorname{conv}(X).

An arbitrary sequence of functions {gk}\{g_{k}\} from N(h,X,ε)\mathcal{N}(h,X,\varepsilon) satisfies the conditions of Lemma 3.11 with nk≡nn_{k}\equiv n and the same ε\varepsilon and ρ\rho constructed above. Therefore, the sequence {gk}\{g_{k}\} is bounded below by some constant greater than y∞y_{\infty}. Thus, the family of functions N(h,X,ε)\mathcal{N}(h,X,\varepsilon) is uniformly bounded below by some m>y∞m>y_{\infty}.

Consider any g∈N(h,X,ε)g\in\mathcal{N}(h,X,\varepsilon). Let MgM_{g} be the supremum of gg on dom⁡h\mathop{\operatorname{dom}}h. By Theorem 32.2 in Rockafellar (1970), the supremum is obtained at some XM∈XX_{M}\in X and therefore Mg<y0M_{g}<y_{0}. Let mgm_{g} be the minimum of gg on XX. We have h(mg)n−1h(Mg)≥enεh(m_{g})^{n-1}h(M_{g})\geq e^{n\varepsilon} and

Thus, we have obtained an upper bound MM which depends only on mm and XX.

Finally, we have to add the “almost surely” clause since we assumed that the points XiX_{i} are in general position.

3 Proofs for consistency results

We begin with proofs for some technical results which we will use in the consistency arguments for both increasing and decreasing models. The main argument for proving Hellinger consistency proceeds along the lines of the proof given in the case of d=1d=1 by Pal, Woodroofe and Meyer (2007) and in the log-concave case for d>1d>1 by Schuhmacher and Duembgen (2010).

Consider a monotone model P(h)\mathcal{P}(h). Suppose that the true density h∘g0h\circ g_{0} and the sequence of MLEs {g^n}\{\hat{g}_{n}\} have the following properties:

for ε>0\varepsilon>0 small enough. The sequence of the MLEs is then Hellinger consistent: H(h∘g^n,h∘g0)→a.s.0H(h\circ\hat{g}_{n},h\circ g_{0})\to_{a.s.}0.

The next lemma allows us to obtain pointwise consistency once Hellinger consistency is proved.

Let us denote by La0L^{0}_{a} and LakL^{k}_{a} the sublevel sets La0=lev⁡ag0L^{0}_{a}=\operatorname{lev}_{a}g_{0} and Lan=lev⁡ag^nL^{n}_{a}=\operatorname{lev}_{a}\hat{g}_{n}, respectively. Consider Ω0\Omega_{0} such that Pr⁡[Ω0]=1\Pr[\Omega_{0}]=1 and H2(h∘g^nω,h∘g0)→0H^{2}(h\circ\hat{g}_{n}^{\omega},h\circ g_{0})\to 0, where g^nω\hat{g}_{n}^{\omega} is the MLE for ω∈Ω0\omega\in\Omega_{0}. For all ω∈Ω0\omega\in\Omega_{0}, we have

We need a general property of the bracketing entropy numbers.

Finally, we prove consistency for decreasing models. We need a general property of convex sets.

Let DD be a convex compact set. By Theorem 8.4.2 in Dudley (1999), the class A∩D\mathcal{A}\cap D has a finite set of ε\varepsilon-brackets. Since the class A\mathcal{A} is invariant under rescaling, the result follows from Lemma 3.15.

For a decreasing model P(h)\mathcal{P}(h), the sequence of MLEs g^n\hat{g}_{n} is almost surely uniformly bounded below.

We will apply Lemma 3.11 to the sequences g^n\hat{g}_{n} and {n}\{n\}. By the strong law of large numbers and Lemma 3.4, we have

Choose some a∈(0,d/nd)a\in(0,d/n_{d}). Then, for any set SS such that μ[S]=ρ≡a/h(min⁡g0)\mu[S]=\rho\equiv a/h(\min g_{0}), where min⁡g0\min g_{0} is attained by Lemma 3.1, we have

Now, let An=lev⁡ang^nA_{n}=\operatorname{lev}_{a_{n}}\hat{g}_{n} be sets such that μ[An]=ρ\mu[A_{n}]=\rho. Then, by Lemma 3.16, we have

By Lemma 3.17, we have inf⁡g^n≥A\inf\hat{g}_{n}\geq A for some A>y∞A>y_{\infty}. Therefore, by Lemma 3.2 applied to the decreasing transformation log⁡[ε+h(y)]−log⁡ε\log[\varepsilon+h(y)]-\log\varepsilon, it follows that

The closure Aˉ\bar{A} is compact and thus, for nn large enough, we have, with probability one, sup⁡Aˉ∣g^n−g0∣<δ{\sup_{\bar{A}}}|\hat{g}_{n}-g_{0}|<\delta, which implies that sup⁡Aˉ∣h∘g^n−h∘g0∣<ε{\sup_{\bar{A}}}|h\circ\hat{g}_{n}-h\circ g_{0}|<\varepsilon since the range of values of g0g_{0} on Aˉ\bar{A} is [m,a][m,a]. The set ∂A\partial A is compact and therefore g^n\hat{g}_{n} attains its minimum mnm_{n} on this set at some point xnx_{n}. By construction,

We have x0∈A∩lev⁡a−δg^nx_{0}\in A\cap\operatorname{lev}_{a-\delta}\hat{g}_{n} and g^n≥mn>a−δ\hat{g}_{n}\geq m_{n}>a-\delta on ∂A\partial A. Thus, by convexity, we have lev⁡a−δg^n⊂A\operatorname{lev}_{a-\delta}\hat{g}_{n}\subset A and for x∉Aˉx\notin\bar{A}, we have

This shows that for any ε>0\varepsilon>0 small enough, we will have

with probability one as n→∞n\to\infty. This concludes the proof.

4 Proofs for lower bound results

We will use the following lemma for computing the Hellinger distance between a function and its local deformation.

In order to apply Corollary 2.21, we need to construct deformations so that they still belong to the class G\mathcal{G}. The following lemma provides a technique for constructing such deformations.

Note that gθ,δg_{\theta,\delta} is not a local deformation of gg. {pf*}Proof of Theorem 2.23 Our statement is nontrivial only if the curvature curv⁡x0g>0\operatorname{curv}_{x_{0}}g>0 or, equivalently, there exists a positive definite d×dd\times d matrix GG such that the function gg is locally GG-strongly convex. Then, by Lemma S.A.17, this means that there exists a convex function qq such that, in some neighborhood O(x0)O(x_{0}) of x0x_{0}, we have

The plan of the proof is as follows: we introduce families of functions {Dε(g;x0,v)}\{D_{\varepsilon}(g;x_{0},v)\} and {Dε∗(g;x0)}\{D_{\varepsilon}^{*}(g;x_{0})\} and prove that these families are local deformations. Using these deformations as building blocks, we construct two types of deformations, {h∘gε+}\{h\circ g^{+}_{\varepsilon}\} and {h∘gε−}\{h\circ g^{-}_{\varepsilon}\}, of the density h∘gh\circ g, which belong to P(h)\mathcal{P}(h). These deformations represent positive and negative changes in the value of the function gg at the point x0x_{0}. We then approximate the Hellinger distances using Lemma 3.18. Finally, applying Corollary 2.21, we obtain lower bounds which depend on GG. We complete the proof by taking the supremum of the obtained lower bounds over all G∈SC(g;x0)G\in\mathcal{SC}(g;x_{0}). Under the mild assumption of strong convexity of the function gg, both deformations give the same rate and structure of the constant C(d)C(d). However, it is possible to obtain a larger constant C(d)C(d) for the negative deformation if we assume that gg is twice differentiable. Note that, by the definition of P(h)\mathcal{P}(h), the function gg is a closed proper convex function.

Let us define a function Dε(g;x0,v0)D_{\varepsilon}(g;x_{0},v_{0}) for a given ε>0\varepsilon>0, x0∈dom⁡gx_{0}\in\mathop{\operatorname{dom}}g and v0∈∂g(x0)v_{0}\in\partial g(x_{0}) as follows: Dε(g;x0,v0)(x)=max⁡(g(x),l0(x)+ε)D_{\varepsilon}(g;x_{0},v_{0})(x)=\max(g(x),l_{0}(x)+\varepsilon), where l0(x)=⟨v0,x−x0⟩+g(x0)l_{0}(x)=\langle v_{0},x-x_{0}\rangle+g(x_{0}) is a support plane to gg at x0x_{0} (see Figure 1). Since l0+εl_{0}+\varepsilon is a support plane to g+εg+\varepsilon, we have g≤Dε(g;x0,v0)≤g+εg\leq D_{\varepsilon}(g;x_{0},v_{0})\leq g+\varepsilon and thus dom⁡Dε(g;x0,v0)=dom⁡g\mathop{\operatorname{dom}}D_{\varepsilon}(g;x_{0},v_{0})=\mathop{\operatorname{dom}}g. As a maximum of two closed convex functions, Dε(g;x0,v0)D_{\varepsilon}(g;x_{0},v_{0}) is a closed convex function. For a given x1x_{1}, we have Dε(g;x0,v0)(x1)=g(x1)D_{\varepsilon}(g;x_{0},v_{0})(x_{1})=g(x_{1}) if and only if

see Figure 2. Both functions Dε(g;x0,v0)D_{\varepsilon}(g;x_{0},v_{0}) and Dε∗(g;x0)D^{*}_{\varepsilon}(g;x_{0}) are convex by construction and, as the next lemma shows, have similar

properties. However, the argument for Dε∗(g;x0)D^{*}_{\varepsilon}(g;x_{0}) is more complicated.

Dε∗(g;x0)D^{*}_{\varepsilon}(g;x_{0}) is a closed proper convex function such that g−ε≤Dε∗(g;x0)≤gg-\varepsilon\leq D^{*}_{\varepsilon}(g;x_{0})\leq g and dom⁡Dε∗(g;x0)=dom⁡g\mathop{\operatorname{dom}}D^{*}_{\varepsilon}(g;x_{0})=\mathop{\operatorname{dom}}g;

if v0∈∂g(x0)v_{0}\in\partial g(x_{0}), then x0∈∂g∗(v0)x_{0}\in\partial g^{*}(v_{0}) and Dε(g;x0,v0)=(Dε∗(g∗;v0))∗D_{\varepsilon}(g;x_{0},v_{0})=(D^{*}_{\varepsilon}(g^{*};v_{0}))^{*}.

Therefore, v∈∂g(x1)v\in\partial g(x_{1}). In particular,

We can represent Dε∗(g∗;x0)D_{\varepsilon}^{*}(g^{*};x_{0}) as the maximal convex minorant of gg defined by g=min⁡(g,g(x0)−ε+δ(⋅∣x0))g=\min(g,g(x_{0})-\varepsilon+\delta(\cdot|x_{0})). For x∈dom⁡gx\in\mathop{\operatorname{dom}}g, by Lemma S.A.10, g∗(v0)+g(x0)=⟨v0,x0⟩g^{*}(v_{0})+g(x_{0})=\langle v_{0},x_{0}\rangle. Thus,

for some v∈∂g(x0)v\in\partial g(x_{0}). By Lemma S.A.7, we have

Therefore, for the point x1x_{1} in the neighborhood O(x0)O(x_{0}) where the decomposition (16) is true, condition (17) is equivalent to

Since ⟨w0,x1−x0⟩+q(x0)\langle w_{0},x_{1}-x_{0}\rangle+q(x_{0}) is a support plane to q(x)q(x), the inequality (17) is satisfied if 2−1(x1−x0)TG(x1−x0)≥ε2^{-1}(x_{1}-x_{0})^{T}G(x_{1}-x_{0})\geq\varepsilon, which is the complement of an open ellipsoid BG(x0,2ε)B_{G}(x_{0},\sqrt{2\varepsilon}) defined by GG with center at x0x_{0}. For ε\varepsilon small enough, this ellipsoid will belong to the neighborhood O(x0)O(x_{0}). Since ∣Dε(g;x0,v0)−g∣≤ε|D_{\varepsilon}(g;x_{0},v_{0})-g|\leq\varepsilon, this proves that the family Dε(g;x0,v0)D_{\varepsilon}(g;x_{0},v_{0}) is a local deformation.

In the same way, the condition (18) is equivalent to

or 2−1(x1−x0)TG(x1−x0)+q(x0)−ε≥⟨w1,x0−x1⟩+q(x1)2^{-1}(x_{1}-x_{0})^{T}G(x_{1}-x_{0})+q(x_{0})-\varepsilon\geq\langle w_{1},x_{0}-x_{1}\rangle+q(x_{1}), which is satisfied if we have 2−1(x1−x0)TG(x1−x0)≥ε2^{-1}(x_{1}-x_{0})^{T}G(x_{1}-x_{0})\geq\varepsilon. Since ∣Dε∗(g;x0)−g∣≤ε|D_{\varepsilon}^{*}(g;x_{0})-g|\leq\varepsilon, this proves that the family Dε∗(g;x0)D_{\varepsilon}^{*}(g;x_{0}) is also a local deformation. Thus, we have proven the following.

For r>0r>0 small enough, h′∘g(x)h^{\prime}\circ g(x) is nonzero and the decomposition (16) is true on B(x0;r)B(x_{0};r). Let us fix some v0∈∂g(x0)v_{0}\in\partial g(x_{0}), some x1∈B(x0;r)x_{1}\in B(x_{0};r) such that x1≠x0x_{1}\neq x_{0} and some y1∈∂g(x1)y_{1}\in\partial g(x_{1}). We fix δ\delta such that equation (14) of Lemma 3.19 is true for the transformation h\sqrt{h} and r=2r=2, and also x0∉BG(x1;2δ)‾x_{0}\notin\overline{B_{G}(x_{1};\sqrt{2\delta})}. Then, by Lemma 3.21, for all ε>0\varepsilon>0 small enough, the support sets supp⁡[Dε(g;x0,v0)−g]\operatorname{supp}[D_{\varepsilon}(g;x_{0},v_{0})-g] and supp⁡[Dδ∗(g;x1)−g]\operatorname{supp}[D_{\delta}^{*}(g;x_{1})-g] do not intersect; that is, these two deformations do not interfere.

We can now prove Theorem 2.23. The argument below is identical for gε+g^{+}_{\varepsilon} and gε−g^{-}_{\varepsilon}, so we will give the proof only for gε+g^{+}_{\varepsilon}. We define deformations gε+g^{+}_{\varepsilon} and gε−g^{-}_{\varepsilon} by means of the following lemma.

For all ε>0\varepsilon>0 small enough, there exist θε+,θε−∈(0,1)\theta_{\varepsilon}^{+},\theta_{\varepsilon}^{-}\in(0,1) such that the functions gε+g^{+}_{\varepsilon} and gε−g^{-}_{\varepsilon} defined by

Next, we will show that θε+\theta_{\varepsilon}^{+} goes to zero fast enough so that gε+g^{+}_{\varepsilon} is very close to Dε(g;x0,v0)D_{\varepsilon}(g;x_{0},v_{0}). Since supports do not intersect, we have

where both integrals have the same sign. For the first integral, by Lemma 3.18, we have

The second integral is monotone in θε+\theta^{+}_{\varepsilon} and, by Lemma 3.19, we have

Thus, we have θε+=O(ε1+d/2)\theta_{\varepsilon}^{+}=O(\varepsilon^{1+d/2}) and

where S(0,1)S(0,1) is the dd-dimensional sphere of radius 11.

For the second part, by Lemma 3.19, we obtain

Taking the supremum over all G∈SC(g;x0)G\in\mathcal{SC}(g;x_{0}), we obtain the statement of the theorem.

5 Indications of proofs for conjectured rates

for all (small) ϵ>0\epsilon>0, for a constant KK depending only on CC and dd. Then, after an argument to transfer this covering number bound to a bracketing entropy bound for P(h)\mathcal{P}(h) with respect to Hellinger distance HH, it follows from oscillation bounds for empirical processes [cf. van der Vaart and Wellner (1996), Theorems 3.4.1 and 3.4.4] that rates of convergence of p^n\hat{p}_{n} with respect to Hellinger distance are determined by rn2ϕ(1/rn)≍nr_{n}^{2}\phi(1/r_{n})\asymp\sqrt{n} with

Assuming that the bound of (20) can be carried over to log⁡N[](ϵ,P(h),H)\log N_{[]}(\epsilon,\mathcal{P}(h),H) sufficiently closely, routine calculations show that the expected rates of convergence of p^n\hat{p}_{n} to p0=h(g0)p_{0}=h(g_{0}) with respect to Hellinger distance HH are

Based on these heuristics, we expect that the MLE p^n\hat{p}_{n} will be rate efficient if d≤3d\leq 3, but rate inefficient (not attaining the optimal rate n2/(d+4)n^{2/(d+4)}) if d≥4d\geq 4.

Case 1: d≤3d\leq 3. In this case, we find that

where M1≡(K(1+L)d/2)1/2/(1−d/4)M_{1}\equiv(K(1+L)^{d/2})^{1/2}/(1-d/4). Solving the relation rn2ϕ(1/rn)≍nr_{n}^{2}\phi(1/r_{n})\asymp\sqrt{n} for rnr_{n} yields rn=n2/(d+4)r_{n}=n^{2/(d+4)} up to a constant.

Case 2: d=4d=4. In this case, we find that

where M2≡(K(1+L)d/2)1/2M_{2}\equiv(K(1+L)^{d/2})^{1/2}. Solving the relation rn2ϕ(1/rn)≍nr_{n}^{2}\phi(1/r_{n})\asymp\sqrt{n} for rnr_{n} yields rn=(n/(log⁡n)2)1/4r_{n}=(n/(\log n)^{2})^{1/4} up to a constant.

Case 3: d>4d>4. In this case, we calculate

where M3≡(K(1+L)d/2)1/2/(d/4−1)M_{3}\equiv(K(1+L)^{d/2})^{1/2}/(d/4-1). Solving the relation rn2ϕ(1/rn)≍nr_{n}^{2}\phi(1/r_{n})\asymp\sqrt{n} for rnr_{n} yields rn=n1/dr_{n}=n^{1/d} up to a constant.

Acknowledgments

This research is part of the Ph.D. dissertation of the first author at the University of Washington. We would like to thank two referees for a number of helpful suggestions.

[id=supp] \snameSupplement \stitleOmitted Proofs and Some Facts from Convex Analysis \slink[doi]10.1214/10-AOS840 \sdatatype.pdf \sfilenameConvexTransfSupp-v4.pdf \sdescriptionIn the supplement, we provide omitted proofs and some basic facts from convex analysis used in this paper.

References