Theoretical properties of the log-concave maximum likelihood estimator of a multidimensional density
Madeleine Cule, Richard Samworth
Introduction
Although work on shape-constrained density estimation dates back to the celebrated paper of Grenander (1956) on the estimation of a decreasing density on the positive half-line, it is in recent years that the area has enjoyed its most significant interest. This is partly because algorithmic and technological advances now allow the computation of estimators that would not previously have been feasible, and partly because statisticians now have more tools at their disposal for the study of the theoretical properties of these estimators.
The attraction of the use of these estimators is that, in contrast to alternative nonparametric density estimation methods such as those based on kernels or wavelets, they provide fully automatic procedures, with no smoothing parameters to choose. Such procedures are particularly desirable when the data are multidimensional, and the choice of (often multiple) smoothing parameters is particularly problematic.
The properties of the Grenander estimator are now quite well understood, thanks to the works of Marshall and Proschan (1965), Prakasa Rao (1969), Devroye (1987), Birgé (1989), van de Geer (1993) and Balabdaoui et al. (2009). Other examples of shape constraints on univariate densities that have been studied in the literature include convexity (Groeneboom, Jongbloed and Wellner, 2001; Dümbgen, Rufibach and Wellner, 2007) and -monotonicity (Balabdaoui and Wellner, 2008). It is also known that a maximum likelihood estimator does not exist over the class of unimodal densities – cf. Birgé (1997).
In this paper, we study the statistical properties of the estimator. Importantly, our results do not assume that the underlying density is log-concave. To the best of our knowledge, such results have not been obtained before even for univariate data, but are of interest because in practice it is impossible to tell from a sample of data whether the assumption of log-concavity is satisfied. It is therefore natural to seek assurance that the estimator will behave sensibly if the condition is violated. In our main result (cf. Theorem 4 below), we prove under very mild conditions the existence and uniqueness of a log-concave density that minimises the Kullback–Leibler divergence from and show that there is an interval of values of for which
Convergence of log-concave densities
Notice in particular that if has density , then Lemma 1 implies that the moment generating function of is finite in an open neighbourhood of the origin.
, -almost everywhere
(a) This part of the proposition can be deduced from Theorems 2.8 and 2.10 of Dharmadhikari and Joag-Dev (1988). Their proof relies on a non-trivial correspondence between log-concave probability measures and log-concave densities, which in turn depends on several other facts about log-concavity – cf. Dharmadhikari and Joag-Dev (1988, pp.46–54). We give an alternative proof because it is perhaps a little more direct, and because it forms part of the proof of part (b) below.
Observe by the concavity of that if then . It follows that . This means that we can apply Lebesgue’s differentiation theorem to choose small enough that for every ,
We conclude that , a contradiction. Hence .
Thus, without loss of generality, we may assume . But by Fatou’s lemma,
so in fact we may assume . Since is log-concave, this proves (a).
We may also assume that for all . We conclude that for all ,
so , a contradiction.
Fix , and let . By Lemma 1, we can find such that \frac{1}{\|x\|}\{\phi(x)-\phi(0)\}\leq-\bigl{(}a+\frac{3\delta}{4}\bigr{)} for all . We claim there exists such that
Now suppose that is continuous and let . Choose large enough that for all . Then certainly,
Theoretical properties of the log-concave maximum likelihood estimator
There exists a constant such that, with probability one,
(a) Let , where the normalisation constant is chosen to ensure is a density, so that
The result follows immediately from this claim. To establish the claim, write , and observe that
The first term on the right-hand side of (3.1) is zero, by the strong law of large numbers.
by Hoeffding’s inequality. The first Borel–Cantelli lemma then completes the proof of (a).
again by Hoeffding’s inequality. By the first Borel–Cantelli lemma, and arguing as in the proof of Lemma 3(a) above, we conclude that
Our next theorem is the main result in this paper and establishes desirable performance properties of the log-concave maximum likelihood estimator. We first recall that the Kullback–Leibler divergence of a density from is given by
Jensen’s inequality shows that the Kullback–Leibler divergence is non-negative, and equal to zero if and only if (almost everywhere). Thus in the case where is log-concave, Theorem 4 below shows that the log-concave maximum likelihood estimator is strongly consistent in certain exponentially weighted total variation metrics. Convergence in exponentially weighted supremum norms also follows if is continuous. The theorem strengthens known results even in the univariate case, which include Corollary 1 of Pal, Woodroofe and Meyer (2007), where it was proved that is strongly consistent in Hellinger distance, and Corollary 4.2 of Dümbgen and Rufibach (2009), where (weak) consistency of in the unweighted total variation distance was established. (The observation that the mode of convergence in the univariate consistency result of Corollary 4.2 of Dümbgen and Rufibach (2009) could be strengthened was also made independently at around the same time in Schumacher, Hüsler and Dümbgen (2009).)
In the case where the model is misspecified, i.e. is not log-concave, we prove that the existence and uniqueness of a log-concave density that minimises the Kullback–Leibler divergence from . Moreover, we show that the log-concave maximum likelihood estimator converges in the same senses as in the previous paragraph to . The natural practical interpretation is that provided is not too far from being log-concave, the estimator is still sensible.
We write and recall that denotes the support of .
By the two integrability conditions, the log-concave density , where is a normalisation constant, satisfies . We can therefore pick a minimising sequence of log-concave densities for the Kullback–Leibler divergence from ; in other words, the sequence satisfies
Thus is absolutely continuous with respect to , and we may write for its density with respect to . By Proposition 2(a), is log-concave, and by Proposition 2(b), almost everywhere. Finally, by Fatou’s lemma, we have
Thus does indeed minimise the Kullback–Leibler divergence from over the class of log-concave densities.
Suppose now that both and minimise the Kullback–Leibler divergence from over the class of log-concave densities. Defining
we see that is a log-concave density with
by the Cauchy–Schwarz inequality, with equality if and only if , -almost everywhere. This proves the claimed uniqueness property of .
Combining this result with an application of the strong law of large numbers to the fourth term on the right-hand side of (3), we deduce that with probability one,
It follows by the monotone convergence theorem that with probability one,
Appendix
A log-concave density is bounded above and the version of that is closed attains its maximum.
for all . Since is bounded (by Lemma 5), the result follows by choosing sufficiently large.
The following lemma is used in the proof of Theorem 4. The conclusion can be immediately strengthened using Proposition 2, and is stated in this way only for brevity.
as . Here, the penultimate inequality uses (4.1). We deduce that the sequence in the statement of the lemma satisfies the condition that there exists such that
Then writing , we have for all that
We deduce that there exists such that
As in the proof of Theorem 4, we have from (4.2) that the sequence is tight. Thus if is an arbitrary subsequence of , then there exists a further subsequence and a log-concave density such that . But then, by the dominated and monotone convergence theorems,
with equality if and only if almost everywhere. By the hypothesis of the lemma, we must have . Since every subsequence of has a further subsequence converging in total variation norm to , we must have . ∎
Acknowledgements: The second author is very grateful for helpful conversations with Richard Nickl and Michael Tehranchi.