Density Level Sets: Asymptotics, Inference, and Visualization
Yen-Chi Chen, Christopher R. Genovese, Larry Wasserman
Introduction
Estimating the level sets of a probability density function has a wide range of applications, including anomaly detection (outlier detection) (Breunig et al., 2000; Hodge and Austin, 2004), two-sample comparison (Duong et al., 2009), binary classification (Mammen and Tsybakov, 1999), and clustering (Rinaldo and Wasserman, 2010; Rinaldo et al., 2012). In this paper, we study the problem of estimating the level set
where is the expected kernel density estimator with bandwidth , a smoothed version of the underlying density . Using (and thus ) has several advantages, which we discuss in detail in Section 2.2. Figures 5 and 10 illustrate the kind of confidence sets and visualizations we will develop in this paper.
A commonly used estimator of the density level-set is the plug-in estimator , where is the kernel density estimator or some other density estimator. There is a large literature for level sets (and upper level sets, which replace with ) that focuses on the consistency, rates of convergence (Polonik, 1995; Tsybakov, 1997; Walther, 1997; Cadre, 2006; Cuevas et al., 2006) and minimaxity (Singh et al., 2009) of such estimators under various error loss functions.
Recent results on statistical inference for level sets include Jankowski and Stanberry (2012) and Mammen and Polonik (2013). Statistical inference is challenging in this setting because the estimand is a set and the estimator is a random set (Molchanov, 2005). Mason and Polonik (2009) establish asymptotic normality for upper level sets when the loss function is the measure of the set difference. However, it is unclear how to derive a confidence set from this result.
Another challenge of level set estimation is that we cannot directly visualize the level sets when the dimension of the data is larger than 3. One approach is to construct a level-set tree, which shows how the connected components for the upper level sets bifurcate when we gradually increase (Stuetzle, 2003; Klemelä, 2004, 2006, 2009; Stuetzle and Nugent, 2010; Kent et al., 2013). The level-set tree reveals topological information about the level sets but loses geometric information.
In this paper, we propose solutions to all of these problems. Our main contributions can be summarized as follows.
We derive the limiting distribution of (Theorem 3).
We develop two bootstrap-based methods to construct confidence regions for (Section 4).
We prove that both bootstrap methods are valid (Theorem 28 and 5).
We devise a visualization technique that preserves the geometric information for density level sets (Section 5).
Related Work. Early work on density level set focuses on proving the consistency or the rate of convergence under various metrics. See e.g. Polonik (1995); Tsybakov (1997); Walther (1997); Cuevas et al. (2006); Rinaldo and Wasserman (2010). However, none of these derives a limiting distribution for the density level sets. To our knowledge, the only paper that considers limiting distributions is Mason and Polonik (2009), proving asymptotic normality under a generalized integrated distance. However, this metric cannot be used to construct a confidence set for density level sets since the asymptotic distribution involves the true density, which is unknown. Estimating the level set is also related to support estimation, see e.g. Cuevas and Rodríguez-Casal (2004) and Cuevas (2009).
Jankowski and Stanberry (2012) and Mammen and Polonik (2013), both provide methods for constructing confidence sets for the density level sets using the variation of the density function. Our approach is similar to theirs but is based on Hausdorff distance. We will compare their methods to ours in Section 4.2.
Outline. We begin with a short introduction to density level sets along with some useful geometric concepts in Section 2. In Section 3, we derive the limiting distribution of the Hausdorff distance between the estimated and true level sets. In Section 4, we construct a valid confidence set for density level sets. In Section 5, we devise a visualization method for density level sets that is simple to interpret and efficient to compute. (We provide an R package that implements our visualization method.) We summarize our results and discuss related problems in Section 6.
Technical Background
Let be a random sample from an unknown, continuous density . We define the density level set by
for some . Note that in the literature the term level sets is sometimes used for the set ; we call the latter the upper level set to distinguish it from the level set in equation 2. Thus, the level sets (in our terminology) are the boundaries to the upper level sets, under mild smoothness assumptions (e.g., assumption G below).
We assume that is a fixed, positive value. A plug-in estimate for is
where is the kernel density estimator (KDE),
2 Smoothed Density Level Set
In this paper, we focus on inference for the level sets of a smoothed version of , specifically:
where and denotes convolution. We denote the level set of by
Note that although we focus on estimating , we allow as .
Here, we argue that (and hence ) is a better target for level-set estimation than (and hence ). For simplicity, in this section we focus on the upper level sets and , but the gist of the argument remains the same in either case.
Our arguments are: (i) always exists while may not even exist; (ii) can be naturally viewed as an estimator for ; (iii) has all the salient structure of but is easier to estimate than ; (iv) While the bias may be analyzed theoretically, in practice, it cannot be accurately estimated. It is better to just be clear that we are really estimating .
As long as , , and hence we can uniformly consistently estimate , whether we keep fixed or let it tend to 0. The same is not true for . In fact, may not even have a density . The bias cannot be uniformly estimated in a distribution-free way.
Regarding (iii): The left plot in Figure 1 shows a density . The blue points at the bottom show the upper level set . The right plot shows for and the blue points at the bottom show the upper level set . The smoothed out density is biased and the upper level set loses the small details of . But these small details are the least estimable, and captures the principal structure of . In addition, when is smooth (having positive -reach; Chazal and Lieutier 2005; Chazal et al. 2009) and is sufficiently small, and will be topologically similar, in the sense that a small expansion of is homotopic to (Chazal and Lieutier 2005; Chazal et al. 2009; Genovese et al. 2014). The -reach and the concept of being nearly homotopic can be found in section 3.2 of Genovese et al. (2014) and section 4 of Chazal et al. (2009).
As a second example, let where is a Normal density and is a point mass at 0. Of course, this distribution does not even have a density. The left plot in Figure (2) shows the density of the absolutely continuous part of with a vertical line to who the point mass. The right plot shows , which is a smooth, well-defined density. Again the blue points show the level sets. As before and are slightly biased, but they are also well-defined. And and are accurate estimators of and , respectively. Moreover, captures the most important qualitative information about , namely, that there are three connected components, one of which is small. These examples show that — and hence — is a sensible target of inference.
Lastly, we would like to point out that the idea of viewing as the estimand is not new. The “scale space” approach to smoothing explicitly argues that we should view as an estimate of , and is then regarded as a view of at a particular resolution. This idea is discussed in detail in Chaudhuri et al. (2000); Chaudhuri and Marron (1999); Godtliebsen et al. (2002).
If one really wants to focus on making inference for , the level set of the original density, then we need to undersmoothUndersmoothing the density estimate to make statistical inferences is a common practice in nonparametric statistics; see e.g. page 89 of Wasserman (2006). so that the bias will not affect the limiting distribution. This leads to estimates of that are highly variable. We believe that an accurate confidence set for is more useful than a poor confidence set for .
These arguments explain why we regard rather than as the estimand. But these arguments do not tell us how to choose . Bandwidth selection is always a challenge and in this paper we mainly use Silverman’s rule of thumb (Silverman (1986)).
3 Geometric Concepts
Let be the projection of a point onto a set . Note that may not be unique. The distance induced by the projection is
where and . The Hausdorff distance is a generalized version of the metric for sets.
The reach of a set (Federer 1959; Cuevas 2009, also known as condition number Niyogi et al. 2008 or minimal feature size Chazal and Lieutier 2005) is the largest distance from such that every point within this distance to has a unique projection onto . i.e.
Note that is unique if . Another way to understand the reach is as the largest radius of a ball that can freely move along ; see Figure 3 for an example. In some cases, the reach is the same as the smallest radius of curvature on . The reach plays a key role in relating the Hausdorff distance to the empirical process. Note that the reach is closely related to ‘rolling properties’ and ‘-convexity’; see Cuevas (2009), Cuevas et al. (2012) and appendix A of Pateiro-López (2008).
Finally, two smooth sets and are called normal compatible (Chazal et al., 2007) if the projection between and are one to one and onto; see Figure 4 for an example. When and are normal compatible, the Hausdorff distance has the simpler form
Asymptotic Theory
In this section, we derive the asymptotic theory for . Note first that for two sets and , the Hausdorff distance satisfies the inclusion property
From this, it follows that when we have a set estimator and a parameter of interest , then for any , the set
is a confidence set for , where is the -quantile of random variable . Thus, whenever we can approximate the distribution of , we can construct a confidence set for . Note that Chen et al. (2014b) and Chen et al. (2014a) have used this property to construct confidence sets, but neither paper used this property to full effect.
We define the sup norm using derivatives up to -th order by:
Let be an multi-index such that each is a non-negative integer and . We define
We now state are main assumptions, for an arbitrary density . When we use the assumptions in what follows, we will take to be .
The kernel function and is symmetric, non-negative, and
for all multi-indices satisfying .
The kernel function and its partial derivative satisfies condition in Giné and Guillou (2002). Specifically, let
Assumption (G) appears in Molchanov (1991); Tsybakov (1997); Walther (1997); Molchanov (1998); Cadre (2006); Mammen and Polonik (2013); Laloe and Servien (2013). For a smooth density , (G) holds whenever the specified level does not coincide with the density value for a critical point.
Assumption (K1) is to guarantee that the variance of the KDE is bounded and to ensure that . This assumption is very common in statistical literature, see e.g. Wasserman (2006). Assumption (K2) is to regularize the complexity of the kernel function so that the supremum norm for kernel functions and their derivatives can be bounded in probability. Similar assumption appears in Einmahl and Mason (2005) and Genovese et al. (2014). The Gaussian kernel and many compactly supported kernels satisfy both assumptions.
An immediate result from assumption (G) is the smoothness of the density level set. This smoothness property will be used to understand the distribution of .
Assume a density satisfies condition (G) and let denote the level set for at . Then
Moreover, let be another density function and define as the level set for at level . When is sufficiently small,
.
and are normal compatible.
The proof is given in the supplementary materials. Lemma 1 is very similar to Theorem 1 and 2 in Walther (1997). Essentially, this lemma shows that the level set is smooth and that whenever two smooth densities are sufficiently close, their level sets will both be smooth, be close to each other, and have one-to-one and onto normal projections between them.
Assume (K1–K2) and (G) hold for . Let and be the density level sets with level for and . Define the function
with . If , then
A key element for the proof of Lemma 19 is the smoothness of and , which relies on Lemma 1. This smoothness allows us to approximate the local difference by an empirical process.
Lemma 19 shows that the projected distance to the level sets can be approximated by a stochastic process (the empirical process) defined on a smooth manifold. Specifically, Lemma 19 shows that the projection distance can be approximated by an empirical process on certain functions , where . The level sets now acts as an index set. Thus, we define the function space
Assume (K1–K2) and (G) holds for . Let and be the density level sets with level for and . Then when , the Hausdorff distance satisfies
The proof of Theorem 3 depends on two geometric observations. First, the empirical approximation in Lemma 19, shows that the local difference is approximately the same as an empirical process, and hence the maximum local difference is approximated by the maximum of the empirical process. Second, the normal compatibility between and guaranteed by Lemma 1, which implies that maximal of local difference is the same as the Hausdorff distance.
Theorem 3 shows that the Hausdorff distance can be approximated by a maximum over a certain Gaussian process. Note that we cannot directly use this theorem to construct a confidence set for since the Gaussian process is defined on , which is unknown. Later we will use the bootstrap to approximate this limiting distribution and construct a confidence set.
Statistical Inference
We now show that we can construct valid confidence sets for with the bootstrap. A set is called an asymptotically valid confidence set for if
where as . We propose two methods for constructing a confidence set, and we will show that they are both asymptotically valid.
The first approach is to use the Hausdorff distance between the level sets. Let and define
where denotes the cdf for a random variable . Then, it is easy to see that
We use the bootstrap to estimate .
Let be a bootstrap sample from the empirical measure. Let denote the KDE using the bootstrap sample, and the corresponding level set. We define and
Then the bootstrap confidence set is .
Assume (K1–K2) and (G) holds for and . Let and and be the density level set with level for and and . Then there exist such that
An intuitive explanation for Theorem 28 is that as goes to infinity, the bootstrap process converges to the same Gaussian process as the empirical process – thus, they share the same Berry-Esseen bound.
2 Method 2: Supremum Loss
The second approach is to use the supremum norm of the KDE and impose an upper and lower bound around the density level.
Assume (K1–K2) and (G) holds for and . Let and and be the density level sets with level for and and . Then
The proof of this Theorem is similar to the proof of Theorem 28, so we omit the details. The basic idea is as follows. By equation (30), the quantile of can be used to construct confidence sets. We then show that is approximated by the maximum of a Gaussian process (similar to Theorem 3 and made explicit in Chernozhukov et al. 2014a). Finally, we show that the bootstrap converges to as in Theorem 28.
This method embodied in Theorem 5 is very similar to the methods in Jankowski and Stanberry (2012) and Mammen and Polonik (2013). Jankowski and Stanberry (2012) proposes to construct a confidence set of the form
where is the symmetric difference between sets. Then they use the upper quantile of to construct a confidence set of a similar form to (29) and apply the bootstrap to estimate the quantile. Their bootstrap consistency relies on Neumann’s method (Proposition 3.1 in Neumann (1998)) and they assume that converges fast enough so that one can ignore the bias for estimating the original density . Actually, under their assumptions, our proposed bootstrap confidence sets (from both methods 1 and 2) are also consistent for the original level set since the bias converges faster than the stochastic variation. The method in Mammen and Polonik (2013) should have higher power than our method 2 since they consider taking the supremum over a smaller region.
We may use a variance stabilizing transform to obtain an adaptive confidence set using similar idea to Chernozhukov et al. (2012). The variance of is proportional to . Thus, we may use
and set as an adaptive threshold for constructing the confidence set. Namely, the adaptive confidence set is given by
Using the same approach as in the proof to Theorem 5, we can show that has asymptotically coverage.
The rate may not be optimal. In Chernozukov et al. (2014), they apply a induction technique that gives a rate of order for the Gaussian approximation. Despite not being mentioned explicitly in that paper, we believe that similar technique applies to the empirical process. If this is true, the rate in Theorem 3, 28 and 5 can be further refined to .
3 Comparing Methods 1 and 2
Despite the fact that method 2 yields a much larger confidence sets, it has some nice properties. First, method 2 is very simple: all we need is to compute the bootstrap distribution of supremum loss. Second, method 2 is connected to the methods proposed in Mammen and Polonik (2013) and Jankowski and Stanberry (2012). Third, the confidence sets produced in method 2 are related to level sets with level . The last property makes it easy to visualize the confidence sets; see Section 5.3.
Here, we consider two simulated datasets to compare the coverage for confidence sets constructed using Hausdorff loss (method 1), loss (supremum loss; method 2), and scaled loss (remark 3).
The first dataset is a three-Gaussian mixture. The data is generated from the following distribution:
where is the density to bivariate Gaussian with mean vector and Covariance matrix . is the identity matrix and , , and . We use density level and smoothing parameter . The corresponding level set is the red curve in the left panel of Figure 7. We consider three different sample sizes, and compare the coverage of confidence sets using the three methods. The corresponding coverages are given in Table 1. As can be seen from Table 1, all the three methods have the desire nominal coverage.
The second dataset is a four-mixture dataset from Cadre et al. (2009). The data is generated from the following distribution:
with and , and , and
We use density level and smoothing parameter . The corresponding level set is the red curves in the right panel of Figure 7. Again, we consider three different sample sizes and compare the coverage of confidence sets using the three methods. The corresponding coverages are given in Table 2. It is clear that all the three methods have the desire nominal coverage. Moreover, it can be seen from both Table 1 and 2 that the supremum loss method over covers.
4 Pointwise Hypothesis Tests for Level Sets
The confidence sets developed in the previous section are related to two types of local hypothesis tests. For fixed density level and an arbitrary point , consider the tests of whether is greater or less than :
When we only want to test just a few points, we can do local tests for each point and control the family-wise error rate to control the type 1 error rate. Usually, however, we are interested conducting the local test at many or even an infinite number of points (like a region), making it difficult to control type 1 error simultaneously.
Inverting the confidence sets of the previous subsection gives a solution to this problem. Let be a confidence set for and let and be, respectively, the upper and lower lambda level sets. Then the decision rules
In addition, inverting the regions where we cannot reject and in (52) yields confidence sets for and . In Figure 6, a confidence regions for the upper level set is the union of yellow and blue regions (). And a confidence regions for , the lower level set, is the union of green and blue regions (). Thus, with confidence, all yellow regions are above and the true high density regions should be contained by the yellow and blue regions.
The two local tests described in (50) and (51) are relevant to the problems of level-set clustering (Hartigan, 1975; Polonik, 1995; Rinaldo and Wasserman, 2010; Rinaldo et al., 2012) and anomaly detection (Desforges et al., 1998; Breunig et al., 2000; He et al., 2003; Chandola et al., 2009). Rejecting can be viewed as evidence that belongs to a level-set cluster. And a point where is rejected can be viewed as anomalous.
Note that one can modify the local testing procedure to control the False Discovery Rate (Benjamini and Hochberg, 1995) rather than familywise error.
Visualization for Multivariate Level Sets
Level-set estimators can reveal useful information about a distribution, but beyond three dimensions, we cannot directly visualize the level sets, making the results difficult to use.
In this section, we propose a novel visualization technique density upper level sets in multidimensions. Any visualization entails some loss of information, but our goals are to preserve important geometric information about the sets, make the overall visualization easy to interpret, and give a method that is efficient to compute. Our method exploits the relationship between level-set clustering and mode clustering (see Figure 9).
A current and commonly used visualization method for level sets is the density tree (Stuetzle, 2003; Klemelä, 2004, 2006; Kent et al., 2013; Balakrishnan et al., 2013). This method considers several density levels, , and computes the number of connected components for the upper density level set at each level. As we increase the density level, some connected components may vanish and others may split into additional components. In the typical case when the underlying density function is a Morse function (i.e., Hessian at critical points is non-degenerate) (Morse, 1925, 1930; Milnor, 1963), the connected component disappears when the density level is above the maximum density value over the component and splits only if the density level passes the density value of some saddle points within the component (Klemelä, 2009). The vanishing and splitting of components as the level changes produces a tree structure. The density tree uses this tree structure as a visualization of the level sets. We refer to Klemelä (2009) for more details.
Density trees display primarily topological information about the underlying density function (Stuetzle, 2003; Kent et al., 2013) but need not preserve or impart other features that may be of interest. Extensions have been proposed that would endow the tree with additional information about the distribution (Klemelä, 2004, 2006, 2009), but this is difficult to do with multiple features and can make the visualization difficult to interpret. See Figure 8 for an example.
In this section, we propose a novel technique that visualizes several density (upper) level sets that preserve some geometric information and that are very easy to understand. Our method is based on the relationship between level set clustering and mode clustering (see Figure 9). Note that in this section, we will focus on density upper level sets.
Our method complements existing tree-based methods. It combine two clustering techniques – level-set and mode clustering – to produce a simple and intuitive visualization.
That is, starts at and moves along the gradient of . We define the destination for as . Let be the collection of all local modes of . It can be shown that except for a set of ’s in a set with Lebesgue measure (this set corresponds to the boundaries of clusters). For each mode , we define its basin of attraction as
The regions are the clusters generated by mode clustering.
Now we recall three facts about an upper density level set (c.f. Figure 9 left and middle panels):
Thus, the upper level sets are covered by the basins of attraction of the local modes (the uncovered regions have Lebesgue measure so we ignore them for visualization).
2 Visualization Algorithm
Given several density levels, , we can overlay the visualization from the previous paragraph from to to create a “tomographic” visualization of the clusters. This gives a visualization for the density level sets. Figure 10 shows an example for visualizing level sets for a -dimensional and a -dimensional simulation datasets at different density levels. This dataset is from Chen et al. (2014c).
3 Visualizing Level Sets with Confidence
A modified version of our algorithm allows us to visualize multivariate confidence sets for an upper level set at a given level . In particular, we visualize the confidence set for the level sets produced by Method 2 (supremum loss, see Section 4.2) . The advantage of Method 2 here is that this confidence set under different ’s is just the density level set at different ’s. This makes it easy to visualize distinct confidence sets for pre-specified levels .
and use the visualization algorithm to create a tomographic visualization. Figure 11 provides an example for visualizing confidence sets for upper level set using the 6 dimensional simulation data in Figure 10.
At the cost of additional computation, we can also use Method 1 (Hausdorff loss) instead of Method 2, combining it with the basins of attractions to visualize the confidence sets. We use method 1 to construct the inner/outer confidence sets and then we find the connected components and partition it by the basins of attraction for local modes and use multidimensional scaling to visualize it in low dimensions.
Discussion
In this paper, we derived the limiting distribution for smoothed density level sets under Hausdorff loss. This result immediately allows us to construct confidence sets for the smoothed level set. We developed two bootstrapping methods to construct the confidence sets, and we showed that both methods are consistent. These confidence sets can be inverted to construct multiple local tests of whether a point’s density value is above or below a given level, which has application to related problems such as anomaly detection. Finally, we developed a new visualization method that is informative and interpretable, even in multidimensions.
Although we focused on density level sets in this paper, our methods, including confidence sets and visualization, can be applied to a more general class of problems such as kernel classifier, two-sample tests, and generalized level sets (See Mammen and Polonik 2013 for more details.).
References
Proofs
Assume (K1–2), then for each there exists some such that whenever , we have
Proof for Lemma 1. We first prove the lower bound for and then we will prove the additional assertions.
Part 1: Lower bound on reach. We prove this by contradiction. Take near such that
We assume that has two projections onto , denoted as and .
Since , so that . Now by Taylor’s theorem
Since both and are projection points from onto ,
Recall that and by Taylor’s theorem,
so that . Note that the lower bound in the last inequality is because so it follows from assumption (G). Plugging in this result into the last equality of (60), we conclude that . This shows so that we have a unique projection. Thus, whenever , we have a unique projection onto and thus we have proved the lower bound on reach.
Part 2: The three assertions. The first assertion is trivially true when is sufficiently small since assumption (G) only involves gradients (first derivatives).
The second assertion follows from the lower bound on reach. By assertion 1, (G) holds for . And the lower bound on reach is bounded by gradient and second derivatives so that we have the prescribed bound.
The third assertion follows from Theorem 1 in Chazal et al. (2007) which states that if two dimensional smooth manifolds and have Hausdorff distance being less than , then and are normal compatible to each other. Now by Theorem 8, the Hausdorff distance between and is at rate so that this assertion is true when is sufficiently small.
Proof of Lemma 2. Let . We define to be the projected point onto . By Lemma 1 and Theorem 8, when , so that is unique. Thus, we assume is unique.
Now since and , . Thus, by Taylor’s theorem
Note that is normal to at so that it points toward the same direction as . Thus, (62) can be rewritten as
By Taylor’s theorem, is close to in the sense that
In addition, is bounded by which is at rate due to Theorem 8. Putting this together with (63), we conclude
Note that the left hand side can be written as
This holds uniformly for all and note that the definition of is
Proof for Theorem 3. The proof for Theorem 3 follows the same procedure as the proof of Theorem 6 in Chen et al. (2014b). The proof contains two parts: Gaussian approximation and anti-concentration.
Part 1: Gaussian approximation. Basically, we will show that
First, when is sufficiently small, and are normal compatible to each other by Lemma 1. Then by the property of normal compatible,
Note that this result basically follows from the same derivation of Proposition 3.1 in Chernozhukov et al. (2014c) with the fact that in their definition.
Combining equations (70) and (71) and pick , we have that for is sufficiently large and ,
Part 2: Anti-concentration. To obtain the desired Berry-Esseen bound, we apply the anti-concentration inequality in Chernozhukov et al. (2014c) and Chernozhukov et al. (2014a).
From Lemma 10 and equation (72), there exists some constant such that
Now pick and use the fact that and converges faster than the other terms; we obtain the desired rate.
Proof for Theorem 4. This proof follows the same strategy for the proof of Theorem 7 in Chen et al. (2014b). We prove the Berry-Esseen type bound first and then show that the coverage is consistent. We prove the Berry-Esseen bound in two simple steps: Gaussian approximation and support approximation.
Thus, if we sample from and consider estimating by , we are doing exactly the same procedure of estimating by . Therefore, Lemma 2 and Theorem 3 hold for approximating by a maxima for a Gaussian process. The difference is that the Gaussian process is defined on
since the “parameter (level sets)” being estimated is (the estimator is ). Note that is very similar to except the denominator is slightly different and the support is also different from . That is, we have
Step 2: Support approximation. In this step, we will show that
The first approximation can be shown by using the Gaussian comparison lemma (Theorem 2 in Chernozhukov et al. (2014b); also see Lemma 17 in Chen et al. (2014b)). We do the same thing as Step 3 in the proof of Theorem 8 in Chen et al. (2014b) so we omit the details. Essentially, given any , we can construct a pair of balanced -nets for both and , denoted as and so that . Then this -net leads to
Now comparing the above result to Theorem 3 and using the fact that the first big-O term dominates the second term (the first is of rate for but the second term is at rate by Theorem 9), we conclude the result for first assertion.
For the coverage, let and . Since , we have
Now by the first assertion, the difference for and the bootstrap estimate differs at rate , which completes the proof.