Uniformly valid confidence intervals post-model-selection
François Bachoc, David Preinerstorfer, Lukas Steinberger
Introduction
Fitting a statistical model to data is often preceded by a model selection step, and practically always has to face the possibility that the candidate set of models from which a model is selected does not contain the true distribution. The construction of valid statistical procedures in such situations is quite challenging, even if the candidate set of models does contain the true distribution (cf. Leeb and Pötscher (2005, 2006, 2008), Kabaila and Leeb (2006) and Pötscher (2009), and the references given in that literature), and has recently attained a considerable amount of attention. In a Gaussian homoskedastic location model and fitting possibly misspecified linear candidate models to data, Berk et al. (2013) have shown how one can obtain valid confidence intervals post-model-selection for (non-standard) model-dependent targets of inference in finite samples (cf. also the discussion in Leeb, Pötscher and Ewald (2015), and related results obtained for prediction post-model-selection in Bachoc, Leeb and Pötscher (2014)). In this setup, their approach leads to valid confidence intervals post-model-selection regardless of the specific model selection procedure applied. This aspect is of fundamental importance, because many model selection procedures used in practice are almost impossible to formalize: researchers typically use combinations of visual inspection and numerical algorithms, and sometimes they simply select models that let them reject many hypotheses, i.e., they are hunting for significance. These often unreported and informal practices of model selection prior to conducting the actual analysis may also play a key role in the current crisis of reproducibility. Thus, to establish and popularize statistical methods that are in some sense robust to ‘bad practice’ is highly desirable.
The methods discussed in Berk et al. (2013) are based on the assumption that the true distribution is Gaussian and homoskedastic, and the authors consider only situations where linear models are fit to data. It is of substantial interest to generalize this approach, and to obtain generic methods for constructing confidence intervals post-model-selection that are widely applicable beyond the Gaussian homoskedastic model considered in Berk et al. (2013). We develop a general asymptotic theory for the construction of uniformly valid confidence sets post-model-selection. These results are applicable whenever the estimation error can be expanded as the sum of independent centered random vectors and a remainder term that is negligible relative to the variance of the leading term. Such a representation typically follows from standard first order linearization arguments, and can therefore be obtained in many situations.
Our confidence intervals can be based on either consistent estimators of the variance of the previously mentioned sum, or, more importantly, if such estimators are not available (which is usually the case when all working models are misspecified), can be based on variance estimators that consistently overestimate their targets. We also present results that allow one to obtain such estimators in general and demonstrate their construction in specific applications, where they often coincide with well known sandwich-type estimators. This overcomes another limitation present in Berk et al. (2013), namely the assumption that there exists an unbiased (and chi-square distributed) or uniformly consistent estimator of the variance of the observations (cf. the discussion in Remark 2.1 of Leeb, Pötscher and Ewald (2015) and in Appendix A of Bachoc, Leeb and Pötscher (2014)). The usage of variance estimators that overestimate their targets, while leading to more conservative inference, renders the approach applicable to the fully misspecified setting. Moreover, the suggested conservative estimators usually have the property that their bias vanishes if the selected model is correct (cf. Remark 2.8 and Subsection 3.1.2).
Another important aspect of the results obtained is that they are valid uniformly over wide classes of potential underlying distributions, which is particularly important as this guarantees that the results provide a better description of finite sample properties than ‘pointwise’ asymptotic results (cf. Leeb and Pötscher (2003), Leeb and Pötscher (2005) and Tibshirani et al. (2015) for a discussion of related issues in a model selection context).
Moreover, we apply our general theory to three important modeling situations: First, we consider the case where linear homoskedastic models are fitted to non-Gaussian homoskedastic data. This provides an extension of the results of Berk et al. (2013) to the non-Gaussian case, without requiring a consistent variance estimator. Next, we study the problem of fitting heteroskedastic linear models to non-Gaussian heteroskedastic data. This scenario necessitates a more careful choice of variance estimators and leads to an extension of the influential results of Eicker (1967) to the misspecified post-model-selection context. Our third application then considers the problem of fitting binary regression models to binary data. In this case, also the link function may be chosen in a data driven way. On a technical level, the third example is quite different from the previous ones, because here non-trivial existence and uniqueness questions concerning the targets of inference and the (quasi-)maximum likelihood estimators have to be addressed.
Our confidence intervals obtained in these specific situations are particularly convenient for practitioners, because they are structurally very similar to the confidence sets one would use in practice following the naive (and invalid (see, e.g., Leeb, Pötscher and Ewald, 2015; Bachoc, Leeb and Pötscher, 2014)) approach that ignores that the model has been selected using the same data set. The main difference of our construction to the naive (and invalid) approach is the choice of a critical value: Quantiles from a standard normal or -distribution are replaced by so-called POSI-constants (cf. Berk et al. (2013) and Section 2.5 below). Thus, the procedures are conceptually simple and easy to implement. Moreover, we provide mild and easily verifiable regularity conditions on observable quantities (e.g., the design or the link functions) under minimal restrictions on the unknown data generating process.
Finally, in a series of numerical examples, we illustrate that the proposed confidence intervals are valid also in small samples while their lengths appear to be practically reasonable when compared to naive (and invalid) procedures. Furthermore, we compare our methods to those of Tibshirani et al. (2015) and Taylor and Tibshirani (2017), and find that our intervals are often shorter than their competitors, even when we study the exact same scenarios for which those competing methods were tailored for and even though our confidence intervals offer much stronger theoretical guarantees.
The structure of the present article is as follows: We first develop a general asymptotic theory for the construction of uniformly valid confidence sets post-model-selection in Section 2. In Section 3, we apply our theoretical results to the three previously mentioned modeling scenarios. Of course, the selection of examples in Section 3 is by no means exhaustive. But besides covering three very important modeling frameworks, Section 3 serves as an illustration of how the general theory developed in Section 2 can be applied. An outline of the numerical results is presented in Section 4. In Section 5 we conclude and discuss possible extensions of the results obtained in this paper that are currently under investigation. Details of the simulations as well as all the proofs are collected in Sections A, B, C and D of the appendix.
The present article is devised in the spirit of Berk et al. (2013), in the sense that we aim at inference post-model-selection that is valid irrespective of the employed model selection procedure. Very recently, Rinaldo et al. (2016) have investigated a classical sample spitting procedure that is also independent of the underlying selection method. However, they consider only the i.i.d. case, thereby excluding, for instance, fixed design regression. Several other authors have proposed inference procedures post-model-selection that are tailored towards specific selection methods and for specific modeling situations. In the context of fitting linear regression models to Gaussian data, methods that provide valid confidence sets post-model-selection, and that are constructed for specific model selection procedures (e.g., forward stepwise, least-angle-regression or the lasso) and for targets of inference similar to those considered in the present article, have been recently obtained by Tibshirani et al. (2016), Lee and Taylor (2014), Fithian, Sun and Taylor (2015) and Lee et al. (2016). Tibshirani et al. (2015) extended the approach of Tibshirani et al. (2016) to non-Gaussian data by obtaining uniform asymptotic results. Furthermore, valid inference post-model-selection on conventional regression parameters under sparsity conditions was considered, among others, by Belloni, Chernozhukov and Hansen (2011, 2014); van de Geer et al. (2014) and Zhang and Zhang (2014).
Inference post-model-selection: A general asymptotic theory
From the coverage property in Part 1 we obtain
As already discussed in the introduction, the fact that our approach does not restrict the model selection procedure used is important. It is precisely this aspect that allows practitioners to obtain valid confidence intervals post-model-selection in situations where a wide variety of (formal or informal) mechanisms have been incorporated to select the model.
2 Discussion
i.e., asymptotic validity of the constructed confidence sets uniformly over . That the development of results that hold uniformly over large classes of distributions is important, in particular so in the context of inference post-model-selection, is well understood (see Leeb and Pötscher (2003) and Leeb and Pötscher (2005)). One recent article that studies uniform coverage properties post-model-selection is Tibshirani et al. (2015). Merits of uniform results in contrast to pointwise asymptotic results are discussed in their Section 1.1. Tibshirani et al. (2015) consider a setup similar to the example we consider in Section 3.1 and for specific model selectors, but compared to our results uniform validity is established only over substantially smaller sets of distributions, and they need to impose stronger conditions on the design matrices, which rule out some important cases our results allow for, e.g., polynomial trends. See also Section 4.1 for numerical results and comparisons.
3 Notation
4 Main assumption
which does not depend on . The condition is as follows:
where, writing , it holds for every and every that
Furthermore, for every coordinate we have
Clearly, an expansion as in Equation (2.1) of Condition 1 is satisfied in many applications, and can typically be obtained by a standard linearization argument (see Subsection 3.3 for an example and further discussion). We emphasize that the two last assumptions in Condition 1 are formulated in terms of rescaled summands, which, in applications, can be exploited to circumvent restrictive compactness assumptions on moments of the distribution generating the data or the design (e.g., in Subsections 3.1 and 3.2 we do not need to restrict variance parameters to a compact set - as opposed to the conditions used by, e.g., Eicker (1967) or Tibshirani et al. (2015); and in Subsection 3.3, we do not require the smallest singular value of the design matrix to diverge to infinity - as opposed to, e.g., Lv and Liu (2014)).
Before proceeding to the main results, we briefly highlight the most important consequence of Condition 1 for our method of constructing confidence sets post-model-selection. The first step of our approach outlined in Subsection 2.1 required the construction of confidence intervals for each coordinate of the stacked vector of targets . Naturally, such confidence intervals will be centered at the respective coordinates of . Now, as a first step towards the construction of such confidence intervals, Condition 1 can be used to provide a useful asymptotic approximation to . More specifically, the first part of the subsequent Lemma 2.2 provides an asymptotic approximation to the distribution
and, for every , it holds that
Using, e.g., Gnedenko and Kolmogorov (1954) Theorem 3 in Paragraph 21, one obtains that Equation (2.3) in Condition 1 can be equivalently phrased as
In some applications it might be easier to check these two conditions directly (for every ), in particular in case one can use existing results in the literature on misspecified models without model selection as indicated above.
5 Confidence intervals post-model-selection
where sums over an empty index set are to be interpreted as .
converges to , or equivalently, that for every
Then the convergence in (2.7) is satisfied for
or equivalently, that for every it holds that
Asymptotic approximations of for large and are provided in Bachoc, Leeb and Pötscher (2014), Berk et al. (2013) and Zhang (2017). In particular, as ,
from Proposition 2.10 in Bachoc, Leeb and Pötscher (2014), itself building on results from Berk et al. (2013) and Zhang (2017).
An often useful upper bound on with a -dimensional correlation matrix is provided in the following lemma:
For every and a correlation matrix we have
In a particular application it might of course be possible to obtain better upper bounds by exploiting structural properties of the specific correlation matrix at hand, cf. Subsection 3.1. Using the upper bound of Lemma 2.9 can also be very useful in situations where the computation of is infeasible.
Applications
In this section we now apply the general results obtained in Section 2 to some important special cases that are frequently encountered in practice. As already mentioned in Section 2.2, we now consider situations of the following type:
We assume that satisfies the following condition, where we denote the -th row of by :
Eventually , and for every ,
Condition X1 particularly holds if , eventually, and . Moreover, it also holds in case is bounded and is bounded away from , which is typically the case in sufficiently balanced factorial designs, but Condition X1 is obviously much more general. For example, it also covers the important cases of polynomial regressors, trigonometric regressors, or mixed polynomial and trigonometric regressors (cf. the discussion in Eicker (1967), pp. 64). Finally, we point out that the condition
where the block-matrix is defined via its -th block of dimension given by
It is also worth noting that up to the choice of the last multiplicative factor in the definition of the confidence intervals above, i.e., the POSI constant, this is just the usual confidence interval for the -th coordinate of the coefficient vector one would typically use in practice working with homoskedastic linear models, and by following the naive way of ignoring the data-driven model selection step. The crucial difference, however, is that the naive approach is invalid (see, e.g., Leeb, Pötscher and Ewald, 2015; Bachoc, Leeb and Pötscher, 2014).
with equality holding only if , i.e., no model selection. That is, the confidence intervals for individual parameters constructed here are smaller than the ones guaranteeing simultaneous coverage. The following can now be said about their asymptotic coverage properties.
We finally mention that similar arguments can be used to construct confidence intervals for model-dependent linear combinations of regression coefficients (i.e., contrasts). Furthermore, analogous constructions can be used to obtain confidence intervals for individual coefficients (or, more generally, model-dependent contrasts) in the examples discussed in Sections 3.2 and 3.3 below. Due to space constraints we do not provide details.
1.2 The POSI-intervals automatically adapt to misspecification
for every and for every .
2 Inference post-model-selection when fitting fixed design linear models to heteroskedastic data
Note, similarly as in Subsection 3.1 above, that up to the choice of the last multiplicative factor , an upper bound for the corresponding POSI-constant, this is just the usual confidence interval for the -th coordinate of the coefficient vector one would typically use in practice working with heteroskedastic linear models by following the naive way of ignoring the data-driven model selection step. Our construction delivers an adjustment to that approach, which turns it, regardless of the (measurable) model selection procedure applied, into an asymptotically valid statistical procedure. The main result of this subsection is as follows:
3 Inference post-model-selection when fitting binary regression models to binary data
We need to impose some regularity conditions on the possible response functions and the design .
, for every ;
;
The elements have the following properties:
Condition X2 is a strengthened version of Condition X1. It is still satisfied if is bounded and is bounded away from , as is typically the case in factorial designs. Note, however, that Condition X2 is invariant under scaling of , so that, in particular, it does not require that , a condition commonly used to prove consistency of the MLE. Furthermore, Condition X2(ii) is implied by
The Conditions H(i) and H(iv) are rather natural and essential for parameter identification. Condition H(iii) is also classical and used to ensure continuity of the Hessian of the log-likelihood (cf. Fahrmeir and Kaufmann, 1985; Fahrmeir, 1990). Finally, Condition H(ii), which is implied by Condition H(iii), ensures strict concavity of the log-likelihood, which, in turn, guarantees uniqueness of pseudo parameters and the MLE (see Lemma 3.9 and Lemma 3.10 below). It is easy to see that Condition H is satisfied, e.g., for response functions corresponding to the classical logit, probit, log-log and complementary log-log link functions discussed in McCullagh and Nelder (1989, p.108).
It is important to note that if one decides a priori to use only the canonical link function, which, in the present case of binary regression, corresponds to the logistic response function , then Theorem 3.11 holds with the POSI-constant decreased to . See Corollary D.3 in Section D.10 of the supplement.
We point out that similar principles used to derive Theorem 3.11 can also be employed to treat other quasi-maximum likelihood or general M-, and Z-estimation problems (see, e.g., Fahrmeir, 1990, for a more general treatment of generalized linear models). The general theory of M- and Z-estimation as presented, e.g., in van der Vaart and Wellner (1996, Sections 3.2 and 3.3), usually also leads to expansions of the form required in Conditon 1 (cf. van der Vaart and Wellner, 1996, Theorem 3.2.16 and Theorem 3.3.1). These results are stated in a pointwise fashion but can be made uniform over large classes of data generating processes by using ideas from Section 2.8 of the same reference. However, in more specific examples, such as the present binary regression setting, conditions can be directly imposed on the design and the link functions and can be optimized for this setup.
Simulation study
In this section, we present the main findings of an extensive simulation study, the details of which can be found in Section A of the appendix.
The “POSI” confidence intervals always have target-specific and simultaneous coverage above the nominal level. The coverage proportions are large, which is so because these confidence intervals offer strong guarantees: they are valid for any model selection procedure, and simultaneously over all the variables in the selected model. Turning to the “TG” confidence intervals, we observe that these intervals have coverage probabilities approximately equal to the nominal level when the three targets are considered separately but their median lengths are often larger and never much smaller than the lengths of the “POSI” intervals. Finally, the quantiles are always larger for the “TG” intervals, for which they can be very large. In Table 3, we sometimes report infinite quantiles for the “TG” intervals. This is because, although the confidence intervals in Tibshirani et al. (2015) always have finite length in theory, the numerical implementation in the R package selectiveInference can return lower or upper bounds equal to . In contrast, the confidence intervals suggested in this paper are more robust, in the sense that their quantile lengths are always less than twice as large as their median lengths.
We believe that the numerical results of Table 3 favor the “POSI” confidence intervals suggested in this paper over the “TG” procedure. Indeed, we have seen that, even though the LAR model selector is used, the “POSI” confidence intervals have larger coverage proportions, remain valid when considered simultaneously, generally have smaller median lengths, and never exhibit very large quantile lengths. On top of this, the “POSI” confidence intervals are much more broadly applicable, as they have theoretical guarantees for any model selection procedure.
One needs to mention here that Tibshirani et al. (2015) also discuss a bootstrap version of their “TG” intervals. These bootstrap confidence intervals have similar coverage properties as the “TG” intervals, but much smaller median width. Their width seems to be comparable to the width of our “POSI” confidence intervals (cf. Tables 1 and 2 in Tibshirani et al. (2015)). All the advantages of the “POSI” method discussed in the preceding paragraph (besides the comments concerning their smaller width) also apply to the bootstrapped “TG” intervals. Furthermore, this suggests that our “POSI” intervals could potentially be improved by using suitable bootstrap methods as well. However, answering this question goes beyond the scope of the present article.
2 The case of ‘significance hunting’
among the models with largest penalized log-likelihood. We set , , and consider two settings for . In the “zero” setting, we set . In the “non-zero” setting, we set . We consider the values and . The errors are normally generated. In Table 4, the coverage proportions are significantly lower than in Table 3, and closer to the nominal level. Hence, the confidence intervals suggested in this paper may have conservative coverage proportions for some model selection procedures (such as LAR) but this is somehow necessary, since there exist other model selection procedures (such as “significance hunting”) for which the coverage proportions are close to the nominal level.
3 Further results
In Section A of the appendix we provide all the details of the previous simulations as well as further discussions of the results. We also present simulations for the binary regression problem of Section 3.3, comparing our methods to a procedure suggested by Taylor and Tibshirani (2017) and to naive intervals that ignore the data driven model selection step. Furthermore, we investigate the effect of misspecification and we also consider the “significance hunting” procedure in the binary regression case. The overall picture is similar to the results for the linear model, with the additional aspect that the “POSI” intervals remain valid also under misspecification, whereas the coverage probabilities of the methods of, e.g., Taylor and Tibshirani (2017) can be substantially below the nominal level in that case.
Conclusion
We have presented a general theory for the construction of asymptotically valid confidence sets post-model-selection. Our methods can be used in a wide number of situations, because they are only based on a standard representation that can often be obtained by simple linearization arguments. We have also applied our theory to construct valid confidence sets after selecting and fitting fixed design linear models to (possibly non-Gaussian) homoskedastic or heteroskedastic data. Moreover, we have investigated the practically very important case when binary regression models are fit to binary data. In this case, in addition to selecting variables from a given design matrix, also the choice of an appropriate link function can be made in a data driven way. The general theory and the proposed methods are applicable irrespective of whether any of the candidate models under consideration is correctly specified, leading to more or less conservative inference depending on the severity of misspecification (see Remark 2.8). This feature is also present in the applications of Section 3.1 (see Subsection 3.1.2), 3.2 and 3.3. In simulation experiments we have illustrated that the confidence intervals constructed in the examples compare favorably to existing procedures (typically offering higher coverage with the confidence intervals having comparable or much smaller length), even though they are not tailored towards specific model selection procedures.
Open questions that go beyond the scope of this article, but are currently under investigation, include the extension of the approach discussed here to dependent data; the applicability and performance of bootstrap procedures; and the development of procedures in the spirit of Berk et al. (2013) in the challenging situation when the number of models fitted can grow with sample size. In ongoing work we apply our methods to real data and investigate if they can prevent spurious findings while detecting true reproducible effects.
Acknowledgements
Results related to the present article were presented in the Statistics and Econometrics Research Seminar at the Department of Statistics and Operations Research at the University of Vienna, and we would like to thank the participants, in particular Hannes Leeb, Benedikt M. Pötscher and Ulrike Schneider, for helpful comments and suggestions. We are also grateful for the comments and suggestions of two anonymous referees who helped to produce a considerably improved version of the paper.
Appendix A Simulation study
In this section, we investigate the confidence intervals suggested in this paper in a numerical study. It is an extended and more detailed version of Section 4 in the main article. We consider linear models and binary regression. When studying linear models, we first address the least angle regression (LAR) model selector Efron et al. (2004) and compare the confidence intervals of Theorem 3.2 with those developed in Tibshirani et al. (2015). The latter intervals are specifically tailored for the LAR model selector. Then, we investigate a model selection procedure which we call “significance hunting” and which is arguably representative of a certain practice of data mining.
When studying binary regression, we consider the lasso model selector, with a fixed regularization parameter . We compare the confidence intervals of Theorem 3.11 with those suggested by Taylor and Tibshirani (2017) and with “naive” confidence intervals. The confidence intervals of Taylor and Tibshirani (2017) are specific to the lasso model selector. The “naive” confidence intervals ignore the model selection step. We also investigate the significance hunting procedure in this case.
In Table 3, we report the coverage proportions, the median lengths and the quantiles of the lengths for each of the six procedures (“POSI” and “TG” for ), in different settings. We also report the proportions of times where the three targets are simultaneously contained by the three confidence intervals. We first observe that the results are approximately the same for the four types of error distributions, so that the Gaussian asymptotic approximation is accurate for these values of . The “POSI” confidence intervals always have target-specific and simultaneous coverage above the nominal level. The coverage proportions are large, which is so because these confidence intervals offer strong guarantees: they are valid for any model selection procedure, and simultaneously over all the variables in the selected model.
Turning to the “TG” confidence intervals, we observe that these intervals have coverage probabilities approximately equal to the nominal level when the three targets are considered separately. This is in agreement with the asymptotic guarantees obtained in Tibshirani et al. (2015). However, the simultaneous coverage is between and and thus always below the nominal level (). In Tibshirani et al. (2015), no asymptotic results are given concerning simultaneous coverage. This can be a practical limitation. If one wished to use the “TG” confidence intervals simultaneously they would have to increase their lengths, for instance by a Bonferroni correction.
We observe in Table 3 that the confidence intervals we suggest in this paper have slightly larger median length than the “TG” intervals (by about ) only in the “independent” design case and for the first step of the LAR procedure. In all the other cases, the “POSI” confidence intervals have smaller median lengths than the “TG” intervals. The difference of median lengths in these cases can be very significant. For instance, in the “correlated” design case, at the third step of LAR, the median length for the “TG” intervals is about times as large as for the “POSI” ones.
Finally, the quantiles are always larger for the “TG” intervals, for which they can be very large. In Table 3, we sometimes report infinite quantiles for the “TG” intervals. This is because, although the confidence intervals in Tibshirani et al. (2015) (in the two-sided case as is considered here) always have finite length in theory, the numerical implementation in the R package selectiveInference can return lower or upper bounds equal to . [When, say, the lower bound is equal to and the upper bound is larger than the target value, we consider the target to be covered.] In contrast, the confidence intervals suggested in this paper are more robust, in the sense that their quantile lengths are always less than twice as large as their median lengths.
We believe that the numerical results of Table 3 favor the “POSI” confidence intervals suggested in this paper over the “TG” procedure. Indeed, we have seen that, even though the LAR model selector is used, the “POSI” confidence intervals have larger coverage proportions, remain valid when considered simultaneously, generally have smaller median lengths, and never exhibit very large quantile lengths. On top of this, the “POSI” confidence intervals are significantly more broadly applicable, as they have theoretical guarantees for any model selection procedure.
One needs to mention here that Tibshirani et al. (2015) also discuss a bootstrap version of their “TG” intervals. These bootstrap confidence intervals have similar coverage properties as the “TG” intervals, but much smaller median width. Their width seems to be comparable to the width of our “POSI” confidence intervals (cf. Tables 1 and 2 in Tibshirani et al. (2015)). All the advantages of the “POSI” method discussed in the preceding paragraph (besides the comments concerning their smaller width) also apply to the bootstrapped “TG” intervals. Furthermore, this suggests that our “POSI” intervals could potentially be improved by using suitable bootstrap methods as well. However, answering this question goes beyond the scope of the present article.
A.1.2 The case of a significance hunting procedure
In Table 4, the coverage proportions are significantly lower than in Table 3, and closer to the nominal level. Hence, the confidence intervals suggested in this paper may have conservative coverage proportions for some model selection procedures (such as LAR) but this is somehow necessary, since there exist other model selection procedures (such as “significance hunting”) for which the coverage proportions are close to the nominal level.
A.2 Binary regression
The results are reported in Table 5. We observe that all the confidence interval lengths decrease when increases, which is natural. The “POSI” confidence intervals always have coverage proportions above the nominal level. In fact, the coverage proportions are quite large, which is explained by a similar argument as for Table 3: since the “POSI” intervals are valid for any model selection procedure, and simultaneously over the selected coefficients, they become conservative when applied specifically to the lasso and only for the first coefficient. In Table 5, the “POSI” intervals compare favorably with the “LASSO” ones. Indeed, the median lengths are generally comparable between the “POSI” and “LASSO” intervals (less than a factor between the two median lengths). Depending on the situation, any of these two intervals can have the smallest median length. On the other hand, the quantile lengths are always larger for the “LASSO” intervals. In some cases, they can be up to times as large as the “POSI” ones (for instance in the “well-specified”, “scaled” , “large” setting with ). Also, more importantly in our opinion, the “LASSO” intervals can have coverage proportions way below the nominal level (down to instead of ), even though the lasso model selector is used here. In fact, the “LASSO” intervals have sufficient coverage proportion only in the cases where has up to one non-zero coefficient, and the coverage proportion becomes too small otherwise. In particular, in the “misspecified” setting, the “LASSO” coverage becomes very small.
Turning to the “naive” intervals, we observe that they are smaller than the “POSI” ones, by approximately a factor . For both the “POSI” and “naive” intervals, the quantile lengths are moderately above the median lengths and are never very large. Importantly, the “naive” intervals can have coverage proportions significantly below the nominal level (down to instead of ). This is in agreement with the fact that these intervals have no asymptotic guarantee, in the post-model-selection context. Hence, these intervals are smaller than the “POSI” ones at the price of not offering reliable coverage properties. [Bachoc, Leeb and Pötscher (2014) and Leeb, Pötscher and Ewald (2015) reach a similar conclusion in the linear regression context.]
A.2.2 The case of a significance hunting procedure
with the notation of (A.1) and otherwise proceed as in Section A.1.2.
In Table 6, we report the results for the significance hunting procedure, where we have repeated data generations, model selections and confidence interval computations as for Table 4. We set , and . The design matrices are randomly generated as for Table 5 but with . We consider the “well-specified” case with equal to (“zero”) or to (“non-zero”). We set or and or .
We observe that for both intervals, similarly as in Table 5, the lengths decrease when goes from to and the quantile lengths are above the median lengths by a factor less than . Also, the “naive” intervals are about half of the length of the “POSI” ones. For the same reasons as for Table 4, the coverage proportions decrease when or when . The “POSI” intervals can have smaller coverage proportions than for the lasso model selector, down to (for a nominal level equal to ). Finally, the coverage proportions of the “naive” intervals are always much too small, with a minimum of . Hence we have another illustration, more pronounced than in Table 5, that these intervals do not offer reliable guarantees for post-model selection inference.
To conclude the simulation study in the binary case, we believe that the results in Tables 5 and 6 provide a complimentary picture of the confidence intervals suggested in this paper. Indeed, these intervals can be much shorter than the “LASSO” ones and they are never more than twice as large as the “LASSO” or “naive” intervals. Furthermore, only they have sufficient coverage proportions in all the settings studied. The two other types of intervals can yield significant under-coverage. More precisely, the ‘LASSO” intervals can exhibit strong under-coverage even though the lasso model selector is used. The “naive” intervals can yield small coverage proportions for the lasso model selector, and yield even smaller ones for the significance hunting procedure. Finally, the “POSI” intervals are asymptotically valid for any model selection procedure, while the “LASSO” ones can only be used in conjunction with the lasso model selector, and the “naive” ones do not have asymptotic guarantees in the post-model-selection context at all.
Appendix B Auxiliary results
Furthermore, for every we have
The first statement in the subsequent lemma is essentially Corollary 2 in Pollak (1972) combined with a tightness argument. The second statement is obtained via an application of Raikov’s theorem (Raikov (1938), cf. the statement given in Gnedenko and Kolmogorov (1954) on p. 143).
Furthermore, for every it holds that
The statement in Equation (B.8) is an immediate consequence of the triangle inequality and the first two statements. ∎
Suppose Condition 3 holds. Then, for every , we have
Furthermore, for every , satisfies
For the statements in Equations (B.17) and (B.18) we first note that the triangular array , and the corresponding quantities and , satisfy Condition 2. Hence, Lemma B.1 is applicable, and shows, in particular, for every and with the abbreviation , that
which we obtain (as above) from Lemma B.1. ∎
Let be a sequence of covariance matrices converging to . By definition, is the -quantile of the distribution of , where is a Gaussian random vector with mean and covariance matrix . By the continuous mapping theorem, converges weakly to , where is a Gaussian random vector with mean and covariance matrix . In case it is easy to see that the distribution function of is everywhere continuous and strictly increasing on , and the result then follows, because weak convergence of distribution functions is equivalent to weak convergence of the corresponding quantile functions. Consider now the case where . Fix . Let be a random variable taking values in , with continuous and strictly increasing (on ) distribution function and -quantile equal to . Clearly, converges weakly to . Hence , say, the quantile of converges to . From it then follows that
Therefore, . ∎
For the first part note that the quotient under consideration is well defined with probability converging to one, that
and that by the Cauchy-Schwarz inequality
For the second part note that the quotients are well defined with probability converging to 1 (by applying Part 1), and write
For the third part we note that (eventually)
where the second equality for follows from the last assumption appearing in Part 3 together with (B.25), and the second equality for follows from (B.25). By the Cauchy-Schwarz inequality and the last assumption appearing in Part 3
and it holds, using uncorrelatedness of for , that
We need to verify that for every it holds that
where we used and Assumption (B.26) to obtain the limit. ∎
Appendix C Proofs for Section 2
We actually prove the following more detailed statement.
Under Condition 1, for we have
C.2 Proof of Theorem 2.4
It hence suffices to verify that the lower bound converges to . And for that (cf. Equation (C.11) below, and Condition 1) it suffices to verify that the following quantity converges to :
Lemma C.1 shows that for every we have
This also shows that the two conditions imposed on in the statement of the theorem are indeed equivalent. Furthermore, we immediately see that from any of these two assumptions, together with the previous display, it follows that for every we have
Since the diagonal elements of are ones (by its definition together with Condition 1), one can easily show that the -probability of being equal to is . It hence follows from the Portmanteau theorem, together with the definition of and the previous display, that
C.3 Proof of Proposition 2.5
Next, we can, in a similar way, apply the second part of Lemma B.4 to obtain
C.4 Proof of Theorem 2.6
Similarly as in the proof of Theorem 2.4 we now need to verify that
which, in turn, is bounded from below (using that is positive on , and Equation (2.11)) by
Now, we argue by contradiction, and suppose that (C.21) is false: Then there exists a , that can be chosen independently of , so that the limit inferior in (C.23) is an element of . Next, let denote a subsequence along which (C.23) is attained. Arguing as in the proof of Theorem 2.4 (borrowing some of its notation) we can obtain a subsequence of along which the sequence of probabilities in the preceding display converges to
Note that was arbitrary, and let converge to . Assume (otherwise pass to a subsequence) that the sequence of correlation matrices (with diagonal entries equal to ) converges to , say. It is then not difficult to obtain (by a weak convergence argument involving Portmanteau theorem) the contradiction
The remaining part follows immediately from what we have already established. ∎
C.5 Proof of Proposition 2.7
Fix and note that with the same notation and argumentation as in the beginning of the proof of Proposition 2.5, and with the convention of Remark 2.1, for every , it holds that
Equation (2.9), or equivalently (equivalence being due to Equation (C.27) above) Equation (2.10), together with Part 3 of Lemma B.4 (applied with: , and ) now shows that
implying the claimed statement. Note that Equation (B.26) in Lemma B.4 is satisfied here because Condition 1 (in particular the Lindeberg condition in Equation (2.3)) implies the corresponding Feller condition
All remaining assumptions in Part 3 of Lemma B.4 can be easily checked using Condition 1, (2.9) and (C.27). ∎
C.6 Proof of Lemma 2.9
Let and let . Since is a correlation matrix of rank , by the spectral decomposition, we can find a -dimensional matrix so that . In particular if it holds that , and hence the -quantiles of the distributions of and of coincide, denoting the -th row of . Since is a correlation matrix it furthermore holds that each row of has Euclidean norm less than or equal to . From the discussion after the definition of it then follows that , the -quantile of the distributions of , is not greater than . ∎
Appendix D Proofs for Section 3
D.2 Proof of Theorem 3.2
for . The -th coordinate of the estimation error
together with Condition X1. We now verify that for every it holds that
An application of Hölder’s inequality (with and ) shows that the quantity to the left in the previous display is bounded from above by
which, using the bound (D.2) and Markov’s inequality, does not exceed
Finally, since the term within brackets coincides with
and since this quantity is not greater than
D.3 Proof of Theorem 3.3
D.4 Proof of Proposition 3.5
D.5 Proof of Theorem 3.6
To verify Condition 1 we replace (D.6) in the proof of Theorem 3.2 by
which, replacing each by , is seen to be eventually positive by Condition X1. For the verification of the Lindeberg condition (D.7) we use essentially the same argument as in the proof of Theorem 3.2, now using (D.18) above. Hence, Condition 1 holds. Next, we verify (2.13) in Theorem 2.6 by means of Proposition 2.7. Let
D.6 Proof of Lemma 3.9
D.7 Auxiliary results for Section 3.3
Suppose that Conditions X2(i,ii) and H(i,ii) hold and fix . There exists a finite positive constant , depending only on and the constant from Condition X2(ii), such that eventually
Suppose that Conditions X2(i,ii) and H(i,ii,iii) hold and fix .
Suppose that, in addition, also Condition X2(iii) holds. Then there exists a positive finite constant , depending only on and the constant from Condition X2, such that eventually
D.8 Proof of Lemma 3.10
for all . Therefore, we have the inclusion
Choosing and using Lemma D.2(ii), we conclude that for every ,
D.9 Proof of Theorem 3.11
But the numerator of the second fraction on the right of the previous display can be bounded by
where we have omitted another constant that depends only on and . We have thus verified Condition 1.
by Lemma D.2(iii). Furthermore, for every and ,
D.10 Canonical link function
In the setting of Theorem 3.11, if contains only the canonical link function , then the confidence intervals