On overfitting and post-selection uncertainty assessments

Liang Hong, Todd A. Kuffner, Ryan Martin

Introduction

Consider the classical multiple linear regression model

can be used for inference on the mean response at the given xx. However, as is now well-known (Berk et al. 2013), the properties that these classical procedures enjoy for a fixed/true SS may not hold for a data-dependent choice, S^\hat{S}. For example, Cα(x;S^)C_{\alpha}(x;\hat{S}) may not have coverage probability equal to 1−α1-\alpha.

This note provides an explanation of this lack-of-validity phenomenon by showing that, when the sub-model is selected according to information criteria such as aic and bic, if the selected sub-model overfits, i.e., contains a superset of the explanatory variables in the true model, then the corresponding estimate of the error variance will be smaller than that for the true model. This explains the empirical findings in Hong et al. (2017), where prediction intervals based on the sub-model minimizing aic tend to be too short compared to those based on the true model and, consequently, they tend to undercover; see Section 3. Moreover, our Theorem 1 together with the dilation phenomenon described in Efron (2003), explains why bootstrap may not correct the selection effect for methods that tend to overfit.

Result

For a given sub-model Θ(S)\Theta(S), corresponding to a subset S⊆{1,…,p}S\subseteq\{1,\ldots,p\}, let (β^S,σ^S)(\hat{\beta}_{S},\hat{\sigma}_{S}) denote the least squares estimators of the Θ(S)\Theta(S)-specific parameters (βS,σS)(\beta_{S},\sigma_{S}). We consider a selection procedure that chooses the subset SS by minimizing the function

where \scsse(S)=∥y−XSβ^S∥2\text{\sc sse}(S)=\|y-X_{S}\hat{\beta}_{S}\|^{2} is the error sum of squares for sub-model Θ(S)\Theta(S), which is proportional to the corresponding least squares estimator σ^S2\hat{\sigma}_{S}^{2}, cn=o(n)c_{n}=o(n) is a user-specified sequence of constants, and ∣S∣|S| denotes the cardinality of the set SS. The aic and bic set cn≡2c_{n}\equiv 2 and cn=log⁡nc_{n}=\log n, respectively.

Suppose that there exists a subset S⋆S^{\star} corresponding to the truly non-zero regression coefficients, i.e., βi≠0\beta_{i}\neq 0 for i∈S⋆i\in S^{\star} and βi=0\beta_{i}=0 for i∉S⋆i\not\in S^{\star}. We write (β^S⋆(\hat{\beta}_{S^{\star}}, σ^S⋆)\hat{\sigma}_{S^{\star}}) for the oracle estimators, those based on knowledge of the true sub-model Θ(S⋆)\Theta(S^{\star}). Of course, if S^\hat{S} is the subset chosen by minimizing γn\gamma_{n} in (3), then γn(S^)≤γn(S⋆)\gamma_{n}(\hat{S})\leq\gamma_{n}(S^{\star}) or, equivalently,

if S^≠S⋆\hat{S}\neq S^{\star}, then the inequality in (4) would be strict.

For the purpose of inference or prediction, it is common to naively use the classical normal linear model theory, based on the selected subset S^\hat{S}, to derive uncertainty assessments. However, using the data to select S^\hat{S} introduces bias, violating the assumptions of that classical theory, and thereby invalidating the conclusions. The next result provides an explanation for this general phenomenon in cases where the selected sub-model Θ(S^)\Theta(\hat{S}) overfits in the sense that S^⊃S⋆\hat{S}\supset S^{\star}. In such cases, we find that σ^S^\hat{\sigma}_{\hat{S}} is smaller than the oracle estimator σ^S⋆\hat{\sigma}_{S^{\star}}. Since the error variance estimate is involved in all uncertainty assessment calculations, and since it is common for selection methods to overfit, especially those based on aic (Hurvich & Tsai 1989), this systematic under-estimation explains the general lack of validity of the classical inferential tools applied naively in a post-selection context.

where an=(cn/n)(n−∣S⋆∣−1)a_{n}=(c_{n}/n)(n-|S^{\star}|-1) and Dn=(∣S^∣−∣S⋆∣)/(n−∣S⋆∣−1)D_{n}=(|\hat{S}|-|S^{\star}|)/(n-|S^{\star}|-1), then σ^S^<σ^S⋆\hat{\sigma}_{\hat{S}}<\hat{\sigma}_{S^{\star}}.

To gain some intuition about the condition (5), first note that anDna_{n}D_{n} will tend to be small. In particular, a very conservative bound is anDn≤cnp/na_{n}D_{n}\leq c_{n}p/n, which is small for moderate cnc_{n} and n≫pn\gg p. Next, since x↦1−exp⁡(−ax)x\mapsto 1-\exp(-ax) is convex for x>0x>0 and a>0a>0, we have 1−exp⁡(−anDn)>anDn1-\exp(-a_{n}D_{n})>a_{n}D_{n} for all DnD_{n} in an interval (0,d)(0,d), where d=d(an)∈[0,1)d=d(a_{n})\in[0,1). So, to meet (5) we need an>1a_{n}>1 and, again, we have a conservative bound an≥cn(n−p−1)/na_{n}\geq c_{n}(n-p-1)/n, which itself is greater than 1 for n≫pn\gg p and cnc_{n} not too small. In particular, if n≫pn\gg p and cn≡2c_{n}\equiv 2 as in the aic, then (5) holds.

Start by writing \scsse(S^)\text{\sc sse}(\hat{S}) in terms of \scsse(S⋆)\text{\sc sse}(S^{\star}). Let XS^X_{\hat{S}} and XS⋆X_{S^{\star}} denote the sub-matrices corresponding to the indicated subsets, and write PS^P_{\hat{S}} and PS⋆P_{S^{\star}} for the respective projections onto their column spaces. Then Pythagoras’ theorem implies that

and Fn(S⋆,S^)F_{n}(S^{\star},\hat{S}) is the usual F-statistic for testing the larger Θ(S^)\Theta(\hat{S}) against the smaller Θ(S⋆)\Theta(S^{\star}). Consequently, we choose S^\hat{S} over the strictly smaller S⋆S^{\star}, according to (4), if and only if rn>1−exp⁡(−anDn)r_{n}>1-\exp(-a_{n}D_{n}).

Then the above connection between \scsse(S^)\text{\sc sse}(\hat{S}) and \scsse(S⋆)\text{\sc sse}(S^{\star}) immediately gives a comparison between the corresponding variance estimates:

As above, we find that σ^S^<σ^S⋆\hat{\sigma}_{\hat{S}}<\hat{\sigma}_{S^{\star}} if and only if rn>Dnr_{n}>D_{n}. By condition (5), it follows that the lower bound on rnr_{n} derived from over-fitting is greater than that derived from the under-estimation. Therefore, over-fitting implies under-estimation, proving the claim. ∎

Illustration

Consider the model (1), with n=50n=50 and p=10p=10, and variance σ2=1\sigma^{2}=1. Set S⋆={1,2,3}S^{\star}=\{1,2,3\}, with corresponding coefficients β1⋆=1\beta_{1}^{\star}=1, β2⋆=2\beta_{2}^{\star}=2, and β3⋆=3\beta_{3}^{\star}=3. The rows of the XX matrix are independent, pp-variate normal, with mean zero, AR(1) dependence structure, and one-step correlation ρ=0.5\rho=0.5. We simulated 1000 data sets and, for each, evaluated σ^S^\hat{\sigma}_{\hat{S}} and σ^S⋆\hat{\sigma}_{S^{\star}}, where S^\hat{S} is chosen based on the aic. The scatterplot shown in Figure 1(a) demonstrates the systematic under-estimation based on the AIC-selected sub-model, as predicted by Theorem 1. In all 1000 cases, we have S^⊇S⋆\hat{S}\supseteq S^{\star}, and those on the diagonal line correspond to S^=S⋆\hat{S}=S^{\star}. To further illustrate the difference between the estimates, Figure 1(b) plots a histogram of the ratio σ^S⋆/σ^S^\hat{\sigma}_{S^{\star}}/\hat{\sigma}_{\hat{S}}, only for the strict over-fit cases. In particular, the mean from this histogram is 1.06.

While the relative difference between the two estimates does not seem remarkable, even this small of a difference can impact the quality of inference. For example, consider using the confidence interval (2) for inference on the mean response at a particular setting xx of the explanatory variables; here, xx is an independent sample from the distribution that generated the rows of XX. The oracle 95% confidence interval C0.05(x;S⋆)C_{0.05}(x;S^{\star}) has coverage exactly equal to 0.950.95 but, in the 1000 simulations above, the coverage probability of Cα(x;S^)C_{\alpha}(x;\hat{S}) is roughly 0.86. It happens that the S^\hat{S}-based intervals tend to be shorter than the oracle, suggesting that valid post-selection inference on the mean response requires σ^S^\hat{\sigma}_{\hat{S}} to be strictly larger than σ^S⋆\hat{\sigma}_{S^{\star}}, which is impossible given Theorem 1 and the aic’s tendency to over-fit.

Acknowledgement

The authors are grateful to the Editor and two referees whose comments greatly enhanced the clarity of our presentation. Kuffner was supported by the National Science Foundation, U.S.A.

References