Supplementary Appendix for "Inference on Treatment Effects After Selection Amongst High-Dimensional Controls"

Alexandre Belloni, Victor Chernozhukov, Christian Hansen

Intuition for the Importance of Double Selection

To build intuition, we discuss the case where there is only one control; that is, p=1p=1. This scenario provides the simplest possible setting where variable selection might be interesting. In this case, Lasso-type methods act like conservative tt-tests which allows the properties of selection methods to be explained easily.

For simplicity, all errors and controls are taken as normal,

The standard post-single-selection method for inference proceeds by applying model selection methods – ranging from standard tt-tests to Lasso-type selectors – to the first equation only, followed by applying OLS to the selected model. In the model selection stage, standard selection methods would necessarily omit xix_{i} wp →1\to 1 if

where σn∗2=σζ2(σd2)−1\sigma^{*2}_{n}=\sigma^{2}_{\zeta}(\sigma^{2}_{d})^{-1}. The variance σn∗2\sigma^{*2}_{n} is smaller than the variance σn2\sigma^{2}_{n} from estimation with xix_{i} included if βm≠0\beta_{m}\neq 0. The potential reduction in variance is often used as a “motivation” for the standard selection procedure. The estimator is super-efficient, achieving a variance smaller than the semi-parametric efficiency bound under homoscedasticity. That is, the estimator is “too good.”

That is, the standard post-selection estimator is not asymptotically normal and even fails to be uniformly consistent at the rate of n\sqrt{n}. This poor behavior occurs because the omitted variable bias created by dropping xix_{i} may be large even when the magnitude of the regression coefficient, ∣βm∣|\beta_{m}|, in the confounding equation (1.2) is small but is not exactly zero. To see this, note

The term i∗i^{*} has standard behavior; namely i∗⇝N(0,1)i^{*}\rightsquigarrow N(0,1). The term iiii generates the omitted variable bias, and it may be arbitrarily large, since wp →1\to 1,

In contrast to the standard approach, our post-double-selection method for inference proceeds by applying model selection methods, such as standard tt-tests or Lasso-type selectors, to both equations and taking the selected controls as the union of controls selected from each equation. This selection is than followed by applying OLS to the selected controls. Thus, our approach drops xix_{i} only if the omitted variable bias term iiii is small. To see this, note that the double-selection-methods include xix_{i} wp →1\to 1 if its coefficient in either (1.1) or (1.2) is not very small. Mathematically, xix_{i} is included if

Given the discussion in the preceding paragraph, it is immediate that the post-double selection estimator satisfies

i.e. coefficients in front of xix_{i} in both equations are small. In this case,

Once again, the term iiii is due to omitted variable bias, and it obeys wp →1\to 1 under (1.13)

since (σv/σd)2=1−ρ2(\sigma_{v}/\sigma_{d})^{2}=1-\rho^{2}. Moreover, we can show i∗−i=oP(1)i^{*}-i=o_{P}(1) under such sequences, so the first order asymptotics of αˇ\check{\alpha} is the same whether xix_{i} is included or excluded.

Extensions: Other Problems and Heterogeneous Treatment Effects

In order to discuss extensions in a very simple manner, we assume i.i.d sampling as well as assume away approximation errors, namely g(zi)=xi′βg0g(z_{i})=x_{i}^{\prime}\beta_{g0} and m(zi)=xi′βm0m(z_{i})=x^{\prime}_{i}\beta_{m0}, where parameters βg0\beta_{g0} and βm0\beta_{m0} are high-dimensional and that xi=P(zi)x_{i}=P(z_{i}) as before. In this paper we considered a moment condition:

where φ(u)=u\varphi(u)=u and viv_{i} are measurable functions of ziz_{i}, and the target parameter is α0\alpha_{0}. We selected the instrument viv_{i} such that the equation is first-order insensitive to the parameter β\beta at β=βg0\beta=\beta_{g0}:

Note that φ(u)=u\varphi(u)=u and vi=di−m(zi)v_{i}=d_{i}-m(z_{i}) implement this condition. If (2.15) holds, the estimator of α0\alpha_{0} gets “immunized” against nonregular estimation of β0\beta_{0}, for example, via a post-selection procedure or other regularized estimators. Such immunization ideas are in fact behind the classical Frisch-Waugh-Robinson partialling out technique in the linear setting and the ?’s C(α)C(\alpha) test in the nonlinear setting. Our contribution here is to recognize the importance of this immunization in the context of post-selection inference, To the best of our knowledge, all prior theoretical and empirical work uses the standard post-selection approach based on the outcome equation alone, which is highly non-robust way of conducting inference, as shown in extensive monte-carlo, in Section 2.4, and in a sequence of fundamental critiques by Leeb and Potscher, see ? . to develop a robust post-selection approach to inference on the target parameter, and characterize the uniformity regions of this procedure. Our approach uses modern selection methods to estimate gg and the function mm defining the intstrument vv. In an ongoing work, we explore other regularization methods, such as the ridge method or combination of ridge method with Lasso methods, and characterize uniformity regions of the resulting procedures. Also, generalizations to nonlinear models, where φ\varphi is non-linear and can correspond to a likelihood score or quantile check function are given in ? and ?; in these generalizations achieving (2.15) is also critical.

Within the context of this paper, a potentially important extension is to consider a general treatment effect model, where did_{i} is interacted with transformations of ziz_{i}. As long as the interest lies in a particular regression coefficient, the current framework covers this implicitly since xix_{i} could contain interactions of did_{i} with transformations of controls ziz_{i}. In the case a fixed number of such regression coefficients is of interest, we can estimate each of the coefficients by re-labeling the corresponding regressor as did_{i} and other regressors as xix_{i} and then applying our procedure the fixed number of times. Such component-wise procedure is valid as long as our regularity conditions hold for each of the resulting regression models in this manner.

A related research direction being pursued is the study of estimation of average treatment effects when treatment effects are fully heterogeneous. When the treatment variable di∈{0,1}d_{i}\in\{0,1\} is binary (or discrete more generally) our approach is readily amenable to this problem. In this case the parameter of interest is the average treatment effect

where φ(α,y,d,g0,g1,m)=α−d(y−g1)m+(1−d)(y−g0)1−m−(g1−g0).\varphi(\alpha,y,d,g_{0},g_{1},m)=\alpha-\frac{d(y-g_{1})}{m}+\frac{(1-d)(y-g_{0})}{1-m}-(g_{1}-g_{0}). It is straightforward to check that for each j∈{1,2,3}j\in\{1,2,3\}:

where the latter is the semiparametric efficiency bound of ?. Formal result is proven and stated below, and other formal results along these lines are given in an ongoing work that studies this as well as other types of effects.

2. Theoretical Results on ATE with Heterogeneity

The confounding factors ziz_{i} affect the policy variable via the propensity score m(zi)m(z_{i}) and the outcome variable via the function g(di,zi)g(d_{i},z_{i}). Both of these functions are unknown and potentially complicated. As in the main text, we use linear combinations of control terms xi=P(zi)x_{i}=P(z_{i}) to approximate g(zi)g(z_{i}) and m(zi)m(z_{i}), writing (2.19) and (2.20) as

where rgir_{gi} and rmir_{mi} are the approximation errors, and

where xi′βg0,1x_{i}^{\prime}\beta_{g0,1}, xi′βg0,0x_{i}^{\prime}\beta_{g0,0}, and xi′βm0x_{i}^{\prime}\beta_{m0} are approximations to g(1,zi)g(1,z_{i}), g(0,zi)g(0,z_{i}), and m(zi)m(z_{i}), and Λ(u)=u\Lambda(u)=u for the case of linear link and Λ(u)=eu/(1+eu)\Lambda(u)=e^{u}/(1+e^{u}) for the case of the logistic link. In order to allow for a flexible specification and incorporation of pertinent confounding factors, the vector of controls, xi=P(zi)x_{i}=P(z_{i}), we can have a dimension p=pnp=p_{n} which can be large relative to the sample size.

The efficient moment condition, derived by ?, for parameter α0\alpha_{0} is as follows:

The post-double-selection estimator αˇ\check{\alpha} that solves

where g^(di,zi)\widehat{g}(d_{i},z_{i}) and m^(zi)\widehat{m}(z_{i}) are post-Lasso estimators of functions gg and mm based upon equations (2.21)-(2.22). In case of the logistic link Λ\Lambda, Lasso for logistic regression is as defined in ? and ?, and the associated post-Lasso estimators are as those defined in ?.

These conditions are simple high-level conditions, which encode both the approximate sparsity of the models as well as impose some reasonable behavior on the post-selection estimators of mm and gg (or other sparse estimators). These conditions are implied by other more primitive conditions in the literature. Sufficient conditions for the equivalence between population and empirical sparse eigenvalues are given in ? and ?. The boundedness conditions are made to simplify arguments, and they could be dealt away with more complicated proofs, under more stringent side conditions.

The next target parameter is the average treatment effect on the treated:

The efficient moment condition, derived by ?, for parameter γ0\gamma_{0} is as follows:

In this case the post-double-selection estimator γˇ\check{\gamma} that solves

3. Proof of Theorems 1 and 2.

The two results have identical structure and have nearly the same proof, and so we present the proof of the Proof of Theorems 1 only.

Step 1. In this step we establish claim (1).

(a) We begin with a preliminary observation. Define, for t=(t1,t2,t3)t=(t_{1},t_{2},t_{3}),

where LL depends only on c′c^{\prime} and CC, ∣k∣=∑j=13kj,|k|=\sum_{j=1}^{3}k_{j}, and

We observe that with probability no less than 1−Δn1-\Delta_{n},

for β=β^g\beta=\widehat{\beta}_{g}, with evaluation after computing the norms, and noting that for any β\beta

under condition (iii). Furthermore, for n⩾n0=min⁡{j:δj⩽1/2}n\geqslant n_{0}=\min\{j:\delta_{j}\leqslant 1/2\}:

for β=β^m0\beta=\widehat{\beta}_{m0}, with evaluation after computing the norms.

Hence with probability at least 1−Δn1-\Delta_{n},

with hh evaluated at h=h^h=\widehat{h}. By Liapunov central limit theorem,

(d) Note that for Δi=h(zi)−h0(zi)\Delta_{i}=h(z_{i})-h_{0}(z_{i}),

(with hh evaluated at h=h^h=\widehat{h}). By the law of iterated expectations and because

Moreover, uniformly for any h∈Hnh\in\mathcal{H}_{n} we have that

Since h^∈Hn\widehat{h}\in\mathcal{H}_{n} with probability 1−Δn1-\Delta_{n}, we have that once n⩾n0n\geqslant n_{0},

The class of functions Gd\mathcal{G}_{d} for d∈{0,1}d\in\{0,1\} is a union of at most (pCs)\binom{p}{Cs} VC-subgraph classes of functions with VC indices bounded by C′sC^{\prime}s. The class of functions M\mathcal{M} is a union of at most (pCs)\binom{p}{Cs} VC-subgraph classes of functions with VC indices bounded by C′sC^{\prime}s (monotone transformation Λ\Lambda preserve the VC-subgraph property). These classes are uniformly bounded and their entropies therefore satisfy

Finally, the class Fn={fh−fh0:h∈Hn}\mathcal{F}_{n}=\{f_{h}-f_{h_{0}}:h\in\mathcal{H}_{n}\} is a Lipschitz transform of Hn\mathcal{H}_{n} with bounded Lipschitz coefficients and with a constant envelope. Therefore, we have that

We shall invoke the following lemma derived in ?.

Let F\mathcal{F} be a measurable function class on a sample space. Let F=sup⁡f∈F∣f∣F=\sup_{f\in\mathcal{F}}|f|, and suppose that there exist some constants ωn>3\omega_{n}>3 and υ>1\upsilon>1, such hat

Then for every δ∈(0,1/6)\delta\in(0,1/6) we have

with probability at least 1−δ1-\delta for some constant that CυC_{\upsilon}.

Then by Lemma 1 together and some simple calculations, we have that

Step 3. Claim (3) is immediate from claims (2) and (3) by the way of contradiction. ∎

Deferred Proofs: Proof of Lemma 1

We establish the result for Lasso (the proof for other feasible Lasso estimators is similar).

By Condition SE, with probability 1−o(1)1-o(1) for nn sufficiently large we have κcˉ>κ′/2∥Ψ^∥∞\kappa_{\bar{c}}>\kappa^{\prime}/2\|\widehat{\Psi}\|_{\infty} so that with the same probability

Moreover, by condition RF we have with probability 1−o(1)1-o(1) that

Finally, since λ≳nlog⁡(p∨n)\lambda\gtrsim\sqrt{n\log(p\vee n)} we have

since cs≲Ps/nc_{s}\lesssim_{P}\sqrt{s/n} by condition ASM and Chebyshev inequality.

To show the second statement in (i), note that

where β^\widehat{\beta} is the Lasso estimator. Again by Lemma 7 in ? we have that the assumptions of Lemma 6 in ? hold with probability 1−o(1)1-o(1). Using Condition SE to bound κcˉ\kappa_{\bar{c}} from below and Condition RF to bound ∥Ψ^∥∞\|\widehat{\Psi}\|_{\infty} from above with probability 1−o(1)1-o(1) as before, and λ≲σnlog⁡(p∨n)\lambda\lesssim\sigma\sqrt{n\log(p\vee n)}, it follows from Lemma 6 in ? that with probability 1−o(1)1-o(1) that

The results regarding Post-Lasso in (ii) follow similarly by invoking Lemma 8 in ?.

Split-Sample Estimation and Inference

In this section we discuss a variant of the double selection estimator based on sample splitting. The motivation for the split-sample estimator is that its use allows us to relax the requirement s2log⁡2(p∨n)=o(n)s^{2}\log^{2}(p\vee n)=o(n) that is assumed in the full-sample counterpart to the milder condition

For each subsample k=a,bk=a,b, the model I^k\widehat{I}^{k} is selected based on the subsample kk independently from the subsample kck^{c}. In what follows the model I^k\widehat{I}^{k} is used to fit the subsample kck^{c}. A constructive way to obtain I^a\widehat{I}^{a} and I^b\widehat{I}^{b} is to apply the double selection method for each subsample to select the sets of controls I^a:=I^1a∪I^2a∪I^3a\widehat{I}^{a}:=\widehat{I}_{1}^{a}\cup\widehat{I}_{2}^{a}\cup\widehat{I}_{3}^{a} and I^b:=I^2b∪I^2b∪I^3b\widehat{I}^{b}:=\widehat{I}_{2}^{b}\cup\widehat{I}_{2}^{b}\cup\widehat{I}_{3}^{b}.

Then we form estimates in the two subsamples

For an index ii in the subsample kk, we define the residuals

Finally, we combine the estimates into the split-sample estimator based on I^a\widehat{I}^{a} and I^b\widehat{I}^{b} is defined as

where Υk=Dk′MI^kcDk/nk\Upsilon^{k}=D^{k}{}^{\prime}\mathcal{M}_{\widehat{I}^{k^{c}}}D^{k}/n_{k}.

We state below sufficient conditions for the analysis of the split-sample method.

The Conditions ASTESS(i)-(iii) agree with the corresponding conditions in ASTE. The remaining conditions ASTESS(iv)-(v) are implied by Condition ASTE. We note that Condition ASTESS(vi) is needed only for obtaining consistent estimates of the asymptotic variance. Such conditions are mild since they do not require uniform estimation of the functions gg and mm.

The next result establishes that the split-sample estimator αˇab\check{\alpha}_{ab} has similar large sample properties to the full-sample double-selection estimator under weaker growth condition.

Step 0.(Combining) In this step we combine both subsample estimators. Letting Υk=Dk′MI^kcDk/nk\Upsilon^{k}=D^{k}{}^{\prime}\mathcal{M}_{\widehat{I}^{k^{c}}}D^{k}/n_{k}, for k=a,bk=a,b, so that we have

which follows similarly to the proofs given in Step 5.

where zi,n=σn−1viζi/nz_{i,n}=\sigma_{n}^{-1}v_{i}\zeta_{i}/\sqrt{n} are i.n.i.d. with mean zero. We have that for some small enough δ>0\delta>0

This condition verifies the Lyapunov condition and thus implies that Zn→dN(0,1)Z_{n}\to_{d}N(0,1).

Step 1.(Main) For the subsample k=a,bk=a,b write αˇk=[Dk′MI^kcDk/nk]−1[Dk′MI^kcYk/nk]\check{\alpha}_{k}=\left[D^{k}{}^{\prime}\mathcal{M}_{\widehat{I}^{k^{c}}}D^{k}/n_{k}\right]^{-1}[D^{k}{}^{\prime}\mathcal{M}_{\widehat{I}^{k^{c}}}Y^{k}/n_{k}] so that

First, note that by Condition ASTESS we have

Second, by the split sample construction, we have that I^kc\widehat{I}^{k^{c}} is independent from ζk\zeta^{k}, and by assumption of the model mkm^{k} is also independent of ζk\zeta^{k}. Thus by Chebyshev inequality

where the last relation follows by ASTESS.

Third, using similar independence arguments, by Chebyshev and Condition ASTESS, conclude

Fourth, using that s^kc≲Ps\widehat{s}^{k^{c}}\lesssim_{P}s by ASTESS so that ϕmin⁡−1(s^kc)≲P1\phi^{-1}_{\min}(\widehat{s}^{k^{c}})\lesssim_{P}1 by condition SE, we have that

Step 3.(Behavior of iikii_{k}.) Since iik=(mk+Vk)′MI^kc(mk+Vk)/nkii_{k}=(m^{k}+V^{k})^{\prime}\mathcal{M}_{\widehat{I}^{k^{c}}}(m^{k}+V^{k})/n_{k}, decompose

Then ∣iik,a∣=oP(1)|ii_{k,a}|=o_{P}(1) by Condition ASTESS, ∣iik,b∣=oP(1)|ii_{k,b}|=o_{P}(1) by reasoning similar to deriving the bound for ∣ik,b∣|i_{k,b}|, and ∣iik,c∣=oP(1)|ii_{k,c}|=o_{P}(1) by reasoning similar to deriving the bound for ∣ik,d∣|i_{k,d}|.

By condition ASTESS ∥MI^kcgk∥=oP(n1/4)\|\mathcal{M}_{\widehat{I}^{k^{c}}}g^{k}\|=o_{P}(n^{1/4}) and by condition SM(ii) we have ∥PI^kcDk/nk∥⩽∥Dk/nk∥≲P1\|\mathcal{P}_{\widehat{I}^{k^{c}}}D^{k}/\sqrt{n_{k}}\|\leqslant\|D^{k}/\sqrt{n_{k}}\|\lesssim_{P}1, and by Step 1 we have ∣αˇk−α0∣≲Pn−1/2|\check{\alpha}_{k}-\alpha_{0}|\lesssim_{P}n^{-1/2}. Moreover,

We have ϕmax⁡,k(s^kc)/ϕmin⁡,k(s^kc)≲P1\sqrt{\phi_{\max,k}(\widehat{s}^{k^{c}})}/\phi_{\min,k}(\widehat{s}^{k^{c}})\lesssim_{P}1 by condition SE, and ∥Xk[Ikc]′ζk/nk∥≲Ps^kc\|X^{k}[I^{k^{c}}]^{\prime}\zeta^{k}/\sqrt{n_{k}}\|\lesssim_{P}\sqrt{\widehat{s}^{k^{c}}} by condition SM(ii), the independence between the selected components I^kc\widehat{I}^{k^{c}} and ζk\zeta^{k} since they are based on different subsamples, and applying Chebyshev inequality.

Similarly, we have ∥mk−Xkβ^k∥/nk≲Po(n−1/4)+s^kc/nk\|m^{k}-X^{k}\widehat{\beta}_{k}\|/\sqrt{n^{k}}\lesssim_{P}o(n^{-1/4})+\sqrt{\widehat{s}^{k^{c}}/n_{k}}.

Step 5.(Variance Estimation.) Since s^k≲Ps=o(n)\widehat{s}^{k}\lesssim_{P}s=o(n), (nk−s^k−1)/nk=oP(1)(n_{k}-\widehat{s}^{k}-1)/n_{k}=o_{P}(1), so we can use nn as the denominator. Recall the definitions ζ^io=yi−diαˇk−xi′βˇk\widehat{\zeta}_{i}^{o}=y_{i}-d_{i}\check{\alpha}_{k}-x_{i}^{\prime}\check{\beta}_{k}, v^i=di−xi′β^k\widehat{v}_{i}=d_{i}-x_{i}^{\prime}\widehat{\beta}_{k} and ζ^i=ζ^io1{∣ζ^io∣∨∣v^i∣⩽Hk}\widehat{\zeta}_{i}=\widehat{\zeta}_{i}^{o}1\{|\widehat{\zeta}_{i}^{o}|\vee|\widehat{v}_{i}|\leqslant H_{k}\} if ii belongs to subsample kk where Hk=Cn/[(s^kc∨n1/2)log⁡n]H_{k}=C\sqrt{n/[(\widehat{s}^{k^{c}}\vee n^{1/2})\log n]}. For notational convenience let Ai={∣ζ^io∣∨∣v^i∣⩽Hk}A_{i}=\{|\widehat{\zeta}_{i}^{o}|\vee|\widehat{v}_{i}|\leqslant H_{k}\}. Since q>4q>4, s^kc≲Ps\widehat{s}^{k^{c}}\lesssim_{P}s, and n2/qslog⁡(n∨p)=o(n)n^{2/q}s\log(n\vee p)=o(n), we have n1/q=oP(Hk)n^{1/q}=o_{P}(H_{k}). Hence consider

Finally, since max⁡i⩽n∥1{Ai}(v^i,ζ^i,ζi,vi)′∥∞2≲P(Hk2∨n2/q)≲PHk2\max_{i\leqslant n}\|1\{A_{i}\}(\widehat{v}_{i},\widehat{\zeta}_{i},\zeta_{i},v_{i})^{\prime}\|_{\infty}^{2}\lesssim_{P}(H_{k}^{2}\vee n^{2/q})\lesssim_{P}H_{k}^{2}, we have

Step 6.(Controlling large terms) By definition of the event AiA_{i} we have

Additional Simulation Results

In this section, we present additional simulation results. All of the simulation results are based on the structural model

where p=dim⁡(xi)=200p=\dim(x_{i})=200, the covariates x∼N(0,Σ)x\sim N(0,\Sigma) with Σkj=(0.5)∣j−k∣\Sigma_{kj}=(0.5)^{|j-k|}, α0=.5\alpha_{0}=.5, and the sample size nn is set to 100100. In each design, we generate

with E[ζivi]=0\zeta_{i}v_{i}]=0. Inference results for all designs are based on conventional t-tests with standard errors calculated using the heteroscedasticity consistent jackknife variance estimator discussed in ?. We set λ\lambda according to the algorithm outlined in Appendix A with 1−γ=.951-\gamma=.95. We draw new xx’s, ζ\zeta’s and vv’s at every replication and draw new β0\beta_{0}’s and β1\beta_{1}’s at every replication in the random coefficient designs.

In the first thirteen designs, β1=β0\beta_{1}=\beta_{0}. We set the constants cyc_{y} and cdc_{d} to generate desired population values for the reduced form R2R^{2}’s, i.e. the R2R^{2}’s for equations (5.41) and (5.42). Let Ry2R^{2}_{y} be the desired R2R^{2} for the regression of yy on xx and Rd2R^{2}_{d} be the desired R2R^{2} from the regression of dd on xx. For each equation, we choose cyc_{y} and cdc_{d} to generate R2=0,.2,.4,.6,R^{2}=0,.2,.4,.6, and .8.8. In the heteroscedastic and binary designs discussed below, we choose cyc_{y} and cdc_{d} based on R2R^{2} as if (5.41) held with di=di∗d_{i}=d_{i}^{*} and viv_{i} and ζi\zeta_{i} were homoscedastic with variance equal to the average variance and label the results by R2R^{2} as in the other cases. In the homoscedastic cases, we set σy=σd=1\sigma_{y}=\sigma_{d}=1; and in the heteroscedastic cases, the average of σd(xi)\sigma_{d}(x_{i}) and the average of σy(di,xi)\sigma_{y}(d_{i},x_{i}) are both one. We set

Design 1. di=di∗d_{i}=d_{i}^{*}, β0=(1,1/2,1/3,1/4,1/5,0,0,0,0,0,1,1/2,1/3,1/4,1/5,0,...,0)′\beta_{0}=(1,1/2,1/3,1/4,1/5,0,0,0,0,0,1,1/2,1/3,1/4,1/5,0,...,0)^{\prime}, σy=σd=1\sigma_{y}=\sigma_{d}=1.

Design 2. di=di∗d_{i}=d_{i}^{*}, β0=(1,1/4,1/9,1/16,1/25,0,0,0,0,0,1,1/4,1/9,1/16,1/25,0,...,0)′\beta_{0}=(1,1/4,1/9,1/16,1/25,0,0,0,0,0,1,1/4,1/9,1/16,1/25,0,...,0)^{\prime}, σy=σd=1\sigma_{y}=\sigma_{d}=1.

Design 22. di=di∗d_{i}=d_{i}^{*}, β0,j=(1/j)2\beta_{0,j}=(1/j)^{2}, σy=σd=1\sigma_{y}=\sigma_{d}=1.

Design 3. di=di∗d_{i}=d_{i}^{*}, β0=(1,1/2,1/3,1/4,1/5,0,0,0,0,0,1,1/2,1/3,1/4,1/5,0,...,0)′\beta_{0}=(1,1/2,1/3,1/4,1/5,0,0,0,0,0,1,1/2,1/3,1/4,1/5,0,...,0)^{\prime}, σd=(1+xi′β0)21n∑i=1n(1+xi′β0)2\sigma_{d}=\sqrt{\frac{(1+x_{i}^{\prime}\beta_{0})^{2}}{\frac{1}{n}\sum_{i=1}^{n}(1+x_{i}^{\prime}\beta_{0})^{2}}}, σy=(1+α0di+xi′β0)21n∑i=1n(1+α0di+xi′β0)2\sigma_{y}=\sqrt{\frac{(1+\alpha_{0}d_{i}+x_{i}^{\prime}\beta_{0})^{2}}{\frac{1}{n}\sum_{i=1}^{n}(1+\alpha_{0}d_{i}+x_{i}^{\prime}\beta_{0})^{2}}}.

Design 4. di=di∗d_{i}=d_{i}^{*}, β0=(1,1/4,1/9,1/16,1/25,0,0,0,0,0,1,1/4,1/9,1/16,1/25,0,...,0)′\beta_{0}=(1,1/4,1/9,1/16,1/25,0,0,0,0,0,1,1/4,1/9,1/16,1/25,0,...,0)^{\prime}, σd=(1+xi′β0)21n∑i=1n(1+xi′β0)2\sigma_{d}=\sqrt{\frac{(1+x_{i}^{\prime}\beta_{0})^{2}}{\frac{1}{n}\sum_{i=1}^{n}(1+x_{i}^{\prime}\beta_{0})^{2}}}, σy=(1+α0di+xi′β0)21n∑i=1n(1+α0di+xi′β0)2\sigma_{y}=\sqrt{\frac{(1+\alpha_{0}d_{i}+x_{i}^{\prime}\beta_{0})^{2}}{\frac{1}{n}\sum_{i=1}^{n}(1+\alpha_{0}d_{i}+x_{i}^{\prime}\beta_{0})^{2}}}.

Design 44. di=di∗d_{i}=d_{i}^{*}, β0,j=(1/j)2\beta_{0,j}=(1/j)^{2}, σd=(1+xi′β0)21n∑i=1n(1+xi′β0)2\sigma_{d}=\sqrt{\frac{(1+x_{i}^{\prime}\beta_{0})^{2}}{\frac{1}{n}\sum_{i=1}^{n}(1+x_{i}^{\prime}\beta_{0})^{2}}}, σy=(1+α0di+xi′β0)21n∑i=1n(1+α0di+xi′β0)2\sigma_{y}=\sqrt{\frac{(1+\alpha_{0}d_{i}+x_{i}^{\prime}\beta_{0})^{2}}{\frac{1}{n}\sum_{i=1}^{n}(1+\alpha_{0}d_{i}+x_{i}^{\prime}\beta_{0})^{2}}}.

Design 5. di=1{di∗>0}d_{i}=\mathbf{1}\{d_{i}^{*}>0\}, β0=(1,1/2,1/3,1/4,1/5,0,0,0,0,0,1,1/2,1/3,1/4,1/5,0,...,0)′\beta_{0}=(1,1/2,1/3,1/4,1/5,0,0,0,0,0,1,1/2,1/3,1/4,1/5,0,...,0)^{\prime}, σy=σd=1\sigma_{y}=\sigma_{d}=1.

Design 6. di=di∗d_{i}=d_{i}^{*}, β0,j∼N(0,1)\beta_{0,j}\sim N(0,1), σy=σd=1\sigma_{y}=\sigma_{d}=1.

Design 7. di=di∗d_{i}=d_{i}^{*}, β~0=(1,1/2,1/3,1/4,1/5,0,0,0,0,0,1,1/2,1/3,1/4,1/5,0,...,0)′\widetilde{\beta}_{0}=(1,1/2,1/3,1/4,1/5,0,0,0,0,0,1,1/2,1/3,1/4,1/5,0,...,0)^{\prime}, β0,j∼N(0,β~0,j2)\beta_{0,j}\sim N(0,\widetilde{\beta}_{0,j}^{2}), σy=σd=1\sigma_{y}=\sigma_{d}=1.

Design 72. di=di∗d_{i}=d_{i}^{*}, β~0=(1,1/4,1/9,1/16,1/25,0,0,0,0,0,1,1/4,1/9,1/16,1/25,0,...,0)′\widetilde{\beta}_{0}=(1,1/4,1/9,1/16,1/25,0,0,0,0,0,1,1/4,1/9,1/16,1/25,0,...,0)^{\prime}, β0,j∼N(0,β~0,j2)\beta_{0,j}\sim N(0,\widetilde{\beta}_{0,j}^{2}), σy=σd=1\sigma_{y}=\sigma_{d}=1.

Design 722. di=di∗d_{i}=d_{i}^{*}, β~0,j=(1/j)2\widetilde{\beta}_{0,j}=(1/j)^{2}, β0,j∼N(0,β~0,j2)\beta_{0,j}\sim N(0,\widetilde{\beta}_{0,j}^{2}), σy=σd=1\sigma_{y}=\sigma_{d}=1.

Design 8. di=di∗d_{i}=d_{i}^{*}, β~0,j=ujz1,j+(1−uj)z2,j\widetilde{\beta}_{0,j}=u_{j}z_{1,j}+(1-u_{j})z_{2,j}, uj∼Bernoulli(.05)u_{j}\sim\textnormal{Bernoulli}(.05), z1,j∼N(0,25)z_{1,j}\sim N(0,25), z2,j∼N(0,.0025)z_{2,j}\sim N(0,.0025), σy=σd=1\sigma_{y}=\sigma_{d}=1

Design 1001. di=di∗d_{i}=d_{i}^{*}, β0,j=1{j∈{2,4,6,...,38,40}}\beta_{0,j}=\mathbf{1}\{j\in\{2,4,6,...,38,40\}\}, σy=σd=1\sigma_{y}=\sigma_{d}=1.

In the last thirteen designs, we set the constants cyc_{y} and cdc_{d} according to

for Rd2=0,.2,.4,.6,R^{2}_{d}=0,.2,.4,.6, and .8.8 and Ry2=0,.2,.4,.6,R^{2}_{y}=0,.2,.4,.6, and .8.8.

Design 1a. di=di∗d_{i}=d_{i}^{*}, β0=(1,1/2,1/3,1/4,1/5,0,0,0,0,0,1,1/2,1/3,1/4,1/5,0,...,0)′\beta_{0}=(1,1/2,1/3,1/4,1/5,0,0,0,0,0,1,1/2,1/3,1/4,1/5,0,...,0)^{\prime}, β1=(1,1/2,1/3,1/4,1/5,1/6,1/7,1/8,1/9,1/10,0,...,0)′\beta_{1}=(1,1/2,1/3,1/4,1/5,1/6,1/7,1/8,1/9,1/10,0,...,0)^{\prime}, σy=σd=1\sigma_{y}=\sigma_{d}=1.

Design 2a. di=di∗d_{i}=d_{i}^{*}, β0=(1,1/4,1/9,1/16,1/25,0,0,0,0,0,1,1/4,1/9,1/16,1/25,0,...,0)′\beta_{0}=(1,1/4,1/9,1/16,1/25,0,0,0,0,0,1,1/4,1/9,1/16,1/25,0,...,0)^{\prime}, β1=(1,1/4,1/9,1/16,1/25,1/36,1/49,1/64,1/81,1/100,0,...,0)′\beta_{1}=(1,1/4,1/9,1/16,1/25,1/36,1/49,1/64,1/81,1/100,0,...,0)^{\prime}, σy=σd=1\sigma_{y}=\sigma_{d}=1.

Design 22a. di=di∗d_{i}=d_{i}^{*}, β0,j=(1/j)2\beta_{0,j}=(1/j)^{2}, β1,j=(1/j)2\beta_{1,j}=(1/j)^{2}, σy=σd=1\sigma_{y}=\sigma_{d}=1.

Design 3a. di=di∗d_{i}=d_{i}^{*}, β0=(1,1/2,1/3,1/4,1/5,0,0,0,0,0,1,1/2,1/3,1/4,1/5,0,...,0)′\beta_{0}=(1,1/2,1/3,1/4,1/5,0,0,0,0,0,1,1/2,1/3,1/4,1/5,0,...,0)^{\prime}, β1=(1,1/2,1/3,1/4,1/5,1/6,1/7,1/8,1/9,1/10,0,...,0)′\beta_{1}=(1,1/2,1/3,1/4,1/5,1/6,1/7,1/8,1/9,1/10,0,...,0)^{\prime}, σd=(1+xi′β1)21n∑i=1n(1+xi′β1)2\sigma_{d}=\sqrt{\frac{(1+x_{i}^{\prime}\beta_{1})^{2}}{\frac{1}{n}\sum_{i=1}^{n}(1+x_{i}^{\prime}\beta_{1})^{2}}}, σy=(1+α0di+xi′β0)21n∑i=1n(1+α0di+xi′β0)2\sigma_{y}=\sqrt{\frac{(1+\alpha_{0}d_{i}+x_{i}^{\prime}\beta_{0})^{2}}{\frac{1}{n}\sum_{i=1}^{n}(1+\alpha_{0}d_{i}+x_{i}^{\prime}\beta_{0})^{2}}}.

Design 4a. di=di∗d_{i}=d_{i}^{*}, β0=(1,1/4,1/9,1/16,1/25,0,0,0,0,0,1,1/4,1/9,1/16,1/25,0,...,0)′\beta_{0}=(1,1/4,1/9,1/16,1/25,0,0,0,0,0,1,1/4,1/9,1/16,1/25,0,...,0)^{\prime}, β1=(1,1/4,1/9,1/16,1/25,1/36,1/49,1/64,1/81,1/100,0,...,0)′\beta_{1}=(1,1/4,1/9,1/16,1/25,1/36,1/49,1/64,1/81,1/100,0,...,0)^{\prime}, σd=(1+xi′β1)21n∑i=1n(1+xi′β1)2\sigma_{d}=\sqrt{\frac{(1+x_{i}^{\prime}\beta_{1})^{2}}{\frac{1}{n}\sum_{i=1}^{n}(1+x_{i}^{\prime}\beta_{1})^{2}}}, σy=(1+α0di+xi′β0)21n∑i=1n(1+α0di+xi′β0)2\sigma_{y}=\sqrt{\frac{(1+\alpha_{0}d_{i}+x_{i}^{\prime}\beta_{0})^{2}}{\frac{1}{n}\sum_{i=1}^{n}(1+\alpha_{0}d_{i}+x_{i}^{\prime}\beta_{0})^{2}}}.

Design 44a. di=di∗d_{i}=d_{i}^{*}, β0,j=(1/j)2\beta_{0,j}=(1/j)^{2}, β1,j=(1/j)2\beta_{1,j}=(1/j)^{2}, σd=(1+xi′β1)21n∑i=1n(1+xi′β1)2\sigma_{d}=\sqrt{\frac{(1+x_{i}^{\prime}\beta_{1})^{2}}{\frac{1}{n}\sum_{i=1}^{n}(1+x_{i}^{\prime}\beta_{1})^{2}}}, σy=(1+α0di+xi′β0)21n∑i=1n(1+α0di+xi′β0)2\sigma_{y}=\sqrt{\frac{(1+\alpha_{0}d_{i}+x_{i}^{\prime}\beta_{0})^{2}}{\frac{1}{n}\sum_{i=1}^{n}(1+\alpha_{0}d_{i}+x_{i}^{\prime}\beta_{0})^{2}}}.

Design 5a. di=1{di∗>0}d_{i}=\mathbf{1}\{d_{i}^{*}>0\}, β0=(1,1/2,1/3,1/4,1/5,0,0,0,0,0,1,1/2,1/3,1/4,1/5,0,...,0)′\beta_{0}=(1,1/2,1/3,1/4,1/5,0,0,0,0,0,1,1/2,1/3,1/4,1/5,0,...,0)^{\prime}, β1=(1,1/2,1/3,1/4,1/5,1/6,1/7,1/8,1/9,1/10,0,...,0)′\beta_{1}=(1,1/2,1/3,1/4,1/5,1/6,1/7,1/8,1/9,1/10,0,...,0)^{\prime}, σy=σd=1\sigma_{y}=\sigma_{d}=1.

Design 6a. di=di∗d_{i}=d_{i}^{*}, β0,j∼N(0,1)\beta_{0,j}\sim N(0,1), β1,j∼N(0,1)\beta_{1,j}\sim N(0,1), E[β0,jβ1,j]=.8[\beta_{0,j}\beta_{1,j}]=.8, σy=σd=1\sigma_{y}=\sigma_{d}=1.

Design 7a. di=di∗d_{i}=d_{i}^{*}, β~0=(1,1/2,1/3,1/4,1/5,0,0,0,0,0,1,1/2,1/3,1/4,1/5,0,...,0)′\widetilde{\beta}_{0}=(1,1/2,1/3,1/4,1/5,0,0,0,0,0,1,1/2,1/3,1/4,1/5,0,...,0)^{\prime}, β~1=(1,1/2,1/3,1/4,1/5,1/6,1/7,1/8,1/9,1/10,0,...,0)′\widetilde{\beta}_{1}=(1,1/2,1/3,1/4,1/5,1/6,1/7,1/8,1/9,1/10,0,...,0)^{\prime}, β0,j=β~0,jz0,j\beta_{0,j}=\widetilde{\beta}_{0,j}z_{0,j}, β1,j=β~1,jz1,j\beta_{1,j}=\widetilde{\beta}_{1,j}z_{1,j}, z0,j∼N(0,1)z_{0,j}\sim N(0,1), z1,j∼N(0,1)z_{1,j}\sim N(0,1), E[z0,jz1,j]=.8[z_{0,j}z_{1,j}]=.8, σy=σd=1\sigma_{y}=\sigma_{d}=1.

Design 72a. di=di∗d_{i}=d_{i}^{*}, β~0=(1,1/4,1/9,1/16,1/25,0,0,0,0,0,1,1/4,1/9,1/16,1/25,0,...,0)′\widetilde{\beta}_{0}=(1,1/4,1/9,1/16,1/25,0,0,0,0,0,1,1/4,1/9,1/16,1/25,0,...,0)^{\prime}, β~1=(1,1/4,1/9,1/16,1/25,1/36,1/49,1/64,1/81,1/100,0,...,0)′\widetilde{\beta}_{1}=(1,1/4,1/9,1/16,1/25,1/36,1/49,1/64,1/81,1/100,0,...,0)^{\prime}, β0,j=β~0,jz0,j\beta_{0,j}=\widetilde{\beta}_{0,j}z_{0,j}, β1,j=β~1,jz1,j\beta_{1,j}=\widetilde{\beta}_{1,j}z_{1,j}, z0,j∼N(0,1)z_{0,j}\sim N(0,1), z1,j∼N(0,1)z_{1,j}\sim N(0,1), E[z0,jz1,j]=.8[z_{0,j}z_{1,j}]=.8, σy=σd=1\sigma_{y}=\sigma_{d}=1.

Design 722a. di=di∗d_{i}=d_{i}^{*}, β~0,j=(1/j)2\widetilde{\beta}_{0,j}=(1/j)^{2}, β~1,j=(1/j)2\widetilde{\beta}_{1,j}=(1/j)^{2}, β0,j=β~0,jz0,j\beta_{0,j}=\widetilde{\beta}_{0,j}z_{0,j}, β1,j=β~1,jz1,j\beta_{1,j}=\widetilde{\beta}_{1,j}z_{1,j}, z0,j∼N(0,1)z_{0,j}\sim N(0,1), z1,j∼N(0,1)z_{1,j}\sim N(0,1), E[z0,jz1,j]=.8[z_{0,j}z_{1,j}]=.8, σy=σd=1\sigma_{y}=\sigma_{d}=1.

Design 8a. di=di∗d_{i}=d_{i}^{*}, β~0,j=5ujz11,j+.05(1−uj)z12,j\widetilde{\beta}_{0,j}=5u_{j}z_{11,j}+.05(1-u_{j})z_{12,j}, β~1,j=5ujz21,j+.05(1−uj)z22,j\widetilde{\beta}_{1,j}=5u_{j}z_{21,j}+.05(1-u_{j})z_{22,j}, uj∼Bernoulli(.05)u_{j}\sim\textnormal{Bernoulli}(.05), z11,j∼N(0,1)z_{11,j}\sim N(0,1), z12,j∼N(0,1)z_{12,j}\sim N(0,1), z21,j∼N(0,1)z_{21,j}\sim N(0,1), z22,j∼N(0,1)z_{22,j}\sim N(0,1), σy=σd=1\sigma_{y}=\sigma_{d}=1

Design 1001a. di=di∗d_{i}=d_{i}^{*}, β0,j=1{j∈{2,4,6,...,38,40}}\beta_{0,j}=\mathbf{1}\{j\in\{2,4,6,...,38,40\}\}, β1,j=1{j∈{1,3,5,...,37,39}}\beta_{1,j}=\mathbf{1}\{j\in\{1,3,5,...,37,39\}\}, σy=σd=1\sigma_{y}=\sigma_{d}=1.

Results are summarized in figures and tables below. In the tables, we report results for the four estimators considered in the main text (Oracle, Double-Selection Oracle, Post-Lasso, and Double-Selection). We also report results for regular Lasso (Lasso), the union of the Double-Selection interval with the Post-Lasso interval (Double-Selection Union ADS), using the union of the set of variables selected by Double-Selection and the set of variables selected by running Lasso of yy on dd and xx without penalizing dd (Double-Selection + I3), and the split-sample procedure discussed in the text (Split-Sample). For Double-Selection Union ADS, the point estimate is taken as the midpoint of the union of the intervals.

References