To build intuition, we discuss the case where there is only one control; that is, p=1. This scenario provides the simplest possible setting where variable selection might be interesting. In this case, Lasso-type methods act like conservative t-tests which allows the properties of selection methods to be explained easily.
For simplicity, all errors and controls are taken as normal,
The standard post-single-selection method for inference proceeds by applying model selection methods – ranging from standard t-tests to Lasso-type selectors – to the first equation only, followed by applying OLS to the selected model. In the model selection stage, standard selection methods would necessarily omit xi wp →1 if
where σn∗2=σζ2(σd2)−1. The variance σn∗2 is smaller than the variance σn2 from estimation with xi included if βm=0. The potential reduction in variance is often used as a “motivation” for the standard selection procedure. The estimator is super-efficient, achieving a variance smaller than the semi-parametric efficiency bound under homoscedasticity. That is, the estimator is “too good.”
That is, the standard post-selection estimator is not asymptotically normal and even fails to be uniformly consistent at the rate of n. This poor behavior occurs because the omitted variable bias created by dropping xi may be large even when the magnitude of the regression coefficient, ∣βm∣, in the confounding equation (1.2) is small but is not exactly zero. To see this, note
The term i∗ has standard behavior; namely i∗⇝N(0,1). The term ii generates the omitted variable bias, and it may be arbitrarily large, since wp →1,
In contrast to the standard approach, our post-double-selection method for inference proceeds by applying model selection methods, such as standard t-tests or Lasso-type selectors, to both equations and taking the selected controls as the union of controls selected from each equation. This selection is than followed by applying OLS to the selected controls. Thus, our approach drops xi only if the omitted variable bias term ii is small. To see this, note that the double-selection-methods include xi wp →1 if its coefficient in either (1.1) or (1.2) is not very small. Mathematically, xi is included if
Given the discussion in the preceding paragraph, it is immediate that the post-double selection estimator satisfies
i.e. coefficients in front of xi in both equations are small. In this case,
Once again, the term ii is due to omitted variable bias, and it obeys wp →1 under (1.13)
since (σv/σd)2=1−ρ2. Moreover, we can show i∗−i=oP(1) under such sequences, so the first order asymptotics of αˇ is the same whether xi is included or excluded.
Extensions: Other Problems and Heterogeneous Treatment Effects
In order to discuss extensions in a very simple manner, we assume i.i.d sampling as well as assume away approximation errors, namely g(zi)=xi′βg0 and m(zi)=xi′βm0, where parameters βg0 and βm0 are high-dimensional and that xi=P(zi) as before. In this paper we considered a moment condition:
where φ(u)=u and vi are measurable functions of zi, and the target parameter is α0. We selected the instrument vi such that the equation is first-order insensitive to the parameter β at β=βg0:
Note that φ(u)=u and vi=di−m(zi) implement this condition. If (2.15) holds, the estimator of α0 gets “immunized” against nonregular estimation of β0, for example, via a post-selection procedure or other regularized estimators. Such immunization ideas are in fact behind the classical Frisch-Waugh-Robinson partialling out technique in the linear setting and the ?’s C(α) test in the nonlinear setting. Our contribution here is to recognize the importance of this immunization in the context of post-selection inference, To the best of our knowledge, all prior theoretical and empirical work uses the standard post-selection approach based on the outcome equation alone, which is highly non-robust way of conducting inference, as shown in extensive monte-carlo, in Section 2.4, and in a sequence of fundamental critiques by Leeb and Potscher, see ? . to develop a robust post-selection approach to inference on the target parameter, and characterize the uniformity regions of this procedure. Our approach uses modern selection methods to estimate g and the function m defining the intstrument v. In an ongoing work, we explore other regularization methods, such as the ridge method or combination of ridge method with Lasso methods, and characterize uniformity regions of the resulting procedures. Also, generalizations to nonlinear models, where φ is non-linear and can correspond to a likelihood score or quantile check function are given in ? and ?; in these generalizations achieving (2.15) is also critical.
Within the context of this paper, a potentially important extension is to consider a general treatment effect model, where di is interacted with transformations of zi. As long as the interest lies in a particular regression coefficient, the current framework covers this implicitly since xi could contain interactions of di with transformations of controls zi. In the case a fixed number of such regression coefficients is of interest, we can estimate each of the coefficients by re-labeling the corresponding regressor as di and other regressors as xi and then applying our procedure the fixed number of times. Such component-wise procedure is valid as long as our regularity conditions hold for each of the resulting regression models in this manner.
A related research direction being pursued is the study of estimation of average treatment effects when treatment effects are fully heterogeneous. When the treatment variable di∈{0,1} is binary (or discrete more generally) our approach is readily amenable to this problem. In this case the parameter of interest is the average treatment effect
where φ(α,y,d,g0,g1,m)=α−md(y−g1)+1−m(1−d)(y−g0)−(g1−g0). It is straightforward to check that for each j∈{1,2,3}:
where the latter is the semiparametric efficiency bound of ?. Formal result is proven and stated below, and other formal results along these lines are given in an ongoing work that studies this as well as other types of effects.
2. Theoretical Results on ATE with Heterogeneity
The confounding factors zi affect the policy variable via the propensity score m(zi) and the outcome variable via the function g(di,zi). Both of these functions are unknown and potentially complicated. As in the main text, we use linear combinations of control terms xi=P(zi) to approximate g(zi) and m(zi), writing (2.19) and (2.20) as
where rgi and rmi are the approximation errors, and
where xi′βg0,1, xi′βg0,0, and xi′βm0 are approximations to g(1,zi), g(0,zi), and m(zi), and Λ(u)=u for the case of linear link and Λ(u)=eu/(1+eu) for the case of the logistic link. In order to allow for a flexible specification and incorporation of pertinent confounding factors, the vector of controls, xi=P(zi), we can have a dimension p=pn which can be large relative to the sample size.
The efficient moment condition, derived by ?, for parameter α0 is as follows:
The post-double-selection estimator αˇ that solves
where g(di,zi) and m(zi) are post-Lasso estimators of functions g and m based upon equations (2.21)-(2.22). In case of the logistic link Λ, Lasso for logistic regression is as defined in ? and ?, and the associated post-Lasso estimators are as those defined in ?.
These conditions are simple high-level conditions, which encode both the approximate sparsity of the models as well as impose some reasonable behavior on the post-selection estimators of m and g (or other sparse estimators). These conditions are implied by other more primitive conditions in the literature. Sufficient conditions for the equivalence between population and empirical sparse eigenvalues are given in ? and ?. The boundedness conditions are made to simplify arguments, and they could be dealt away with more complicated proofs, under more stringent side conditions.
The next target parameter is the average treatment effect on the treated:
The efficient moment condition, derived by ?, for parameter γ0 is as follows:
In this case the post-double-selection estimator γˇ that solves
3. Proof of Theorems 1 and 2.
The two results have identical structure and have nearly the same proof, and so we present the proof of the Proof of Theorems 1 only.
Step 1. In this step we establish claim (1).
(a) We begin with a preliminary observation. Define, for t=(t1,t2,t3),
where L depends only on c′ and C, ∣k∣=∑j=13kj, and
We observe that with probability no less than 1−Δn,
for β=βg, with evaluation after computing the norms, and noting that for any β
under condition (iii). Furthermore, for n⩾n0=min{j:δj⩽1/2}:
for β=βm0, with evaluation after computing the norms.
Hence with probability at least 1−Δn,
with h evaluated at h=h. By Liapunov central limit theorem,
(d) Note that for Δi=h(zi)−h0(zi),
(with h evaluated at h=h). By the law of iterated expectations and because
Moreover, uniformly for any h∈Hn we have that
Since h∈Hn with probability 1−Δn, we have that once n⩾n0,
The class of functions Gd for d∈{0,1} is a union of at most (Csp) VC-subgraph classes of functions with VC indices bounded by C′s. The class of functions M is a union of at most (Csp) VC-subgraph classes of functions with VC indices bounded by C′s (monotone transformation Λ preserve the VC-subgraph property). These classes are uniformly bounded and their entropies therefore satisfy
Finally, the class Fn={fh−fh0:h∈Hn} is a Lipschitz transform of Hn with bounded Lipschitz coefficients and with a constant envelope. Therefore, we have that
We shall invoke the following lemma derived in ?.
Let F be a measurable function class on a sample space. Let F=supf∈F∣f∣, and suppose that there exist some constants ωn>3 and υ>1, such hat
Then for every δ∈(0,1/6) we have
with probability at least 1−δ for some constant that Cυ.
Then by Lemma 1 together and some simple calculations, we have that
Step 3. Claim (3) is immediate from claims (2) and (3) by the way of contradiction. ∎
Deferred Proofs: Proof of Lemma 1
We establish the result for Lasso (the proof for other feasible Lasso estimators is similar).
By Condition SE, with probability 1−o(1) for n sufficiently large we have κcˉ>κ′/2∥Ψ∥∞ so that with the same probability
Moreover, by condition RF we have with probability 1−o(1) that
Finally, since λ≳nlog(p∨n) we have
since cs≲Ps/n by condition ASM and Chebyshev inequality.
To show the second statement in (i), note that
where β is the Lasso estimator. Again by Lemma 7 in ? we have that the assumptions of Lemma 6 in ? hold with probability 1−o(1). Using Condition SE to bound κcˉ from below and Condition RF to bound ∥Ψ∥∞ from above with probability 1−o(1) as before, and λ≲σnlog(p∨n), it follows from Lemma 6 in ? that with probability 1−o(1) that
The results regarding Post-Lasso in (ii) follow similarly by invoking Lemma 8 in ?.
Split-Sample Estimation and Inference
In this section we discuss a variant of the double selection estimator based on sample splitting. The motivation for the split-sample estimator is that its use allows us to relax the requirement s2log2(p∨n)=o(n) that is assumed in the full-sample counterpart to the milder condition
For each subsample k=a,b, the model Ik is selected based on the subsample k independently from the subsample kc. In what follows the model Ik is used to fit the subsample kc. A constructive way to obtain Ia and Ib is to apply the double selection method for each subsample to select the sets of controls Ia:=I1a∪I2a∪I3a and Ib:=I2b∪I2b∪I3b.
Then we form estimates in the two subsamples
For an index i in the subsample k, we define the residuals
Finally, we combine the estimates into the split-sample estimator based on Ia and Ib is defined as
where Υk=Dk′MIkcDk/nk.
We state below sufficient conditions for the analysis of the split-sample method.
The Conditions ASTESS(i)-(iii) agree with the corresponding conditions in ASTE. The remaining conditions ASTESS(iv)-(v) are implied by Condition ASTE. We note that Condition ASTESS(vi) is needed only for obtaining consistent estimates of the asymptotic variance. Such conditions are mild since they do not require uniform estimation of the functions g and m.
The next result establishes that the split-sample estimator αˇab has similar large sample properties to the full-sample double-selection estimator under weaker growth condition.
Step 0.(Combining) In this step we combine both subsample estimators. Letting Υk=Dk′MIkcDk/nk, for k=a,b, so that we have
which follows similarly to the proofs given in Step 5.
where zi,n=σn−1viζi/n are i.n.i.d. with mean zero. We have that for some small enough δ>0
This condition verifies the Lyapunov condition and thus implies that Zn→dN(0,1).
Step 1.(Main) For the subsample k=a,b write αˇk=[Dk′MIkcDk/nk]−1[Dk′MIkcYk/nk] so that
First, note that by Condition ASTESS we have
Second, by the split sample construction, we have that Ikc is independent from ζk, and by assumption of the model mk is also independent of ζk. Thus by Chebyshev inequality
where the last relation follows by ASTESS.
Third, using similar independence arguments, by Chebyshev and Condition ASTESS, conclude
Fourth, using that skc≲Ps by ASTESS so that ϕmin−1(skc)≲P1 by condition SE, we have that
Step 3.(Behavior of iik.) Since iik=(mk+Vk)′MIkc(mk+Vk)/nk, decompose
Then ∣iik,a∣=oP(1) by Condition ASTESS, ∣iik,b∣=oP(1) by reasoning similar to deriving the bound for ∣ik,b∣, and ∣iik,c∣=oP(1) by reasoning similar to deriving the bound for ∣ik,d∣.
By condition ASTESS ∥MIkcgk∥=oP(n1/4) and by condition SM(ii) we have ∥PIkcDk/nk∥⩽∥Dk/nk∥≲P1, and by Step 1 we have ∣αˇk−α0∣≲Pn−1/2. Moreover,
We have ϕmax,k(skc)/ϕmin,k(skc)≲P1 by condition SE, and ∥Xk[Ikc]′ζk/nk∥≲Pskc by condition SM(ii), the independence between the selected components Ikc and ζk since they are based on different subsamples, and applying Chebyshev inequality.
Similarly, we have ∥mk−Xkβk∥/nk≲Po(n−1/4)+skc/nk.
Step 5.(Variance Estimation.) Since sk≲Ps=o(n), (nk−sk−1)/nk=oP(1), so we can use n as the denominator. Recall the definitions ζio=yi−diαˇk−xi′βˇk, vi=di−xi′βk and ζi=ζio1{∣ζio∣∨∣vi∣⩽Hk} if i belongs to subsample k where Hk=Cn/[(skc∨n1/2)logn]. For notational convenience let Ai={∣ζio∣∨∣vi∣⩽Hk}. Since q>4, skc≲Ps, and n2/qslog(n∨p)=o(n), we have n1/q=oP(Hk). Hence consider
Finally, since maxi⩽n∥1{Ai}(vi,ζi,ζi,vi)′∥∞2≲P(Hk2∨n2/q)≲PHk2, we have
Step 6.(Controlling large terms) By definition of the event Ai we have
Additional Simulation Results
In this section, we present additional simulation results. All of the simulation results are based on the structural model
where p=dim(xi)=200, the covariates x∼N(0,Σ) with Σkj=(0.5)∣j−k∣, α0=.5, and the sample size n is set to 100. In each design, we generate
with E[ζivi]=0. Inference results for all designs are based on conventional t-tests with standard errors calculated using the heteroscedasticity consistent jackknife variance estimator discussed in ?. We set λ according to the algorithm outlined in Appendix A with 1−γ=.95. We draw new x’s, ζ’s and v’s at every replication and draw new β0’s and β1’s at every replication in the random coefficient designs.
In the first thirteen designs, β1=β0. We set the constants cy and cd to generate desired population values for the reduced form R2’s, i.e. the R2’s for equations (5.41) and (5.42). Let Ry2 be the desired R2 for the regression of y on x and Rd2 be the desired R2 from the regression of d on x. For each equation, we choose cy and cd to generate R2=0,.2,.4,.6, and .8. In the heteroscedastic and binary designs discussed below, we choose cy and cd based on R2 as if (5.41) held with di=di∗ and vi and ζi were homoscedastic with variance equal to the average variance and label the results by R2 as in the other cases. In the homoscedastic cases, we set σy=σd=1; and in the heteroscedastic cases, the average of σd(xi) and the average of σy(di,xi) are both one. We set
Results are summarized in figures and tables below. In the tables, we report results for the four estimators considered in the main text (Oracle, Double-Selection Oracle, Post-Lasso, and Double-Selection). We also report results for regular Lasso (Lasso), the union of the Double-Selection interval with the Post-Lasso interval (Double-Selection Union ADS), using the union of the set of variables selected by Double-Selection and the set of variables selected by running Lasso of y on d and x without penalizing d (Double-Selection + I3), and the split-sample procedure discussed in the text (Split-Sample). For Double-Selection Union ADS, the point estimate is taken as the midpoint of the union of the intervals.