Selective inference with a randomized response
Xiaoying Tian, Jonathan E. Taylor
Introduction
Tukey (1980) promoted the use of exploratory data analysis to examine the data and possibly formulate hypotheses for further investigation. Nowadays, many statistical learning methods allow us to perform these exploratory data analyses, based on which we can posit a model on the data generating distribution. Since this model is not given a priori, classical statistical inference will not provide valid tests that control the Type-I errors.
Selective inference seeks to address this problem, see Lee et al. (2013a); Lockhart et al. (2014); Lee & Taylor (2014); Fithian et al. (2014). Loosely speaking, there are two stages in selective inference. The first is the selection stage that explores the data and formulates a plausible model for the data distribution. Then we enter the inference stage that seeks to provide valid inference under the selected model which is proposed after inspecting the data. Inference under different models have been studied, notably the Gaussian families Lee et al. (2013a); Tian et al. (2015); Lee & Taylor (2014) as well as other exponential families Fithian et al. (2014).
In this work, we consider selective inference in a general setting that include nonparametric settings. In addition, we introduced the use of randomized response in model selection. A most common example of randomized model selection is probably the practice of data splitting. Assuming independent sampling, we can divide the data into two subsets, using the first for model selection and the second subset for inference. Though not emphasized, this split is often random. Hence, data splitting can be thought of as a special case of randomized model selection. To motivate the use of randomized selection and introduce the inference problem that ensues, we consider the following example.
Publication bias, (also called the “file drawer effect” by Rosenthal (1979)) is a bias introduced to scientific literature by failure to report negative or non-confirmatory results. We formulate the problem in the simple example below.
Suppose that we are interested in discovering positive effects and would only report the sample mean if it survives the file drawer effect, i.e.
Then what is the “correct” -value to report for an observation that exceeds the threshold?
where is the CDF of an random variable. Therefore, we get a pivotal quantity
for the ’s surviving the file drawer effect (1).
Randomized selection circumvents this problem. In the following, we propose a randomized version of the “file drawer problem”.
We assume the same setup of a triangular array of observations as in Example 1. But instead of reporting when it survives the file drawer effect (1), we independently draw , and only report if
To compute the exact form of , we have to compute the convolution of and which has explicit forms for many distributions . Moreover, when is Logistic or Laplace distribution, we have
which leads to three scenarios for selection.
In this case, the dominant term for selection is , and since we have a big positive effect, we would always report the sample mean when is big. This corresponds to the selection event having probability tending to and the selective likelihood ratio goes to as well. In this case, there is very little selection bias, and the original law is a good approximation to the selective distribution for valid inference.
, for some .
In this case, the dominant term is also , but in the negative direction. As , the selection probability vanishes and the selective likelihood becomes degenerate. We almost never report the sample mean in this scenario, but in the rare event where we do, by no means can we use the original distribution for inference.
, for some .
This corresponds to local alternatives. In this case, the selective likelihood neither converges to or becomes degenerate. Rather, it becomes an indicator function of a half interval. Proper adjustment is needed for valid inference in this case.
It is in the second scenario that pivotal quantity (2) will not converge to . Different distributions will have different behaviors in the tail. Since the conditioning event becomes a large-deviations event, we cannot expect it to behave like the normal distribution in the tail.
where is the survival function of . When for some , and is the Laplace or Logistic distribution so that has an exponential tail, the dominant term in both the numerator and the denominator will cancel out, making the selective likelihood ratio properly behaved in this difficult scenario.
It turns out that this selective likelihood ratio is fundamental to formalizing asymptotic properties of selective inference procedures. Its behavior determines not only the asymptotic convergence of the pivotal quantities like in (4), but also whether consistent estimation of the population parameters is possible with large samples.
Again in the negative mean scenario where , the sample mean surviving the non-randomized “file drawer effect” cannot be a consistent estimator for the underlying means because it will always be positive. But if is reported as in Example 2, it will be consistent for even if is negative and bounded away from . For detailed discussion, see Section 3.
In general, the behavior of the selective likelihood ratio can be used to study the asymptotic properties of selective inference procedures. We study consistent estimation and weak convergence for selective inference procedures in Section 3 and Section 5 respectively.
We are especially inspired by the field of differential privacy (c.f. Dwork et al. (2014) and references therein) to study the use of randomization in selective inference. Privatized algorithms purposely randomize reports from queries to a database in order to allow valid interactive data analysis. To our understanding, our results are the first results related to weak convergence in privatized algorithms, as most guarantees provided in the differentially private literature are consistency guarantees. Some other asymptotic results in selective inference have also been considered in Tibshirani et al. (2015); Tian & Taylor (2015), though these have a slightly different flavor in that they marginalize over choices of models.
We conclude this section with some more examples.
2 Linear regression
Randomized selection in this setting is a natural extension of these works. Fithian et al. (2014) proposed to use a subset of data for model selection, which yields a significant increase in power. In this work, we study general randomized selection procedures. Consider the following example.
Due to the sparsity of the solution of LASSO Tibshirani (1996)
3 Nonparametric selective inference
All the previous works on selective inference assume a parametric model like the Gaussian family or the exponential family. In this work, we allow selective inference in a non-parametric setting. Consider the following examples.
Suppose in a classification problem, we observe independent samples,
4 Outline of the paper
There are three main advantages of applying randomization for selective inference,
Consistent estimation under the selective distribution
Weak convergence of selective inference procedures
In the following sections, Section 2 gives the setup of selective inference and introduced selective likelihood ratio, which is the key for studying consistent estimation and weak convergence of selective inference procedures. Section 4 focuses on linear regression models with different randomization schemes, demonstrating the increase in power. Section 5 proposes an asymptotic test for the nonparametric settings. Theorem 9 proves that the central limit theorem holds under the selective distribution with mild conditions. Applications to the two examples in Section 1.3 are discussed. This is a result for fixed dimension . Finally, Section 6 discusses the possibility of extending our work to the setting, when multiple selection procedures are performed on different randomizations of the original data. One application is selective inference after cross validation for the square-root LASSO Belloni et al. (2011).
Selective Likelihood Ratio
where is loosely defined as being made up of “potentially interesting statistical questions”.
Since we use the data to choose the model , it is only fair to consider the conditional distribution for inference,
Therefore, we seek to control the selective Type-I error:
where is the selected family of distributions in the range of and is the null hypothesis. Selective intervals for parametric models can then be constructed by inverting such selective hypothesis tests, though only the one-parameter case has really been considered to date.
Similar to the selective inference we defined above, we seek to control the selective Type-I error,
Moreover, we also want to achieve good estimation, which makes
2 Selective likelihood ratio
with the same sufficient statistic and natural parameters .
Furthermore, to test , we consider the following law,
The first claim of the lemma is quite straight-forward using the relationship in (13). The second claim is a Lehmann–Scheffe (c.f. Chapter 4.4 in Lehmann (1986)) construction which was proposed in Fithian et al. (2014), to construct tests for one of the natural parameters treating the others as nuisance parameters. For detailed construction of such tests in the linear regression setting, see Section 4.
Consistent Estimation After Model Selection
In this section, we leave the parametric setup and consider general models . In particular, we study the consistency of estimators under the selective distribution for arbitrary models. We first introduce the framework of asymptotic analysis under the selective model. Then we state conditions for consistent estimation in Lemma 3 and conclude with examples.
For any model , which is a collection of distributions, we define its corresponding selective model, which is the collection of corresponding selective distributions,
In order to make meaningful asymptotic statements, we consider a sequence of randomized selection procedures and models with each in the range of .
The following lemma states the conditions for consistency of under the sequence of corresponding selective models ,
Consider a sequence of randomized selection procedures and models. Suppose the selective likelihood ratios satisfies, for some ,
Further, if is uniformly consistent for in probability, then is uniformly consistent for in probability under the sequence .
We illustrate the application of Lemma 3 through our “file drawer effect” examples in Section 1.1.
where we independently draw .
By law of large numbers, we easily see that if we always report , it will be an unbiased estimator for . However, since we only observe the sample means surviving the file drawer effect. Will still be consistent for ?
In the most difficult scenario discussed in Section 1.1, where for some , cannot be a consistent estimator for in Example 1. This is easy to see as Example 1 will only report positive sample means. A remarkable feature of randomized selection is that consistent estimation of the population parameters is possible even when the selection event has vanishing probabilities. In fact, the following lemma states that when is a Logistic distribution, is consistent for after the randomized file drawer effect in Example 2.
Before we prove the lemma, we want to point out that although the selection procedure in Example 2 is different from that in Example 1 because of randomization, is still the dominant term in selection. Note that
Since both and are random variables, the dominant term , would ensure that the selection event has vanishing probabilities in Example 2 as well. Thus it is particularly impressive that Example 2 gives consistent estimation where Example 1 cannot. The proof of Lemma 4 is deferred to the appendix.
We also verified this theory of consistent estimation through simulations. Figure 1 shows the empirical distributions of the sample mean after the file drawer effect in Example 1 or the “randomized” file drawer effect in Example 2. They are marked with “blue” colors or “red” colors respectively. We set the true underlying mean to be and mark it with the dotted vertical line in Figure 1. The upper panel Figure 1(a) is simulated with and the lower panel Figure 1(b) is simulated with . We notice that in both simulations, the sample mean in Example 1 concentrates around the thresholding boundary, which is positive. Thus, these sample means can not be possibly for the underlying mean . However, the existence of randomization allows us to report negative sample means. As a result, the sample mean in Example 2 will be consistent for . We see that as we increase sample size , the sample means concentrates closer to .
Inference in linear regression models
with known or unknown or the saturated model,
with known variance. Now we consider some randomized selection procedures and inference after selection.
In the introduction, we introduced data splitting Cox (1975) as a special case of randomized selective inference. In Fithian et al. (2014), the term data carving was introduced to demonstrate that data splitting is inadmissible. In data splitting (and data carving) inference makes most sense in the selected model , hence we should think of as returning a subset of variables selected.
However, there are two disadvantages with this randomization scheme. First, it is computationally difficult to aggregate over all random splits. Second, it seems difficult to consider the saturated model for inference, which is more robust to model misspecifications. To overcome those difficulties, we introduce other randomization schemes below.
2 Additive noise and more powerful tests
One major advantage of using a randomized response for selective inference is that these procedures yield much more powerful tests, at a small cost of on the quality of the selected models. In other words, small amount of randomization is cause a small loss in the model selection stage, but we gain much more power in the inference stage.
and is the non-selective Fisher information for in or . The parameters depend on which of the two models we are considering.
In the saturated model , the score statistic is Since is measurable with respect to ,
Since and are both normal distributions with covariance matrices,
In the selected model , the score statistic is Similarly,
When there is no randomization , we potentially have no leftover Fisher information. This corresponds to a very rare selection event. However after randomization, even with very extreme selection, there is always leftover Fisher information, which makes the selective tests more powerful. Consider the following examples.
Moreover, the increase in leftover Fisher information with randomization is not specific to Gaussian randomizations. For example, in Figure 1 when we use Logistic randomization, we also observe that under the selective distribution with randomization, has a much bigger variance than without randomization. As discussed above, this variance multiplied by is exactly the leftover Fisher information, which explains why selective procedures after randomization will have better performances than without.
We investigate the relationship between the leftover Fisher information and the length of confidence intervals constructed by inverting the pivot in (4). Specifically, in Example 2, after observing a reported sample mean, we want to report confidence intervals for the underlying mean .
Figure 2 demonstrates the selective intervals (solid lines) after (3) with being either Gaussian or Logistic noises. The sample size . Unlike the nominal confidence intervals (dashed lines), the selective intervals are valid with coverage for the underlying mean. Since Lemma 3 gives a lower bound of , we would intuitively expect the selective confidence intervals to be the length of the nominal intervals. This is verified in Figure 2(a), when we observe really negative sample means. (The sample means can be negative because we added randomization.) On the other hand, for Logistic randomization in Figure 2(b), the intervals are slightly wider than the nominal intervals around the , but narrow to roughly the nominal size on both sides of the truncation point. This indicates that added logistic noise might preserve more information than Gaussian additive noise. Both additive noises improve significantly over a non-randomization scheme (c.f. Figure 3 in Fithian et al. (2014)).
Of course, the increase in power and shortening of selective confidence intervals does not come without a price. Because we select with a randomized response, we are likely to select a worse model. But the trade-off between model quality and power is highly in favor of randomization. See the following example.
2.2 Linear regression with added noise
Back to the general setup of linear regression models, we select a model by solving LASSO with the randomized response and return the active set of the solution (as in (7)). Then per Lemma 2, we can construct valid selective tests in both and . For instance, in , we can construct tests for the hypothesis based on the law,
where , is the -th column of the identity matrix, is the projection matrix onto the column space of but orthogonal to , and are the appropriate matrix and vector corresponding to LASSO selection. This is a UMPU test due to the Lehmann–Scheffe construction (Fithian et al., 2014) and controls the selective Type-I error (11). Although, we cannot compute the explicit forms of (20), the selection events in (20) are polyhedrons and thus a hit-and-run or Hamiltonian Monte Carlo algorithm Pakman & Paninski (2012) can be used for sampling.
Figure 3 compares inference in the additive Gaussian noise scheme to the data carving procedure proposed in Fithian et al. (2014) as well as data splitting. In , the probability of screening (i.e. selecting including all the nonzero ’s) is a surrogate for the quality of the model. As additive noise uses a different randomization scheme than data splitting and data carving, we vary the amount of randomization used in each scheme and match on the probability of screening. Thus Figure 3 is like an ROC curve for the trade-off between model quality and power of tests. The -axis goes in the direction of increased randomization, with the left most point corresponding to no randomization at all. We see even with a small randomization that barely affects model selection, we can substantially lower the Type-II error from to less than . The trade-off is highly in favor of (small) randomization. We see in Figure 3 that additive noise lowers the Type-II error by almost half than data carving for the same screening probability and they both clearly dominate data splitting. For the concrete setup of the simulation, see Chapter 7 of Fithian et al. (2014).
Weak convergence and selective inference for statistical functionals
Throughout this section, we assume the dimension is fixed. We are interested in establishing a pivotal quantity for like (4) in Example 2 where is the sample mean after the randomized “file drawer effect”. It turns out we have an exact pivotal quantity if is normally distributed. To lighten notation, we suppress the script in the following lemma, which is a finite sample result valid for any . We prove the lemma in Section 7.
Of course the pivot in (22) is very difficult to compute explicitly, and we need to use sampling schemes like in (20). But in a nutshell, is simply a CDF transform of the law
After introducing the null statistic, Lemma 7 is agnostic to the selected model , where or the saturated model , where the parameter is simply . The nuances between the two models in terms of sampling is that the saturated model condition on (treating it as part of ), but selected model integrate over .
Lemma 7 is written with implicitly being the approximate average of i.i.d variables, hence the distribution . Linearizable statistics are of particular interest as they converge to due to central limit theorem. In the following, we seek to establish conditions under which the pivot will be asymptotically .
In other work on asymptotics of selective inference Tian & Taylor (2015); Tibshirani et al. (2015), the setup considered is usually the saturated model . These works considered asymptotics of selective inference marginalized over the range of . In contrast, we consider the convergence for any particular selected model , under the conditional law of the selection event . Specifically, we allow weak convergence of the pivot in (22) in the sequence of selected models . As explained above, selected models integrate over the null statistics while saturated models condition on those, thus the selective tests should have more power provided that the selected model is believable. In the saturated model, our result provides a finer measure of convergence than in Tian & Taylor (2015). On the other hand, Tian & Taylor (2015) allows high-dimensional setting in some cases while we consider fixed dimension .
Similar to the asymptotic setting in Section 3, we consider the convergence of under a sequence of models selected by a sequence of selection procedures . is a sequence of linearizable statistics defined in Definition 6, with asymptotic mean and asymptotic covariance matrix .
It will be convenient to rewrite the likelihood ratio in terms of the normalized vector
where denotes the k-fold differentiation with respect to the -dimensional vector , denotes element wise maximum.
Now we state our selective central limit theorem, which we prove in Section 7.
Moreover, assume has uniformly bounded moment generating function in some neighbourhood of . Namely, , such that
Then, for any with uniformly bounded derivatives up to third order
where depends only on the bounds on the derivatives of , the constants and the dimension . Thus the convergence is uniform in for models satisfying (28), (29) and (30).
2 Revisit the “file drawer problem”
In Examples 1 and 2, we considered only reporting an interval or a -value about when or . This is an example where we do not really select a model, but rather select only a proportion of the data to report. The selective distribution simply refers to the law of the reported sample means, which pass the threshold.
The data we observe is with the linearizable statistic simply being the sample mean . Example 1 corresponds to the degenerate randomization of adding 0 to . Work of Tian & Taylor (2015) show that in order for the corresponding pivot to converge weakly we can take, for fixed
That is, will satisfy a selective CLT when the population mean is not too negative.
On the other hand, in Example 2, the pivot in (22) is of the form,
When is the Logistic noise, then condition (28) and (30) can be verified. Formally, we have the following lemma whose proof we defer to the appendix,
If , with being the scale parameter, then if centered ’s have moment generating functions in the neighbourhood of zero, then the pivot is asymptotically .
In other words, with Logistic randomization noise, we can take the sequence of models to be
Requiring exponential moments is stricter than the third moment condition in (32), but we would have a stronger conclusion, namely weak convergence uniformly over all ’s.
3 Two-sample median problem
with with probability .
Our (randomized) selection algorithm reports
This pivot strikes a similarity with the pivot in (33) for Example 2 with the truncation threshold being replaced by and plugging in the appropriate means and variances of the medians. A result similar to Lemma 10 can be established, which ensures convergence of the pivot uniformly for any underlying medians .
In order to construct the above pivot, we need knowledge of the variance . Without selection, there are natural estimates of this variance. One may ask, how will inference be affected if we plug this estimate into our pivot? We revisit this question in Section 5.5.
4 Affine selection events
In this section, we discuss the special case of affine selection events (regions). This combined with the asymptotic result in Theorem 9 applies to more general settings. In particular, it allows us to approximate non-affine regions. For a concrete example, see Section 5.4.1.
We again normalize to be , then the selection event can be rewritten as
where , converges to .
Lower bound: We assume there is some norm , such that
Smoothness: Suppose has density , we assume the first 3 derivatives of are integrable,
where the norm on the left-hand side is the maximum element-wise of the partial derivatives.
The above two conditions essentially require to be differentiable and have heavier tails than (or equal to) exponential tails. In fact we prove that the lower bound and smoothness conditions ensure that (28) are satisfied under the local alternatives introduced below.
For the sequence of selected model , we define the local alternatives of radius of to be the set all sequences , such that
where is the distance induced by the norm .
The notion of local alternatives is natural in the asymptotic setting as we expect even a small effect size will be more prominent when we collect more and more data.
Formally, we have the following lemma, whose proof is deferred to the appendix.
Suppose , satisfy the lower bound and smoothness conditions, then condition (28) are satisfied under the local alternatives.
Now, we are left to verify conditions (29) and (30). Condition (29) is essentially a moment condition on the centered statistics , which we have to assume. Condition (30) can be verified using the well known results in multivariate CLT (see Gotze (1991)). To be rigorous, we state the following lemma, which we also prove in the appendix.
Unlike the sample mean and sample median examples, the pivot is difficult to compute explicitly in this case. However, as we discuss in the beginning of Section 5, the pivot is essentially the CDF transform of the conditional law (23), which we can sample from. As discussed above, we can just take to be from a Logistic distribution.
Now we apply the above theory to logistic regression.
Selective inference in this setting has not been considered before. Without the Gaussian assumptions Lee et al. (2013a) does not apply. The parametric setting of this problem has been discussed in Fithian et al. (2014), but computation of the selective tests are mostly infeasible for general . Finally, the asymptotic result by Tian & Taylor (2015) does not apply here as the framework require exactly affine selection regions, which is not the case in this setting.
Suppose the solution to (38) has nonzero entry set , then our target of inference , the unique population minimizer which satisfies
Selective inference in this setting is carried out conditioned on , the active set and its signs. We first introduce the following notations,
where is the feature matrix, and , is the columns corresponding to the active set and inactive set respectively. By law of large numbers, we have
Now we introduce our linearizable statistics and show that the conditioning event can be expressed as affine regions of these statistics.
Suppose is the active set of the solution of (38), and we denote
as the unpenalized MLE restricted to the selected variables .
The following statistic is linearizable with asymptotic mean and variance ,
The proof of this lemma is also deferred to the appendix.
Thus using Lemma 12 and Lemma 13, we can conclude under local alternatives, the pivot (22) converges to . To test , we take , and sample
In Lemma 14, we assume the covariance matrix is known. In applications, we can bootstrap it. But is it valid to plug in the bootstrap estimate of ?
5 Plugging in variance estimates
In Section 5.3 we derived quantities that were asymptotically pivotal for the best median, up to an unknown variance. In the sample median case, by (35), the variance of the sample median is approximately , where is the PDF evaluated at the median . A simple consistent estimator for is to take quantiles and , then
is consistent for based on which we get a consistent estimator for .
Figure 4 is some simulation results for the two-sample medians problem. In each case, we take the sample size for each treatment group to be , and generate the noise from a skewed distribution . We standardize it such that the noise has median under the null hypothesis. We use additive logistic noise with scale for randomization. The better group is decided using the randomized sample median, and selective inference is carried out. In Figure 4(a), the pivot with plugin variance estimate in (41) is plotted under both the null hypothesis and the . The pivot has reasonable power even for identifying local alternatives. The pivot is almost exactly under the null hypothesis with the sample size . In fact, it is very close at a relatively small sample size justifying the application of asymptotics in the nonparametric setting. Figure 4(b) further illustrates the difference in the unselective v.s. selective distribution and its convergence to its theoretical limit. We see that there is a clear shift in selective distribution that calls for adjustment for the selection. For sample size , the empirical selective distribution converges to our theoretical distribution.
Multiple Randomizations of the Data
Most of the examples above focus on a single randomization on the data, which we use for model selection. We naturally want to extend it to multiple randomizations, and multiple randomized selections, which will collectively suggest a model for inference. In this section, we allow multiple randomizations in a possibly sequential fashion and discuss how inference can be carried out.
Consider the case where we first choose a regularization parameter by cross-validation, and then fit the square-root LASSO problem Belloni et al. (2011) at this parameter,
where is picked from a fixed grid . The discussion below is not specific to selection by square-root LASSO.
The model selected by cross-validated square-root LASSO involves two steps of selection. We denote by the response for selecting the randomization parameter, and the response vector for fitting the square-root LASSO at the selected regularization parameter . Both vectors are randomized version of the original vector . Inference after cross validation requires combining two steps of randomized selection. Consider the following procedure.
First, we randomize to get the vector and
Note the intermediate vector is introduced convenience of sampling. The above is just one of the plausible randomization schemes.
After having randomized, we select with -fold cross-validation using :
where is the usual -fold cross-validation score with coefficients estimated by the square-root LASSO. Alternatively, one could compute the cross-validation score using the OLS estimators of the selected variables. Note that we have left implicit the randomization that splits observations into groups. That is in (44) above is a function of where is a random partition of into groups. When we sample below, we redraw each time.
The subset of variables and signs is selected using the square-root LASSO with response :
After seeing the selected variables , we perform inference in the selected model . Since , we will still have an exponential family after selection. Per Lemma 2, we sample from the following law,
The additional conditioning on the signs are for computational reasons. In fact, recent development in Harris et al. (2016) proposes sampling schemes that overcome these difficulties, so that we do not need to condition on this additional information.
To sample from the above law, we use a Gibbs-type sampler, which iterate over , , and , conditional on the other three and the selection event. It includes the following steps.
Using the conditional independence of and given , we have
This is the computational bottleneck, as we do not have good description for the selection event for cross validation. A brute-force sampling scheme will be computationally expensive, as we need to refit the model over a grid of ’s. Thus, we do not update too often.
The conditional independence of and given implies,
Tian et al. (2015) has given an explicit description of the selection event
Thus hit-and-run sampling provides a tractable sampling scheme.
This is a simple step. Because the selection event is based on and , we have
This is also simple with our randomization scheme. Note that is conditionally independent of and given ,
Since we condition on , we essentially take and project out the update on the space orthogonal to that of .
A chain that iterates through the above four steps will give us samples from the desired distribution for inference.
2 Collaborative selective inference
One of the motivations of the reusable holdout described in Dwork et al. (2014) is that it allows a data analyst to repeatedly query a database yet still be able to approximately estimate expectations even after asking many questions about the data. Another version of this model may be that several groups wish to model the same data and then, as a consortium, decide on a final model and be able to approximately estimate expectations in this final model. We might call this collaborative selective inference.
Now suppose that the groups choose models and convene to discuss what the best model is . For every choice of models and final model , the following selective distribution can be used for valid selective inference
When the ’s are conditionally independent given then it is clear that
It is possible that the consortium has beforehand decided on an algorithm that will choose a best model automatically, determined by some function . In this case, one should use the selective distribution
When the models in question are parametric, perhaps Gaussian distributions, and the randomization is additive Gaussian noise the central data bank can explicitly lower bound the leftover information by
This quantity is expressible in terms of the marginal variance of and the central data bank’s noise generating distribution for . By maintaining a lower bound on the above quantity, the central data bank can maintain a minimum prescribed information in the data for final estimation and/or inference. In a sequential setting, where valid inference is desired at each step, maintaining a lower bound may involve releasing noisier and noisier versions of . Sampling under this scheme seems quite difficult, and we leave it as an area of interesting future research.
Proof
To prove Theorem 9, we first prove the following lemma, which might be of independent interest.
for some , where is a constant only dependent on the dimension.
Lemma 15 can be seen as an extension of the result by Chatterjee (2005) in the sense that the author in Chatterjee (2005) established result for the case . The proof is also an adaptation of the technique in Chatterjee (2005).
We also define for any and ,
Let be the derivative with respect to the -th row . Using Taylor’s expansion at , we have
where the precise form of the Taylor remainder depends on realizing the laws and on the same probability space. In order to not introduce new notation, we have avoided explicitly writing out this construction, directing readers to Chatterjee (2005) for details. Nevertheless,
where are centered version of and is some dimension dependent constant.
Let be the constant s.t , only depends on the dimension . Thus, using the independence of the ’s,
Now we bound these two expectations. By the exponential moment condition (29) and Lemma 17, it is easy to conclude the first term is bounded by
The second expectation is bounded by , an upper bound on the third moment of ,
and summing over terms, we have the conclusion of the lemma. ∎
Now we prove the main theorem, Theorem 9.
per condition (30) and is a bound on . ∎
2 Proof of Lemma 7
References
Appendix A Proof of Lemma 1
First we normalize the sample mean as and rewrite the pivot as
and is the CDF of the standard normal distribution. As , we can use Mills ratio to approximate the normal tail. Specifically, denote ,
We study the behavior of for . By studying its distribution, we will also see that , for , thus the term
Now we study the distribution of conditioning on . Since is a translation of a binomial distribution divided by , we can rewrite in terms of a Binomial distribution, which will be useful for calculating the conditional distribution of . Specifically,
Noticing that , for any , thus
Now let , and use the above inequality times, we have
We can draw two conclusions from (50). First, conditional on , , which implies the first term in the pivot approximation (49) . Moreover, (50) shows that the overshoot is not distributed in the limit. In fact, we can conclude its limit (if existed) is strictly stochastically dominated by an . Thus,
and hence the pivot does not converge to . ∎
Appendix B Proof of Lemma 14
We first prove that is in fact a linearizable statistic. Since is the restricted MLE, we see that
Thus we can conclude that is a linearizable statistic with
Now we rewrite the selection event in terms of . Using the KKT conditions of (38),
Plugging in the equalities in the KKT conditions, we will have,
Using the inequalities in the KKT conditions, we have the selection event is with , and defined in the lemma.
Appendix C Proofs related to Logistic noise
Throughout the article, logistic noise has played an important role in all the examples.
The following lemma on the tail behavior of the logistic distribution is crucial to all the proofs with added logistic noise. Let be the CDF of , with being the scale parameter. is the PDF of .
where ’s are universal constants.
where . For , define . By induction, I claim that for each , is rational such that the polynomial in the numerator is of order 2 less than the denominator, and the denominator polynomial is bounded below by 1. Hence, ’s are bounded on the interval $$. Now, it is not hard to see that
for universal ’s and . ∎
Now we state the following lemmas which are foundations of the proofs of various lemmas in the article.
Assume is a decomposable statistic and has mean , variance , and centered exponential moments in a neighbourhood of zero, i.e satisfies condition (29). Denote , then
In Example 2, if we normalize the sample mean , we can rewrite the selective likelihood ratio and the pivot as
for some only depending on .
Noticing the lower bound in (52), we have
On the other hand, using the upper bounds in (53), we have for ,
Since is convex on the positive axis, it is hard to see
The term cancels with the one in the denominator, thus (55) holds for as well.
Analogously, similar bounds can be derived for the derivatives of as well, thus we have the conclusion of the lemma. ∎
The proof of Lemma 4 is a simple application of Lemma 17 and Lemma 18.
By law of large numbers, we know that is consistent for unselectively. Thus, using the result by Lemma 3, we only need to verify that the selective likelihood is integrable in . For simplicity, we take .
First notice from (54) that the selective likelihood ratio is bounded by a multiple of . Then by Lemma 17,
The proof of Lemma 10 uses results in Lemma 18 and Lemma 15
It follows simply from (54) that condition (28) are satisfied with the norm function simply being the absolute value function. Therefore, we only need to verify (30). Note for
Appendix D Proofs related to affine selection regions
The quantity that appears in both the pivot and the selective likelihood ratio is
where . The associated selective likelihood in terms of is
We first rewrite the pivot in terms of .
If we assume the lower bound condition, then under the local alternatives with radius , i.e. , we have
where is a constant only depending on the normal distribution and the norm in the local alternatives condition.
We first see that the lower bound condition gives the following lower bound.
Finally, since the has uniformly bounded derivatives up to the third order, we have
Suppose the smoothness and the lower bound conditions are satisfied, then for local alternatives with radius ,
The smoothness condition implies the following upper bound. For a multi-index , we have
Therefore, from the smoothness condition,
This combined with Lemma 19 gives the conclusion of the lemma. ∎
Next, we derive the exponential bounds on the derivatives of the pivot with respect to .
Assuming the conditions of Lemma 12, for a multi-index up to the order of ,
To get a lower bound on the denominator, note (57)
Therefore, the denominator will be lower bounded by
On the other hand, the upper bound (59) ensures,
Note the derivatives of the pivot will be a polynomial in terms of the form,
and therefore, it is easy to get the conclusion of the lemma. ∎
D.2 Proof of Lemma 13
Using Lemma 19 and the following lemma, we can easily prove Lemma 13.
Proof of Lemma 22 uses the well known results of Berry-Esseen Theorem. A multivariate extension can be found in Gotze (1991).
Thus the difference in the two probabilities is
where only depends on the dimension . The last inequality is a direct application of equation (1.5) in Gotze (1991). ∎