Semiparametrically efficient inference based on signed ranks in symmetric independent component models
Pauliina Ilmonen, Davy Paindaveine
Introduction
In multivariate statistics, concepts of location and scatter are usually defined through affine transformations of a noise vector. To be more specific, assume that the observation is obtained through
where is a -vector, is a full-rank matrix and is some standardized random vector. The exact nature of the resulting location parameter , mixing matrix parameter , and scatter parameter crucially depends on the standardization adopted.
This paper focuses on an alternative standardization of , for which has mutually independent marginals with common median zero. The resulting model in (1)—the independent component (IC) model, say—is more flexible than the elliptical model, even if one restricts, as we will do, to vectors with symmetrically distributed marginals. The IC model indeed allows for heterogeneous marginal distributions for , whereas, in contrast, marginals in the elliptical model all share—up to location and scale—the same distribution, hence also the same tail weight. This severely affects the relevance of elliptical models for practical applications, particularly so for moderate to large dimensions, since it is then very unlikely that all variables share, for example, the same tail weight.
The IC model provides the most standard setup for independent component analysis (ICA), in which the mixing matrix is to be estimated on the basis of independent copies of , the objective being to recover (up to a translation) the original unobservable independent signals by premultiplying the ’s with the resulting . It is well known in ICA, however, that is severely unidentified: for any permutation matrix and any full-rank diagonal matrix , one can always write
This paper considers inference on the mixing matrix . More precisely, because of the identifiability issues above, we rather consider a normalized version of , where is a well-defined representative of the class of mixing matrices that are equivalent to . This parameter is actually the parameter of interest in ICA: an estimate of will indeed allow one to recover the independent signals equally well as an estimate of any other with . Interestingly, the situation is extremely similar when considering inference on in the elliptical model. There, is only identified up to a positive scalar factor, and it is often enough to focus on inference about the well-defined shape parameter (e.g., in PCA, principal directions, proportions of explained variance, etc. can be computed from ). Just as is a normalized version of in the IC model, is a normalized version of in the elliptical model, and in both classes of models, the normalized parameters actually are the natural parameters of interest in many inference problems. The similarities further extend to the semiparametric nature of both models: just as the density of in the elliptical model, the pdf of the various independent components , , in the IC model, can hardly be assumed to be known in practice.
These strong similarities motivate the approach we adopt in this paper: we plan to conduct inference on (hypothesis testing and point estimation) in the IC model by adopting the methodology that proved extremely successful in Ha06c , Ha06 for inference on in the elliptical model. This methodology combines semiparametrically efficient inference and invariance arguments. In the IC model, the fixed- nonparametric submodels (indexed by ) indeed enjoy a strong invariance structure that is parallel to the one of the corresponding elliptical submodels (indexed by ). As in Ha06c , Ha06 , we exploit this invariance structure through a general result from HW03 that allows one to derive invariant versions of efficient central sequences, on the basis of which one can define semiparametrically efficient (at fixed target densities , ) invariant procedures. As the maximal invariant associated with the invariance structure considered turns out to be the vector of marginal signed ranks of the residuals, the proposed procedures are of a signed-rank nature and do not require to estimate densities. While they achieve semiparametric efficiency under correctly specified densities, they remain valid (correct asymptotic size under the null, for hypothesis testing, and root- consistency, for point estimation) under misspecified densities.
We will consider the problem of estimating and that of testing the null against the alternative , for some fixed . While point estimation is undoubtedly of primary importance for applications (e.g., in blind source separation), one might question the practical relevance of the testing problem considered, especially when is not the -dimensional identity matrix. Solving this generic testing problem, however, is the main step in developing tests for any linear hypothesis on , and we will explicitly describe the resulting tests in the sequel. An extensive study of these tests is beyond the scope of the present paper, though; we refer to IC for an extension of our tests to the particular case of testing the (linear) hypothesis that is block-diagonal, a problem that is obviously important in practice (nonrejection of the null would indeed allow practitioners to proceed with two separate, lower-dimensional, analyses). Testing linear hypotheses on includes many other testing problems of high practical relevance, such as testing that a given column of is equal to some fixed -vector, and testing that a given entry of is zero—the practical importance of these two testing problems, in relation, for example, with functional magnetic resonance imaging (fMRI), is discussed in Ol11 .
The paper is organized as follows. In Section 2, we fix the notation and describe the model (Section 2.1), state the corresponding uniformly locally and asymptotically normal (ULAN) property that allows us to determine semiparametric efficiency bounds (Section 2.2) and then introduce, in relation with invariance arguments, rank-based efficient central sequences (Section 2.3). In Sections 3 and 4, we develop the resulting rank tests and estimators for the mixing matrix , respectively. Our estimators actually require the delicate estimation of “cross-information coefficients,” an issue we solve in Section 4.2 by generalizing the method recently developed in r12 . In Section 5, simulations are conducted both to compare the proposed estimators with some competitors and to investigate the validity of asymptotic results—simulation results for hypothesis testing are provided in the supplementary article IP11a . Finally, the Appendix states some technical results (Appendix A) and reports proofs (Appendix B).
The model, the ULAN property and invariance arguments
where is the positive definite diagonal matrix that makes each column of have Euclidean norm one, is the permutation matrix for which the matrix satisfies for all and is the diagonal matrix such that all diagonal entries of are equal to one.
If one restricts to the collection of mixing matrices for which no ties occur in the permutation step above, it can easily be shown that, for any , we have that iff , so that this mechanism succeeds in identifying a unique representative in each class of equivalence (this is ensured with the double scaling scheme above, which may seem a bit complicated at first). Besides, is then a continuously differentiable mapping from onto . While ties may always be taken care of in some way (e.g., by basing the ordering on subsequent rows of the matrix ), they may prevent the mapping to be continuous, which would cause severe problems and would prevent us from using the Delta method in the sequel. It is clear, however, that the restriction to only gets rid of a few particular mixing matrices, and will not have any implications in practice.
The parametrization of the IC model we consider is then associated with
The resulting semiparametric model is then
Performing semiparametrically efficient inference on , at a fixed , typically requires that the corresponding parametric submodel satisfies the uniformly locally and asymptotically normal (ULAN) property.
2 The ULAN property
As always, the ULAN property requires technical regularity conditions on . In the present context, we need that each corresponding univariate pdf , , is absolutely continuous (with derivative , say) and satisfies
where , and full-rank information matrix
where and
and converges in distribution to a -variate normal distribution with mean zero and covariance matrix .
The performance of semiparametrically efficient tests on can similarly be characterized in terms of : a test of is semiparametrically efficient at (at asymptotic level ) if its asymptotic powers under local alternatives of the form , where is an arbitrary matrix with zero diagonal entries, are given by
where stands for the -upper quantile of the distribution, and denotes the cumulative distribution function of the noncentral distribution with noncentrality parameter .
3 Invariance arguments
Such an invariance structure actually exists and the relevant group collects all transformations
with and , where each , , is continuous, odd, monotone increasing and fixes . It is easy to check that is invariant under (and is generated by) , and that the corresponding maximal invariant is the vector of signed ranks
Theorem 2.1(ii) then follows from (8) and Theorem 2.1(i).
Inference procedures based on , unlike those (from CB06 ) based on the efficient central sequence obtained through tangent space projections, are measurable with respect to signed ranks, hence enjoy all nice properties usually associated with rank methods: robustness, ease of computation, validity without density estimation (and, for hypothesis testing, even distribution-freeness), etc.
Hypothesis testing
We now consider the problem of testing the null hypothesis against the alternative , with unspecified underlying density . Beyond their intrinsic interest, the resulting tests will play an important role in the construction of the -estimators of Section 4 below, and they pave the way to testing linear hypotheses on .
The objective here is to define a test that is semiparametrically efficient at some target density , yet that remains valid—in the sense that it meets asymptotically the level constraint—under a very broad class of densities . As we will show, this objective is achieved by the signed-rank test—, say—that rejects at asymptotic level whenever
where was introduced on Page 2.1 (an explicit expression is given below) and where is based on a sequence of estimators that is locally asymptotically discrete (see Appendix A for a precise definition) and root- consistent under the null.
In order to state this theorem, we need to define
We also let and , that involve (see Section 2.2) and . We then have the following result (see Appendix B for a proof).
This is to be compared to the semiparametric approach of Chen and Bickel CB06 —these authors focus on point estimation, but their methodology also leads to tests that enjoy the same properties as their estimators. Their procedures achieve uniform (in ) semiparametric efficiency, while our methods achieve semiparametric efficiency at the target density only—more precisely, at any corresponding . However, it turns out that the performances of our procedures do not depend much on the target density , so that our procedures are close to achieving uniform (in ) semiparametric efficiency; see the simulations in the supplemental article IP11a . As any uniformly semiparametrically efficient procedures (see Am02 ), Chen and Bickel’s procedures require estimating , hence choosing various smoothing parameters. In contrast, our procedures, by construction, are invariant (here, signed-rank) ones. As such, they do not require us to estimate densities, and they are robust, easy to compute, etc.
One might still object that the choice of is quite arbitrary. This choice should be based on the practitioner’s prior belief on the underlying densities. If he/she has no such prior belief, a kernel estimate of could be used. The resulting test would then enjoy the same properties as any in terms of validity, since kernel density estimators, in the symmetric case considered, typically are measurable with respect to the order statistics of the ’s, that, asymptotically, are stochastically independent of the signed ranks used in ; see HW03 for details. The test would further achieve uniform semiparametric efficiency.
Further results on the proposed tests are given in the supplemental article IP11a . More precisely, a simple explicit expression of the test statistics, local asymptotic powers of the corresponding tests, and simulation results can be found there.
for some fixed . If one forgets about the tacitly assumed constraint that in (14), the null hypothesis above imposes a set of linear constraints on . This clearly includes all testing problems mentioned in the Introduction: testing that a given column of is equal to a fixed vector, testing that a given (off-diagonal) entry of is zero and testing block-diagonality of .
Inspired by the tests from Le86 (Section 10.9), the analog of our signed-rank test above then rejects for large values of
with where denotes the Moore–Penrose pseudoinverse of , and where is an estimator of that is locally and asymptotically discrete, root- consistent under the null, and constrained—in the sense that satisfies the linear constraints in .
Point estimation
We turn to the problem of estimating , which is of primary importance for applications. Denoting by the signed-rank test statistic for in (10), a natural signed-rank estimator of is obtained by “inverting the corresponding test,”
This estimator, however, is not satisfactory: as any signed-rank quantity, the objective function is piecewise constant, hence discontinuous and nonconvex, which makes it very difficult to derive the asymptotic properties of . It is also virtually impossible to compute in practice, since this lack of smoothness and convexity essentially forces computing the estimator by simply running over a grid of possible values of the -dimensional parameter —a strategy that cannot provide a reasonable approximation of , even for moderate values of . Finally, there is no way to estimate the asymptotic covariance matrix of , which rules out the possibility to derive confidence zones for , hence drastically restricts the practical relevance of this estimator.
In order to avoid the aforementioned drawbacks, we propose adopting a one-step approach that was first used in Ha06c for the problem of estimating the shape of an elliptical distribution or in Ha08b in a more general context. The resulting one-step signed-rank estimators—in the sequel, we simply speak of one-step rank estimators or one-step -estimators—can easily be computed in practice, their asymptotic properties can be derived explicitly, and their asymptotic covariance matrix can be estimated consistently.
Describing our one-step -estimators requires:
For any target density , we propose the one-step -estimator , with values in , defined by
The following result states the asymptotic properties of this estimator (see Appendix B for a proof).
as , where is defined in Theorem A.1 (see Appendix A). (ii) The estimator is semiparametrically efficient at .
For , define and as the statistics obtained by plugging the estimators and from Assumption (A) in
and let , . The estimator then admits the following explicit expression (see Appendix B for a proof).
where stands for the diagonal matrix with the same diagonal entries as .
As shown above, the estimator enjoys very nice properties: its asymptotic behavior is completely characterized, it is semiparametrically efficient under correctly specified densities, yet remains root- consistent and asymptotically normal under a broad range of densities , its asymptotic covariance matrix can easily be estimated consistently, etc.
However, requires estimates and that fulfill Assumption (A). We now provide such estimates.
2 Estimation of cross-information coefficients
Therefore, we rather propose a solution that is based on ranks and avoids estimating the underlying nuisance . The method, that relies on the asymptotic linearity—under —of an appropriate rank-based statistic , was first used in Ha06c , where there is only one cross-information coefficient to be estimated. There, it is crucial that is involved as a scalar factor in the asymptotic covariance matrix, under , between the rank-based efficient central sequence and the parametric central sequence . In r12 , the method was extended to allow for the estimation of a cross-information coefficient that appears as a scalar factor in the linear term of the asymptotic linearity, under , of a (possibly vector-valued) rank-based statistic .
In all cases, thus, this method was only used to estimate a single cross-information coefficient that appears as a scalar factor in some structural—typically, cross-information—matrix. In this respect, our problem, which requires us to estimate cross-information quantities appearing in various entries of the cross-information matrix , is much more complex. Yet, as we now show, it allows for a solution relying on the same basic idea of exploiting the asymptotic linearity, under , of an appropriate -score rank-based statistic.
Simulations
Here we report simulation results for point estimation only—simulation results for hypothesis testing can be found in the supplemental article IP11a . Our aim is to both compare the proposed estimators with some competitors and to investigate the validity of asymptotic results.
We focused on the bivariate case , and we generated, for three different setups indexed by , independent random samples , , of size . Denoting by the common pdf of , , , the marginal densities and were chosen as follows: {longlist}
In Setup , is the pdf of the standard normal distribution (), and is the pdf of the Student distribution with degrees of freedom ();
In Setup , is the pdf of the logistic distribution with scale parameter one (log), and is ;
In Setup , is and is . We chose to use and , so that the observations are given by (other values of and led to extremely similar results).
Figure 1 reports, for each setup , a boxplot of the squared errors
for each of the twelve estimators considered (the nine -estimators and their three competitors).
As a conclusion, for practical sample sizes, the proposed -estimators outperform the standard competitors considered, and their behavior is very well in line with our asymptotic results.
Finally, we illustrate the proposed method for estimating cross-information coefficients. We consider again the first 50 replications of our simulation with , and focus on Setup 1 () and the target density (. The cross-information coefficients to be estimated then are , , and . The upper left picture in Figure 3 shows 150 graphs of the mapping (based on ), among which the 50
Appendix A Rank-based efficient central sequences
In this first Appendix, we study the asymptotic behavior of the rank-based efficient central sequences . The main result is the following (see Appendix B for a proof).
Both for hypothesis testing and point estimation, we had to replace in the parameter with some estimator (, say). The asymptotic behavior of the resulting (so-called aligned) rank-based efficient central sequence is given in the following result.
Since the sequence of estimators is assumed to be locally asymptotically discrete [which means that the number of possible values of in balls with radius centered at is bounded as ], this result is a direct consequence of Theorem A.1(iii) and Lemma 4.4 from kreiss . Local asymptotic discreteness is a concept that goes back to Le Cam and is quite standard in one-step estimation; see, for example, Bi82 or kreiss .
for some arbitrary constant . In practice, however, one can safely forget about such discretizations: irrespective of the accuracy of the computer used, the discretization constant can always be chosen large enough to make discretization be irrelevant at the fixed sample size at hand—hence also at any .
Appendix B Proofs
as , where stands for the common cdf of the ’s and denotes the rank of among . The quantities in (25) and (26) are linear signed-rank quantities that are said to be based on approximate and exact scores, respectively.
We go on with the proof of Theorem 2.1, for which it is important to note that, by proceeding as in the proof of Theorem A.1(i) but with (26) instead of (25), we further obtain that
B.2 Proof of Theorem 3.1
B.3 Proofs of Lemma 4.1, Theorems 4.1 and 4.2
Now, by using the fact that for any matrix with only zero diagonal entries, we have that , so that (16), (17) and (18) follow from (34), (B.3) and (36), respectively.
To prove Theorem 4.2, we will need the following result.
where denotes the entry of .
Proof of Theorem 4.2 By using again the fact that for any matrix with only zero diagonal entries, and then Lemma B.1, we obtain
The identity then yields
which, by using the fact that for any matrix with only zero diagonal entries, leads to
Acknowledgments
We would like to express our gratitude to the Co-Editor, Professor Peter Bühlmann, an Associate Editor and one referee. Their careful reading of a previous version of the paper and their comments and suggestions led to a considerable improvement of the present paper. We are also grateful to Klaus Nordhausen for sending to us the R code for FastICA authored by Abhijit Mandal.
[id=suppA] \stitleFurther results on tests and a proof of Theorem 4.3 \slink[doi]10.1214/11-AOS906SUPP \sdatatype.pdf \sfilenameaos906_supp.pdf \sdescriptionThis supplement provides a simple explicit expression for the proposed test statistics, derives local asymptotic powers of the corresponding tests, and presents simulation results for hypothesis testing. It also gives a proof of Theorem 4.3.