Semiparametrically efficient inference based on signed ranks in symmetric independent component models

Pauliina Ilmonen, Davy Paindaveine

Introduction

In multivariate statistics, concepts of location and scatter are usually defined through affine transformations of a noise vector. To be more specific, assume that the observation XX is obtained through

where μ\mu is a pp-vector, Λ\Lambda is a full-rank p×pp\times p matrix and ZZ is some standardized random vector. The exact nature of the resulting location parameter μ\mu, mixing matrix parameter Λ\Lambda, and scatter parameter Σ=ΛΛ′\Sigma=\Lambda\Lambda^{\prime} crucially depends on the standardization adopted.

This paper focuses on an alternative standardization of ZZ, for which ZZ has mutually independent marginals with common median zero. The resulting model in (1)—the independent component (IC) model, say—is more flexible than the elliptical model, even if one restricts, as we will do, to vectors ZZ with symmetrically distributed marginals. The IC model indeed allows for heterogeneous marginal distributions for XX, whereas, in contrast, marginals in the elliptical model all share—up to location and scale—the same distribution, hence also the same tail weight. This severely affects the relevance of elliptical models for practical applications, particularly so for moderate to large dimensions, since it is then very unlikely that all variables share, for example, the same tail weight.

The IC model provides the most standard setup for independent component analysis (ICA), in which the mixing matrix Λ\Lambda is to be estimated on the basis of nn independent copies X1,…,XnX_{1},\ldots,X_{n} of XX, the objective being to recover (up to a translation) the original unobservable independent signals Z1,…,ZnZ_{1},\ldots,Z_{n} by premultiplying the XiX_{i}’s with the resulting Λ^−1\hat{\Lambda}^{-1}. It is well known in ICA, however, that Λ\Lambda is severely unidentified: for any p×pp\times p permutation matrix PP and any full-rank diagonal matrix DD, one can always write

This paper considers inference on the mixing matrix Λ\Lambda. More precisely, because of the identifiability issues above, we rather consider a normalized version LL of Λ\Lambda, where LL is a well-defined representative of the class of mixing matrices that are equivalent to Λ\Lambda. This parameter LL is actually the parameter of interest in ICA: an estimate of LL will indeed allow one to recover the independent signals Z1,…,ZnZ_{1},\ldots,Z_{n} equally well as an estimate of any other Λ\Lambda with Λ∼L\Lambda\sim L. Interestingly, the situation is extremely similar when considering inference on Σ\Sigma in the elliptical model. There, Σ\Sigma is only identified up to a positive scalar factor, and it is often enough to focus on inference about the well-defined shape parameter V=Σ/(det⁡Σ)1/pV=\Sigma/(\det\Sigma)^{1/p} (e.g., in PCA, principal directions, proportions of explained variance, etc. can be computed from VV). Just as LL is a normalized version of Λ\Lambda in the IC model, VV is a normalized version of Σ\Sigma in the elliptical model, and in both classes of models, the normalized parameters actually are the natural parameters of interest in many inference problems. The similarities further extend to the semiparametric nature of both models: just as the density g∥⋅∥g_{\|\cdot\|} of ∥Z∥\|Z\| in the elliptical model, the pdf grg_{r} of the various independent components ZrZ_{r}, r=1,…,pr=1,\ldots,p, in the IC model, can hardly be assumed to be known in practice.

These strong similarities motivate the approach we adopt in this paper: we plan to conduct inference on LL (hypothesis testing and point estimation) in the IC model by adopting the methodology that proved extremely successful in Ha06c , Ha06 for inference on VV in the elliptical model. This methodology combines semiparametrically efficient inference and invariance arguments. In the IC model, the fixed-(μ,Λ)(\mu,\Lambda) nonparametric submodels (indexed by g1,…,gpg_{1},\ldots,g_{p}) indeed enjoy a strong invariance structure that is parallel to the one of the corresponding elliptical submodels (indexed by g∥⋅∥g_{\|\cdot\|}). As in Ha06c , Ha06 , we exploit this invariance structure through a general result from HW03 that allows one to derive invariant versions of efficient central sequences, on the basis of which one can define semiparametrically efficient (at fixed target densities gr=frg_{r}=f_{r}, r=1,…,pr=1,\ldots,p) invariant procedures. As the maximal invariant associated with the invariance structure considered turns out to be the vector of marginal signed ranks of the residuals, the proposed procedures are of a signed-rank nature and do not require to estimate densities. While they achieve semiparametric efficiency under correctly specified densities, they remain valid (correct asymptotic size under the null, for hypothesis testing, and root-nn consistency, for point estimation) under misspecified densities.

We will consider the problem of estimating LL and that of testing the null H0\dvtxL=L0\mathcal{H}_{0}\dvtx L=L_{0} against the alternative H1\dvtxL≠L0\mathcal{H}_{1}\dvtx L\neq L_{0}, for some fixed L0L_{0}. While point estimation is undoubtedly of primary importance for applications (e.g., in blind source separation), one might question the practical relevance of the testing problem considered, especially when L0L_{0} is not the pp-dimensional identity matrix. Solving this generic testing problem, however, is the main step in developing tests for any linear hypothesis on LL, and we will explicitly describe the resulting tests in the sequel. An extensive study of these tests is beyond the scope of the present paper, though; we refer to IC for an extension of our tests to the particular case of testing the (linear) hypothesis that LL is block-diagonal, a problem that is obviously important in practice (nonrejection of the null would indeed allow practitioners to proceed with two separate, lower-dimensional, analyses). Testing linear hypotheses on LL includes many other testing problems of high practical relevance, such as testing that a given column of LL is equal to some fixed pp-vector, and testing that a given entry of LL is zero—the practical importance of these two testing problems, in relation, for example, with functional magnetic resonance imaging (fMRI), is discussed in Ol11 .

The paper is organized as follows. In Section 2, we fix the notation and describe the model (Section 2.1), state the corresponding uniformly locally and asymptotically normal (ULAN) property that allows us to determine semiparametric efficiency bounds (Section 2.2) and then introduce, in relation with invariance arguments, rank-based efficient central sequences (Section 2.3). In Sections 3 and 4, we develop the resulting rank tests and estimators for the mixing matrix LL, respectively. Our estimators actually require the delicate estimation of 2p(p−1)2p(p-1) “cross-information coefficients,” an issue we solve in Section 4.2 by generalizing the method recently developed in r12 . In Section 5, simulations are conducted both to compare the proposed estimators with some competitors and to investigate the validity of asymptotic results—simulation results for hypothesis testing are provided in the supplementary article IP11a . Finally, the Appendix states some technical results (Appendix A) and reports proofs (Appendix B).

The model, the ULAN property and invariance arguments

where D1+D^{+}_{1} is the positive definite diagonal matrix that makes each column of ΛD1+\Lambda D^{+}_{1} have Euclidean norm one, PP is the permutation matrix for which the matrix B=(bij)=ΛD1+PB=(b_{ij})=\Lambda D^{+}_{1}P satisfies ∣bii∣>∣bij∣|b_{ii}|>|b_{ij}| for all i<ji<j and D2D_{2} is the diagonal matrix such that all diagonal entries of Π(Λ)=ΛD1+PD2\Pi(\Lambda)=\Lambda D^{+}_{1}PD_{2} are equal to one.

If one restricts to the collection Mp\mathcal{M}_{p} of mixing matrices Λ\Lambda for which no ties occur in the permutation step above, it can easily be shown that, for any Λ1,Λ2∈Mp\Lambda_{1},\Lambda_{2}\in\mathcal{M}_{p}, we have that Λ1∼Λ2\Lambda_{1}\sim\Lambda_{2} iff Π(Λ1)=Π(Λ2)\Pi(\Lambda_{1})=\Pi(\Lambda_{2}), so that this mechanism succeeds in identifying a unique representative in each class of equivalence (this is ensured with the double scaling scheme above, which may seem a bit complicated at first). Besides, Π\Pi is then a continuously differentiable mapping from Mp\mathcal{M}_{p} onto M1p:=Π(Mp)\mathcal{M}_{1p}:=\Pi(\mathcal{M}_{p}). While ties may always be taken care of in some way (e.g., by basing the ordering on subsequent rows of the matrix BB), they may prevent the mapping Π\Pi to be continuous, which would cause severe problems and would prevent us from using the Delta method in the sequel. It is clear, however, that the restriction to Mp\mathcal{M}_{p} only gets rid of a few particular mixing matrices, and will not have any implications in practice.

The parametrization of the IC model we consider is then associated with

The resulting semiparametric model is then

Performing semiparametrically efficient inference on ϑ\vartheta, at a fixed f∈Ff\in\mathcal{F}, typically requires that the corresponding parametric submodel Pf(n)\mathcal{P}_{f}^{(n)} satisfies the uniformly locally and asymptotically normal (ULAN) property.

2 The ULAN property

As always, the ULAN property requires technical regularity conditions on ff. In the present context, we need that each corresponding univariate pdf frf_{r}, r=1,…,pr=1,\ldots,p, is absolutely continuous (with derivative fr′f^{\prime}_{r}, say) and satisfies

where Zi=Zi(ϑ)=L−1(Xi−μ)Z_{i}=Z_{i}(\vartheta)=L^{-1}(X_{i}-\mu), and full-rank information matrix

where ΓL,f;1:=(L−1)′IfL−1\Gamma_{L,f;1}:=(L^{-1})^{\prime}\mathcal{I}_{f}L^{-1} and

and Δϑn,f\Delta_{\vartheta_{n},f} converges in distribution to a p2p^{2}-variate normal distribution with mean zero and covariance matrix ΓL,f\Gamma_{L,f}.

The performance of semiparametrically efficient tests on LL can similarly be characterized in terms of ΓL,f;2∗\Gamma^{*}_{L,f;2}: a test of H0\dvtxL=L0\mathcal{H}_{0}\dvtx L=L_{0} is semiparametrically efficient at ff (at asymptotic level α\alpha) if its asymptotic powers under local alternatives of the form H1(n)\dvtxL=L0+n−1/2H\mathcal{H}_{1}^{(n)}\dvtx L=L_{0}+n^{-1/2}H, where HH is an arbitrary p×pp\times p matrix with zero diagonal entries, are given by

where χp(p−1),1−α2\chi^{2}_{p(p-1),1-\alpha} stands for the α\alpha-upper quantile of the χp(p−1)2\chi^{2}_{p(p-1)} distribution, and Ψp(p−1)(⋅;δ)\Psi_{p(p-1)}(\cdot;\delta) denotes the cumulative distribution function of the noncentral χp(p−1)2\chi^{2}_{p(p-1)} distribution with noncentrality parameter δ\delta.

3 Invariance arguments

Such an invariance structure actually exists and the relevant group Gϑ\mathcal{G}^{\vartheta} collects all transformations

with zi(ϑ):=L−1(xi−μ)z_{i}(\vartheta):=L^{-1}(x_{i}-\mu) and h((z1,…,zp)′)=(h1(z1),…,hp(zp))′h((z_{1},\ldots,z_{p})^{\prime})=(h_{1}(z_{1}),\ldots,h_{p}(z_{p}))^{\prime}, where each hrh_{r}, r=1,…,pr=1,\ldots,p, is continuous, odd, monotone increasing and fixes +∞+\infty. It is easy to check that Pϑ(n)\mathcal{P}_{\vartheta}^{(n)} is invariant under (and is generated by) Gϑ\mathcal{G}^{\vartheta}, and that the corresponding maximal invariant is the vector of signed ranks

Theorem 2.1(ii) then follows from (8) and Theorem 2.1(i).

Inference procedures based on Δ‾ϑ,f;2∗\underline{\Delta}^{*}_{\vartheta,f;2}, unlike those (from CB06 ) based on the efficient central sequence Δϑ,f;2∗\Delta^{*}_{\vartheta,f;2} obtained through tangent space projections, are measurable with respect to signed ranks, hence enjoy all nice properties usually associated with rank methods: robustness, ease of computation, validity without density estimation (and, for hypothesis testing, even distribution-freeness), etc.

Hypothesis testing

We now consider the problem of testing the null hypothesis H0\dvtxL=L0\mathcal{H}_{0}\dvtx L=L_{0} against the alternative H1\dvtxL≠L0\mathcal{H}_{1}\dvtx L\neq L_{0}, with unspecified underlying density gg. Beyond their intrinsic interest, the resulting tests will play an important role in the construction of the RR-estimators of Section 4 below, and they pave the way to testing linear hypotheses on LL.

The objective here is to define a test that is semiparametrically efficient at some target density ff, yet that remains valid—in the sense that it meets asymptotically the level constraint—under a very broad class of densities gg. As we will show, this objective is achieved by the signed-rank test—ϕ‾f\underline{\phi}_{f}, say—that rejects H0\mathcal{H}_{0} at asymptotic level α∈(0,1)\alpha\in(0,1) whenever

where ΓL,f;2∗\Gamma^{*}_{L,f;2} was introduced on Page 2.1 (an explicit expression is given below) and where ϑ^0=(μ^′,(vecd⁡∘L0)′)′\hat{\vartheta}_{0}=(\hat{\mu}^{\prime},(\operatorname{vecd}^{\circ}L_{0})^{\prime})^{\prime} is based on a sequence of estimators μ^\hat{\mu} that is locally asymptotically discrete (see Appendix A for a precise definition) and root-nn consistent under the null.

In order to state this theorem, we need to define

We also let ΓL,f;2∗:=ΓL,f,f;2∗\Gamma^{*}_{L,f;2}:=\Gamma^{*}_{L,f,f;2} and Gf:=Gf,fG_{f}:=G_{f,f}, that involve γrs(f,f)=γrs(f)\gamma_{rs}(f,f)=\gamma_{rs}(f) (see Section 2.2) and ρrs(f,f)=1\rho_{rs}(f,f)=1. We then have the following result (see Appendix B for a proof).

This is to be compared to the semiparametric approach of Chen and Bickel CB06 —these authors focus on point estimation, but their methodology also leads to tests that enjoy the same properties as their estimators. Their procedures achieve uniform (in gg) semiparametric efficiency, while our methods achieve semiparametric efficiency at the target density ff only—more precisely, at any corresponding fσf_{\sigma}. However, it turns out that the performances of our procedures do not depend much on the target density ff, so that our procedures are close to achieving uniform (in gg) semiparametric efficiency; see the simulations in the supplemental article IP11a . As any uniformly semiparametrically efficient procedures (see Am02 ), Chen and Bickel’s procedures require estimating gg, hence choosing various smoothing parameters. In contrast, our procedures, by construction, are invariant (here, signed-rank) ones. As such, they do not require us to estimate densities, and they are robust, easy to compute, etc.

One might still object that the choice of ff is quite arbitrary. This choice should be based on the practitioner’s prior belief on the underlying densities. If he/she has no such prior belief, a kernel estimate f^\hat{f} of ff could be used. The resulting test ϕ‾f^\underline{\phi}_{\hat{f}} would then enjoy the same properties as any ϕ‾f\underline{\phi}_{f} in terms of validity, since kernel density estimators, in the symmetric case considered, typically are measurable with respect to the order statistics of the ∣Zir(ϑ^0)∣|Z_{ir}(\hat{\vartheta}_{0})|’s, that, asymptotically, are stochastically independent of the signed ranks Sir(ϑ^0),Rir+(ϑ^0)S_{ir}(\hat{\vartheta}_{0}),R^{+}_{ir}(\hat{\vartheta}_{0}) used in ϕ‾f\underline{\phi}_{f}; see HW03 for details. The test ϕ‾f^\underline{\phi}_{\hat{f}} would further achieve uniform semiparametric efficiency.

Further results on the proposed tests are given in the supplemental article IP11a . More precisely, a simple explicit expression of the test statistics, local asymptotic powers of the corresponding tests, and simulation results can be found there.

for some fixed L0∈M1pL_{0}\in\mathcal{M}_{1p}. If one forgets about the tacitly assumed constraint that L∈M1pL\in\mathcal{M}_{1p} in (14), the null hypothesis above imposes a set of linear constraints on LL. This clearly includes all testing problems mentioned in the Introduction: testing that a given column of LL is equal to a fixed vector, testing that a given (off-diagonal) entry of LL is zero and testing block-diagonality of LL.

Inspired by the tests from Le86 (Section 10.9), the analog of our signed-rank test ϕ‾f\underline{\phi}_{f} above then rejects H0(L0,Ω)\mathcal{H}_{0}(L_{0},\Omega) for large values of

with PΩ:=(ΓL^,f;2∗)−−Ω(Ω′ΓL^,f;2∗Ω)−Ω′,P_{\Omega}:=(\Gamma^{*}_{\hat{L},f;2})^{-}-\Omega(\Omega^{\prime}\Gamma^{*}_{\hat{L},f;2}\Omega)^{-}\Omega^{\prime}, where B−B^{-} denotes the Moore–Penrose pseudoinverse of BB, and where ϑ^=(μ^′,(vecd⁡∘L^)′)′\hat{\vartheta}=(\hat{\mu}^{\prime},(\operatorname{vecd}^{\circ}\hat{L})^{\prime})^{\prime} is an estimator of ϑ\vartheta that is locally and asymptotically discrete, root-nn consistent under the null, and constrained—in the sense that L^\hat{L} satisfies the linear constraints in H0(L0,Ω)\mathcal{H}_{0}(L_{0},\Omega).

Point estimation

We turn to the problem of estimating LL, which is of primary importance for applications. Denoting by Q‾f=Q‾f(L0)\underline{Q}_{f}=\underline{Q}_{f}(L_{0}) the signed-rank test statistic for H0\dvtxL=L0\mathcal{H}_{0}\dvtx L=L_{0} in (10), a natural signed-rank estimator of LL is obtained by “inverting the corresponding test,”

This estimator, however, is not satisfactory: as any signed-rank quantity, the objective function L↦Q‾f(L)L\mapsto\underline{Q}_{f}(L) is piecewise constant, hence discontinuous and nonconvex, which makes it very difficult to derive the asymptotic properties of L‾^f;arg⁡min⁡\hat{\underline{L}}_{f;\arg\min}. It is also virtually impossible to compute L‾^f;arg⁡min⁡\hat{\underline{L}}_{f;\arg\min} in practice, since this lack of smoothness and convexity essentially forces computing the estimator by simply running over a grid of possible values of the p(p−1)p(p-1)-dimensional parameter LL—a strategy that cannot provide a reasonable approximation of L‾^f;arg⁡min⁡\hat{\underline{L}}_{f;\arg\min}, even for moderate values of pp. Finally, there is no way to estimate the asymptotic covariance matrix of L‾^f;arg⁡min⁡\hat{\underline{L}}_{f;\arg\min}, which rules out the possibility to derive confidence zones for LL, hence drastically restricts the practical relevance of this estimator.

In order to avoid the aforementioned drawbacks, we propose adopting a one-step approach that was first used in Ha06c for the problem of estimating the shape of an elliptical distribution or in Ha08b in a more general context. The resulting one-step signed-rank estimators—in the sequel, we simply speak of one-step rank estimators or one-step RR-estimators—can easily be computed in practice, their asymptotic properties can be derived explicitly, and their asymptotic covariance matrix can be estimated consistently.

Describing our one-step RR-estimators requires:

For any target density ff, we propose the one-step RR-estimator L‾^f\hat{\underline{L}}_{f}, with values in M1p\mathcal{M}_{1p}, defined by

The following result states the asymptotic properties of this estimator (see Appendix B for a proof).

as n→∞n\to\infty, where Δϑ,f,g;2∗\Delta^{*}_{\vartheta,f,g;2} is defined in Theorem A.1 (see Appendix A). (ii) The estimator L‾^f\hat{\underline{L}}_{f} is semiparametrically efficient at ff.

For r≠s∈{1,…,p}r\neq s\in\{1,\ldots,p\}, define α^rs(f)\hat{\alpha}_{rs}(f) and β^rs(f)\hat{\beta}_{rs}(f) as the statistics obtained by plugging the estimators γ^rs(f)\hat{\gamma}_{rs}(f) and ρ^rs(f)\hat{\rho}_{rs}(f) from Assumption (A) in

and let α^rr(f):=0=:β^rr(f)\hat{\alpha}_{rr}(f):=0=:\hat{\beta}_{rr}(f), r=1,…,pr=1,\ldots,p. The estimator L‾^f\hat{\underline{L}}_{f} then admits the following explicit expression (see Appendix B for a proof).

where diag⁡(A)=A−odiag⁡(A)\operatorname{diag}(A)=A-\operatorname{odiag}(A) stands for the diagonal matrix with the same diagonal entries as AA.

As shown above, the estimator L‾^f\hat{\underline{L}}_{f} enjoys very nice properties: its asymptotic behavior is completely characterized, it is semiparametrically efficient under correctly specified densities, yet remains root-nn consistent and asymptotically normal under a broad range of densities gg, its asymptotic covariance matrix can easily be estimated consistently, etc.

However, L‾^f\hat{\underline{L}}_{f} requires estimates γ^rs(f)\hat{\gamma}_{rs}(f) and ρ^rs(f)\hat{\rho}_{rs}(f) that fulfill Assumption (A). We now provide such estimates.

2 Estimation of cross-information coefficients

Therefore, we rather propose a solution that is based on ranks and avoids estimating the underlying nuisance gg. The method, that relies on the asymptotic linearity—under gg—of an appropriate rank-based statistic S‾ϑ,f\underline{S}_{\vartheta,f}, was first used in Ha06c , where there is only one cross-information coefficient J(f,g)J(f,g) to be estimated. There, it is crucial that J(f,g)J(f,g) is involved as a scalar factor in the asymptotic covariance matrix, under gg, between the rank-based efficient central sequence Δ‾ϑ,f∗\underline{\Delta}^{*}_{\vartheta,f} and the parametric central sequence Δϑ,g{\Delta}_{\vartheta,g}. In r12 , the method was extended to allow for the estimation of a cross-information coefficient that appears as a scalar factor in the linear term of the asymptotic linearity, under gg, of a (possibly vector-valued) rank-based statistic S‾ϑ,f\underline{S}_{\vartheta,f}.

In all cases, thus, this method was only used to estimate a single cross-information coefficient that appears as a scalar factor in some structural—typically, cross-information—matrix. In this respect, our problem, which requires us to estimate 2p(p−1)2p(p-1) cross-information quantities appearing in various entries of the cross-information matrix ΓL,f,g;2∗\Gamma^{*}_{L,f,g;2}, is much more complex. Yet, as we now show, it allows for a solution relying on the same basic idea of exploiting the asymptotic linearity, under gg, of an appropriate ff-score rank-based statistic.

Simulations

Here we report simulation results for point estimation only—simulation results for hypothesis testing can be found in the supplemental article IP11a . Our aim is to both compare the proposed estimators with some competitors and to investigate the validity of asymptotic results.

We focused on the bivariate case p=2p=2, and we generated, for three different setups indexed by d∈{1,2,3}d\in\{1,2,3\}, M=2\mbox,000M=2\mbox{,}000 independent random samples Zi(d,m)=(Zi1(d,m),Zi2(d,m))′Z^{(d,m)}_{i}=(Z^{(d,m)}_{i1},Z^{(d,m)}_{i2})^{\prime}, i=1,…,ni=1,\ldots,n, of size n=4\mbox,000n=4\mbox{,}000. Denoting by g(d)(z)=g1(d)(z1)g2(d)(z2)g^{(d)}(z)=g^{(d)}_{1}(z_{1})g^{(d)}_{2}(z_{2}) the common pdf of Zi(d,m)Z^{(d,m)}_{i}, i=1,…,ni=1,\ldots,n, m=1,…,Mm=1,\ldots,M, the marginal densities g1(d)g^{(d)}_{1} and g2(d)g^{(d)}_{2} were chosen as follows: {longlist}

In Setup d=1d=1, g1(d)g^{(d)}_{1} is the pdf of the standard normal distribution (N\mathcal{N}), and g2(d)g^{(d)}_{2} is the pdf of the Student distribution with 55 degrees of freedom (t5t_{5});

In Setup d=2d=2, g1(d)g^{(d)}_{1} is the pdf of the logistic distribution with scale parameter one (log), and g2(d)g^{(d)}_{2} is t5t_{5};

In Setup d=3d=3, g1(d)g^{(d)}_{1} is t8t_{8} and g2(d)g^{(d)}_{2} is t5t_{5}. We chose to use L=I2L=I_{2} and μ=(0,0)′\mu=(0,0)^{\prime}, so that the observations are given by Xi(d,m)=LZi(d,m)+μ=Zi(d,m)X_{i}^{(d,m)}=LZ^{(d,m)}_{i}+\mu=Z^{(d,m)}_{i} (other values of LL and μ\mu led to extremely similar results).

Figure 1 reports, for each setup dd, a boxplot of the MM squared errors

for each of the twelve estimators L^\hat{L} considered (the nine RR-estimators and their three competitors).

As a conclusion, for practical sample sizes, the proposed RR-estimators outperform the standard competitors considered, and their behavior is very well in line with our asymptotic results.

Finally, we illustrate the proposed method for estimating cross-information coefficients. We consider again the first 50 replications of our simulation with n=4\mbox,000n=4\mbox{,}000, and focus on Setup 1 (g=g(1)g=g^{(1)}) and the target density f=f(3)f=f^{(3)} (≠\neqg(1))g^{(1)}). The cross-information coefficients to be estimated then are γ12(f,g)≈1.478\gamma_{12}(f,g)\approx 1.478, γ21(f,g)≈0.862\gamma_{21}(f,g)\approx 0.862, ρ12(f,g)≈1.149\rho_{12}(f,g)\approx 1.149 and ρ21(f,g)≈0.887\rho_{21}(f,g)\approx 0.887. The upper left picture in Figure 3 shows 150 graphs of the mapping λ↦hγ12(λ)\lambda\mapsto h^{\gamma_{12}}(\lambda) (based on f=f(3)f=f^{(3)}), among which the 50

Appendix A Rank-based efficient central sequences

In this first Appendix, we study the asymptotic behavior of the rank-based efficient central sequences Δ‾ϑ,f;2∗\underline{\Delta}^{*}_{\vartheta,f;2}. The main result is the following (see Appendix B for a proof).

Both for hypothesis testing and point estimation, we had to replace in Δ‾ϑ,f;2∗\underline{\Delta}^{*}_{\vartheta,f;2} the parameter ϑ\vartheta with some estimator (ϑˇ(n)\check{\vartheta}^{(n)}, say). The asymptotic behavior of the resulting (so-called aligned) rank-based efficient central sequence Δ‾ϑˇ(n),f;2∗\underline{\Delta}^{*}_{\check{\vartheta}^{(n)},f;2} is given in the following result.

Since the sequence of estimators ϑˇ(n)\check{\vartheta}^{(n)} is assumed to be locally asymptotically discrete [which means that the number of possible values of ϑˇ(n)\check{\vartheta}^{(n)} in balls with O(n−1/2)O(n^{-1/2}) radius centered at ϑ\vartheta is bounded as n→∞n\to\infty], this result is a direct consequence of Theorem A.1(iii) and Lemma 4.4 from kreiss . Local asymptotic discreteness is a concept that goes back to Le Cam and is quite standard in one-step estimation; see, for example, Bi82 or kreiss .

for some arbitrary constant c>0c>0. In practice, however, one can safely forget about such discretizations: irrespective of the accuracy of the computer used, the discretization constant cc can always be chosen large enough to make discretization be irrelevant at the fixed sample size n0n_{0} at hand—hence also at any n>n0n>n_{0}.

Appendix B Proofs

as n→∞n\to\infty, where G+G_{+} stands for the common cdf of the ∣Yi∣|Y_{i}|’s and Ri+R^{+}_{i} denotes the rank of ∣Yi∣|Y_{i}| among ∣Y1∣,…,∣Yn∣|Y_{1}|,\ldots,|Y_{n}|. The quantities in (25) and (26) are linear signed-rank quantities that are said to be based on approximate and exact scores, respectively.

We go on with the proof of Theorem 2.1, for which it is important to note that, by proceeding as in the proof of Theorem A.1(i) but with (26) instead of (25), we further obtain that

B.2 Proof of Theorem 3.1

B.3 Proofs of Lemma 4.1, Theorems 4.1 and 4.2

Now, by using the fact that C′(vecd⁡∘H)=(vec⁡H)C^{\prime}(\operatorname{vecd}^{\circ}H)=(\operatorname{vec}H) for any p×pp\times p matrix HH with only zero diagonal entries, we have that nvec⁡(L‾^f−L)=nC′vecd⁡∘(L‾^f−L)\sqrt{n}\operatorname{vec}(\hat{\underline{L}}_{f}-L)=\sqrt{n}C^{\prime}\operatorname{vecd}^{\circ}(\hat{\underline{L}}_{f}-L), so that (16), (17) and (18) follow from (34), (B.3) and (36), respectively.

To prove Theorem 4.2, we will need the following result.

where LrsL_{rs} denotes the entry (r,s)(r,s) of LL.

Proof of Theorem 4.2 By using again the fact that C′(vecd⁡∘H)=(vec⁡H)C^{\prime}(\operatorname{vecd}^{\circ}H)=(\operatorname{vec}H) for any p×pp\times p matrix HH with only zero diagonal entries, and then Lemma B.1, we obtain

The identity (C′⊗A)(vec⁡B)=vec⁡(ABC)(C^{\prime}\otimes A)(\operatorname{vec}B)=\operatorname{vec}(ABC) then yields

which, by using the fact that C′(vecd⁡∘H)=(vec⁡H)C^{\prime}(\operatorname{vecd}^{\circ}H)=(\operatorname{vec}H) for any p×pp\times p matrix HH with only zero diagonal entries, leads to

Acknowledgments

We would like to express our gratitude to the Co-Editor, Professor Peter Bühlmann, an Associate Editor and one referee. Their careful reading of a previous version of the paper and their comments and suggestions led to a considerable improvement of the present paper. We are also grateful to Klaus Nordhausen for sending to us the R code for FastICA authored by Abhijit Mandal.

[id=suppA] \stitleFurther results on tests and a proof of Theorem 4.3 \slink[doi]10.1214/11-AOS906SUPP \sdatatype.pdf \sfilenameaos906_supp.pdf \sdescriptionThis supplement provides a simple explicit expression for the proposed test statistics, derives local asymptotic powers of the corresponding tests, and presents simulation results for hypothesis testing. It also gives a proof of Theorem 4.3.

References