Asymptotic power of sphericity tests for high-dimensional data
Alexei Onatski, Marcelo J. Moreira, Marc Hallin
Introduction
Recently, there has been much interest in testing sphericity in a high-dimensional setting. Various tests have been proposed and analyzed in Ledoit and Wolf (2002), Srivastava (2005), Birke and Dette (2005), Schott (2006), Bai et al. (2009), Fisher, Sun and Gallagher (2010), Chen, Zhang and Zhong (2010) and Berthet and Rigollet (2012). In many studies, a distinct interesting alternative to the null of sphericity is the existence of a low-dimensional structure or signal in the data. Detecting such a structure has been the focus of recent studies in various applied fields including population and medical genetics [Patterson, Price and Reich (2006)], econometrics [Onatski (2009, 2010)], wireless communication [Bianchi et al. (2011)], chemometrics [Kritchman and Nadler (2008)] and signal processing [Perry and Wolfe (2010)].
Most of the existing sphericity tests are based on the eigenvalues of the sample covariance matrix, which constitute the maximal invariant statistic with respect to orthogonal transformations of the data. The asymptotic power of such tests depends on the asymptotic behavior of the sample covariance eigenvalues under the alternative hypothesis. When the alternative is a rank- perturbation of the null, the corresponding population covariance matrix is proportional to a sum of the identity matrix and a matrix of rank . Johnstone (2001) calls such a situation “spiked covariance.”
The asymptotic behavior of the sample covariance eigenvalues in “spiked covariance” models of increasing dimension is well studied. Consider the simplest case, when . If the largest population covariance eigenvalue is above the “phase transition” threshold studied in Baik, Ben Arous and Péché (2005), then the largest sample covariance eigenvalue remains separated from the rest of the eigenvalues, which are asymptotically “packed together as in the support of the Marchenko–Pastur density” [Baik and Silverstein (2006)]. Since the largest eigenvalue separates from the “bulk,” it is easy to detect a signal.
If the largest population covariance eigenvalue is at or below the threshold, the empirical distribution of the sample covariance eigenvalues still converges to the Marchenko–Pastur distribution, but the largest sample covariance eigenvalue now converges to the upper boundary of its support, both under the null of sphericity and the “spiked” alternative [Silverstein and Bai (1995) and Baik and Silverstein (2006)]. Hence, the signal detection becomes problematic. At the threshold, the null and the alternative hypotheses lead to different asymptotic distributions for the centered and normalized largest sample covariance eigenvalue [Bloemendal and Virág (2012) and Mo (2012)], which implies some asymptotic detection power. However, below the threshold, the difference disappears with the joint distribution of any finite number of the centered and normalized largest sample covariance eigenvalues converging to the multivariate Tracy–Widom law under both the null and the alternative [Johnstone (2001), Baik, Ben Arous and Péché (2005), El Karoui (2007) and Féral and Péché (2009)].
This similarity in the asymptotic behavior of covariance eigenvalues under the null and the alternative prompts Nadakuditi and Edelman (2008) and Nadakuditi and Silverstein (2010) to call the transition threshold “the fundamental asymptotic limit of sample-eigenvalue-based detection.” They claim that no reliable signal detection is possible below that limit in the asymptotic sense. This asymptotic impossibility is also pointed out and discussed in several other recent studies, including Patterson, Price and Reich (2006), Hoyle (2008), Nadler (2008), Kritchman and Nadler (2009) and Perry and Wolfe (2010).
In this paper, we analyze the capacity of statistical tests to detect a one-dimensional signal with the corresponding population covariance eigenvalue below the “impossibility threshold,” showing that the terminology “impossibility threshold” is overly pessimistic. We establish that the eigenvalue region below the threshold actually is the region of mutual contiguity [in the sense of Le Cam (1960)] of the joint distributions of the sample covariance eigenvalues under the null and under the alternative. We obtain the limit in distribution of the log likelihood ratio process inside this contiguity region and derive the asymptotic power envelope for sample-eigenvalue-based detection tests.
The power envelope is larger than size for local alternatives and monotonically tends to one as the signal’s population eigenvalue approaches the threshold from below. Hence, the detection of a signal with high asymptotic probability is quite possible even in cases where the largest population covariance eigenvalue is smaller than the threshold, especially when the distance from the threshold remains small.
In the contiguity region, the log likelihood ratio is asymptotically equivalent to a simple statistic related to the Stieltjes transform of the empirical distribution of the sample covariance eigenvalues. The reason the asymptotic behavior of this statistic differs under the null and under the alternative despite the apparent similarity of eigenvalue behaviors just mentioned is that it is not based merely on a contrast between the largest and the rest of the eigenvalues. The information about the presence of the signal exploited by this statistic is hidden in the small deviations of the empirical distribution of the eigenvalues from its Marchenko–Pastur limit.
Let us examine our setting and our results in more detail. Suppose that data consist of independent observations of -dimensional real-valued vectors distributed according to the Gaussian law with mean zero and covariance matrix , where is the -dimensional identity matrix, and are scalars and is a -dimensional vector with Euclidean norm one. We are interested in the asymptotic power of the tests of the null hypothesis against the alternative based on the eigenvalues of the sample covariance matrix of the data when both and go to infinity. The vector is an unspecified nuisance parameter indicating the direction of the perturbation of sphericity. In contrast to Berthet and Rigollet (2012), who study signal detection in a similar setting where the vector is sparse, we do not constrain in any way except normalizing its Euclidean norm to one.
We consider the cases of known and unknown . For the sake of brevity, in the rest of this Introduction, we discuss only the case of unknown , which, in practice, is also more relevant. Let be the th largest sample covariance eigenvalue, let be its normalized version and let , where . We begin our analysis with a study of the asymptotic properties of the likelihood ratio process defined as the ratio of the density of when to that when . We represent in the form of an integral over a contour in the complex plane and use the Laplace approximation method and recent results from the large random matrix theory to derive an asymptotic expansion of as so that , which we throughout abbreviate into .
We show that, for any such that , converges in distribution under the null to a Gaussian process on with
By Le Cam’s first lemma [see van der Vaart (1998), page 88], this implies that the joint distributions of the normalized sample covariance eigenvalues under the null and under the alternative are mutually contiguous for any . We also show that these joint distributions are not mutually contiguous for any .
Since , as a likelihood ratio process, is not of the LAN Gaussian shift type, local asymptotic normality does not hold, and the asymptotic optimality analysis of tests of against is difficult. However, an asymptotic power envelope is easy to construct using the Neyman–Pearson lemma along with Le Cam’s third lemma. We show that, for tests of asymptotic size , the maximum achievable power against a specific alternative is , where , as usual, denotes the standard normal distribution function.
Using our result on the limiting distribution of and Le Cam’s third lemma, we compute the asymptotic powers of several previously proposed tests of sphericity and of the likelihood ratio (LR) test based on . We find that the power of the LR test comes close to the asymptotic power envelope. The LR test outperforms the test proposed by John (1971) and studied in Ledoit and Wolf (2002), as well as Srivastava (2005) and the test proposed by Bai et al. (2009). The asymptotic powers of the tests based on the largest sample covariance eigenvalue, such as the tests proposed by Bejan (2005), Patterson, Price and Reich (2006), Kritchman and Nadler (2009), Onatski (2009), Bianchi et al. (2011) and Nadakuditi and Silverstein (2010), equals the tests’ asymptotic size for alternatives in the contiguity region.
The rest of the paper is organized as follows. Section 2 provides a representation of the likelihood ratio in terms of a contour integral. Section 3 applies Laplace’s method to obtain an asymptotic approximation to the contour integral. Section 4 uses that approximation to establish the convergence of the log likelihood ratio process to a Gaussian process. Section 5 provides an analysis of the asymptotic power of various sphericity tests and derives the asymptotic power envelope. Section 6 concludes. Proofs are given in the Appendix; the more technical ones are relegated to the Supplementary Appendix [Onatski, Moreira and Hallin (2013)].
Likelihood ratios as contour integrals
Let be a matrix with i.i.d. real Gaussian columns. Let be the ordered eigenvalues of and let , where . Finally, let , where .
As explained in the Introduction, our goal is to study the asymptotic power of the eigenvalue-based tests of against . If is known, the model is invariant with respect to orthogonal transformations, and the maximal invariant statistic is . Therefore, we consider tests based on . If is unknown (which, strictly speaking, is what is meant by “sphericity”), the model is invariant with respect to orthogonal transformations and multiplications by nonzero scalars, and the maximal invariant is . Hence, we consider tests based on . Note that the distribution of does not depend on , whereas if is known, we can always normalize dividing it by . Therefore, in what follows, we will assume that without loss of generality.
Let us denote the joint density of as and that of as . The following proposition gives explicit formulas for and .
where and depend only on and , and on and , respectively.
The spherical integrals in (1) and (1) can be represented in the form of a confluent hypergeometric function of matrix argument [Hillier (2001), page 4]. For example, for the integral in (1),
Butler and Wood (2002) develop Laplace approximations to functions but do not analyze the asymptotic behavior of the approximation errors. The next lemma derives an alternative representation of the spherical integrals in Proposition 1. This representation has the form of a contour integral of a single complex variable, and our asymptotic analysis will be based on the Laplace approximation to such an integral.
Let , where are arbitrary complex numbers. Further, let be a contour in the complex plane starting at
, encircling counter-clockwise the points , and going back to . Such a contour is shown in Figure 1. We have
Now, expanding the exponent in the latter expression into power series and taking expectations term by term yields
The Dirichlet average of is well studied. By Theorem 3.1 of Dickey (1983),
where is Pochhammer’s notation for the shifted factorial.
where the last equality is the definition of the confluent form of the Lauricella function, denoted as . The functions were introduced by Erdelyi (1937) and are discussed by Srivastava and Karlsson (1985). In probability and statistics, they were recently used to study the mean of a Dirichlet process [see Lijoi and Regazzini (2004) and references therein].
Erdelyi (1937), formula (8,6), establishes the following contour integral representation of :
Lemma 2 follows from equalities (2) and (2).
The contour integral representation given in Lemma 2 has been derived independently by Mo (2012) and Wang (2012), who use it to study the largest sample covariance eigenvalue when the corresponding population eigenvalue equals the critical threshold or lies above it. Our proof effectively takes advantage of old results of Dickey (1983) and Erdelyi (1937), and thus is different from the proofs in the above mentioned papers.
Using Lemma 2 and Proposition 1, we derive contour integral representations for the likelihood ratios and . The quantity is the likelihood ratio based on as opposed to the entire data . Similarly, is the likelihood ratio based on .
where and .
Close inspection of the proof of Lemma 3 reveals that the right-hand side of (3) depends on only through . Although it is possible to express as an explicit function of , the implicit form given in (3) is convenient because it allows us to use similar methods for the asymptotic analysis of the two likelihood ratios.
In the next two sections, we perform an asymptotic analysis of and that relies on the Laplace approximation of the contour integrals in Lemma 3 after those contours have been suitably deformed without changing the value of the integrals.
Laplace approximation
The contour integrals in (9) and (3) can be represented in the Laplace form with a deterministic function and a random function that converges to a log-normal random process on the contour as . To see this, note that the logarithm of the multiple product in (9) and (3) equals . For each , this expression is a special form of the linear spectral statistic studied by Bai and Silverstein (2004). According to the central limit theorem (Theorem 1.1) established in that paper, the random variable
converges in distribution to a normal random variable when . Here is the cumulative distribution function of the Marchenko–Pastur distribution with a mass of at zero and density
where , and .
Such a convergence suggests the following choices of and in the Laplace forms of the integrals in (9) and (3):
where the branch of the square root is chosen so that the real and the imaginary parts of have the same signs as the real and the imaginary parts of , respectively.
Substituting (16) into (15) and solving for when , we get
A proof of the following technical lemma is relegated to the Supplementary Appendix [Onatski, Moreira and Hallin (2013)].
Suppose that our null hypothesis is true, and let be any fixed number such that . Deforming contour into leaves the value of the integrals (9) and (3) in Lemma 3 unchanged for all with probability approaching one as .
We now derive, uniform (over ), Laplace approximations to the integrals (9) and (3) in Lemma 3. First, we introduce additional notation. When and are analytic at , let and with be the coefficients in the power series representations
When and are not analytic at , let the coefficients and be arbitrary numbers for all .
The following lemma is a generalization of the well-known Watson lemma for contour integrals; see Olver (1997), page 118. Theorem 7.1 in Olver (1997), page 127, derives a similar generalization for the case when and are fixed deterministic analytic functions. In contrast to Olver’s theorem, our lemma allows to be a random function, and to depend on parameter , and obtains a uniform approximation over . The proof is relegated to the Supplementary Appendix [Onatski, Moreira and Hallin (2013)].
Under the conditions of Lemma 4, for any and any positive integer , as , we have
where is uniform in . The coefficients in (21) can be expressed through and defined above. In particular, we have
Neither Lemma 5 nor Lemma 6 addresses interesting cases with in a neighborhood of . In such cases, would be close to the upper boundary of the support of the Marchenko–Pastur distribution. This may lead to the nonanalyticity of and on and a more complicated asymptotic behavior of . We leave the analysis of cases where may approach for future research.
Guionnet and Maïda (2005) study the asymptotic behavior of spherical integrals using large deviation techniques. Their Theorems 3 and 6 imply Lemma 6 and can be used to obtain the first term in the asymptotic expansion of Lemma 5.
Asymptotic behavior of the likelihood ratios
In this section, we discuss the asymptotic behavior of the likelihood ratios and . First, let us focus on the case where . In the Appendix, we use Lemmas 4 and 5 to derive the following theorem.
Suppose that the null hypothesis is true (). Let be any fixed number such that and let be the space of real-valued continuous functions on equipped with the supremum norm. Then as , , we have, uniformly in
Furthermore, and , viewed as random elements of converge weakly to and with Gaussian finite-dimensional distributions such that, for any ,
The log likelihood ratio processes studied in Theorem 7 are not of the standard locally asymptotically normal form. This is because they cannot be represented as , where and are some deterministic functions of , and is a standard normal random variable. Indeed, had the representation been possible, the covariance of the limiting log likelihood process at and would have been . Hence, for , for instance, we would have had and , which cannot be true for all and .
The quantity plays an important role in the limits of experiments. The likelihood ratio processes are well approximated by simple functions of and , which are easy to compute from the data and are asymptotically Gaussian by the central limit theorem of Bai and Silverstein (2004). Recalling the definition (11) of , we see that asymptotically, all statistical information about parameter is contained in the deviations of the sample covariance eigenvalues from . Although the latter limit does not have an obvious interpretation when , it is the probability limit of under alternatives with ; see, for example, Baik and Silverstein (2006).
Let us now consider cases where . We prove the following theorem in the Appendix.
Suppose that the null hypothesis is true (), and let be any fixed number such that . Then as , the following holds. For any , the likelihood ratios and converge to zero; more precisely, there exists that depends only on such that
Note that Theorem 7 and Le Cam’s first lemma [see van der Vaart (1998), page 88] imply that the joint distributions of (as well as those of ) under the null and under the alternative are mutually contiguous for any . In contrast, Theorem 8 shows that mutual contiguity is lost for . For such , consistent tests (as ) exist at any probability level .
In a similar setting, Nadakuditi and Edelman (2008) call the number of “signal eigenvalues” of the population covariance matrix that exceed the “effective number of identifiable signals” [see also Nadakuditi and Silverstein (2010)]. Theorems 7 and 8 shed light on the formal statistical content of this concept. The “identifiable signals” are detected with probability approaching one in large samples (irrespective of the probability level at which identification tests are performed). Other signals still can be detected, but the probability of detecting them will never approach one (whatever the probability level ).
Asymptotic power analysis
Theorem 7 can be used to study “local” powers of the tests for detecting signals in noise. The nonstandard form of the limit of log likelihood ratio processes in our setting makes it hard to develop tests with optimal local power properties. However, using the Neyman–Pearson lemma and Le Cam’s third lemma, we can analytically derive the local asymptotic power envelope and compare local asymptotic powers of specific tests to this envelope.
It is convenient to reparametrize our problem to . As varies in the region of contiguity , spans the entire half-line . Note that the asymptotic mean and autocovariance functions of the log likelihood ratios derived in the previous section depend on only through . Therefore, under the new parametrization, they depend only on . Loosely speaking, and play the classical roles of a “local parameter” and a contiguity rate, respectively.
Let and be the asymptotic powers of the asymptotically most powerful - and -based tests of size of the null against the alternative . The following proposition is proven in the Appendix.
Let denote the standard normal distribution function. Then
Plots of the asymptotic power envelopes and against for asymptotic size are shown in the left panel of Figure 3. The power loss of the -based tests relative to the -based tests is due to the nonspecification of . In contrast to -based tests, -based tests may achieve the corresponding power envelope even when is unknown.
The right panel of Figure 3 shows the envelopes as functions of the original parameter normalized by . We see that the alternatives that can theoretically be detected with high probability are concentrated near the threshold . The strong nonlinearity of the -parametrization should be kept in mind while interpreting the figures that follow.
Therefore, to numerically evaluate the asymptotic power function of the -based LR test, we simulate 500,000 observations of on a grid of 1000 equally spaced points in , where is chosen as the upper limit of the grid because it is large enough for the power envelopes to rich the value of 99%. For each observation, we save its supremum on the grid, and use the empirical distribution of two times the suprema as the approximate asymptotic distribution of the likelihood ratio statistic under the null. We denote this distribution as . Its 95% quantile equals .
The asymptotic powers of the LR and WAP tests both come close to the power envelope. The LR and WAP power functions are so close that they are difficult to distinguish clearly. The asymptotic power of the WAP test appears to be larger than that of the LR test for all in the $\theta_{1}$. Hence, the LR test still may be admissible. More accurate numerical analysis is needed to shed further light on this issue.
In the remaining part of this section, we consider some of the tests that have been proposed previously in the literature, and, in Proposition 10, derive their asymptotic power functions. We focus on four examples. Three of them are inspired by the “classical” fixed- theory, while the fourth is more directly based on results from the large random matrix theory.
The problem of testing the hypothesis of sphericity has a long history, and has generated a considerable body of literature, which we only very briefly summarize here. The classical fixed- Gaussian analysis of the various problems considered here goes back to Mauchly (1940), who first derived the Gaussian likelihood ratio test for sphericity. The (Gaussian) locally most powerful invariant (under shift, scale and orthogonal transformations) test was obtained by John (1971, 1972) and by Sugiura (1972), with adjusted versions resisting elliptical violations of the Gaussian assumptions proposed in Hallin and Paindaveine (2006), where a Le Cam approach is adopted under a general elliptical setting. Ledoit and Wolf (2002) propose two extensions (for the unknown and known scale problems, resp.) of John’s test, while Bai et al. (2009) adapt Mauchly’s (1940) likelihood ratio test.
John (1971) proposes testing the sphericity hypothesis against general alternatives using the test statistic , where is the sample covariance matrix of the data. He shows that, when , such a test is locally most powerful invariant. Studying John’s test when , Ledoit and Wolf (2002) prove that, under the null, . Hence, the test with asymptotic size rejects the null hypothesis of sphericity if .
Ledoit and Wolf (2002) propose using as a test statistic for testing the hypothesis that the population covariance matrix is a unit matrix. Under the null, . As in the previous example, the null hypothesis is rejected at asymptotic size if .
More directly inspired by the asymptotic theory of random matrices, several authors have recently proposed and studied various tests based on or : see Bejan (2005), Patterson, Price and Reich (2006), Kritchman and Nadler (2009), Onatski (2009), Bianchi et al. (2011) and Nadakuditi and Silverstein (2010). We refer to these tests, which reject for large values of or , as Tracy–Widom-type tests.
Asymptotic critical values of such tests are obtained using the fact, established by Johnstone (2001), that under the null,
where TW denotes the Tracy–Widom law of the first kind. The null hypothesis is rejected when or exceeds the adequate Tracy–Widom quantile.
Denote as . The asymptotic power functions of the tests described in Examples 1–4 satisfy, for any ,
With the important exception of Srivastava (2005), (34)–(36) are the first results on the asymptotic power of those tests against contiguous alternatives. Srivastava (2005) analyzes the asymptotic power of tests similar to those in Examples 1 and 2. His Theorems 3.1 and 4.1 can be used to establish (35).
From Proposition 10, we see that the local asymptotic power of the Tracy–Widom-type tests is trivial. As shown by Baik, Ben Arous and Péché (2005) in the complex data case and by Féral and Péché (2009) in the real data case, the convergence (33) holds not only under the null, but also under any alternative of the form . Under the “local” parametrization adopted in this section, such alternatives have the form . It can be shown that the Tracy–Widom-type tests are consistent against noncontiguous alternatives . However, such a consistency is likely to be also a property of the LR tests based on or on . If this holds true, the LR tests asymptotically dominate the Tracy–Widom-type tests. A more detailed analysis of the optimality properties of LR tests is the subject of ongoing research.
The left panel of Figure 5 shows that the power function of John’s test is very close to the power envelope in the vicinity of . Such behavior is consistent with the fact that John’s test is locally most powerful invariant. However, for large , the asymptotic power functions of all the tests from Examples 1, 2 and 3 are lower than the corresponding asymptotic power envelopes. We should stress here that these tests have power against general alternatives as opposed to the “spiked” alternatives that maintain the assumption that the population covariance matrix of data has the form .
For the “spiked” alternatives, the - and -based LR tests may be more attractive. However, implementing these tests requires some care. A “quick-and-dirty” approach would be to approximate and by the simple but asymptotically equivalent expressions from (24) and (25), compute two times their maxima on a grid over , and compare them with critical values obtained by simulation as for the construction of Figure 4. Unfortunately, in finite samples, this simple approach will lead to a numerical breakdown whenever happens to be less than the largest sample covariance eigenvalue for some . In addition, since the asymptotic approximation derived in Theorem 7 is not uniform over entire half-line , its quality will depend on the choice of . For relatively large , the asymptotic behavior of the LR test implemented as above may poorly match its finite sample behavior.
Instead, we recommend implementing the LR tests without using the asymptotic approximations. The finite sample log likelihood ratios and can be computed using the contour integral representations (9) and (3). Choosing the contour of integration so that the sample covariance eigenvalues remain to its left will eliminate the numerical breakdown problem associated with the asymptotic tests. Furthermore, under the Gaussianity assumption, the finite sample distributions of the log likelihood ratios are pivotal. Hence, the exact critical values can be computed via Monte Carlo simulations as follows: simulate many replications of data under the null. For each replication, compute the log likelihood ratio and store two times its maximum. Use the 95% quantile of the empirical distribution of the stored values as a numerical approximation for the exact critical value of the test. The finite sample properties of such a test are left as an important topic for future research.
Conclusion
In this paper, we study the asymptotic power of tests for the existence of rank-one perturbations of sphericity as both the dimensionality of the data and the number of observations go to infinity. Focusing on tests that are invariant with respect to orthogonal transformations and rescaling, we establish the convergence of the log ratio of the joint densities of the sample covariance eigenvalues under the alternative and null hypotheses to a Gaussian process indexed by the norm of the perturbation.
When the perturbation norm is larger than the phase transition threshold studied in Baik, Ben Arous and Péché (2005), the limiting log-likelihood process is degenerate and the joint eigenvalue distributions under the null and alternative hypotheses are asymptotically mutually singular, so that the discrimination between the null and the alternative is asymptotically certain. When the norm is below the threshold, the limiting log-likelihood process is nondegenerate and the joint eigenvalue distributions under the null and alternative hypotheses are mutually contiguous. Using the asymptotic theory of statistical experiments, we obtain power envelopes and derive the asymptotic size and power for various eigenvalue-based tests in the region of contiguity.
Several questions are left for future research. First, we only considered rank-one perturbations of the spherical covariance matrices. It would be desirable to extend the analysis to finite-rank perturbations. Such an extension will require a more complicated technical analysis. Second, it would be interesting to extend our analysis to the asymptotic regime with or . In the context of sphericity tests, such asymptotic regimes have been recently studied in Birke and Dette (2005). Third, a thorough analysis of the finite sample properties of the proposed LR tests would clarify the related practical implementation issues. Fourth, our Lemma 5 can be used to derive higher-order asymptotic approximations to the likelihood ratios, which may improve finite-sample performances of asymptotic tests. Finally, it would be of considerable interest to relax the Gaussian assumptions, for example, into elliptical ones, preferably with unspecified radial densities, on the model (in a fixed- context) of Hallin and Paindaveine (2006).
Appendix
For the joint density of , we have
Let be a matrix. Since , we have , and we can rewrite (.1) as
Note that , where is the first column of . When is uniformly distributed over , its first column is uniformly distributed over . Therefore, we have
which establishes (1). Now, let so that . Note that , , and that the Jacobian of the coordinate change from to equals . Changing variables in (.1), and integrating out, we obtain (1).
.2 Proof of Lemma 3
Using (3) in the ratio of the right-hand side of (1) with to that with , and changing the variable of integration from to , we get (9). Further, from (1), we have
where is a contour starting at , encircling counter-clockwise the points , , and going back to . In addition, for any , . Such a choice of guarantees that the integrand in the above double integral is absolutely integrable on , so that Fubini’s theorem can be used to justify the interchange of the order of the integrals. Changing the order of the integrals and setting , we obtain (3).
.3 Proof of Theorem 7
First, let us formulate the following technical lemma. Its proof is in the Supplementary Appendix [Onatski, Moreira and Hallin (2013)].
(i) If ,
(ii) If ,
Below, we prove Theorem 7 for . The proof for is similar but simpler, and we omit it to save space. As follows from Lemmas 4 and 5, the integral in (3) can be represented as uniformly in . Therefore, and since , we can write
where . Using Stirling’s approximation with , and , and the fact that , we find, after algebraic simplifications, that
which, together with the fact that , implies that
Now, as can be verified using (13) and (16), if , then
Using (14), (.3), (45) and Lemma 11(i) in (41), after algebraic simplifications and rearrangements of terms, we get
Finally, using the fact that , we obtain and
The latter two equalities, (.3) and the fact that entail
which, together with (43), imply formula (25).
Now, let us prove the convergence of to . By (25), the joint convergence of with to a Gaussian vector is equivalent to the convergence of to a Gaussian vector. A proof of the following technical lemma, based on Theorem 1.1 of Bai and Silverstein (2004), is given in the Supplementary Appendix [Onatski, Moreira and Hallin (2013)].
Suppose that the null hypothesis holds. Then, as , the vector converges in distribution to a Gaussian vector with
To complete the proof of Theorem 7, we need to note that the tightness of , viewed as a random element of the space , as , follows from formula (25) and the fact that and , are , uniformly in . This uniformity is a consequence of Lemma A2 proven in the Supplementary Appendix [Onatski, Moreira and Hallin (2013)].
.4 Proof of Theorem 8
.5 Proof of Proposition 9
For brevity, we derive only the asymptotic power envelope for the case of -based tests. According to the Neyman–Pearson lemma, the most powerful test of the null against a particular alternative is the test which rejects the null when is larger than some critical value . It follows from Theorem 7 that, for such a test to have asymptotic size , must be , where and are obtained from (28) and (29) by the re-parametrization . Now, according to Le Cam’s third lemma and Theorem 7, under , . Therefore, the asymptotic power of the asymptotically most powerful test of against is (32).
.6 Proof of Proposition 10
As shown by Baik, Ben Arous and Péché (2005) in the complex case and by Féral and Péché (2009) in the real case, the convergence (33) takes place not only under the null, but also under alternatives with , yielding under the parametrization . Hence, (34) follows.
Formulas (35) and (36) can be established using conceptually similar steps. To save space, below we only establish formula (36). The following technical lemma is proven in the Supplementary Appendix [Onatski, Moreira and Hallin (2013)].
Acknowledgments
This work started when the first two authors worked at and the third author visited Columbia University. We would like to thank Tony Cai, the Associate Editor, Nick Patterson and an anonymous referee for helpful and encouraging comments.
Supplementary Appendix \slink[doi]10.1214/13-AOS1100SUPP \sdatatype.pdf \sfilenameaos1100_supp.pdf \sdescriptionThe Supplementary Appendix contains proofs of Lemmas 4, 5, 6, 11, 12 and 13.