Large Covariance Estimation by Thresholding Principal Orthogonal Complements

Jianqing Fan, Yuan Liao, Martina Mincheva

Introduction

Information and technology make large data sets widely available for scientific discovery. Much statistical analysis of such high-dimensional data involves the estimation of a covariance matrix or its inverse (the precision matrix). Examples include portfolio management and risk assessment (Fan, Fan and Lv, 2008), high-dimensional classification such as Fisher discriminant (Hastie, Tibshirani and Friedman, 2009), graphic models (Meinshausen and Bühlmann, 2006), statistical inference such as controlling false discoveries in multiple testing (Leek and Storey, 2008; Efron, 2010), finding quantitative trait loci based on longitudinal data (Yap, Fan, and Wu, 2009; Xiong et al. 2011), and testing the capital asset pricing model (Sentana, 2009), among others. See Section 5 for some of those applications. Yet, the dimensionality is often either comparable to the sample size or even larger. In such cases, the sample covariance is known to have poor performance (Johnstone, 2001), and some regularization is needed.

Realizing the importance of estimating large covariance matrices and the challenges brought by the high dimensionality, in recent years researchers have proposed various regularization techniques to consistently estimate Σ\mathbf{\Sigma}. One of the key assumptions is that the covariance matrix is sparse, namely, many entries are zero or nearly so (Bickel and Levina, 2008, Rothman et al, 2009, Lam and Fan 2009, Cai and Zhou, 2010, Cai and Liu, 2011). In many applications, however, the sparsity assumption directly on Σ\mathbf{\Sigma} is not appropriate. For example, financial returns depend on the equity market risks, housing prices depend on the economic health, gene expressions can be stimulated by cytokines, among others. Due to the presence of common factors, it is unrealistic to assume that many outcomes are uncorrelated. An alternative method is to assume a factor model structure, as in Fan, Fan and Lv (2008). However, they restrict themselves to the strict factor models with known factors.

A natural extension is the conditional sparsity. Given the common factors, the outcomes are weakly correlated. In order to do so, we consider an approximate factor model, which has been frequently used in economic and financial studies (Chamberlain and Rothschild, 1983; Fama and French 1993; Bai and Ng, 2002, etc):

We emphasize that in model (1.1), only yity_{it} is observable. It is intuitively clear that the unknown common factors can only be inferred reliably when there are sufficiently many cases, that is, p→∞p\to\infty. In a data-rich environment, pp can diverge at a rate faster than TT. The factor model (1.1) can be put in a matrix form as

does not grow too fast as p→∞.p\rightarrow\infty. In particular, this includes the exact sparsity assumption (q=0q=0) under which mp=max⁡i≤p∑j≤pI(σu,ij≠0)m_{p}=\max_{i\leq p}\sum_{j\leq p}I_{(\sigma_{u,ij}\neq 0)}, the maximum number of nonzero elements in each row.

This paper focuses on the high-dimensional static factor model (1.2), which is innately related to the principal component analysis (PCA), as clarified in Section 2. This feature makes it different from the classical factor model with fixed dimensionality (e.g., Lawley and Maxwell 1971). In the last ten years, much theory on the estimation and inference of the static factor model has been developed, for example, Stock and Watson (1998, 2002), Bai and Ng (2002), Bai (2003), Doz, Giannone and Reichlin (2011), among others. Our contribution is on the estimation of covariance matrices and their inverse in large factor models.

There has been extensive literature in recent years that deals with sparse principal components, which has been widely used to enhance the convergence of the principal components in high-dimensional space. d’Aspremont, Bach and El Ghaoui (2008), Shen and Huang (2008), Witten, Tibshirani, and Hastie (2009) and Ma (2011) proposed and studied various algorithms for computations. More literature on sparse PCA is found in Johnstone and Lu (2009), Amini and Wainwright (2009), Zhang and El Ghaoui (2011), Birnbaum et al. (2012), among others. In addition, there has also been a growing literature that theoretically studies the recovery from a low-rank plus sparse matrix estimation problem, see for example, Wright et al. (2009), Lin et al. (2009), Candès et al. (2011), Luo (2011), Agarwal, Nagahban, Wainwright (2012), Pati et al. (2012). It corresponds to the identifiability issue of our problem.

There is a big difference between our model and those considered in the aforementioned literature. In the current paper, the first KK eigenvalues of Σ\mathbf{\Sigma} are spiked and grow at a rate O(p)O(p), whereas the eigenvalues of the matrices studied in the existing literature on covariance estimation are usually assumed to be either bounded or slowly growing. Due to this distinctive feature, the common components and the idiosyncratic components can be identified, and in addition, PCA on the sample covariance matrix can consistently estimate the space spanned by the eigenvectors of Σ\mathbf{\Sigma}. The existing methods of either thresholding directly or solving a constrained optimization method can fail in the presence of very spiked principal eigenvalues. However, there is a price to pay here: as the first KK eigenvalues are “too spiked”, one can hardly obtain a satisfactory rate of convergence for estimating Σ\mathbf{\Sigma} in absolute term, but it can be estimated accurately in relative term (see Section 3.3 for details). In addition, Σ−1\mathbf{\Sigma}^{-1} can be estimated accurately.

We would like to further note that the low-rank plus sparse representation of our model is on the population covariance matrix, whereas Candès et al. (2011), Wright et al. (2009), Lin et al. (2009)We thank a referee for reminding us these related works. considered such a representation on the data matrix. As there is no Σ\mathbf{\Sigma} to estimate, their goal is limited to producing a low-rank plus sparse matrix decomposition of the data matrix, which corresponds to the identifiability issue of our study, and does not involve estimation and inference. In contrast, our ultimate goal is to estimate the population covariance matrices as well as the precision matrices. For this purpose, we require the idiosyncratic components and common factors to be uncorrelated and the data generating process to be strictly stationary. The covariances considered in this paper are constant over time, though slow-time-varying covariance matrices are applicable through localization in time (time-domain smoothing). Our consistency result on Σu\mathbf{\Sigma}_{u} demonstrates that the decomposition (1.3) is identifiable, and hence our results also shed the light of the “surprising phenomenon” of Candès et al. (2011) that one can separate fully a sparse matrix from a low-rank matrix when only the sum of these two components is available.

Regularized Covariance Matrix via PCA

We illustrate this in the following example.

Let {λj}j=1p\{\lambda_{j}\}_{j=1}^{p} be the eigenvalues of Σ\Sigma in a descending order and \{\mbox{\boldmath\xi}_{j}\}_{j=1}^{p} be their corresponding eigenvectors. Then, an application of Weyl’s eigenvalue theorem (see the appendix) yields that

Using Proposition 2.1 and the sin⁡θ\sin\theta theorem of Davis and Kahn (1970, see the appendix), we have the following:

Propositions 2.1 and 2.2 state that PCA and factor analysis are approximately the same if \|\mbox{\boldmath\Sigma}_{u}\|=o(p). This is assured through a sparsity condition on Σu=(σu,ij)p×p\mathbf{\Sigma}_{u}=(\sigma_{u,ij})_{p\times p}, which is frequently measured through

The intuition is that, after taking out the common factors, many pairs of the cross-sectional units become weakly correlated. This generalized notion of sparsity was used in Bickel and Levina (2008) and Cai and Liu (2011). Under this generalized measure of sparsity, we have

if the noise variances {σu,ii2}\{\sigma_{u,ii}^{2}\} are bounded. Therefore, when mp=o(p)m_{p}=o(p), Proposition 2.1 implies that we have distinguished eigenvalues between the principal components {λj}j=1K\{\lambda_{j}\}_{j=1}^{K} and the rest of the components {λj}j=K+1p\{\lambda_{j}\}_{j=K+1}^{p} and Proposition 2.2 ensures that the first KK principal components are approximately the same as the columns of the factor loadings.

The aforementioned sparsity assumption appears reasonable in empirical applications. Boivin and Ng (2006) conducted an empirical study and showed that imposing zero correlation between weakly correlated idiosyncratic components improves forecastWe thank a referee for this interesting reference.. More recently, Phan (2012) empirically estimated the level of sparsity of the idiosyncratic covariance using the UK market data.

Recent developments on random matrix theory, for example, Johnstone and Lu (2009) and Paul (2007), have shown that when p/Tp/T is not negligible, the eigenvalues and eigenvectors of Σ\mathbf{\Sigma} might not be consistently estimated from the sample covariance matrix. A distinguished feature of the covariance considered in this paper is that there are some very spiked eigenvalues. By Propositions 2.1 and 2.2, in the factor model, the pervasiveness condition

With assumption (2.3), the standard literature on approximate factor models has shown that the PCA on the sample covariance matrix Σ^sam\widehat{\mathbf{\Sigma}}_{{}_{\text{sam}}} can consistently estimate the space spanned by the factor loadings (e.g., Stock and Watson (1998), Bai (2003)). Our contribution in Propositions 2.1 and 2.2 is that we connect the high-dimensional factor model to the principal components, and obtain the consistency of the spectrum in the population level Σ\mathbf{\Sigma} instead of the sample level Σ^sam\widehat{\mathbf{\Sigma}}_{{}_{\text{sam}}}. The spectral consistency also enhances the results in Chamberlain and Rothschild (1983). This provides the rationale behind the consistency results in the factor model literature.

2 POET

Sparsity assumption directly on Σ\mathbf{\Sigma} is inappropriate in many applications due to the presence of common factors. Instead, we propose a nonparametric estimator of Σ\mathbf{\Sigma} based on the principal component analysis. Let λ^1≥λ^2≥⋯≥λ^p\widehat{\lambda}_{1}\geq\widehat{\lambda}_{2}\geq\cdots\geq\widehat{\lambda}_{p} be the ordered eigenvalues of the sample covariance matrix \widehat{\mbox{\boldmath\Sigma}}_{{}_{\text{sam}}} and \{\widehat{}\mbox{\boldmath\xi}_{i}\}_{i=1}^{p} be their corresponding eigenvectors. Then the sample covariance has the following spectral decomposition:

where sij(⋅)s_{ij}(\cdot) is a generalized shrinkage function of Antoniadis and Fan (2001), employed by Rothman et al. (2009) and Cai and Liu (2011), and τij>0\tau_{ij}>0 is an entry-dependent threshold. In particular, the hard-thresholding rule sij(x)=xI(∣x∣≥τij)s_{ij}(x)=xI(|x|\geq\tau_{ij}) (Bickel and Levina, 2008) and the constant thresholding parameter τij=δ\tau_{ij}=\delta are allowed. In practice, it is more desirable to have τij\tau_{ij} be entry-adaptive. An example of the adaptive thresholding is

The estimator of Σ\mathbf{\Sigma} is then defined as:

We will call this estimator the Principal Orthogonal complEment thresholding (POET) estimator. It is obtained by thresholding the remaining components of the sample covariance matrix, after taking out the first KK principal components. One of the attractiveness of POET is that it is optimization-free, and hence is computationally appealing. We have written an R package for POET, which outputs the estimated Σ\mathbf{\Sigma}, Σu\mathbf{\Sigma}_{u}, KK, the factors and loadings.

With the choice of τij\tau_{ij} in (2.6) and the hard thresholding rule, our estimator encompasses many popular estimators as its specific cases. When τ=0\tau=0, the estimator is the sample covariance matrix and when τ=1\tau=1, the estimator becomes that based on the strict factor model (Fan, Fan, and Lv , 2008). When K=0K=0, our estimator is the same as the thresholding estimator of Bickel and Levina (2008) and (with a more general thresholding function) Rothman et al. (2009) or the adaptive thresholding estimator of Cai and Liu (2011) with a proper choice of τij\tau_{ij}.

In practice, the number of diverging eigenvalues (or common factors) can be estimated based on the sample covariance matrix. Determining KK in a data-driven way is an important topic, and is well understood in the literature. We will describe the POET with a data-driven KK in Section 2.4.

3 Least squares point of view

The constraints (2.9) correspond to the normalization (2.1). Here we assume that the mean of each variable {yit}t=1T\{y_{it}\}_{t=1}^{T} has been removed, that is, Eyit=Efjt=0Ey_{it}=Ef_{jt}=0 for all i≤p,j≤Ki\leq p,j\leq K and t≤T.t\leq T. Putting it in a matrix form, the optimization problem can be written as

It is easy to verify that sij(⋅)s_{ij}(\cdot) includes many interesting thresholding functions such as the hard thresholding (sij(z)=zI(∣z∣≥τij)s_{ij}(z)=zI_{(|z|\geq\tau_{ij})}), soft thresholding (sij(z)=sign(z)(∣z∣−τij)+s_{ij}(z)=\text{sign}(z)(|z|-\tau_{ij})_{+}), SCAD, and adaptive lasso (See Rothman et al. (2009)).

Analogous to the decomposition (1.3), we obtain the following substitution estimators

Suppose that the entry-dependent threshold in (2.5) is the same as the thresholding parameter used in (2.11). Then for any K1≤pK_{1}\leq p, the estimator (2.7) is equivalent to the substitution estimator (2.12), that is,

In this paper, we will use a data-driven K^\widehat{K} to construct the POET (see Section 2.4 below), which has two equivalent representations according to Theorem 2.1.

4 POET with Unknown K𝐾K

Determining the number of factors in a data-driven way has been an important research topic in the econometric literature. Bai and Ng (2002) proposed a consistent estimator as both pp and TT diverge. Other recent criteria are proposed by Kapetanios (2010), Onatski (2010), Alessi et al. (2010), etc.

Our method also allows a data-driven K^\widehat{K} to estimate the covariance matrices. In principle, any procedure that gives a consistent estimate of KK can be adopted. In this paper we apply the well-known method in Bai and Ng (2002). It estimates KK by

Throughout the paper, we let K^\widehat{K} be the solution to (2.14) using either IC1 or IC2. The asymptotic results are not affected regardless of the specific choice of g(T,p)g(T,p). We define the POET estimator with unknown KK as

The procedure is as stated in Section 2.2 except that K^\widehat{K} is now data-driven.

Asymptotic Properties

The first assumption has been one of the most essential ones in the literature of approximate factor models. Under this assumption and other regularity conditions, the number of factors, loadings and common factors can be consistently estimated (e.g., Stock and Watson (1998, 2002), Bai and Ng (2002), Bai (2003), etc.).

It implies from Proposition 2.1 in Section 2 that the first KK eigenvalues of Σ\mathbf{\Sigma} grow at rate O(p)O(p). This unique feature distinguishes our work from most of other low-rank plus sparse covariances considered in the literature, e.g., Luo (2011), Pati et al. (2012), Agarwal et al. (2012), Birnbaum et al. (2012). To our best knowledge, the only other papers that estimate large covariances with diverging eigenvalues (growing at the rate of dimensionality O(p)O(p)) are Fan et al. (2008, 2011) and Bai and Shi (2011). While Fan et al. (2008, 2011) assumed the factors are observable, Bai and Shi (2011) considered the strict factor model in which Σu\mathbf{\Sigma}_{u} is diagonal.

As to be illustrated in Section 3.3 below, due to the fast diverging eigenvalues, one can hardly achieve a good rate of convergence for estimating Σ\mathbf{\Sigma} under either the spectral norm or Frobenius norm when p>Tp>T. This phenomenon arises naturally from the characteristics of the high-dimensional factor model, which is another distinguished feature compared to those convergence results in the existing literature.

In addition, we impose the following regularity conditions.

These conditions are needed to consistently estimate the transformed common factors as well as the factor loadings. Similar conditions were also assumed in Bai (2003), and Bai and Ng (2006). The number of factors is assumed to be fixed. Our conditions in Assumption 3.4 are weaker than those in Bai (2003) as we focus on different aspects of the study.

2 Convergence of the idiosyncratic covariance

Throughout the paper, we apply the adaptive threshold

where C>0C>0 is a sufficiently large constant, though the results hold for other types of thresholding. As in Bickel and Levina (2008) and Cai and Liu (2011), the threshold chosen in the current paper is in fact obtained from the optimal uniform rate of convergence of max⁡i≤p,j≤p∣σ^ij−σu,ij∣.\max_{i\leq p,j\leq p}|\widehat{\sigma}_{ij}-\sigma_{u,ij}|. When direct observation of uitu_{it} is not available, the effect of estimating the unknown factors also contributes to this uniform estimation error, which is why p−1/2p^{-1/2} appears in the threshold.

The following theorem gives the rate of convergence of the estimated idiosyncratic covariance. Let γ−1=3r1−1+1.5r2−1+r3−1+1\gamma^{-1}=3r_{1}^{-1}+1.5r_{2}^{-1}+r_{3}^{-1}+1. In the convergence rate below, recall that mpm_{p} and qq are defined in the measure of sparsity (2.2).

Suppose log⁡p=o(Tγ/6)\log p=o(T^{\gamma/6}), T=o(p2)T=o(p^{2}), and Assumptions 3.1-3.4 hold. Then for a sufficiently large constant C>0C>0 in the threshold (3.2), the POET estimator Σ^u,K^T\widehat{\mathbf{\Sigma}}_{u,\widehat{K}}^{\mathcal{T}} satisfies

If further ωT1−qmp=o(1)\omega_{T}^{1-q}m_{p}=o(1), then the eigenvalues of Σ^u,K^T\widehat{\mathbf{\Sigma}}_{u,\widehat{K}}^{\mathcal{T}} are all bounded away from zero with probability approaching one, and

When estimating Σu\mathbf{\Sigma}_{u}, pp is allowed to grow exponentially fast in TT, and Σ^u,K^T\widehat{\mathbf{\Sigma}}_{u,\widehat{K}}^{\mathcal{T}} can be made consistent under the spectral norm. In addition, Σ^u,K^T\widehat{\mathbf{\Sigma}}_{u,\widehat{K}}^{\mathcal{T}} is asymptotically invertible while the classical sample covariance matrix based on the residuals is not when p>T.p>T.

Consistent estimation of Σu\mathbf{\Sigma}_{u} indicates that Σu\mathbf{\Sigma}_{u} is identifiable in (1.3), namely, the sparse Σu\mathbf{\Sigma}_{u} can be separated perfectly from the low-rank matrix there. The result here gives another proof (when assuming ωT1−qmp=o(1)\omega_{T}^{1-q}m_{p}=o(1)) of the “surprising phenomenon” in Candès et al (2011) under different technical conditions.

When KK is known and grows with pp and TT, with slightly weaker assumptions, our working paper (Fan et al. 2011) shows that under the exactly sparse case (that is, q=0q=0), the result continues to hold with convergence rate mp(K2log⁡pT+K3p)m_{p}(K^{2}\sqrt{\frac{\log p}{T}}+\frac{K^{3}}{\sqrt{p}}).

3 Convergence of the POET estimator

Since the first KK eigenvalues of Σ\mathbf{\Sigma} grow with pp, one can hardly estimate Σ\mathbf{\Sigma} with satisfactory accuracy in the absolute term. This problem arises not from the limitation of any estimation method, but is due to the nature of the high-dimensional factor model. We illustrate this using a simple example.

Consider an ideal case where we know the spectrum except for the first eigenvector of Σ\mathbf{\Sigma}. Let \{\lambda_{j},\mbox{\boldmath\xi}_{j}\}_{j=1}^{p} be the eigenvalues and vectors, and assume that the largest eigenvalue λ1≥cp\lambda_{1}\geq cp for some c>0c>0. Let \widehat{}\mbox{\boldmath\xi}_{1} be the estimated first eigenvector and define the covariance estimator \widehat{\mathbf{\Sigma}}=\lambda_{1}\widehat{}\mbox{\boldmath\xi}_{1}\widehat{}\mbox{\boldmath\xi}_{1}^{\prime}+\sum_{j=2}^{p}\lambda_{j}\mbox{\boldmath\xi}_{j}\mbox{\boldmath\xi}_{j}^{\prime}. Assume that \widehat{}\mbox{\boldmath\xi}_{1} is a good estimator in the sense that \|\widehat{}\mbox{\boldmath\xi}_{1}-\mbox{\boldmath\xi}_{1}\|^{2}=O_{p}(T^{-1}). However,

which can diverge when T=O(p2)T=O(p^{2}). □\square

In the presence of very spiked eigenvalues, while the covariance Σ\mathbf{\Sigma} cannot be consistently estimated in absolute term, it can be well estimated in terms of the relative error matrix

which is more relevant for many applications (see Example 5.2). The relative error matrix can be measured by either its spectral norm or the normalized Frobenius norm defined by

In the last equality, there are pp terms being added in the trace operation and the factor p−1p^{-1} plays the role of normalization. The loss (3.3) is closely related to the entropy loss, introduced by James and Stein (1961). Also note that

Compared to estimating Σ\mathbf{\Sigma}, in a large approximate factor model, we can estimate the precision matrix with a satisfactory rate under the spectral norm. The intuition follows from the fact that Σ−1\mathbf{\Sigma}^{-1} has bounded eigenvalues.

The following theorem summarizes the rate of convergence under various norms.

Under the assumptions of Theorem 3.1, the POET estimator defined in (2.15) satisfies

In addition, if mpωT1−q=o(1)m_{p}\omega_{T}^{1-q}=o(1), then Σ^K^\widehat{\mathbf{\Sigma}}_{\widehat{K}} is nonsingular with probability approaching one, with

When estimating Σ−1\mathbf{\Sigma}^{-1}, pp is allowed to grow exponentially fast in TT, and the estimator has the same rate of convergence as that of the estimator Σ^u,K^T\widehat{\mathbf{\Sigma}}_{u,\widehat{K}}^{\mathcal{T}} in Theorem 3.1. When pp becomes much larger than TT, the precision matrix can be estimated at the same rate as if the factors were observable.

As in Remark 3.2, when K>0K>0 is known and grows with pp and TT, the working paper Fan et al. (2011) proves the following results (when q=0q=0) The assumptions in the working paper Fan et al. (2011) are slightly weak than those presented here, in that it required λmax⁡(Σu)\lambda_{\max}(\mathbf{\Sigma}_{u}) instead of ∥Σu∥1\|\mathbf{\Sigma}_{u}\|_{1} be bounded.:

The results state explicitly the dependence of the rate of convergence on the number of factors.

4 Convergence of unknown factors and factor loadings

As a consequence of Theorem 3.3, we obtain the following: (recall that the constant r2r_{2} is defined in Assumption 3.2.)

Choice of Threshold

Recall that the threshold value τij=Cθ^ijωT\tau_{ij}=C\sqrt{\hat{\theta}_{ij}}\omega_{T}, where CC is determined by the users. To make POET operational in practice, one has to choose CC to maintain the positive definiteness of the estimated covariances for any given finite sample. We write Σ^u,K^T(C)=Σ^u,K^T\widehat{\mathbf{\Sigma}}_{u,\widehat{K}}^{\mathcal{T}}(C)=\widehat{\mathbf{\Sigma}}_{u,\widehat{K}}^{\mathcal{T}}, where the covariance estimator depends on CC via the threshold. We choose CC in the range where λmin⁡(Σ^u,K^T)>0\lambda_{\min}(\widehat{\mathbf{\Sigma}}_{u,\widehat{K}}^{\mathcal{T}})>0. Define

When CC is sufficiently large, the estimator becomes diagonal, while its minimum eigenvalue must retain strictly positive. Thus, Cmin⁡C_{\min} is well defined and for all C>Cmin⁡C>C_{\min}, Σ^u,K^T(C)\widehat{\mathbf{\Sigma}}_{u,\widehat{K}}^{\mathcal{T}}(C) is positive definite under finite sample. We can obtain Cmin⁡C_{\min} by solving λmin⁡(Σ^u,K^T(C))=0,C≠0.\lambda_{\min}(\widehat{\mathbf{\Sigma}}_{u,\widehat{K}}^{\mathcal{T}}(C))=0,C\neq 0. We can also approximate Cmin⁡C_{\min} by plotting λmin⁡(Σ^u,K^T(C))\lambda_{\min}(\widehat{\mathbf{\Sigma}}_{u,\widehat{K}}^{\mathcal{T}}(C)) as a function of CC, as illustrated in Figure 1. In practice, we can choose CC in the range (Cmin⁡+ϵ,M)(C_{\min}+\epsilon,M) for a small ϵ\epsilon and large enough M.M. Choosing the threshold in a range to guarantee the finite-sample positive definiteness has also been previously suggested by Fryzlewicz (2012).

2 Multifold Cross-Validation

It is possible to derive the rate of convergence for Σ^u,K^T(C∗)\widehat{\mathbf{\Sigma}}_{u,\widehat{K}}^{\mathcal{T}}(C^{*}) under the current model setting, but it ought to be much more technically involved than the regular sparse matrix estimation considered by Bickel and Levina (2008) and Cai and Liu (2011). To keep our presentation simple we do not pursue it in the current paper.

Applications of POET

We give four examples to which the results in Theorems 3.1–3.3 can be applied. Detailed pursuits of these are beyond the scope of the paper.

Controlling the false discovery rate in large-scale hypothesis testing based on correlated test statistics is an important and challenging problem in statistics (Leek and Storey, 2008; Efron, 2010; Fan, et al., 2012). Suppose that the test statistic for each of the hypothesis

If the covariance Σ\Sigma admits the approximate factor structure (1.3), then the test statistics can be stochastically decomposed as

By the principal factor approximation (Theorem 1, Fan, Han, Gu, 2012)

Now suppose that we have nn repeated measurements from the model (5.1). Then, by Corollary 3.1, {ηi}\{\eta_{i}\} can be uniformly consistently estimated, and hence p−1V(x)p^{-1}V(x) and \mboxFDP(x)\mbox{FDP}(x) can be consistently estimated. Efron (2010) obtained these repeated test statistics based on the bootstrap sample from the original raw data. Our theory (Theorem 3.3) gives a formal justification to the framework of Efron (2007, 2010).

The maximum elementwise estimation error ∥Σ^K^−Σ∥max⁡\|\widehat{\mathbf{\Sigma}}_{\widehat{K}}-\mathbf{\Sigma}\|_{\max} appears in risk assessment as in Fan, Zhang and Yu (2012). For a fixed portfolio allocation vector w, the true portfolio variance and the estimated one are given by \mboxw′Σ\mboxw\mbox{\bf w}^{\prime}\mathbf{\Sigma}\mbox{\bf w} and \mboxw′Σ^K^\mboxw\mbox{\bf w}^{\prime}\widehat{\mathbf{\Sigma}}_{\widehat{K}}\mbox{\bf w} respectively. The estimation error is bounded by

where ∥\mboxw∥1\|\mbox{\bf w}\|_{1}, the L1L_{1}-norm of w, is the gross exposure of the portfolio. Usually a constraint is placed on the total percentage of the short positions, in which case we have a restriction ∥\mboxw∥1≤c\|\mbox{\bf w}\|_{1}\leq c for some c>0.c>0. In particular, c=1c=1 corresponds to a portfolio with no-short positions (all weights are nonnegative). Theorem 3.2 quantifies the maximum approximation error.

The above compares the absolute error of perceived risk and true risk. The relative error is bounded by

for any allocation vector w. Theorem 3.2 quantifies this relative error.

Consider the following panel regression model

and wishes to test H_{0}:\mbox{\boldmath\alpha}=0. The F-test statistic involves the estimation of the covariance matrix Σu\mathbf{\Sigma}_{u}, whose estimates are degenerate without regularization when p≥Tp\geq T. Therefore, in the literature (Sentana, 2009, and references therein), one focuses on the case pp is relatively small. The typical choices of parameters are T=60T=60 monthly data and the number of assets p=5p=5, 10 or 25. However, the CAPM should hold for all tradeable assets, not just a small fraction of assets. With our regularization technique, non-degenerate estimate Σ^u,K^T\widehat{\mathbf{\Sigma}}_{u,\widehat{K}}^{\mathcal{T}} can be obtained and the F-test or likelihood-ratio test statistics can be employed even when p≫Tp\gg T.

Since under the null hypothesis \hat{\mbox{\boldmath\alpha}}\sim N(0,\mathbf{\Sigma}_{u}/c_{T}), we have c_{T}\|\mathbf{\Sigma}_{u}^{-1/2}\hat{\mbox{\boldmath\alpha}}\|^{2}=O(p). Thus, it follows from boundness of ∥Σu∥\|\mathbf{\Sigma}_{u}\| that ∣W^−W∣=O(p)∥(Σ^u,K^T)−1−Σu−1∥.|\hat{W}-W|=O(p)\|(\widehat{\mathbf{\Sigma}}_{u,\widehat{K}}^{\mathcal{T}})^{-1}-\mathbf{\Sigma}_{u}^{-1}\|. Theorem 3.1 provides the rate of convergence for the above difference. Detailed development is out of the scope of the current paper, and we will leave it as a separate research project.

Monte Carlo Experiments

In this section, we will examine the performance of the POET method in a finite sample. We will also demonstrate the effect of this estimator on the asset allocation and risk assessment. Similarly to Fan, et al. (2008, 2011), we simulated from a standard Fama-French three-factor model, assuming a sparse error covariance matrix and three factors. Throughout this section, the time span is fixed at T=300T=300, and the dimensionality pp increases from 11 to 600600. We assume that the excess returns of each of pp stocks over the risk-free interest rate follow the following model:

For each value of pp, we generate a sparse covariance matrix Σu\mathbf{\Sigma}_{u} of the form:

2 Simulation

For the simulation, we fix T=300T=300, and let pp increase from 11 to 600600. For each fixed pp, we repeat the following steps N=200N=200 times, and record the means and the standard deviations of each respective norm.

Set hard-thresholding with threshold 0.5θ^ij(log⁡pT+1p)0.5\sqrt{\hat{\theta}_{ij}}(\sqrt{\frac{\log p}{T}}+\frac{1}{\sqrt{p}}). Estimate KK using Bai and Ng (2002)’s IC1. Calculate covariance estimators using the POET method. Calculate the sample covariance matrix Σ^sam\widehat{\mathbf{\Sigma}}_{sam}.

In the graphs below, we plot the averages and standard deviations of the distance from Σ^K^\widehat{\mathbf{\Sigma}}_{\widehat{K}} and Σ^sam\widehat{\mathbf{\Sigma}}_{{}_{\text{sam}}} to the true covariance matrix Σ\mathbf{\Sigma}, under norms ∥.∥Σ\|.\|_{\Sigma}, ∥.∥\|.\| and ∥.∥max⁡\|.\|_{\max}. We also plot the means and standard deviations of the distances from (Σ^K^)−1(\widehat{\mathbf{\Sigma}}_{\widehat{K}})^{-1} and Σ^sam−1\widehat{\mathbf{\Sigma}}_{{}_{\text{sam}}}^{-1} to Σ−1\mathbf{\Sigma}^{-1} under the spectral norm. The dimensionality pp ranges from 2020 to 600600 in increments of 2020. Due to invertibility, the spectral norm for Σ^sam−1\widehat{\mathbf{\Sigma}}_{{}_{\text{sam}}}^{-1} is plotted only up to p=280p=280. Also, we zoom into these graphs by plotting the values of pp from 11 to 100100, this time in increments of 11. Notice that we also plot the distance from Σ^obs\widehat{\mathbf{\Sigma}}_{obs} to Σ\mathbf{\Sigma} for comparison, where Σ^obs\widehat{\mathbf{\Sigma}}_{obs} is the estimated covariance matrix proposed by Fan et al. (2011), assuming the factors are observable.

3 Results

In a factor model, we expect POET to perform as well as Σ^obs\widehat{\mathbf{\Sigma}}_{obs} when pp is relatively large, since the effect of estimating the unknown factors should vanish as pp increases. This is illustrated in the plots below.

From the simulation results, reported in Figures 2-5, we observe that POET under the unobservable factor model performs just as well as the estimator in Fan et al. (2011) if the factors are known, when pp is large enough. The cost of not knowing the factors is approximately of order Op(1/p)O_{p}(1/\sqrt{p}). It can be seen in Figures 2 and 3 that this cost vanishes for p≥200p\geq 200. To give a better insight of the impact of estimating the unknown factors for small pp, a separate set of simulations is conducted for p≤100p\leq 100. As we can see from Figures 2 (bottom panel) and 3 (middle and bottom panels), the impact decreases quickly. In addition, when estimating Σ−1\mathbf{\Sigma}^{-1}, it is hard to distinguish the estimators with known and unknown factors, whose performances are quite stable compared to the sample covariance matrix. Also, the maximum absolute elementwise error (Figure 4) of our estimator performs very similarly to that of the sample covariance matrix, which coincides with our asymptotic result. Figure 5 shows that the performances of the three methods are indistinguishable in the spectral norm, as expected.

4 Robustness to the estimation of K𝐾K

The POET estimator depends on the estimated number of factors. Our theory uses a consistent esimator K^\widehat{K}. To assess the robustness of our procedure to K^\widehat{K} in finite sample, we calculate Σ^u,KT\widehat{\mathbf{\Sigma}}_{u,K}^{\mathcal{T}} for K=1,2,...,10K=1,2,...,10. Again, the threshold is fixed to be 0.5θ^ij(log⁡pT+1p)0.5\sqrt{\hat{\theta}_{ij}}(\sqrt{\frac{\log p}{T}}+\frac{1}{\sqrt{p}}).

The simulation setup is the same as before where the true K0=3K_{0}=3. We calculate ∥Σ^u,KT−Σu∥\|\widehat{\mathbf{\Sigma}}_{u,K}^{\mathcal{T}}-\mathbf{\Sigma}_{u}\| , ∥(Σ^u,KT)−1−Σu−1∥\|(\widehat{\mathbf{\Sigma}}_{u,K}^{\mathcal{T}})^{-1}-\mathbf{\Sigma}_{u}^{-1}\|, ∥Σ^K−1−Σ−1∥\|\widehat{\mathbf{\Sigma}}^{-1}_{K}-\mathbf{\Sigma}^{-1}\| and ∥Σ^K−Σ∥Σ\|\widehat{\mathbf{\Sigma}}_{K}-\mathbf{\Sigma}\|_{\Sigma} for K=1,2,...,10K=1,2,...,10. Figure 6 plots these norms as pp increases but with a fixed T=300T=300. The results demonstrate a trend that is quite robust when K≥3K\geq 3; especially, the estimation accuracy of the spectral norms for large pp are close to each other. When K=1K=1 or 2, the estimators perform badly due to modeling bias. Therefore, POET is robust to over-estimated KK, but not to under-estimation.

4.2 Design 2

We also simulated from a new data generating process for the robustness assessment. Consider a banded idiosyncratic matrix

We still consider a K0=3K_{0}=3 factor model, where the factors are independently simulated as

Table 3 summarizes the average estimation error of covariance matrices across KK in the spectral norm. Each simulation is replicated 50 times and T=200T=200.

5 Comparisons with Other Methods

We compare POET with related methods that address low-rank plus sparse covariance estimation, specifically, LOREC proposed by Luo (2012), the strict factor model (SFM) by Fan, Fan and Lv (2008), the Dual Method (Dual) by Lin et al. (2009), and finally, the singular value thresholding (SVT) by Cai, Candès and Shen (2008). In particular, SFM is a special case of POET which employs a large threshold that forces Σ^u\widehat{\mathbf{\Sigma}}_{u} to be diagonal even when the true Σu\mathbf{\Sigma}_{u} might not be. Note that Dual, SVT and many others dealing with low-rank plus sparse, such as Candès et al. (2011) and Wright et al. (2009), assume a known Σ\mathbf{\Sigma} and focus on recovering the decomposition. Hence they do not estimate Σ\mathbf{\Sigma} or its inverse, but decompose the sample covariance into two components. The resulting sparse component may not be positive definite, which can lead to large estimation errors for Σ^u−1\widehat{\mathbf{\Sigma}}_{u}^{-1} and Σ^−1\widehat{\mathbf{\Sigma}}^{-1}.

5.2 Comparison with direct thresholding

This section compares POET with direct thresholding on the sample covariance matrix without taking out common factors (Rothman et al. 2009, Cai and Liu 2011. We denote this method by THR). We also run simulations to demonstrate the finite sample performance when Σ\mathbf{\Sigma} itself is sparse and has bounded eigenvalues, corresponding to the case K=0K=0. Three models are considered and both POET and THR use the soft thresholding. We fix T=200T=200. Reported results are the average of 100 replications.

Model 1: one-factor. The factors and loadings are independently generated from N(0,1)N(0,1). The error covariance is the same banded matrix as Design 2 in Section 6.4. Here Σ\mathbf{\Sigma} has one diverging eigenvalue.

Model 2: sparse covariance. Set K=0K=0, hence Σ=Σu\mathbf{\Sigma}=\mathbf{\Sigma}_{u} itself is a banded matrix with bounded eigenvalues.

Model 3: cross-sectional AR(1). Set K=0K=0, but Σ=Σu=(0.85∣i−j∣)p×p\mathbf{\Sigma}=\mathbf{\Sigma}_{u}=(0.85^{|i-j|})_{p\times p}. Now Σ\mathbf{\Sigma} is no longer sparse (or banded), but is not too dense either since Σij\Sigma_{ij} decreases to zero exponentially fast as ∣i−j∣→∞|i-j|\rightarrow\infty. This is the correlation matrix if {yit}i=1p\{y_{it}\}_{i=1}^{p} follows a cross-sectional AR(1) process: yit=0.85yi−1,t+εity_{it}=0.85y_{i-1,t}+\varepsilon_{it}.

For each model, POET uses an estimated K^\widehat{K} based on IC1 of Bai and Ng (2002), while THR thresholds the sample covariance directly. We find that in Model 1, POET performs significantly better than THR as the latter misses the common factor. For Model 2, IC1 estimates K^=0\widehat{K}=0 precisely in each replication, and hence POET is identical to THR. For Model 3, POET still outperforms. The results are summarized in Table 5.

6 Simulated portfolio allocation

We demonstrate the improvement of our method compared to the sample covariance and that based on the strict factor model (SFM), in a problem of portfolio allocation for risk minimization purposes.

the estimated (minimum variance) portfolio. Then the actual risk of the estimated portfolio is defined as R(^\mboxw)=^\mboxw′Σ^\mboxwR(\widehat{}\mbox{\bf w})=\widehat{}\mbox{\bf w}^{\prime}\mathbf{\Sigma}\widehat{}\mbox{\bf w}, and the estimated risk (also called empirical risk) is equal to R^(^\mboxw)=^\mboxw′Σ^^\mboxw\hat{R}(\widehat{}\mbox{\bf w})=\widehat{}\mbox{\bf w}^{\prime}\widehat{\mathbf{\Sigma}}\widehat{}\mbox{\bf w}. In practice, the actual risk is unknown, and only the empirical risk can be calculated.

Real Data Example

We demonstrate the sparsity of the approximate factor model on real data, and present the improvement of the POET estimator over the strict factor model (SFM) in a real-world application of portfolio allocation.

The data were obtained from the CRSP (The Center for Research in Security Prices) database, and consists of p=50p=50 stocks and their annualized daily returns for the period January 1st,20101^{st},2010-December 31st,201031^{st},2010 (T=252T=252). The stocks are chosen from 55 different industry sectors, (more specifically, Consumer Goods-Textile &\& Apparel Clothing, Financial-Credit Services, Healthcare-Hospitals, Services-Restaurants, Utilities-Water utilities), with 1010 stocks from each sector. We made this selection to demonstrate a block diagonal trend in the sparsity. More specifically, we show that the non-zero elements are clustered mainly within companies in the same industry. We also notice that these are the same groups that show predominantly positive correlation.

The largest eigenvalues of the sample covariance equal 0.0102,0.00450.0102,0.0045 and 0.00390.0039, while the rest are bounded by 0.00200.0020. Hence K=0,1,2,3K=0,1,2,3 are the possible values of the number of factors. Figure 9 shows the heatmap of the thresholded error correlation matrix (for simplicity, we applied hard thresholding). The threshold has been chosen using the cross validation as described in Section 4. We compare the level of sparsity (percentage of non-zero off-diagonal elements) for the 5 diagonal blocks of size 10×1010\times 10, versus the sparsity of the rest of the matrix. For K=2K=2, our method results in 25.8%25.8\% non-zero off-diagonal elements in the 55 diagonal blocks, as opposed to 7.3%7.3\% non-zero elements in the rest of the covariance matrix. Note that, out of the non-zero elements in the central 55 blocks, 100%100\% are positive, as opposed to a distribution of 60.3%60.3\% positive and 39.7%39.7\% negative amongst the non-zero elements in off-diagonal blocks. There is a strong positive correlation between the returns of companies in the same industry after the common factors are taken out, and the thresholding has preserved them. The results for K=1,2K=1,2 and 33 show the same characteristics. These provide stark evidence that the strict factor model is not appropriate.

2 Portfolio Allocation

We extend our data size by including larger industrial portfolios (p=100p=100), and longer period (ten years): January 1st1^{st},2000 to December 31st31^{st}, 2010 of annualized daily excess returns. Two portfolios are created at the beginning of each month, based on two different covariance estimates through approximate and strict factor models with unknown factors. At the end of each month, we compare the risks of both portfolios.

We can see from Figure 10 that the minimum-risk portfolio created by the POET estimator performs significantly better, achieving lower variance 76%76\% of the time. Amongst those months, the risk is decreased by 48.63%48.63\%. On the other hand, during the months that POET produces a higher-risk portfolio, the risk is increased by only 17.66%17.66\%.

Next, we demonstrate the impact of the choice of number of factors and threshold on the performance of POET. If cross-validation seems computationally expensive, we can choose a common soft-threshold throughout the whole investment process. The average constant in the cross-validation was 0.530.53, close to our suggested constant 0.50.5 used for simulation. We also present the results based on various choices of constant C=0.5C=0.5,0.750.75,11 and 1.251.25, with soft threshold Cθ^ijωTC\sqrt{\hat{\theta}_{ij}}\omega_{T}. The results are summarized in Table 6. The performance of POET seems consistent across different choices of these parameters.

Conclusion and Discussion

We study the problem of estimating a high-dimensional covariance matrix with conditional sparsity. Realizing unconditional sparsity assumption is inappropriate in many applications, we introduce a latent factor model that has a conditional sparsity feature, and propose the POET estimator to take advantage of the structure. This expands considerably the scope of the model based on the strict factor model, which assumes independent idiosyncratic noise and is too restrictive in practice. By assuming sparse error covariance matrix, we allow for the presence of the cross-sectional correlation even after taking out the common factors. The sparse covariance is estimated by the adaptive thresholding technique.

It is found that the rates of convergence of the estimators have an extra term approximately Op(p−1/2)O_{p}(p^{-1/2}) in addition to the results based on observable factors by Fan et al. (2008, 2011), which arises from the effect of estimating the unobservable factors. As we can see, this effect vanishes as the dimensionality increases, as more information about the common factors becomes available. When pp gets large enough, the effect of estimating the unknown factors is negligible, and we estimate the covariance matrices as if we knew the factors.

The proposed POET also has wide applicability in statistical genomics. For example, Carvalho et al. (2008) applied a Bayesian sparse factor model to study the breast cancer hormonal pathways. Their real-data results have identified about two common factors that have highly loaded genes (about half of 250 genes). As a result, these factors should be treated as “pervasive” (see the explanation in Example 2.1), which will result in one or two very spiked eigenvalues of the gene expressions’ covariance matrix. The POET can be applied to estimate such a covariance matrix and its network model.

Appendix A Estimating a sparse covariance with contaminated data

We can estimate Σu\mathbf{\Sigma}_{u} using the adaptive thresholding proposed by Cai and Liu (2011): for the threshold τij=Cθ^ijωT,\tau_{ij}=C\sqrt{\hat{\theta}_{ij}}\omega_{T}, define

Suppose (log⁡p)6α=o(T)(\log p)^{6\alpha}=o(T), and Assumptions 3.2 and 3.3 hold. In addition, suppose there is a sequence aT=o(1)a_{T}=o(1) so that max⁡i≤p1T∑t=1T∣uit−u^it∣2=Op(aT2),\max_{i\leq p}\frac{1}{T}\sum_{t=1}^{T}|u_{it}-\hat{u}_{it}|^{2}=O_{p}(a_{T}^{2}), and max⁡i≤p,t≤T∣uit−u^it∣=op(1);\max_{i\leq p,t\leq T}|u_{it}-\hat{u}_{it}|=o_{p}(1); Then there is a constant C>0C>0 in the adaptive thresholding estimator (A.1) with

If further ωTmp=o(1)\omega_{T}m_{p}=o(1), then Σ^uT\widehat{\mathbf{\Sigma}}_{u}^{\mathcal{T}} is invertible with probability approaching one, and

By Assumptions 3.2 and 3.3, the conditions of Lemmas A.3 and A.4 of Fan, Liao and Mincheva (2011, Ann. Statist, 39, 3320-3356) are satisfied. Hence for any ϵ>0\epsilon>0, there are positive constants M,θ1M,\theta_{1} and θ2\theta_{2} such that each of the events

occurs with probability at least 1−ϵ1-\epsilon. By the condition of threshold function, sij(t)=sij(t)I∣t∣>CωTθ^ijs_{ij}(t)=s_{ij}(t)I_{|t|>C\omega_{T}\sqrt{\hat{\theta}_{ij}}}. Now for C=θ2−12M,C=\theta_{2}^{-1}2M, under the event A1∩A2,A_{1}\cap A_{2},

Let M1=(Cθ1+M)(M−q+(Cθ1+M)−q).M_{1}=(C\theta_{1}+M)(M^{-q}+(C\theta_{1}+M)^{-q}). Then with probability at least 1−2ϵ1-2\epsilon, ∥Σ^uT−Σu∥≤mpωT1−qM1.\|\widehat{\mathbf{\Sigma}}_{u}^{\mathcal{T}}-\mathbf{\Sigma}_{u}\|\leq m_{p}\omega_{T}^{1-q}M_{1}. Since ϵ\epsilon is arbitrary, we have ∥Σ^uT−Σu∥=Op(ωT1−qmp)\|\widehat{\mathbf{\Sigma}}_{u}^{\mathcal{T}}-\mathbf{\Sigma}_{u}\|=O_{p}(\omega_{T}^{1-q}m_{p}). If in addition, ωTmp=o(1)\omega_{T}m_{p}=o(1), then the minimum eigenvalue of Σ^uT\widehat{\mathbf{\Sigma}}_{u}^{\mathcal{T}} is bounded away from zero with probability approaching one since λmin⁡(Σu)>c1\lambda_{\min}(\mathbf{\Sigma}_{u})>c_{1}. This then implies ∥(Σ^uT)−1−Σu−1∥=Op(ωT1−qmp).\|(\widehat{\mathbf{\Sigma}}_{u}^{\mathcal{T}})^{-1}-\mathbf{\Sigma}_{u}^{-1}\|=O_{p}\left(\omega_{T}^{1-q}m_{p}\right).

Appendix B Proofs for Section 2

We first cite two useful theorems, which are needed to prove propositions 2.1 and 2.2. In Lemma B.1 below, let {λi}i=1p\{\lambda_{i}\}_{i=1}^{p} be the eigenvalues of Σ\Sigma in descending order and \{\mbox{\boldmath\xi}_{i}\}_{i=1}^{p} be their associated eigenvectors. Correspondingly, let {λ^i}i=1p\{\widehat{\lambda}_{i}\}_{i=1}^{p} be the eigenvalues of Σ^\widehat{\mathbf{\Sigma}} in descending order and \{\widehat{}\mbox{\boldmath\xi}_{i}\}_{i=1}^{p} be their associated eigenvectors.

(Weyl’s Theorem) |\widehat{\lambda}_{i}-\lambda_{i}|\leq\|\widehat{\mathbf{\Sigma}}-\mbox{\boldmath\Sigma}\|.

(sin⁡θ\sin\theta Theorem, Davis and Kahan, 1970)

Appendix C Proofs for Section 3

We will proceed by subsequently showing Theorems 3.3, 3.1 and 3.2.

The following results are to be used subsequently. The proofs of Lemmas C.1,C.2 and C.3 are found in Fan, Liao and Mincheva (2011).

Suppose that the random variables Z1,Z2Z_{1},Z_{2} both satisfy the exponential-type tail condition: There exist r1r_{1}, r2∈(0,1)r_{2}\in(0,1) and b1,b2>0b_{1},b_{2}>0, such that ∀s>0\forall s>0,

Then for some r3r_{3} and b3>0b_{3}>0, and any s>0s>0,

Under the assumptions of Theorem 3.1, (i) max⁡i,j≤K∣1T∑t=1Tfitfjt−Efitfjt∣=Op(1/T)\max_{i,j\leq K}|\frac{1}{T}\sum_{t=1}^{T}f_{it}f_{jt}-Ef_{it}f_{jt}|=O_{p}(\sqrt{1/T}). (ii) max⁡i,j≤p∣1T∑t=1Tuitujt−Euitujt∣=Op((log⁡p)/T)\max_{i,j\leq p}|\frac{1}{T}\sum_{t=1}^{T}u_{it}u_{jt}-Eu_{it}u_{jt}|=O_{p}(\sqrt{(\log p)/T}) (iii) max⁡i≤K,j≤p∣1T∑t=1Tfitujt∣=Op((log⁡p)/T)\max_{i\leq K,j\leq p}|\frac{1}{T}\sum_{t=1}^{T}f_{it}u_{jt}|=O_{p}(\sqrt{(\log p)/T})

First of all, by Proposition 2.1, under Assumption 3.1, the KthK{th} largest eigenvalue λK\lambda_{K} of Σ\mathbf{\Sigma} satisfies: for some c>0,c>0,

C.2 Proof of Theorem 3.3

Our derivation below relies on a result obtained by Bai and Ng (2002), which showed that the estimated number of factors is consistent, in the sense that K^\widehat{K} equals the true KK with probability approaching one. Note that under our Assumptions 3.1-3.4, all the assumptions in Bai and Ng (2002) are satisfied. Thus immediately we have the following Lemma.

Using (A.1) in Bai (2003), we have the following identity:

(i) We have ∀i\forall i, ∑s=1Tf^is2=T\sum_{s=1}^{T}\hat{f}_{is}^{2}=T. By the Cauchy-Schwarz inequality,

Note that E(∑s=1T∑l=1T(∑t=1Tζstζlt)2)=T2E(∑t=1Tζstζlt)2≤T4max⁡stE∣ζst∣4.E(\sum_{s=1}^{T}\sum_{l=1}^{T}(\sum_{t=1}^{T}\zeta_{st}\zeta_{lt})^{2})=T^{2}E(\sum_{t=1}^{T}\zeta_{st}\zeta_{lt})^{2}\leq T^{4}\max_{st}E|\zeta_{st}|^{4}. By Assumption 3.4, max⁡stEζst4=O(p−2)\max_{st}E\zeta_{st}^{4}=O(p^{-2}), which implies that ∑s,l(∑t=1Tζstζlt)2=Op(T4/p2)\sum_{s,l}(\sum_{t=1}^{T}\zeta_{st}\zeta_{lt})^{2}=O_{p}(T^{4}/p^{2}), and yields the result.

(iv) Similar to part (iii), noting that ξst\xi_{st} is a scalar, we have:

where the third line follows from the Cauchy-Schwarz inequality. ∎

The result then follows from Assumption 3.3.

It follows from Assumption 3.4 that E(1T∑s=1Tζst2)2≤max⁡s,t≤TEζst4=O(1p2).E(\frac{1}{T}\sum_{s=1}^{T}\zeta_{st}^{2})^{2}\leq\max_{s,t\leq T}E\zeta_{st}^{4}=O(\frac{1}{p^{2}}). It then follows from the Chebyshev’s inequality and Bonferroni’s method that max⁡t1T∑s=1Tζst2=Op(T/p)\max_{t}\frac{1}{T}\sum_{s=1}^{T}\zeta_{st}^{2}=O_{p}(\sqrt{T}/p).

We prove this lemma conditioning on the event K^=K.\widehat{K}=K. Once this is done, due to P(K^≠K)=o(1)P(\widehat{K}\neq K)=o(1), it then implies the unconditional arguments.

Each of the four terms on the right hand side above are bounded in Lemma C.7, which then yields the desired result.

Part (iii) is implied by (C.2) and Lemma C.8. ∎

C.3 Proof of Theorem 3.1

max⁡i≤p1T∑t=1T∣uit−u^it∣2=Op(ωT2),\max_{i\leq p}\frac{1}{T}\sum_{t=1}^{T}|u_{it}-\hat{u}_{it}|^{2}=O_{p}\left(\omega_{T}^{2}\right), and max⁡i,t∣uit−u^it∣=op(1).\max_{i,t}|u_{it}-\hat{u}_{it}|=o_{p}(1).

The first part of the lemma then follows from Theorem 3.3 and Lemma C.9. The second part follows from Corollary 3.1.

Proof of Theorem 3.1: The theorem follows immediately from Theorem A.1 and Lemma C.11.

C.4 Proof of Theorem 3.2

(ii) The result follows from part (i) and Lemma C.13. Part (iii) and (iv) follow from a similar argument of part (i) and Lemma C.10.

We derive the rate for ∥Σ^K^−1−Σ−1∥.\|\widehat{\mathbf{\Sigma}}^{-1}_{\widehat{K}}-\mathbf{\Sigma}^{-1}\|. Define

One the other hand, using Sherman-Morrison-Woodbury formula again implies

Proof of Theorem 3.2: ∥Σ^T−Σ∥max⁡\|\widehat{\mathbf{\Sigma}}^{\mathcal{T}}-\mathbf{\Sigma}\|_{\max}

On the other hand, let σu,ij\sigma_{u,ij} be the (i,j)(i,j) entry of Σu\mathbf{\Sigma}_{u}. Then max⁡ij∣σ^ij−σu,ij∣=Op(ωT)\max_{ij}|\widehat{\sigma}_{ij}-\sigma_{u,ij}|=O_{p}(\omega_{T}).

Hence ∥Σ^u,K^T−Σu∥max⁡=Op(ωT).\|\widehat{\mathbf{\Sigma}}_{u,\widehat{K}}^{\mathcal{T}}-\mathbf{\Sigma}_{u}\|_{\max}=O_{p}(\omega_{T}). The result then follows immediately. ∎

References