Large Covariance Estimation by Thresholding Principal Orthogonal Complements
Jianqing Fan, Yuan Liao, Martina Mincheva
Introduction
Information and technology make large data sets widely available for scientific discovery. Much statistical analysis of such high-dimensional data involves the estimation of a covariance matrix or its inverse (the precision matrix). Examples include portfolio management and risk assessment (Fan, Fan and Lv, 2008), high-dimensional classification such as Fisher discriminant (Hastie, Tibshirani and Friedman, 2009), graphic models (Meinshausen and Bühlmann, 2006), statistical inference such as controlling false discoveries in multiple testing (Leek and Storey, 2008; Efron, 2010), finding quantitative trait loci based on longitudinal data (Yap, Fan, and Wu, 2009; Xiong et al. 2011), and testing the capital asset pricing model (Sentana, 2009), among others. See Section 5 for some of those applications. Yet, the dimensionality is often either comparable to the sample size or even larger. In such cases, the sample covariance is known to have poor performance (Johnstone, 2001), and some regularization is needed.
Realizing the importance of estimating large covariance matrices and the challenges brought by the high dimensionality, in recent years researchers have proposed various regularization techniques to consistently estimate . One of the key assumptions is that the covariance matrix is sparse, namely, many entries are zero or nearly so (Bickel and Levina, 2008, Rothman et al, 2009, Lam and Fan 2009, Cai and Zhou, 2010, Cai and Liu, 2011). In many applications, however, the sparsity assumption directly on is not appropriate. For example, financial returns depend on the equity market risks, housing prices depend on the economic health, gene expressions can be stimulated by cytokines, among others. Due to the presence of common factors, it is unrealistic to assume that many outcomes are uncorrelated. An alternative method is to assume a factor model structure, as in Fan, Fan and Lv (2008). However, they restrict themselves to the strict factor models with known factors.
A natural extension is the conditional sparsity. Given the common factors, the outcomes are weakly correlated. In order to do so, we consider an approximate factor model, which has been frequently used in economic and financial studies (Chamberlain and Rothschild, 1983; Fama and French 1993; Bai and Ng, 2002, etc):
We emphasize that in model (1.1), only is observable. It is intuitively clear that the unknown common factors can only be inferred reliably when there are sufficiently many cases, that is, . In a data-rich environment, can diverge at a rate faster than . The factor model (1.1) can be put in a matrix form as
does not grow too fast as In particular, this includes the exact sparsity assumption () under which , the maximum number of nonzero elements in each row.
This paper focuses on the high-dimensional static factor model (1.2), which is innately related to the principal component analysis (PCA), as clarified in Section 2. This feature makes it different from the classical factor model with fixed dimensionality (e.g., Lawley and Maxwell 1971). In the last ten years, much theory on the estimation and inference of the static factor model has been developed, for example, Stock and Watson (1998, 2002), Bai and Ng (2002), Bai (2003), Doz, Giannone and Reichlin (2011), among others. Our contribution is on the estimation of covariance matrices and their inverse in large factor models.
There has been extensive literature in recent years that deals with sparse principal components, which has been widely used to enhance the convergence of the principal components in high-dimensional space. d’Aspremont, Bach and El Ghaoui (2008), Shen and Huang (2008), Witten, Tibshirani, and Hastie (2009) and Ma (2011) proposed and studied various algorithms for computations. More literature on sparse PCA is found in Johnstone and Lu (2009), Amini and Wainwright (2009), Zhang and El Ghaoui (2011), Birnbaum et al. (2012), among others. In addition, there has also been a growing literature that theoretically studies the recovery from a low-rank plus sparse matrix estimation problem, see for example, Wright et al. (2009), Lin et al. (2009), Candès et al. (2011), Luo (2011), Agarwal, Nagahban, Wainwright (2012), Pati et al. (2012). It corresponds to the identifiability issue of our problem.
There is a big difference between our model and those considered in the aforementioned literature. In the current paper, the first eigenvalues of are spiked and grow at a rate , whereas the eigenvalues of the matrices studied in the existing literature on covariance estimation are usually assumed to be either bounded or slowly growing. Due to this distinctive feature, the common components and the idiosyncratic components can be identified, and in addition, PCA on the sample covariance matrix can consistently estimate the space spanned by the eigenvectors of . The existing methods of either thresholding directly or solving a constrained optimization method can fail in the presence of very spiked principal eigenvalues. However, there is a price to pay here: as the first eigenvalues are “too spiked”, one can hardly obtain a satisfactory rate of convergence for estimating in absolute term, but it can be estimated accurately in relative term (see Section 3.3 for details). In addition, can be estimated accurately.
We would like to further note that the low-rank plus sparse representation of our model is on the population covariance matrix, whereas Candès et al. (2011), Wright et al. (2009), Lin et al. (2009)We thank a referee for reminding us these related works. considered such a representation on the data matrix. As there is no to estimate, their goal is limited to producing a low-rank plus sparse matrix decomposition of the data matrix, which corresponds to the identifiability issue of our study, and does not involve estimation and inference. In contrast, our ultimate goal is to estimate the population covariance matrices as well as the precision matrices. For this purpose, we require the idiosyncratic components and common factors to be uncorrelated and the data generating process to be strictly stationary. The covariances considered in this paper are constant over time, though slow-time-varying covariance matrices are applicable through localization in time (time-domain smoothing). Our consistency result on demonstrates that the decomposition (1.3) is identifiable, and hence our results also shed the light of the “surprising phenomenon” of Candès et al. (2011) that one can separate fully a sparse matrix from a low-rank matrix when only the sum of these two components is available.
Regularized Covariance Matrix via PCA
We illustrate this in the following example.
Let be the eigenvalues of in a descending order and \{\mbox{\boldmath\xi}_{j}\}_{j=1}^{p} be their corresponding eigenvectors. Then, an application of Weyl’s eigenvalue theorem (see the appendix) yields that
Using Proposition 2.1 and the theorem of Davis and Kahn (1970, see the appendix), we have the following:
Propositions 2.1 and 2.2 state that PCA and factor analysis are approximately the same if \|\mbox{\boldmath\Sigma}_{u}\|=o(p). This is assured through a sparsity condition on , which is frequently measured through
The intuition is that, after taking out the common factors, many pairs of the cross-sectional units become weakly correlated. This generalized notion of sparsity was used in Bickel and Levina (2008) and Cai and Liu (2011). Under this generalized measure of sparsity, we have
if the noise variances are bounded. Therefore, when , Proposition 2.1 implies that we have distinguished eigenvalues between the principal components and the rest of the components and Proposition 2.2 ensures that the first principal components are approximately the same as the columns of the factor loadings.
The aforementioned sparsity assumption appears reasonable in empirical applications. Boivin and Ng (2006) conducted an empirical study and showed that imposing zero correlation between weakly correlated idiosyncratic components improves forecastWe thank a referee for this interesting reference.. More recently, Phan (2012) empirically estimated the level of sparsity of the idiosyncratic covariance using the UK market data.
Recent developments on random matrix theory, for example, Johnstone and Lu (2009) and Paul (2007), have shown that when is not negligible, the eigenvalues and eigenvectors of might not be consistently estimated from the sample covariance matrix. A distinguished feature of the covariance considered in this paper is that there are some very spiked eigenvalues. By Propositions 2.1 and 2.2, in the factor model, the pervasiveness condition
With assumption (2.3), the standard literature on approximate factor models has shown that the PCA on the sample covariance matrix can consistently estimate the space spanned by the factor loadings (e.g., Stock and Watson (1998), Bai (2003)). Our contribution in Propositions 2.1 and 2.2 is that we connect the high-dimensional factor model to the principal components, and obtain the consistency of the spectrum in the population level instead of the sample level . The spectral consistency also enhances the results in Chamberlain and Rothschild (1983). This provides the rationale behind the consistency results in the factor model literature.
2 POET
Sparsity assumption directly on is inappropriate in many applications due to the presence of common factors. Instead, we propose a nonparametric estimator of based on the principal component analysis. Let be the ordered eigenvalues of the sample covariance matrix \widehat{\mbox{\boldmath\Sigma}}_{{}_{\text{sam}}} and \{\widehat{}\mbox{\boldmath\xi}_{i}\}_{i=1}^{p} be their corresponding eigenvectors. Then the sample covariance has the following spectral decomposition:
where is a generalized shrinkage function of Antoniadis and Fan (2001), employed by Rothman et al. (2009) and Cai and Liu (2011), and is an entry-dependent threshold. In particular, the hard-thresholding rule (Bickel and Levina, 2008) and the constant thresholding parameter are allowed. In practice, it is more desirable to have be entry-adaptive. An example of the adaptive thresholding is
The estimator of is then defined as:
We will call this estimator the Principal Orthogonal complEment thresholding (POET) estimator. It is obtained by thresholding the remaining components of the sample covariance matrix, after taking out the first principal components. One of the attractiveness of POET is that it is optimization-free, and hence is computationally appealing. We have written an R package for POET, which outputs the estimated , , , the factors and loadings.
With the choice of in (2.6) and the hard thresholding rule, our estimator encompasses many popular estimators as its specific cases. When , the estimator is the sample covariance matrix and when , the estimator becomes that based on the strict factor model (Fan, Fan, and Lv , 2008). When , our estimator is the same as the thresholding estimator of Bickel and Levina (2008) and (with a more general thresholding function) Rothman et al. (2009) or the adaptive thresholding estimator of Cai and Liu (2011) with a proper choice of .
In practice, the number of diverging eigenvalues (or common factors) can be estimated based on the sample covariance matrix. Determining in a data-driven way is an important topic, and is well understood in the literature. We will describe the POET with a data-driven in Section 2.4.
3 Least squares point of view
The constraints (2.9) correspond to the normalization (2.1). Here we assume that the mean of each variable has been removed, that is, for all and Putting it in a matrix form, the optimization problem can be written as
It is easy to verify that includes many interesting thresholding functions such as the hard thresholding (), soft thresholding (), SCAD, and adaptive lasso (See Rothman et al. (2009)).
Analogous to the decomposition (1.3), we obtain the following substitution estimators
Suppose that the entry-dependent threshold in (2.5) is the same as the thresholding parameter used in (2.11). Then for any , the estimator (2.7) is equivalent to the substitution estimator (2.12), that is,
In this paper, we will use a data-driven to construct the POET (see Section 2.4 below), which has two equivalent representations according to Theorem 2.1.
4 POET with Unknown K𝐾K
Determining the number of factors in a data-driven way has been an important research topic in the econometric literature. Bai and Ng (2002) proposed a consistent estimator as both and diverge. Other recent criteria are proposed by Kapetanios (2010), Onatski (2010), Alessi et al. (2010), etc.
Our method also allows a data-driven to estimate the covariance matrices. In principle, any procedure that gives a consistent estimate of can be adopted. In this paper we apply the well-known method in Bai and Ng (2002). It estimates by
Throughout the paper, we let be the solution to (2.14) using either IC1 or IC2. The asymptotic results are not affected regardless of the specific choice of . We define the POET estimator with unknown as
The procedure is as stated in Section 2.2 except that is now data-driven.
Asymptotic Properties
The first assumption has been one of the most essential ones in the literature of approximate factor models. Under this assumption and other regularity conditions, the number of factors, loadings and common factors can be consistently estimated (e.g., Stock and Watson (1998, 2002), Bai and Ng (2002), Bai (2003), etc.).
It implies from Proposition 2.1 in Section 2 that the first eigenvalues of grow at rate . This unique feature distinguishes our work from most of other low-rank plus sparse covariances considered in the literature, e.g., Luo (2011), Pati et al. (2012), Agarwal et al. (2012), Birnbaum et al. (2012). To our best knowledge, the only other papers that estimate large covariances with diverging eigenvalues (growing at the rate of dimensionality ) are Fan et al. (2008, 2011) and Bai and Shi (2011). While Fan et al. (2008, 2011) assumed the factors are observable, Bai and Shi (2011) considered the strict factor model in which is diagonal.
As to be illustrated in Section 3.3 below, due to the fast diverging eigenvalues, one can hardly achieve a good rate of convergence for estimating under either the spectral norm or Frobenius norm when . This phenomenon arises naturally from the characteristics of the high-dimensional factor model, which is another distinguished feature compared to those convergence results in the existing literature.
In addition, we impose the following regularity conditions.
These conditions are needed to consistently estimate the transformed common factors as well as the factor loadings. Similar conditions were also assumed in Bai (2003), and Bai and Ng (2006). The number of factors is assumed to be fixed. Our conditions in Assumption 3.4 are weaker than those in Bai (2003) as we focus on different aspects of the study.
2 Convergence of the idiosyncratic covariance
Throughout the paper, we apply the adaptive threshold
where is a sufficiently large constant, though the results hold for other types of thresholding. As in Bickel and Levina (2008) and Cai and Liu (2011), the threshold chosen in the current paper is in fact obtained from the optimal uniform rate of convergence of When direct observation of is not available, the effect of estimating the unknown factors also contributes to this uniform estimation error, which is why appears in the threshold.
The following theorem gives the rate of convergence of the estimated idiosyncratic covariance. Let . In the convergence rate below, recall that and are defined in the measure of sparsity (2.2).
Suppose , , and Assumptions 3.1-3.4 hold. Then for a sufficiently large constant in the threshold (3.2), the POET estimator satisfies
If further , then the eigenvalues of are all bounded away from zero with probability approaching one, and
When estimating , is allowed to grow exponentially fast in , and can be made consistent under the spectral norm. In addition, is asymptotically invertible while the classical sample covariance matrix based on the residuals is not when
Consistent estimation of indicates that is identifiable in (1.3), namely, the sparse can be separated perfectly from the low-rank matrix there. The result here gives another proof (when assuming ) of the “surprising phenomenon” in Candès et al (2011) under different technical conditions.
When is known and grows with and , with slightly weaker assumptions, our working paper (Fan et al. 2011) shows that under the exactly sparse case (that is, ), the result continues to hold with convergence rate .
3 Convergence of the POET estimator
Since the first eigenvalues of grow with , one can hardly estimate with satisfactory accuracy in the absolute term. This problem arises not from the limitation of any estimation method, but is due to the nature of the high-dimensional factor model. We illustrate this using a simple example.
Consider an ideal case where we know the spectrum except for the first eigenvector of . Let \{\lambda_{j},\mbox{\boldmath\xi}_{j}\}_{j=1}^{p} be the eigenvalues and vectors, and assume that the largest eigenvalue for some . Let \widehat{}\mbox{\boldmath\xi}_{1} be the estimated first eigenvector and define the covariance estimator \widehat{\mathbf{\Sigma}}=\lambda_{1}\widehat{}\mbox{\boldmath\xi}_{1}\widehat{}\mbox{\boldmath\xi}_{1}^{\prime}+\sum_{j=2}^{p}\lambda_{j}\mbox{\boldmath\xi}_{j}\mbox{\boldmath\xi}_{j}^{\prime}. Assume that \widehat{}\mbox{\boldmath\xi}_{1} is a good estimator in the sense that \|\widehat{}\mbox{\boldmath\xi}_{1}-\mbox{\boldmath\xi}_{1}\|^{2}=O_{p}(T^{-1}). However,
which can diverge when .
In the presence of very spiked eigenvalues, while the covariance cannot be consistently estimated in absolute term, it can be well estimated in terms of the relative error matrix
which is more relevant for many applications (see Example 5.2). The relative error matrix can be measured by either its spectral norm or the normalized Frobenius norm defined by
In the last equality, there are terms being added in the trace operation and the factor plays the role of normalization. The loss (3.3) is closely related to the entropy loss, introduced by James and Stein (1961). Also note that
Compared to estimating , in a large approximate factor model, we can estimate the precision matrix with a satisfactory rate under the spectral norm. The intuition follows from the fact that has bounded eigenvalues.
The following theorem summarizes the rate of convergence under various norms.
Under the assumptions of Theorem 3.1, the POET estimator defined in (2.15) satisfies
In addition, if , then is nonsingular with probability approaching one, with
When estimating , is allowed to grow exponentially fast in , and the estimator has the same rate of convergence as that of the estimator in Theorem 3.1. When becomes much larger than , the precision matrix can be estimated at the same rate as if the factors were observable.
As in Remark 3.2, when is known and grows with and , the working paper Fan et al. (2011) proves the following results (when ) The assumptions in the working paper Fan et al. (2011) are slightly weak than those presented here, in that it required instead of be bounded.:
The results state explicitly the dependence of the rate of convergence on the number of factors.
4 Convergence of unknown factors and factor loadings
As a consequence of Theorem 3.3, we obtain the following: (recall that the constant is defined in Assumption 3.2.)
Choice of Threshold
Recall that the threshold value , where is determined by the users. To make POET operational in practice, one has to choose to maintain the positive definiteness of the estimated covariances for any given finite sample. We write , where the covariance estimator depends on via the threshold. We choose in the range where . Define
When is sufficiently large, the estimator becomes diagonal, while its minimum eigenvalue must retain strictly positive. Thus, is well defined and for all , is positive definite under finite sample. We can obtain by solving We can also approximate by plotting as a function of , as illustrated in Figure 1. In practice, we can choose in the range for a small and large enough Choosing the threshold in a range to guarantee the finite-sample positive definiteness has also been previously suggested by Fryzlewicz (2012).
2 Multifold Cross-Validation
It is possible to derive the rate of convergence for under the current model setting, but it ought to be much more technically involved than the regular sparse matrix estimation considered by Bickel and Levina (2008) and Cai and Liu (2011). To keep our presentation simple we do not pursue it in the current paper.
Applications of POET
We give four examples to which the results in Theorems 3.1–3.3 can be applied. Detailed pursuits of these are beyond the scope of the paper.
Controlling the false discovery rate in large-scale hypothesis testing based on correlated test statistics is an important and challenging problem in statistics (Leek and Storey, 2008; Efron, 2010; Fan, et al., 2012). Suppose that the test statistic for each of the hypothesis
If the covariance admits the approximate factor structure (1.3), then the test statistics can be stochastically decomposed as
By the principal factor approximation (Theorem 1, Fan, Han, Gu, 2012)
Now suppose that we have repeated measurements from the model (5.1). Then, by Corollary 3.1, can be uniformly consistently estimated, and hence and can be consistently estimated. Efron (2010) obtained these repeated test statistics based on the bootstrap sample from the original raw data. Our theory (Theorem 3.3) gives a formal justification to the framework of Efron (2007, 2010).
The maximum elementwise estimation error appears in risk assessment as in Fan, Zhang and Yu (2012). For a fixed portfolio allocation vector w, the true portfolio variance and the estimated one are given by and respectively. The estimation error is bounded by
where , the -norm of w, is the gross exposure of the portfolio. Usually a constraint is placed on the total percentage of the short positions, in which case we have a restriction for some In particular, corresponds to a portfolio with no-short positions (all weights are nonnegative). Theorem 3.2 quantifies the maximum approximation error.
The above compares the absolute error of perceived risk and true risk. The relative error is bounded by
for any allocation vector w. Theorem 3.2 quantifies this relative error.
Consider the following panel regression model
and wishes to test H_{0}:\mbox{\boldmath\alpha}=0. The F-test statistic involves the estimation of the covariance matrix , whose estimates are degenerate without regularization when . Therefore, in the literature (Sentana, 2009, and references therein), one focuses on the case is relatively small. The typical choices of parameters are monthly data and the number of assets , 10 or 25. However, the CAPM should hold for all tradeable assets, not just a small fraction of assets. With our regularization technique, non-degenerate estimate can be obtained and the F-test or likelihood-ratio test statistics can be employed even when .
Since under the null hypothesis \hat{\mbox{\boldmath\alpha}}\sim N(0,\mathbf{\Sigma}_{u}/c_{T}), we have c_{T}\|\mathbf{\Sigma}_{u}^{-1/2}\hat{\mbox{\boldmath\alpha}}\|^{2}=O(p). Thus, it follows from boundness of that Theorem 3.1 provides the rate of convergence for the above difference. Detailed development is out of the scope of the current paper, and we will leave it as a separate research project.
Monte Carlo Experiments
In this section, we will examine the performance of the POET method in a finite sample. We will also demonstrate the effect of this estimator on the asset allocation and risk assessment. Similarly to Fan, et al. (2008, 2011), we simulated from a standard Fama-French three-factor model, assuming a sparse error covariance matrix and three factors. Throughout this section, the time span is fixed at , and the dimensionality increases from to . We assume that the excess returns of each of stocks over the risk-free interest rate follow the following model:
For each value of , we generate a sparse covariance matrix of the form:
2 Simulation
For the simulation, we fix , and let increase from to . For each fixed , we repeat the following steps times, and record the means and the standard deviations of each respective norm.
Set hard-thresholding with threshold . Estimate using Bai and Ng (2002)’s IC1. Calculate covariance estimators using the POET method. Calculate the sample covariance matrix .
In the graphs below, we plot the averages and standard deviations of the distance from and to the true covariance matrix , under norms , and . We also plot the means and standard deviations of the distances from and to under the spectral norm. The dimensionality ranges from to in increments of . Due to invertibility, the spectral norm for is plotted only up to . Also, we zoom into these graphs by plotting the values of from to , this time in increments of . Notice that we also plot the distance from to for comparison, where is the estimated covariance matrix proposed by Fan et al. (2011), assuming the factors are observable.
3 Results
In a factor model, we expect POET to perform as well as when is relatively large, since the effect of estimating the unknown factors should vanish as increases. This is illustrated in the plots below.
From the simulation results, reported in Figures 2-5, we observe that POET under the unobservable factor model performs just as well as the estimator in Fan et al. (2011) if the factors are known, when is large enough. The cost of not knowing the factors is approximately of order . It can be seen in Figures 2 and 3 that this cost vanishes for . To give a better insight of the impact of estimating the unknown factors for small , a separate set of simulations is conducted for . As we can see from Figures 2 (bottom panel) and 3 (middle and bottom panels), the impact decreases quickly. In addition, when estimating , it is hard to distinguish the estimators with known and unknown factors, whose performances are quite stable compared to the sample covariance matrix. Also, the maximum absolute elementwise error (Figure 4) of our estimator performs very similarly to that of the sample covariance matrix, which coincides with our asymptotic result. Figure 5 shows that the performances of the three methods are indistinguishable in the spectral norm, as expected.
4 Robustness to the estimation of K𝐾K
The POET estimator depends on the estimated number of factors. Our theory uses a consistent esimator . To assess the robustness of our procedure to in finite sample, we calculate for . Again, the threshold is fixed to be .
The simulation setup is the same as before where the true . We calculate , , and for . Figure 6 plots these norms as increases but with a fixed . The results demonstrate a trend that is quite robust when ; especially, the estimation accuracy of the spectral norms for large are close to each other. When or 2, the estimators perform badly due to modeling bias. Therefore, POET is robust to over-estimated , but not to under-estimation.
4.2 Design 2
We also simulated from a new data generating process for the robustness assessment. Consider a banded idiosyncratic matrix
We still consider a factor model, where the factors are independently simulated as
Table 3 summarizes the average estimation error of covariance matrices across in the spectral norm. Each simulation is replicated 50 times and .
5 Comparisons with Other Methods
We compare POET with related methods that address low-rank plus sparse covariance estimation, specifically, LOREC proposed by Luo (2012), the strict factor model (SFM) by Fan, Fan and Lv (2008), the Dual Method (Dual) by Lin et al. (2009), and finally, the singular value thresholding (SVT) by Cai, Candès and Shen (2008). In particular, SFM is a special case of POET which employs a large threshold that forces to be diagonal even when the true might not be. Note that Dual, SVT and many others dealing with low-rank plus sparse, such as Candès et al. (2011) and Wright et al. (2009), assume a known and focus on recovering the decomposition. Hence they do not estimate or its inverse, but decompose the sample covariance into two components. The resulting sparse component may not be positive definite, which can lead to large estimation errors for and .
5.2 Comparison with direct thresholding
This section compares POET with direct thresholding on the sample covariance matrix without taking out common factors (Rothman et al. 2009, Cai and Liu 2011. We denote this method by THR). We also run simulations to demonstrate the finite sample performance when itself is sparse and has bounded eigenvalues, corresponding to the case . Three models are considered and both POET and THR use the soft thresholding. We fix . Reported results are the average of 100 replications.
Model 1: one-factor. The factors and loadings are independently generated from . The error covariance is the same banded matrix as Design 2 in Section 6.4. Here has one diverging eigenvalue.
Model 2: sparse covariance. Set , hence itself is a banded matrix with bounded eigenvalues.
Model 3: cross-sectional AR(1). Set , but . Now is no longer sparse (or banded), but is not too dense either since decreases to zero exponentially fast as . This is the correlation matrix if follows a cross-sectional AR(1) process: .
For each model, POET uses an estimated based on IC1 of Bai and Ng (2002), while THR thresholds the sample covariance directly. We find that in Model 1, POET performs significantly better than THR as the latter misses the common factor. For Model 2, IC1 estimates precisely in each replication, and hence POET is identical to THR. For Model 3, POET still outperforms. The results are summarized in Table 5.
6 Simulated portfolio allocation
We demonstrate the improvement of our method compared to the sample covariance and that based on the strict factor model (SFM), in a problem of portfolio allocation for risk minimization purposes.
the estimated (minimum variance) portfolio. Then the actual risk of the estimated portfolio is defined as , and the estimated risk (also called empirical risk) is equal to . In practice, the actual risk is unknown, and only the empirical risk can be calculated.
Real Data Example
We demonstrate the sparsity of the approximate factor model on real data, and present the improvement of the POET estimator over the strict factor model (SFM) in a real-world application of portfolio allocation.
The data were obtained from the CRSP (The Center for Research in Security Prices) database, and consists of stocks and their annualized daily returns for the period January -December (). The stocks are chosen from different industry sectors, (more specifically, Consumer Goods-Textile Apparel Clothing, Financial-Credit Services, Healthcare-Hospitals, Services-Restaurants, Utilities-Water utilities), with stocks from each sector. We made this selection to demonstrate a block diagonal trend in the sparsity. More specifically, we show that the non-zero elements are clustered mainly within companies in the same industry. We also notice that these are the same groups that show predominantly positive correlation.
The largest eigenvalues of the sample covariance equal and , while the rest are bounded by . Hence are the possible values of the number of factors. Figure 9 shows the heatmap of the thresholded error correlation matrix (for simplicity, we applied hard thresholding). The threshold has been chosen using the cross validation as described in Section 4. We compare the level of sparsity (percentage of non-zero off-diagonal elements) for the 5 diagonal blocks of size , versus the sparsity of the rest of the matrix. For , our method results in non-zero off-diagonal elements in the diagonal blocks, as opposed to non-zero elements in the rest of the covariance matrix. Note that, out of the non-zero elements in the central blocks, are positive, as opposed to a distribution of positive and negative amongst the non-zero elements in off-diagonal blocks. There is a strong positive correlation between the returns of companies in the same industry after the common factors are taken out, and the thresholding has preserved them. The results for and show the same characteristics. These provide stark evidence that the strict factor model is not appropriate.
2 Portfolio Allocation
We extend our data size by including larger industrial portfolios (), and longer period (ten years): January ,2000 to December , 2010 of annualized daily excess returns. Two portfolios are created at the beginning of each month, based on two different covariance estimates through approximate and strict factor models with unknown factors. At the end of each month, we compare the risks of both portfolios.
We can see from Figure 10 that the minimum-risk portfolio created by the POET estimator performs significantly better, achieving lower variance of the time. Amongst those months, the risk is decreased by . On the other hand, during the months that POET produces a higher-risk portfolio, the risk is increased by only .
Next, we demonstrate the impact of the choice of number of factors and threshold on the performance of POET. If cross-validation seems computationally expensive, we can choose a common soft-threshold throughout the whole investment process. The average constant in the cross-validation was , close to our suggested constant used for simulation. We also present the results based on various choices of constant ,, and , with soft threshold . The results are summarized in Table 6. The performance of POET seems consistent across different choices of these parameters.
Conclusion and Discussion
We study the problem of estimating a high-dimensional covariance matrix with conditional sparsity. Realizing unconditional sparsity assumption is inappropriate in many applications, we introduce a latent factor model that has a conditional sparsity feature, and propose the POET estimator to take advantage of the structure. This expands considerably the scope of the model based on the strict factor model, which assumes independent idiosyncratic noise and is too restrictive in practice. By assuming sparse error covariance matrix, we allow for the presence of the cross-sectional correlation even after taking out the common factors. The sparse covariance is estimated by the adaptive thresholding technique.
It is found that the rates of convergence of the estimators have an extra term approximately in addition to the results based on observable factors by Fan et al. (2008, 2011), which arises from the effect of estimating the unobservable factors. As we can see, this effect vanishes as the dimensionality increases, as more information about the common factors becomes available. When gets large enough, the effect of estimating the unknown factors is negligible, and we estimate the covariance matrices as if we knew the factors.
The proposed POET also has wide applicability in statistical genomics. For example, Carvalho et al. (2008) applied a Bayesian sparse factor model to study the breast cancer hormonal pathways. Their real-data results have identified about two common factors that have highly loaded genes (about half of 250 genes). As a result, these factors should be treated as “pervasive” (see the explanation in Example 2.1), which will result in one or two very spiked eigenvalues of the gene expressions’ covariance matrix. The POET can be applied to estimate such a covariance matrix and its network model.
Appendix A Estimating a sparse covariance with contaminated data
We can estimate using the adaptive thresholding proposed by Cai and Liu (2011): for the threshold define
Suppose , and Assumptions 3.2 and 3.3 hold. In addition, suppose there is a sequence so that and Then there is a constant in the adaptive thresholding estimator (A.1) with
If further , then is invertible with probability approaching one, and
By Assumptions 3.2 and 3.3, the conditions of Lemmas A.3 and A.4 of Fan, Liao and Mincheva (2011, Ann. Statist, 39, 3320-3356) are satisfied. Hence for any , there are positive constants and such that each of the events
occurs with probability at least . By the condition of threshold function, . Now for under the event
Let Then with probability at least , Since is arbitrary, we have . If in addition, , then the minimum eigenvalue of is bounded away from zero with probability approaching one since . This then implies
Appendix B Proofs for Section 2
We first cite two useful theorems, which are needed to prove propositions 2.1 and 2.2. In Lemma B.1 below, let be the eigenvalues of in descending order and \{\mbox{\boldmath\xi}_{i}\}_{i=1}^{p} be their associated eigenvectors. Correspondingly, let be the eigenvalues of in descending order and \{\widehat{}\mbox{\boldmath\xi}_{i}\}_{i=1}^{p} be their associated eigenvectors.
(Weyl’s Theorem) |\widehat{\lambda}_{i}-\lambda_{i}|\leq\|\widehat{\mathbf{\Sigma}}-\mbox{\boldmath\Sigma}\|.
( Theorem, Davis and Kahan, 1970)
Appendix C Proofs for Section 3
We will proceed by subsequently showing Theorems 3.3, 3.1 and 3.2.
The following results are to be used subsequently. The proofs of Lemmas C.1,C.2 and C.3 are found in Fan, Liao and Mincheva (2011).
Suppose that the random variables both satisfy the exponential-type tail condition: There exist , and , such that ,
Then for some and , and any ,
Under the assumptions of Theorem 3.1, (i) . (ii) (iii)
First of all, by Proposition 2.1, under Assumption 3.1, the largest eigenvalue of satisfies: for some
C.2 Proof of Theorem 3.3
Our derivation below relies on a result obtained by Bai and Ng (2002), which showed that the estimated number of factors is consistent, in the sense that equals the true with probability approaching one. Note that under our Assumptions 3.1-3.4, all the assumptions in Bai and Ng (2002) are satisfied. Thus immediately we have the following Lemma.
Using (A.1) in Bai (2003), we have the following identity:
(i) We have , . By the Cauchy-Schwarz inequality,
Note that By Assumption 3.4, , which implies that , and yields the result.
(iv) Similar to part (iii), noting that is a scalar, we have:
where the third line follows from the Cauchy-Schwarz inequality. ∎
The result then follows from Assumption 3.3.
It follows from Assumption 3.4 that It then follows from the Chebyshev’s inequality and Bonferroni’s method that .
We prove this lemma conditioning on the event Once this is done, due to , it then implies the unconditional arguments.
Each of the four terms on the right hand side above are bounded in Lemma C.7, which then yields the desired result.
Part (iii) is implied by (C.2) and Lemma C.8. ∎
C.3 Proof of Theorem 3.1
and
The first part of the lemma then follows from Theorem 3.3 and Lemma C.9. The second part follows from Corollary 3.1.
Proof of Theorem 3.1: The theorem follows immediately from Theorem A.1 and Lemma C.11.
C.4 Proof of Theorem 3.2
(ii) The result follows from part (i) and Lemma C.13. Part (iii) and (iv) follow from a similar argument of part (i) and Lemma C.10.
We derive the rate for Define
One the other hand, using Sherman-Morrison-Woodbury formula again implies
Proof of Theorem 3.2:
On the other hand, let be the entry of . Then .
Hence The result then follows immediately. ∎