Optimal whitening and decorrelation
Agnan Kessy, Alex Lewin, Korbinian Strimmer
Introduction
Whitening, or sphering, is a linear transformation that converts a -dimensional random vector with mean and positive definite covariance matrix into a new random vector
of the same dimension and with unit diagonal “white” covariance . The square matrix is called the whitening matrix. As orthogonality among random variables greatly simplifies multivariate data analysis both from a computational and a statistical standpoint, whitening is a critically important tool, most often employed in preprocessing but also as part of modeling (e.g. Zuber and Strimmer,, 2009; Hao et al.,, 2015).
Whitening can be viewed as a generalization of standardizing a random variable which is carried out by
where the matrix contains the variances . This results in but it does not remove correlations. Often, standardization and whitening transformations are also accompanied by mean-centering of or to ensure , but this is not actually necessary for producing unit variances or a white covariance.
The whitening transformation defined in Equation (1) requires the choice of a suitable whitening matrix . Since it follows that and thus , which is fulfilled if satisfies the condition
However, unfortunately, this constraint does not uniquely determine the whitening matrix . Quite the contrary, given there are in fact infinitely many possible matrices that all satisfy Equation (3), and each leads to a whitening transformation that produces orthogonal but different sphered random variables.
This raises two important issues: first, how to best understand the differences among the various sphering transformations, and second, how to select an optimal whitening procedure for a particular situation. Here, we propose to address these questions by investigating the cross-covariance and cross-correlation matrix between and . As a result, we identify five natural whitening procedures, of which we recommend two particular approaches for general use.
Notation and useful identities
In the following, we will make use of a number of covariance matrix identities: the decomposition of the covariance matrix into the correlation matrix and the diagonal variance matrix , and the eigendecomposition of the covariance matrix and the eigendecomposition of the correlation matrix , where , contain the eigenvectors and , the eigenvalues of , respectively. We will frequently use , the unique inverse matrix square root of , as well as , the unique inverse matrix square root of the correlation matrix.
Following the standard convention we assume that the eigenvalues are sorted in order from largest to smallest value. In addition, we recall that by construction all eigenvectors are defined only up to a sign, i.e. the columns of and can be multiplied with a factor of and the resulting matrix is still valid. Indeed, using different numerical algorithms and software will often result in eigendecompositions with and showing diverse column signs.
Rotational freedom in whitening
The constraint Equation (3) on the whitening matrix does not fully identify but allows for rotational freedom. This becomes apparent by writing in its polar decomposition
where is an orthogonal matrix with . Clearly, satisfies Equation (3) regardless of the choice of .
This implies a geometrical interpretation of whitening as a combination of multivariate rescaling by and rotation by . It also shows that all whitening matrices have the same singular values , which follows from the singular value decomposition with orthogonal. This highlights that the fundamental rescaling is via the square root of the eigenvalues . Geometrically, the whitening transformation with is a rotation followed by scaling, possibly followed by another rotation (depending on the choice of ).
Since in many situations it is desirable to work with standardized variables another useful decomposition of that also directly demonstrates the inherent rotational freedom is
where is a further orthogonal matrix with . Evidently, this also satisfies the constraint of Equation (3) regardless of the choice of .
In this view, with , the variables are first scaled by the square root of the diagonal variance matrix, then rotated by , then scaled again by the square root of the eigenvalues of the correlation matrix, and possibly rotated once more (depending on the choice of ).
For the above two representations to result in the same whitening matrix two different rotations and are required. These are linked by where the matrix is itself orthogonal. Since the eigendecompositions of the covariance and the correlation matrix are not readily related to each other, the matrix can unfortunately not be further simplified.
Cross-covariance and cross-correlation
For studying the properties of the different whitening procedures we will now focus on two particularly useful quantities, namely the cross-covariance and cross-correlation matrix between the whitened vector and the original random vector . As it turns out, these are closely linked to the rotation matrices and encountered in the above two decompositions of (Equation (4) and Equation (5)).
The cross-covariance matrix between and is given by
Likewise, the cross-correlation matrix is
Thus, we find that the rotational freedom inherent in , which is represented by the matrices and , is directly reflected in the corresponding cross-covariance and cross-correlation between and . This provides the leverage that we will use to select and discriminate among whitening transformations by appropriately choosing or constraining or .
As can be seen from Equation (6) and Equation (7), both and are in general not symmetric, unless or , respectively. Note that the diagonal elements of the cross-correlation matrix need not be equal to 1.
Furthermore, since each is perfectly explained by a linear combination of the uncorrelated , and hence the squared multiple correlation between and equals 1. Thus, the column sum over the squared cross-correlations is always 1. In matrix notation, . In contrast, the row sum of over the squared cross-correlations varies for different whitening procedures, and is, as we will see below, highly informative for choosing relevant transformations.
Five natural whitening procedures
In practical application of whitening there are a handful of sphering procedures that are most commonly used (e.g. Li and Zhang,, 1998). Accordingly, in Table (1) we describe the properties of five whitening transformations, listing the respective sphering matrix , the associated rotation matrices and , and the resulting cross-covariances and cross-correlations . All five methods are natural whitening procedures arising from specific constraints on or , as we will show further below.
The ZCA whitening transformation employs the sphering matrix
where ZCA stands for “zero-phase components analysis” (Bell and Sejnowski,, 1997). This procedure is also known as Mahalanobis whitening. With it is the unique sphering method with a symmetric whitening matrix.
PCA whitening is based on scaled principal component analysis (PCA) and uses
(e.g. Friedman,, 1987). This transformation first rotates the variables using the eigenmatrix of the covariance as is done in standard PCA. This results in orthogonal components, but with in general different variances. To achieve whitened data the rotated variables are then scaled by the square root of the eigenvalues . PCA whitening is probably the most widely applied whitening procedure due to its connection with PCA.
It can be seen that the PCA and ZCA whitening transformations are related by a rotation , so ZCA whitening can be interpreted as rotation followed by scaling followed by the rotation back to the original coordinate system. The ZCA and the PCA sphering methods both naturally follow the polar decomposition of Equation (4), with equal to and respectively.
Due to the sign ambiguity of eigenvectors the PCA whitening matrix given by Equation (9) is still not unique. However, adjusting column signs in such that , i.e. that all diagonal elements are positive, results in the unique PCA whitening transformation with positive diagonal cross-covariance and cross-correlation (cf. Table (1)).
Another widely known procedure is Cholesky whitening which is based on Cholesky factorization of the precision matrix . This leads to the sphering matrix
where is the unique lower triangular matrix with positive diagonal values. The same matrix can also be obtained from a QR decomposition of .
A further approach is the ZCA-cor whitening transformation, which is used, e.g., in the CAT (correlation-adjusted -score) and CAR (correlation-adjusted marginal correlation) variable importance and variable selection statistics (Zuber and Strimmer,, 2009; Ahdesmäki and Strimmer,, 2010; Zuber and Strimmer,, 2011; Zuber et al.,, 2012). ZCA-cor whitening employs
as its sphering matrix. It arises from first standardizing the random variable by multiplication with and subsequently employing ZCA whitening based on the correlation rather than covariance matrix. The resulting whitening matrix differs from , and unlike the latter it is in general asymmetric.
In a similar fashion, PCA-cor whitening is conducted by applying PCA whitening to standardized variables. This approach uses
as its sphering matrix. Here, the standardized variables are rotated by the eigenmatrix of the correlation matrix, followed by scaling using the correlation eigenvalues. Note that differs from .
PCA-cor whitening has the same relation to the ZCA-cor transformation as does PCA whitening to the ZCA transformation. Specifically, ZCA-cor whitening can be interpreted as PCA-cor whitening followed by a rotation back to the frame of the standardized variables. Both the ZCA-cor and the PCA-cor transformation naturally follow the decomposition of Equation (5), with equal to and respectively.
Similarly as in PCA whitening, the PCA-cor whitening matrix given by Equation (12) is subject to sign ambiguity of the eigenvectors in . As above, setting leads to the unique PCA-cor whitening transformation with positive diagonal cross-covariance and cross-correlation (cf. Table (1)).
Finally, we may also apply the Cholesky whitening transformation to standardized variables. However, this does not lead to a new whitening procedure, as the resulting sphering matrix remains identical to since the Cholesky factor of the inverse correlation matrix is , and therefore .
Optimal whitening
We now demonstrate how an optimal sphering matrix , and hence an optimal whitening approach, can be identified by evaluating suitable objective functions computed from the cross-covariance and cross-correlation . Intriguingly, for each of the five natural whitening transforms listed in Table (1) we find a corresponding optimality criterion.
In many applications of whitening it is desirable to remove correlations with minimal additional adjustment, with the aim that the transformed variable remains as similar as possible to the original vector .
One possible implementation of this idea is to find the whitening transformation that minimizes the total squared distance between the original and whitened variables (e.g. Eldar and Oppenheim,, 2003). Using mean-centered random vectors and with and this least squares objective can be expressed as
Since the dimension and sum of the variances do not depend on the whitening matrix minimizing Equation (13) is equivalent to maximizing the trace of the cross-covariance matrix
Maximization of uniquely determines the optimal whitening matrix to be the symmetric sphering matrix .
since is diagonal. As and are both orthogonal is also orthogonal. This implies diagonal entries , with equality signs for all occurring only if , hence the maximum of is assumed at , or equivalently at . From Equation (4) it follows that the corresponding optimal sphering matrix is . ∎
For related proofs see also Johnson, (1966), Genizi, (1993, p. 412) and Garthwaite et al., (2012, p. 789).
As a result, we find that ZCA-Mahalanobis whitening is the unique procedure that maximizes the average cross-covariance between each component of the whitened and original vectors. Furthermore, with it is also the unique whitening procedure with a symmetric cross-covariance matrix .
2 ZCA-cor whitening
In the optimization using Equation (13) the underlying similarity measure is the cross-covariance between the whitened and original random variables. This results in an optimality criterion that depends on the variances and hence on the scale of the original variables. An alternative scale-invariant objective can be constructed by comparing the centered whitened variable with the centered standardized vector . This leads to the minimization of
Equivalently, we can maximize instead the trace of the cross-correlation matrix
Maximization of uniquely determines the whitening matrix to be the asymmetric sphering matrix .
Completely analogous to Proposition 1, we can write where is orthogonal. By the same argument as before it follows that maximizes . From Equation (5) it follows that . ∎
As a result, we identify ZCA-cor whitening as the unique procedure that ensures that the components of the whitened vector remain maximally correlated with the corresponding components of the original variables . In addition, with it is also the unique whitening transformation exhibiting a symmetric cross-correlation matrix .
3 PCA whitening
Another frequent aim in whitening is the generation of new uncorrelated variables that are useful for dimension reduction and data compression. In other words, we would like to construct components such that the first few components in represent as much as possible the variation present in the all original variables .
One way to formalize this is to use the row sum of squared cross-covariances between each individual and all as a measure of how effectively each integrates, or compresses, the original variables. Note that here, unlike in ZCA-Mahalanobis whitening, the objective function links each component in simultaneously with all components in . In vector notation the can be more elegantly written as
Our aim is to find a whitened vector such that the are maximized with .
Maximization of subject to monotonically decreasing is achieved by the whitening matrix .
The vector can be written as . Setting we arrive at , i.e. for this choice the corresponding to each component of are equal to the corresponding eigenvalues of . As the eigenvalues are already sorted in decreasing order, we find (cf. Table (1)) that whitening with leads to a sphered variable with monotonically decreasing . For general the -th element of is where is orthogonal. This is maximized when , or equivalently, . ∎
As a result, PCA whitening is singled out as the unique sphering procedure that maximizes the integration, or compression, of all components of the original vector in each component of the sphered vector based on the cross-covariance as underlying measure. Thus, the fundamental property of PCA that principal components are optimally ordered with respect to dimension reduction (Jolliffe,, 2002) carries over also to PCA whitening.
4 PCA-cor whitening
For reasons of scale-invariance we prefer to optimize cross-correlations rather than cross-covariances for whitening with compression in mind. This leads to the row sum of squared cross-correlation as measure of integration and compression, and correspondingly to the objective function
Maximization of subject to monotonically decreasing is achieved by using as the sphering matrix.
Analogous to the proof of Proposition 3 we find to yield optimal and decreasing and with Equation (5) we arrive at . ∎
Hence, the PCA-cor whitening transformation is the unique transformation that maximizes the integration, or compression, of all components of the original vector in each component of the sphered vector employing the cross-correlation as underlying measure.
5 Cholesky whitening
Finally, we investigate the connection between Cholesky whitening and corresponding characteristics of the cross-covariance and cross-correlation matrices. Unlike the other four whitening methods listed in Table (1), which result from optimization, Cholesky whitening is due to a symmetry constraint.
Specifically, the whitening matrix leads to a cross-covariance matrix that is lower-triangular with positive diagonal elements as well as to a cross-correlation matrix with the same properties. This is a consequence of the Cholesky factorization with being subject to the same constraint. Crucially, as is unique the converse argument is valid as well, and hence Cholesky whitening is the unique whitening procedure that results from lower-triangular positive diagonal cross-covariance and cross-correlation matrices.
A consequence of using Cholesky factorization for whitening is that we implicitly assume an ordering of the variables. This can be useful specifically in time course analysis to account for auto-correlation (cf. Pourahmadi,, 2011, and references therein).
Application
In the sections above, we have discussed the theoretical background of whitening in terms of random variables and and using the population covariance to guide the construction of a suitable sphering matrix .
In practice, however, we frequently need to whiten data rather than random variables. In this case we have an data matrix whose rows are assumed to be drawn from a distribution with expectation and covariance matrix . In this setting the transformation of Equation (1) from original to whitened data matrix becomes .
A further complication is that the covariance matrix is often unknown. Accordingly, it needs to be learned from data, either from or from another suitable data set, yielding a covariance matrix estimate . Typically, for large sample size and small dimension the standard unbiased empirical covariance with is used. In high-dimensional cases with the empirical estimator breaks down, and the covariance matrix needs to be estimated by a suitable regularized method instead (e.g. Schäfer and Strimmer,, 2005; Pourahmadi,, 2011). Finally, from the spectral decomposition of the estimated covariance or corresponding correlation matrix we then obtain the desired estimated whitening matrix .
2 Iris flower data example
For an illustrative comparison of the five natural whitening transforms discussed in this paper and listed in Table (1) we applied them on the well-known iris flower data set of Anderson reported in Fisher, (1936), which comprises correlated variables (: sepal length, : sepal width, : petal length, : petal width) and observations.
The results are shown in Table (2) with all estimates based on the empirical covariance . For the PCA and PCA-cor whitening transformation we have set and , respectively. The upper half of Table (2) shows the estimated cross-correlations between each component of the whitened and original vector for the five methods, and the lower half the values of the various objective functions discussed above.
As expected, the ZCA and the ZCA-cor whitening produce sphered variables that are most correlated to the original data on a component-wise level, with the former achieving the best fit for the covariance-based and the latter for the correlation-based objective.
In contrast, the PCA and PCA-cor methods are best at producing whitened variables that are maximally simultaneously linked with all components of the original variables. Consequently, as can be seen from the top half of Table (2), for PCA and PCA-cor whitening only the first two components and are highly correlated with their respective counterparts and , whereas the subsequent pairs and are effectively uncorrelated. Furthermore, the last line of Table (2) shows that PCA-cor whitening achieves higher maximum total squared correlation of the first component with all components of than PCA whitening, indicating better compression.
Interestingly, Cholesky whitening always assumes third place in the rankings, either behind ZCA and ZCA-cor whitening, or behind PCA and PCA-cor whitening. Moreover, it is the only approach where by construction one pair () perfectly correlates between whitened and original data.
Conclusion
In this note we have investigated linear transformations for whitening of random variables. These methods are commonly employed in data analysis for preprocessing and to facilitate subsequent analysis.
In principle, there are infinitely many possible whitening procedures all satisfying the fundamental constraint of Equation (3) for the underlying whitening matrix. However, as we have demonstrated here, the rotational freedom inherent in whitening can be broken by considering cross-covariance and cross-correlations between whitened and original variables.
Specifically, we have studied five natural whitening transforms, cf. Table (1), all of which can be interpreted as either optimizing a suitable function of or , or satisfying a symmetry constraint on or . As a result, this not only leads to a better understanding of the differences among whitening methods, but also enables an informed choice.
In particular, selecting a suitable whitening transformation depends on the context of application, specifically whether minimal adjustment or compression of data is desired. In the former the whitened variables remain highly correlated to the original variables, and thus maintain their original interpretation. This is advantageous for example in the context of variable selection where one would like to understand the resulting selected submodel. In contrast, in a compression context the whitened variables by construction bear no interpretable relation to the original data but instead reflect their intrinsic effective dimension.
In general, we advocate using scale-invariant optimality functions and thus recommend using cross-correlation as a basis for optimization. Consequently, we particularly endorse two specific whitening approaches. If the aim is to obtain sphered variables that are maximally similar to the original ones, we suggest to employ the ZCA-cor whitening procedure of Equation (11). Conversely, if maximal compression is desirable we recommend to use the PCA-cor whitening approach of Equation (12).