RobustNet: Improving Domain Generalization in Urban-Scene Segmentation via Instance Selective Whitening

Sungha Choi, Sanghun Jung, Huiwon Yun, Joanne Kim, Seungryong Kim, Jaegul Choo

Introduction

When deploying deep neural networks (DNNs) trained on a given dataset (i.e., source domain) in real-world unseen data (i.e., target domain), DNNs often fail to perform properly due to the domain shift. Overcoming this issue is crucial, especially for safety-critical applications such as autonomous driving. In particular, real-world data consist of unexpected and unseen samples, for example, those images taken under diverse illumination, adverse weather conditions, or from different locations. It is generally impossible to model such a full data distribution with limited training data, so reducing the domain gap between source and target domains has been a long-standing problem in computer vision.

Domain adaptation (DA) is an approach to mitigate the performance degradation caused by such a domain gap . Generally, DA focuses on adapting the source domain distribution to that of the target domain, but it requires access to the samples in the target domain, which limits their applicability. When we set the entire real world as a target domain, it is difficult in pactice to obtain data samples that fully cover the target domain.

Domain generalization (DG) overcomes this limitation by improving the robustness of DNNs to arbitrary unseen domains. In general, most DG methods accomplish this through the learning of a shared representation across multiple source domains. However, collecting such multi-domain datasets is costly and labor-intensive, and furthermore, the performance highly depends on the number of source datasets.

A recent study has shown that the DG problem can be addressed by exploiting instance normalization layers instead of relying on multiple source domains, leading to a simple and cost-effective training process. The instance normalization just standardizes features while not considering the correlation between channels. However, a number of studies claim that feature covariance contains domain-specific style such as texture and color. This implies that applying instance normalization to the networks may not be sufficient for domain generalization, because the feature covariance is not considered. A whitening transformation is a technique that removes feature correlation and makes each feature have unit variance. It has been proven that the feature whitening effectively eliminates domain-specific style information as shown in image translation , style transfer , and domain adaptation , and thus it may improve the generalization ability of the feature representation, but not yet fully explored in DG. However, simply adopting the whitening transformation to improve the robustness of DNNs is not straightforward, since it may eliminate domain-specific style and domain-invariant content at the same time. Decoupling the two factors and selectively removing the domain-specific style is the main scope of this paper.

In this paper, we present an instance selective whitening loss that alleviates the limitations of the existing whitening transformation for domain generalization, by selectively removing information that causes a domain shift while maintaining a discriminative power of feature within DNNs. Our method does not rely on an explicit closed-form whitening transformation, but implicitly encourage the networks to learn such a whitening transformation through the proposed loss function, thus requiring negligible computational cost. As illustrated in Fig. 1, our method selectively removes only those feature covariances that respond sensitively to photometric augmentation such as color transformation. Our experiments on urban-scene segmentation in DG settings, performed using several backbone networks, show evidence that our approach consistently boosts the DG performance.

The main contributions include the following:

We propose an instance selective whitening loss for domain generalization, which disentangles domain-specific and domain-invariant properties from higher-order statistics of the feature representation and selectively suppresses domain-specific ones.

Our proposed loss can easily be used in existing models and significantly improves the generalization ability with negligible computational cost.

We apply the proposed loss to urban-scene segmentation in a DG setting and show the superiority of our approach over existing approaches in both a qualitative and quantitative manner.

Related Work

It is well known that significant labeling efforts are required so as to ensure the reliable performance of various tasks such as semantic segmentation . To tackle this challenge, domain adaptation (DA) methods were proposed to transfer the knowledge learned from abundant labeled data (i.e., a source domain) to a target domain where labeled data are scarce. In contrast to DA, domain generalization (DG) methods assume that the model cannot access the target domain during training and aim to improve the generalization ability to perform well in an unseen target domain. Various approaches such as meta-learning , adversarial training , autoencoder , metric learning , data augmentation have been proposed to learn domain-agnostic feature representations. Recently, several studies have shown the effectiveness of exploiting both batch normalization (BN) and instance normalization (IN) within DNNs to solve the DG problem. These studies show that BN improves discriminative ability on features, while IN prevents overfitting on training data, so that generalization performance is improved on unseen domains by combining BN and IN. Especially, IBN-Net shows a significant performance improvement with the marginal architectural modification that incorporates the IN layers through training on a single source domain, unlike most DG methods that require multiple source domains. This normalization based DG method is attractive because it can be applied as a complement to other DG methods based on multiple source domains.

Based on the synthetic data such as GTAV and SYNTHIA , numerous DA studies have been proposed in semantic segmentation, but only a few DG studies address semantic segmentation, as the majority of the DG methods mainly focused on image classification. DA, which can access the target domains, generally has better performance than DG, but DG methods that can handle an arbitrary unseen domain without access to the target domain are mandatory in the real world. This paper focuses on the DG method practically helpful in semantic segmentation where various conditions exist such as adverse weather, diverse illumination, location differences, and so on.

The seminal studies have demonstrated that feature correlations (i.e., a gram matrix or covariance matrix) take style information of images. Since then, numerous studies exploit the feature correlation in style transfer , image-to-image translation , domain adaptation and networks architecture . Especially, the whitening transformation that removes feature correlation and makes each feature have unit variance, has been known to help to remove the style information from the feature representations . Our work explores the whitening transformation to improve domain generalization performance. To the best of our knowledge, this is the first attempt to apply whitening to DG.

Preliminaries

It has been known that WT can effectively remove style information by being applied to each instance in style transfer .

We can compute the whitening transformation matrix Σμ−12\mathbf{\Sigma}_{\mu}^{-\frac{1}{2}} analytically through Eq. (4), but eigenvalue decomposition is computationally expensive, leading to slow training and inference speed and prevents the gradient back-propagation . To alleviate these problems, previous studies have shown that the goal of WT can be achieved without the eigen-decomposition through the whitening loss or approximating the whitening transformation matrix using Newton’s iteration .

Especially, GDWCT proposes the deep whitening transformation (DWT) that implicitly makes the covariance matrix Σμ\mathbf{\Sigma}_{\mu} close to the identity matrix I\mathbf{I} by means of the loss defined as

Proposed Method

This section presents our approach to solve the domain generalization problem through whitening the feature representation by mitigating undesirable effects of a whitening transformation. Our method disentangles the covariance into the encoded style and content so that only the style information can be selectively removed, thus increasing the domain generalization ability. We firstly propose an instance whitening and instance-relaxed loss in Section 4.1 and then finally propose our novel instance selective whitening loss in Section 4.3.

This subsection describes a series of steps to transform the input feature into the whitening transformed feature as shown in Fig. 2. Note that our method is applied to each instance, not to a mini-batch. Let Σμ (i,i)\mathbf{\Sigma}_{\mu\,(i,i)} denote a diagonal element (i,ii,i) and Σμ (i,j)\mathbf{\Sigma}_{\mu\,(i,j)} denote an off-diagonal element (i,ji,j) of the covariance matrix Σμ\mathbf{\Sigma}_{\mu} of the intermediate feature map, where 0≤i,j<C0\leq i,j<C, i≠ji\neq j. The DWT loss in Eq. (5) can be decomposed as

To address this issue, the feature map X\mathbf{X} can first be standardized into Xs\mathbf{X_{s}} through an instance normalization :

After standardization of the intermediate feature map, the covariance matrix is calculated as

where Xs\mathbf{X_{s}} is the standardized feature map. Thanks to the standardization process, diagonal elements of the covariance matrix are already set as unit values. Thus, we only need to make the off-diagonals of the covariance matrix close to zero, which makes it easy to optimize for the whitening process, and the aforementioned conflict can thus be resolved. Since the covariance matrix is symmetric, the loss can be applied only to the strict upper triangular part. Our instance whitening (IW) loss is formulated as

2 Margin-based relaxation of whitening loss

The instance whitening loss (Eq. (10)) suppresses all covariance elements to zero, so it can adversely affect the discriminative power of features within DNNs. To address this issue, we propose an instance-relaxed whitening (IRW) loss to sustain the covariance elements essential in maintaining the discriminative power. The IRW loss is designed so that the expected value of the total covariance lies within a specified margin δ\delta rather than being close to zero, i.e.,

The loss LIRW\mathcal{L}_{\text{IRW}} allows the covariance to have a certain level of values, so it gives room to keep discriminative features intact. The empirical effect of the IRW loss can be found in Section 5.2.1. It shows better performance compared to the IW loss not including margin δ\delta (Eq. (10)). Nonetheless, it may not be sufficient because we cannot guarantee that only the covariance useful for generalization performance remains through the margin relaxation.

3 Separating Covariance Elements

To further improve our approach, we need to separate the covariance terms into two groups: domain-specific style and domain-invariant content. We propose to selectively suppress only the style-encoded covariances that cause the domain shift. Assuming that the domain shift includes changes in color and blurriness, we simulate the domain shift through photometric augmentation such as color jittering and Gaussian blurring.

from mean μΣi\boldsymbol{\mu}_{\mathbf{\Sigma}_{i}} and variance σi2\boldsymbol{\sigma}_{i}^{2} for each element from two different covariance matrices of the ii-th image, i.e.,

where NN denotes the number of image samples, xix_{i} is the ii-th image sample, τ\tau is a photometric transformation, and Σs(⋅)\mathbf{\Sigma_{s}}(\cdot) extracts the covariance matrix of the intermediate feature map from an input image. As a result, V\mathbf{V} consists of elements of the variance of each covariance element across various photometric transformations.

We assume that the variance matrix V implies the sensitivity of the corresponding covariance to the photometric transformation. This means that the covariance elements with high variance value contain the domain-specific style such as color and blurriness. To identify such elements, we apply k-means clustering on the strict upper triangular elements Vi,j  (i<j)\mathbf{V}_{i,j}\;(i<j) of the variance matrix V\mathbf{V} to assign the elements into kk clusters C={c1,c2,…,ck}C=\{c_{1},c_{2},\dots,c_{k}\} with respect to the value. Next, we split the kk clusters into two groups, Glow={c1,…,cm}G_{low}=\{c_{1},\dots,c_{m}\} with low variance value and Ghigh={cm+1,…,ck}G_{high}=\{c_{m+1},\dots,c_{k}\} with high variance value. The hyper-parameters kk and mm are empirically set to 3 and 1, respectively. More details can be found in the supplementary Section A.2. We assume that GhighG_{high} contains the domain-specific style and GlowG_{low} contains domain-invariant content.

The networks continue training for the remaining epochs incorporating the proposed ISW loss.

4 Network architecture with proposed ISW loss

IBN-Net has explored a number of ResNet -based architectures to combine instance normalization with batch normalization and proposed several IBN blocks based on a residual block (Fig. 4(a)). Among the proposed blocks, IBN-b, which adds an instance normalization layer right after the addition operation of a residual block (Fig. 4(b)), shows the best generalization performance on semantic segmentation tasks. After all, they add three instance normalization layers after the first three convolution groups (i.e., conv1, conv2_x, and conv3_x). We follow this architectural approach as our baseline. As shown in Fig. 4(d), we simply add our proposed ISW loss to the instance normalization layer. Our loss in total is described as

where λ\lambda denotes the weight of our ISW loss and is empirically set to 0.6, Ltask\mathcal{L}_{\text{task}} is the task loss (e.g., a per-pixel cross-entropy loss for semantic segmentation), ii indicates the layer index, and LL is the number of layers to which the ISW loss is applied. The hyper-parameter λ\lambda is analyzed in the supplementary Section A.2. LL is set to three by following IBN-Net. An affine transformation is not used since the subsequent convolution operation after a whitening transformation can do the equivalent job, and empirically, we found no performance gain by explicitly adding the affine transformation.

Experiments

This section describes the experimental setup and presents evaluation results to assess the effectiveness of our proposed methods on semantic segmentation with comparison to other methods. Furthermore, we provide an in-depth analysis of our results including the covariance matrices.

We train our model on several datasets (e.g., Cityscapes) and show its performance on other datasets (e.g., BDD-100K, Mapillary, GTAV, and SYNTHIA) to measure the generalization capability on unseen domains. For fair comparisons with other normalization techniques, we re-implement IBN-Net and IterNorm on our baseline models and compare them with our methods. As described in Section 4.4, our proposed loss can easily be added to existing models, so we apply our methods to various backbone networks such as ResNet , ShuffleNetV2 and MobileNetV2 and show wide applicability of the proposed methods. For all the quantitative experiments, mean Intersection over Union (mIoU) is used to measure the segmentation performance.

We adopt DeepLabV3+ for a semantic segmentation architecture, and SGD optimizer with an initial learning rate of 1e-2 and momentum of 0.9 is used. Besides, we follow the polynomial learning rate scheduling with the power of 0.9. We train all the models for 40K iteration, except for multi-source models, which are trained for 110K iterations. To prevent the model from overfitting, color and positional augmentations such as color jittering, Gaussian blur, random cropping, random horizontal flipping, and random scaling with the range of [0.5, 2.0] are conducted. For the photometric transformation in ISW, we apply color jittering and Gaussian blur. Also, as suggested by IBN-Net, we add three instance normalization layers after the first three convolution groups and apply our proposed loss. Further details are provided in the supplementary Section A.3.

1.2 Datasets

To verify the generalization capability of our methods, we conduct the experiments on five different datasets.

Cityscapes is a large-scale dataset containing high-resolution (e.g., 2048×\times1024) urban scene images collected from 50 different cities in primarily Germany. It provides 3,450 finely-annotated images and 20,000 coarsely-annotated images. We use only a finely-annotated set for training and validation. BDD-100K is another real-world dataset that contains diverse urban driving scene images with the resolution of 1280×\times720. The images are collected from various locations in the US. For a semantic segmentation task, 7,000 training and 1,000 validation images are provided. The last real-world dataset we use is Mapillary , a diverse street-view dataset consisting of 25,000 high-resolution images with a minimum resolution of 1920×\times1080 collected from all around the world.

GTAV is a large-scale dataset containing 24,966 driving-scene images generated from Grand Theft Auto V game engine. It has 12,403, 6,382, and 6,181 images of size 1914×\times1052 for a train, a validation, and a test set, respectively. It has 19 object categories compatible with Cityscapes. Also, we use SYNTHIA , composed of photo-realistic synthetic images containing 9,400 samples with a resolution of 960×\times720.

2 Quantitative Evaluation

This subsection provides ablation studies, the comparisons of our results against other normalization methods, the evaluation on multiple source domains, and the analysis of computational cost. Since the experiments follow domain generalization settings, the model cannot access any datasets other than the source data.

To verify the effectiveness of our methods, we conduct comparisons with other normalization methods and ablation studies on instance whitening (IW), instance-relaxed whitening (IRW), and instance selective whitening (ISW). Note that all the experiments in this subsection are performed three times and averaged for fair comparisons.

Table 1 shows the generalization performance of the models trained on GTAV dataset. ISW outperforms other methods on all datasets except the source dataset (i.e., GTAV). Especially, ISW shows a significant improvement on real-world datasets (i.e., Cityscapes, BDD-100K, and Mapillary). Table 2 shows the generalization performance of those models trained on Cityscapes dataset. Although IterNorm outperforms our models on GTAV, the performance gap is minimal. ISW outperforms other normalization and baseline models on BDD-100K, Mapillary, and SYNTHIA datasets.

Baseline, Switchable Whitening (SW), and IBN-Net, which are less generalizable than our method, tend to overfit the source domain, suffering from performance degradation on the target domain due to the large domain shift. Our method may sacrifice the performance on the source domains (i.e., training and evaluating on the same dataset) as shown in the last column in Table 1 and 2. However, our models shows good generalizability, which is critical when deployed in the wild, where large domain-shift is expected.

Table 3 explains the wide applicability of our work. The first group is reported by adopting ShuffleNetV2, and the second group is using MobileNetV2 as backbone networks. In both cases, our model with ISW outperforms the baseline and IBN-Net on real-world datasets. To further validate the capability of our method, we present the comparison with baselines trained on multiple synthetic domains, GTAV, and SYNTHIA. For the training, we aggregate the training domains without any joint training methodologies. Learning domain-invariant features across multiple datasets is essential to optimize the model on different distributions of multiple datasets. Table 4 shows our model trained on multiple datasets performs better than other models due to its generalization ability by extracting domain-invariant features during training.

2.2 Comparison with other DG and DA methods

This subsection compares our method with two existing DG methods on semantic segmentation task, based on the results reported in the papers . DRPC proposes a domain randomization method, which maps the synthetic images to multiple auxiliary real domains using image-to-image translation with the style of real images (e.g., ImageNet). As shown in Table 5, our model gains the largest performance increase on average, compared to other methods such as IBN-Net and DRPC . Our method shows a large amount of performance improvement on BDD-100K and Mapillary datasets that involve significantly more diverse driving scenes than Cityscapes.

In addition, we compare the result of our method with those reported from several domain adaptation methods. See the supplementary Section A.1.

2.3 Computational cost analysis

To ensure our method requires no additional computational cost, we report the number of parameters, GFLOPS, and inference time. As seen in Fig. 4, all the models in Table 6 share the same network architecture, but with different normalization methods. As shown in Table 6, our approach performs a whitening transformation without additional computational cost.

3 Qualitative Analysis

To show how the covariance matrix is selectively whitened, we visualize the covariance matrix of intermediate feature maps from IBN-Net and our model with ISW. As shown in Fig. 5, the first pair of covariance matrices are from the first convolution layer and the others are from the second convolution layer. Note that the style information mainly exists in the early layers of the network as pointed out in IBN-Net. Moreover, the style information is encoded as a form of the features covariance as revealed in previous studies . Hence, the covariance matrices are sparser at the second pair, compared to the first ones. By comparing the covariance maps from IBN-Net and ISW, we can find the ones from ours are whitened but a small number of covariance elements remain large, showing our ISW selectively eliminates the covariance.

For in-depth analysis, we reconstruct input images from the whitened feature maps of our ISW model. For the experiment, we adopt U-Net as reconstruction networks. To newly train a decoder, we append the decoder to the backbone of a pre-trained baseline and train the decoder. We then replace the backbone network with the pre-trained ISW model. As seen in Fig. 6, generated images preserve the relevant content information for segmentation while the style information such as illumination and colors is suppressed. These examples support the validity of our approach that selectively suppresses the style information.

Discussions

In this section, we discuss potential issues and improvements of our approach for further research.

Most of the normalization layers contain affine parameters to recover the original distribution and enhance the representation of a network. We attempted to deploy this by adding affine parameters or a 1×\times1 convolution layer after the normalization layer incorporating our proposed whitening loss. Despite our effort, this approach did not improve our method. We conjecture it is because affine parameters or a 1×\times1 convolution layer do not have sufficient complexity in recovering the original distribution.

Our method adopted photometric transformation to separate the style and content information, where we found that applying color transform and Gaussian blur does not harm the content information. We expect our approach can be further improved by exploring various photometric augmentation techniques.

Conclusions

This paper proposed a novel instance selective whitening (ISW) loss, which facilitates disentangling the covariances of the intermediate features into the style- and content-related ones and suppressing only the former to learn the domain-invariant feature representation. We focused on solving the domain generalization problem in urban-scene segmentation, which has practical impact when deployed in the wild but has not been studied much. In this regard, we strive to promote the importance of the domain generalization and inspire new research paths in this area.

Acknowledgments This work was partially supported by Institute of Information & communications Technology Planning & Evaluation (IITP) grant funded by the Korea government(MSIT) (No. 2019-0-00075, Artificial Intelligence Graduate School Program(KAIST) and No. 2020-0-00368, A Neural-Symbolic Model for Knowledge Acquisition and Inference Techniques), the National Research Foundation of Korea (NRF) grant funded by the Korean government (MSIT) (No. NRF-2019R1A2C4070420).

References

A Supplementary Material

This supplementary section provides additional quantitative results to examine hyper-parameter impacts, further implementation details, and qualitative results.

Comparison of segmentation results is shown in Fig. 7. Our method makes reasonable predictions, while the baseline completely fails on them.

We compare the result of our method with those reported from several domain adaptation (DA) methods under various settings. Fig. 8 shows the increase in mIoU from the baseline for each method. Although our method may not be the top performer, it shows comparable results to other DA methods. Note that DA methods require access to the target domain to solve DA problems. In contrast, our method is designed to improve generalization performance on an arbitrary unseen domain under the assumption of no access to the target domain, so we believe a comparison with DA methods under the same setting is impossible. However, we expect to solve DA by extending our key idea of selectively removing style-sensitive covariances to selectively matching such covariances between source and target domain.

A.2 Hyper-parameter Impacts

We adopt kk-means clustering to separate covariance elements into two groups, domain-specific style and domain-invariant content, according to the variance of each covariance element across various photometric transformations such as color jittering and Gaussian blur. As specified in Section 4.3, after dividing the covariance elements into kk clusters by the magnitude of the variance, the clusters from the first to the mm-th are considered to be insensitive, and the remaining clusters are considered sensitive to photometric transformation. We set mm to one and search the optimal kk through the hyper-parameter search. Fig. 9 shows the threshold where the covariances are divided into two groups depending on the kk value. Table 7 shows the changes in mIoU performance according to the kk values, suggesting the optimal kk as 3. Also, we can see that ours (ISW) performs better than IBN-Net or ours (IW) for all kk values. Note that ours (IW) applies instance whitening loss to all covariance elements, while ours (ISW) applies it to a part of the covariance elements according to the kk value.

As described in Section 4.2, we propose margin-based relaxation of whitening loss. Table 8 shows the performance of ours (IRW) according to the margin δ\delta.

As described in Section 4.4, we empirically set the weight γ\gamma of the proposed ISW loss as 0.6. Table 9 shows the impact of changing γ\gamma.

A.3 Further Implementation Details

Fig. 10 shows the detailed architecture of the semantic segmentation networks based on ResNet and DeepLabV3+. We adopt the auxiliary per-pixel cross-entropy loss proposed in PSPNet and concatenate the low-level features from the ResNet stage 1 to the high-level features according to the encoder-decoder architecture proposed in DeepLabV3+. Instance normalization (IN) with ISW loss replaces batch normalization (BN) in the input convolutional layer, and these ones are added after the skip-connection of the last residual block for each ResNet stage. As IBN-Net pointed out, earlier layers tend to encode the style information, hence we only adopt the ISW loss to the input convolutional layer and ResNet stage 1 and 2. In the end, the final loss LTotal\mathcal{L}_{\text{Total}} is formulated as,

where the γ1\gamma_{1} is 0.4 and the γ2\gamma_{2} is 0.6. We set the batch size to 8 for Cityscapes and 16 for GTA. For the photometric transformation, we apply Gaussian blur and color jittering implemented in Pytorch with a brightness of 0.8, contrast of 0.8, saturation of 0.8, and hue of 0.3.

A.4 Additional Qualitative Results

This section demonstrates additional qualitative results. We first present the comparison of the segmentation results on a seen domain (i.e., Cityscapes) and diverse driving conditions in BDD-100K, and then show the failure cases of our method. Besides, we show the effects of the whitening by comparing the reconstructed images from our proposed approach and the baseline. Finally, we provide the tendency of images from the most sensitive and insensitive covariance elements to the photometric transformation.

To qualitatively describe the effect of our method, we compare the segmentation results from the baseline and ours. Fig. 12 presents the segmentation results on a seen domain (i.e., Cityscapes). Similar to the quantitative results reported in Section 5, even with qualitative results, our model shows comparable performance to the baseline model on the seen domain. Fig. 13 shows the segmentation results under illumination changes on an unseen domain (i.e., BDD-100K). Note that Cityscapes dataset only contains images taken at the daytime. The first group images are taken at the dusk. We can see that the baseline model is vulnerable to these changes, but in contrast, our model outputs less damaged maps and reasonably predicts roads and cars. In extreme cases such as at night, both models fail to predict the sky, but our method still finds key components such as roads and cars well. In addition, our method produces reasonable segmentation results even for drastic changes in lighting such as shadows, as seen in the third group. Fig. 14 shows the segmentation results under the adverse weather conditions, unseen structures, and lush vegetation. Our model successfully predicts a partially snowy sidewalk, whereas the baseline model incorrectly predicts it as a building. The second case in the first group shows a foggy urban scene. The baseline fails to cope with these weather changes, while ours still shows fair results. Under the structural changes as shown in the second group, our method finds the road and sidewalk better than the baseline. Moreover, the baseline totally fails to detect the parking lot. In the last case, which is lush vegetation, the baseline produces noisy segmentation results and confused the road as a car. On the other hand, our model shows reasonable performance in both cases. Fig. 11 shows the failure cases caused by a large domain shift.

To reveal the information that the covariance represents, we first identify the most sensitive and insensitive covariances to the photometric transformation. Then, we sort the BDD-100K images according to the magnitude of the identified covariances. The results are described in Fig. 15. In the left group, the images are getting dark as the most sensitive covariance is getting smaller. We conjecture that the corresponding covariance tends to represent the illumination information. On the other hand, the right group shows the sorted images along with the most insensitive covariance. The scenes are getting simpler as the covariance gets smaller, which implies that the most insensitive covariance tends to represent the scene complexity.