Are Vision Transformers Robust to Spurious Correlations?
Soumya Suvra Ghosal, Yifei Ming, Yixuan Li
Introduction
A key challenge in building robust image classification models is the existence of spurious correlations: misleading heuristics imbibed within the training dataset that are correlated with majority examples but do not hold in general. Prior works have shown that convolutional neural networks (CNNs) can rely on spurious features to achieve high average test accuracy. Yet, such models lead to low accuracy on rare and untypical test samples lacking those heuristics . In Figure 1, we illustrate a model setup that exploits the spurious correlation between the water background and label waterbird for prediction. Consequently, a model that relies on spurious features performs poorly on test samples where the correlation no longer holds, such as waterbird on land background.
While the robustness of CNNs has been widely studied, it remains underexplored how spurious correlation is manifested in the recent development of vision transformers (ViT) . As with the paradigm shift to attention-based architectures, it becomes increasingly critical to understand their behavior under ill-conditioned data. From a network architecture perspective, ViTs lack the inductive bias in CNNs, such as translational equivariance and spatial locality, and may be more prone to overfitting . For this reason, one may expect the fully-connected dependencies in ViT models may exacerbate capturing the spurious correlations in the training data. In this paper, we seek to answer the following question: Are Vision Transformers more robust to spurious correlations compared to CNNs? Motivated by the question, we systematically investigate how and when ViT models exhibit robustness to spurious correlations on challenging benchmarks. Our findings reveal that for transformers, larger models and more pre-training data yield a significant improvement in robustness to spurious correlations. The key reason for success can be attributed to the ability to generalize better from those examples where spurious correlations do not hold, while fine-tuning. However, despite better generalization capability, ViT models suffer high errors on challenging benchmarks when these counterexamples are scarce in the training set. On the other hand, when pre-trained on a relatively smaller dataset such as ImageNet-1k, the performance of transformer-based models are much worse as compared to CNN counterparts. This indicates that in smaller pre-training data regimes, transformers have a higher propensity to overfit the spurious features and are less robust than CNNs of comparable size.
Going beyond, we perform extensive ablations and experiments to understand the role of self-attention mechanism in providing robustness to ViT models. Our findings reveal that the self-attention mechanism in ViTs plays a crucial role in guiding the model to focus on spatial locations in an image which are essential for accurately predicting the target label. Interestingly, we also found that restricting the attention to be local can result in sharp degradation in model robustness to spurious associations. Thus, the global attention in ViT models is indeed important for providing additional robustness to spurious correlations.
Our key contributions are summarized below:
To the best of our knowledge, we provide a first systematic study on the robustness of Vision Transformers when learned on datasets containing spurious correlations. Our work sheds light on the effectiveness of pre-training on ViT’s robustness to spurious correlations.
We perform extensive experiments and ablations to understand the effect of model architectures, model capacity, pre-training dataset, data imbalance, fine-tuning, etc.
We provide insights on ViT’s robustness by analyzing the attention matrix, which encapsulates important information about the interaction among image patches. We hope that our work will inspire future research on further understanding the robustness of ViT models.
Preliminaries
Spurious features refer to statistically informative features that work for majority of training examples but do not capture essential cues related to the labels . We illustrate a few examples in Figure 1. In waterbird vs landbird classification problem, majority of the training images has the target label (waterbird or landbird) spuriously correlated with the background features (water or land background). Sagawa et al. showed that deep neural networks can rely on these statistically informative yet spurious features to achieve high test accuracy on average, but fail significantly on groups where such correlations do not hold such as waterbird on land background.
Here represents a function transformation from the feature space to the pixel space . Considering the example of waterbird vs landbird classification, invariant features would refer to signals which are essential for classifying as , such as the feather color, presence of webbed feet, and fur texture of birds, to mention a few. Environmental features , on the other hand, are cues not essential but correlated with target label . For example, many waterbird images are taken in water habitats, so water scenes can be considered as . Under the data model, we form groups that are jointly determined by the label and environment . For this study, we consider the binary setting where and , resulting in four groups. The concrete meaning for each environment and label will be instantiated in corresponding tasks, which we describe in Section 3.
2 Transformers
Similar to the Transformer architecture in , ViT model expects the input as a 1D sequence of token embeddings. An input image is first partitioned into non-overlapping fixed-size square patches of resolution , resulting in a sequence of flattened 2D patches. For example, given an image of size and patch size , the image is divided into patches of resolution , resulting in image patches. Next, these patches are mapped to constant size embeddings with a trainable linear projection. In the previous example, the output of the projection layer will be embedding vectors of fixed dimension.
Following , ViT prepends a learnable embedding (class token) to the sequence of embedded patches, and this class token is used as image representation at the output of the transformer. To imbibe relative positional information of patches, position embeddings are further added to the patch embeddings.
The core architecture of ViT mainly consists of multiple stacked encoder blocks, where each block primarily consists of: (1) multi-headed self-attention layers, which learn and aggregate information across various spatial locations of an image by processing interactions between different patch embeddings in a sequence; and (2) a feed-forward layer. See an expansive discussion in related work (Section 5).
3 Model Zoo
In this study, we aim to understand the robustness of ViT models when trained on a dataset containing spurious correlations and how they fare against popular CNNs. We contrast ViT with Big Transfer (BiT) models that are primarily based on the ResNet-v2 architecture. For both ViT and BiT models, we consider different variants that differ in model capacity and pre-training dataset, as summarized in Table 1. Specifically, we use model variants pre-trained on both ImageNet-1k and on ImageNet-21k datasets.
Table 1 summarizes the size and pre-training dataset of different models used in our study. Note that the DeiT architecture is identical to ViT variant of comparable size with the only difference lying in the pre-training dataset and data augmentations.
Notation: To indicate input patch size in ViT models, we append “/x” to model names. We prepend -B, -S, -Ti to indicate Base, Small and Tiny version of the corresponding architecture. For instance: ViT-B/16 implies the Base variant with an input patch resolution of . In this paper, we use a input patch size for computational simplicity.
Robustness to Spurious Correlation
In this section, we systematically measure the robustness performance of ViT models when trained on datasets containing spurious correlations, and compare how their robustness fares against popular CNNs. For evaluation benchmarks, we adopt the same setting as in . Specifically, we consider the following three classification datasets to study the robustness of ViT models in a spurious correlated environment: Waterbirds (Section 3.1), CelebA (Section 3.2), and ColorMNIST. Due to space constraints, results on ColorMNIST are in the Supplementary.
Introduced in , this dataset contains spurious correlation between the background features and target label {waterbird, landbird}. The dataset is constructed by selecting bird photographs from the Caltech-UCSD Birds-200-2011 (CUB) dataset and then superimposing on either of background selected from the Places dataset . The spurious correlation is injected by pairing waterbirds on water background and landbirds on land background more frequently, as compared to other combinations. The dataset consists of training examples, with the smallest group size 56 (i.e, waterbird on land background).
Results and insights on generalization performance Table 2 compares worst-group accuracies of different models when fine-tuned on Waterbirds using empirical risk minimization. Note that all the compared models are pre-trained on ImageNet-21k. This allows us to isolate the effect of model architectures, in particular, ViT vs. BiT models. The worst-group test accuracy reflects the model’s generalization performance for groups where the correlation between the label and environment does not hold. A high worst-group accuracy is indicative of less reliance on the spurious correlation in training. Our results suggest that: (1) ViTs are relatively more robust to spurious associations between background feature and target label than convolution-based BiTs. Interestingly, ViT-B/16 attains a significantly higher worst-group test accuracy (89.3%) than BiT-M-R50x3 despite having a considerably smaller capacity (86.1M vs. 211M). (2) Furthermore, these results reveal a correlation between generalization performance and model capacity. With an increase in model capacity, both ViTs and BiTs tend to generalize better, measured by both average accuracy and worst-group accuracy. The relatively poor performance of ViT-Ti/16 can be attributed to its failure to learn the intricacies within the dataset due to its compact capacity.
Figure 2 provides a visual illustration of the experimental setup (left), along with the evaluation results (right). Our operating hypothesis is that a robust model should predict same class label and for a given pair , as they share exactly the same foreground object (i.e., invariant feature). Our results in Figure 2 show that ViT models achieve overall higher consistency measures than BiT counterparts. For example, the best model ViT-B/16 obtains consistent predictions for 93.9% of image pairs. Overall, using ViT pre-trained models yields strong generalization and robustness performance on Waterbirds.
2 CelebA
Results We see from Table 3 that ViT models achieve higher test accuracy (both average and worst-group) as opposed to BiTs. In particular, ViT-B/16 achieves higher worst-group test accuracy than BiT-M-R50x3, despite having a considerably smaller capacity (86.1M vs. 211M). These findings along with our observations in Section 3.1 demonstrate that ViTs are not only more robust when there are strong associations between the label and background features, but also avoid learning spurious correlations between demographic features and target label.
Discussion: A Closer Look at ViT Under Spurious Correlation
In this section, we perform extensive ablations and experiments to understand the role of ViT models under spurious correlations. For consistency, we present the analyses below based on the Waterbirds dataset.
In this section, we aim to understand the role of large-scale pre-training on the model’s robustness to spurious correlations. Specifically, we compare pre-trained models of different capacities, architectures, and sizes of pre-training data. To understand the importance of the pre-training dataset, we compare models pre-trained on ImageNet-1k ( million images) and ImageNet-21k ( million images). We report results for transformer-based models and BiT models in Table 4. For detailed ablation results on other benchmark datasets, please refer to the Appendix. Based on these results, we highlight the following observations:
First, large-scale pre-training improves the performance of the models on challenging benchmarks. For transformers, larger models (base and small) and more pre-training data (ImageNet-21k) yields a significant improvement in all reported metrics. Hence, larger pre-training data and increasing model size play a crucial role in improving model robustness to spurious correlations. We also see a similar trend in the case of BiT models.
Second, when pre-trained on a relatively smaller dataset such as ImageNet-1k, the performance of transformer-based DeiT models are much worse as compared to BiT-S models. Interestingly, although increasing size of DeiT models leads to improved average test accuracy but suffers high error on worst-group samples. This indicates that in smaller pre-training data regimes, transformers have a higher propensity of memorizing training samples and are less robust compared to CNNs of comparable size. From a network architecture perspective, this may be due to fully-connected layers in transformer models which capture spurious correlations occurring in the target task in case of limited pre-training data. Our findings corroborate reportings in that inductive bias in convolutional neural networks plays a crucial role without strong pre-training.
2 Understanding role of self-attention mechanism for improved robustness in ViT models
Given the results above, a natural question arises: what makes ViT particularly robust in the presence of spurious correlations? In this section, we aim to understand the role of ViT by looking into the self-attention mechanism. The attention matrix in ViT models encapsulates crucial information about the interaction between different image patches.
Latent pattern in attention matrix To gain insights, we start by analyzing the attention matrix, where each element in the matrix represents attention values with which an image patch focuses on another patch . For example: consider an input image of size and patch resolution of , then we have a attention matrix (excluding the class token). To compute final attention matrix, we use Attention Rollout which recursively multiplies attention weight matrices in all layers below. Our analysis here is based on the ViT-B/16 model fine-tuned on Waterbirds.
Intriguingly, we observe that each image patch, irrespective of its spatial location, provides maximum attention to the patches representing essential cues for accurately identifying the foreground object.
Figure 3 exhibits this interesting pattern, where we mark (in red) the top patches being attended by every image patch. To do so, for every image patch , where , we find the top patches receiving the highest attention values and mark (in red) on the original input image. This would give us patches, which we overlay on the original image. Note that different patches may share the same top patches, hence we observe the sparse pattern. In Figure 3, we can see that the patches receiving the highest attention represent important signals such as the shape of the beak, claw, and fur color—all of which are essential for the classification task waterbird vs landbird.
It is particularly interesting to note the last row in Figure 3, which is an example from the minority group (waterbird on land background). This is a challenging case where the spurious correlations between and do not hold. A non-robust model would utilize the background environmental features for predictions. In contrast, we notice that each patch in the image correctly attends to the foreground patches.
Masked attention The attention matrix in ViT models encapsulates crucial information about the interaction between different image patches resulting in access to more global information. Inspired by , we use a spatial mask to study the effect of restricting image patches to attend only those lying within a certain distance. However, the class token is allowed to interact and attend to all other image patches. Note, while fine-tuning we do not use any spatial mask and allow the model to leverage information from the complete attention matrix. Masking is done only during inference time. Figure 4 depicts the results of our study on ViT-B/16 when fine-tuned on Waterbirds (left) and CelebA (right). For both datasets, we see a monotonic decrease in worst-group test accuracy and Consistency Measure, as we increase the restriction on allowable attention distance. In the extreme case, when the constrained attention distance equals 2, the model completely fails to correctly classify the test images in the smallest group indicating high reliance on spurious features while making the prediction. In other words, limiting the attention to be local results in degradation of model robustness to spurious correlations. Thus, we conclude that global attention in ViT models indeed plays a crucial role in providing additional robustness to spurious correlations.
3 Investigating model performance under data imbalance
Recall that model robustness to spurious correlations is correlated with its ability to generalize from the training examples where spurious correlations do not hold. We hypothesize that this generalization ability varies depending on the inherent data imbalance. In this section, we investigate the effect of data imbalance on the model’s performance. In the extreme case, the model only observes 5 samples from the underrepresented group.
Setup Considering the problem of waterbird vs landbird classification, these examples correspond to those in the groups: waterbird on land background and landbird on water background. We refer to these examples that do not include spurious associations with label as minority samples. For this study, we remove varying fraction of minority samples from the smallest group( waterbird on land background ), while fine-tuning. We measure the effect based on the worst-group test accuracy and model consistency defined in Section 3.1.
Takeaways In Figure 5, we report results for ViT-S/16 and BiT-M-R50x1 model when finetuned on Waterbirds dataset . We find that as more minority samples are removed, there is a graceful degradation in the generalization capability of both ViT and BiT models. However, the decline is more prominent in BiTs with the model performance reaching near-random when we remove 90% of minority samples. From this experiment, we conclude that additional robustness of ViT models to spurious associations stems from their better generalization capability from minority samples. However, they still suffer from spurious correlations when minority examples are scarce.
4 Does longer fine-tuning in ViT improve robustness to spurious correlations?
Recent studies in the domain of natural language processing have shown that the performance of BERT models on smaller datasets can be significantly improved through longer fine-tuning. In this section, we investigate if longer fine-tuning also plays a positive role in the performance of ViT models in spuriously correlated environments.
Takeaways Figure 6 reports the loss (left) and accuracy (right) at each epoch for ViT-S/16 model fine-tuned on Waterbirds dataset . To better understand the effect of longer fine-tuning on worst-group accuracy, we separately plot the model loss and accuracy on all examples and minority samples. From the loss curve, we observe that the training loss for minority examples decreases at a much slower rate as compared to the average loss. Specifically, the average train loss takes 20 epochs of fine-tuning to reach near-zero values, while training loss on minority group plateaus after 40 epochs. Similarly, we see that although the average test accuracy of the model stops increasing after 30 epochs, the accuracy of minority samples reaches a stationary state after 50 epochs of fine-tuning. These results reveal two key observations: (1) While longer fine-tuning does not benefit the average test accuracy, it plays a positive role in improving model performance on minority samples, and (2) ViT models do not overfit with longer fine-tuning.
5 Spurious Out-of-Distribution Detection
Finally, we study the performance of ViT models in out-of-distribution setting. Introduced in , spurious out-of-distribution (OOD) data is defined as samples that do not contain the invariant features essential for accurate classification, but contain the spurious features . Hence, these samples are denoted as where is an out-of-class label, such that . In the problem of waterbird vs landbird classification, an image of a person standing in forest would be an example of spurious OOD, since it contains different semantic class person , yet has the environmental features of land background. A non-robust model relying on the background feature may classify such OOD data as an in-distribution class with high confidence. Hence, we aim to understand if self-attention based ViT models can mitigate this problem and if so, to what extent.
Setup To investigate the performance of different models against spurious OOD examples, we use the setup introduced in . Specifically, for Waterbirds we test on subset of images of land and water sampled from the Places dataset . Considering, CelebA as in-distribution, our test suite consists of images of bald male as spurious OOD, since they contain environmental features (gender) without invariant features (hair). For CMNIST, the in-distribution data contains digits = and the background colors, = {red, green, purple, pink}. We use digits {5, 6, 7, 8, 9} with background color red and green as test OOD samples.
Takeaways We report our findings in Table 5. Clearly, ViT models achieve better OOD evaluation metrics as compared to BiTs. Specifically, ViT-B/16 achieves higher AUROC than BiT-M-R50x3, considering Waterbirds as in-distribution.
Related Works
Pre-training and robustness Recently, there has been an increasing amount of interest in studying the effect of pre-training . Specifically, when the target dataset is small, generalization can be significantly improved through pre-training and then finetuning . Findings of Hendrycks et al. reveal that pre-training provides significant improvement to model robustness against label corruption, class imbalance, adversarial examples, out-of-distribution detection, and confidence calibration. In this work, we focus distinctly on robustness to spurious correlation, and how it can be improved through large-scale pretraining.
Vision transformer Since the introduction of transformers by Vaswani et al. in 2017, there has been a deluge of studies adopting the attention-based transformer architecture for solving various problems in natural language processing . In the domain of computer vision, Dosovitskiy et al. first introduced the concept of Vision Transformers (ViT) by adapting the transformer architecture in for image classification tasks. Subsequent studies have shown that when pre-trained on sufficiently large datasets, ViT achieves superior performance on downstream tasks, and outperforms state-of-art CNNs such as residual networks (ResNets) of comparable sizes. Since coming to the limelight, multiple variants of ViT models have been proposed. Touvron et al. showed that it is possible to achieve comparable performance in small pre-training data regimes using extensive data augmentation and novel distillation strategy. Further improvements on ViT include enhancement in tokenization module , efficient parameterization for scalability and building multi-resolution feature maps on transformers . In this paper, we provide a first systematic study on the robustness of vision transformers when learned on datasets containing spurious correlations.
Robustness of transformers Naseer et al. provides a comprehensive understanding of the working principle of ViT architecture through extensive experimentation. Some notable findings in reveal that transformers are highly robust to severe occlusions, perturbations, and distributional shifts. Recently, performance of ViT models in the wild has been extensively studied using a set of robustness generalization benchmarks, e.g., ImageNet-C , Stylized-ImageNet , ImageNet-A , etc. Different from prior works, we focus on robustness performance on challenging datasets, which are designed to expose spurious correlations learned by the model. Our analysis reveals that pre-training improves robustness by better generalizing on examples from under-represented groups. Our findings are also complementary to robustness studies in the domain of natural language processing, which reported that transformer-based BERT models improve robustness to spurious correlations.
Conclusion
In this paper, we investigate the robustness of ViT models when learned on datasets containing spurious associations between target label and environmental features. Our findings can be summarized as: 1) ViTs are more robust to spurious correlations than CNNs under large-scale pre-training data regime. However, when the pre-training dataset is relatively small, transformer models perform much worse as compared to CNNs of comparable size; 2) We find that global attention in ViT architecture plays a crucial role in providing improved robustness. Further, restricting the attention to be local results in degradation of model performance; 3) Improved robustness of ViT models can be attributed to better generalization capability from the counterexamples where spurious correlations do not hold. However, when such samples become scarce ViT models tend to overfit to spurious associations. We hope that our work will inspire future research on understanding the robustness of ViT models.
References
Appendix A Implementation Details
Transformers. For both ViT and DeiT models, we obtain the pre-trained checkpoints from the timm libraryhttps://github.com/rwightman/pytorch-image-models/tree/master/timm. For downstream fine-tuning on Waterbirds and CelebA dataset, we scale up the resolution to 384 × 384 by adopting 2D interpolation of the pre-trained position embeddings proposed in . Note, for CMNIST we keep the resolution as during fine-tuning. We fine-tune models using SGD with a momentum of 0.9 with an initial learning rate of 3e-2. As described in , we use a fixed batch size of 512, gradient clipping at global norm 1 and a cosine decay learning rate schedule with a linear warmup. We fine-tune tiny & small versions of models (i.e., ViT-Ti/16 and ViT-S/16) for 1000 steps, whereas base version (i.e., ViT-B/16) is fine-tuned for 2000 steps.
BiT. We obtain the pretrained checkpoints from the official repositoryhttps://github.com/google-research/big_transfer. For downstream fine-tuning, we use SGD with an initial learning rate of 0.003, momentum 0.9, and batch size 512. We fine-tune models with various capacity for 500 steps, including BiT-M-R50x1, BiT-M-R50x3, and BiT-M-R101x1.
Appendix B Extension: How does the size of pre-training dataset affect robustness to spurious correlations?
In this section, to further validate our findings on the importance of large-scale pre-training dataset, we show results on CelebA dataset. We report our findings in Table 6. We also observe a similar trend for this setup that larger model capacity and more pre-training data yields significant improvement in worst-group accuracy for ViT models. Further, when pre-trained on a relatively smaller dataset such as ImageNet-1k, the performance of transformer-based DeiT models are poor as compared to the corresponding CNN counterpart.
Also, compared to BiT models, the robustness of ViT models benefits more with a large pre-training dataset. For example, compared to ImageNet-1k, fine-tuning ViT-B/16 pre-trained on ImageNet-21k improves the worst-group accuracy by 6%. On the other hand, for BiT models, fine-tuning with a larger pre-trained dataset yields marginal improvement. Specifically, BiT-M-R50x3 only improves the worst-group accuracy by 1.5% with ImageNet-21k.
Appendix C Extension : Color Spurious Correlation
Results and insights on robustness performance We compare model predictions on samples with same class label but different background & foreground colors. Given a data point (), we modify the background and foreground color of randomly to generate a new test image with the constraint of having the same semantic label. During evaluation, the background color is chosen uniform-randomly from the set of colors: {#ecf02b, #f06007, #0ff5f1, #573115, #857d0f, #015c24, #ab0067, #fbb7fa, #d1ed95, #0026ff} and the foreground color is selected randomly from the set . For evaluation purpose, we form a dataset consisting of samples and the results reported are averaged over 50 random runs. Figure 7 depicts the distribution of training samples in CMNIST dataset (left) and few representative examples after transformation (right).
We report our findings in Figure 8. Our operating hypothesis is that a robust model should predict same class label and for a given pair , as they share exactly the same target label (i.e., the invariant feature is approximately the same). We can observe from Figure 8 that the best model ViT-B/16 obtains consistent predictions for 100% of image pairs. After extensive experimentation over all combinations, we find that setting the foreground color as black and the background as white caused the models to be most vulnerable. We see a significant decline in model consistency when the foreground color is set as black and the background as white (indicated as BW) as compared to random setup.
Appendix D Visualization
In Figure 9, we visualize attention maps obtained from ViT-B/16 model for some samples images from Waterbirds and CMNIST dataset. We use Attention Rollout to obtain the attention matrix. We can observe that the model successfully attends spatial locations representing invariant features while making predictions.
D.2 The Attention Matrix of CMNIST
In the main text, we provide visualizations in which each image patch, irrespective of its spatial location, provides maximum attention to the patches representing essential cues for accurately identifying the foreground object. In Figure 10, we show visualizations for ViT-B/16 fine-tuned on CMNIST dataset to further validate our findings.