On Feature Learning in the Presence of Spurious Correlations

Pavel Izmailov, Polina Kirichenko, Nate Gruver, Andrew Gordon Wilson

Introduction

In classification problems, a feature is spurious if it is predictive of the label without being causally related to it. Models that exploit the predictive power of spurious features can achieve strong average performance on training and in-distribution test data but often perform poorly on sub-groups of the data where the spurious correlation does not hold . For example, neural networks trained on ImageNet are known to rely on backgrounds or texture , which are often correlated with labels without being causally significant. Similarly in natural language processing, models often rely on specific words and syntactic heuristics when predicting the sentiment of a sentence or the relationship between a pair of sentences . In extreme cases, neural networks completely ignore task-relevant core features and only use spurious features in their predictions , achieving zero accuracy on the subgroups of the data where the spurious correlation does not hold.

In recent work, Kirichenko et al., showed that, surprisingly, standard Empirical Risk Minimization (ERM) learns a high-quality representation of the core features on datasets with spurious correlations, even when the model primarily relies on spurious features to make predictions. Moreover, they showed that it is often possible to recover state-of-the-art performance on benchmark spurious correlation problems by simply retraining the last layer of the model on a small held-out dataset where the spurious correlation does not hold. This procedure is called Deep Feature Reweighting (DFR).

In this paper, we provide an in-depth study of the factors that affect the quality of learned representations in the presence of spurious correlations: how accurately we can decode the core features from the learned representations. Following Kirichenko et al., , we break the problem of training a robust classifier into two tasks: extracting feature representations and training a linear classifier on these features. In order to study the feature learning in isolation, we use the DFR procedure to learn an optimal linear classifier on the feature representations, and evaluate the features learned with different training methods, neural network architectures, and hyper-parameters.

First, on a range of problems with spurious correlations we show that while specialized group robustness methods such as group distributionally robust optimization (group DRO) can significantly outperform the standard ERM training, the quality of the features learned by ERM is highly competitive: by applying the DFR procedure to the features learned by ERM and group DRO we achieve similar performance. Furthermore, we show that the performance improvements of group DRO are largely explained by the better weighting of the learned features in the last classification layer, and not by learning a better representation of the core features. This observation has high practical significance, as the problem of training the last layer of the model is much simpler both conceptually and computationally than training the full model to avoid spurious correlations .

Next, focusing on the ERM training, we explore the effect of model class, pretraining strategy and regularization on feature learning. We find a linear dependence between the in-distribution accuracy of the model and the worst group accuracy after applying DFR, meaning that on natural datasets good generalization typically implies good feature learning, even in the presence of spurious features. Further, we show that the pre-training strategy has a very significant effect on the quality of the learned features, while strong regularization does not significantly improve the feature representations on most benchmarks.

Finally, by finetuning a pretrained state-of-the-art ConvNext model , we significantly outperform the best reported results on the popular Waterbirds , CelebA hair color and WILDS FMOW spurious correlation benchmarks, using only simple ERM training followed by DFR.

Our code is available at github.com/izmailovpavel/spurious_feature_learning.

Related Work

Numerous works describe how neural networks can rely on spurious correlations in real world problems. In vision, neural networks can learn to rely on an image’s background , secondary objects , object textures and other semantically irrelevant features . Spurious correlations are especially problematic in high-risk domains such as medical imaging, where it was shown that neural networks can use hospital-specific metal tokens or cues of disease treatment rather than symptoms to perform automated diagnosis on chest X-ray images. Spurious features are also extremely prevalent in NLP, where models can achieve good performance on benchmarks without properly solving them, e.g. by using simple syntactic heuristics such as lexical overlap between the two sentences in order to classify the relationship between them . For a comprehensive survey of the area, see Geirhos et al., .

Because of the high practical significance of spurious correlations, many group robustness methods have been proposed. These methods aim to reduce the reliance of deep learning models on spurious correlations and improve worst group performance. Group DRO is the state-of-the-art group robustness method, which minimizes the worst-group loss instead of the average loss. Other works focus on automatically identifying the minority group examples , learning several diverse classifiers that use different features or using partially available group labels . Group subsampling was shown to be a strong baseline for some benchmarks .

In this work, we focus on feature learning in the presence of spurious correlations. Hermann and Lampinen, perform a conceptually similar study, but focusing on synthetic datasets. Similar to their work, we explore how well the different features of the data can be decoded from the features learned by deep neural networks, but on large-scale natural datasets. Hermann et al., explore the feature learning in the context of texture bias , finding that data augmentation has a profound effect on the texture bias while architectures and training objectives have a relatively small effect. Ghosal et al., show that Vision Transformer models pretrained on ImageNet22k significantly outperform standard CNN models on several spurious correlation benchmarks.

Lovering et al., explore the factors which affect the extractability of features after pre-training and fine-tuning of NLP models. Kaushik et al., construct counterfactually augmented sentiment analysis and naural language inference datasets (CAD) and show that combining CAD with the original data reduces the reliance on spurious correlations on the corresponding benchmarks. Kaushik et al., explain the efficacy of CAD and show that while adding noise to causal features degrades in-distribution and out-of-distribution performance, adding noise to non-causal features improves robustess. Eisenstein, and Veitch et al., formally define and study spurious features in NLP from the perspective of causality.

Kirichenko et al., show that models trained with standard ERM training often learn high-quality representations of the core features, and propose the DFR procedure (see Section 3) which we use extensively in this paper. Related observations have also been reported in other works in the context of spurious correlations , domain generalization and long-tail classification . While we build on the observations of Kirichenko et al., , our work provides profound new insights and greatly expands on the scope of their work. In particular, we investigate the feature representations learned by methods beyond standard ERM, and the role of model architecture, pre-training, regularization and data augmentation on learning semantic structure. We also extend our analysis beyond the standard spurious correlation benchmarks studied by Kirichenko et al., , by considering the challenging real world satellite imaging and chest X-ray datasets.

In an independent and concurrent work, Shi et al., also propose an evaluation framework for out-of-distribution generalization based on last layer retraining, inspired by the observations of Kirichenko et al., and Kang et al., . They focus on the comparison of supervised, self-supervised and unsupervised training methods, providing complementary observations to our work.

Background

Preliminaries. We consider classification tasks with inputs x∈Xx\in\mathcal{X} and classes y∈Yy\in\mathcal{Y}. We assume that the data distribution consists of groups G\mathcal{G} which are not equally represented in the training data. The distribution of groups can change between the training and test distributions, with majority groups becoming less common or minority groups becoming more common. Because of the imbalance in training data, models trained with ERM often have a gap between average and worst group performance on test. Throughout this paper, we will be studying worst group accuracy (WGA), i.e. the lowest test accuracy across all the groups G\mathcal{G}. For most problems considered in this paper, we assume that each data point has an attribute s∈Ss\in\mathcal{S} which is spuriously correlated with the label yy, and the groups are defined by a combination of the label and spurious attribute: G∈Y×S\mathcal{G}\in\mathcal{Y}\times\mathcal{S}. In test distribution we might find that ss is no longer correlated with yy, and thus a model that has learned to rely on the spurious feature ss during training will perform poorly at test time. Models that rely on the spurious features will typically achieve poor worst group accuracy, while models that rely on core features will have more uniform accuracies across the groups. In Appendix A, we describe the groups, spurious and core features in the datasets that we use in this paper.

In order to perform controlled experiments, we assume that we have access to the spurious attributes ss (or group labels) for training or validation data, which we use for training of group robustness baselines and feature quality evaluation. However, we emphasize that our results on the features learned by ERM hold generally, even when spurious attributes are unknown, as ERM does not use the information about the spurious features: we only use the spurious attributes to perform analysis.

Deep feature reweighting. Suppose we are given a model m:X→Cm:\mathcal{X}\rightarrow\mathcal{C}, where X\mathcal{X} is the input space and C\mathcal{C} is the set of classes. Kirichenko et al., assume that the model mm consists of a feature extractor (typically, a sequence of convolutional or transformer layers) followed by a classification head (typically, a single linear layer): m=h∘em=h\circ e, where e:X→Fe:\mathcal{X}\rightarrow\mathcal{F} is a feature extractor and h:F→Ch:\mathcal{F}\rightarrow\mathcal{C} is a classification head. They discard the classification head, and use the feature extractor ee to compute the set of embeddings D^e={(e(xi),yi)}i=1n\hat{\mathcal{D}}_{e}=\{(e(x_{i}),y_{i})\}_{i=1}^{n} of all the datapoints in the reweighting dataset D^\hat{\mathcal{D}}; the reweighting dataset is used to retrain the last layer of the model, and contains group-balanced data where the spurious correlation does not hold. Finally, they train a logistic regression classifier l:F→Cl:\mathcal{F}\rightarrow\mathcal{C} on the dataset D^e\hat{\mathcal{D}}_{e} For stability, logistic regression models are trained 1010 times on different random group-balanced subsets of the reweighting dataset D^\hat{\mathcal{D}}, and the weights of the learned logistic regression models are averaged. See Appendix B of Kirichenko et al., for full details on the DFR procedure.. Then, the final model used on new test data is given by ml=l∘em_{l}=l\circ e. Thoughout this paper, we use a group-balanced held-out dataset (subset of the validation dataset where each group has the same number of datapoints) as the reweighting dataset D^\hat{\mathcal{D}}; Kirichenko et al., denote this variation of the method as DFRTrVal{}_{\text{Tr}}^{\text{Val}} .

Experimental Setup and Evaluation Procedure

In this section, we describe the datasets, models and evaluation procedure that we use throughout the paper.

Datasets. In order to cover a broad range of practical scenarios, we consider four image classification and two text classification problems.

Waterbirds is a binary image classification problem, where the class corresponds to the type of the bird (landbird or waterbird), and the background is spuriously correlated with the class. Namely, most landbirds are shown on land, and most waterbirds are shown over water.

CelebA hair color is a binary image classification problem, where the goal is to predict whether a person shown in the image is blond; the gender of the person serves as a spurious feature, as 94%94\% of the images with the “blond” label depict females.

WILDS-FMOW is a satellite image classification problem, where the classes correspond to one of 6262 land use or building types, and the spurious attribute ss corresponds to the region (Africa, Americas, Asia, Europe, Oceania or Other; the “Other” region is not used in the evaluation). We note that for the FMOW datasets the groups Gs\mathcal{G}_{s} are defined by the value of the spurious attribute, and not the combination of the spurious attribute and the class label Gy,s\mathcal{G}_{y,s}, as described in Section 3. Moreover, on FMOW there is also a domain shift: the images for test and validation data (used for last layer retraining) are collected in 2016 and 2017, while the training data is collected before 2016. For more details, please see Appendix A.

CXR-14 is a dataset with chest X-ray images for which we focus on a binary classification problem of pneumothorax prediction. Oakden-Rayner et al., showed that there is a hidden stratification in the dataset such that most images from the positive class contain a chest drain, which is a non-causal feature related to treatment of the disease. While for all other benchmarks we report WGA on test data, for this dataset, following prior work , we report worst group AUC because of the heavy class imbalance.We take the minimum out of the two scores on test: AUC for classifying negative class against positive examples with chest drain and AUC for negative class against positive without chest drain.

MultiNLI is a text classification problem, where the task is to classify the relationship between a given pair of sentences as a contradiction, entailment or neither of them. In this dataset, the presence of negation words (e.g. “never”) in the second sentence is spuriously correlated with the “contradiction” class.

CivilComments is a text classification problem, where the goal is to classify whether a given comment is toxic. We follow Idrissi et al., and use the coarse version of the dataset both for training and evaluation, where the spurious attribute is s=1s=1 if the comment mentions at least one of the following categories: male, female, LGBT, black, white, Christian, Muslim, other religion; otherwise, the spurious label is . The presence of the eight categories above is spuriously correlated with the comment being classified as toxic.

The Waterbirds, CelebA, CivilComments and MultiNLI datasets are commonly used to benchmark the performance of group robustness methods [see e.g. 34, 50, 63]. The FMOW and CXR-14 datasets present challenging real-world problems with spurious correlations. In these datasets, the inputs do not resemble natural images from datasets such as ImageNet , so models have to learn the relevant features from data to achieve good performance, and cannot simply rely on feature transfer. We provide detailed descriptions of the data and show example datapoints in Appendix A, Figures 7 and 8.

Models. Following prior work we use a ResNet-50 model pretrained on ImageNet1k on Waterbirds, CelebA and FMOW. For the NLP problems, we use a BERT model pre-trained on Book Corpus and English Wikipedia data. On CXR-14, following prior work [e.g. 70, 65, 48] we use a DenseNet-121 model pretrained on ImageNet1k. In Section 6, we provide an extensive study of the effect of architecture and pretraining on the image classification problems, and in Appendix E we perform a similar study on the MultiNLI text classification problem.

Evaluation strategy. We use DFR to evaluate the quality of the learned feature representations, as described in Section 3: we measure how well the core features can be decoded from the learned representations with last layer retraining. In some of the experiments, we also train a linear classifier to predict the spurious attribute ss instead of the class label yy. Using this classifier, we can evaluate the decodability of the spurious feature from the learned feature representation. We refer to this procedure as ss-DFR and the corresponding worst group accuracy (in predicting the spurious attribute ss) as DFR ss-WGA. Additionally, we evaluate the worst group accuracy and mean Following Sagawa et al., and Kirichenko et al., , we evaluate the mean accuracy according to the group distribution in the training set. This way, mean accuracy represents the in-distribution generalization of the model. accuracy of the base model without applying DFR, which we refer to as base WGA and base accuracy respectively.

ERM vs Group Robustness Training

Multiple methods have been proposed for training classifiers which are more robust to spurious correlations, with significant improvements in worst group accuracy compared to standard training. In this section, we use DFR to investigate whether the improvements of group robustness methods are caused by better feature representations or by better weighting of the learned features.

Methods. We consider 4 methods for learning the features. ERM or Empirical Risk Minimization is the standard training on the original training data, without any techniques targeted at improving worst group performance. RWG reweights the loss on each of the groups according to the size of the group and RWY reweights the loss on each class according to the size of the class . Group DRO is a state-of-the-art method which uses the group information on the training data to minimize the worst group loss instead of the average loss. Group DRO is often considered as an oracle method or upper-bound on the worst group performance under spurious correlations .

On the CXR dataset the group labels are not available on the train data, so we cannot apply RWG or group DRO; on this dataset we compare ERM to RWY. On several datasets, the performance of RWY and RWG methods deteriorates during training. For these datasets, we additionally report the results for the checkpoint obtained with early stopping (RWY-ES and RWG-ES). For group DRO, we report the performance with early stopping on all datasets except for CXR (GDRO-ES). In all cases, early stopping is performed based on the worst-group accuracy on the validation set.

Hyper-parameters selection. We train ERM models, RWG and RWY with the same hyper-parameters shared between all the image datasets (apart from batch size which is set to 3232 on Waterbirds, and 100100 on the other datasets) and between the natural language datasets. We do not tune the hyper-parameters of these methods for worst group accuracy. For group DRO, we run a grid search over the values of the generalization adjustment CC, weight decay and learning rate hyper-parameters, and select the best combination according to the worst-group accuracy on validation data with early stopping. For details, please see Appendix B.

We compare feature learning methods on all datasets in Figure 1. As expected, the worst group accuracy of group robustness methods is significantly better than the ERM worst group accuracy on most datasets. For example, on Waterbirds, ERM only gets 68.8%68.8\% WGA, while group DRO with early stopping gets 90.6%90.6\%. However, after applying DFR the performance of ERM and group DRO is very close, with a slight advantage for ERM (91.1%91.1\% for ERM and 90%90\% for group DRO), and similar observations hold on all datasets.

The results for RWG and RWY are analogous. Namely, when combined with early stopping, these methods outperform ERM on base model performance. Once we apply DFR, however, the gap in performance between the methods becomes very small. In fact, on all the datasets and for all the methods the improvement in worst group accuracy from using any of the considered group robustness methods compared to ERM does not exceed 11–2%2\% after applying DFR.

These results suggest that the improvements over ERM in base model performance for methods such as group DRO and RWG are largely the result of better weighting of the learned features rather than learning better representations of the core features. Indeed, if the core feature was better represented by group robustness methods, we would expect to see a significant improvement over ERM after applying DFR.

This observation is significant both practically and scientifically. Practically, the problem of training the last layer is simpler, more data efficient and less compute intensive than retraining the full model . Our results suggest that for many problems, practitioners can primarily focus on retraining the last layer, as training the feature extractor model with group robustness methods does not provide significant improvements. Scientifically, robustness to spurious correlations is often implicitly or explicitly associated with the quality of learned feature representations [e.g. 3, 4, 74, 100, 48, 99]. Our results suggest that the quality of feature representations, i.e. the decodability of the core feature, is not significantly improved by group DRO, refining our understanding of group robustness training and representation learning in the presence of spurious correlations.

Effect of early stopping. Early stopping is crucial to achieving strong base model performance with RWY, RWG and group DRO on many of the datasets. In Figure 1 we report the results both with and without early stopping on datasets where the validation worst group accuracy significantly degrades over the course of training. Generally, early stopping does not appear to significantly improve the quality of the learned feature representations even in these problems: after applying DFR, methods with and without early stopping achieve similar performance. This observation suggests that late in training, neural networks may start to assign a higher weight to the spurious features, but the information about the core features is still preserved in the learned representations. For ERM and Group DRO, we explore the DFR WGA performance as a function of the training iteration in Appendix D, also finding that the length of training has a relatively small effect on the final DFR WGA.

Group DRO analysis. In Figure 2 we report the worst group accuracy of multiple group DRO runs before and after applying DFR. For each run, we evaluate the best checkpoint according to validation accuracy (i.e. the checkpoint selected by early stopping) and the last checkpoint saved after a fixed number of epochs. We observe that while the last checkpoints perform poorly in terms of base WGA in most runs, DFR can significantly improve their performance, removing the need for early stopping. Interestingly, we find that the best performing group DRO models cannot be improved by DFR on each of the datasets In Figure 1 DFR improves the results of GDRO-ES only on the FMOW dataset, where DFR adapts the model to the domain shift, as validation and test images were taken after 2016, while the training images were taken before 2016. . This result suggests that the weighting of the features learned by group DRO is already close to optimal, again indicating that the success of group DRO can largely be attributed to learning a better weighting for the features in the last linear layer, rather than learning better features.

Effect of the Base Model

Most of the prior work on spurious correlation considers a fixed model class for each problem: for example, on Waterbirds and CelebA datasets almost all the papers use a ResNet-50 base model pre-trained on ImageNet1k [e.g. 76, 50, 48, 40, 84, 63, 100]. Recently, Ghosal et al., showed that vision transformer models may provide better robustness to spurious correlations if pretrained on a large dataset. Here, we perform a systematic large-scale evaluation of the effect of base model choice on the quality of learned feature representations.

We repeat the experiments on the effect of base model, pretraining, and training on the target dataset presented in this section on the MultiNLI text classification task in Appendix E, with similar observations.

In Figure 3, we plot the base model mean and worst group accuracy and DFR worst group accuracy for a wide range of models and pretraining strategies on each of the four image classification datasets. We provide model descriptions and training hyper-parameters in Appendix C. We train a total of 78 models on Waterbirds, 78 on CelebA, 40 on FMOW and 40 on CXR.

Accuracy on the line. Miller et al., showed that for many distribution shifts in practice the out-of-distribution performance is highly correlated with the in-distribution generalization performance. In Figure 3 (top row), for each dataset we show scatter plots of base model mean accuracy vs base model worst group accuracy, analogously to Miller et al., . For FMOW (which was also considered by Miller et al., ) we observe a linear correlation between the base model mean and worst group accuracies. On CXR, the dependence is also roughly linear. However, both on CelebA and on Waterbirds, the correlation does not appear entirely linear. In particular, on Waterbirds there is a large number of models that have similar worst group accuracy ≈20%\approx 20\%, for which there appears to be little correlation between the base model WGA and mean accuracy. On CelebA, the same phenomenon occurs for the best performing models, with WGA between 40%40\% and 50%50\%.

DFR accuracy is on the line. Next, in the bottom panels of Figure 3, for each of the datasets we report the base model mean accuracy vs DFR WGA. On all datasets other than CXR On CXR, we found that the base model is often hard to improve with DFR, suggesting that it already learns to weight the features correctly. We discuss the possible reasons and the nature of spurious features on this dataset in Appendix A.2. , we observe a high linear correlation between the metrics, including the Waterbirds and CelebA. In particular, the models with the best mean (in-distribution) accuracy also achieve the best DFR WGA. Note that this is not the case for the base model WGA on CelebA, where the best base model WGA is 51%51\% achieved by a ResNet-101 model pretrained on ImageNet1k; this model only achieves mean accuracy of 95.45%95.45\% compared to 96.2%96.2\% accuracy for the ConvNext XLarge model. The results in Figure 3 confirm that models that achieve better accuracy on the training data distribution generally learn better representations of the core features, and provide better worst group accuracies with DFR.

The difference between the results for base model WGA and DFR WGA suggests that, analogously to the results in Section 5, some models achieve better worst group performance by choosing a better weighting of the features in the last linear layer, rather than by learning a better representation of the core features.

Are VITs more robust than CNNs? Ghosal et al., noted that vision transformers pre-trained on ImageNet22k achieved better worst group accuracies on benchmarks with spurious correlations than popular CNN models. In particular, with a VIT-B/16 model, they achieve 89.3%89.3\% worst group accuracy on Waterbirds. With the ConvNext Large model pretrained on ImageNet22k with ImageNet1k finetuning, we achieve 88.9%88.9\% worst group accuracy on the Waterbirds dataset. Notably, ConvNext is a CNN model and not a vision transformer. Generally, we observe that the models that provide the best in-distribution performance also provide better WGA. In our experiments, we did not observe qualitative differences between the results for vision transformers and CNN models.

ERM features are sufficient for SOTA performance. The DFR WGA results for the Waterbirds, CelebA and FMOW datasets significantly improve upon the previous best reported results, to the best of our knowledge. In particular, the ConvNext Large model pretrained on ImageNet22k with ImageNet1k finetuning achieves 97.2% DFR WGA on Waterbirds and 92.2% on CelebA; on FMOW, we only considered smaller ConvNext versions due to computational constraints, still achieving 50.6% DFR WGA with ConvNext Small pretrained on ImageNet22k; the current best results on the WILDS leaderboard for this dataset are 47.6%47.6\% followed by 35.5%35.5\% . We note that our DFR evaluation uses the validation set to train the last layer of the model, similarly to e.g. Nam et al., , and unlike most standard group robustness methods which only use the validation set to tune the parameters. However, the results presented in this section prove that standard ERM with a strong pretrained model can achieve outstanding results on the robustness benchmarks, significantly improving upon specialized group robustness methods using a weaker model.

2 Does training on the target data improve features?

Above, we have shown that the choice of the base model architecture and pretraining has a large effect on the quality of the learned feature representations, as measured by DFR WGA. It is then natural to ask how much of the feature learning happens during training on the target data, and how much is simply transferred from the pretraining task. To answer this question, for the models trained in the previous section, we run the DFR evaluation on the initial weights of the models, without training the feature extractor on the target data. We report the DFR WGA and DFR ss-WGA results in Figure 4. We repeat this experiment on the MultiNLI dataset in Appendix Table 5.

Surprisingly, we find that on Waterbirds, for most models the improvement from training on the target (Waterbirds) data is very small, if any. For example, for the ImageNet1k-pretrained ResNet-50 model that was not trained on Waterbirds data at all This model is used as initialization by most works that report experiments on the Waterbirds dataset. , we get 88.2%88.2\% worst group accuracy by simply training the last layer on the validation data with DFR. If we finetune the feature extractor on the Waterbirds training data, we can achieve 92.9%92.9\% DFR WGA. For reference, the state-of-the-art group DRO method achieves 91%91\% WGA on Waterbirds with this architecture.

Furthermore, with the ConvNext Large model pretrained on ImageNet22k, we get 94%94\% DFR WGA without training the feature extractor on the Waterbirds data, exceeding the best results previously reported in the literature, to the best of our knowledge. This strong performance is not particularly surprising, as ImageNet22k has many of the Waterbirds bird types as classes. The performance is almost unchanged by training on the target data: DFR WGA after training is 94.3%94.3\%. From these results, we can conclude that Waterbirds performance should not be used as a primary metric for feature learning performance, especially if large-scale pretraining is used! Indeed, it is possible to achieve outstanding performance, exceeding the previously reported state-of-the-art, without training the feature extractor on the target data.

On the other datasets, the feature learning is more pronounced: the DFR WGA improves after training for all the models considered. However, on CelebA it is still possible to achieve 88.3%88.3\% WGA without training on CelebA data, with ConvNext XLarge pretrained on ImageNet22k, while the best result that we were able to achieve with feature extractors trained on CelebA is 92.2%92.2\%.

Is the spurious feature representation improved during training? We additionally explore the quality of representation of the spurious feature via DFR ss-WGA. We show the results in Figure 4 (bottom row). Interestingly, on CelebA the spurious gender feature becomes less predictable during training for many of the models. On FMOW and Waterbirds there is no consistent trend, and the spurious feature does not become significantly more or less predictable during training.

3 Effect of pretraining strategy

Finally, using a fixed ResNet-50 model architecture we evaluate the effect of pretraining strategy. On each of the four datasets, we train a ResNet-50 model initialized with (1) random initialization, (2) supervised pretraining, (3) DINO pretraining , (4) SimCLR pretraining and (5) Barlow Twins pretraining on ImageNet1k. We report the results in Figure 5(a). On Waterbirds and FMOW datasets, the randomly initialized model does not provide competitive performance. On CelebA, it still underperforms the pretrained models, but the gap is much smaller. Among all pretraining methods, supervised is preferable, but the contrastive methods are highly competitive.

In Appendix Table 4 on MultiNLI with BERT models, we also show that pretraining is crucial for strong performance, while the specific choice of pretraining data has a smaller effect.

Effect of Regularization

Regularization is used to combat overfitting, including reliance on spurious features. We consider two regularization techniques: weight decay and data augmentation.

Effect of weight decay. Using a ResNet-50 model pretrained on ImageNet1k, we run training with a range of weight decay values on each of the four datasets, and report the results in Figure 6. For the base model WGA performance, it is generally helpful to set the weight decay to non-zero values. Indeed, on CelebA the base model WGA for no weight decay is the worst across all weight decay values, losing to the best weight decay by ≈5%\approx 5\%. However, the no weight decay model is in fact the best model according to DFR WGA! On these datasets, strong weight decay allows the model to rely less on the spurious features in the last layer (leading to higher base WGA), but does not improve the learned feature representations (similar or worse DFR WGA). On other datasets no weight decay also provides competitive performance. FMOW is the only dataset considered where non-zero weight appears to be helpful for learning high quality representations of the core features. In Appendix Table 6 we report analogous results on the MultiNLI text classification problem.

Generally, weight decay affects both the feature extractor and the last classification layer. While in some cases higher weight decay is helpful for learning better features, zero weight decay is competitive across the board. This observation is especially interesrting given that Sagawa et al., showed that group DRO requires stronger than usual weight decay to achieve good performance on Waterbirds and CelebA datasets.

Effect of Data Augmentation. Next, we consider 5 data augmentation policies: (1) no augmentations, (2) default augmentations, i.e. random crops and horizontal flips, (3) MixUp combined with default augmentations, (4) Random Erasing , and (5) AugMix . We train ImageNet1k-pretrained DenseNet-121 models on CXR and ResNet-50 models on Waterbirds, CelebA and FMOW with each of the augmentation policies, and report the results in Figure 5(b). AugMix provides the best performance on each dataset. However, we find that data augmentation is generally not required to achieve strong performance on any of the datasets with the exception of CXR: the model trained without augmentations is competitive across the board. Moreover, data augmentation can hurt the learned features. For example, while MixUp is helpful on Waterbirds and CelebA, it significantly hurts the preformance on FMOW, with 30%30\% DFR WGA compared to 42%42\% for the model trained without augmentation. We hypothesize that on the FMOW dataset, which contains highly detailed satellite images, mixing the images makes training overly challenging.

Similarly, Random Erasing hurts the DFR WGA on Waterbirds. We hypothesize that the randomly erased image block is more likely to fully cover the bird features than the background features, as the bird occupies a small fraction of the image relative to the background. Consequently, the model is incentivised to focus on the spurious feature: the model trained with Random Erasing learned the highest quality representation of the spurious feature across all augmentations with DFR ss-WGA of 91.8±0.1%91.8\pm 0.1\% across 33 runs, compared to e.g. 91.4±0.5%91.4\pm 0.5\% for the model trained with no augmentation.

Balestriero et al., report a related observation for models trained on ImageNet: in some cases data augmentation affects different classes disproportionately, and best data augmentation policies for mean accuracy can lead to poor worst-class accuracy.

Discussion

The worst group performance of a model is affected by two factors: the quality of the representation of the core features produced by the feature extractor and the weight assigned to the core features in the last classification layer. In contrast to prior work, we consider the quality of the feature extractor in isolation, focusing on realistic datasets and large-scale models. We find that many of the popular group robustness methods improve the worst group performance primarily by learning a better last layer and not by learning a better feature representation. Similarly, regularization techniques such as early stopping and strong weight decay can improve the worst group accuracy by learning a better last layer, but do not lead to a consistent improvement in terms of the quality of the learned feature representations. On the other hand, the base model architecture and pre-training strategy have a major effect on the quality of the feature representations.

Our observations suggest an important open question: is it possible to significantly improve upon standard ERM in terms of the quality of the learned representations for a given base model? In future work, it would be interesting to evaluate methods such as Rich Feature Construction , gradient starvation , ensembling and other techniques for increasing feature diversity [e.g. 48]. We hope that our work will also inspire new group robustness methods targeted specifically at improving the quality of the core feature representations.

References

Appendix Outline

This appendix is structured as follows. In Section A we describe the datasets, augmentation policies and models used in this paper. In Section B we provide details on the methods, implementations and hyper-parameters, as well as detailed results for the experiments in Section 5. In Section C we provide additional details on the experiments in Section 6. In Section D we provide additional details and results for the experiments in Section 7. In Section E we provide additional results on the MultiNLI dataset. Finally, in Section F we describe the limitation, broader impact, compute and licenses.

Tools and packages. During the work on this paper, we used the following tools and packages: NumPy , SciPy , PyTorch , TorchVision , Jupyter notebooks , Matplotlib , Pandas , Weights&Biases , timm , transformers , vissl .

Appendix A Data and Models

In this section, we describe the datasets, data augmentation policies and models used throughout the paper.

We perform experiments on 44 image classification and 22 text classification problems. We illustrate the image datasets in Figure 7 and the text datasets in Figure 8.

Waterbirds. The Waterbirds dataset is described in Figure 7. The dataset contains images of birds from the CUB dataset pasted on the backgrounds from the Places dataset . The spurious attribute ss describes the type of background (water or land) and the core feature associated with the target yy is the type of the bird (waterbird or landbird). For a detailed description of the data generating process, see Sagawa et al., . The background is spuriously correlated with the bird type in the training datasets: waterbirds are more likely to be placed on a water background, and landbirds are more likely to be placed on a land background. There are 44 groups defined to the tuples (y,s)(y,s).

CelebA. The CelebA hair color dataset is described in Figure 7. The dataset contains photos of celebrities from the CelebA dataset Liu et al., . The core attribute associated with the target yy is the hair color (blond vs non-blond). The gender serves as a spurious feature ss: the vast majority of blond people in CelebA are female. There are 44 groups defined to the tuples (y,s)(y,s).

FMOW. The WILDS-FMOW dataset is described in Figure 7. This dataset is a part of the WILDS benchmark , and was originally collected by Christie et al., . The dataset contains satellite images, and the target yy describes the type of building or land use shown in the image. There are 6262 classes. The spurious attribute ss corresponds to the region (Asia, Europe, Africa, America, Oceania) shown in the image. The training data additionally contains another group Other, which is dropped during evaluation. For this dataset, following the WILDS benchmark, we define groups by just the value of the spurious attribute: g=sg=s. In particular, worst group accuracy corresponds to the worst accuracy across regions. The regions are represented unequally in the data, leading to unequal performance. Moreover, the test images were taken several years later than the train images, constituting an additional type of distribution shift. We note that for the DFR evaluation, we use the validation data to train the last layer of the model, which means that our results would be classified as non-standard submissions to the WILDS leaderboard.

CXR. The CXR dataset is described in Figure 7. The images are taken from the CXR-14 dataset . CXR-14 is a multi-label dataset, where each label corresponds to a disease, and one image can show multiple diseases. We perform the pneumothorax classification task, i.e. all images that have the pneumothorax label have y=1y=1 and all images that do not have y=0y=0. The dataset contains several images for some of the patients; there is no patient overlap between the train, test and validation splits. Oakden-Rayner et al., identified a hidden stratification in this dataset: a lot of images showing patients with pneumothorax showed a chest drain, which is a treatment for the pneumothorax disease. The neural networks trained on this dataset are using the chest drain as a shortcut feature, and perform much worse when classifying images without the chest drain. The labels for the spurious feature are only available on the validation and test datasets, and only for the images showing sick patients, i.e. there is no (y=0,s=1)(y=0,s=1) group. There are 33 groups corresponding to available pairs (y,s)(y,s). Following prior work [e.g. 65, 48, 70], we compute three AUC values: for classifying group against group 11, group against group 22 and group against the combined groups and 11; instead of worst group accuracy we report the lower of the first two AUC values, and instead of mean accuracy we report the last AUC value throughout the experiments.

Civil Comments The Civil Comments Coarse dataset is described in Figure 8. The dataset was originally collected in Borkan et al., and is a part of the WILDS benchmark . This dataset contains comments that are classified as toxic or not toxic. We use the coarse version of the dataset, following Idrissi et al., . The spurious attribute ss determines whether or not the comment mentions one of the following protected attributes: male, female, LGBT, black, white, Christian, Muslim, other religions. These protected attributes are mentioned more frequently in toxic comments compared to neutral comments, constituting a spurious correlation. There are 44 groups corresponding to the pairs (y,s)(y,s).

MultiNLI The MultiNLI dataset is described in Figure 8. It contains pairs of sentences, and the label yy describes the relationship between the sentences: contradiction, entailment or neutral. The spurious attribute ss describes the presence of negation words, which appear much more frequently in the examples from the negation class. There are 66 groups corresponding to the pairs (y,s)(y,s).

A.2 The nature of the spurious correlations

The nature of the spurious correlation differs between the datasets that we consider. On Waterbirds, CelebA, Civil Comments and MultiNLI, the spurious attribute is correlated with the target on train, but is not generally predictive of the class label across the groups. In particular, we wish to train a model that ignores the spurious attribute in its predictions, as using the spurious attribute hurts the performance on some of the groups. By applying DFR on a group-balanced validation set, we try to train a model that ignores the spurious feature.

On the FMOW dataset, the groups correspond to the spurious attribute (region), and not pairs (y,s)(y,s). As the regions are not represented equally, standard ERM is incentivised to perform well on the majority groups, with less weight on the minority groups. In this case, we do not wish to remove the reliance on the spurious attribute, but instead we wish to find a model that performs well on the minority groups. We apply DFR to this dataset by retraining the last layer on a group-balanced validation set, analogously to the other datasets.

Finally, in the CXR dataset, there are no examples showing healthy patients with a chest drain (to the best of our knowledge). Consequently, the reliance on the spurious attribute ss does not necessarily hurt performance on the other groups, unlike e.g. on the Waterbirds dataset. In this case, we do not wish to remove ignore the feature ss, but instead we wish to find a model that performs well on the images with s=0s=0 (no chest drain). Applying DFR to this dataset is not straightforward, as (1) we do not have examples from the (y=0,s=1)(y=0,s=1) group, and (2) we do not wish to remove the reliance on the ss attribute. We found that the best approach in this case was to simply train a logistic regression model on all of the available validation data, without any balancing. DFR generally provides a smaller improvement on this dataset, and behaves less consistently, as we see e.g. in Figure 3.

A.3 Data augmentation and preprocessing

On the image datasets, we consider several data augmentation policies. In Figure 9, we provide the code implementing the No augmentation and Default augmentation policies in torchvision. The other policies are more involved, so we do not include the code here. All policies normalize the data analogously to the code in Figure 9.

No augmentation. This policy resizes the data to a fixed resolution and applies channel normalization.

Default policy. This policy additionally applies random crops and horizontal flips.

MixUp. For this policy, we use the Default policy to initially preprocess the images, and then apply MixUp with the mixing parameter α=0.2\alpha=0.2.

Random Erasing. This policy is described in Zhong et al., . It randomly erases rectangular blocks of the image, and replaces them with uniform grey blocks. We use the implementation in timm.data.random_erasing in the timm package.

Augmix. We adapt the official implementation of the AugMix policy available here.

Text models. For data preprocessing in the experiments on text data, we use the BERT tokenizer: BertTokenizer.from_pretrained("bert-base-uncased") from the transformers package.

A.4 Models

Here, we list the models used in this paper.

ResNet-50. On the Waterbirds, CelebA and FMOW by default we use the ResNet-50 model He et al., pretrained on ImageNet. We use the model implemented in the torchvision package: torchvision.models.resnet50(pretrained=True). In Figure 5, we additionally consider the randomly initialized model torchvision.models.resnet50(pretrained=False), and the models pre-trained with SimCLR and Barlow Twins contrastive learning methods, imported from the vissl package (models available here).

DenseNet-121. On CXR, by default we use the DenseNet-121 model implemented in the torchvision package: torchvision.models.densenet121(pretrained=True). We also consider this model without ImageNet pretraining: torchvision.models.densenet121(pretrained=False).

Other image models. In Section 6, we additionally consider a broad range of architectures and pretraining methods, which we briefly list here. We use the following models from torchvision, all pretrained on ImageNet1k: ResNet-18, ResNet-34, ResNet-50, ResNet-101, ResNet-152 ; Wide-ResNet-50-2 ; ResNext-50-32×4d32\times 4d ; DenseNet-121 ; VGG-16, VGG-19 ; AlexNet . We also use the following models from the timm package: ConvNext-Small, ConvNext-Base, ConvNext-Large, ConvNext-XLarge , pretrained on either ImageNet1k or ImageNet22k; ViT-Small, ViT-Base, ViT-Large, ViT-Huge , pretrained on either ImageNet1k or ImageNet22k; BEiT-Base, BEiT-Large , pretrained on either ImageNet1k or ImageNet22k; DEiT-Small, DEiT-Base , pretrained ImageNet1k. We also use ViT-Small, ViT-Base models with DINO pretraining on ImageNet1k available here. Finally, we use ViT-Base, ViT-Large and ViT-Huge models with MAE pretraining on ImageNet1k , with or without supervised finetuning on ImageNet1k, available here.

BERT model. On the text classification problems, we use the BERT for classification model from the transformers package: BertForSequenceClassification.from_pretrained(’bert-base-uncased’, num_labels=num_classes).

Appendix B Details: ERM vs Group Robustness

ERM, RWY and RWG. We report the hyper-parameters used on each of the datasets in Table 2. We did not tune the hyper-parameters for ERM, RWY and RWG aside from the learning rate for the text classification problems. For all vision datasets, we used the default data augmentation policy (see Section A.3). We describe the default model choices in Section A.4. The RWY method is implemented by providing a sampler to the DataLoader in PyTorch, which samples the datapoints from different classes with the same frequency. The RWG method is implemented analogously, and samples datapoints from different groups with the same frequency.

Group DRO. The training objective used in Sagawa et al., is the following:

where fθ(⋅)f_{\theta}(\cdot) is a neural network model with parameters θ\theta, l(⋅,⋅)l(\cdot,\cdot) is a loss function (cross-entropy for classification), G\mathcal{G} is a set of all groups, ngn_{g} is the size of the group gg, and CC is a generalization adjustment hyper-parameter. We follow Sagawa et al., to choose hyper-parameter combinations for tuning group DRO on Waterbirds, CelebA and MultiNLI.

In particular, on Waterbirds we considered the following combinations of the initial learning rate lrlr and weight decay wdwd: (lr=10−3,wd=10−4),(lr=10−4,wd=0.1){(lr=10^{-3},wd=10^{-4}),(lr=10^{-4},wd=0.1)} and (lr=10−5,wd=1)(lr=10^{-5},wd=1). We varied the parameter CC in the range {0,1,2,3,4,5}\{0,1,2,3,4,5\} for each combination of learning rate and weight decay. Additionally, we varied weight decay in range {0,10−4,10−4,3⋅10−4,10−3,3⋅10−3,10−2,3⋅10−2,0.1,0.3,1}\{0,10^{-4},10^{-4},{3\cdot 10^{-4}},10^{-3},{3\cdot 10^{-3}},10^{-2},{3\cdot 10^{-2}},0.1,0.3,1\} for fixed values of learning rate 10−310^{-3} and C=0C=0.

On CelebA, we considered the following combinations of lrlr and wdwd: (lr=10−4,wd=10−4),(lr=10−4,wd=10−2){(lr=10^{-4},wd=10^{-4})},{(lr=10^{-4},wd=10^{-2})} and (lr=10−5,wd=0.1){(lr=10^{-5},wd=0.1)}, varying CC in the same range as on Waterbirds. For weight decay ablation we considered values in {0,10−5,3⋅10−5,10−4,3⋅10−4,10−3,3⋅10−3,10−2}\{0,10^{-5},3\cdot 10^{-5},10^{-4},3\cdot 10^{-4},10^{-3},3\cdot 10^{-3},10^{-2}\}.

On MultiNLI, we use learning rate 2⋅10−52\cdot 10^{-5}, and vary weight decay in the range {0.,0.01,0.1,1.0}\{0.,0.01,0.1,1.0\} and CC in the range {0,1,3,5}\{0,1,3,5\}.

On CivilComments, we considered the following combinations of lrlr and wdwd: (lr=10−5,wd=0.1),(lr=10−5,wd=0.01),(lr=10−4,wd=0.01){(lr=10^{-5},wd=0.1)},{(lr=10^{-5},wd=0.01)},{(lr=10^{-4},wd=0.01)}; for each combination we varied CC in {0,3,5}\{0,3,5\}.

On FMOW, the learning rate and weight decay pairs were (lr=10−3,wd=10−3){(lr=10^{-3},wd=10^{-3})}, (lr=10−4,wd=10−3){(lr=10^{-4},wd=10^{-3})}, (lr=10−4,wd=10−2){(lr=10^{-4},wd=10^{-2})}, (lr=10−5,wd=10−1){(lr=10^{-5},wd=10^{-1})}. For each lrlr and wdwd combination we varied CC in the same range as on Waterbirds and CelebA datasets.

For all image datasets, we use the default data augmentation policy (see Section A.3). In all runs we used the same batch size and train for the same number of epochs as the corresponding hyper-parameters in ERM (see Table 2).

B.2 DFR implementation and hyper-parameters

As explained in Section A.2, on CXR we train the logistic regression model on all of the validation set without group balancing. Further, on CXR we tune the regularization strength parameter according to the worst AUC, and not worst group accuracy.

We additionally compute DFR ss-WGA by using DFR (with the same hyper-parameters and implementation) to predict the spurious attribute ss (instead of the class label yy) from the learned features.

B.3 Results

We provide detailed results for all methods on all datasets in Table 1. Group-DRO significantly improves the base model performance compared to all other methods across the board. After applying DFR, the performance across the different methods is very similar, although Group-DRO with early stopping still typically provides a small improvement. The improvement is, however, very small compared to the improvement from using a better base model (see Section 6).

Appendix C Details: Effect of the Base Model

For all experiments in Section 6, we use the default data augmentation policy (see Section A.3). We consider a broad range of models and architectures (see Section A.4). We ran all models with the default hyper-parameters provided in Table 2. For CXR-14 dataset, we used class reweighting due to heavy class imbalance present in the train data (95%95\% of train images are from the negative class).

We note that the default hyper-parameters are suboptimal for some of the ViT models, which lead to poor performance for some of the models. With more tuning, we expect that it should be possible to improve the results further for all the considered models, especially ViT-based.

Appendix D Details: Effect of Regularization

Effect of weight decay. For the experiments on the effect of weight decay we use the default models described in Section A.4, and default hyper-parameters in Table 2, and vary the weight decay strength in the range {0,10−5,10−4,3⋅10−4,10−3,3⋅10−3,10−2}\{0,10^{-5},10^{-4},3\cdot 10^{-4},10^{-3},3\cdot 10^{-3},10^{-2}\}. We report the results in Figure 6.

Effect of data augmentation. For the experiments on the effect of data augmentation, we use the default models described in Section A.4, and default hyper-parameters in Table 2, and apply the data augmentation policies desribed in A.3. We report the results in Figure 5(b).

Effect of training length on ERM. We plot the DFR WGA, DFR ss-WGA, as well as base model WGA and mean accuracy as a function of training epoch for ERM training in Figure 10. On all datasets, we observe that the DFR WGA quickly converges and stays roughly constant throughout training. On all datasets, 55 epochs is sufficient for near-optimal performance. Longer training generally does not help or hurt DFR WGA, even when it hurts the base model.

Effect of training length on Group-DRO. We repeat the same experiment, but with Group-DRO instead of ERM training in Figure 11. Again, we observe that 55 epochs are generally sufficient for near-optimal performance. Interestingly, for Group-DRO, DFR WGA does deteriorate over time on Waterbirds but not as significantly as the base model WGA.

Appendix E Additional results on MultiNLI

For all experiments in this section, we train the models for 55 epochs with learning rate 10−510^{-5} and weight decay.

Effect of base model. In Table 3, we report the results for BERT-Base, BERT-Large, DeBERTa-Base, and DeBERTa-Large models . For BERT models, following Sagawa et al., , we use the cached tokenizer outputs with maximum sequence length 150150. For DeBERTa models, we re-tokenize the dataset with the corresponding tokenizer with a maximum sequence length of 220220. We train all models for 55 epochs. The advanced DeBERTa-Large model provides the best base performance and DFR WGA.

Effect of pretraining. In Table 4 we evaluate the results of BERT-Base models with different types of pretraining on MultiNLI. This experiment is analogous to the experiment for image classification problems presented in Figure 5(a). We find that pretraining is necessary to achieve strong performance on this dataset, but different pretraining datasets lead to competitive results.

Effect of training on target data. In Table 5 we evaluate the effect of training on the MultiNLI dataset for the BERT and DeBERTa pretrained models. This experiment is analogous to the experiment for image classification problems presented in Figure 4. We find that after training on the target data both the core features (DFR WGA) and the spurious features (DFR ss-WGA) become significantly more decodable. This result is in contrast to the results on Waterbirds in Figure 4, where DFR WGA is not significantly improved from training.

Effect of weight decay. In Table 6 we report the results of the weight decay ablation for the BERT-Base model on MultiNLI; this model uses the AdamW optimizer , so we consider larger values of weight decay.

Appendix F Broader Impact and Limitations

Limitations. While we consider a wide range of factors that affect the feature learning under spurious correlations, we inevitably do not cover all the possible factors. In particular, it would be interesting to consider the effect of regularization methods beyond weight decay and early stopping, and methods for diverse feature learning such as DivDis , or the method of Teney et al., . As another limitation, while DFR performs well in our experiments, it is not guaranteed to learn an optimal linear classifier with the given features; further improvements in learning the last layer can be used to refine the results of our study. Despite these limitations, we believe that our work provides a comprehensive analysis of the feature learning under spurious correlations.

Broader impact. Research on spurious correlations is closely related to ML Fairness . We hope that our work can motivate further research in fairness, where techniques similar to DFR can be considered to improve the fairness of ML models. A potential negative outcome that can result from misinterpretation of our analysis is if the practitioners assume that spurious correlations are not an important issue, as ERM learns high quality representation of the core features. We emphasize that ERM still performs suboptimally (see Figure 1), as it does not provide a correct weighting for the features in the final classification layer. Spurious correlations are a significant practical issue that should be considered carefully in real-world applications.

Compute. We estimate the total compute used in the process of working on this paper at roughly 30003000 GPU hours. The compute usage is dominated by the experiments presented in Figure 3, where we trained a large number of large-scale models on 44 vision datasets. The tuning of Group-DRO hyper-parameters was also relatively compute-heavy. The experiments were run on GPU clusters on Nvidia Tesla V100, Titan RTX, RTX8000, 3080 and 1080Ti GPUs.

Licenses. The Civil Comments dataset is distributed under the CC0 license. The FMOW dataset is under the FMoW Challenge Public License. The Places dataset is under the CC BY license. For the details of the license for the MultiNLI dataset, see Williams et al., .