Don't Judge an Object by Its Context: Learning to Overcome Contextual Bias

Krishna Kumar Singh, Dhruv Mahajan, Kristen Grauman, Yong Jae Lee, Matt Feiszli, Deepti Ghadiyaram

Introduction

Visual context serves as a valuable auxiliary cue for the human visual system for scene interpretation and object recognition . Context can either be a co-occurrence of objects and scenes (e.g., “boat” is often present in “outdoor waters”) or of two or more objects in a given scene (e.g., “skis” often co-occur with a “skier”). Context becomes especially crucial for our visual system when the visual signal is ambiguous or incomplete (e.g., due to occlusion, viewpoint of the scene capture, etc.). Past research explicitly models context and shows benefits on standard visual tasks such as classification and detection . Meanwhile, convolution networks by design implicitly capture context.

Deep networks rely on the availability of large-scale annotated datasets for training. As highlighted in , despite the best efforts of its creators, most prominent vision datasets are afflicted with several forms of biases. Let us consider an object category “microwave.” A significant portion of images belonging to this category are likely to be captured in kitchen environments, where other objects such as “refrigerator,” “kitchen sink,” and “oven” frequently co-occur. This may inadvertently induce contextual bias in these datasets, which would consequently seep into models trained on them. Specifically, in the process of learning features that separate positive and negative instances in such a (biased) training dataset, a deep discriminative model can very often also strongly capture the context co-occurring with the category of interest. This issue is exacerbated in a setting where we do not have explicit location annotations (e.g., bounding boxes and segmentation masks) of such biased categories, and a model being trained has to rely solely on image-level annotations to perform multi-label classification. Having a model implicitly learn to localize such context-biased categories in the absence of location annotations is challenging.

Does it even matter if a model inadvertently learns such correlations? We believe this can cause problems on two fronts: (1) failing to identify “microwave” in a different context such as an “outdoor” scene or in the absence of “refrigerator” and (2) hallucinating “refrigerator” even in an indoor kitchen scene containing only “microwave.” The issue of co-occurring bias is also prevalent in visual attributes . For example, in the Deep Fashion dataset , the attribute “trapeze” strongly co-occurs with “striped.” This results in a less credible classifier that has a hard time recognizing “trapeze” in clothes with “floral.” Recent research has identified far more serious mistakes made by trained models due to inherent biases in both language and vision datasets – learning correlations between ethnicity and certain sport activities , gender and profession , and age and gender of celebrities . Such grave confusion caused due to biases in the data impedes the deployment of these models in real-world applications.

Given these issues, our goal is to train an unbiased visual classifier that can accurately recognize a category both in the presence and absence of its context. Specifically, given two categories with a strong co-occurring bias, our aim is to accurately recognize them when either one occurs exclusively, and at the same time not hurt the performance when they co-occur. To this end, we propose two key ideas. First, we hypothesize that a network should learn about a category by relying more on its corresponding pixel regions than those of its context. Since we only have class labels, we use class activation maps (CAM) as “weak” location annotations and minimize their mutual spatial overlap.

Building on this, we devise a second method that learns feature representations to decorrelate a category from its context. While the entire feature space learned by the network jointly represents category and context, we explicitly carve out a subspace to represent categories that occur away from typical context. We learn this feature subspace only from training instances where a biased category occurs in the absence of its context. In all other cases, the model should also leverage context and thus the entire feature space. At test time, we make no such distinction and the entire feature space is equally leveraged. Therefore, in the example from Fig. 1, our goal is to learn a feature subspace to represent “skateboard” while the entire feature space jointly represents “skateboard” and “person.”

Through extensive evaluation, we demonstrate significant performance gains for the hard cases where a category occurs away from its typical context. Crucially, we show that our framework does not adversely effect recognition performance when categories and context co-occur. To summarize, we make the following contributions:

With an aim to teach the network to “learn from the right thing,” we propose a method that minimizes the overlap between the class activation maps (CAM) of the co-occurring categories (Sec. 4.1).

Building on the insights from the CAM-based method, we propose a second method that learns feature representations that decorrelate context from category (Sec. 4.2).

We apply both methods on two tasks: object and attribute classification, and 44 datasets, and achieve significant boosts over strong baselines for the hard cases where a category occurs away from its typical context (Sec. 5).

Related work

Addressing biases: Prior work has shown that existing datasets suffer from bias and are not perfectly representative of the real world. Hence, a model trained on such data will have difficulty generalizing to non-biased cases. Attempts to reduce dataset bias include domain adaptation techniques and data re-sampling , e.g., so that minority class instances are better represented. One limitation of data re-sampling is that it can involve reducing the dataset, leading to sub-optimal models. Recent adversarial learning approaches try to mitigate bias from the learned feature representations while optimizing performance for the task at hand (e.g., removing gender bias while classifying age). However, these methods would not be directly applicable for mitigating contextual bias, as context (the bias factor) can still be useful for recognition—so it cannot be simply removed. Others study various forms of bias in the context of image captioning (e.g., gender bias) , image classification (e.g., ethnicity bias) , and object recognition (e.g., socio-economic bias) . Overall, contextual bias in visual recognition remains relatively under explored.

Co-occurring-bias: Contextual bias is a well-studied problem in the field of natural language processing , however, it is much less studied in the computer vision community. In vision, most efforts consider context as a useful cue . A few efforts have shown that a recognition model will fail to recognize an object without its co-occurring context, but do not propose a solution .

A recent method reduces contextual bias in video action recognition , but it relies on temporal information and thus cannot be applied to the image recognition problems we tackle in this work. A pre-deep learning approach reduces the correlation (bias) between visual attributes by leveraging additional knowledge in the form of semantic groupings of attributes. Recently tried to reduce contextual bias for object detection by learning focused foreground features, but they require expensive bounding-box annotations. In contrast, our deep learning approach does not require any additional supervision apart from the object/attribute class labels. Most importantly, to our knowledge, there is no prior work focusing on mitigating contextual bias for object classification as we do in this paper. Relation to few-shot learning: Lastly, contextual bias could also be formulated as a few-shot or class imbalance problem, since images in which objects appear without their usual co-occurring context (e.g., keyboard without a mouse next to it) are relatively rare. However, treating such rare (exclusive) images as a separate class or simply assigning them higher weight can be sub-optimal, as we show in our experiments.

Problem setup

Our method operates on the premise that the training data distribution corresponding to a few categories suffers from co-occurring bias. We henceforth refer to them as biased categories. We make no such assumptions about the test data distribution. For example, COCO-Stuff has 22092209 images where “ski” co-occurs with “person,” but only has 2929 images where “ski” occurs without “person.” A model trained on such skewed data may fail to recognize when “ski” occurs in isolation. Our goal is to learn a feature space that is robust to such training data biases. In particular, given a (presumably) unbiased test dataset, our goal is to (1) correctly identify “ski” when it occurs in isolation and (2) not lose performance when “ski” co-occurs with “person.” A key aspect of our approach is to identify most biased categories for a given dataset, which we describe next.

Approach

Our first method relies on class activation maps (CAM) as “weak” automatically inferred location annotations and minimizes their spatial overlap between biased categories (Sec. 4.1). Building on the observations from this CAM-based approach, we propose a second method which learns a feature space by encouraging context sharing when a biased category co-occurs with context while suppressing context when it occurs in isolation (Sec. 4.2).

CAM offers two nice properties: (1) it is learned only through class labels without requiring any annotation effort and (2) it is fully differentiable, and thus can be integrated in an end-to-end network during training.

Fig. 3 for the entire approach. As we show in results (Sec. 5), our CAM-based method successfully learns to rely more on the biased category’s pixel regions thereby improving recognition performance. Our method yields large gains when a biased category occurs in the absence of its typical context. However, it sometimes hurts performance when biased category co-occurs with context (discussed later in Fig. 7). One reason could be that the pixel regions surrounding the co-occurring category also offer useful complementary information for recognizing the biased category. By discouraging mutual spatial overlap, CAM-based approach may not be able to leverage this information. This key insight led to the formulation of our next approach, which splits the feature space into two and separately represents context and category, while posing no constraints on their spatial extents.

2 Feature splitting and selective context suppression

Rather than optimizing CAMs, we propose to learn a feature space that is robust to the inherent co-occurring biases in the training data. We observe that cases when a biased category co-occurs with context are often visually distinct from those where it occurs exclusively (see Fig. 1). This motivates us to learn a dedicated feature (sub) space to represent biased categories occurring away from their typical context. While the entire feature space learned by the model jointly represents context and category, this dedicated subspace should decouple the representations of a category from its context. We learn this feature subspace only from training instances where biased categories occur in the absence of their typical context. These modifications only affect training; at inference time the architecture is identical to the standard model.

Figure 4 illustrates the proposed method. While a standard classifier jointly encodes category and context, it fails to recognize biased categories occurring without context. By contrast, our approach splits the feature space and represents biased categories occurring without context in a dedicated subspace. As we will show in results, due to selective context suppression, this feature subspace successfully captures category-specific information. Furthermore, in the second subspace, our method effectively leverages context when available and jointly encodes it with category.

3 Training setup

Experiments

In this section, we study the effectiveness of our approach across two tasks: object and attribute classification. We first describe our evaluation setup then report qualitative and quantitative performance on four image datasets against competitive baselines.

Aside from a standard classifier trained with a binary cross-entropy loss for each category, we compare with the following state-of-the-art methods that tackle the issue of co-occurring bias: (1) class balancing loss by treating the scenarios where biased categories occur exclusively as tail classes and (2) attribute decorrelation approach , where we replace the hand-crafted features with deep network features (conv5 features of ResNet-50) for a fairer comparison. To further test the strength of our method, we designed the following competitive baselines:

remove co-occur images shares the same motivation as (2) but instead we remove training instances where the biased category and context co-occur.

weighted loss, where we apply 1010 times higher weight to the loss when biased categories occur exclusively.

negative penalty, where we assign a large negative penalty if the network predicts co-occurring category in cases where a biased category occurs exclusively.

1 Object Classification Performance

Next, we observe that both ours-CAM and ours-feature-split outperform standard by 1.9%\mathbf{1.9\%} and 4.3%\mathbf{4.3\%} respectively on the exclusive test set. ours-feature-split has a very marginal drop of 0.2%0.2\% on the co-occurring split, compared to standard, while the performance drop is higher for ours-CAM. On categories such as “ski” and “skateboard” which have a very high co-occurrence bias with “person”, the mAP boost from ours-feature-split is 24.2%\mathbf{24.2\%} and 19.5%\mathbf{19.5\%} respectively (per-class mAP for both methods in supp. material).

Comparison with other baselines: We note that remove co-occur images approach performs poorly as it relies only on the exclusive images of the biased categories and do not take advantage of the vast amount of co-occurring images which supply complementary visual information. weighted loss improves performance on the exclusive test split compared to ours-feature-split (30.4% vs. 28.8%), but significantly hurts performance on co-occurring split (60.8% vs. 66.0%). negative penalty does not hurt co-occurring split, but has inferior performance compared to our methods on the exclusive split. We also note that performance trends exhibited by these methods are consistent across all other datasets we test on; for all future experiments, we compare our methods with standard and class balancing loss.

Performance on the non-biased categories: We evaluate on the 6060 non-biased object categories of COCO-Stuff and observe that both ours-CAM and ours-feature-split perform on par with standard, with a very mild drop of  0.2%~{}0.2\% overall mAP (details in supp. material). This indicates that our methods, while successfully improving performance for the biased categories, do not adversely effect the rest of the (non-biased) categories.

1.2 Qualitative Analysis

2 Cross dataset experiment on UnRel

3 Attribute Classification

Here, we show that our approach of reducing contextual bias generalizes to attributes. Our CAM-based approach is not applicable to attributes, as they lack well-defined spatial extents (details in Sec. 4.1). As noted in Sec 5.1, the inherent contextual bias and difficulty in recognizing biased categories in the absence of their context leads to low scores on exclusive test split for all methods and datasets.

Results on DeepFashion: As is the common practice, we report per class top-33 recall on DeepFashion . From Table 4, we note that ours-feature-split outperforms standard by a significant margin on both test splits. For attributes like trapeze and bell which exhibit strong co-occurrence with striped and lace respectively, ours-feature-split yields a boost of 21.2%\mathbf{21.2\%} and 17.4%\mathbf{17.4\%} top-3 recall respectively compared to standard classifier. We present per-attribute results and comparisons with other baselines in the suppl. material.

Results on Animals with Attributes: Animals with Attributes suffers from severe bias among attributes, e.g. blue and spots are highly correlated to coastal and long leg respectively. In this task, the goal is to learn an attribute classifier on “seen” animal categories (e.g “spots” attribute from the animal category “dalmatian”) and evaluate the model’s generalizability on unseen animal categories (e.g. “spots” attribute on the unseen animal category “leopard”). From Table 4, we observe that ours-feature-split offers gains on the exclusive test split over other methods without hurting the co-occurring case. In particular, we outperform attribute decorrelation , which was specifically designed to decorrelate attributes.

Conclusion

We demonstrated the problem of contextual bias in popular object and attribute datasets by showing that standard classifiers perform poorly when biased categories occur away from their typical context. To tackle this issue, we proposed two simple yet effective methods to decorrelate feature representations of a biased category from its context. Both methods perform better at recognizing biased classes occurring away from their co-occurring context while maintaining the overall performance. More importantly, our methods generalize to new unseen datasets and perform significantly better than standard methods. Our current framework tackles contextual bias between pairs of categories; future efforts should leverage more available (scene or category) information and model relationships between them. Extending proposed methods to tasks like object detection and video action recognition is a worthy future direction.

This work was supported in part by NSF CAREER IIS-1751206.

References

Appendix

Additional implementation Details

More results

Comparison with split biased: Results in Table 5 shows that ours-feature-split outperforms split biased with a significant margin on COCO-Stuff (28.828.8 vs. 19.119.1). Also, ours-CAM gives much better performance than split biased (26.426.4 vs. 19.119.1). Given that split biased cannot take full advantage of the co-occurring images (and vice-versa), it has inferior performance compared to both our methods.

Performance on non-biased classes: In Table 6, we show the mAP of our approach and standard classifier on the non-biased object classes (6060 classes) and on the entire COCO-Stuff dataset (object + stuff, 171 classes). We can see that our approach very marginally ( 0.02%~{}0.02\%) reduces the performance on non-biased object and stuff classes, while improving performance when biased categories occur away from their context.

Per class mAP and co-occurrence bias for 20 biased classes: In the Table 10, we show per class results for the COCO-Stuff for the top 20 biased classes. We also show the co-occurrence bias value for each class computed according to Eq. 1 in the main paper. From these results, we may observe that when a category occurs out of its context ours-feature-split gives better performance compared to standard classifier while maintaining the performance when a category co-occurs with context. ours-CAM performs better than standard when a category occurs away from its context, but struggles when categories co-occur.

Ablation study of ours-feature-split by varying fraction of biased category images: Here, we study the performance of our method as we vary the fraction of training images with biased categories occurring away from their typical context for COCO-Stuff. Specifically, for each of the 2020 biased categories in COCO-Stuff, we fix the total number of training images and vary the fraction of exclusive images. From Fig. 10, we note that standard performs rather poorly at lower fractions compared to both approaches (ours-CAM and ours-feature-split). Thus, both proposed methods achieve higher boosts at a fraction of 0.050.05 compared to 0.250.25. We also observe that a higher fraction of exclusive images benefits all the approaches, yet, our methods consistently outperform standard. This indicates that our approaches are more robust than the baseline especially on heavily skewed training data.

2 Comparison with other baselines for attribute classification

Table 8 reports performance on DeepFashion . We outperform all baselines by a significant margin on the exclusive test set. Although remove co-occur labels has slightly higher performance when attributes co-occur (20.420.4 vs. 20.120.1), ours-feature-split performs significantly better when attributes occur exclusively (6.06.0 vs. 9.29.2).

From Table 9, we observe that ours-feature-split offers gains on the exclusive test split compared to most methods for Animals with Attributes dataset. Though remove co-occur images yields higher gains on the exclusive test split, unlike ours-feature-split, it severely hurts the performance of co-occurring cases. Meanwhile ours-feature-split achieves good gains in exclusive cases without hurting co-occurring cases.

Finally, in Table 11 and 12, we show per category performance for the top 2020 biased categories for two datasets: DeepFashion and Animals with Attributes. These results show that ours-feature-split gives better performance than the standard classifier when attributes occur exclusively without their co-occurring context. At the same time, ours-feature-split maintains performance when biased attribute categories appear with co-occurring context.