Visual Commonsense R-CNN

Tan Wang, Jianqiang Huang, Hanwang Zhang, Qianru Sun

Introduction

“On the contrary, Watson, you can see everything. You fail, however, to reason from what you see.”

–Sherlock Holmes, The Adventure of the Blue Carbuncle

Today’s computer vision systems are good at telling us “what” (e.g., classification , segmentation ) and “where” (e.g., detection , tracking ), yet bad at knowing “why”, e.g., why is it dog? Note that the “why” here does not merely mean by asking for visual reasons — attributes like furry and four-legged — that are already well-addressed by machines; beyond, it also means by asking for high-level commonsense reasons — such as dog barks — that are still elusive, even for us human philosophers , not to mention for machines.

It is not hard to spot the “cognitive errors” committed by machines due to the lack of common sense. As shown in Figure 1, by using only the visual features, e.g., the prevailing Faster R-CNN based Up-Down , machine usually fails to describe the exact visual relationships (the captioning example), or, even if the prediction is correct, the underlying visual attention is not reasonable (the VQA example). Previous works blame this for dataset bias without further justification , e.g., the large concept co-occurrence gap in Figure 1; but here we take a closer look at it by appreciating the difference between the “visual” and “commonsense” features. As the “visual” only tells “what”/“where” about person or leg per se, it is just a more descriptive symbol than its correspondent English word; when there is bias, e.g., there are more person than leg regions co-occur with the word “ski”, the visual attention is thus more likely to focus on the person region. On the other hand, if we could use the “commonsense” features, the action of “ski” can focuses on the leg region because of the common sense: we ski with legs.

We are certainly not the first to believe that visual features should include more commonsense knowledge, rather than just visual appearances. There is a trend in our community towards weakly-supervised learning features from large-scale vision-language corpus . However, despite the major challenge in trading off between annotation cost and noisy multimodal pairs, common sense is not always recorded in text due to the reporting bias , e.g., most may say “people walking on road” but few will point out “people walking with legs”. In fact, we humans naturally learn common sense in an unsupervised fashion by exploring the physical world, and we wish that machines can also imitate in this way.

A successful example is the unsupervised learning of word vectors in our sister NLP community : a word representation XX is learned by predicting its contextual word YY, i.e., P(Y∣X)P(Y|X) in a neighborhood window. However, its counterpart in our own community, such as learning by predicting surrounding objects or parts , is far from effective in down-stream tasks. The reason is that the commonsense knowledge, in the form of language sentences, has already been recorded in discourse; in contrast, once an image has been taken, the explicit knowledge why objects are contextualized will never be observed, so the true common sense that causes the existence of objects XX and YY might be confounded by the spurious observational bias, e.g., if keyboard and mouse are more often observed with table than any other objects, the underlying common sense that keyboard and mouse are parts of computer will be wrongly attributed to table.

Intrigued, we perform a toy MS-COCO experiment with ground-truth object labels — by using a mental apparatus, intervention, that makes us human — to screen out the existence of confounders and then eliminate their effect. We compare the difference between association P(Y∣X)P(Y|X) and causal intervention P(Y∣do(X))P(Y|do(X)) . Before we formally introduce do in Section 3.1, you can intuitively understand it as the following deliberate experiment illustrated in Figure 2: 1) “borrow” objects ZZ from other images, 2) “put” them around XX and YY, then 3) test if XX still causes the existence of YY given ZZ. The “borrow” and “put” is the spirit of intervention, implying that the chance of ZZ is only dependent on us (probably subject to a prior), but independent on XX or YY. By doing so, as shown in Figure 3, P(sink∣do(dryer))P(\texttt{sink}|do(\texttt{dryer})) is lower because the most common restroom context such as towel is forced to be seen as fair as others. Therefore, by using P(Y∣do(X))P(Y|do(X)) as the learning objective, the bias from the context will be alleviated.

More intrigued, P(person∣do(toilet))P(\texttt{person}|do(\texttt{toilet})) is higher. Indeed, person and toilet co-occur rarely due to privacy. However, human’s seeing is fundamentally different from machine’s because our instinct is to seek the causality behind any association — and here comes the common sense. As opposed to the passive observation P(Y∣X)P(Y|X): “How likely I see person if I see toilet”, we keep asking “Why does seeing toilet eventually cause seeing person?” by using P(Y∣do(X))P(Y|do(X)). Thanks to intervention, we can increase P(Y∣do(X))P(Y|do(X)) by “borrowing” non-local context that might not be even in this image, for the example in Figure 2, objects usable by person such as chair and handbag — though less common in the restroom context — will be still fairly “borrowed” and “put” in the image together with the common sink. We will revisit this example formally in Section 3.1.

So far, we are ready to present our unsupervised region feature learning method: Visual Commonsense R-CNN (VC R-CNN), as illustrated in Figure 4, which uses Region-based Convolutional Neural Network (R-CNN) as the visual backbone, and the causal intervention as the training objective. Besides its novel learning fashion, we also design a novel algorithm for the dodo-operation, which is an effective approximation for the imaginative intervention (cf. Section 3.2). The delivery of VC R-CNN is a region feature extractor for any region proposal, and thus it is fundamental and ready-to-use for many high-level vision tasks such as Image Captioning , VQA , and VCR . Through extensive experiments in Section 5, VC R-CNN shows significant and consistent improvements over strong baselines — the prevailing methods in each task. Unlike the recent “Bert-like” methods that require huge GPU computing resource for pre-training features and fine-tuning tasks, VC R-CNN is light and non-intrusive. By “light”, we mean that it is just as fast and memory-efficient as Faster R-CNN ; by “non-intrusive”, we mean that re-writing the task network is not needed, all you need is numpy.concatenate and then ready to roll.

We apologize humbly to disclaim that VC R-CNN provides a philosophically correct definition of “visual common sense”. We only attempt to step towards a computational definition in two intuitive folds: 1) common: unsupervised learning from the observed objects, and 2) sense-making: pursuing the causalities hidden in the observed objects. VC R-CNN not only re-thinks the conventional likelihood-based learning in our CV community, but also provides a promising direction — causal inference — via practical experiments.

Related Work

Multimodal Feature Learning. With the recent success of pre-training language models (LM) in NLP, several approaches seek weakly-supervised learning from large, unlabelled multi-modal data to encode visual-semantic knowledge. However, all these methods suffer from the reporting bias of language and the great memory cost for downstream fine-tuning. In contrast, our VC R-CNN is unsupervised learning only from images and the learned feature can be simply concatenated to the original representations.

Un-/Self-supervised Visual Feature Learning . They aim to learn visual features through an elaborated proxy task such as denoising autoencoders , context & rotation prediction and data augmentation . The context prediction is learned from correlation while image rotation and augmentation can be regarded as applying the random controlled trial , which is active and non-observational (physical); by contrast, our VC R-CNN learns from the observational causal inference that is passive and observational (imaginative).

Visual Common Sense. Previous methods mainly fall into two folds: 1) learning from images with commonsense knowledge bases and 2) learning actions from videos . However, the first one limits the common sense to the human-annotated knowledge, while the latter is essentially, again, learning from correlation.

Causality in Vision. There has been a growing amount of efforts in marrying complementary strengths of deep learning and causal reasoning and have been explored in several contexts, including image classification , reinforcement learning and adversarial learning . Lately, we are aware of some contemporary works on visual causality such as visual dialog , image captioning and scene graph generation . Different from their task-specific causal inference, VC R-CNN offers a generic feature extractor.

Sense-making by Intervention

We detail the core technical contribution in VC R-CNN: causal intervention and its implementation.

As shown in Figure 5 (left), our visual world exists many confounders z∈Zz\in Z that affects (or causes) either XX or YY, leading to spurious correlations by only learning from the likelihood P(Y∣X)P(Y|X). To see this, by using Bayes rule:

where the confounder ZZ introduces the observational bias via P(z∣X)P(z|X). For example, as recorded in Figure 6, when PP(zz=sink∣|XX=toilet) is large while PP(zz=chair∣|XX=toilet) is small, most of the likelihood sum in Eq. (1) will be credited to PP(YY=person∣|XX=toilet,zz=sink), other than PP(YY=person∣|XX=toilet,zz=chair), so, the prediction from toilet to person will be eventually focused on sink rather than toilet itself, e.g., the learned features of a region toilet are merely its surrounding sink-like features.

As illustrated in Figure 5 (right), if we intervene XX, e.g., dodo(XX=toilet), the causal link between ZZ and XX is cut-off. By applying the Bayes rule on the new graph, we have:

Compared to Eq. (1), zz is no longer affected by XX, and thus the intervention deliberately forces XX to incorporate every zz fairly, subject to its prior P(z)P(z), into the prediction of YY. Figure 6 shows the gap between the prior P(z)P(z) and P(z∣toilet)P(z|\texttt{toilet}), z∈Zz\in Z is the set of MS-COCO labels. We can use this figure to clearly explain the two interesting key results by performing intervention. Please note that P(Y∣X,z)P(Y|X,z) remains the same in both Eq. (1) and Eq. (2),

Please recall Figure 3 for the sensible difference between P(Y∣X)P(Y|X) and P(Y∣do(X))P(Y|do(X)). First, PP(person∣|do(toilet))>>PP(person∣|toilet) is probably because the number of classes zz such that PP(z∣z|toilet)>>P(z)P(z) is smaller than those such that of PP(z∣z|toilet) << P(z)P(z), i.e., the left grey area is smaller than the right grey area in Figure 6, making Eq. (1) smaller than Eq. (2). Second, we can see that zz making P(z)<P(z∣X)P(z)<P(z|X) is mainly from the common restroom context such as sink, bottle, and toothbrush. Therefore, by using intervention P(Y∣do(X))P(Y|do(X)) as the feature learning objective, we can adjust between “common” and “sense-making”, thus alleviate the observational bias.

Figure 7LABEL:sub@Fig:R1 visualizes the features extracted from MS-COCO images by using the proposed VC R-CNN. Promisingly, compared to P(Y∣X)P(Y|X) (left), P(Y∣do(X))P(Y|do(X)) (right) successfully discovers some sensible common sense. For example, before intervention, window and leg features in red box are close due to the street view observational bias, e.g., people walking on street with window buildings; after intervention, they are clearly separated. Interestingly, VC R-CNN leg features are closer to head while window features are closer to wall. Furthermore, Figure 7LABEL:sub@Fig:R2 shows the features of ski, snow and leg on same MS-COCO images via Up-Down (left) and our VC R-CNN (right). We can see the ski feature of our VC R-CNN is reasonably closer to leg and snow than Up-Down. Interestingly, VC R-CNN merges into sub-clusters (dashed boxes), implying that the common sense is actually multi-facet and varies from context to context.

X\bm{X} →\rightarrow Y\bm{Y} or Y\bm{Y} →\rightarrow X\bm{X}? We want to further clarify that both two causal directions between XX and YY can be meaningful and indispensable with do calculus. For XX →\rightarrow YY, we want to learn the visual commonsense about XX (e.g., toilet) that causes the existence of YY (e.g., person), and vice versa. Only objects are confounders? No, some confounders are unobserved and beyond objects in visual commonsense learning, e.g., color, attributes, and the nuanced scene contexts induced by them; however, in unsupervised learning, we can only exploit the objects. Fortunately, this is reasonable: 1) we can consider the objects as the partially observed children of the unobserved confounder ; 2) we propose the implementation below to approximate the contexts, e.g., in Figure 8, Stop sign may be the child of the confounder “transportation”, and Toaster and Refrigerator may contribute to “kitchen”.

2 The Proposed Implementation

To implement the theoretical and imaginative intervention in Eq. (2), we propose the proxy task of predicting the local context labels of YY’s RoI. For the confounder set ZZ, since we can hardly collect all confounders in real world, we approximate it to a fixed confounder dictionary Z=[z1,...,zN]\bm{Z}=[\bm{z}_{1},...,\bm{z}_{N}] in the shape of N×dN\times d matrix for practical use, where NN is the category size in dataset (e.g., 80 in MS-COCO) and dd is the feature dimension of RoI. Each entry zi\bm{z}_{i} is the averaged RoI feature of the ii-th category samples in dataset. The feature is pre-trained by Faster R-CNN.

Normalized Weighted Geometric Mean (NWGM). We apply NWGM to approximate the above expectation. In a nutshell, NWGMThe detailed derivation about NWGM can be found in the Supp.. effeciently moves the outer expectation into the Softmax as:

Note that the above approximation is reasonable, because the effect on YY comes from both XX and confounder ZZ (cf. the right Figure 5).

Neural Causation Coefficient (NCC). Due to the fact that the causality from the confounders as the category averaged features are not yet verified, that is, Z\bm{Z} may contain colliders (or v-structure) causing spurious correlations when intervention. To this end, we apply NCC to remove possible colliders from Z\bm{Z}. Given x\bm{x} and z\bm{z}, NCC(x→z)NCC\left(\bm{x}\rightarrow\bm{z}\right) outputs the relative causality intensity from x\bm{x} to z\bm{z}. Then we discard the training samples with strong collider causal intensities above a threshold.

VC R-CNN

Architecture. Figure 4 illustrates the VC R-CNN architecture. VC R-CNN takes an image as input and generates feature map from a CNN backbone (e.g., ResNet101 ). Then, unlike Faster R-CNN , we discard the Region Proposal Network (RPN). The ground-truth bounding boxes are directly utilized to extract the object level representation with the RoIAlign layer. Finally, each two RoI features x\bm{x} and y\bm{y} eventually branch into two sibling predictors: Self Predictor with a fully connected layer to estimate each object class, while Context Predictor with the approximated do-calculus in Eq. (3) to predict the context label.

Training Objectives. The Self-Predictor outputs a discrete probability distribution p=(p,...,p[N])p=(p,...,p[N]) over NN categories (note that we do not have the “background” class). The loss can be defined as Lself(p,xc)=−log(p[xc])L_{self}(p,x^{c})=-log(p[x^{c}]), where xcx^{c} is the ground-truth class of RoI XX. The Context Predictor loss LcxtL_{cxt} is defined for each two RoI feature vectors. Considering XX as the center object while YiY_{i} is one of the KK context objects with ground-truth label yicy_{i}^{c}, the loss is Lcxt(pi,yic)=−log(pi[yic])L_{cxt}(p_{i},y_{i}^{c})=-log(p_{i}[y_{i}^{c}]), where pip_{i} is calculated by pi=P(Yi∣do(X))p_{i}=P(Y_{i}|do(X)) in Eq. (3) and pi=(pi,...,pi[N])p_{i}=(p_{i},...,p_{i}[N]) is the probability over NN categories. Finally, the overall mulit-task loss for each RoI XX is:

Feature Extractor. We consider VC R-CNN as a visual commonsense feature extractor for any region proposal. Then the extracted features are directly concatenated to the original visual feature utilized in any downstream tasks. It is worth noting that we do NOT recommend early concatenations for some models that contain a self-attention architecture such as AoANet . The reasons are two-fold. First, as the computation of these models are expensive, early concatenation significantly slows down the training. Second, which is more crucial, the self-attention essentially and implicitly applies P(Y∣X)P(Y|X), which contradicts to causal intervention. We will detail this finding in Section 5.4.

Experiments

We used the two following datasets for unsupervised learning VC R-CNN.

MS-COCO Detection . It is a popular benchmark dataset for classification, detection and segmentation in our community. It contains 82,783, 40,504 and 40,775 images for training, validation and testing respectively with 80 annotated classes. Since there are 5K images from downstream image captioning task which can be also found in MS-COCO validation split, we removed those in training. Moreover, recall that our VC R-CNN relies on the context prediction task, thus, we discarded images with only one annotated bounding box.

Open Images . We also used a much larger dataset called Open Images, a huge collection containing 16M bounding boxes across 1.9M images, making it the largest object detection dataset. We chose images with more than three annotations from the official training set, results in about 1.07 million images consisting of 500 classes.

2 Implementation Details

We trained our VC R-CNN on 4 Nvidia 1080Ti GPUs with a total batch size of 8 images for 220K iterations (each mini-batch has 2 images per GPU). The learning rate was set to 0.0005 which was decreased by 10 at 160K and 200K iteration. ResNet-101 was set to the image feature extraction backbone. We used SGD as the optimizer with weight decay of 0.0001 and momentum of 0.9 following . To construct the confounder dictionary Z\bm{Z}, we first employed the pre-trained official ResNet-101 model on Faster R-CNN with ground-truth boxes as the input to extract the RoI features for each object. For training on Open Images, we first trained a vanilla Faster R-CNN model. Then Z\bm{Z} is built by making average on RoIs of the same class and is fixed during the whole training stage.

3 Comparative Designs

To evaluate the effectiveness of our VC R-CNN feature (VC), we present three representative vision-and-language downstream tasks in our experiment. For each task, a classic model and a state-of-the-art model were both performed for comprehensive comparisons. For each method, we used the following five ablative feature settings: 1) Obj: the features based on Faster R-CNN, we adopted the popular used bottom-up feature ; 2) Only VC: pure VC features; 3) +Det: the features from training R-CNN with single self detection branch without Context Predictor. “+” denotes the extracted features are concatenated with the original feature, e.g., bottom-up feature; 4) +Cor: the features from training R-CNN by predicting all context labels (i.e., correlation) without the intervention; 5) +VC: our full feature with the proposed implemented intervention, concatenated to the original feature. For fair comparisons, we retained all the settings and random seeds in the downstream task models. Moreover, since some downstream models may have different settings in the original papers, we also quoted their results for clear comparison. For each downstream task, we detail the problem settings, dataset and evaluation metrics as below.

Image Captioning. Image captioning aims to generate textual description of an image. We trained and evaluated on the most popular “Karpathy” split built on MS-COCO dataset, where 5K images for validation, 5K for testing, and the rest for training. The sentences were tokenized and changed to lowercase. Words appearing less than 5 times were removed and each caption was trimmed to a maximum of 16 words. Five standard metrics were applied for evaluating the performances of the testing models: CIDEr-D , BLEU , METROT , ROUGE and SPICE .

Visual Question Answering (VQA). The VQA task requires answering natural language questions according to the images. We evaluated the VQA model on VQA2.0 . Compared with VQA1.0 , VQA2.0 has more question-image pairs for training (443,757) and validation (214,354), and all the question-answer pairs are balanced. Before training, we performed standard text pre-processing. Questions were trimed to a maximum of 14 words and candidate answer set was restricted to answers appearing more than 8 times. The evaluation metrics consist of three pre-type accuracies (i.e., “Yes/No”, “Number” and “Other”).

Visual Commonsense Reasoning (VCR). In VCR, given a challenging question about an image, machines need to present two sub-tasks: answer correctly (Q→\rightarrowA) and provide a rationale justifying its answer (QA→\rightarrowR). The VCR dataset contains over 212K (training), 26K (validation) and 25K (testing) derived from 110K movie scenes. The model was evaluated in terms of 4-choice accuracy and the random guess accuracy on each sub-task is 25%.

4 Results and Analysis

Results on Image Captioning. We compared our VC representation with ablative features on two representative approaches: Up-Down and AoANet . For Up-Down model shown in Table 1, we can observe that with our +VC trained on MS-COCO, the model can even outperform current SOTA method AoANet over most of the metrics. However, only utilizing the pure VC feature (i.e., Only VC) would hurt the model performance. The reason can be obvious. Even for human it is insufficient to merely know the common sense that “apple is edible” for specific tasks, we also need visual features containing objects and attributes (e.g., “what color is the apple”) which are encoded by previous representations. When comparing +VC with the +Det and +Cor without intervention, results also show absolute gains over all metrics, which demonstrates the effectiveness of our proposed causal intervention in representation learning. AoANet proposed an “Attention on Attention” module on feature encoder and caption decoder for refining with the self-attention mechanism. In our experiment, we discarded the AoA refining encoder (i.e., AoANet†) rather than using full AoANet since the self-attentive operation on feature can be viewed as an indiscriminate correlation against our do-expression. From Table 1 we can observe that our +VC with AoANet† achieves a new SOTA performance. We also evaluated our feature on the online COCO test server in Table 2. We can find our model also achieves the best single-model scores across all metrics outperforming previous methods significantly.

Moreover, since the existing metrics fall short to the dataset bias, we also applied a new metric CHAIR to measure the object hallucination (e.g., “hallucinate” objects not in image). The lower is better. As shown in Table 3, we can see that our VC feature performs the best on both standard and CHAIR metrics, thanks to our proposed intervention that can encode the visual commonsense knowledge.

Results on VQA. In Table 4, we applied our VC feature on classical Up-Down and recent state-of-the-art method MCAN . From the results, our proposed +VC outperforms all the other ablative representations on three answer types, achieving the state-of-the-art performance. However, compared to the image captioning, the gains on VQA with our VC feature are less significant. The potential reason lies in the limited ability of the current question understanding, which cannot be resolved by “visual” common sense. Table 5 reports the single model performance of various models on both test-dev and test-standard sets. Although our VC feature is limited by the question understanding, we still receive the absolute gains by just feature concatenation compared to previous methods with complicated module stack, which only achieves a slight improvement.

Results on VCR. We present two representative methods R2C and ViLBERT in this emerging task on the validation set. Note that as the R2C applies the ResNet backbone for residual feature extraction, here for fair comparison we switched it to the uniform bottom-up features. Moreover, for ViLBERT, since our VC features were not involved in the pretraining process on Conceptual Captions, here we utilized the ViLBERT† rather than the full ViLBERT model. From the comparison with ablative visual representations in Table 6, our +VC feature still shows the superior performances similar to the above two tasks.

Results on Open Images. To evaluate the transfer ability and flexibility of the learned visual commonsense feature, we also performed our proposed VC R-CNN on a large image detection collection. The results can be referred to Table 1&4&6. We can see that the performances are extremely close to the VC feature trained on MS-COCO, indicating the stability of our learned semantically meaningful representation. Moreover, while performing VCR with the dataset of movie clip, which has quite diverse distributions compared to the captioning and VQA built on MS-COCO, our VC R-CNN trained on Open Images achieves the reasonable better results.

5 Qualitative Analysis

We visualize several examples with our VC feature and previous Up-Down feature for each task in Figure 9. Any other settings except for feature kept the same. We can observe that with our VC, models can choose more precise, reasonable attention area and explicable better performance.

6 Ablation Study

Conclusions

We presented a novel unsupervised feature representation learning method called VC R-CNN that can be based on any R-CNN framework, supporting a variety of high-level tasks by using only feature concatenation. The key novelty of VC R-CNN is that the learning objective is based on causal intervention, which is fundamentally different from the conventional likelihood. Extensive experiments on benchmarks showed impressive performance boosts on almost all the strong baselines and metrics. In future, we intend to study the potential of our VC R-CNN applied in other modalities such as video and 3D point cloud.

Acknowledgments We would like to thank all reviewers for their constructive comments. This work was partially supported by the NTU-Alibaba JRI and the Singapore Ministry of Education (MOE) Academic Research Fund (AcRF) Tier 1 grant.

References