Invariant Grounding for Video Question Answering

Yicong Li, Xiang Wang, Junbin Xiao, Wei Ji, Tat-Seng Chua

Introduction

Video Question Answering (VideoQA) is growing in popularity and importance to interactive AI, such as vision-language navigation for in-home robots and personal assistants . It is the task of multi-modal reasoning, which answers the natural language question about the content of a given video. Clearly, inferring a reliable answer requires a deep understanding of visual scenes, linguistic semantics, and more importantly, the visual-linguistic alignments.

Towards this end, a number of VideoQA models have emerged . Scrutinizing these models, we summarize their common paradigm as a combination of two modules: (1) video-question encoder, which encapsulates the visual scenes of video and the linguistic semantics of question as representations; and (2) answer decoder, which exploits these representations to model the visual-linguistic alignment and yield an answer. Consequently, the criterion of empirical risk minimization (ERM) is widely adopted as the learning objective to optimize these modules — that is, minimizing the loss between the predictive answer and the ground-truth answer.

However, the ERM criterion is prone to over-exploiting the superficial correlations between video-question pairs and answers. Specifically, we use the metric of local mutual information (LMI) to quantify the correlations between the “track” scene and answers. As Figure 1(a) shows, most videos with “track” scene are associated with the “race” answer. Instead of inspecting the visual-linguistic alignments (i.e. which scene is critical to answer the question), ERM blindly captures all statistical relations. As Figure 1(b) shows, it makes VideoQA model naively link the “track”-relevant videos with the strongly-correlated “race” answer, instead of the gold “jump” answer. Taking a causal look at VideoQA (see Section 3), we partition the visual scenes into two parts: (1) causal scene, which holds the question-critical information, and (2) its complement, which is irrelevant to the answer. We scrutinize that the complement is spuriously correlated with the answer, thus ERM hardly differentiates the effects of causal and complement scenes on the answer. Worse still, the unsatisfactory reasoning obstacles the VideoQA model to own the intriguing properties:

Visual-explainability to exhibit “Which visual scene are the right reasons for the right answering?” . Taking Figure 1(b) as an example to answer “What is the man doing?”, the model should attend the “jump” event present in the last three clips, rather than referring to the “track” complement in the first two clips. One straightforward solution is “learning to attend” to ground some scenes via the attentive mechanism. Nonetheless, guided by ERM, such attentive grounding still suffers from the spurious correlations, thus making the highly-correlated complement grounded.

Introspective learning to double-check “How would the predictive answer change if the causal scenes were absent?”. On top of attentive grounding, the model needs to introspect whether the learned knowledge (i.e. attended scene) reliably and faithfully reflects the logic behind the answering. Briefly put, it should fail to answer the question if the causal scenes were removed.

Generalization ability to enquire “How would the predictive answer response to the change of spurious correlations?”. As spurious correlations poorly generalize to open-world scenarios, the model should instead latch on the causal visual-linguistic relations that are stable across different environments.

Inspired by recent invariant learning , we conjecture that invariant grounding is the key to distinguishing causal scenes from the complements and overcoming these limitations. By “invariant”, we mean that the relations between question-critical scenes and answers are invariant regardless of changes in complements. Towards this end, we propose a new learning framework, Invariant Grounding for VideoQA (IGV). Concretely, it integrates two additional modules with into the VideoQA backbone model: a grounding indicator, a scene intervener. Specifically, the grounding indicator learns to attend the causal scenes for a given question and leaves the rest as the complement. Then, we collect visual clips from other training videos to compose a memory bank of complement stratification. For the causal part of interest, the scene intervener conducts the causal interventions on its complement — that is, replace it with the stratification sampled from the memory bank and compose the “intervened videos”. After pairing the casual, complement, and intervened scenes with the question, we feed them into the backbone model to obtain the corresponding predictions: (1) causal prediction, which approaches the gold answer, so as to achieve visual explainability; (2) complement prediction, which contains no critical clues to the ground-truth answer, thus enforces the backbone model to perform introspective reasoning; and (3) intervened prediction, which is consistent with the causal prediction across different intervened complements. Jointly learning these predictions enables the backbone model to alleviate the negative influence of multi-modal data bias. It is worthwhile emphasizing that IGV is a model-agnostic strategy, which trains the VideoQA backbones in a plug-and-play fashion.

Our contributions are summarized as follows:

We highlight the importance of grounding causal scenes from the complements to visual-explainability, generalization, and introspective learning of VideoQA models.

We propose a new model-agnostic training scheme, IGV, which incorporates invariant grounding into the VideoQA models, to mitigate the negative influence of multi-modal data bias and enhance the multi-modal reasoning ability.

On three benchmark datasets (i.e. MSRVTT-QA , MSVD-QA , NExT-QA ), we conduct extensive experiments to justify the superiority of IGV in training the VideoQA backbones. In particular, IGV significantly outperforms the state-of-the-art models.

Preliminaries

In this section, we summarize the common paradigm of VideoQA models. Throughout the paper, we denote the random variables and their deterministic values by upper-cased (e.g. VV) and lower-cased (e.g. vv) letters, respectively.

Modeling. Given the video-question pair (V,Q)(V,Q), the primer task of VideoQA is to generate an answer A^\hat{A} as:

where fA^f_{\hat{A}} is the VideoQA model, which is typically composed of two modules: video-question encoder, and answer decoder. Specifically, the encoder includes two components: (1) a video encoder, which encodes visual scenes of the target video as a visual representation, such as motion-appearance memory design , structural graph representation , hierarchical architecture ; and (2) a question encoder, which encapsulates linguistic semantics of the question into a linguistic representation, such as global/local representation of textual content , graph representation of grammatical dependencies . On top of these representations, the decoder learns the visual-linguistic alignments to generate the answer. In particular, the alignments are modeled via cross-modal interaction like graph alignment, cross-attention and co-memory , etc.

Learning. To optimize these modules, most of the leading VideoQA models cast the multi-modal reasoning problem as a supervised learning task and adopt the learning objective of empirical risk minimization (ERM) as:

where LERM\mathbf{\mathop{\mathcal{L}}}_{\text{ERM}} is the risk function to measure the loss between the predictive answer A^\hat{A} and ground-truth answer AA, which is usually set as cross-entropy loss or hinge loss . In essence, ERM encourages these VideoQA modules to capture the statistical correlations between the video-question pairs and answers.

Causal Look at VideoQA

From the perspective of causal theory , we revisit the VideoQA scenario to show superficial correlations between video-question pairs and answers. We then analyze ERM’s suffering from the spurious correlations.

In general, multiple visual scenes are present in a video. But only part of the scenes are critical to answering the question of interest, while the rest hardly offers information relevant to the question. Moreover, the linguistic variations in different questions should activate different scenes of a video. These facts inspire us to split the video into the causal and complement parts in terms of the question. Here we use a causal graph to exhibit the relationships among five variables: input video VV, input question QQ, causal scene CC, complement scene TT, ground-truth answer AA. Figure 2 illustrates the causal graph, where each link is a cause-and-effect relationship between two variables:

C←V→TC\leftarrow V\to T. The input video VV consists of CC and TT. For example, the video in Figure 1(b) is the combination of the first two clips (i.e. CC) and the last three clips (i.e. TT).

V→C←QV\to C\leftarrow Q. The causal scene CC is conditional upon the video-question pair (V,Q)(V,Q), which distills QQ-relevant information from VV. For a given VV, the variations in QQ result in different CC.

Q→A←CQ\to A\leftarrow C. The answer AA is determined by the question QQ and causal scene CC, reflecting the visual-linguistic alignments. Considering the example in Figure 1(b) again, CC is the oracle scene that perfectly explains why “jump” is labeled as the ground truth to answer the question.

T⇠⇢CT\dashleftarrow\dashrightarrow C. The dashed arrow summarizes the additional probabilistic dependencies between CC and TT. Such dependencies are usually caused by the selection bias or inductive bias during the process of data collection or annotation . For example, one mostly collects the videos with the “jump” events on the “track”. Here we list three typical scenarios: (1) CC is independent of TT (i.e. T⊥CT\bot C); (2) CC is the direct cause of TT (i.e. C→TC\rightarrow T), or vise versa (i.e. C←TC\leftarrow T); (3) CC and TT have a common cause EE (i.e. C←E→TC\leftarrow E\rightarrow T). See Appendix A for details.

2 Spurious Correlations

Taking a closer look at the causal graph, we find that the complement scene TT and the ground-truth answer AA can be spuriously correlated. Specifically, as the confounder between TT and AA, QQ and VV open the backdoor paths: T←V→C→AT\leftarrow V\to C\to A and T←V→C←Q→AT\leftarrow V\to C\leftarrow Q\to A, which make TT and AA spuriously correlated even though there is no direct causal path from TT to AA. Worse still, T⇠⇢CT\dashleftarrow\dashrightarrow C can amplify this issue. Assuming C→TC\to T, CC becomes an additional confounder to yield another backdoor path T←C→AT\leftarrow C\to A. Such spurious correlations can be summarized as the probabilistic dependence: A̸ ⁣⊥TA\not\!\bot T.

As ERM naively captures the statistical correlations between video-question pairs and answers, it fails to distinguish the causal scene CC and complement scene TT, thus failing to mitigate the negative influence of spurious correlations. As a result, it limits the reasoning ability of VideoQA models, especially in the following aspects: (1) visual-explainability to reason about “Which visual scenes are the supporting evidence to answer the question?”; (2) introspective learning to answer “How would the answer change if the causal scenes were absent?”; and (3) generalization ability to enquire “How would the answer response to the change of spurious correlations?”.

Methodology

We get inspiration from invariant learning and argue that invariant grounding of causal scenes is the key to reducing the spurious correlations and overcoming the foregoing limitations. We then present a new learning framework, Invariant Grounding for VideoQA (IGV).

Upon closer inspection on the causal graph, we notice that the ground-truth answer AA is independent of the visual complement TT, only when conditioned on the question QQ and the causal scene CC, more formally:

This probabilistic independence indicates the invariance — that is, the relations between the (C,Q)(C,Q) pair and AA are invariant regardless of changes in TT. The causal relationship Q→A←CQ\rightarrow A\leftarrow C is invariant across different TT. Taking Figure 1(b) as an example, if the question and the causal scene (i.e. the last three clips) remain unchanged, the answer should arrive at “jump”, no matter how the complement variesNote that the complement substitutes will not involve the question-relevant scenes, in order to avoid creating additional paths from TT to AA. (e.g. substitute the “track” clips by the “cloud”- or “sea”-relevant ones). This highlights that the (C,Q)(C,Q) pair is the key to shielding AA from the influence of TT.

Modeling. However, only the (V,Q)(V,Q) pair and AA are available in the training set, while neither CC nor the grounding function towards CC is known. This motivates us to incorporate visual grounding into the VideoQA modeling, where the grounded scene C^\hat{C} aims to estimate the oracle CC and guide the prediction of answer A^\hat{A}. More formally, instead of the conventional modeling (cf. Equation (1)), we systematize the modeling process as:

where fC^f_{\hat{C}} is the grounding model, and fA^f_{\hat{A}} is the VideoQA model that relies on the (C^,Q)(\hat{C},Q) pair instead. See Section 4.2 for our implementations of fC^f_{\hat{C}} and fA^f_{\hat{A}}.

Learning. Nonetheless, simply integrating visual grounding with the VideoQA model falls into the “learning to attend” paradigm, which still suffers from the spurious correlations and erroneously attends to the complement scenes as C^\hat{C}. To this end, we exploit the invariance property of CC (cf. Equation (3)) and reformulate the learning objective of invariant grounding as:

where LIGV\mathbf{\mathop{\mathcal{L}}}_{\text{IGV}} is the loss function to our IGV; T^=V∖C^\hat{T}=V\setminus\hat{C} is the complement of C^\hat{C}. In the next section, we will elaborate how to implement LIGV\mathbf{\mathop{\mathcal{L}}}_{\text{IGV}} and achieve invariant grounding.

2 IGV Framework

Figure 3 displays our IGV framework, which involves two additional modules, the grounding indicator and scene intervener, beyond the VideoQA backbone model fA^f_{\hat{A}}.

For a video-question pair instance (v,q)(v,q), at the core of the grounding indicator is to split the video instance vv into two parts, c^\hat{c} and t^\hat{t}, according to the question qq. Towards this end, it first employs two independent LSTMs to encode the visual and linguistic characteristics of vv and qq, respectively:

where I0kI_{0k} and I1kI_{1k} suggests that the kk-th clip belongs to the causal and complement scenes, respectively.

2.2 Scene Intervener

It is challenging to learn the grounding indicator, owing to the lack of supervisory signals of clip-level importance. To remedy this issue, we propose the scene intervener, which preserves the estimated causal scene c^\hat{c} but intervenes the estimated complement t^\hat{t} to create the “intervened videos”, as Figure 4 shows.

Specifically, for the observed video-question pairs during training, the scene intervener first collects visual clips from other training videos as a memory bank of complement stratification, T^={t^}\hat{\mathcal{T}}=\{\hat{t}\}. Then, for the video of interest v=c^∪t^v=\hat{c}\cup\hat{t}, the intervener conducts causal interventions on its t^\hat{t} — that is, random sample a complement stratification t^∗∈T^\hat{t}^{*}\in\hat{\mathcal{T}} to replace t^\hat{t} and combine it with c^\hat{c} at hand as a new video v∗=c^∪t^∗v^{*}=\hat{c}\cup\hat{t}^{*}.

It is worthwhile mentioning that, distinct from the current invariant learning studies that only partition the training set into different environments, our scene intervener exploits the interventional distributions instead. The interventional distribution (i.e., the videos with the same interventions) can be viewed as one environment.

2.3 VideoQA Backbone Model

Inspired by , we design a simple yet effective architecture as our backbone predictor, where the video encoder is shared with the grounding indicator. It embodies convolutional graph networks (GCN) to propagate clip-level visual messages, then integrates cross-modal fused local and global representations via BLOCK fusion . See Appendix B for the detailed architecture.

2.4 Joint Training

For a video-question pair instance (v,q)(v,q), we have established the causal scene c^\hat{c}, complement scene t^\hat{t}, and intervened video v∗v^{*} via the grounding indicator and scene intervener. Pairing them with qq synthesizes three new instances: (c^,q)(\hat{c},q), (t^,q)(\hat{t},q), (v∗,q)(v^{*},q). We next feed these instances into the backbone VideoQA model fA^f_{\hat{A}} to obtain three predictions:

Causal prediction. As the causal scene c^\hat{c} is expected to be sufficient and necessary to answer the question qq, we leverage its predictive answer fA^(c^,q)f_{\hat{A}}(\hat{c},q) to approach the ground-truth answer aa solely:

Complement prediction. As no critical clues should exist in the complement scene t^\hat{t} to answer the question qq, we encourage its predictive answer fA^(t^,q)f_{\hat{A}}(\hat{t},q) to evenly predict all answers. This uniform loss is formulated as:

where KL denotes KL-divergence, and uu is the uniform distribution over all answer candidates.

Intervened prediction. According to the invariant constraint (cf. Equation (3)), the causal relationship between the causal scene and the answer is stable across different complements. To parameterize this constraint, we enforce all vv’s intervened versions to hold the consistent predictions:

Aggregating the foregoing risks, we attain the learning objective of IGV:

where O+\mathcal{O}^{+} is the training set of the video-question pair (v,q)(v,q) and the ground-truth answer aa; λ1\lambda_{1} and λ2\lambda_{2} are the hyper-parameters to control the strengths of invariant learning. Jointly learning these predictions enables the VideoQA backbone model to uncover the question-critical scene, so as to mitigate the negative influence of spurious correlations between the question-irrelevant complement scene and answer. In the inference phase, we use the causal prediction fA^(c^,q)f_{\hat{A}}(\hat{c},q) to answer the question.

Experiments

We conduct extensive experiments to answer the following research questions:

RQ1: How effect is IGV in training VideoQA backbones as compared with the State-of-the-Art (SoTA) models?

RQ2: How do the loss component and feature setting affect the performance?

RQ3: What are the learning patterns and insights of IGV training?

Settings: We compare IGV with seven baselines from families of Memory, GNN and Hierarchy (Appendix C) on three VideoQA datasets: NExT-QA which features causal and temporal action interactions among multiple objects. It contains about 47.7K manually annotated questions for multi-choice QA collected from 5.4K videos with an average length of 44s. MSVD-QA and MSRVTT-QA are two prevailing datasets that focus on the description of video elements. They respectively contain 50K and 243K QA pairs with open answer space over 1.6K and 6K. For all three datasets, we follow their official data splits for experiments and report accuracy as evaluation metric.

Implementation Details: For the visual feature, we follow previous works and extract video feature as a combination of motion and appearance representations by using the pre-trained 3D ResNeXt-101 and ResNet-101, respectively. Specifically, each video is uniformly sampled into KK=16 clips, where each clip is represented by a combined feature vector vkdvv_{k}^{d_{v}}, where dvd_{v} equals 4096. Similar to , we obtain the contextualized word representation from the finetuned BERT model, and the feature dim dqd_{q} is 768. For our model, the dimension of the hidden states are set to dd = 512, and the number of graph layers in IGV backbone predictor is 2. During training, IGV is optimized by Adam optimizer with the initial learning rate of 1e-4, which will be halved if no validation improvements in 5 epochs. We set the batch size to 256 and a maximum of 60 epochs. (See Appendix D for more details and complexity analysis).

As shown in Table 1 and Table 2, our method outperforms SoTAs with questions of all sub-types surpassing their competitors. Specifically, we have two major observations:

First, on NExT-QA, IGV gains remarkable improvement on temporal type (+1.65%), the underlying explanation are: 1) temporal question generally corresponds to video content with a longer time span, which requires more introspective grounding of the causal scene. Fortunately, IGV’s design philosophy comfort such demand by wiping out the trivial scenes, which takes up a huge proportion in temporal type, thus making the predicting faithful. 2) temporal questions tend to include a temporal indicative phase (e.g. ”at the end of the video”) that serves as a strong signal for grounding indicator to locate the target window.

Second, along with descriptive questions on NExT-QA, the result on MSRVTT-QA and MSVD-QA (both emphases on question of descriptive type) demonstrate the superiority in descriptive question across all three datasets (+1.85% on NExT-QA, +1.4% on MSRVTT-QA and MSVD-QA). Such improvement is underpinned by logic that answering descriptive questions requires scrutiny on the scene of interest, instead of a holistic view of the entire sequence. Accordingly, targeted prediction inducted by IGV concentrates reasoning on keyframes, thus achieving better performance. As a consequence, such improvement strongly validates that IGV generalizes better over various environments.

1.2 Backbone Agnostic

By nature, our IGV principle is orthogonal to backbone design, thus helping to boost any off-the-shelf SoTAs without compromising the underlying architecture. We therefore experimentally testify the generality and effectiveness of our learning strategy by marrying the IVG principle with methods from two different categories: Co-Mem from memory-based architecture and HGA from Graph-based method. Table 3 shows the results on three backbone predictors (including ours). Our findings are:

1. Better improvement for severe bias. We notice that the improvement on MSVD-QA (+3.1%∼\sim4.7%) is considerably larger than that on MSRVTT-QA (+1.4%∼\sim2%). Such expected discrepancy is caused by the fact that, although identical in question type, MSRVTT-QA is almost 5 times larger than MSVD-QA (#QA pairs 243K vs 50K). As a result, the baseline model trained on MSRVTT-QA is gifted with better generalization ability, whereas the model on MSVD-QA still suffers from severe shortcut correlation. For the same reason, the IGV framework achieves much better improvement in the severe-shortcut situation (e.g. MSVD-QA). Such discrepancy validates our motivation of eliminating statistic dependency.

2. Constant improvement for each method. Through row-wise inspection, we notice that for each benchmark, IGV can bring considerable improvement across different backbone models (+3.1%∼\sim4.7% for MSVD-QA, +1.4%∼\sim2% for MSRVTT-QA). Such stable enhancement strongly verifies our modal-agnostic statement.

2 In-Depth Study (RQ2)

An in-depth comprehension of IGV framework requires careful scrutiny on its components. Alone this line, we exhaust the combination of IGV loss components and design three variants: Lc^,\mathbf{\mathop{\mathcal{L}}}_{\hat{c}}, Lc^+Lt^\mathbf{\mathop{\mathcal{L}}}_{\hat{c}}+\mathbf{\mathop{\mathcal{L}}}_{\hat{t}} and Lc^+Lv∗\mathbf{\mathop{\mathcal{L}}}_{\hat{c}}+\mathbf{\mathop{\mathcal{L}}}_{v^{*}}. Table 4 shows the result of the above variants on two benchmarks across two backbone predictors. Our observations are as follow:

Using Lc^\mathbf{\mathop{\mathcal{L}}}_{\hat{c}} solely, which can be viewed as a special case of ERM-guided attention, hardly outperforms the baseline, because grounding indicators can not identify the causal scene without clip-level supervision. Such an expected result reflects our motivation in interventional design.

Lc^+Lt^\mathbf{\mathop{\mathcal{L}}}_{\hat{c}}+\mathbf{\mathop{\mathcal{L}}}_{\hat{t}} and Lc^+Lv∗\mathbf{\mathop{\mathcal{L}}}_{\hat{c}}+\mathbf{\mathop{\mathcal{L}}}_{v^{*}} matched equally in accuracy that consistently surpass baseline and Lc^\mathbf{\mathop{\mathcal{L}}}_{\hat{c}} in all cases. Such progress shows the effectiveness of intervention strategy and introspective regularization imposed on complement.

In all cases, Lc^+Lt^+Lv∗\mathbf{\mathop{\mathcal{L}}}_{\hat{c}}+\mathbf{\mathop{\mathcal{L}}}_{\hat{t}}+\mathbf{\mathop{\mathcal{L}}}_{v^{*}} further boosts the performance significantly, which shows Lt^\mathbf{\mathop{\mathcal{L}}}_{\hat{t}} and Lv∗\mathbf{\mathop{\mathcal{L}}}_{v^{*}} contribute in different aspects and their benefits are mutually reinforcing.

2.2 Study of Feature

By convention, we study the effect of the input condition by ablation on the visual feature. Particularly, we denote APP for tests that adopt only appearance feature as input and MOT for tests that utilize motion feature alone. Figure 6(a) delivers results on two benchmarks, where we observe:

First, IGV can improve the performance significantly for all input conditions, which generalizes the effectiveness of our framework. Similar to Table 3, the improvement on MSVD-QA is larger than that on MSRVTT-QA, which solidifies our finding in Section 5.1.2.

Second, compared to motion feature, IGV brings distinctively larger improvements using appearance feature. Considering the causal nature of IGV, we conclude that static correlation tends to bias more in appearance feature.

2.3 Study of Hyper-parameter

For MSVD-QA, we observe consistent peaks around 0.8 for both λ1\lambda_{1} and λ2\lambda_{2}. Comparatively, fluctuation on MSRVTT-QA is more moderate, where tuning on λ\lambda only causes a 1.5% difference in their accuracy. It’s noteworthy that IGV outperforms the baseline by a large margin (+3%) under all tests, which indicates IGV’s robustness against variation of hyper-parameters. Additionally, comparing to λ2\lambda_{2}, IGV is more sensitive to λ1\lambda_{1}. Typically, the performance suffers a drastic degradation for λ1\lambda_{1} larger than 5 on both datasets. Whereas λ2\lambda_{2} maintain above 39% (MSVD-QA) and 37.5% (MSRVTT-QA) for all tests.

3 Qualitative analysis (RQ3)

As mentioned in Section 1 , IGV is empowered with visual-explainability, and is apt to account for the right scene for its prediction. Following this essence, we grasp the learning insight of IGV by inspecting some correct examples from the NExT-QA dataset and show the visualization in Figure 5. Concretely, each video comes with two questions that emphasize different parts of the video. We notice that, even for the same video, our grounding window is question-sensitive to enclose the explainable content with correct prediction. Nonetheless, we also observe results of insufficient-grounding on the third row Q2, where the girl starts to bend down before the last two frames, even though the most informative last two frames are encompassed.

Related works

Video Question Answering (VideoQA). Aiming to answer the question in a video scenario, VideoQA is defined as an escalation of imageQA, because the temporal nature of the input has enriched its reasoning process as well as the answer space. Previous efforts towards VideoQA establish their contribution on either a better multi-modal interaction or stronger video representation. Specifically, early studies tend to impose sophisticated cross-modal fusion via attention or dynamic memory , while more recent approaches perform relation reasoning through visual or textual graph . In addition, current efforts that model video as a hierarchical structure also intrigue wide interest. Among them, HCRN stack conditional relation blocks in different feature granularity, whereas HOSTR employs a spatio-temporal graph for multilevel reasoning. Despite their effectiveness, their visual-explainability still dwells on ERM-guided attention weights, which only reflect the intensity of feature-prediction correlation.

Invariant Learning. Multi-modal datasets tend to display inherent bias in some forms . In contrast to overarching reality, the collection process degrades its generalization ability by introducing undesirable correlations between the inputs and the ground truth annotations.

To overcome such correlation, invariant learning is developed to discover causal relations from the causal factors to the response variable, which remains constant across distributions. As the most prevailing formulation, IRM promotes this philosophy from feature level to representation level by finding a data representation Υ\Upsilon, from which the optimal predictor φ\varphi can yield the prediction Υ∘φ\Upsilon\circ\varphi that is stable across all environments. In terms of environment acquisition, previous studies either manually partition the training set by prior knowledge , or generates data partition iteratively via adversarial environment inference . Our method, instead of partitioning the training, assumes no prophets about environments but performs causal intervention to perturb the original distribution. To the best of our knowledge, IGV is the first work that introduces invariant learning as a model-agnostic framework to VideoQA.

Conclusions

In this paper, we pinpoint that the spurious visual-linguistic correlations in VideoQA are triggered by question-irrelevant scenes. We propose a novel invariant grounding framework, IGV, to distinguish the causal scene and emphasize its causal effect on the answer. With the grounding indicator and scene intervener, IGV captures the causal patterns that remain stable across complements. Extensive experiments verify the effectiveness of IGV on different backbone VideoQA models.

Our future work includes two aspects: 1) the spurious correlations can nest in entities, object-level invariant learning is promising to alleviate this issue; 2) as the current intervention strategy might threaten the causal prediction by introducing complement with new shortcuts, we will explore new intervention methods.

References

Appendix A Example of context type

As shown in Figure 7, we classify the relation between causal scene and its complement (e.g. T⇠⇢CT\dashleftarrow\dashrightarrow C) into three types, where each row encompasses a causal graph (left) that depicts typical causal-complement relation demonstrated in the example (right):

In the first row, CC and TT has no causal relation (i.e. T⊥CT\bot C).

The second row shows a scenario that CC is the direct cause of TT (i.e. C→TC\rightarrow T), or vise versa if the question is modified (e.g.’What is the cat doing?’)

Similar to the example in Figure 1, the third row demonstrates how shortcut deviate the prediction from the gold answer (e.g. ”talk”) to false prediction (e.g. ”cook”) via common cause EE (e.g. visual concept ”kitchen”) since LMI between visual concept ”kitchen” and candidate answer ”cook” is much higher than it is with ”talk”.

Appendix B Our backbone

Most VideoQA architectures from the state of the art are compatible with our IGV learning strategy. To testify, we design a simple and effective architecture inspired by . Specifically, fA^f_{\hat{A}} is presented as a combination of a visual-question mixer and an answer classifier. The mixer first encode c^\hat{c}:

Similarly, we obtain the final representation by applying the BLOCK again to global and local factor, which is further decoded into answer space with classifier Ψ\Psi:

Analogously, we can obtain the predictive answer for t^\hat{t} and v∗v^{*} via the shared backbone predictor.

Appendix C Baselines

We compare our design against some existing work, which can be categorized into three categories: 1) Memory-based methods that perform multi-step reasoning via updating the recurrent unit, which refines the cross-modal representation iteratively. Specifically, AMU , Co-Mem apply this module to encode the visual representation, and HME managed better exploitation for both modalities; 2) Graph-based methods like HGA and B2A adopt graph reasoning on the clip-level, whose adjacent matrix is built on node-wise visual similarity. Comparatively, B2A additionally establishes a text graph through question parsing, and abridge two modalities via message passing; 3) Hierarchical-based methods HOSTR and HCRN have similar hierarchical conditional architectures. Their discrepancy lies in the feature granularity, where HCRN grounds the temporal relation between frames, while HOSTR roots in object trajectories.

Appendix D Implementation details

All experiments are conducted on GPU NVIDIA Tesla V100 installed on Ubuntu 18.0.4. In terms of complexity, our algorithm matched equally with the corresponding baseline. As a comparison, the default backbone model is trained for 2 hours till convergence on MSRVTT-QA, whereas IGV takes 2.6 hours. For space complexity, since we use the same predictor for the causal, complement, and intervened prediction, IGV only takes 10% more parameters than the default backbone model.