Can I Trust Your Answer? Visually Grounded Video Question Answering

Junbin Xiao, Angela Yao, Yicong Li, Tat Seng Chua

Introduction

Video Question Answering (VideoQA) has recently emerged as a golden testbed to develop vision-language models (VLMs), especially foundation VLMs pretrained at scale on multi-modal web corpora . Despite significant advancements in QA performance, a fundamental concern arises – whether or to what extent the answers of such techniques are truly grounded on the relevant visual content, versus relying on the language bias for the use of powerful language models or spurious VL correlation captured via cross-modal correlation pretraining .

For example, in Fig. 1(Top), we find that existing VLMs incline to answer the questions with language-biased predictions, e.g., “unwrap (Q1: present)” and “tear (Q2: paper)”. Fig. 1(Bottom) shows that the overall predictions of SoTA VLMs overlap the predictions of standalone language models (BlindQA), i.e., models without visual inputs, by 62.5%. In fact, the BlindQA counterpart contributes to 66% of the correct predictions but also brings 79% of the wrong predictions in SoTA VLMs. The overlap gets much heavier by injecting a coarse visual signal from a single frame into the language model; because our later analysis shows that, most of the time, this frame lies outside the key moments where the correct answers fall.

Given these findings, a natural question arises – to what extent are the predictions of current VLMs grounded on the video content, and more precisely on the relevant parts? To answer this, we propose to study visually grounded VideoQA. Grounded VQA requires VLMs to answer the questions and simultaneously output the relevant video moments to support the answers. Earlier works have explored grounded QA under full supervision , but we target visual explanability in VideoQA and thus define the task under weak-supervision, which is the first of its kind.

To accomplish the goal, we construct NExT-GQA (short for GroundedQA) by extending NExT-QA with 10.5KK temporal labels (start and end timestamps) tied to the questions in the QA pairs for the validation and test sets. The labels are manually annotated and checked to be key to comprehend the questions and drive the correct answers. With this benchmark, we then examine a series of recent high-performing VLMs, including task-specific architectures without pretraining and pretrained models with either image-text or video-text data . Our findings reveal that all these VLMs struggle to predict visually grounded answers, despite their strong QA performance. For example, the SoTA model achieves 69% QA accuracy, but only 16% of the correctly predicted answers are grounded in the video. In contrast, humans can ground 82% out of the 93% of the correctly answered questions. Such clear discrepancy thus underscores the need for continued research efforts.

As a pioneering solution, we further propose a temporal grounding approach which can be easily applied to existing VLMs for visually grounded VideoQA. Specifically, our approach learns differentiable Gaussian masks along the temporal dimension of the videos, by optimizing light-weight transformer layers under both QA and self-supervision, without the need for temporal labels. Experiments with different QA backbones demonstrate that our approach effectively improves video grounding and question answering. The improvement is specially significant on a subset of questions that necessitate video understanding or temporal grounding.

To summarize our contributions: 1) we conduct the first study of weakly grounded VideoQA, and release a related benchmark NExT-GQA, to facilitate the research on more trustable VLMs; 2) we comprehensively analyze a wide range of advanced VLMs and reveal their limitation in performing visually grounded QA; 3) we propose a simple yet effective grounding mechanism which not only enhance existing VLMs in visual grounding, but also contributes to new SoTA QA performance, e.g., 73.1% on NExT-QA test set.

Related Work

Benchmarks. Supervised grounded VQA has been studied in both images and videos . Recently, weakly-supervised grounding has received increasing attention in ImageQA and temporal sentence grounding . Nonetheless, to our best knowledge, there is no work for weakly-grounded VideoQA. Also, existing supervised benchmarks are either biased towards localizing the subtitles in TV shows (e.g. TVQA ) or feature few objects (e.g. VidSTG ). Thus they are unsuitable to study visual evidence grounding.

Techniques. Strong VideoQA methods are predominantly banked on transformer and pre-training . The popular transformer architectures follow either shared , dual or stacked implementations, and pre-training is done with image-text , video-text or both forms of data. Notably, all these VLMs utilize powerful language models (e.g., BERT , T5 , GPT or their successors) for text encoding and focus on improving QA while ignoring visual evidence grounding. Recently, a handful of works have begun to ground key frames or objects to improve VideoQA. Yet still, they aim to improve QA accuracy, and do not evaluate the grounding.

For weakly-supervised video grounding, typical approaches extract temporal proposals and then rank the proposals according to their similarities with the language query . Despite the effectiveness, these two-stage approaches are notorious for inefficiency and sub-optimal for multi-granular temporal modeling. More recent research highlights the superiority of end-to-end Gaussian mask learning. Motivated by this, we design a simple yet effective Gaussian mask learning module for grounding in VideoQA. However, unlike these works to hand-craft negative visual proposals for contrastive learning, we optimize the Gaussian weights via question answering and cross-modal self-supervision.

NExT-GQA Dataset

Data Source. We choose NExT-QA as our data source to augment with temporal labels. Because most of the other VideoQA datasets feature short videos (3 ∼\sim 15s) already trimmed around the relevant content. NExT-QA has three different types of questions: Causal (“why/how”), Temporal (“before/when/after”) and Descriptive (“what/who/where”). We exclude the descriptive questions because they mostly pertain to global content (e.g., “what event?”) or answers can be found almost throughout the whole video (e.g., “where is?”). In addition, we only label the val and test sets since we aim for a weakly-supervised setup. As a result, \fnum11378QA pairs drawn from\fnum1570videos are to be annotated.

Label Collection. We invite undergraduate students for annotation. Each student is trained with our demo annotations together with some toy trial samples following specific criteria (as specified in Appendix A.1) before the actual annotation exercise. To guarantee the quality and reduce subjectiveness, each QA pair is annotated by at least two persons, and the final temporal label is determined by a further check and refinement of the two accepted annotations. The entire exercise lasted around 2 months and involved a team of 30 annotators. Eventually, we collect \fnum10531 valid temporal segments corresponding to \fnum8911 QA pairs and \fnum1557 videos. Detailed statistics are presented in Tab. 1.

Label Analysis. Fig. 2(a)(left) shows that most of the segments last for less than 15s, and with an average duration of 7s (Tab. 1) which is short compared to the video length (∼\sim40s). In fact, the ratio reflected in Fig. 2(a)(right) shows that most of the segments occupy less than half (0.5) length of the videos, and the average ratio is merely 0.2 (Tab. 1. This ratio is relative low, compared to temporal sentence grounding, e.g., 0.3 for both ActivityNet-Caption and Charades-STA . Moreover, Fig. 2(b)(left) shows that the segments are evenly distributed in the left, middle and right parts of the video. Fig. 2(b)(right) indicates that the vast majority of the QAs ground on a single temporal segment. To better understand the dataset, we show two examples in Fig. 3.

2 Comparison with Existing Benchmarks

We highlight the uniqueness of NExT-GQA by comparing it with other relevant benchmarks in Tab. 2.

NExT-GQA vs. NExT-QA. NExT-QA targets the prediction of textual answers. NExT-GQA differs in two major aspects: 1) it goes beyond that to provide visual evidence to support the answers, and 2) it extends the VQA setup by allowing visual answers. This satisfies more real-world applications, and additionally helps to better diagnose the models. For example, is a prediction wrong because the model failed to localize the relevant video contents, or because it could not convert the localized video contents into a textual answer? NExT-GQA is also more challenging, because: 1) the models need to achieve multiple goals (i.e., grounding and QA) and maintain their consistency, and 2) the questions are harder by factoring local video moments in untrimmed long videos. This also differs from major VideoQA benchmarks that focus on trimmed (short) video understanding .

NExT-GQA vs. TSG benchmarks. The benchmarks for temporal sentence grounding (TSG) aim to find a video moment described by a declarative sentence. NExT-GQA shares core challenges, i.e., cross-modal correspondence learning and multi-granular temporal modeling, while featuring some unique aspects. First, the questions feature unmentioned visual content to ground, such as “a baby falls and cries.” vs. “why did the baby cry?”. Thus, to answer the questions, the models not only need to find the described video moments (e.g., “baby cries”), but also should be capable of refining the moment to enclose the answer (e.g., “baby falls”). This may ask for temporal and causal relation reasoning. Second, the video backgrounds are relatively monotonous and accordingly the temporal segments pertaining to QA pairs are often more fine-grained than that in TSG benchmarks. Notably, NExT-GQA prioritizes finding visual evidence to support the answers. This means that any individual frame or moment which sufficiently tells the answer should be considered as a valid grounding, as opposed to retrieving all of the video contents that match with the query. This is reflected in our chosen of intersection over prediction (IoP) as one of the evaluation criteria. That is, a correct grounding depends on whether the predicted segment is falling into the labeled segment but not necessarily an exact match.

NExT-GQA vs. Supervised benchmarks. Fully-supervised benchmarks provide temporal annotations for training data, to resolve the reference ambiguities in the questions or improve QA performance with well-localized visual inputs. NExT-GQA differs from them by seeking to identify visual evidence that explains the answers with QA supervision alone. It is worth mentioning that directly applying the fully-supervised benchmarks for weakly grounding does not suit our goal, because these benchmarks are either biased to text localization or the answers are a limited set of, e.g. 80 objects . Additionally, we focus on weakly-supervised temporal grounding and leave spatio-temporal grounding for future exploration. Our consideration is that fine-grained spatio-temporal grounding is currently more challenging than question-answering, especially in the weak supervision setting , which could derail the main goal of VQA models.

Weakly-Supervised Grounding in VideoQA

VideoQA We first give an overview of the typical approaches to VideoQA, focusing on transformer-based methods for their superior performance. Given a video vv and a question qq , the goal of VideoQA is to predict a correct answer a∗a^{*} from a set of candidate answers AA. Depending on the task setting, AA can be given by multiple choices along with each question or by a global answer set . It is worth noting that SoTA transformer-methods formulate and solve both multi-choice QA and open-ended QA in an unified framework:

in which the mapping Ψ\Psi is typically realized as either a stacked (Fig. 4a) or dual (Fig. 4b) transformer. Thus, we primarily study the behaviors of these two styles of transformer architectures.

Weakly Grounded VideoQA Aside from answering questions, weakly-grounded VideoQA requires the models to explicitly estimate a QA-relevant video segment to serve as visual evidence. We introduce below three model-agnostic solutions to achieve this goal:

Post-hoc (PH). Intuitively, the temporal segment can be found through a post-hoc analysis of the temporal attentions, i.e., identifying the segment or frame with the maximal attention value and then thresholding around it to obtain the time interval. To accomplish this, we use attention-pooling to summarize the outputs from the temporal transformers for dual architectures. For stacked architectures, we directly return the averaged multi-head attention values corresponding to the prediction token.

Naive Gaussian (NG). The post-hoc approach is designed to analyze the models, but not influence their predictions. More favourably, we propose to explicitly incorporate a video grounding mechanism into VideoQA. We illustrate the framework in Fig. 5(a), and reformulate Eqn. 1 as

in which the grounding module Φ\Phi firstly estimates the key moment specified by tt and thereafter the QA module Ψ\Psi takes the more localized video content vtv_{t} for answer prediction. To enable end-to-end learning, tt is represented by differentiable Gaussian weights over the entire video sequence, i.e., t∼N(μ,σ2)t\sim N(\mu,\sigma^{2}), where μ\mu, σ∈\sigma\in are two learnable Gaussian parameters corresponding to the mean value and standard deviation. During inference, the grounding can be achieved by the confidence interval t=(μ−γσ,μ−γσ)∗dt=(\mu-\gamma\sigma,\mu-\gamma\sigma)*d, where γ\gamma is a hyper-parameter to control the width of the confidence interval and dd denotes the duration of the video.

Fig. 4c shows a dual transformer instantiation of this naive solution. The difference with respect to the original VideoQA counterpart (Fig. 4b) lies in a Gaussian mask prediction head, along with a Gaussian weighted token learning and aggregation stage. We find that this approach effectively learns and outputs grounding information. Nevertheless, the improvements over a post-hoc solution are limited due to the weak QA supervision.

NG+. In light of the naive Gaussian results, we further design an auxiliary objective via cross-modal self-supervision, to regulate the VQA objective towards more visually grounded QA. Specifically, for each question q+q^{+}, we treat the corresponding grounding hypothesis vtv_{t} as an anchor point and pull close its distance with q+q^{+} while pushing away its distance with other questions Q−Q^{-} in the feature space. The negative set Q−Q^{-} includes: 1) The other questions defined in the same video as hard negatives. Because different questions often invoke different video moments for answers; 2) The questions sampled from other videos to ensure the sufficiency and diversity of negative samples. Moreover, we enrich 10% of the positive questions by rephrasing each question (using GPT-4 ) with maximal 5 additional questions to form Q+Q^{+}. Thus, our final solution is:

where Q=Q+∪Q−Q=Q^{+}\cup Q^{-} which comprises both the positive and negative questions concerning to vtv_{t}. α\alpha is a trade-off parameter. Note that the Grounding-term coarsely identifies the question-relevant video moment tt, while the GroundedQA-term not only makes the prediction but also helps to refine the moment tt with answer supervision. The overall objective thus enforces the grounded video contents to be relevant to both the answers and the questions.

Experiments

Our experiments majorly answer three research questions: Q1: To which extent are the current VLMs’ predictions grounded on relevant video content? Q2: Does better QA performance imply better grounding and vice versa? Q3: How effective is our Gaussian masking mechanism? We study a wide variety of VLMs, covering different architectures (dual and stacked transformers), vision encoders (task-specific and pretrained with image- or video-text data), and text encoders (BERT, RoBERTa, DeBERTa, Flan-T5):

VGT is a task-specific, dual-style graph transformer model. It encodes spatio-temporal object information for VideoQA. We also investigate VGT with RoBERTa as suggested by .

Temp[Swin] is a dual architecture. The Swin Transformer (SWT) is pre-trained on ImageNet . Temp[CLIP] and Temp[BLIP] follow the same dual architecture, but use ViT pretrained by CLIP and BLIP respectively as vision encoders.

VIOLETv2 adopts a stacked transformer. It uses video Swin Transformer (VSWT) and BERT for vision and text encoding, respectively. The model is pretrained with both image- and video-text data, and achieves SoTA on various VL tasks.

FrozenBiLM applies a stacked transformer. It uses CLIP as vision encoder and highlights the strength of adapting frozen large language models (LLMs) (e.g., DeBERTa-V2-XL (1B) ) for VideoQA.

IGV and SeViLA are additionally reproduced for comparison. Both works explicitly learn to ground key frames for VideoQA. IGV is built on visual graph, whereas SeViLA is founded on BLIP-2 . It exploits ViT-G and frozen LLM (e.g., Flan-T5-XL (3B)) for video localization and QA. In our implementation, we choose the smallest time spans that can enclose the localized key frames as grounded moments.

Experimental Settings. For all models, we uniformly sample 32 frames from each video and freeze the vision encoders. In post-hoc analysis, the temporal attention thresholds are searched around the mean attention values, to maximize the grounded QA accuracy. The number of negative questions in Eqn. 3 is kept the same as the number of distractor answers in MCQA to facilitate joint optimization. The trade-off parameter α\alpha is set to 1 and 0.1 for dual and stacked transformers, respectively. During inference, the hyperparameter γ\gamma for the Gaussian confidence interval is set to 1 or 0.8 depending on different models. Our final results are reported based on a combination of predictions from Gaussian and temporal attention. All hyper-parameters are tuned on the validation set, and unless otherwise indicated, the results are reported on the test set. Other details are described in Appendix A.2.

Evaluation. We report accuracy for QA , which stands for the percentage of correctly answered questions. For visual evidence grounding, we use intersection over prediction (IoP) to measure whether the predicted temporal window lies inside the ground truth. Additionally, we include temporal IoU following TSG benchmarks. For both IoP and IoU, we report the mean values and values with overlap thresholds of 0.3 and 0.5. If a QA pair involves multiple temporal segments, we report the results based on the one with maximal overlap with the prediction. Notably, we define grounded QA accuracy (Acc@GQA) to inspect the percentages of questions that are correctly answered and also visually grounded (i.e., IoP ≥\geq 0.5).

2 Result and Analysis

We focus on Acc@QA, Acc@GQA and IoP@0.5 in the Post-hoc (PH) block of Tab. 3. Generally, the existing VLMs excel at QA but are weak in grounding the answers in the videos. For example, all the methods exceed 50% in QA accuracy, yet cannot reach more than 12-16% for grounded QA accuracy. In fact, the SoTA QA model (FrozenBiLM) achieves an accuracy of 69% for QA compared to a surprisingly low 16% for GQA. The results of IoP@0.5 suggest that the large disparity is mainly due to the models’ poor performance in temporal grounding and partly because of inconsistency between grounding and QA (Not all correct grounding yields correct answers according to Acc@GQA vs. IoP@0.5). We additionally exclude the influence of sparse video sampling by investigating the coverage of QA content w.r.t the number of sampled video frames in Fig. 6(a). The figure shows that the sampled 32 frames can cover almost all QA contents. Moreover, to understand the extent of such poor performance, we estimate the upper-bound performance through a human study on 10% of the test data. The study shows that participants correctly answered 93% of the questions, with 82% being also visually grounded.

Given the above observations, we believe that most of these models’ answers are not grounded on the relevant video content but more likely derived from the language shortcut or spurious correlation of irrelevant visual context.

To investigate the language shortcut, we conduct a BlindQA experiment, in which we train only the language counterparts of the VQA models without the video. Tab. 4(a) shows that BlindQA achieves 80% of the performance of standard VQA (NormalQA), i.e., 50.3% vs. 59.4% for the dual models and 56.7% vs. 69.1% for the stacked models. To study the spurious correlation, we test the VLMs by directly sampling inside (PosQA) or outside (NegQA) the ground-truth video segments. Surprisingly, the models’ QA performances almost remain unaffected compared with a normal uniform sampling (NormalQA) (likely because the image representations are not fine-grained enough to differentiate different frames.). Tab. 4(a) shows that providing the ground-truth temporal segments (PosQA) brings marginal improvement (<<1%) for the dual-style models and even hurts stacked-style transformers (likely due to the distribution shift of visual inputs). Furthermore, excluding the temporal segments (NegQA) only degenerates the performance by less than 1% for both dual and stacked-style models. The above studies reinforce our belief that the current VLM’s answer predictions are hardly grounded on relevant video contents, but instead relied on shortcuts of languages and a superficial vision-language correlation.

2.2 Q2: Better QA implies better grounding?

First, by focusing on Acc@QA, mIoP and mIoU in Tab. 3. we find that better QA is not necessarily established on better grounding, and the results vary among different architectures. For instance, by comparing across different architectures, FrozenBiLM shows the strongest QA performance, yet with surprisingly poor grounding, e.g., the IoP values are even worse than those of VGT which displays the lowest QA results among other transformer models. This could be due to FrozenBiLM’s freezing of the LLMs, causing its predictions to heavily rely on the common sense knowledge of the LLMs rather than the provided videos (similar problem is also found on SeViLA). In contrast, VGT is a task specific model. It focuses on exploiting the fine-grained video information, and thus conditions better on the visual content. By comparing among different instantiations of the same architectures (e.g., Temp[Swin] to Temp[CLIP]) as well as different training epochs of the same models in Fig. 6(b), we find that the grounding performance (mIoP) improves along with the increase of QA accuracy for dual-style architectures yet not for stacked-style ones. Second, regarding the influence of grounding on QA, our conclusion is that having grounding is better than no. Yet, this is not controlled and opts for the underlying short-cuts that the models convey and thus leaves for the models to learn on their own. The conclusion is backed by the observations that PosQA always outperforms NegQA in Tab. 4(a) regardless of model architectures. Moreover, our effort to improve grounding also brings better QA performance (Tab. 3 &\& 4(b)). However, as aforementioned, correct grounding does not guarantee correct answer predictions.

2.3 Q3: Is Gaussian masking solution effective?

We incorporate our Gaussian grounding mechanism (NG and NG+) into the top-performing dual- and stacked-style models and compare with Post-hoc baseline. Despite the weaker performance, we highlight the higher efficiency of dual-style implementation, especially in retrieval-based QA systems as exemplified by multi-choice QA. Tab. 3 shows that both NG and NG+ lead to better grounding and QA performance. Also, NG+ generally outperforms NG, especially for dual-style architectures. Additionally, Tab. 4(b) indicates that our superiority gets enlarged in answering the subset of questions that necessitate videos and temporal grounding.

For better understanding, we analyze two cases in Fig. 6(c). The top example shows that the Gaussian masks (NG and NG+) are more focused on the relevant video moment than temporal attention, thus bringing better grounding, especially for IoU. The bottom example highlights the strength of NG+. In this case, there are multiple visual instances that correspond to the answer “girl stands up”. The correct instance is the one after the “girl takes the green ball”, though the instance after “take the red ball” is more salient. Both the Post-hoc and Naive methods are distracted because they are learned via answer supervision alone. In contrast, NG+ finds the correct grounding since it also optimizes the cross-modal correspondence between questions and video segments. More detailed analyses are presented in Appendix A.3.

2.4 Method Comparison

Tab. 3 shows that compared with a random baseline, all methods effectively perform grounded QA (refer to Acc@GQA and IoP@0.5). More concretely, we find that both IGV and SeViLA obtain lower GQA accuracy than FrozenGQA though they incorporate a sense of grounding in their models. The weakness manifest in both visual evidence grounding (IoP@0.5) and QA. However, we find that SeViLA performs significantly better than other methods in grounding (mIoP and mIoU) alone. We attribute such strength to the fact that SeViLA is pretrained with localization supervisions.

2.5 Other Observations

Tab. 3 also compares the Acc@QA performance on NExT-GQA versus the full (original) NExT-QA test set. There is a consistent 2-3% higher accuracy on the full set, suggesting that the questions rooted in local video moments are harder to answer than those rely on overall video content. Besides, the cross-modal pretrained representations perform better than the uni-modal pretrained ones for both VQA and visual grounding. Also, the image-text pretrained representations outperform those pretrained with video-text data. Moreover, the dual-style architectures tend to have better grounding performance than stacked ones (Note that FrozenBiLM’s high-ranking GQA result is due to its strong QA performance instead of grounding). This is surprising, as there is no cross-modal interaction in dual-style implementations. we speculate that cross-modal transformers likely suffer from a uni-modal bias, which leads to the attention being skewed towards the language side for predicting textual answers. The findings on the one hand consolidate the benefits of harnessing foundation VLMs or LLMs for videoQA. On the other hand, they accentuate the need to coordinate and balance between visual and textual evidence.

Conclusion

We summarize the following points and raise them as open challenges for the rest of the community: First, current VLMs built on powerful language models excel in answering visual questions. Yet, their predictions often lack a strong connection to the pertinent visual information but instead heavily rely short-cut of languages and irrelevant visual context. This calls for more efforts towards the interpretability and trustability. Second, while temporal sentence grounding has progressed a lot, our experiments show that localizing the questions, especially those featuring temporal actions and events in weak supervision, is much more challenging. Our studies indicate that solving this problem would largely benefit visually-grounded VideoQA. Third, although our solution improves grounding as well as QA, there is still large gap compared to human performance. This leaves ample opportunity for follow-up works. Last but not least, we highlight the significance of NExT-GQA and hope it can contribute towards the advancement in these areas.

Limitations. Our analyses are focused on multi-choice QA. Furthermore, the additional grounding objective in our solution demands more memory and time to train.

This research is supported by NUS NExT++. This research is also supported by the National Research Foundation, Singapore under its NRF Fellowship for AI (NRF-NRFFAI1-2019-0001). Any opinions, findings and conclusions or recommendations expressed in this material are those of the author(s) and do not reflect the views of National Research Foundation, Singapore.

References

Appendix A Appendix

Our annotation is facilitated by Elan . A screenshot is shown in Fig. 7.

Criteria. We set clear criteria to limit the ambiguities and subjectiveness. 1) For each question, the annotation should encompass the entire temporal segment which features the answer and also sufficient context to interpret the question. 2) If the visual content mentioned in the question is not simultaneous or contiguous in time with the answer, then the annotation should focus on the answer. 3) If visual evidence for an answer appears multiple times in the video, then all relevant video moments (individual segment) should be annotated. 4) If the answer of the question can be seen throughout the entire video, the question is omitted. Yet, to ensure we can collect sufficient labels, we pay annotators on a per annotated segment basis.

A.2 Implementation Details

Post-hoc. For dual-style transformer, we have tried both attention-pooling the visual tokens as well as prepending a summary token and then averaging the multi-head attention of transformer. We find that the two methods bring similar QA performance. Yet, the prepending approach demands much more training epochs and thus we chose attention-pooling as our final solution. Moreover, to obtain a reasonable time span from the learned temporal attention, we treat the frame of maximal attention value as the pivot location, and search around it to enclose the frames whose attention values satisfying certain criteria. Before that, the attention values are normalized to using min-max method. The criteria of whether a frame should be enclosed are jointly determined by its attention score and its distance with the pivot frame. In our implementation, we also smooth the attention values and the distance threshold is set to 10s. Finally, the minimal frame id and the maximal frame id are mapped to the time seconds to obtain the temporal span. It is worth mentioning that the frame of maximal attention value will always be selected.

Naive Gaussian (NG). For both dual- and stacked-style architectures, the Gaussian prediction head is implemented with a lightweight transformer layer followed by linear projectors. Notably, as there is no independent visual stream in stacked-style transformer, we pick the tokens belonging to the visual inputs and go through the Gaussian-weighted transformer. The resultant tokens are then preprended back into the multi-modal token sequence for answer prediction.

Video-Question Correspondence Learning (NG+). We find that a two-stage training paradigm to pretrain with the Grounding-term and then finetune with both objectives in Eqn. 3 of the main text brings better performance than one-stage training. In both stage, the negative questions are selected from the same videos as the positive question at a chance of 0.3. Note that we exclude the descriptive questions because their answers usually appear throughout the video. Also at a chance of 0.3, we replace the positive question with a rephrased one. During generation, we prompt GPT-4 to focus on the nouns and actions in the questions, so as to ensure the generated questions share the same video moment with the original one. We show in Fig. 8 some examples of the generated questions.

Others. We train all models by 10∼\sim20 epochs with an initial learning rate 1e-5 and earlier stopping is adopted if the results on the validation set do not increase in 5 epochs. The batch size is set to 64 for dual-style models and 4∼\sim6 for stacked-style models respectively. Unless otherwise specified, our major experiments are conducted on 4 A5000 GPUs with each 24G memories.

A.3 Additional Experiments

We take Temp[CLIP] with Naive Gaussian (NG) grounding approach to study the effect of using different number of Gaussian masks. The results in Tab. 5 show that using multiple Gaussian masks will hurt the QA accuracy though it increases the grounding performance according to IoU value. The best grounded QA (Acc@GQA) result is achieved by using 5 Gaussian masks. Nonetheless, the improvement over a single Gaussian mask is negligible, e.g., 0.3% from 15.5% to 15.8%. Therefore, we by default use a single Gaussian mask in major experiments. This also brings higher efficiency.

A.3.2 Does the generated questions help?

We do additional studies on the use of extended positive questions in the NG+ method. As shown in Tab. 6, we find that it can slightly improves the QA results (Acc@QA) but not for grounded QA (Acc@GQA). In terms of grounding, it bring slightly higher IoU result yet lower IoP compared with the models without using the generated questions.

A.3.3 Model Efficiency

We discuss the efficiency of Temp[CLIP] and FrozenBiLM in the visually-grounded QA task. For Temp[CLIP], all results are obtained with 1 A5000 GPU. For FrozenBiLM without NG+, the experiment was conducted on 4 A5000 GPUs; for FronzenBilM with NG+, we run with 4 R8000 GPUs as the model needs about 46G per GPU memory. The time is reported based on 1 epoch over the training and validation data respectively. The results in Tab. 7 show that our grounding module introduces little additional parameters for training and inference compared with the respective backbone models. Yet, the NG+ method takes more time to train. Another observation is that the Temp[CLIP] has much higher training and inference speed than FrozenBiLM.

A.3.4 Result Visualization

We show some prediction cases in Fig. 9. Both models predict the correct answer with reasonable visual grounding results for Q1 and Q2. From the 3rd question, we show that the models suffer a lot in either correctly answering the questions (e.g., Q5, Q6 FrozenGQA and Q8) or providing the right visual evidence for the correct answers (e.g., Q3 FrozenGQA, Q4 and Q7). From the failure examples, we find that when the the visual concepts in the answers present throughout the videos (e.g. “grass” and “snow” in Q4 and Q7 respectively), the models can easily predict the correct answers without the need to truly localizing the questioned video segments. Furthermore, the models are still weak in 1) answering the questions which involve small visual objects and 2) substantiating the answers when the visual evidence only takes small portion of the videos (Q4 ∼\sim Q8).