A Simple LLM Framework for Long-Range Video Question-Answering

Ce Zhang, Taixi Lu, Md Mohaiminul Islam, Ziyang Wang, Shoubin Yu, Mohit Bansal, Gedas Bertasius

Introduction

Recent years have witnessed remarkable progress in short video understanding (5-15s in length) . However, extending these models to long videos (e.g., several minutes or hours in length) is not trivial due to the need for sophisticated long-range temporal reasoning capabilities. Most existing long-range video models rely on costly and complex long-range temporal modeling schemes, which include memory queues , long-range feature banks , space-time graphs , state-space layers and other complex long-range modeling modules .

Recently, Large Language Models (LLMs) have shown impressive capability for long-range reasoning on a wide range of tasks such as document understanding and long-horizon planning . Motivated by these results in the natural language and decision-making domain, we explore using LLMs for long-range video question answering (LVQA). Specifically, we propose LLoVi, a simple language-based framework for long-range video understanding. Unlike prior long-range video models, our approach does not require specialized long-range video modules (e.g., memory queues, state-space layers, spatiotemporal graphs, etc.) but instead uses a short-term visual captioner coupled with an LLM, thus exploiting the long-range temporal reasoning ability of LLMs. Our simple two-stage approach tackles the LVQA task by decomposing it into short and long-range modeling subproblems:

First, given a long video input, we segment it into multiple short clips and convert them into one-sentence textual descriptions using a pre-trained frame/clip-level visual captioner (e.g., BLIP2, LaViLa).

Afterward, we concatenate the temporally ordered captions from Step 1 and feed them into an LLM (e.g., GPT-3.5, GPT-4) with proper prompts to perform long-range video reasoning for LVQA.

Recent works have explored incorporating visual captioning models into LLMs for video understanding. Several methods proposed using LLMs for various short-term video understanding tasks. For example, Video ChatCaptioner generates enriched video descriptions by learning to answer questions from ChatGPT. Similarly, ChatVideo stores the detected tracklets in a database, which are then used for interacting with the user. Additionally, Socratic Models , VidIL , and VideoChat all use pretrained visual models to extract low-level video concepts (e.g., objects, actions, scenes, etc.) and then leverage LLM to perform various short-term video understanding tasks. Beyond short-term video understanding tasks, we note that the concurrent work in applies an LLM-based framework for long-range video understanding. However, their analysis is limited to movie-based datasets that rely heavily on non-visual inputs such as speech and subtitles, thus requiring limited visual analysis . Compared to this concurrent work, we conduct our experiments on EgoSchema, the newly introduced LVQA benchmark that addresses many limitations of prior long-range video benchmarks. Furthermore, unlike prior LLM-based video understanding frameworks, many of which perform only qualitative analysis of their models , we conduct a thorough empirical study investigating the effectiveness of our framework.

Specifically, we investigate (i) the selection of a visual captioner, (ii) the choice of an LLM, (iii) the LLM prompt design, (iv) few-shot in-context learning, (v) optimal video processing configurations (i.e., clip length, sampling rate, etc.), and (vi) the generalization of our framework to other datasets and tasks. Our key empirical findings include:

A multi-round prompt that first asks the LLM to summarize the noisy/redundant short-term visual captions and then answer a given question based on the summarized captions leads to the most significant boost in performance (+5.8%) among the prompts we have tried (e.g., zero-shot CoT, Self-Consistency).

GPT-3.5 provides the best tradeoff between accuracy, computational cost, and sufficiently large context length.

Using LaViLa visual captioner leads to best results (51.8%) followed by BLIP-2 (46.7%) and EgoVLP (46.6%).

Few-shot in-context learning leads to a large improvement on both the variant of our model with the simplest prompt (+4.7%) and our best-performing variant with a multi-round prompt (+4.1%).

Densely extracting visual captions from consecutive 1-second video clips of the long video input leads to the strongest performance on EgoSchema.

Our final method, built using the above-listed empirical insights, achieves 50.3% zero-shot LVQA accuracy on the full test set of EgoSchema, outperforming the previous best approach by 18.1% as shown in Figure 1.

Our LLoVi also outperforms prior approaches on NeXT-QA and IntentQA by 4.1% and 3.1%, and it also achieves state-of-the-art performance on NeXT-GQA, a recently introduced grounded LVQA benchmark.

We hope that our simple, training-free method will encourage new ideas and a simpler model design in LVQA. We will release our code to enable the community to build on our work.

Related Work

Long-range Video Understanding. Modeling long-range videos (e.g., several minutes or longer) typically requires models with sophisticated temporal modeling capabilities, often leading to complex model design. LF-VILA proposes a Temporal Window Attention (HTWA) mechanism to capture long-range dependency in long-form video. MeMViT and MovieChat adopt a memory-based design to store information from previously processed video segments. Several prior methods use space-time graphs or relational space-time modules to capture spatiotemporal dependencies in long videos. Lastly, the recently introduced S4ND , ViS4mer and S5 use Structured State-Space Sequence (S4) layers to capture long-range dependencies in the video. Unlike these prior approaches, we do not propose any complex long-range temporal modeling modules but instead develop a simple and strong LLM-based framework for zero-shot LVQA.

LLMs for Video Understanding. The recent surge in large language models (LLMs) has inspired many LLM-based applications in video understanding. Models like Socratic Models and VideoChat integrate pretrained visual models with LLMs for extracting visual concepts and applying them to video tasks. Video ChatCaptioner and ChatVideo leverage LLMs for video representation and dialog-based user interaction, respectively. VidIL employs LLMs for adapting image-level models to video tasks using few-shot learning. Beyond short-term understanding, studies have explored LLMs for long-range video modeling. The work in uses GPT-4 for video summarization but lacks quantitative evaluation. Meanwhile, focuses on movie datasets, requiring limited visual analysis . In contrast, we conduct our experiments on the EgoSchema benchmark for long-range LVQA and provide an extensive empirical analysis of various design choices of our simple approach.

Video Question Answering. Unlike image question-answering, video question-answering (VidQA) presents unique challenges, requiring both spatial and temporal reasoning. Most existing VidQA methods, either using pretraining-finetuning paradigms , zero-shot , or few-shot learning , focus on short-term video analysis (5-30s). To overcome the limitations of short-term VidQA, new benchmarks have been proposed covering longer video durations: NextQA averages 44s, while ActivityNet-QA , TVQA , How2QA , MovieQA , and DramaQA range from 100s to several minutes. Despite these lengths, works in found that many benchmarks can be solved by analyzing only brief clips, not requiring extensive video modeling, and exhibited language biases, being solvable using pure text-only methods that ignore visual content. To address these issues, the EgoSchema benchmark was recently introduced, requiring at least 100 seconds of video analysis and not exhibiting any language biases. Thus, our work focuses on the LVQA analysis on EgoSchema.

Method

Recently, LLMs have been shown to excel on a wide range of long-range modeling tasks . Motivated by these long-range modeling capabilities, we propose to tackle the long-range video question-answering (LVQA) task by decomposing it into two subtasks: 1) short-term video clip captioning, and 2) long-range text-based video understanding. Our motivation for such decomposition is to separate the short-term and long-range video modeling so that we could leverage the strong existing short-term visual captioners (e.g., LaViLa, BLIP2) for the first sub-task, and powerful zero-shot LLMs (e.g., GPT-3.5, GPT-4, LLaMA, T5) for the second sub-task, thus, exploiting the power of LLMs for long-range modeling.

Our decomposed LVQA framework, named LLoVi (Language-based Long-range Video Question-answering), brings important advantages. First, our approach is simple as it does not rely on complex/specialized long-range video modeling operators (e.g., memory queues, state-space layers, space-time graphs, etc.). Second, our framework is training-free, which makes it easy to apply it to LVQA in zero-shot settings. Third, our method is agnostic to the exact choices of a visual captioner and LLM, and thus, it can benefit from future improvements in visual captioning/LLM model design. Figure 2 presents a detailed illustration of our high-level approach. Below, we provide details about each component of our framework.

2 Long-range Reasoning with an LLM

We want to leverage foundational LLMs for holistic long-range video understanding rather than relying on complex long-range temporal modeling schemes (e.g., memory queues, state-space layers, space-time graphs, etc.) used by prior methods .

Formally, given short-term visual captions {cm}m=1Nv\{c_{m}\}_{m=1}^{N_{v}} for all NvN_{v} short video clips, we first concatenate the clip captions into the full video captions C=[c1,…,cNv]C=[c_{1},\ldots,c_{N_{v}}] in the same order as the captions appear in the original video. Afterward, the concatenated video captions CC are fed into an LLM for long-range video reasoning. Specifically, given the concatenated video captions CC, the question QQ, and the answer candidates AA, we prompt the LLM to select the correct answer using the following prompt template: “Please provide a single-letter answer (A, B, C, D, E) to the following multiple-choice question {Q}\{Q\}. You are given language descriptions of a video. Here are the descriptions: {C}\{C\}. Here are the choices {A}\{A\}.”. The full prompt is included in Supplementary Materials.

Our experiments in Section 4.3 suggest that this simple approach works surprisingly well for LVQA. However, we also discovered that many modern LLMs (e.g., GPT-3.5, LLaMA) may struggle when provided with long (>>1K words), noisy, and potentially redundant/irrelevant caption sequences. To address these issues, we investigate more specialized LLM prompts that ask an LLM first to summarize the noisy short-term visual captions (first round of prompting) and then answer a given question about the video (second round of prompting). Specifically, we formulate such a multi-round prompt as follows: given the video captions CC, the question QQ, and the answer candidates AA, instead of directly feeding the {C,Q,A}\{C,Q,A\} triplet into LLM for LVQA, we first ask the LLM to provide a summary of the captions in the first round, which we denote as SS using the following prompt template: “You are given language descriptions of a video: {CC}. Please give me a {NwN_{w}} word summary.” NwN_{w} denotes the desired number of words in the summary SS. Afterward, during the second round of prompting, instead of using the captions CC, we use the summary SS as input for the LLM to select one of the answer candidates. Conceptually, this may be beneficial, as the LLM-generated summary SS filters out potentially irrelevant/noisy information from the initial set of captions CC, making LLM inputs for the subsequent QA process more succinct and cleaner. A detailed illustration of our multi-round prompt is shown in Figure 3.

3 Implementation Details

For the experiments on EgoSchema, we use LaViLa as our short-term captioner. We note that LaViLa is pretrained on Ego4D, the same dataset that the authors of EgoSchema use to build their benchmark. Thus, to avoid any data intersection between training and testing sets, we retrain our variant of the LaViLa model on a subset of 6K Ego4D videos that do not include any EgoSchema videos. To process video inputs, we segment each video into multiple 1s clips with a stride of 1s, resulting in a list of consecutive clips that cover the entire video. We use GPT-3.5 as the LLM for long-range reasoning on EgoSchema. For NeXT-QA, IntentQA and NeXT-GQA, we use LLaVA-1.5 as the visual captioner and GPT-4 as the LLM. We downsample the videos to 0.5 FPS and prompt LLaVA to generate captions that contain roughly 30 words for each frame. We use LaViLa as our visual captioner on EgoSchema because compared to other captioners (e.g. BLIP-2, LLaVA), LaViLa works better on first-person view videos. For the third-person view video datasets, we use LLaVA because we found its performance to be better than LaViLa. We also observe that GPT-4 generally performs better than other language models. However, on EgoSchema, we found the output of GPT-4 inconclusive for many questions (e.g. an output message that the provided information is insufficient to answer the given question). In contrast, we did not experience such issues with GPT-3.5. We also did not observe these issues with GPT-4 on other datasets except EgoSchema. We provide more implementation details in the Supplementary Material.

Experiments

Unlike short-term video question-answering, long-range video question-answering (LVQA) lacks robust and universally agreed-upon benchmarks. Many prior movie-based long-range video understanding benchmarks exhibit significant language biases, i.e., pure text-based approaches can achieve excellent performance by using information from subtitles/speech while ignoring the video content entirely (See the example in the top row of Figure 4). Furthermore, while many existing long-range video understanding benchmarks contain long video inputs, prior work have shown that all of these benchmarks can be solved by analyzing only several second-long video clips within the longer video inputs, without requiring any long-range video modeling capabilities (See the example in the middle row of Figure 4).

To address these limitations, recent work introduced EgoSchema , a new long-range video question-answering benchmark, consisting of 5K multiple choice question-answer pairs, spanning 250 hours of video, and covering a wide range of human activities. Unlike prior benchmarks in this area, the authors in have manually verified that to answer a given question correctly, one must watch at least 100 seconds of video content, which is orders of magnitude longer than any existing benchmark (See the example in the bottom row of Figure 4). Furthermore, unlike prior LVQA benchmarks, EgoSchema has no textual inputs such as speech, subtitles, or storylines. This means the questions must be answered based only on video inputs, preventing language-based biases. The EgoSchema benchmark consists of 5,000 questions, each requiring the correct answer to be selected between 5 given options based on a three-minute-long video clip. The entire dataset is designed for zero-shot evaluation and has no training set. For validation, the authors released a subset of 500 questions with ground truth answers (EgoSchema Subset). By default, our experiments are conducted on the EgoSchema Subset. The metric we use is QA accuracy, i.e., the percentage of correctly answered questions among all questions.

In addition to EgoSchema, we also perform zero-shot LVQA experiments on three other LVQA benchmarks:

NExT-QA contains 5,440 videos with an average duration of 44s and 48K multi-choice questions and 52K open-ended questions. There are 3 different question types: Temporal, Causal, and Descriptive. Following common practice, we perform zero-shot evaluation on the validation set, which contains 570 videos and 5K multiple-choice questions.

IntentQA contains 4,303 videos and 16K multiple-choice question-answer pairs focused on reasoning about people’s intent in the video. We perform a zero-shot evaluation on the test set containing 2K questions.

NExT-GQA is an extension of NExT-QA with 10.5K temporal grounding annotations associated with the original QA pairs. The dataset was introduced to study whether the existing LVQA models can temporally localize video segments needed to answer a given question. We evaluate all methods on the test split, which contains 990 videos with 5,553 questions, each accompanied by a temporal grounding label. The metrics we used include: 1) Intersection over Prediction (IoP) , which measures whether the predicted temporal window lies inside the ground truth temporal segment, 2) temporal Intersection over Union (IoU), and 3) Acc@GQA, which depicts the percentage of accurately answered and grounded predictions. For IoP and IoU, we report the mean values and values with the overlap thresholds of 0.5.

To study the most important factors in our framework design, we first conduct an empirical study on the released subset of EgoSchema. Afterward, we present our main results on the full EgoSchema test set. Lastly, we present our results on other datasets and tasks.

2 Empirical Study on EgoSchema

Before presenting our main results, we first study the effectiveness of different components within our LLoVi framework, including (i) the visual captioner, (ii) the LLM, (iii) the LLM prompt design, and (iv) few-shot in-context learning. The experiments are conducted on the EgoSchema Subset with 500 multi-choice questions. We discuss our empirical findings below. We also include additional experiments in the supplementary material.

In Table 1, we study the effectiveness of various clip-level video captioners, including LaViLa , EgoVLP , and VideoBLIP . In addition to video captioners, we also try the state-of-the-art image captioner, BLIP-2 . Lastly, to study the upper bound of our visual captioning results, we include the ground truth Oracle captioning baseline obtained from the Ego4D dataset. All baselines in Table 1 use similar experimental settings, including the same LLM model, i.e., GPT-3.5. The results are reported as LVQA accuracy on the EgoSchema Subset.

Based on the results in Table 1, we observe that LaViLa is the best captioning model, outperforming BLIP-2, EgoVLP, and VideoBLIP. We also observe that despite not being pre-trained on Ego4D , BLIP-2 performs reasonably well (46.7%) and even outperforms other strong Ego4D-pretrained baselines, EgoVLP and VideoBLIP. Lastly, the Oracle baseline with ground truth captions outperforms LaViLa captions by a large margin (14.0%). This demonstrates that our framework can benefit from future improvements in visual captioning models.

2.2 Large Language Model

In Table 2, we analyze the performance of our LLoVi framework using different LLMs while fixing the visual captioner to be LaViLa. Based on these results, we observe that GPT-4 achieves the best performance (58.3%), followed by GPT-3.5 (51.8%). These results suggest that stronger LLMs (GPT-4) are better at long-range modeling, as indicated by a significant margin in LVQA accuracy between GPT-4 and all other LLMs (>>6.5%). We also note that the Llama2 performs reasonably well with its 70B variant (50.6%), but its performance drastically degrades with smaller capacity LLMs (i.e., Llama2-7B, Llama2-13B). Due to the tradeoff between accuracy and cost, we use GPT-3.5 for all of our remaining experiments unless noted otherwise.

2.3 LLM Prompt Analysis

In this section, we (1) analyze several variants of our summarization-based prompt (described in Section 3), and (2) experiment with other commonly used prompt designs, including Zero-shot Chain-of-Thought (Zero-shot CoT) , Plan-and-Solve , and Self-Consistency . Below, we present a detailed analysis of these results.

A Multi-round Summarization-based Prompt. As discussed in Section 3, we found that a specialized multi-round prompt that first asks the LLM to summarize the noisy short-term captions, and then answer the question using the LLM-generated summary performs better than our base approach that uses caption inputs directly for LVQA. Given a concatenated set of captions CC, an input question QQ, and a set of candidate answers AA, there are several input combinations that we can use to obtain the summary SS. Thus, here, we investigate three distinct variants of obtaining summaries SS:

(C) →\rightarrow S: the LLM uses caption-only inputs CC to obtain summaries SS in the first round of prompting.

(C, Q) →\rightarrow S: the LLM uses captions CC and a question QQ as inputs for generating summaries SS. Having additional question inputs is beneficial as it allows the LLM to generate a summary SS specifically tailored for answering an input question QQ.

(C, Q, A) →\rightarrow S: the LLM takes captions CC, a question QQ, and the answer candidates AA as its inputs to produce summaries SS. Like above, having additional answer candidate inputs enables the LLM to generate a summary SS most tailored to particular question-answer pairs.

In Table 3, we explore the effectiveness of these three prompt variants. Our results show that all three variants significantly outperform our simple, yet already strong baseline that uses a standard LVQA prompt (described in Section 3). Specifically, we note that the variant (C) →\rightarrow S that uses caption-only inputs to obtain the summaries outperforms the standard baseline by 1.8%. Furthermore, we observe that incorporating a given question as an input (i.e., the (C, Q) →\rightarrow S variant) leads to the best performance (57.6%) with a significant 5.8% boost over the standard LVQA prompt baseline. This confirms our earlier intuition that having additional question QQ inputs enables the LLM to generate a summary SS specifically tailored for answering that question, thus leading to a big boost in LVQA performance. Lastly, we observe that adding answer candidates AA as additional inputs (i.e., the (C, Q, A) →\rightarrow S variant) leads to a drop in performance (-1.7%) compared with the (C, Q) →\rightarrow S variant. We conjecture that this might be because the wrong answers in the candidate set AA may mislead the LLM, leading to a suboptimal summary SS.

We also investigate the optimal length of the generated summary SS, and present these results in Table 4. Specifically, for these experiments, we ask the LLM to generate a summary SS using a different number of words (as part of our prompt). We use the best performing (C, Q) →\rightarrow S variant for these experiments. Our results indicate that using a very small number of words (e.g., 50) leads to a drop in performance, indicating that compressing the caption information too much hurts the subsequent LVQA performance. Similarly, generating summaries that are quite long (e.g., 700 words) also leads to worse results, suggesting that the filtering of the potentially noisy/redundant information in the captions is important for good LVQA performance. The best performance is obtained using 500-word summaries.

Comparison with Commonly Used Prompts. Next, in Table 5, we compare our multi-round summarization-based prompt with other commonly used prompts such as Zero-shot Chain-of-Thought , Plan-and-Solve , and Self-Consistency . From these results, we observe that all of these prompts outperform the base variant of our model that uses a standard prompt. In particular, among these commonly used prompts, the self-consistency prompting technique achieves the best results (55.4%). Nevertheless, our multi-round summarization-based prompt achieves the best performance (57.6%).

2.4 Few-shot In-Context Learning

In-context learning with LLMs has shown strong few-shot performance in many NLP tasks . In Table 5, we evaluate the few-shot in-context learning capabilities of our LLoVi framework. Our results show that our LLoVi framework greatly benefits from few-shot in-context learning. Specifically, the few-shot in-context learning leads to a 4.7% boost on the variant of our framework that uses a standard prompt and 4.1% boost on our advanced framework using a multi-round summarization-based prompt. We used 6 few-shot examples for our experiments as we found this configuration to produce the best performance.

3 Main Results on EgoSchema

In Table 6, we evaluate our best-performing LLoVi framework, developed using our empirical insights on the full EgoSchema test set containing 5K video samples. We compare our approach with prior state-of-the-art methods including InternVideo , mPLUG-Owl , VIOLET , FrozenBiLM . Based on these results, we observe that the best-performing zero-shot variant of our LLoVi framework achieves 50.3% accuracy, outperforming the previous best-performing InternVideo model by a large margin (18.2%). Additionally, we show that using few-shot in-context learning, our best variant gets further improvement. These results validate our design choice of leveraging the long-range modeling capabilities of LLMs for the LVQA task. Furthermore, since our proposed LLoVi framework is agnostic to the visual captioning model and an LLM it uses, we believe that in the future, we could further improve these results by leveraging more powerful visual captioners and LLMs.

4 Results on Other Datasets

Next, we demonstrate that our simple framework generalizes well to other LVQA benchmarks.

NExT-QA. In Table 7, we evaluate LLoVi on the NExT-QA validation set in a zero-shot setting. We compare our approach with prior methods: VFC , InternVideo , ViperGPT , and SeViLA . We observe that LLoVi outperforms the previous best-performing method, SeViLA by 4.1%. Notably, in the Causal category, LLoVi achieves 8.2% improvement. We conjecture this improvement comes from the simple 2-stage design of our LLoVi framework: captioning followed by LLM reasoning. By captioning the video, we are able to directly leverage the reasoning ability of the powerful LLMs and thus achieve good causal reasoning performance.

IntentQA. In Table 8, we evaluate our method on the IntentQA test set. In our comparisons, we include several supervised methods (HQGA , VGT , BlindGPT , CaVIR ) and one recent zero-shot approach, SeViLA. From the results in Table 8, we observe that our method greatly outperforms all prior approaches, both in the fully supervised and zero-shot settings.

NExT-GQA. In Table 9, we extend our framework to the grounded LVQA task and evaluate it on the NExT-GQA test set. We compare LLoVi with the weakly-supervised methods: IGV , Temp[CLIP](NG+) , FrozenBiLM(NG+) and SeViLA . These baselines are first trained on NExT-GQA to maximize the QA accuracy, and then use ad-hoc methods to estimate a relevant video segment for question-answering. Although LLoVi is not trained on NExT-GQA, it still outperforms these weakly-supervised methods by a large margin on all evaluation metrics. These results demonstrate that in addition to LVQA, our framework can also be used to temporally ground its predictions for more explainable long-range reasoning.

Conclusion

In this work, we present a simple, yet highly effective LLM-based framework for long-range video question-answering (LVQA). We thoroughly evaluate various design choices of our approach and use our empirical insights to develop a method that improves upon prior LVQA approaches by a significant margin on the newly introduced EgoSchema benchmark. We also demonstrate that our framework generalizes to other LVQA benchmarks such as NeXT-QA, IntentQA, and it can be extended to grounded LVQA tasks. While our paper does not provide any major technical contributions, we hope that our simple LVQA framework will help inspire new ideas and simplify model design in long-range video understanding.

References

Appendix A Additional Analysis

In this section, we provide additional analysis on the EgoSchema Subset using the standard prompt.

In Figure 5, we investigate the sensitivity of LVQA performance on EgoSchem with different video sampling configurations. Specifically, in Subfigure 5(a), we experiment with 4 different clip lengths: 0.5s, 1s, 4s, and 8s. For each clip length, we use the stride that would be sufficient to cover the entire long video input. For these experiments, we use a LaViLa visual captioner and a GPT-3.5 LLM. Our results indicate that LVQA performance is the best when the sampled clip length is 1s. We observe that using an even shorter video clip length (i.e., 0.5s) produces many repetitive/redundant captions, which leads to 2% drop in LVQA performance. Furthermore, we also note that increasing the clip length to longer durations (e.g., 2s-8s) makes the accuracy lower since the extracted captions start to lack detailed visual information needed to answer the question. In addition, in Subfigure 5(b), we fix the clip length to 1s and experiment with 4 different stride values: 1s, 2s, 4s, and 8s. Note that the 1s clips sampled using a 1s stride will cover the entire video without overlap. Our results suggest a gradual decrease in LVQA accuracy when increasing the stride from 1s to 8s. This indicates that having gaps in long video coverage leads to suboptimal LVQA performance. Thus, for the rest experiments, we use 1s video clips sampled with a 1s stride.

A.2 Accuracy on Different Question Types

To better understand the strengths and limitations of our LVQA framework, we manually categorize questions in the EgoSchema Subset into 5 categories: (1) Purpose/Goal Identification, (2) Tools and Materials Usage, (3) Key Action/Moment Detection, (4) Action Sequence Analysis, (5) Character Interaction (see Supplementary Materials for details). Note that some questions belong to more than one category. Based on this analysis, we observe that almost half of the questions relate to purpose/goal identification, which makes intuitive sense as inferring human goals/intent typically requires a very long video analysis. We also observe that a significant portion of the questions relate to tool usage, key action detection, and action sequence analysis. Lastly, the smallest fraction of the questions belong to character interaction analysis.

In Table 10, we break down our system’s performance according to each of the above-discussed question categories. Our results indicate that our system performs the best in the Character Interaction category (63.8%). One possible explanation is that the LaViLa model, which we use as our visual captioner, is explicitly pretrained to differentiate the camera wearer from other people, making it well-suited for understanding various interactions between characters in the video. We also observe that our framework performs much worse in the Key Action/Moment Detection category (43.5%). We conjecture that this might be caused by the limitations in the visual captioning model, i.e., if the key action fails to appear in any of the visual captions, the question will be almost impossible to answer. Lastly, we note that our model performs quite well on the remaining categories (>>50%). It is especially encouraging to see strong results (54.9%) in the Purpose/Goal Identification category since inferring human intentions/goals from the video inherently requires very long-form video analysis.

Appendix B Additional Implementation Details

For most experiments on EgoSchema, we use LaViLa as the visual captioner. For other pre-trained visual captioners, we use off-the-shelf pre-trained models, e.g., BLIP2 , EgoVLIP .

LaViLa is trained on the Ego4D dataset. The original LaViLa train set has 7743 videos with 3.9M video-text pairs and the validation set has 828 videos with 1.3M video-text pairs. The EgoSchema dataset is cropped from Ego4D. Since EgoSchema is designed for zero-shot evaluation and the original LaViLa train set includes EgoSchema videos, we retrain LaViLa on Ego4D videos that do not have any overlap with EgoSchema videos to avoid unfair comparison with other methods. After removing the EgoShema videos, the train set consists 6100 videos with 2.3M video-text pairs, and the validation set has 596 videos with 0.7M video-text pairs. We retrain LaViLa on this reduced train set to prevent data leakage. LaViLa training consists of two stages: 1) dual-encoder training and 2) narrator training. Below we provide more details.

Dual-encoder. We use TimeSformer base model as the visual encoder and a 12-layer Transformer as the text encoder. The input to the visual encoder comprises 4 RGB frames of size 224×224224\times 224. We randomly sample 4 frames from the input video clip and use RandomResizedCrop for data augmentation. The video-language model follows a dual-encoder architecture as CLIP and is trained contrastively. Following LaViLa , we use 1024 as batch size. We train at a 3×10−53\times 10^{-5} learning rate for 5 epochs on 32 NVIDIA RTX 3090 GPUs.

Narrator is a visually conditioned autoregressive Language Model. It consists of a visual encoder, a resampler module, and a text encoder. We use the visual encoder (TimeSformer base model) from the pretrained dual-encoder (See the previous paragraph). The resampler module takes as input a variable number of video features from the visual encoder and produces a fixed number of visual tokens (i.e. 256). The text decoder is the pretrained GPT-2 base model with a cross-attention layer inserted in each transformer block which attends to the visual tokens of the resampler module. We freeze the visual encoder and the text decoder, while only training the cross-attention layers of the decoder and the resampler module. Following the design in LaViLa , we use a batch size of 256 and a learning rate of 3×10−53\times 10^{-5}. We use AdamW optimizer with (β1,β2)=(0.9,0.999)(\beta_{1},\beta_{2})=(0.9,0.999) and weight decay 0.01. We train the model on 8 NVIDIA RTX 3090 GPUs for 5 epochs.

Narrating video clips. We use nucleus sampling with p=0.95p=0.95 and return K=5K=5 candidate outputs. Then we take the narration with the largest confidence score as the final caption of the video clip.

For NExT-QA, IntentQA and NExT-GQA datasets, we use LLaVA1.5 as the visual captioner and GPT-4 as the LLM. Specifically, we use the llava-1.5-7b-hf variant with the prompt “USER: <<image>>. Describe the image in 30 words. ASSISTANT: ”.

B.2 LLMs

For most experiments on EgoSchema we use GPT-3.5 as the LLM. Specifically, we use the gpt-3.5-turbo-0613 variant which has 4K context. When the context length is not enough, we use the gpt-3.5-turbo-16k variant. We use 0 as temperature for all experiments.

We use Llama-2-7b-chat-hf, Llama-2-13b- chat-hf, and Llama-2-70b-chat-hf variants as Llama2 models. For all Llama2 models, we use greedy sampling to generate the output.

For NExT-QA, IntentQA and NExT-GQA datasets, we use GPT-4 as the LLM with the variant gpt-4-1106-preview.

B.3 Prompting Techniques Implementation

Prompt Details. We provide detailed prompts for our standard prompt in Table 11, multi-round summarization-based prompt in Table 12, Zero-shot Chain of Thought in Table 13, and Plan-and-Solve prompting in Table 14. To implement Self-Consistency, we set the temperature of GPT3.5 to 0.70.7 and run Zero-shot Chain of Thought for 5 times following the design in Self-Consistency . Each run of the model provides a result, and the final output is determined by a majority vote. The prompt for the grounded LVQA benchmark is shown in Table 15.

Output Processing. When answering multiple choice questions, GPT3.5 usually outputs complete sentences instead of a single-letter answer, i.e. A, B, C, D, or E. One way to obtain the single-character response is to perform post-processing on the output, which usually requires substantial engineering efforts. In our work, however, we observe that GPT3.5 is very sensitive to the starting sentences of the prompts. Therefore, we explicitly prompt it as in Table 11 to force GPT3.5 to generate a single character as response. In practice, we take out the first character of the output as the final answer.

Appendix C Qualitative Analysis

In Table 16 we compare different captions generated by BLIP2 and LaViLa on EgoSchema. LaViLa captions are generally more concise than BLIP2 captions, focusing more on the actions while BLIP2 focuses more on describing the objects. We also observe that LaViLa is better at differentiating the camera wearer and other person. As shown in the second image in Table 16, LaViLa tends to focus more on the action of the other person when the camera wearer and other person both appear in the video.

C.2 LLoVi with Standard Prompt

We show two examples of our method with standard prompt, including a successful one and a failed one in Figure 6. Our method performs long-range modeling from short-term video captions through LLM to understand the video. In the success case demonstrated in Subfigure 6(a), the captions describe the camera wearer’s action in a short period of time, such as the interation with the tape measure and the wood. With the short-term captions, LLM understand the long video and answers the question correctly.

In the failure case shown in Subfigure 6(b), although the video captioner identifies the object in the video correctly as a tablet, LLM understands the action of the camera wearer as watching TV rather than using an iPad. This might be caused by misguidance from the redundant captions that are not related to the question.

C.3 LLoVi with Multi-round Summarization-based Prompt

Figure 7 illustrates two EgoSchema questions that our framework with multi-round summarization-based prompt answers correctly. In Subfigure 7(a), the question asks for the primary function of a tool that the video taker uses. However, shown in the first two images, the long video contains descriptions that are not related to the question, such as operating a machine and rolling a dough. As a result, the generated text captions would contain a large section that is not our direction of interest. By summarizing the captions with awareness to the question, LLM extracts key information and cleans redundant captions to provide clearer textual background for answering the question. The same pattern is observed in Subfigure 7(b).

Figure 8 shows two questions that our method fails to answer. In the summarization stage, the LLM answers the question directly instead of using the question to guide the summarization. For example, in Subfigure 8(a), all the frames show the camera wearer engaging in actions related to washing dishes, but LLM infers that the person is cleaning the kitchen in the summarization stage. This wrong inference further misdirects the following question answering stage, which leads to an incorrect answer. In Subfigure 8(b), LLM concludes that the cup of water is used to dilute the paint because the camera wearer dips the brush into water before dipping it into the paint palette.

In Figure 9, we also show a question which the standard prompt fails to answer, but the multi-round summarization-based prompt answers correctly. In the video in the example question, we observe the camera wearer involving in activities related to laundry, such as picking up clothes from the laundry basket and throwing them into the washing machine. However, the short-term video captions shown in Subfigure 9(a) demonstrate the redundancy of actions. The repetitive actions complexes extracting and comprehending the information presented in the caption. For example, excessive captions on picking up clothes can make LLM think that the camera wearer is packing something. Our multi-round summarization-based prompt mitigate this problem by first ask LLM to provide a summary of the captions. The summary shown in Subfigure 9(b) states clearly that the camera wearer is doing laundry. With the cleaner and more comprehensive summary, the LLM answer the question correctly.

C.4 Question Categories

We provide detailed descriptions of each question category in Table 17. Note that each question can be classified into multiple categories.