EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding
Karttikeya Mangalam, Raiymbek Akshulakov, Jitendra Malik
Introduction
We introduce EgoSchema, a diagnostic benchmark for assessing very long-form video-language understanding capabilities of modern multimodal systems. Understanding long natural videos requires a host of interconnected abilities such as action and scene understanding, perceiving and tracking object states, long-term visual memory, abstract reasoning, hierarchical information aggregation, and more. Shown in Fig. 1 is an exemplar of the curated EgoSchema dataset. Consider the visual cognitive faculties involved in answering the question: ‘What is the overarching behavior of C and the man in the video?’. First, is the spatial recognition capabilities for disambiguating the referred character ‘C’ (camera wearer) and ‘the man’ as well as the present objects such as ‘cards’, ‘notebook’, deck as so on. Next is short-term temporal recognition capabilities of understanding the atomic actions and movement of the characters such as ‘playing’, ‘taking notes’, ‘shuffling’ etc. Built upon these are the capabilities for visually understanding the mental states such ‘distracted’, ‘attention’ and social dynamics such as ‘teaching’, ‘showing’. Next are medium-term actions such as ‘organizing the deck’ or ‘keeping track’. Finally, long-term reasoning capabilities need to be employed for abstracting the ‘overarching behavior’ of the video from all the low-level signals to be able to rule out all the other wrong options and conclude option 3 to be correct. Note that even for humans, it is impossible to answer the illustrated questions with only the shown 9 uniformly sampled frames from the three-minute video (Fig. 1).
While there have been some prior attempts to formulate long-form video tasks , they broadly tend to fall into two failure modes. The first failure mode stems from the difficulty of capturing the explosive diversity of human behavior in narrow pre-defined label spaces that leading unduly narrow and oddly specific tasks, such as like ratio or relationship prediction . Hence, we propose to probe video systems capturing the rich complexity of long-form video with something just as rich and complex – natural language. However, natural language outputs are notoriously difficult to evaluate with popular metrics such as BLEU and ROUGE having well-known shortcomings . Hence, we propose to evaluate language understanding as a multiple-choice question-answering task, thereby using the well-defined benchmark metric of overall question-answering accuracy.
The second failure mode for a long-term video task is that the proposed task happens to actually be a short-term one - only disguised as a long-term task. To measure the intrinsic "long-term" nature of a video understanding task, we propose the notion of temporal certificate length . Intuitively, certificate length (§3.2) is the length of the video a human verifier needs to observe to be convinced of the veracity of the marked annotation. The idea of temporal certificates is not limited only to question-answering or vision-language tasks but is applicable to several video understanding tasks, including pure vision tasks such as action classification, detection, or even temporal action localization.
Based on the length of the temporal certificate, we propose the following temporal understanding taxonomy for video tasks: Datasets with certificate length in the order of second are termed short video tasks. Next, we name datasets with certificate length in the order of seconds as, long-form video tasks. Finally, datasets with certificate length in the order of seconds are termed as, very long-form video tasks. Fig. 3 presents estimates of the certificate lengths for a variety of datasets plotted against the temporal length of the video clip. We observe that the temporal certificate length is quite weakly correlated with the length of the video clip. This is due to the intentional design choice in defining the certificate set, which decouples the task of searching or retrieving the relevant sub-clip from a bigger clip from the task of visually understanding the retrieved sub-clip. And in this manner, using temporal certificate length as a metric for measuring the intrinsic temporal hardness of a dataset, avoids the failure mode of formulating an implicitly short-term task disguised as a long-term one. Section 3.2 details precise operationalizations for estimating the temporal certificate sets.
In summary, our contributions are three-fold. First, we propose the notion of temporal certificates, a broadly applicable notion that measures the intrinsic temporal hardness of clips in a video understanding dataset. We estimate temporal certificate lengths for a broad variety of existing datasets and show that EgoSchema has a median temporal certificate of about seconds, which is longer than the dataset with the second longest certificate length , and to longer than all other existing video understanding datasets (with or without language). Second, building upon the notion of temporal certificates, we introduce EgoSchema, a diagnostic benchmark for assessing the very long-form video understanding capability of multimodal video-language systems. Third, we benchmark both state-of-the-art video-language systems and humans in Zero-shot settings on EgoSchema to find that even the most advanced current video-language understanding systems consisting of billion of parameters achieve very low accuracy in long-from multiple-choice question-answering (< 33%) while humans achieve about accuracy in the unconstrained setting.
Related Works
Video Question-Answering Datasets. Visual Question-Answering is a popular video-language task with several large internet-scale datasets for video-language pre-training such as Ego4D , HowTo100M and HowToVQA69M . However, as the scope and size of pre-training datasets and models soar, it becomes critical to construct evaluations for assessing the model capabilities on various axes. Hence, many smaller datasets have been proposed for evaluating different aspects of video-language understanding such as compositional reasoning , causal and common scene comprehension , instruction understanding , video description ability , dynamic environments understanding , complex web video understanding , situated reasoning , spatiotemporal reasoning , social intelligence , dynamic neuro-symbolic reasoning , external knowledge-based reasoning and many more . How2VQA69M and iVQA have leveraged HowTo100M ASR text for generating questions. However, unlike Ego4D narrations that are used in EgoSchema, ASR text does not necessarily describe the visual elements in the scene. Hence, questions can suffer from biases where a key required information is visually absent. Additionally, generated question-answers also have quite short certificate lengths (iVQA in Fig. 3) due to the local nature of the ASR text.
Long-form Video Understanding Datasets have been very sparsely explored in prior works. posits a long-form video understanding benchmark but the proposed tasks are unduly narrow and specific, such as the ‘like’ ratio and view count prediction. Also, average certificate length is about smaller than EgoSchema.
proposes a dataset for benchmarking efficient video inference consisting of frame-wise object mask annotations from Mask-RCNN but without any long-term annotations. introduces a dataset of about 111 hours of video sourced from Kinetics-400 for generic event boundary detection. While the task itself requires comprehensive understanding, the video clip length is only 10 seconds long, with temporal certificates (§3.2) being much shorter. proposes a question-answering dataset based on long movie clips but due to the open-ended nature of questions, successful approaches tend to neglect the visual data and are biased purely with approaches using additional text such as story lines. proposes MAD, a language grounding dataset with an average clip of minutes. However, the length of the retrieved clip is quite short (average seconds) thereby resulting in a temporal certificate (§3.2) only a few seconds long. Further, MAD and several other movie-based datasets do not release any video data because of copyright issues. In contrast, EgoSchema has an average certificate length of about seconds. Further, EgoSchema will be publicly released under the Ego4D license, which allows direct public use of the video and text data for both research and commercial purposes.
Collecting EgoSchema
Collecting video and language datasets, even without a focus on very long-form video is quite challenging. Manually collecting, observing, and annotating videos with free-form language, in contrast to using images and pre-defined label categories, is both labor-intensive and time-consuming and thereby quite expensive. In addition to burgeoning cost, ensuring visual data diversity and minimizing visual and linguistic bias while ensuring high quality of marked annotations also contribute to the overall difficulty. All these factors get severely more challenging for long-form videos.
In this work, we propose a staged data collection pipeline (Fig. 4) utilizing existing large-scale but short-term video datasets, rule-based filtering procedures, and exciting new capabilities afforded by LLMs to significantly lighten the burden on human annotators. We use the proposed pipeline for curating EgoSchema, a high-quality and diverse very long-form video question-answering dataset. Associated datasheets and data cards for EgoSchema are provided in the supplementary.
Ego4D has over 3670 hours of RGB video spread consisting of over 3.85 million narration instances covering over 1,772 unique verbs (activities) and 4,336 unique nouns (objects) . The narrators are instructed to continuously pause and describe everything that the camera wearer (‘C’) does. This creates dense and precise narrations that accurately describe the visuals.
Naturally, the collected video has non-uniform length and narration density. Since we would like to standardize the clip length for evaluation and have sufficiently rich narrations to allow interesting question-answer pairs to form in later stages, we filter the data based on the length and narration density. We choose to filter for non-overlapping three-minute clips each with at least 30 human annotated narrations (each narration is a timestamped sentence) to build EgoSchema. Detailed statistic of the number of viable clips for different possible length and narration density choices is discussed in supplementary.
1.2 Stage II: Question Answer Generation
The filtered narrations are processed with a capable LLM to generate Question-Answer triplets (), each consisting of the question , the correct answer , and wrong answers , per clip. To achieve this, we experimented with several LLM inference call chaining procedures with trade-offs between quality and cost of generation that are briefly described next.
One-shot is the simplest prompting procedure to prompt for all instances of in one inference call. This is the most cost-efficient option but we found the generations to be of significantly low quality. The generated often are very similar to each other and the generated have a very high false positive rate for the correct answers as well as a false negative rate for the wrong answers.
N-shot is the next natural prompting procedure where we generate one per LLM inference call. This significantly improves the false positive and false negative rates but since the generated are independent and generated with the same prompt, they still tend to be very similar (comparable to one-shot), even at higher sampling temperatures. Further, the cost of generation also scales with .
QAW-shot generates each of the questions in one inference call, followed by another inference call for generating correct answer and finally, wrong answers, . Since each of the is generated jointly, they can be forced to be distinct with appropriate prompting. Similarly, the generated and can also be made distinct. However, this requires 3 chained LLM inference calls, and generation failures in earlier calls cascade steeply.
Q(AW)-shot generates each of the questions in one inference call, followed by a final inference call for generating all the correct and incorrect answers in one go . It enjoys the same uniqueness properties as QAW-shot while having just two chained calls, making it both cheaper and less prone to generation failure cascading. Further, between Q(AW)-shot and QAW-shot, we observe Q(AW)-shot to have a higher generated quality, perhaps since LLM can jointly model while generating . We choose this to be our main method of choice for generating .
Prompt for imputing narrations into the LLM has a tremendous effect on the quality of generated . We experiment with several seed prompts for each of which we inspect the quality of the generated for clips. Based on this we iteratively improve the seed prompts manually in a zeroth order optimization fashion. In total, we experiment with a total of about prompts in this fashion to arrive at our final EgoSchema prompts – for generating questions and for generating all remaining options . While we fix the prompt, we use multiple prompts so as to avoid any unintended bias in the options. Fig. 5 shows an abridged example of and , full versions available in supplementary material.
Choice of LLM is extremely crucial for obtaining interesting long-form and generating hard negatives for . With weaker LLMs, the diversity across video clips remains narrow, and tends to be either obviously wrong or, too similar to and thus a false negative. While we experimented with both GPT-3 and ChatGPT but only found good quality generated at a high enough rate with GPT-4 , Bard , and Claude . For details please see supplementary.
We generate questions per three-minute clip as well as wrong answers to every question in addition to the correct answer. We observe that larger or tends to generate similar questions and wrong answers putting unnecessary pressure on Stages III and IV for filtering.
1.3 Stage III: Generated Question Answer Filtering
While Stage II produces several high-quality , even the best LLM generations are prone to output format aberrations, hallucinations, and sometimes plain false outputs. Further, despite specific pinpointed prompts (Fig. 5), LLMs can fail to comply. Since, we want to ensure EgoSchema to be extremely high-quality and accurate, we set up several filtering rounds to ensure the correctness and high difficulty of questions.
Rule-based filtering. Keywords from the prompts such as ‘long-term’, ‘narrations’, ‘timestamp’ etc. can sometimes bleed into the generated which are then discarded. The output generations can also fail to parse according to a specified format and are also then discarded and the concerned is regenerated.
LLM-based filtering. While rule-based filtering weeds out logic errors, we would like to further enrich before employing human labor. For example, we aim to ensure EgoSchema requires grounded visual reasoning to solve, and hence questions should not be answerable ungrounded, without carefully observing the video. Hence, we develop a "blind" baseline.
Blind filtering baseline employs LLM to guess the correct answer based on the question, without having access to the video narrations conditioned on the shown filtering prompt (Fig. 5). All such ungrounded questions that can be answered blindly are filtered out. This also ensures that generated are indeed relevant and plausible answers to , since otherwise, the LLM would be able to guess based only on the setting of . Note that this is overly restrictive since it is possible that a question is guessed correctly through chance and is not necessarily ungrounded. However, we choose to optimize precision over recall since the amount of filtered is still large enough.
No- baseline. We also experimented with a No- baseline, where the LLM is prompted to guess the correct answer using the narrations but without the question . This ensures that the wrong answers are relevant and plausible to the video clip. However, we found this baseline to have near random accuracy (), highlighting the efficacy of Stage II. Hence, we decided to not use this filter in the final pipeline. Additional details including the full prompt are in supplementary.
1.4 Stage IV: Manual 𝒬𝒜𝒲𝒬𝒜𝒲\mathcal{QAW} Curation
While LLM filtering ensures that the generated relates to the video content, it’s also necessary to ensure the veracity and a long temporal certificate length for every generated . This is achieved through a two-step manual curation process.
In the first round of curation, annotators are tasked with three primary responsibilities: (A) First, they verify that is well-formed and is indeed the correct answer to . (B) Next, they confirm that all the distractors, , are indeed wrong answers to . (C) Finally, they ensure that the temporal certificate length for answering is at least 30 seconds.
A is discarded if any of these three conditions are not met. This reduces the number of admissible questions by a factor of about to within the first round itself. Next is a second round of re-curation, to reinforce the conditions and guarantee data of the highest quality. We find that more than of the questions that pass the first round also pass the second round, speaking to the efficacy of the curation process. A crucial aspect of ensuring that the question assesses very long-form video-language understanding capabilities is the notion of temporal certificate length (condition (C) above), which we describe next. The detailed procedures for onboarding and training the human annotators, as well as the instructions for the curation process are provided in the supplementary.
2 Temporal Certificates
We define the temporal certificate of a given video in a video understanding task to be the minimum set of subclips of the video that are both necessary and sufficient to convince a human verifier that the marked annotation for that data (such as timestamps in temporal activity localization, class label in activity recognition or, the correct option in multiple-choice question-answering) is indeed correct, without having to watch the rest of the clip outside of the certificate set (Fig. 3). Naturally, we define certificate length to be the sum of the temporal lengths of the sub-clips present in the certificate set.
Meta-rules. Datasets often have implicit rules that apply uniformly across the entire dataset. We call these conventions meta-rules and allow the human verifier to be well aware of them. For example, in temporal action localization datasets , an implicit assumption is that the action to be localized in a contiguous sub-clip and hence can be uniquely determined by the start and end timestamps. Since this rule is valid for all data, we consider it to be a meta-rule.
A comprehensive understanding of meta-rules of a dataset is necessary for accurate estimation of the certificate set, and hence the certificate length. Otherwise, a spuriously long certificate might be necessary to ensure the veracity of the marked annotations. For example, consider the task of action classification on Kinetics-400. A valid meta-rule to be made available to the human verifier in this case is the mutual exclusivity of action classes i.e., each data point can belong only to one of the 400 classes present in Kinetics-400. Without this understanding, given, say a 10-second clip of a human skiing, the certificate set needs to necessarily encompass the entire 10 seconds since otherwise the human verifier might not be convinced that all of the other 399 actions are not occurring in the clip. However, with the knowledge of the label exclusivity meta-rule, the certificate length will be drastically reduced to just a fraction of a second since just observing the action of skiing in a few frames is sufficient for the human verifier to out-rule all other action classes.
Certificate Conventions. For small certificate lengths, it is difficult for humans to estimate the exact sub-clip timestamps to be included in the certificate set. Hence, we choose to have a minimum length of second for a certificate. Further, in the case of two non-contiguous certificates, we collapse them into one if their closest ends are seconds apart. In cases where a fact needs to be verified at several places throughout the video, we let the annotator make a reasonable judgment for the length of the certificate to be included as long as it follows the above conditions.
Benchmarking EgoSchema
Fig. 3 presents certificate lengths for a spectrum of tasks spread across different datasets such as, action classification (Kinetics , Something-Something , UCF101 , HVU-Action ), detection (AVA ), relationship classification (LVU ), concept classification (HVU-Concept ), video classification (Youtube-8M ), Question-Answering (NextQA , AGQA , NextQA , IVQA , MSRVTT , ActivityNet-QA , EgoSchema). For EgoSchema we benchmark the certificate length for 5 hours of video data () chosen randomly. For each other dataset, we ensure that (A) each annotated label class (if applicable) has at least 1 data sample evaluated and, (B) at least two hours of human effort is applied. Fig. 3 shows the histogram of estimated EgoSchema temporal certificate lengths for the 100 clips.
Fig. 3 plots the certificate length against the actual clip length. We observe that EgoSchema has temporal certificate length longer than the second longest certificate length dataset, and to longer than all other video understanding datasets.
2 Evaluating Multiple-choice Question Answering on EgoSchema
In Table 6, We benchmark several state-of-the-art video-language models, with the intention of adding more models in the future, in a Zero-shot question-answering setting on EgoSchema. We evaluate each model in at least two settings. First is the conventional inference setting, where the model is assessed based on the same number of frames it was trained with. And second is a less challenging setting, where the model is tested on the maximum number of frames possible to execute inference with, using an 80G A100, without exceeding the GPU memory capacity. In both settings, frames are sampled uniformly from the input video clip.
FrozenBiLM adapts frozen multi-modal encoders trained on web-scale data for the task of question answering and achieves state-of-the-art zero-shot QA accuracy across video question-answering datasets. We choose the How2QA FrozenBilM model under both and frames.
VIOLET a masked token modeling-based video language transformer that performs competitively on a variety of video-language tasks. We evaluate four of the best VIOLET models that are finetuned on different tasks for both and frames and choose the model with the best overall accuracy. More details are in supplementary.
mPLUG-Owl proposes a training strategy to add image & video modality to pretrained large language models. We adapt mPLUG to facilitate the multiple choice QA by prompting the model with each of the options individually in the format: ‘Given question
InternVideo proposes training video-language models jointly with masked video modeling and contrastive learning objectives. By default, InternVideo does not directly support multiple-choice video QA. We adapt the MSRVTT finetuned InternVideo model, which performs zero-shot multiple-choice tasks, by incorporating the question with each answer choice in the format: ’Question:
Human. We also benchmark human performance on multiple-choice question answering task on EgoSchema in Table 7. First, are time pressure settings where the annotators are asked to choose the correct answer under one (‘In <1 min’) and three (‘In <3 min’) minutes. Humans can already achieve an impressive 67.0% accuracy, in under 1 minute! Interestingly, this only slightly increases (+1.0%) when allowed three minutes. We believe that this can inform about performance on EgoSchema in limited model inference capacities. We believe this could inform about the frame rate needed for long-form video understanding in future models. Second, we also benchmark human performance using only 1 fps video (‘180 frames’). Surprisingly, we observe that just with 1 fps humans can achieve an impressive 67.2%.
Third, we evaluate human performance in a restrictive setting where the annotator is forced to first watch the video without reading the text, and then answer the question without re-watching the video (‘Video Text’). Curiously, this achieves better accuracy than the ‘No constraint’ setting where the annotators are asked to simply answer without any constraints (76.2% vs. 75.0%). A possible hypothesis is that watching the video without text allows the annotator to focus more closely on the video, thereby benefiting performance than the setting where the attention is somewhat divided between the text and video. We believe this will help us understand the performance trade-offs in the early vs. late fusion of video and text modalities for long-form video-language models. Accuracy for ‘No constraint’ setting is estimated over 9 hours of video. All other accuracies are estimated over 5 hours of video.
Conclusion
We present EgoSchema, a novel diagnostic benchmark designed for assessing very long-form video-language understanding capabilities of modern multimodal models. We also introduce the notion of a temporal certificate set, a probe that can be applied to a wide array of video tasks and benchmarks for understanding their intrinsic temporal lengths. We estimate temporal certificates of 15 varied datasets and demonstrate EgoSchema to exhibit temporal certificate length approximately longer than the next longest dataset and to longer than all other video understanding datasets. We also benchmark several state-of-the-art models on EgoSchema and find their Zero-shot question-answering accuracy to be less than while humans achieve 76%. We believe that EgoSchema will play a key role in the development and evaluation of future very long-form video-language models.
Limitations. EgoSchema RGB clips are sourced from Ego4D and inherit Ego4D egocentric video biases. Further, the text is carefully curated for veracity, there are inevitable text data distribution biases that can occur in LLM-generated outputs due to biases present in web-scale LLM training data. Finally, human curation itself is far from perfect and while we perform two rounds of curation to minimize false positives, the collected EgoSchema is most likely to inevitably contain some small mislabelled or ill-formed question-answer sets. We plan to host a crowd-sourced errata board to minimize human curation error over time with the support of the open-source research community.
References
EgoSchema Datasheet
For what purpose was the dataset created? Was there a specific task in mind? Was there a specific gap that needed to be filled? Please provide a description.
EgoSchema is a diagnostic benchmark for assessing very long-form video-language understanding capabilities of modern multimodal systems. While some prior works have proposed video datasets with long clip lengths, we posit that merely the length of the video clip does not truly capture the temporal difficulty of the video task that is being considered. To remedy this, we introduce temporal certificate sets, a general notion for capturing the intrinsic temporal understanding length associated with a broad range of video understanding tasks & datasets. Please see Section 3.2 in the main paper for more details.
Who created this dataset (e.g., which team, research group) and on behalf of which entity (e.g., company, institution, organization)?
The authors created the dataset within the Malik Group at Berkeley AI Research, UC Berkeley. The authors created it for the public at large without reference to any particular organization or institution.
What do the instances that comprise the dataset represent (e.g., documents, photos, people, countries)? Are there multiple types of instances (e.g., movies, users, and ratings; people and interactions between them; nodes and edges)? Please provide a description.
Each instance in the dataset represents a 3-minute video and text that contains a question and five answer options.
How many instances are there in total (of each type, if appropriate)?
EgoSchema has a total of 5063 instances each containing one video, one question, and five answer options. You can see further statistics on the whole data on our website egoschema.github.io.
Does the dataset contain all possible instances or is it a sample (not necessarily random) of instances from a larger set? If the dataset is a sample, then what is the larger set? Is the sample representative of the larger set (e.g., geographic coverage)? If so, please describe how this representativeness was validated/verified. If it is not representative of the larger set, please describe why not (e.g., to cover a more diverse range of instances, because instances were withheld or unavailable).
The video component of our dataset derives from the broader Ego4D dataset. For our research, we selectively extracted non-overlapping three-minute segments from the Ego4D video data, each segment consisting of a minimum of 30 human-annotated narrations (where each narration refers to a timestamped sentence). Detailed statistic of the number of viable clips for different possible length and narration density choices is discussed in Supplementary Section 6. The selected subset is very diverse in human behavior as can be seen by the activity statistics presented on egoschema.github.io.
What data does each instance consist of? “Raw” data (e.g., unprocessed text or images) or features? In either case, please provide a description.
Each instance in our dataset comprises raw mp4 video data, captured at a rate of 30 frames per second and with a high resolution. Accompanying this video data, there are six text elements - one question and five corresponding answer options one of which is marked as the correct answer to the question.
Is there a label or target associated with each instance? If so, please provide a description.
Each instance is associated with a label ranging from 1 to 5 that indicates which of the five answer options is correct.
Is any information missing from individual instances? If so, please provide a description, explaining why this information is missing (e.g. because it was unavailable). This does not include intentionally removed information but might include, e.g., redacted text.
Are relationships between individual instances made explicit (e.g., users’ movie ratings, social network links)? If so, please describe how these relationships are made explicit.
Some instances may have the same video but different questions and answers. It will be indicated by a clip unique identifier in the final dataset.
Are there recommended data splits (e.g., training, development/validation, testing)? If so, please provide a description of these splits, explaining the rationale behind them.
EgoSchema is designed specifically for zero-shot testing. Its primary purpose is to be able to asses the out of the box long-term video-language understanding capabilities of modern multimodal models.
Are there any errors, sources of noise, or redundancies in the dataset? If so, please provide a description.
The dataset was very carefully manually curated to mitigate any incidence of errors within the questions and answers. Although different questions may be posed for the same clip, it is ensured that there is no overlap between any two distinct clips. Further related details are also discussed in the limitations section in the main paper.
Is the dataset self-contained, or does it link to or otherwise rely on external resources (e.g., websites, tweets, other datasets)? If it links to or relies on external resources, a) are there guarantees that they will exist, and remain constant, over time; b) are there official archival versions of the complete dataset (i.e., including the external resources as they existed at the time the dataset was created); c) are there any restrictions (e.g., licenses, fees) associated with any of the external resources that might apply to a future user? Please provide descriptions of all external resources and any restrictions associated with them, as well as links or other access points, as appropriate.
Entirety of the dataset will be made publicly available at our project website egoschema.github.io. We will also provide a download tool for preprocessing all the videos such as cutting clips, associating the question/answer text etc. Text will be released in a JSON format, hosted on our github repository. EgoSchema will be publicly released under the Ego4D license, which allows public use of the video and text data for both research and commercial purposes.
Does the dataset contain data that might be considered confidential (e.g., data that is protected by legal privilege or by doctor-patient confidentiality, data that includes the content of individuals non-public communications)? If so, please provide a description.
Does the dataset contain data that, if viewed directly, might be offensive, insulting, threatening, or might otherwise cause anxiety? If so, please describe why.
Does the dataset relate to people? If not, you may skip the remaining questions in this section.
Some videos do contain people. However, the Ego4D authors employed an array of de-identification procedures primarily centered on ensuring a controlled environment with informed consent from all participants, and, where applicable, in public spaces with faces and other personally identifiable information suitably obscured. We strictly import all RGB information from Ego4D without any addition of our own.
Does the dataset identify any subpopulations (e.g., by age, gender)? If so, please describe how these subpopulations are identified and provide a description of their respective distributions within the dataset.
Is it possible to identify individuals (i.e., one or more natural persons), either directly or indirectly (i.e., in combination with other data) from the dataset? If so, please describe how.
No, Ego4D has employed an array of deidentification procedures in order to obscure any personally identifiable information such as people’s faces.
Does the dataset contain data that might be considered sensitive in any way (e.g., data that reveals racial or ethnic origins, sexual orientations, religious beliefs, political opinions or union memberships, or locations; financial or health data; biometric or genetic data; forms of government identification, such as social security numbers; criminal history)? If so, please provide a description.
How was the data associated with each instance acquired? Was the data directly observable (e.g., raw text, movie ratings), reported by subjects (e.g., survey responses), or indirectly inferred/derived from other data (e.g., part-of-speech tags, model-based guesses for age or language)? If data was reported by subjects or indirectly inferred/derived from other data, was the data validated/verified? If so, please describe how.
The video data, which is directly observable, was procured from the publicly accessible Ego4D dataset. In contrast, the text data was generated through the use of Large Language Models (LLMs) including GPT4, BARD, and Claude. These LLMs employed visual narrations from each video within the Ego4D dataset to generate the corresponding text.
What mechanisms or procedures were used to collect the data (e.g., hardware apparatus or sensor, manual human curation, software program, software API)? How were these mechanisms or procedures validated?
The video and narration data were downloaded in accordance with the official Ego4D guidelines for data access: https://ego4d-data.org/docs/start-here. For the generation of the text data within our dataset, we utilized API access for GPT4 via OpenAI, for BARD via Google, and for Claude via Anthropic. This allowed us to generate three distinct questions for each video clip sampled from the Ego4D dataset. Upon the generation of these questions for each sampled video clip, we implemented a series of filtering procedures including Rule-based filtering, Blind filtering, and Manual curation. See Section 3.1.2 in the main paper for a more detailed explanation.
If the dataset is a sample from a larger set, what was the sampling strategy (e.g., deterministic, probabilistic with specific sampling probabilities)?
The video component of our dataset derives from the broader Ego4D dataset. For our research, we selectively extracted non-overlapping three-minute segments from the Ego4D video data, each segment consisting of a minimum of 30 human-annotated narrations (where each narration refers to a timestamped sentence). Detailed statistic of the number of viable clips for different possible length and narration density choices is discussed in Supplementary Section 6.
Who was involved in the data collection process (e.g., students, crowdworkers, contractors) and how were they compensated (e.g., how much were crowdworkers paid)?
Our research utilized the services of Quantigo, a specialized data labelling company. The teams of Quantigo employees that were based in Bangladesh were compensated at a rate of 5 dollars per hour, at a wage significantly higher than the market hourly rate in Bangladesh. This was done to ensure fair compensation for the complex tasks performed while also contributing to the highest quality of the work delivered. It’s important to note that our collaboration with Quantigo followed ethical guidelines, with the fair treatment of all employees involved and the appropriate respect for their expertise and labor. For exact instructions for human curation, see Supplementary Section 7.
Over what timeframe was the data collected? Does this timeframe match the creation timeframe of the data associated with the instances (e.g., recent crawl of old news articles)? If not, please describe the timeframe in which the data associated with the instances was created.
The original videos within the Ego4D dataset were collected across various occasions spanning from 2019 to 2021. As for the EgoSchema, the textual information was collected over several sprints during the first half of 2023 based on the Ego4D narrations.
Were any ethical review processes conducted (e.g., by an institutional review board)? If so, please provide a description of these review processes, including the outcomes, as well as a link or other access point to any supporting documentation.
Does the dataset relate to people? If not, you may skip the remaining questions in this section.
Did you collect the data from the individuals in question directly, or obtain it via third parties or other sources (e.g., websites)?
The video and narration data were acquired in accordance with the official Ego4D guidelines for data access: https://ego4d-data.org/docs/start-here/. The Ego4D authors had in turn ensured consent of the people involved.
Were the individuals in question notified about the data collection? If so, please describe (or show with screenshots or other information) how notice was provided, and provide a link or other access point to, or otherwise reproduce, the exact language of the notification itself.
Ego4d paper followed several procedures to ensure the preservation of privacy and the upholding of ethical standards. Notably, these procedures included obtaining informed consent from those wearing the cameras and adhering to de-identification requirements for personally identifiable information (PII). Given that the video collection was conducted by Ego4D, we are not in a position to provide specific instructions that were given to the camera wearers. The Ego4D privacy statement is available at https://ego4d-data.org/pdfs/Ego4D-Privacy-and-ethics-consortium-statement.pdf
Did the individuals in question consent to the collection and use of their data? If so, please describe (or show with screenshots or other information) how consent was requested and provided, and provide a link or other access point to, or otherwise reproduce, the exact language to which the individuals consented.
Ego4d paper privacy procedures have included obtaining informed consent from those wearing the cameras. Given that the video collection was conducted by Ego4D, we are not in a position to provide specific instructions that were given to the camera wearers. See Ego4D privacy statement.
If consent was obtained, were the consenting individuals provided with a mechanism to revoke their consent in the future or for certain uses? If so, please provide a description, as well as a link or other access point to the mechanism (if appropriate).
Ego4d paper privacy procedures have included allowing camera users to ask questions and withdraw at any time. Additionally, they were free to review and redact their own video. Given that the video collection was conducted by Ego4D, we are not in a position to provide specific instructions that were given to the camera wearers. You can find the Ego4D privacy statement at https://ego4d-data.org/pdfs/Ego4D-Privacy-and-ethics-consortium-statement.pdf.
Has an analysis of the potential impact of the dataset and its use on data subjects (e.g., a data protection impact analysis) been conducted? If so, please provide a description of this analysis, including the outcomes, as well as a link or other access point to any supporting documentation.
While we recognize the importance of this topic, we would, once more, refer to the Ego4D paper for an in-depth discussion. Ego4D acknowledges the potential privacy risks associated with the use of wearable devices in data collection and has taken several steps such as depersonalizing any sensitive information, blurring out faces and bodies, etc. towards maintaining privacy. The same carries over to the video data in EgoSchema as well. Broadly, very long-form video understanding is a core capability for agents that are to perceive the natural visual world. Hence, developing datasets such as EgoSchema will be critical to unlocking this key AI capability. Additionally, according to Ego4D privacy statement, all videos from Ego4D were reviewed by an approved member of one of the participant’s universities or institutes to identify and assess potential privacy concerns.
Was any preprocessing/cleaning/labeling of the data done (e.g., discretization or bucketing, tokenization, part-of-speech tagging, SIFT feature extraction, removal of instances, processing of missing values)? If so, please provide a description. If not, you may skip the remainder of the questions in this section.
The set of generated questions and answers from output was filtered by those LLMs and finally curated by humans. A detailed description can be found in Section 3. There was no preprocessing done on the video clips sampled from Ego4D.
Was the “raw” data saved in addition to the preprocessed/cleaned/labeled data (e.g., to support unanticipated future uses)? If so, please provide a link or other access point to the “raw” data.
Human curation was employed to rectify errors in the question-answer sets, particularly cases where the identified correct answer was wrong or a wrong answer was actually correct. Given the crucial role of this step in ensuring the accuracy of our dataset, we do not find it necessary to release a version of the dataset prior to human curation. However, all the discarded "raw" data is indeed also saved.
Is the software used to preprocess/clean/label the instances available? If so, please provide a link or other access point.
The APIs for the Large Language Models (LLMs) are publicly accessible. The prompts for filtering and instructions for human curation are provided in Supplementary Section Full Prompts and Supplementary Section 7 respectively. Additionally all necessary code for generation, filtering etc. is provided in the supplementary materials.
Will the dataset be distributed to third parties outside of the entity (e.g., company, institution, organization) on behalf of which the dataset was created? If so, please provide a description.
The dataset will be made publicly available and can be used for both research and commercial purposes under the Ego4D license.
How will the dataset be distributed (e.g., tarball on website, API, GitHub) Does the dataset have a digital object identifier (DOI)?
The dataset will be distributed as a JSON file describing the unique identifier for each clip, the associated question, the five answer options, the label, and additional clip information that facilitates the tracing of the clip back to the original Ego4D data, such as the Ego4D video identification of the clip’s source video, among other details. In addition, download tools to acquire and pre-process the video RGB data will also be provided on our website.
The full dataset will be made available upon the acceptance of the paper before the camera-ready deadline.
Will the dataset be distributed under a copyright or other intellectual property (IP) license, and/or under applicable terms of use (ToU)? If so, please describe this license and/or ToU, and provide a link or other access point to, or otherwise reproduce, any relevant licensing terms or ToU, as well as any fees associated with these restrictions.
EgoSchema will be publicly released under the Ego4D license, which allows direct public use of the video and text data for both research and commercial purposes.
Have any third parties imposed IP-based or other restrictions on the data associated with the instances? If so, please describe these restrictions, and provide a link or other access point to, or otherwise reproduce, any relevant licensing terms, as well as any fees associated with these restrictions.
Do any export controls or other regulatory restrictions apply to the dataset or to individual instances? If so, please describe these restrictions, and provide a link or other access point to, or otherwise reproduce, any supporting documentation.
Who will be supporting/hosting/maintaining the dataset?
The authors of the paper will be maintaining the dataset, pointers to which will be hosted on github repo https://github.com/egoschema/EgoSchema along with the code for download and preprocessing tool, with the actual data hosted either on Amazon AWS as an S3 bucket or as a google drive folder.
How can the owner/curator/manager of the dataset be contacted (e.g., email address)?
We will post the contact information on our website. We will be available through github issues as well as through email.
Is there an erratum? If so, please provide a link or other access point.
We will host an erratum on the Github repo in the future, to host any approved errata suggested by the authors or the video research community.
Will the dataset be updated (e.g., to correct labeling errors, add new instances, delete instances)? If so, please describe how often, by whom, and how updates will be communicated to users (e.g., mailing list, GitHub)?
Yes, we plan to host an erratum publicly. There are no specific plans for a v2 version, but there does seem plenty oppurtunities for exciting future dataset work based on EgoSchema.
If the dataset relates to people, are there applicable limits on the retention of the data associated with the instances (e.g., were individuals in question told that their data would be retained for a fixed period of time and then deleted)? If so, please describe these limits and explain how they will be enforced.
Will older versions of the dataset continue to be supported/hosted/maintained? If so, please describe how. If not, please describe how its obsolescence will be communicated to users.
N/A There are no older versions at the current moment. All updates regarding the current version will be communicated via our website.
If others want to extend/augment/build on/contribute to the dataset, is there a mechanism for them to do so? If so, please provide a description. Will these contributions be validated/verified? If so, please describe how. If not, why not? Is there a process for communicating/distributing these contributions to other users? If so, please provide a description.
Contributions will be made possible using standard open-source tools, submitted as pull requests to the relevant GitHub repository. Moreover, we will provide information on how to trace sampled clips back to their original source within the Ego4D dataset. This will enable users to access additional Ego4D data, such as narrations, summaries, and object detections, as applicable.
Full Prompts
Here are some of the prompts we developed for generating EgoSchema.
1.2 Answer prompt
2 Set B
2.2 Wrong answer prompt
Our clip length and narration density choice
Human curation
Our research utilized the services of a third part company (not MTurk), for specifically training annotators to ensure quality. The process involved two distinct annotation procedures: data curation and human accuracy testing.
Generated data curation was performed by Quantigo employees. These curators were responsible for ensuring that the released EgoSchema dataset is tge highest high quality possible. Here is the exact instructions that was provided to annotators:
Benchmarking details
Violet is a video language model comprised of a visual encoder, text encoder, and multimodal transformer pretrained on a variety of masked visual modeling tasks ranging from simple ones such as RBG pixel values up to more high levels ones such as spatially focussed image features. It performs competitively on a variety of video-language tasks such as Video-QA and Video-Text Retrieval. We evaluate one pre-trained model and 3 models finetuned on lsmdc-mc, msrvtt-qa, and msrvtt-retrieval. We evaluate using both 5 frames and 75 frames and choose the model with the best overall accuracy.
2 mPLUG-Owl
By default, mPLUG-Owl does not possess inherent capabilities for direct video question answering. As such, we undertook several experiments to adapt it to our required format. One approach involved inputting all answer choices in the form of a shuffled test. However, this resulted in a bias towards selecting the first option in most cases. For another approach, FrozenBiLM offered a methodology for frozen zero-shot models to operate in the context of multiple-choice video question answering, which inspired us to adapt this methodology for mPLUG-Owl. As mPLUG-Owl utilizes word-level tokenization, we could extract the confidence score for each generated token, particularly the ’Yes’ token. We recorded the ’Yes’ token confidence score for each answer option. In instances where the ’Yes’ token was absent, we assigned the confidence score as zero, though empirically, in most cases, the model output was positive and contained the ’Yes’ token. Ultimately, we selected the answer option with the highest ’Yes’ confidence score as the model output given the question. In scenarios where multiple options scored the same highest confidence for the ’Yes’ token, we randomly selected the answer from these top-scoring options. It should be noted that mPlug-Owl was originally trained to process a single image, and its capacity to handle additional frames is an emergent ability that has not been thoroughly tested to date."
3 InternVideo
The two most closely aligned formats supported by InternVideo are open-ended Video Question Answering and Zero-shot Multiple Choice tasks. In the case of open-ended Video Question Answering, the task is to predict the answer to a question posed within a video. However, due to the restricted vocabulary of open-ended answers in open-ended Video Questions Answering, we decided to formulate EgoSchema within the context of a Zero-shot Multiple Choice task. This task aims to identify the correct answer from a set of given options, without the inclusion of a question. InternVideo has provided finetuned weights for two datasets: MSRVTT and LSMDC. We selected the model finetuned on MSRVTT because it shares greater contextual similarity with EgoSchema.
4 Human
To conduct human benchmarking, we engaged a distinct team of ten employees within the same data annotation company to carry out human benchmarking on our dataset. The answers were randomized and presented in the form of a test. The following are the precise instructions provided to the annotators: