VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models

Shicheng Li, Lei Li, Shuhuai Ren, Yuanxin Liu, Yi Liu, Rundong Gao, Xu Sun, Lu Hou

Introduction

Many important concepts in human languages contain a temporal dimension , such as human actions, changes in status, and event order, which are beyond the expressive power of individual static images. Such temporal concepts bring great challenges to video-language learning and are crucial for the generalization capability of intelligent systems in real-life scenarios.

Although these temporal concepts are present in existing text-to-video retrieval or video question answering benchmarks, most of these datasets fail to faithfully assess the temporal understanding ability of Video-Language Models (VidLMs) due to the strong correlation between static objects/scenes and temporal information. For example, in the blue box in Fig. 1, each video can be aligned to its description by merely identifying the static objects such as the fire, the microphone, and the PC case. As a consequence, the models may learn to simply rely on static clues to make predictions, leading to failure in real-world applications that require a genuine understanding of temporal concepts, e.g., to distinguish between the action of “connecting something to system” and “disconnecting something from system” as demonstrated by the red box in Fig. 1. Previous works have pointed out similar issues and provided several solutions. However, they do not properly define and categorize different aspects of temporal information. The lack of a clear definition adds to the difficulty of assessing the precise abilities of VidLMs. Additionally, they often construct evaluation datasets by following certain templates or using synthetic scenes, making them unsuitable for more diverse and realistic scenarios.

In light of the drawbacks of current video-language testbeds, we propose a new dataset for VidLMs, VITATECS, to fill the gap for temporal concept understanding evaluation by decoupling temporal information and static information. Inspired by Winoground , to measure the ability of VidLMs to understand and align the temporal concepts, we ask the models to distinguish between the correct caption of a video and a modified version of the caption which contains similar static information and only differs in temporal information. To allow for a more comprehensive and fine-grained evaluation of temporal understanding ability, we summarize several aspects of temporal concepts that are commonly present in video descriptions, including Direction, Intensity, Sequence, Localization, Compositionality and Type, which according to our study cover most of the temporal information in video-language datasets.

Since collecting high-quality video-text pairs is time-consuming and expensive, we follow previous works in dataset construction , and augment existing open-domain video-language datasets by harnessing the world knowledge encoded in pre-trained large language models (LLMs) . Specifically, given an annotated video-text pair in the dataset, we ask the LLM to generate a counterfactual description that only differs from the original description in one given temporal aspect using in-context learning . To prevent potential mismatch when dealing with complex instructions, we design a human-in-the-loop procedure to filter out low-quality generations by iteratively generating counterfactual descriptions, human labeling, and fine-tuning a filter model. In each iteration, the generated samples are used to update the filter model and the in-context learning exemplar set to boost generation and filtering quality. This annotation framework allows us to construct a 13k+ dataset from 231 human-written counterfactuals while maintaining high quality and diversity.

Based on our dataset, we conduct a comprehensive evaluation of state-of-the-art video-language understanding models. Our findings can be summarized as follows.

Existing models barely surpass random guesses in many aspects, confirming their general lack of temporal understanding.

Temporally-adapted image-text models outperform video-text pre-training, but primarily due to better utilization of static clues.

Failure of text encoders to learn temporal concepts during pre-training is partly responsible for low performance on temporal understanding.

Different video-text datasets tend to invoke different temporal understanding abilities.

In summary, our work with VITATECS sheds light on limitations in current VidLMs’ temporal understanding, providing insights for future development.

Related Work

With the great success of end-to-end deep learning models in natural language processing and image-text understanding, the research community has shown a growing interest in the more challenging task of video-language understanding, with promising results achieved on a wide range of tasks including video captioning , video question answering and video-text retrieval . Some research follows the prevalent paradigm in NLP and multi-modal understanding by directly conducting pre-training on video-text pairs. Another line of work adapts powerful image-text pre-trained models like CLIP to transfer their knowledge to the video-language domain. Recent studies have also explored the possibility of integrating LLMs with vision encoders to perform video-language understanding tasks. Despite these valuable efforts, we argue that the apparent prosperity of video-language understanding models still rests upon the power of image-language models and that more attention should be paid to their temporal understanding abilities.

Datasets on Temporal Understanding.

Although the temporal dimension is the primary difference between videos and images, it has not received proper acknowledgment from current model design and dataset construction processes in the video-language community. Previous works have pointed out the lack of emphasis on temporal understanding abilities. Evidence of this negligence includes the insensitivity of models to the frame order of input videos , several order-agnostic architecture designs with state-of-the-art retrieval performance , visualization of intermediate layer features or saliency maps , and even the success of using single frames training to achieve promising results . Although a few datasets have been proposed to address this issue, they do not properly define and categorize different aspects of temporal information in video-language understanding. In addition, their videos are constructed from either computer-rendered synthetic videos or videos focusing on single human actions , which are not representative of real-world videos. We remedy these problems by identifying aspects of temporality and introducing a new dataset for measuring the temporal understanding abilities in VidLMs. See Tab. 1 for a comparison between VITATECS and existing video datasets.

VITATECS: Diagnosing Temporal Concept Understanding

In this section, we propose VITATECS, a new dataset for measuring how well VidLMs capture temporal information across modalities. It consists of (video, caption, counterfactual) triples, where the counterfactual description retains the same static information as the original caption while modifying its temporal information in one of the six fine-grained aspects that we define in Sec. 3.1. We elaborate on the details of our temporal dataset in Sec. 3.2 and the human-in-the-loop annotation framework we devise to facilitate its construction process in Sec. 3.3.

Measuring the temporal understanding ability of VidLMs is a challenging task. On one hand, it is not clear how to define and characterize the temporal information in a video. Previous works draw a rough equivalence between temporal information and the actions in the video. In reality, temporal information can emerge in a variety of forms, such as human actions, changes in object status, dynamics of substances, the order of events, etc., and is widely manifested in daily activities. On the other hand, it is infeasible to completely disentangle the temporal information from the static information. The background scenes, objects, and people’s postures are all highly correlated with the temporal information in open-domain videos. If not properly controlled, such static bias would allow models to rely on static clues as shortcuts for making predictions while seemingly learning to capture the temporal information.

To achieve high coverage of temporal information in video-language datasets and allow for fine-grained diagnosis of temporal understanding abilities, we identify six aspects of temporal concepts commonly reflected in natural language: Direction, Intensity, Sequence, Localization, Compositionality and Type. These aspects of temporal information are disentangled from static information to different degrees and address different facets of the temporal information in video-language datasets, allowing us to pinpoint the temporal understanding abilities of VidLMs. Since our final target is to construct text pairs with aspect-specific modifications, for clarity, we define these aspects in terms of the temporal questions they address and the corresponding modification patterns as follows.

“Direction” measures the model’s ability to answer the following question: “In which direction does the status of objects change?” Examples of this aspect include sentence pairs describing opposite spatial movements or one action reversing the effect of the other.

“Intensity” measures the model’s ability to answer the following question: “How fast or how intense does the change occur?” Examples of this aspect include counterfactual sentences which change the words that modify the verbs or change the verb to a similar action with subtle differences in the manner it is conducted.

“Sequence” measures the model’s ability to answer the following question: “How many events are depicted in the video and in what order?” Examples of this aspect usually involve changing the temporal order or number of occurrences of the events.

“Localization” measures the model’s ability to answer the following question: “On which part of the frame does the change occur?” Examples of this aspect include sentence pairs with the same action conducted either in different absolute spatial locations or in different locations in relation to other objects in the video.

“Compositionality” measures the model’s ability to answer the following question: “Who performed which action and to whom?” Examples of this aspect often include actions with interchanged subjects or objects.

“Type” measures the model’s ability to answer the following question: “What is the action depicted in the video?” This aspect contains general alterations to the actions with a less stringent constraint on the static information contained.

To validate the coverage of our temporal concept categorization, we randomly sample 200 video-text pairs from MSR-VTT and VATEX and inspect the types of temporal information they contain. We find that for 98% of the samples, their temporal information falls in one of our categories, which demonstrates that our taxonomy is able to achieve high coverage while taking into account the disentanglement from static information.

2 Dataset Format

3 Human-in-the-Loop Annotation Framework

Due to the heavy expenses of collecting high-quality (video, caption, counterfactual) triples, we present a human-in-the-loop annotation framework for semi-automatic counterfactual generation based on existing (video, caption) datasets. At the core of our framework is a loop consisting of three stages: generation, filtering, and revision. In stage 1, we use in-context learning to generate candidate counterfactuals based on ground-truth video-text pairs with LLMs. In stage 2, the candidates are filtered using a combination of rules, off-the-shelf language understanding models, and fine-tuned language understanding models. In stage 3, we ask human annotators to verify the quality of the candidates and use the high-quality ones to refine the generation process and the filter model. The three stages are conducted on a small subset and are repeated until the filter model achieves satisfactory precision on a held-out evaluation set. Below, we first lay out the criteria for our counterfactual descriptions and then elaborate on the details of each stage.

During our effort to construct the dataset, we found that LLMs encounter some difficulty in following our instructions when generating counterfactual descriptions, possibly due to the reflective nature of our temporal concepts. To enable consistent and high-quality counterfactual generation, we first identify five major criteria for measuring the quality of generated counterfactuals as follows.

The counterfactual should neither entail nor be entailed by the caption.

The counterfactual should contain roughly the same amount of information as the caption.

The counterfactual should be grammatically correct and semantically plausible.

The counterfactual should retain the static information in the caption and only change the given aspect of temporal information.

The pattern of counterfactual description should be diverse across the entire dataset.

Among these desirable properties, criteria (a)-(d) are instance-level criteria we aim to address in both the generation and filtering stages. In contrast, criterion (e) is a dataset-level criterion dealt with in a finalization step after the filter model has converged.

Exemplar Sets.

Throughout our annotation process, we maintain three sets of exemplars: positive set X+\mathcal{X}^{+} contains sentence pairs that differ only in a given aspect; negative set X−\mathcal{X}^{-} contains sentence pairs that violate one of the aforementioned criteria (a)-(d); N/A set XNA\mathcal{X}^{NA} contains captions that do not describe a certain aspect of the temporal concept. These exemplars serve two purposes: on the one hand, they compose the demonstrations of valid and invalid data samples for in-context learning, which supply the generative language models with clearer and better-informed instructions; on the other hand, they provide supervision signals for the fine-tuning of the filter model. These three sets are initialized with manually annotated examples and expanded semi-automatically to boost the generation and filter model performance as more data samples are generated.

In-Context Learning Generation.

In this stage, we draw upon the generative strength of ChatGPT (gpt-3.5-turbo-0613) to generate counterfactual descriptions given the original caption and the desired aspect of variation. The use of in-context learning allows us to capture the different aspects of temporal concepts through carefully-designed instructions and demonstrations. Specifically, we first randomly sample a small subset (500 for each aspect) of (video, caption) pairs from the test sets of two popular video-text retrieval datasets, MSR-VTT and VATEX . Then, for each (video, caption) pair, we invoke the instruction following and pattern replication abilities of ChatGPT by constructing a prompt consisting of an aspect-specific instruction, demonstrations sampled from the exemplar sets, and the query for which we aim to generate the counterfactual description. The demonstrations are sampled from both X+\mathcal{X}^{+} and XNA\mathcal{X}^{NA} so that the LLM not only learns to generate counterfactual descriptions for valid captions but also learns to recognize which captions do not concern the temporal aspect of interest.

Automatic Filtering.

In view of the uneven quality of generated examples, we propose to filter the candidates and automatize this procedure using natural language understanding models. First, we leverage an off-the-shelf natural language inference (NLI) model, Sentence-BERT , to filter out examples that do not meet criterion (a), i.e., the cases where one description entails the other. Then, to filter out candidates that do not meet criterion (b)-(d), we use a neural network that takes a pair of sentences as input and performs a 7-way classification task, where category 0 corresponds to disqualified generations and categories 1-6 correspond to the six aspects we define. Considering the similarity in task formulation, we initialize the filter model with the same NLI model above. The fine-tuning data consists of samples from both X+\mathcal{X}^{+} and X−\mathcal{X}^{-}. We adopt a rigorous decision mechanism that classifies the given sentence pair into one of the six aspects only if the model makes consistent predictions for the pair and its reversed version with high confidence, as we care more about the precision of the filter model than its recall.

Human Revision.

To guarantee the quality of filtered examples and guide both the in-context learning procedure and the filter model in the right direction, we introduce human supervision to revise the filtering results. We manually check the samples that are predicted to fall in one of the six aspects and correct the wrong predictions. Note that, on the one hand, due to the relatively small size of the sampled subset and the rigorous confidence-based filtering procedure, the number of examples for human revision is reduced significantly; on the other hand, human annotators only need to rectify the predicted labels instead of writing the entire counterfactual description. Therefore, this revision stage does not require excessive human effort and only incurs acceptable annotation costs.

Iterative Procedure.

We repeat the generation, filtering, and revision procedure to iteratively enlarge the exemplar sets and refine the filter model. In each iteration, the previously revised examples are incorporated into X+\mathcal{X}^{+} and X−\mathcal{X}^{-} according to their labels. This simultaneously augments the demonstration set of in-context learning for better generation quality and provides more training data for fine-tuning the filter model. After each iteration, the fine-tuned filter model is evaluated on an independently annotated test set. We terminate the iteration once no significant improvement of the filter model is observed.

Finalization.

After the filter model has converged, we perform generation and filtering on a larger scale (20,000 for each aspect) without human revision. As a finalization step, we address the issue of diversity by favoring generations that involve a less common change of verb throughout the dataset when merging the filtered samples.

Annotation Efficiency of the Framework.

Our framework can be easily scaled to generate larger datasets since no more human efforts are required once the filter model has converged. In our case, it only takes 231 human-written descriptions and around 1500 labeling annotations to obtain the final benchmark with 13k+ samples, showing the efficiency of our annotation framework. The statistics of our dataset are shown in Tab. 2.

Evaluation of Video-Language Models

In this section, we evaluate prevailing VidLMs to examine their temporal understanding ability. We first introduce the evaluation settings and then discuss the findings drawn from our evaluation to facilitate future studies.

In our experiments, we focus on models designed for the video-text retrieval task, which can calculate the similarity score between a video and a text query. We test three pre-trained VidLMs (VIOLET , ALPRO and Singularity ) and three temporally-adapted image-language models (CLIP4Clip , X-Pool and X-CLIP ). We also include two recent video large language models, VideoLLaMA and VideoChat , as well as pure image-text foundation models such as BLIP , which has shown strong performance on zero-shot video-text retrieval.

Evaluation Metric.

A model’s prediction is considered correct if the similarity score of the correct caption is higher than that of the generated counterfactual. We measure the accuracy of the models on each of the six aspects of temporal concepts, and explore a recall-based metric in Sec. 4.3.

Human Baseline.

We randomly choose 100 samples for each aspect from our dataset and ask five volunteers to help establish a human performance baseline. The annotators are shown a video and two text descriptions at a time and are required to choose the text that best describes the video. We report the average accuracy of the five annotators as the human baseline.

2 Evaluation Results

As shown in Tab. 3, although humans can easily match the videos to their correct descriptions with high consistency (κ=0.86\kappa=0.86) and nearly no mistakes, the overall performance of all the evaluated models is still far from expectations. No model achieves an accuracy of over 70% on the temporal aspects other than the relatively easy “Type” aspect, which has the strongest correlation with the static information. Particularly, on the more temporally demanding aspects (“Direction”, “Intensity”, and “Sequence”), the models perform barely over the random baseline (50%). Considering that part of our videos directly comes from MSR-VTT, the poor performance of models fine-tuned on MSR-VTT reaffirms our statement that existing video-language datasets are incapable of assessing the temporal understanding ability of models.

Effects of Vision Encoders.

Among the models we evaluate, the temporally-adapted image-text models based on CLIP generally outperform the models with video-text pre-training. To further investigate how much the temporal aggregation modules contribute to the temporal understanding abilities of the CLIP-based models, we disable the temporal aggregation module in these models and replace it with a simple mean pooling layer. The results are shown in Tab. 4. Contrary to what is expected, disabling the temporal aggregation module only results in a slight drop in performance for X-Pool. It even improves the temporal understanding ability of CLIP4Clip and X-CLIP. This suggests that these temporal aggregation modules are potentially under-trained due to the weak requirement of temporal modeling in video-language datasets like MSR-VTT. Consequently, the superiority of the CLIP-based models mainly stems from the effective utilization of the static information in the video instead of a true understanding of the temporal concepts. For a similar reason, image-text models are able to achieve comparable performance on our dataset without further video-text training.

Similarity of Text Representations.

We calculate the average cosine similarity between the representations of the original captions and the counterfactual descriptions with different text encoders. As shown in Tab. 5, both the CLIP text encoder and Sentence-BERT produce highly similar sentence representations for samples in the “Sequence” aspect, indicating that the struggle of the evaluated models can partly be explained by the inability of text encoders to recognize the temporal distinction between the captions and the counterfactual descriptions. We also notice that the CLIP text encoder generally produces higher similarity scores even after it is fine-tuned on video-text data. This suggests that the ability to identify temporal concepts in natural language may be lost during the image-text pre-training stage and cannot be recovered by fine-tuning on existing video-language datasets.

Effects of Fine-Tuning Data.

We conduct a comparison between the performance of VidLMs fine-tuned on different downstream datasets. The results are shown in Tab. 6. We find that models fine-tuned on different text-to-video retrieval datasets exhibit different temporal understanding abilities. For example, DiDeMo tends to elicit higher accuracy on “Localization” and “Compositionality”, while LSMDC contributes to better understanding of “Intensity”. Also, since SSv2 only depicts single human actions, it brings benefits on the “Direction” aspect but not on “Sequence” understanding, which can be improved by fine-tuning on datasets with longer video duration and dense captions such as YouCook2. This finding advocates the use of diverse videos and captions in the training process.

3 Discussions

Previous work on the challenges of Winoground points out that accuracies based on cosine similarity comparison might be too harsh for the models, and it is possible that they under-perform on Winoground because the image-text pairs are out-of-distribution for them. This is also a concern for our dataset, so we follow them by calculating the Recall at k>1k>1 on the task of video-to-text retrieval on the entire VITATECS dataset for each aspect. Since a video may have multiple caption-counterfactual pairs in our dataset, we choose k=10k=10 and show the recalls for captions, counterfactuals, and both descriptions in Tab. 7. We observe that for both ALPRO and CLIP4Clip, the recalls of captions and counterfactuals are very close. This indicates that the models are able to connect the texts with their corresponding videos through the shared static information, but cannot distinguish between the different temporal information in the caption and the counterfactual.

Ablation Study of Counterfactual Design.

To verify the design of our counterfactual descriptions, we randomly sample 100 instances from each aspect of VITATECS and apply different modification strategies to the original captions. Specifically, we randomly choose 1-3 words in the caption and replace them with its synonym or a random word of the same part of speech. We also experiment with different types of words (nouns, verbs, or adjectives) as the target for replacement. The results are shown in Tab. 8. On the one hand, we can conclude that discriminating between the original caption and these altered ones is much easier when we randomly replace the words in the caption, even when only one word is changed. This margin is greater when we modify the nouns than when we modify the verbs in the captions, which aligns with our observation that current models rely heavily on static clues to make predictions. This demonstrates that the temporal understanding addressed by our VITATECS is more difficult to solve than simple object or action replacement. Also, the accuracy of the model rises quickly as we increase the number of replaced words, while our VITATECS maintains its difficulty despite showing greater lingual diversity. On the other hand, replacing words with their synonyms without contextual information may change their semantics significantly, as evidenced by the relatively high accuracy of models on these counterfactuals compared to VITATECS. This cautions us against the use of purely lexical methods for counterfactual construction. Finally, neither of these replacement methods is able to attach fine-grained labels to the resulting sentence, demonstrating the superiority of our counterfactual design.

Conclusion

This work aims to address the deficiency of temporal understanding evaluation abilities in existing video-language datasets. We present a fine-grained characterization of temporal concepts in video descriptions, and introduce a novel dataset that measures the temporal understanding capabilities of VidLMs by their ability to distinguish between the actual description of a video and its temporally modified alternative. To facilitate dataset construction, we design a human-in-the-loop annotation framework by leveraging LLMs for counterfactual description generation. Evaluation of state-of-the-art models demonstrates their failure to fully grasp temporal concepts. We hope our work can provide valuable insight into the future development of video-language understanding research.

References

Appendix A VITATECS

The videos and captions in our dataset come from two existing video-text datasets: MSR-VTT and VATEX . MSR-VTT is a video description dataset with 10K web video clips, each annotated with 20 natural sentences. Following common practice, we adopt the split introduced by Yu et al. and use 1,000 video-text pairs as the test set. VATEX is a large-scale multilingual video description dataset with comprehensive video content and a rich lexicon. Its public test set contains 6,000 videos, each annotated with 10 English descriptions. Both datasets are open-domain video-text datasets with diverse video content and descriptions.

Coverage.

We validate our categorization of temporal concepts by conducting a preliminary study on the temporal information in open-domain video-language datasets. Specifically, we randomly sample 100 video-text pairs from MSR-VTT and 100 from VATEX, and inspect whether they fall in one of the temporal aspects we define, contain other types of temporal information, or have no temporal information. As shown in Tab. 9, the result indicates the high coverage of our temporal aspect categorization, which allows us to assess the ability of VidLMs in a comprehensive and fine-grained manner.

Diversity.

One of the principles of our dataset is to be as diverse as possible and include open-domain video-text pairs with comprehensive content to simulate real-world scenarios. We verify the diversity of our dataset from two perspectives: videos and descriptions. Since our dataset is built upon existing open-domain video-text datasets, the diversity of videos is naturally guaranteed. Fig. 4 shows some videos from our dataset and other datasets aimed at temporal understanding evaluation. Other datasets often use computer-rendered scenes of simple geometric objects like CLEVRER , or action-centric videos of short duration like Something-Something , while our dataset contains videos of various topics and styles. Besides diversity in video content, VITATECS also involves diversified modification patterns of text descriptions owing to the flexibility of LLM generation. Fig. 5 illustrates the occurrence of verb pairs in our dataset, where the verbs in the inner circle correspond to the original caption, and the verbs in the outer circle correspond to the counterfactual description. Our dataset exhibits high diversity, allowing for a more comprehensive and faithful evaluation of VidLMs.

Quality Check.

We randomly sample 100 instances for each aspect from the final dataset and manually check if they satisfy our criteria. As shown in Tab. 10, the six aspects achieve an average pass rate of 94.8%, demonstrating the quality of our generated counterfactuals.

Discussion

We notice that for some of the instances in VITATECS (mostly in the “Intensity” aspect), it is infeasible to distinguish the caption from the counterfactual solely based on the frame contents. Some instances may require information about the audio track (e.g. “loud music” v.s. “serene music”) or the frame sampling rate (e.g. “slow motion” v.s. “normal speed”). However, most of the prevailing models still focus on the visual contents alone, rendering them incapable of distinguishing between these concepts. Nevertheless, we do not filter out these instances for the following two reasons. First, we believe the perception of audio and time is a vital part of video understanding and such information should be taken account into for future video understanding models. Second, these instances only constitute a minor part of the entire dataset. In fact, we find that for 51 instances, the modified part includes the word “loud” or “quiet”, and for 218 instances it includes the word “fast” or “slow”, most of them belonging to the “Intensity” aspect. This indicates that the conclusions of our evaluation experiments are still valid.

Appendix B Model Details

Our evaluation experiments are conducted on three pre-trained VidLMs (VIOLET , ALPRO , Singularity ), three temporally-adapted image-text model based on CLIP (CLIP4Clip , X-Pool , X-CLIP ), and one image-text foundation model (BLIP ). We also measure the performance of two recently open-sourced video LLMs. This section introduces the architectures and training paradigms of the evaluated models.

All the pre-trained models are composed of a vision encoder, a text encoder, and a cross-modal encoder. VIOLET initializes the vision encoder from VideoSwin-Base pre-trained on Kinetics-400 , and initializes both the text encoder and the cross-modal encoder with BERT-base . ALPRO uses a 12-layer pre-trained TimeSformer as the vision encoder, and initializes the text encoder using the first six layers of BERT-base and the cross-modal encoder using the last six layers. Singularity explores using only single frames for video-language pre-training. Its vision encoder is initialized from BEiT-base , while its text encoder is initialized from the first nine layers of BERT-base and its cross-modal encoder from the last three layers. It also has a temporal version, Singularity-temporal, which adds a 2-layer temporal Transformer encoder following the vision encoder.

Pre-training Objectives.

All the models adopt Masked Language Modeling (MLM) and Video-Text Matching (VTM) as part of their pre-training objectives. MLM randomly masks the input text and predicts the masked tokens based on the video and text context. VTM is a binary classification task that predicts whether the video and the text description are correctly matched with negative samples generated from the same batch. Besides these two objectives, VIOLET also designs the Masked Visual-token Modeling (MVM) objective using pre-trained DALL-E to discretize visual tokens. Both ALPRO and Singularity use Video-Text Contrastive (VTC) loss which aligns the representations from the vision and text encoders. ALPRO further proposes Prompt Entity Modeling (PEM) which predicts the entities in a video by leveraging a pre-trained prompter to generate pseudo entity labels.

Pre-training Data.

All three models are pre-trained using a combination of video-text data and image-text data. VIOLET is first pre-trained on YT-Temporal videos with noisy Automatic Speech Recognition (ASR) texts and then on WebVid-2.5M video-text pairs and CC-3M image-text pairs. ALPRO only utilizes WebVid-2.5M and CC-3M during pre-training. Singularity has two versions trained on different data. Singularity-5M is trained on WebVid-2.5M and CC-3M, while Singularity-17M is also trained using image-text pairs from COCO , Visual Genome , SBU Captions , and CC-12M . Singularity-temporal further performs a second stage pre-training using WebVid-2.5M based on single-frame pre-trained checkpoints.

B.2 CLIP-based models

CLIP-based models initialize the vision and text encoder from the pre-trained image-text model CLIP with an extra temporal aggregation module for video modeling. CLIP4Clip first encodes each video frame into a single vector using the CLIP vision encoder. The frames are then aggregated using mean pooling (“meanP”), an LSTM (“seqLSTM”), or a 4-layer Transformer encoder initialized from CLIP pre-trained weights (“seqTransf”). The video-text similarities are calculated with cosine similarity. CLIP4Clip also experiments with a tight Transformer (“tightTransf”) where the frame and text representations are concatenated and fed into a Transformer to produce the similarity score. X-Pool alters the aggregation of frames in CLIP4Clip by performing text-conditioned pooling over the frame representations. It uses a simple cross-attention module to pool the frames and thus is agnostic to the frame order. X-CLIP modifies the similarity calculation of CLIP4Clip by jointly considering video-sentence, video-word, sentence-frame, and frame-word similarity. The models are directly trained on downstream tasks without further video-text pre-training.

B.3 Zero-shot image-text models

Besides models trained on video-language data, some pre-trained image-language models have also shown strong generalization capabilities and can perform zero-shot transfer to video-language tasks. For example, BLIP is able to outperform many VidLMs on text-to-video retrieval and video question answering datasets despite ignoring all temporal information. We also evaluate BLIP on our dataset in a zero-shot manner and compare its performance with the VidLMs.

B.4 Video LLMs

Recently, with the advance of LLMs and chatbots , some works aim to endow them with the ability of visual understanding by connecting these LLMs with powerful visual encoders . Two pioneering models that introduce this idea into the video domain are Video-LLaMA and VideoChat . Both of the video LLMs first encode the video into visual embeddings, concatenate them with the input text embeddings, and feed the sequence into LLMs for auto-regressive generation. Video-LLaMA first encodes each sampled frame with a frozen ViT-G/14 from EVA-CLIP and a frozen Q-Former from BLIP-2 to obtain frame-level embeddings. It then aggregates them into video-level embeddings with a video Q-Former. Video-LLaMA also comprises an audio encoder based on ImageBind . VideoChat has a similar architecture to Video-LLaMA, except that it uses a GMHRA module for temporal aggregation instead of a video Q-Former. Both models are trained on large-scale image-text pairs and video-text pairs, as well as curated multimodal instruction data.

Appendix C Implementation Details

For BLIP and ALPRO, we use the checkpoints available on the LAVIShttps://github.com/salesforce/LAVIS library for evaluation. We adapt the code from the original GitHub repository for BLIP https://github.com/salesforce/BLIP to evaluate its performance on video-language tasks. For VIOLET and Singularity, we use the checkpoints and code released in their respective GitHub repositories https://github.com/tsujuifu/pytorch_violet https://github.com/jayleicn/singularity for evaluation. We report the performance of the temporal version of Singularity-17M in the main experiment of our paper.

For the CLIP-based models, since no fine-tuned checkpoints are released, we adopt ViT-B/32 as the backbone of the visual encoder and fine-tune the models on MSR-VTT according to the instructions in their respective codebases. Since the GitHub repository of X-CLIPhttps://github.com/xuguohai/X-CLIP is adapted from that of CLIP4Cliphttps://github.com/ArrowLuo/CLIP4Clip, we use the former for fine-tuning and evaluating both models. We experiment with the “meanP” (non-temporal) and “seqTransf” (temporal) aggregation strategies for both models and report the performance of the temporal version in the main experiment of our paper. For X-Pool, we fine-tune the model based on the code in its GitHub repositoryhttps://github.com/layer6ai-labs/xpool. Since the original X-Pool performs text-conditioned pooling directly on the set of all frames without temporal information, we report the performance of this non-temporal version in the main experiment. To enable the analysis of temporal aggregation modules, we follow CLIP4Clip and X-CLIP, and introduce a temporal version that adds a Transformer module before the text-conditioned pooling with the same initialization strategy. The fine-tuning of the temporal X-Pool employs the same hyper-parameters as the non-temporal version. The fine-tuning experiments are conducted on 8 NVIDIA TITAN RTX GPUs.

For the video LLMs, we perform zero-shot evaluation directly using the open-sourced models https://github.com/DAMO-NLP-SG/Video-LLaMA https://github.com/OpenGVLab/Ask-Anything. We use Vicuna-7B-v0 for Video-LLaMA and StableVicuna-13B for VideoChat as the LLMs. We disable the audio input of Video-LLaMA for fair comparison. The instruction we give the video LLMs is shown in Tab. 11, which is similar to the ones we use in human evaluation. The predictions are generated with temperature T=0.2T=0.2.

For human evaluation, the annotators are shown a webpage containing a video and two text descriptions, as shown in Fig. 6. They are required to choose the caption that best describes the video content. The differences between the two descriptions are underlined and emphasized in boldface.

Appendix D Full Evaluation Results

We summarize the results of evaluation experiments in Tab. 13.

Appendix E Details of Dataset Construction

For each of the six temporal aspects, we manually annotate 30 caption-counterfactual pairs as the initialization of the positive exemplar set. We also write 30 negative examples and a total of 21 N/A examples for different aspects. Some examples are shown in Tab. 14.

In-Context Learning Generation.

For the prompt of in-context learning generation, we design the input template as shown in Tab. 15. The instruction is determined according to the given temporal aspect as in Tab. 16. For each input query, we choose six positive exemplars and two N/A exemplars from the specified aspect, and construct the input to ChatGPT with the template, the instruction, the exemplars, and the query. We sample one candidate for each query using top-p sampling with p=0.8.

Filter Model.

The NLI models are implemented using the Sentence-BERT libraryhttps://www.sbert.net/ and are initialized with the pre-trained nli-deberta-v3-base models. We use a simple data augmentation technique during the fine-tuning procedure, i.e., randomly choose one word from the sentence and replace it with its synonym based on WordNet . The filter model is fine-tuned with the Adam optimizer with a learning rate of 5e-5. We adopt a batch size of 16 and fine-tune our filter model on the augmented positive and negative exemplar sets for five epochs. After the fine-tuning procedure is finished, for each candidate sentence pair (t1,t2)(t_{1},t_{2}), we feed both (t1,t2)(t_{1},t_{2}) and (t2,t1)(t_{2},t_{1}) into the NLI model and multiply the corresponding classification probabilities. We then classify the candidate into one of the six aspects if its score is higher than all the other scores by a threshold of 0.3. Consequently, only those candidates for which the filter model produces consistent and highly confident scores will be regarded as positive exemplars, which allows us to filter out most of the wrong predictions and maintain the quality of the exemplar sets.

Iterative Procedure.

We conduct iterative generation, filtering, and human revision until the performance of the filter model converges on a held-out evaluation set with 100 samples per class. The filter model converges after three iterations, as shown in Table 12. Table 12 also shows the valid samples generated in each iteration after human revision. Its increasing number demonstrates the iterative procedure is able to benefit the generation quality. The number of human labeling needed for each iteration is 554, 523, 477, respectively.

Annotation Instructions.

During the human revision stage, we first give the annotators instructions including the definition of the aspect and typical positive and negative examples, as shown in Table 17 and Table 18. Then the annotators are shown the candidate text pairs and need to decide whether the candidate falls in one of the six temporal aspects or should be classified as a negative example.

Appendix F Ethics Consideration & Broader Impact

We are aware that texts generated by LLMs may contain toxic and biased data. To mitigate this issue, we leverage Perspective API https://developers.perspectiveapi.com/ for scoring possible negative impacts of generated sentences, including toxicity, severe toxicity, identity attack, insult, profanity, threat, and sexually explicit content. Data samples with one of these scores exceeding 0.9 are discarded from our dataset. It is still possible, though, that a few harmful samples may exist in the final dataset. Future work that adopts our annotation framework should take similar safety measures to reduce the harm of generated data.

Appendix G Limitations

First, our categorization of temporal aspects does not completely eliminate the correlation between temporal and static information. Since current text-to-video generation models still cannot generate sufficiently natural videos, it is unrealistic for us to control for the static scenes and generate counterfactual videos. As a result, our dataset focuses on manipulating the text description, and our videos are still subject to the correlation between temporal and static information, especially for the “Localization”, “Compositionality” and “Type” aspects. Nevertheless, by picking out video-text pairs with different types of temporal information and generating temporally counterfactual descriptions, our dataset can offer a fine-grained evaluation of these temporality-related concepts, which are disentangled from static information to different degrees. This allows us to pinpoint the abilities of VidLMs at our best while covering various types of temporal information, given the limited resources and the current state of video-language research. Future work may achieve better disentanglement by employing massive human labor or leveraging stronger video generation models.

Second, our dataset may contain wrong labels or unnatural counterfactual descriptions, since it is partly generated by language models. However, this should account for only a small proportion of the entire dataset according to the performance of our filter model and the result of quality check.

In addition, due to limitations in time and resources, we were not able to conduct a full-scale controlled experiment to investigate the impact of specific model design choices on temporal understanding abilities. We hope our dataset can promote research in the temporal understanding of VidLMs and leave a more in-depth analysis of architectures for future work.