MovieChat+: Question-aware Sparse Memory for Long Video Question Answering
Enxin Song, Wenhao Chai, Tian Ye, Jenq-Neng Hwang, Xi Li, Gaoang Wang
Introduction
Recent advances in Large Language Models (LLMs) achieve great success in Natural Language Processing (NLP). It is a natural progression to introduce multi-modality into LLMs and turn it into Multi-modal Large Language Models (MLLMs), which are able to conduct multimodal rationalization and understanding. MLLMs have shown incredible emergent capabilities in various multimodal tasks such as perception (e.g. , existence, count, position, OCR) , commonsense reasoning , embodied agent , and code reasoning , resulting in a potential path to Artificial General Intelligence (AGI). Compared to LLMs and other task-specific models, MLLMs provide a more human-like interpretation of the scenarios, a user-friendly interface for interaction, and a broader range of capabilities.
Existing vision-centric MLLMs follow the paradigm that utilizes pre-trained LLMs and visual encoders with additional learnable modules (Q-former or simple projection layer ). In the video understanding field, some previous works follow this paradigm to build video MLLMs, while works in the other paradigm combine existing visual perception tools (e.g. , tracking and classification) and LLMs through application programming interface to build a system without training. However, these methods either employ complex spatial-temporal modules or heavily rely on additional perception tools to acquire temporal information for video understanding. Additionally, when handling long videos, the computational complexity and memory costs associated with long-term temporal connections are significantly increased, posing additional challenges, as shown in Fig. 1. Furthermore, there is also a lack of a standardized benchmark to evaluate the capabilities of these systems.
To the best of our knowledge, our work is the first to address long video understanding tasks (K frames). We argue that the computational complexity, memory costs, and long-term temporal connection are the main challenges for understanding long videos. Unlike other approaches that require training additional temporal modules in Fig. 2, our MovieChat lifts pre-trained MLLMs to understand long videos without the need for additional trainable temporal modules, employing a zero-shot approach. Inspired by the Atkinson-Shiffrin memory model , we propose a memory mechanism to address long video understanding tasks. This mechanism comprises a rapidly updated short-term memory and a compact, thus, sustained long-term memory. In our updated version, namely MovieChat+, we design a vision-question matching-based memory consolidation mechanism to enhance the compactness of memory. This mechanism significantly anchors the predictions of the visual language models in the relevant visual content. Our MovieChat+ significantly improves upon the initial version and outperforms the state-of-the-art in both short and long video question-answering tasks, surpassing even methods specifically tailored for short video question-answering challenges. As shown in Fig. 1, our approach outperforms other existing methods in terms of Video Random-Access Memory (VRAM) cost. We also release a new benchmark, MovieChat-1K, with 1K long videos and 14K manual question-answering pairs for validation of the effectiveness of our proposed method. Additionally, we expand MovieChat-1K with 2K temporal labels. The contributions of this work are summarized as follows:
We present MovieChat, the first framework designed to support long-term videos (K frames), leveraging pre-trained MLLMs and employing a zero-shot, training-free memory consolidation mechanism.
Our updated version, MovieChat+, enhances memory compactness through the implementation of a vision-question matching-based memory consolidation technique. This improvement significantly surpasses the initial version and outperforms the current state-of-the-art in both short and long video question-answering tasks.
We have released the first long-video understanding benchmark, MovieChat-1K, which now includes an expansion to 2K temporal labels compared to the initial version. We conducted extensive quantitative evaluations and case studies to assess the comparable performance of both understanding capability and inference cost.
Related Works
LLMs have achieved great success in natural language processing (NLP) tasks recently. Many works try to build MLLMs by combining models of other modalities. Flamingo bridges powerful pre-trained vision-only and language-only models and achieves state-of-the-art performance with few-shot learning. BLIP-2 proposes a generic and efficient pre-training strategy that bootstraps vision-language pre-training from an off-the-shelf frozen pre-trained image encoder and a frozen large language model. MiniGPT-4 also aligns a frozen visual encoder with a frozen LLM, Vicuna , using just one projection layer to realize the system. Otter showcases improved instruction-following ability and in-context learning. In the video field, ChatVideo treats tracklets as the basic video unit and allows users to interact with the LLMs. VideoChat integrates video foundation models and LLMs via a learnable neural interface, excelling in spatiotemporal reasoning, event localization, and causal relationship inference. Video-LLaMA further leverages pre-trained models ImageBind and LLaMA , bootstraping cross-modal training in videos following BLIP-2. Yet, these methods fail to handle long video understanding because of high computation complexity, large memory cost, and weak long-term temporal connection. Therefore, our main effort is to introduce an effective memory mechanism to overcome these challenges.
2 Long Video Understanding
Understanding long videos is a challenging task in computer vision. Prior arts use 3D CNN for long-term feature bank , object/human-centric motion , or other forms as video representations. MIST decomposes dense self-attention into a cascade segment and region selection module to increase the computation efficiency for understanding minutes of long videos. Building long-form video understanding datasets is challenging and rarely explored. captures large scale data from Kinetics-400 , but only for generic event boundary detection tasks. creates a language grounding benchmark from audio descriptions of movies, but it lacks long-term understanding evaluation. successfully builds a benchmark contains multiple sources of information (e.g. , video clips, plots, and DVS) for question-answering tasks in the movie field. There are also several datasets of video-caption/description pairs among various domains, such as cooking (e.g. , MPII Cooking and TACoS ), instruction (e.g. , HowTo100M and HiREST ), Ego , and movie (e.g. , MovieQA and MovieNet ) from different sources such as YouTube , Twitter , and Internet . Yet, those datasets lack diverse and fine-grained dense captioning for long videos.
3 Memory Models in Vision Tasks
There are some prior works exploring memory models in various vision tasks in videos, such as video object segmentation (VOS) , multi-object tracking (MOT) , visual object tracking (VOT) , and action understanding . MeMOT builds a large spatiotemporal memory that stores the past observations of the tracked objects. XMem develops an architecture that incorporates multiple independent yet deeply connected feature memory storage to handle long videos with thousands of frames. Recently, propose a memory streaming method to support dense video captioning. We learn from the experience of those prior arts and further adopt an effective memory mechanism in combination with LLMs.
Our method focuses on reducing the redundancy of visual tokens in the video and building a memory mechanism to pass the information among a large temporal range.
MovieChat
Our proposed method, MovieChat, comprises several key components, including the frame-wise visual feature extractor, the short-term memory module, the long-term memory module with question-aware consolidation strategy (our improved version, namely MovieChat+), the video projection layer, and the Large Language Model (LLM), as illustrated in Fig. 3. our approach is designed for ultra-long videos (K frames) understanding through interactive dialogue with the user. To address the impractical storage demands of concurrently storing a vast number of frames in both GPU memory and RAM, we employ a sliding window approach to efficiently process the video. The short-term memory module embeds dense tokens with sliding windows, and the long-term memory module periodically updates based on question-aware consolidation. MovieChat supports two inference modes: breakpoint mode and global mode. Breakpoint mode is used to understand a specific moment in the video, providing insights and answers based on that particular frame or scene. Global mode, on the other hand, is employed to comprehend the entire video as a whole, enabling a comprehensive understanding of the overall content and context.
2 Visual Feature Extraction
3 Short-term Memory
The short-term memory maintains a fixed-length buffer to temporarily hold frame tokens. The previously extracted visual features by sliding window times without further processing are used to construct short-term memory, which can be described by the following formulation:
where is the short-term memory, and is equal to . Note that we configure short-term memory to contain a fixed length of frames. The design stems from the fundamental purpose of short-term memory, which is to aid in the interpretation and understanding of video content by leveraging contextual information from recent, short-term segments of the video.
As a new batch of visual tokens enters, when the short-term memory reaches its capacity, we pop the currently stored frames to the memory consolidation module and clear the short-term memory. The output video feature obtained from the consolidation module augments the long-term memory; on the other hand, it re-initializes the short-term memory with this feature. The initialization aims to communicate information between different sliding windows, thereby achieving more efficient compression.
4 Question-aware Long-term Memory (MovieChat+)
The long-term memory can effectively avoid the problem of catastrophic knowledge forgetting, which is crucial for handling long video understanding tasks. The features stored in short-term memory are dense tokens, but due to the limitations of GPU memory and computation cost, storing all the tokens dropped from short-term memory into long-term memory buffer in sequence is infeasible. Besides, we observe significant temporal redundancy in videos, where activities span multiple frames with minimal visual changes. Additionally, only a small fraction of the entire long video content is relevant to the given question in practice. To this end, in our updated version, MovieChat+, we propose to merge adjacent frames based on their relevance to specific questions, thereby streamlining video feature representation and enhancing encoding efficiency. This method converts dense tokens into sparse memories centered on pertinent questions, which are stored in long-term memory.
To be specific, as shown in Algorithm 1, we first utilize a pre-trained text encoder to encode the specific question to the same embedding space as the visual features, which can be formulated as:
We then calculate the average cosine similarity between each frame feature within the short-term memory and the encoded question , which can be formulated as:
When we watch a video with a question in mind, we tend to skip over largely irrelevant segments. This is the motivation for using the calculated similarity between questions and visual features. If the visual features in short-term memory are highly related to the questions, we merge fewer into long-term memory; otherwise, we merge more during consolidation. We consider the target merging threshold , i.e., the number of consolidated frames after merging using the similarity calculated above.
We compare the average similarity to the threshold to assess its relevance to the question. The target merging threshold is formulated as follows,
For segments with high relevance, we set a base merging threshold . For segments with low relevance, we set a compression coefficient , reducing the target merging threshold to , thus compressing fixed-length segments into fewer ones.
Following the methodology outlined in ToMe , we then periodically perform memory consolidation by merging the most similar tokens in the adjacent frames. Here, we calculate the average cosine similarity among embedded tokens, as the tokens can effectively summarize the information of each frame:
Extend Positional Encoding. For long-term memory, the number of tokens exceeds the maximum length of the positional encoding from the pre-trained model. Thus, our model utilizes the positional encoding mechanism following BERT , which results in a portion exceeding the length threshold without available positional encoding. In order to handle long enough memory, we adopt the hierarchically decomposed positional encoding method proposed by Su , which allows us to extend the absolute positional encoding of length from to .
5 Inference
Previous methods always use the representation of the whole video to conduct understanding and question-answering. While this method provides a broad overview, it often struggles with accurately localizing specific moments or details in long videos. To this end, we propose two inference modes, global and breakpoint, for long video understanding tasks as follows.
Global Mode. Global mode is defined as the understanding and question-answering for the whole video. Under this mode, the focus is on capturing the details of the full duration of the video. Therefore, we only use long-term memory as the video representation .
Breakpoint Mode. Breakpoint mode is defined as understanding specific moments in a video. Since events possess continuity, we need to consider not only the information directly related to the moments stored in short-term memory but also the information indirectly related stored in long-term memory . Therefore, we hypothesize that when querying the movie at a specific moment , the video representation should be the aggregation of , , and the current video frame feature . We observe that straightforward concatenation of these elements delivers outstanding results, and we defer the investigation of alternative aggregation methods to future research.
Subsequently, the video representation goes through a Q-former and a linear projection layer before being fed into the LLM , which can be formulated as:
where is the projection from visual space to text space, represents the answer or instruction of the breakpoint, and is employed to denote the question, respectively.
A New Benchmark: MovieChat-1K
Previous works on building long video understanding benchmarks either focus on non-question-answering tasks (e.g. , language grounding , generic event boundary detection , user engagement and movie metadata prediction , etc. ) or lack long-form understanding evaluation . To better evaluate the performance of MovieChat, we collect a new benchmark for long video understanding tasks, MovieChat-1K, which contains 1K high-quality video clips sourced from various movies and TV series with 14K manual annotations. In our updated version, MovieChat+, we expand by an additional 2K temporal grounding labels.
Video Source. As shown in Fig. 5a, we collect videos from 15 popular categories with varying distribution, including documentary film, detective film, animation film, etc. Among these, each video comprises multiple alternating scenes, contributing to a diverse and dynamic visual narrative within the context of the collection. We further illustrate our improved content-based categorization and analysis of MovieChat-1K questions in Tab. I. The visual representation in Fig. 5b demonstrates the clip duration distribution of MovieChat-1K. Over 90% of the videos exhibit a duration ranging from 10K to 12K frames, while 14.6% of videos extend beyond 12K frames. Only 8.6% of videos have a duration of less than 10k frames. To demonstrate that MovieChat-1K is indeed a long-form dataset, we employ the same method as proposed by EgoSchema to calculate the certificate lengths. As depicted in Fig. 6, our approach results in a certificate length that is 5.1 times longer than that of EgoSchema . Specifically, the annotated captions and questions are used as temporal tags, and the corresponding clip lengths are calculated manually.
Temporal Label Collection. Following , we augment MovieChat-1K with temporal labels. Most VideoQA datasets are unsuitable for exploring how to handle irrelevant redundant frames, as they are composed of short video clips (no more than 15 seconds) that have been pre-trimmed to focus solely on the pertinent content. We apply temporal labels exclusively to questions categorized under breakpoint mode. This is because questions in global mode mostly pertain to global video content (e.g. , ”Where does the video take place?”). Furthermore, the answers to these global mode questions are often discernible across extensive segments of the video, such as “Is there more than five different characters appearing?”. For each question-answer pair in breakpoint mode, we annotate the start and end times of the relevant segments as shown in Fig. 7.
We restrict the labeling of temporal annotations in MovieChat-1K to the validation and test sets, under the premise that these labels are instrumental in assessing the ability of models to identify question-relevant video segments, rather than for training purposes. As a result, 2K question-answer pairs drawn from 200 videos are annotated with temporal labels. Fig. 8 demonstrates that most of the segments last for less than 12 seconds, with an average duration of 6.3 seconds, which is extremely short compared to the video length (approximately 700 seconds).
Annotations Analysis. For each video, we manually set and provide 1 dense caption for the whole video, 3 question-answering pairs for global mode, and 10 question-answering pairs with timestamps for breakpoint mode. Fig. 5c illustrates the distribution of question types in MovieChat-1K. Note that MovieChat-1K is specifically designed for long video comprehension tasks. The majority of questions are open-ended, with only a quarter classified as multiple-choice questions, marked by initiators such as ‘Do,’ ‘Does,’ ‘Is,’ or ‘Are.’ As illustrated in Fig. 9, we also compute the word distributions of the question-answer pairs, which includes common objects (people, clothes, etc.), time (day, night, etc.), scenes (indoor, outdoor, etc.), and so on.
To facilitate a detailed understanding of long videos, we provide a dense caption for each video. As shown in Fig. 10, MovieChat-1K exhibits diverse caption lengths in the segmented clip level. Approximately two-thirds of the clips have captions with 100-149 words, while one-fifth of the clip captions have fewer than 100 words. About 11% of clips have long captions with more than 150 words.
To analyze the word distribution of our generated captions, we compute their distributions. The resulting word distribution is presented in Fig. 11, which includes common objects (man, woman, people, girl, etc.), attributes (detective, various, small, white, etc.), locations (inside, behind, south, next, etc.), scenes (room, house, building, office, etc.), actions/events (talk, enter, leave, take, etc.), and more.
In terms of actions, MovieChat-1K captions contain nearly the same number of verbs as with the WebVid10M dataset . To evaluate this, we use the NLTK toolkit to analyze the number of verbs in captions, focusing on extracting and tagging all unique verbs. We find a total of 109,485 verbs in the WebVid10M caption dataset, while the MovieChat-1K captions contain 102,988 unique instances of verbs. While these counts may not be entirely accurate due to our simple counting method, we believe they provide a rough indication of the actions of the two datasets.
Experiments
We conduct quantitative and qualitative evaluations of our complete MovieChat+ compared to previous methods and the original MovieChat. Additionally, we perform ablation studies to investigate MovieChat+.
We use several widely used open-ended datasets: MSVD-QA , MSRVTT-QA , ActivityNet-QA , and NExT-QA for short video question-answering tasks. The evaluation process is under the assistance of LLM with the default hyper-parameter settings. The accuracy and relative scores on a scale of to are reported. Compared to previous methods , MovieChat achieves comparable performance even it is not specifically designed for short video question-answering tasks, as shown in Tab. II.
We also report the results of our zero-shot evaluation on the test split of the NExT-QA benchmark in Tab. III. NExT-QA divides its questions into three categories: Causal (C), Temporal (T), and Description (D). Compared to prior work, our approach achieves higher accuracy across all aspects, demonstrating its effectiveness at understanding temporal context with the question-aware consolidation.
Following , we employ GPT-assisted evaluation to conduct a more comprehensive comparison of the text generation performance between our appraoch and previous methods on processed ActivityNet-QA . The evaluation pipeline covers crucial metrics (including Correctness of Information, Detailed Orientation, Contextual Understanding, Temporal Understanding and Consistency) and assigns relative scores to the generated predictions on a scale of 0-5. We present the results of the generation performance evaluation in Tab. IV. The results reveal its competitive performance across all key aspects compared to previous methods. It should be noted that in comparison with MovieChat, MovieChat+ does not exhibit substantial enhancements in both question-answering accuracy and generative performance when evaluated on short video datasets. This is because the content of short videos is often closely related to the questions, which is an extreme case for our proposed frame filtering strategy. Similar to the original MovieChat, all video frames are merged with nearly equal consideration.
We further evaluate our approach on the task of action recognition on Seed-Bench to study the effect of MovieChat for short-term temporal understanding tasks. In contrast to the longer setting in the procedure understanding task, the videos in this task generally have duration of around 10 seconds. As shown in Fig. V, we compile the number of frames fed into the LLM decoder for different models, along with the corresponding performance of procedure understanding and action recognition. As the input frames increasing, our question-aware approach yields greater benefits in both the procedure understanding task and the action recognition task. These results suggest that the question-aware merge strategy, which filters more related context for reasoning about spatial-temporal relationships between video segments, may be crucial for fine-grained action understanding. However, when feeding the same number of frames into the LLM decoder, MovieChat shows minimal improvement in procedure understanding. We speculate that when dealing with a limited number of sampling frames, the inclusion of compressed frames containing redundant information could potentially hinder the model’s ability to comprehend procedural sequences.
1.2 Long Video Question-answering
We evaluate the long video question-answering performance of MovieChat with our proposed MovieChat-1K. We split 1,000 videos into training set (800), test set (100), validation set (100) and only use test set for final performance evaluation. We select two non-LLM based video understanding models (e.g. GIT , and mPLUG-2 ) and three recent LLM-based video understanding models (e.g. Video Chat , Video LLaMA , and Video-ChatGPT ) as the baselines. Yet, none of those methods can support such long video (K frames). Therefore, to accommodate their length limitations in global questions, we uniformly sample from the original video up to the maximum frame count which can be officially supported by each individual model. For breakpoint questions, we extend half of the maximum frame count before and after the breakpoint ( placing the breakpoint at the center frame).
To enhance the robustness of the results, we simultaneously employ GPT-3.5 and Claude as LLM assistants, with the additional support of human blind rating. We observe a discrepancy between the accuracy and relative score generated by the previously LLM-assisted evaluation method for video question-answering tasks. However, merely adjusting the prompt for the LLM cannot effectively address this issue. Therefore, after obtaining the accuracy and score from the LLM-assisted evaluation method, we implement manual filtering to remove results with inconsistent values, thus improving the reliability of our outcomes.
As shown in Tab. VI, compared to previous methods, MovieChat reads more video frames. In both global mode and breakpoint mode, our method maintains a performance gain in terms of the average accuracy and score provided by LLM assistants and human blind rating. Compared with MovieChat, our method significantly improves accuracy in the global mode, which fully demonstrates the effectiveness of our question-aware consolidation strategy.
We further compare the quality of answers generated by MovieChat and previous methods in long video question-answering on MovieChat-1K. As shown in Tab. VII and Tab. VIII, with the average score provided by GPT-3.5 , Claude and human bling rating, our complete method, MovieChat+, continues to generate higher-quality answers even as the video contents become more extensive, significantly outperforming the initial and simpler version of MovieChat.
MovieChat-1K contains question-answer pairs of varies types. To better assess the performance of our approach, we conduct evaluations on the long video question answering task using various types of questions. We roughly categorize the question types into multiple-choice questions and open-ended questions. With the average results of GPT-3.5 , Claude and human blind rating, Tab. IX and Tab. X respectively present the accuracy and scores of MovieChat and the baseline across different question categories in both global mode and breakpoint mode. In various research conditions, our approach consistently outperforms the baselines in both open-ended and true-false questions.
1.3 Question-answering on Other Long Video Dataset
We further report zero-shot question-answering results for MovieChat on another commonly used long-form video dataset, EgoSchema in Tab. XI. EgoSchema is a diagnostic benchmark for evaluating long video understanding capabilities of advancing systems, featuring over 5000 human-curated multiple-choice question-answer pairs based on more than 250 hours of real-world video data. Prior works have demonstrated that the sequence in which options are presented can significantly influence the outcomes of tasks involving multiple choices. To mitigate this effect, we provide MovieChat with questions in EgoSchema exclusively. Subsequently, we employ LangChain to assess the similarity between the responses of MovieChat and the provided options. We then select the option that most closely aligns with our anticipated answer as our decision. MovieChat+ produces significantly superior result than its initial version and other leading non-LLM based and LLM-based methods.
2 Ablation Study
As our approach incorporates a memory mechanism including short-term memory and long-term memory, it is imperative to evaluate how the proposed memory mechanism influences the performance. Tab. XII and Tab. XIII provide the memory-dependent performance of our approach for long video question-answering and generative tasks with the average results of GPT-3.5 , Claude , and human blind rating. MovieChat with the memory mechanism significantly outperforms the memory-independent variant, which signifies the importance of memory mechanisms.
We further consider the following approaches to evaluate our memory consolidation strategy:
No memory. Due to memory constraints, we uniformly sample 16 frames from all frames, concatenate all visual tokens, and feed them into the decoder.
Spatial- or temporal-pooling. We pool the visual features, along either the spatial or temporal dimensions to reduce the number of tokens fed to the LLM decoder.
EMA. Following , we use an exponential moving average of frame features at each time step.
Tab. XIV compares the results of the different memory modules. For , where we can feed all the tokens from the vision backbone into the decoder, “No Memory” performs the best because it uses the most tokens. However, it is impractical to use “No Memory” for due to its computational cost. With more frames, naively pooling along the spatial- or temporal-dimensions actually performs worse. This is likely because we are averaging out information over longer temporal durations, and thus losing the important details required for more detailed localization or captioning. Our method, on the other hand, leverages more frames to improve performance by retaining diverse features within the memory.
2.2 Large Language Models Ablations
Most previous video understanding methods primarily employed LLama and its variants as text decoders. With the average results of GPT-3.5 , Claude and human blind rating, Tab. XV and Tab. XVI illustrate how the performance of MovieChat changes when using LLama and LLama2 as the large language model respectively.
Contrary to our initial hypothesis, the performance of MovieChat with LLama2 hardly surpasses those of MovieChat with LLama across various key metrics. The outcome suggests that the advancements incorporated into LLama2 may not translate to significant improvements. We further investigate a specific example to analyze this phenomenon. As shown in Fig. 13, MovieChat with LLama provides answers that are more aligned with the video content. Surprisingly, MovieChat with LLama2 offers an approximation of the time required for each step (indicated in italics). While its time estimates do not precisely match the actual durations, the proportion of time provided is realistic. Even though LLama2 cannot obtain specific time information when processing feature-rich video frames, the memory buffer design allows for dense sampling of video frames, enabling LLama2 to estimate the proportion of time for each scene based on adjacent similar frames. Therefore, we propose that the lower evaluation metric results of MovieChat with LLama2 compared to MovieChat with LLama may be attributed to the question-answer pairs in the dataset.
2.3 Hyper-parameter Ablations
We perform a series of hyper-parameter ablations based on the MovieChat-1K dataset to better understand our approach. Fig. 12 shows the performance when ablating the length of long-term memory buffer , the length of short-term memory buffer , short-term initialization, question-frame similarity , target merging coefficient and judging relevance basis with the average results of GPT-3.5 , Claude , and human blind rating. The performance of MovieChat degrades which shows the validity of our empirically chosen hyper-parameters.
Length of Memory Buffer. The length of different memory buffers has a combined effect on MovieChat’s performance. Since the LLM-based evaluation shows a positive correlation between accuracy and score, we use accuracy to gauge performance. Fig. 12 (top left and middle) demonstrates that information obtained from the video expands with the growing length of memory buffers. However, this benefit is tempered by a more pronounced loss of fine details, a consequence of maintaining a fixed length for consolidation. Therefore, as the lengths of two memory buffers increase, the performance of our approach exhibits a trend of initially rising and then declining. Thus, we set the length of long/short-term memory to 256 and 16, respectively.
Short-term Initialization. As shown in Fig. 12 (top right), using merged tokens for short-term initialization outperforms the last few tokens and uniform sampling. When initializing the next short-term memory with the last few tokens from the previous short-term memory, it is unable to adequately represent the previous information, leading to the final merged tokens being either repetitive or lacking coherence with the previous time step. Uniform sampling faces similar issues, but it manages to capture information with representative frames from the previous time step.
Question-frame Similarity Threshold. Before compressing the short-term memory, we need to assess the relevance of the stored segments to the question by evaluating the similarity between video frames and the question, thereby deciding the degree of compression for the current short-term memory. Fig. 12 (bottom left) illustrates the outcomes of experimenting with various question-frame similarity thresholds, and the optimal performance is achieved at . When the threshold is low, it is difficult to effectively compress and filter out distracting or redundant information. Conversely, an excessively high threshold might lead to the over-compression of valuable information.
Target Merging Coefficient. We further explore the target merging coefficient of weakly to strongly related segments. For fairness, all weakly related short-term memories are compressed into 1 frame, allowing us to focus on how the performance of strongly related segments varies with different compression levels. Fig. 12 (bottom middle) shows that increasing merged frames for strongly related segments initially boosts performance but eventually leads to a decline. We initially assume that less compression of strongly related segments would significantly enhance model performance. Yet, the results hint at a more intricate link between compression intensity and performance. To incorporate more long-term memory frames into the pre-trained model, we extend positional encoding with hierarchical decomposition. However, the approach involves balancing extended input lengths with the integrity of positional representations. A direct extension may not be ideal since most training frames are shorter than those in the pre-trained model, making lower positions well-trained for absolute positions, whereas higher positions are less trained, offering only a rough estimate of relative positions. Thus, interpolating lower positions poses a greater risk of disrupting established positional embeddings compared to interpolating higher positions. When querying the same video with the same question, retaining more merged frames for strongly related segments leads to a noticeably elongated long-term memory, where effective positional encoding becomes challenging, reducing comprehension of long videos. This highlights the necessity of balancing the retention of dense information and the compression for effective long-term video understanding.
Judging Relevance Basis. Determining the relevance of short-term memory to a question based on similarity can be approached in three ways: comparing the highest similarity within a segment, the lowest, or the average with the question-frame similarity threshold . According to Fig. 12 (bottom right), using the minimum or average similarity shows similar performance. However, selecting the maximum similarity as the criterion leads to a performance drop. We believe this is due to the lenient judgment of relevance between the question and frames when choosing the maximum similarity within a segment, which introduces more redundant information. Thus we have elected to utilize the average similarity as the criterion for comparison.
3 Case Study
We perform an extensive case study of MovieChat on a variety of open-ended long video (such as cartoon movie and TV series) including the breakpoint mode (Q#1) and the global mode (Q#2). The evaluation is conducted between our approach and previous methods as shown in Fig. 13. For Q#1 in breakpoint mode, we mark the timestamp when the question is asked. For long videos over K frames, MovieChat is still capable of providing excellent responses to questions regarding both the current moment and the entire video content with less hallucination. We also provide more examples to show the long video scene understanding and temporal understanding ability of MovieChat in Fig. 16, 16 and 16.
Limitation
Although MovieChat has demonstrated impressive abilities in long video understanding, it is still an early-stage prototype and has some limitations, including 1) Limited perception capacities. The performance of our approach is hindered by the pre-trained short video understanding model. 2) Inadequate Time Processing. MovieChat provides only rough estimates of the duration proportions of events within long videos, lacking precision in temporal details.
Conclusion
Conclusively, we present an innovative video understanding system integrating video foundation models and large language models. By incorporating an enhanced memory mechanism represented by tokens in Transformers, our proposed system, MovieChat overcomes challenges associated with analyzing long videos. MovieChat achieves state-of-the-art performance in long video understanding, surpassing existing systems, which are limited to handling videos with few frames.