Streaming Long Video Understanding with Large Language Models
Rui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Shuangrui Ding, Dahua Lin, Jiaqi Wang
Introduction
The evolution of Large Language Models (LLMs) has significantly advanced artificial intelligence, encompassing text generation and reasoning in complex language environments . Later, the community extends LLMs to multi-modal domains, demonstrating promising results in captioning and question-answering tasks that integrate diverse visual signals . Yet, within the domain of video understanding, long video sequences pose a formidable challenge. Incorporating such long visual contents into LLMs requires a substantial number of tokens, which not only amplifies computational demands but also risks early contextual information loss .
Among the recent works on general video understanding with LLMs , a prevalent strategy is using sparse temporal sampling or spatio-temporal pooling to reduce tokens. Unfortunately, this paradigm explicitly loses substantial information in the long time span. To address this limitation, develop frame-wise compression, with LLaMA-VID as a typical example. It compresses each frame into only two tokens but overlooks the inter-frame temporal dynamics which are vital in compressing temporal redundancy within videos. Besides, its question-dependent compression pipeline limits the ability to produce a general representation that can handle diverse instructions. Another line of works employ memory banks to store history information . Whereas, these methods rely on explicit timestamps to recall the historical details, limiting the ability to generate comprehensive responses without specific time indicators.
In this work, we propose VideoStreaming, a novel Memory-Propagated Streaming Encoding architecture with Adaptive Memory Selection to sequentially encode a long video into condensed memories and generate responses referring to relevant timestamps. The core idea behind the memory-propagated streaming encoding is to preserve representative spatial cues and temporal dynamics while reducing temporal redundancy in videos. To achieve this goal, we segment the long video into multiple short clips and sequentially encode each clip. When encoding each clip, we first refer to the encoded results of its preceding clip as historical memory, then concatenate it with the current clip features and feed them into a small decoder-only language model . Due to its autoregressive nature, the information of the sequence naturally accumulates to the last few tokens . Consequently, we take these last few tokens as an updated memory that encapsulates the video information up to the current timestamp. Through this streaming encoding, we explicitly take long-term temporal relations into consideration and maintain a fixed-length memory to represent an arbitrarily long video.
However, this fixed-length memory inevitably loses detailed information, especially in early contexts. To address this problem, we store the historical memories of all clips and select a constant number of subsets that are closely related to the question. To accomplish this, when streaming encoding each clip, we additionally append a summary token at the end of the sequence as a clip indicator that summarizes the clip contents within one token. Then, given a specific question, we concatenate the condensed memory from the final iteration with the question and pass it through the same small language model used in streaming encoding. We take the final token as the question indicator and calculate its similarity with all historical clip indicators, the clip indicator with higher similarity means its corresponding memory is more related to the question. Finally, we feed the adaptively selected memories into the LLM for detailed question answering.
In practice, we realize our VideoStreaming with a carefully designed two-stage progressive training process and long-video data construction strategy. In the first stage, we empower a small language model with the single-clip encoding capability by a specialized prefix task. In the second stage, it serves as the streaming encoder and we jointly train it with the LLM for long video understanding. Due to the lack of long video QA data, we manually constructed a set of long video QA pairs in two ways. On the one hand, we concatenate short videos from existing datasets into longer ones, where the original questions correspond to different segments. On the other hand, we curate a subset of Panda-70M which includes captions for segmented clips as well as the original long videos, and use this to create multi-round long video QA pairs with explicit timestamps. These long video QA data not only optimize the responses from the LLM but also guide accurate memory selection.
In summary, our contributions are as follows: (1) We analyzed the challenge of long video understanding in the vision language area, and pointed out that the problem of current methods lies in the inefficient video encoding. (2) In response to the challenges, we propose two efficient designs: Memory-Propagated Streaming Encoding and Adaptive Memory Selection, which result in our advanced video understanding model VideoStreaming. (3) The extensive experiments demonstrate that our model achieves precise temporal grounding with respect to specific questions, attains superior performance, and exhibits higher inference efficiency on long video benchmarks.
Related Work
Large Language Models (LLMs) have revolutionized natural language processing. Early works establish encoder-decoder models with masked language modeling , while later decoder-only models like GPT showcase remarkable performance and scalability. Recent groundbreaking works, such as PaLM , LLaMA and GPT-4 , have pushed the boundaries by developing significantly larger models with billions of parameters. To harness the full potential of LLMs, a series of works adopt supervised instruction tuning to guide models towards generating more natural and contextually relevant responses. Inspired by the powerful reasoning capacities of LLMs, we explore using LLMs for challenging long video understanding.
Vision Language Models like CLIP employ contrastive learning on image-text pairs to formulate a unified embedding space . Later, integrate image features into LLMs and achieve promising visual reasoning in image domain. Considering video as a prevalent visual signal , some works further expand the application to process more complex spatio-temporal video data. use sparse sampling or simple temporal pooling to obtain compact video tokens for LLMs. employ Q-Former to project frame-wise features into the textual space. To handle longer videos, utilize token merging to reduce redundancy and alleviate computational burden. LLaMA-VID proposes an instruction-aware compression strategy to represent each frame with only two tokens, but it overlooks the temporal relations in the compression step. develop memory banks to accumulate information in long videos and excel in global video comprehension. However, these methods struggle with moment-specific questions without explicit time indicators. To address these limitations, we propose a memory-propagated streaming encoding architecture with adaptive memory selection, which effectively reduces temporal redundancy and accurately selects relevant information for detailed question answering.
Long Video Understanding is a challenging task in computer vision. The most prevalent strategy is to maintain a memory bank to store history information in long videos . To ensure computation efficiency, it is crucial to compress the history into a finite-length memory, which is typically done by parametric or non-parametric compression modules. More recently, use language as a bridge for long-term video understanding. They first divide a long video into short clips, generate textual descriptions for each clip, and then employ an LLM to aggregate the short captions for long video analysis. However, this architecture cannot be trained end-to-end, and the long video understanding quality depends on the short clip captions. In contrast, we employ a trainable small language model to iteratively encode short clips into compact memories, which can be jointly optimized with the subsequent LLM on long video understanding tasks.
VideoStreaming
In this section, we introduce VideoStreaming, a streaming long video understanding framework with LLM. As illustrated in Fig. 1a, given a long video input, VideoStreaming segments it into multiple short clips and iteratively encodes each clip into compact historical memory. To enhance the reasoning ability to specific questions, we design an adaptive memory selection strategy to select a subset of relevant memories and feed them into an LLM to produce detailed responses.
To effectively distill the information within a sequence into a compact set of tokens, we take inspiration from recent advanced decoder-only language models and employ a comparatively small language model, Phi-2 , for efficient encoding. Due to the causal attention and autoregressive nature, the language model spontaneously aggregates the sequence information onto the last few tokens , which naturally serve as a compact representation that provides a high-level summary of the input sequence.
where denotes concatenation operation, is the channel dimension of Phi-2.
To reinforce the visual consolidation ability, we design a prefix task to train the encoder on visual captioning and question-answering tasks. In particular, to guarantee that the clip information is distilled into the summarization tokens, we enforce the language model to generate the response only with reference to these few tokens. To achieve this goal, a straightforward way is to modify the attention mask in each Transformer decoder layer. As depicted in Fig. 3, we take a sequence covering clip feature tokens, summarization tokens, and text response tokens as an example. Based on the standard causal attribute, the binary attention mask is modified as shown in Figure 3:
with the modified attention mask, the text tokens can only get video-related information from the summarization tokens to predict the next token. This encourages the summarization tokens to extract more video information from previous video clip tokens, ie. learns better video encoding.
2 Memory-Propagated Streaming Long Video Encoding
Till this point, we have obtained an encoder capable of distilling short video clips into condensed representations. The next step is to comprehensively consider the long-term temporal relations within the complete videos, leveraging the historical information from previous clips to facilitate the encoding of subsequent segments as depicted in Fig. 1b.
Note that for the first clip encoding, the historical memory is not used. Through this streaming encoding process, not only encompasses the current clip information but encapsulates the overall video content up to the -th clip. To this end, we manage to maintain a fixed length of memory to represent arbitrarily long videos.
Discussion. In this architecture, we use a language model for video encoding, which has the unique advantage that we can flexibly provide the encoder with diverse prompts to guide the encoding process. Hence, the summarization tokens capture not only the core content but also additional contextual information. Typically, the explicit timestamp is an important cue in videos . As shown in Fig. 1b, we incorporate a text prompt indicating the specific timestamps of each clip and historical memory to enhance temporal awareness. Besides, this prompt-based approach also allows the user to tailor the condensed output to better suit the needs of downstream tasks, going beyond a purely extractive summarization.
Another noteworthy point is that in the language model, the feature space of the final decoder layer is designed for the next token prediction, which may not perfectly align with the objective of producing condensed video representations. Considering that we modify the attention masks in each decoder layer to encourage information consolidation, this allows us to leverage the intermediate outputs from partial attention layers as the encoded results. Similar to the techniques in vision domain , this strategy potentially enables the model to capture a richer set of semantic and contextual features as the condensed representations, bridging the gap between the language model’s original training objective and the requirements for video encoding.
3 Adaptive Memory Selection
Through the streaming video encoding, it is feasible to use the encoded results from the final iteration, i.e., , as a compact global memory that concludes the entire video. However, this fixed-length memory inevitably loses details, especially the information from early segments. Hence, this global memory alone is insufficient for comprehensive long video understanding.
Based on , we select the corresponding encoded results from to formulate a subset of memories that are related to the instruction:
where denotes the selected indexes. We concatenate the selected memories in temporal order, resulting in a sequence consisting of tokens. Then, we feed the sequence with instruction texts into an LLM for comprehensive reasoning.
Our adaptive memory selection allows the model to dynamically access historical memories relevant to specific instructions, which mitigates the information loss inherent in the streaming encoding process. By drawing upon fine-grained details across the full video duration, the LLM can provide detailed and informative responses, while preserving high computational efficiency.
4 Progressive Training
To train VideoStreaming, we design a progressive two-stage paradigm. First, we train single clip encoding on image and short video understanding tasks. Next, we train memory-propagated streaming encoding and adaptive memory selection as well as the LLM for long video understanding.
Single Clip Training. In this stage, both image- and video-text pairs are used to train the encoder to handle general visual signals. Following , we employ 790K image and short video caption data to train the MLP projector for modality alignment. After that, we employ 763K image and video instruction data from to finetune the small language model. For video input, we uniformly sample frames with spatial resolution and use a frozen CLIP ViT-L/14 to extract frame-wise features. After adjacent token merging, we obtain tokens as the clip feature representation. Then, the encoder, a two-layer MLP and a small language model Phi-2 2.7B , distills each frame into tokens, resulting in tokens as the condensed representation with a compression ratio of . For image-text pairs, we regard the images as single-frame clips and encode each into tokens. We use standard next token prediction to consolidate visual contents into compact summarization tokens as illustrated in Fig. 3.
Streaming Long Video Training. In the second stage, we use long video QA pairs to finetune the whole architecture, including ViT, the streaming encoder, and the LLM, as shown in Fig. 1a. The long video QA data encompasses three parts. (1) We adopt 25K movie QA pairs from . (2) We curate a subset from Panda-70M , which provides the original long videos and the captions of segmented clips. Based on this subset, we create 300K multi-round long video QA pairs with explicit timestamps. (3) We synthesize 20K long videos by concatenating short videos from existing QA datasets , and the original QA pairs correspond to different segments in the synthesized long videos. For each video, we extract 16-frame clips at 1 FPS, and the number of clips varies with the video duration. In streaming encoding, we employ the intermediate outputs from the first layers of Phi-2 as the condensed memories. Finally, we select most relevant timestamps and feed the selected memories of tokens into the LLM, Vicuna-7B , for long video reasoning. Since our curated long video data could provide pseudo temporal grounding labels of specific questions, we utilize 30K QA pairs to warm up memory selection via a KL divergence loss. Subsequently, we use the rest 315K QA pairs to optimize the responses from the LLM and guide memory selection in a weakly-supervised manner. More training details are included in Appendix A.
Experiments
We evaluate our model on long video QA datasets and present the statistics on the temporal duration of individual datasets in Table. 2. Among them, Next-QA , Next-GQA and VideoChatGPT encompass minute-long videos with thousands of frames. EgoSchema contains over 5K three-minute videos with multiple-choice questions. Each question has a long temporal certificate, requiring more than 100 seconds within a video to produce a correct answer. MovieChat-1K and MovieNet-QA consist of around ten-minute-long or even hour-long movies, posing significant challenges for the model to comprehend the visual contents across such long time spans.
2 Main Results
In this section, we present the results of our 8.3B model (half of Phi-2 2.7B in streaming encoder and Vicuna-7B as the LLM). We omit the comparisons to proprietary LLMs.
VideoChatGPT. Table 2 presents the results on VideoChatGPT in terms of Correctness of Information (CI), Detailed Orientation (DO), Contextual Understanding (CU), Temporal Understanding (TU) and Consistency (CO). Our model outperforms LLM-based video understanding methods on all five metrics, with a significant advantage in temporal understanding. It can be attributed to the memory-propagated streaming encoding architecture that explicitly captures temporal dynamics.
EgoSchema. In Table 4, we report the zero-shot performance on the fullset test split of EgoSchema . MC-ViT consolidates a long-term memory to memorize long contexts but requires finetuning on related dataset . LLM-based methods curate answers from the captions of segmented video clips. However, these short-term captions cannot be optimized end-to-end and inevitably lose some detailed information. In contrast, we use a trainable streaming encoder to produce memory embeddings in long videos and feed them into an LLM to generate responses. Our model outperforms all zero-shot methods and is comparable to the finetuned MC-ViT, demonstrating the effectiveness of our streaming architecture for long-term temporal modeling.
Next-QA. In Table 4, we perform zero-shot evaluation on the validation split of Next-QA covering 5K multiple-choice questions. We respectively report the accuracy on Causal (C), Temporal (T) and Descriptive (D) subsets. Our method consistently surpasses all zero-shot counterparts. Typically, compared to LangRepo with Mixtral-87B , our 8.3B model improves the causal, temporal, and descriptive accuracy by 0.7%, 10.8%, 9.0% with considerably fewer model parameters.
Next-GQA. Besides the evaluation of the generated responses, we also assess the temporal grounding ability on Next-GQA . We calculate the Intersection of Prediction (IoP) and Intersection of Union (IoU), and use Acc@GQA to measure the accuracy of the correctly grounded predictions. According to the comparisons in Table 5, our simple similarity score based selection achieves the highest IoP and comparable IoU to SeViLA with a specialized grounding module. Moreover, the highest Acc@GQA demonstrates the comprehensive capacity for grounding and high-level understanding.
MovieChat-1K. Table 4 shows the results on MovieChat-1K , including a global mode for overall long-term understanding and a breakpoint mode for detailed analysis of specific moments. In breakpoint mode, manually extract segments according to the timestamps in questions, while our model adaptively selects the related historical memories. Fig. 4 reveals that our selected timestamps are close to the ground-truths, and the higher breakpoint accuracy validates our adaptive selection effectively gathers the desired information from long contexts. Meanwhile, we reach significantly superior results in global mode, with the model’s selection concentrated at the beginning and ending parts. On the one hand, the beginning of a movie often contains hints of global information while the middle comprises redundant details. On the other hand, the condensed memories near the end of the video encapsulate the entire video, making them quite suitable for global understanding.
MovieNet-QA. Finally, we show the results on MovieNet-QA consisting of 100 hour-long movies. Inspired by , we use GPT-3.5 to produce scores in range 0-5 to evaluate the performance in overview, plot, and temporal understanding in Table 6. Specifically, LLaMA-VID compresses each frame into two tokens, which are then combined with movie subtitles as input to an LLM. MovieLLM further incorporates more generated data in training. These approaches largely rely on the texts for movie understanding, and only using visual frames leads to dramatic performance drop. Moreover, its frame-wise compression is dependent on specific questions. The model has to reprocess the entire movie to extract visual features for different questions, resulting in a high inference latency of over 10 seconds per question. Conversely, our architecture requires only once streaming encoding to obtain a general condensed representation and adaptively selects significantly fewer tokens as input to LLM to answer specific questions. Therefore, we achieve a higher inference speed of 5.32 seconds per question and attain promising movie understanding without using subtitles.
Qualitative Results. We also present qualitative examples in Fig. 5. Typically, in Fig. 5a, our model accurately captures the detailed descriptions in the question, and precisely selects the relevant segments that contain the corresponding character. Moreover, in Fig. 5b, given a two-hour long movie and a high-level question on the movie plot, without relying on subtitles, VideoStreaming can comprehend the intent of the question and select relevant scenes from the lengthy video. In particular, the model selects the scenes of tightrope walk, team disputes, and equipment setup, clearly illustrating the protagonist’s challenges, thereby contributing to a comprehensive answer generation.
3 Ablation Study
Historical Memory. We explore the influence of memory in the streaming encoding process, i.e., in Eq 2. We report the fullset accuracy on EgoSchema as well as global and breakpoint accuracy on MovieChat-1K in Table 8. Typically, the historical memory significantly improves global understanding by 46.6%. This verifies our intuition that leveraging historical memory enables the model to produce a global representation that summarizes the entire video. Meanwhile, since we select a small portion of the encoded results from the long video as input to LLM, the lack of historical memory limits the temporal respective field and impairs the performance.
Memory Selection. We also validate the effects of our memory selection strategy. For comparison, we directly use the encoded results from the final four iterations as input to LLM and present the result in the first row of Table. 8. The historical memories in streaming encoding process enable the encoded results from the final iterations to provide coarse summarization of the entire video, thus attaining satisfactory results on global understanding. However, for questions regarding detailed analysis of specific moments, the lack of temporal selection leads to 31.9% performance drop in breakpoint mode accuracy. It demonstrates the effectiveness of our adaptive selection in gathering detailed information over the long time span, which facilitates more accurate and informative responses.
Streaming Encoder Architecture. Besides, we ablate the number of layers in Phi-2 used in memory-propagated streaming encoding. We show the results on EgoSchema and Next-GQA as well as the number of encoder parameters in Table 8. Interestingly, using fewer layers leads to better results. We conjecture this is because the language model is originally trained for next token prediction. Its feature space of the final Transformer decoder layer might not align with the objective of visual content condensation. Similar to , the shallower layers might produce feature embeddings that encode richer information and serve as more comprehensive condensed video representations.
Temporal Grounding Supervision. First, we present the studies on the use of temporal grounding supervision. As mentioned in Section 3.4, we employ around one-tenth of long video QA pairs to provide pseudo temporal labels. We compare four training strategies: (1) Fully weakly-supervised manner without any pseudo labels. (2) Using pseudo labels to train a warm-up model, then expanding to large-scale QA pairs. (3) Mixing all long video QA data, where the model uniformly receives temporal supervision in training. (4) Training on mixed data after warm-up initialization. The results on EgoSchema and Next-GQA in Table 6 indicate three key points: First, warm-up training contributes to more powerful grounding ability. The sparse temporal label supervision in mixed mode is overcome by the powerful initialization from warm-up training, which can generalize to large-scale data. Second, reusing the temporal labels after warm-up offers no additional benefits, so we adopt warm-up as the default setting. Third, without using temporal labels, the grounding performance drops, but the QA accuracy remains stable. Fig. 6 reveals that compared to those trained with temporal labels, the weakly-supervised model selects relatively later segments that preserve previous contexts with the help of historical memory, thus maintaining comparable QA capacity.
More ablation studies on the number of summarization tokens and selected timestamps, the time prompts, and the similarity measurement are included in Appendix C.
Conclusion
In this paper, we introduce a novel approach to tackle the complexities of long video understanding with large language models (LLMs). Our proposed memory-propagated streaming encoding architecture segments long videos into short clips and iteratively encodes each clip in sequence. By leveraging historical memory from preceding clips, we incorporate temporal dynamics into the encoding process and produce a fixed-length memory to encapsulate arbitrarily long videos. To further augment the detailed information for handling specific questions, we develop adaptive memory selection that selects relevant timestamps based on given instructions. This approach ensures that the most pertinent historical memories are utilized for question answering, thereby facilitating detailed and informative responses. Our model achieves superior performance with substantially fewer tokens and higher efficiency on extensive long video benchmarks. We demonstrate that memories from the streaming encoding significantly enhance global video understanding, while adaptive selection results in accurate temporal grounding with respect to specific questions.
References
Limitations
One potential limitation is that we simply uniformly sample frames to form a set of short clips for memory-propagated streaming encoding. However, in a long video, different segments possess different amounts of information. The uniform sampling may result in using redundant tokens for clips with bland content. Meanwhile, the number of tokens used to represent clips with abundant visual contents and intensive temporal dynamics may be insufficient, leading to information loss. To address this limitation, we plan to explore adaptive segmentation techniques that dynamically adjust the segmented clip lengths based on the complexity and content of the video.
Impact Statements
Our proposed VideoStreaming, a streaming long video understanding architecture with large language models has various potential impacts for society. On the positive aspect, VideoStreaming contributes to improved intelligent video understanding, especially for long videos. This could be beneficial in education, entertainment, and information retrieval, where users often need to navigate and understand complex video materials. Besides, our technique could lead to advancements in multimedia analytics with applications in areas like video surveillance, market research, and content personalization.
On the negative aspect, the ability to efficiently process and retrieve information from long videos raises potential privacy and security concerns. If misused, this technology could be employed for unauthorized surveillance, personal monitoring, or other unethical purposes that infringe on individual privacy. In addition, the enhanced video understanding capabilities might be exploited for the creation of manipulated or misleading video content, leading to the spread of misinformation and the potential for social manipulation.
In conclusion, despite that VideoStreaming presents advancement in long video comprehensive, its development should be accompanied by careful consideration of ethical and societal implications.
Appendix A More Implementation Details
We use CLIP ViT-L/14 to extract frame-wise features with input resolution , resulting in tokens per frame. Then, we concatenate every four spatially adjacent visual tokens along channel dimension, representing each frame with tokens with channel dimension 4096. The streaming encoder consists of a two-layer MLP projector (channel dimension 4096-2560-2560) with GELU activation and a language model Phi-2 2.7B . In the first training stage, we initially freeze Phi-2, and only tune the MLP projector on 790K caption pairs, including 558K image caption data from CC3M and 232K short video caption data from WebVid 2.5M . Following LLaVA , we use AdamW optimizer with global batchsize 256, initial learning rate with cosine decay to train 1 epoch for modality alignment. Subsequently, we jointly train Phi-2 and the MLP projector on 763K QA pairs, including 625K image QA pairs , 40K text conversations and 98K video QA pairs , with global batchsize 128, initial learning rate with cosine decay.
In the second stage, we jointly train ViT, the streaming encoder and the LLM on long video data. In the memory-propagated streaming encoding process, we insert a brief prompt to indicate the explicit timestamps of the historical memory and the input clip formulated as This contains a history of {start} to {end} seconds, and a clip sampled in {start} to {end} seconds.. We adopt the output of the first 16 layers out of the 32 layers of Phi-2 as the condensed representation. Then, we adaptively select 4 most relevant timestamps and feed the associated memory tokens into a two-layer MLP projector with channel dimension 2560-4096-4096 and an LLM, Vicuna-7B to generate the final responses. We jointly train the whole architecture, including Vicuna, Phi-2, MLP projectors and ViT encoder, on long video QA data with global batchsize 128, initial learning rate with cosine decay. In default, we first use 20K synthesized long videos and sample 10K QA pairs curated from Panda-70M with pseudeo temporal grounding labels to train memory selection as warm-up. The learning objectives contain a standard next token prediction loss and a supervised KL divergence loss that aligns the distribution of the predicted memory selection results and the pseudo temporal labels. Next, based on the warm-up model, we further train on the rest 295 long video QA pairs only with next token prediction loss. The whole training is conducted on 32 A100 (80G) GPUs for around 2.5 days.
Appendix B Long Video QA Data Creation
In addition to the existing 25K long video QA pairs on movies , we create more QA data from two aspects. First, we leverage the existing short video QA dataset and synthesize short videos into minute-long videos with average duration of one minute. The original questions of each short video coarsely correspond to a temporal segment in the synthesized long video. We use this correspondence as noisy labels to supervise the memory selection. Second, recent Panda-70M segments long videos into short clips and produces captions for each clips. This dataset provides the original long videos, the captions of segmented clips as well as the segmentation timestamps. Based on these cues, we produce multi-round QA conversations. Below we show an example in Fig. 7. The produced time-sensitive QA pairs are crucial to enhance the temporal awareness and guide precise memory selection in long videos.
Appendix C More Ablation Studies
We provide more ablation studies on the number of summarization tokens and selected timestamps, the effects of time prompts in the memory-propagated streaming encoding process, and the similarity measurement used in memory selection.
The Number of Summarization Tokens and Selected Timestamps. We compare using different number of summarization tokens and selected timestamps, i.e., in Eq. 1 and in Eq. 3. We compare the performance as well as the number of tokens input to LLM in Table 9. We conclude three observations. First, too few summarization tokens, e.g., , leads to substantial performance drop, since it condenses a 16-frame into only 16 tokens with significant information loss in spatial contexts. Such information loss cannot be compensated by selecting more temporal segments. Second, the performance saturates when improving from 4 to 16. This is because the existing video benchmarks do not place high demands on spatial detail understanding. It is sufficient to represent each frame with 4 tokens on average. Third, increasing the number of selected timestamps only results in minor improvements, which is not proportional to the increased number of tokens. This can be attributed to the historical memory used in the streaming encoding process. The utilization of historical memory enables the condensed representation of each clip to encompass the information in preceding clips, which enlarges the temporal receptive field. Hence, increasing the number of selected timestamps does not proportionally increase the temporal receptive field, resulting in slight performance improvements.
We explore three different formulations of the time prompts used in memory-propagated streaming encoding: (1) Only with the timestamps of the current clip, e.g., This clip is sampled in {start} to {end} seconds. (2) Only with the timestamps of the historical memory, e.g., This contains a history of {start} to {end} seconds. (3) Simultaneously with the timestamps of the historical memory and the current clip, e.g., This contains a history of {start} to {end} seconds, and a clip sampled in {start} to {end} seconds. We report the results of different time prompts in Table 10. It is obvious that the lack of time prompts leads to substantial performance drop in the MovieNet-1K breakpoint mode accuracy, which requires detailed analysis of specific moments. The reason is that the breakpoint mode requires the model to answer questions at specified timestamps, the time prompts provide the model with necessary information in adaptive selection. Meanwhile, incorporating the timestamps of historical memories results in more significant improvements in global understanding. Overall, jointly leveraging the memory and clip timestamps contributes to the best results.
Similarity Measurement. Finally, we present the study on the similarity measurement used in adaptive memory selection. We compare the default cosine similarity against simple dot product without normalization in Table 11. Empirically, we observe that dot product could result in numerical instability, leading to overflow in training. Consequently, the calculated similarity score cannot reflect the correlation between the instruction and different segments and results in poor results on questions that require accurate temporal grounding, e.g., the Acc@GQA metric on Next-GQA and the breakpoint mode accuracy on MovieChat-1K .