LongVLM: Efficient Long Video Understanding via Large Language Models

Yuetian Weng, Mingfei Han, Haoyu He, Xiaojun Chang, Bohan Zhuang

Introduction

Large language models (LLMs) have revolutionized natural language understanding tasks and have demonstrated a remarkable capability to follow human instructions and intentions, emerging as a universal agent for general-purpose assistants. Drawing from the development of LLMs, Multi-modal Large Language Models (MLLMs) have driven the advance in vision-language learning by integrating visual encoders with LLMs and finetuning on language-image instruction-following data. However, developing Video-based Large Language Models (VideoLLMs) still poses a significant challenge due to the necessity of processing a large number of tokens for jointly modeling spatial-temporal dependencies across consecutive video frames. For instance, employing OpenAI CLIP-ViT-L/14 as a visual encoder for a 100-frame video clip necessitates handling 25.6K visual tokens, leading to impractical computational costs with existing LLMs. To address this, recent approaches propose to extract video representation via precompression over visual tokens, utilizing pooling operation or query aggregation over the video token sequence before feeding them into the LLM, as shown in Fig. 1(a). While these models showcase impressive capabilities in providing a meaningful understanding of video content, they still face challenges in achieving a significant advantage in fine-grained understanding of long-term videos. For example, as shown in Fig. 1(b), while all models recognize the overall environment (workshop), the object (bike), and the action (fixing), previous methods may fail to correctly identify details such as the color of the clothes (brown), or the specific component being fixed (bicycle chain).

The main reason is that long-term videos typically involve numerous key actions, complex activities, and camera movements. Consequently, a long video can be divided into a sequence of short-term segments. For instance, in the example depicted in Fig. 1(b), various short-term actions occur, e.g. speaking, displaying spare components, grabbing the bicycle chain, along with the camera moving from the human to the bicycle wheel, and eventually focusing on the broken chain. Similarly, prior methods in video recognition task suggest to decompose complex activities into sequences of sub-activities . These approaches treat the features of each short-range activity as the local information within the videos and emphasize the importance of reasoning over local features to develop a temporal-structural understanding within long-term videos for comprehending fine-grained information. From this perspective, existing VideoLLMs treat all visual tokens equally and aggregate them into compact representations through pooling operations and query aggregation . While they successfully capture the global semantic context spanning the entire long-term videos, they often overlook preserving the local information for the short-term segments and the temporal structure of different short-term components, e.g., the order of events or sub-actions. However, exclusively modeling the temporal structure through the sequence of local features may still lead to inconsistent recognition across different segments and impede the overall understanding of the videos. To comprehend the content in long videos, the human visual system relies on a blend of local and global information . Building on this insight, earlier approaches in video object detection suggest integrating global semantics into local localization descriptors, motivating us to include global semantic information into the sequence of local features for enriching the context understanding for each short-term segment.

In this paper, we present LongVLM, a simple yet effective VideoLLM for efficient long video understanding, as illustrated in Fig. 2. We propose to extract video representations as sequences of short-term local features, and integrate global semantics into each short-term segment feature. Specifically, we begin by uniformly sampling a sequence of video frames from long-term videos and utilize a pretrained visual encoder, e.g., CLIP-ViT-L/14 , to extract visual features for each individual video frame. These frame-level features include the [CLS] tokens from a range of encoder layers and the patch features from the last second layer of the visual encoder. Then, we divide the sequence of patch features along the temporal dimension, resulting in a sequence of short-term segments. Each segment is considered as a local unit in the videos and includes patch features of the video frames within that segment. To reduce computational costs and obtain the compact features for each segment, a token merging module is employed to aggregate these patch features for the specific segment into a condensed set of tokens. In this way, we obtain the local features for each segment. We next concatenate these features sequentially to explicitly preserve the temporal order of the short-term segments in long-term videos. Moreover, we average the [CLS] tokens from each video frame along the temporal dimension to represent the global semantic information of the entire video. To integrate the global information, we sequentially arrange the averaged [CLS] tokens and the segment-level features, and then feed them into the LLM after passing through a projection layer. Benefiting from the causal attention mechanism in the LLM, we simultaneously achieve temporal structure modeling over the sequence of short-term segments and inject global semantics into the local features. Finally, the LLM generates responses based on the input sequence, which is composed of the obtained video representation and the designed system command with the specific user queries.

Overall, our main contributions is threefold:

We propose LongVLM, a simple yet effective VideoLLM for efficient long-term video understanding at a fine-grained level while maintaining affordable computational cost.

We propose decomposing long videos into short segments and extracting local features for each segment to preserve their temporal order. To compactly represent each segment, we propose a hierarchical token merging module to aggregate visual tokens. Additionally, we integrate global semantics into each segment to enhance context understanding.

Extensive experiments on VideoChatGPT benchmark and zero-shot video question-answering datasets demonstrate that our LongVLM surpasses the previous state-of-the-art methods by a significant margin while generating more precise and accurate response at fine-grained level for long-term videos.

Related Work

Large Language Models (LLMs) have revolutionized natural language processing in recent years. Pretrained on large text corpora, LLMs like GPT , OPT , and LLaMA utilize auto-regressive Transformer models to predict subsequent tokens, showcasing remarkable adaptability and generalization. Models such as InstructGPT , ChatGPT , and GPT-4 benefit from instruction-tuning techeque on instructional datasets, leveraging the knowledge of pretrained LLMs and demonstrating improvements in diverse conversational interaction capabilities. This strategy is widely adopted in open-source models like Alpaca and Vicuna , which build upon the advancements made by LLaMA using specially designed instruction pairs. Drawing from the advancement of LLMs, recent Multi-modal Large Language Models (MLLMs), e.g., BLIP-2 , Mini-GPT4 , LLaVA , LLama Adapter v2 , have demonstrated the feasibility of enabling visual conversation capabilities of LLMs over input images through instruction tuning on image-text instruction datasets. Our model aims to utilize existing MLLMs to develop efficient video dialogue model for long-term video understanding.

2 Video-based Large Language Models

Building upon pretrained image-language models, traditional approaches are proposed to address various video-language learning tasks.With the advent of LLMs, recent approaches are exploring the potential of developing Video-based Large Language Models (VideoLLMs) to unify various video-language understanding scenarios through human-video dialogue interactions. Existing VideoLLMs typically follow a common paradigm, which involves using a pretrained visual encoder to encode visual features, a projection layer to convert visual representations into the text latent space of LLMs, and a pretrained LLM for response generation. VideoChatGPT and Valley rely on pooling over visual tokens to obtain compact visual representations. VideoChat utilizes pretrained video foundation models and Q-Former from BLIP-2 to aggregate video representations. Video-LLaMA proposes a Video Q-Former and an Audio Q-Former, enabling multiple modalities for video comprehension, while Video-ChatCaptioner employs ChatGPT to summarize video descriptions in multiple rounds of interactive question-and-answer conversation. Recently, MovieChat proposes an effective memory management mechanism to enable LLMs to reason over hour-long videos. Multiple video-centric instruction datasets have also been proposed to finetune VideoLLMs for better video understanding capacity. Moreover, BT-Adapter proposes a temporal adapter alongside the visual encoder for post-pretraining, while Video-Teller highlights the importance of modality alignment in pretraining. Overall, these VideoLLMs rely on pooling and query aggregation on the whole long videos to extract visual representation for developing VideoLLMs, which lack of local information for fine-grained understanding in long videos. In contrast to existing VideoLLMs, we propose a simple yet effective framework that is feasible for aggregating both local and global information in long-term videos and preserves fine-grained content understanding.

3 Long-term Video Processing

Long-term video understanding poses several challenges due to the need to exploit complicated spatial-temporal dependencies while removing temporal redundancy over extended time duration. Previous studies propose efficient architectures , temporal pooling/aggregation , dynamic clip selection to aggregate video representation while removing redundant information in videos. Other methods in video-language understanding tasks suggest to capture event temporality, causality, and dynamics in long-term videos by designing temporal alignment modules . Memory mechanism is also widely adopted in video dense prediction tasks to capture historical information and maintain temporal coherence, which results in more accurate and consistent prediction over time in long-term videos. In contrast to previous methods, we investigate the techniques for long-term video understanding in VideoLLMs. We propose to aggregate both local segment-level information and global semantic information, empowering MLLMs enhanced fine-grained understanding for long-term videos.

Method

In Sec. 3.1, we introduce the overall architecture and generation pipeline of the proposed LongVLM. In Sec. 3.2, we introduce the process of constructing local representation via short-term feature aggregation. In Sec. 3.3, we discuss the integration of both local segment-level feature and global semantic feature for the video representation.

The overall architecture consists of three components: a visual encoder, a projection layer, and a large language model, as illustrated in Fig. 2.

To enable fine-grained understanding in long videos, we propose to divide long videos into a sequence of short-term segments, where each segment corresponds to the local features in the long videos. Without loss of generality, the input video V{\mathcal{V}} is divided into SS segments, where each segment includes KK frames, i.e., K=TSK=\frac{T}{S}. We collect patch features within the sths^{th} segment, i.e., Vs={Pt}t=(s−1)Kt=sK{\bf V}^{s}=\{{\bf P}^{t}\}_{t=(s-1)K}^{t=sK}, and apply a token merging module G(⋅){\mathcal{G}}(\cdot) to aggregate Vs{\bf V}^{s} into the compact segment-level feature Zs=G(Vs){\bf{Z}}^{s}={\mathcal{G}}({\bf V}^{s}). These segment-level features are sequentially concatenated as the sequence of local representation L{\bf L}, explicitly preserving temporal order of short-term segments in long videos. Furthermore, to integrate global semantic information, we propose to collect the [CLS] tokens for each frame from EE encoder layers and average them in the time dimension, resulting in our global feature G{\bf G}.

We forward the concatenated global features and the sequence of local features into a projection layer FV(⋅;W){\mathcal{F}}_{V}(\cdot;{\bf W}) to obtain the projected visual feature, i.e., HV=[G^,L^]=[FV(G;W),FV(L;W)]{\bf H}_{V}=[\hat{{\bf G}},\hat{{\bf L}}]=[{\mathcal{F}}_{V}({\bf G};{\bf W}),{\mathcal{F}}_{V}({\bf L};{\bf W})]. The projected visual feature HV{\bf H}_{V} are concatenated with the tokenized system command and user queries Hq{\bf H}_{q}, which are inputted into LLM for response generation Ha{\bf H}_{a}, i.e., Ha=FLLM([HV,Hq]){\bf H}_{a}={\mathcal{F}}_{LLM}([{\bf H}_{V},{\bf H}_{q}]), where FLLM(⋅){\mathcal{F}}_{LLM}(\cdot) denotes an LLM model.

2 Local Feature Aggregation

3 Global Semantics Integration

Following the previous studies, a projection layer converts the visual features into the language space, and then the visual features are concatenated with the instruction as the input of LLM. By utilizing the attention mechanism in the LLM, we can easily enable each token in the local feature to attend to the global semantic feature, thereby achieving straightforward injection of global semantics into the local feature.

Remark. To address the risk of overlooking detailed understanding in long-term videos, we propose to divide long videos into multiple short-term segments and aggregate local spatial-temporal representation for each segment and preserving the temporal structure over the sequence of local feature vectors. Moreover, we enrich the local features with context information for better response generation by integrating global semantic information into short-term features. Different from previous approaches that rely solely on global semantics for long video understanding, we present a straightforward yet effective VideoLLM for achieving fine-grained understanding in long-term videos.

Experiments

Datasets and evaluation metrics. We conduct quantitative evaluations of our model using the VideoChatGPT benchmark to assess its performance in generating text from videos. The benchmark comprises 500 videos sampled from ActivityNet-v1.3 dataset , with 2000, 2000, 2000, 500, and 1000 questions in terms of five evaluation aspects: Correctness Information(CI), Detail Orientation(DO), Contextual Understanding(CU), Temporal Understanding(TU) and Consistency(C). Additionally, we evaluate the model on the zero-shot question-answering task using the ANET-QA dataset, which contains 8000 QA pairs for 800 videos sampled from ActivityNet-v1.3 dataset . The videos range from several seconds to minutes long and cover a wide range of daily human activities. We also utilize MSRVTT-QA (72821 QA pairs for 2990 videos) and MSVD-QA (13157 QA pairs for 520 videos) to evaluate the model performance, derived from publicly available video captioning, MSRVTT , and MRVDC , respectively. Following the evaluation protocol outlined in Video-ChatGPT , we employ ChatGPT for response evaluation and report the generation quality scores on VideoChatGPT benchmark and the answer accuracy and quality scores of models on zero-shot video QA tasks.

Implementation details. We employ CLIP-ViT-L/14 as the visual encoder and Vicuna-7B-v1.1 as the LLM. We initialize them with the pretrained weights in LLaVA-7B-v1.1 . We finetune the model on the Video-ChatGPT-100K instruction dataset for 3 epochs, with a learning rate of 2×e−52\times e^{-5} and a batch size of 32. We only finetune the linear projection layer to align the visual features into the input space of the LLM, keeping both the visual encoder and LLM frozen. It takes three hours to train three epochs on 4 A100 80GB GPUs. During training and inference, we sample T=100T=100 video frames for each video, and resize the frames to 224×224224\times 224 resolutions. We set S=10S=10 for each video, and the number of tokens in each segment-level feature is M=30M=30. We collect the [CLS][CLS] tokens from the last five encoder layers and average them along the temporal axis, resulting in E=5E=5 tokens as the global semantic features. Therefore, the length of visual tokens for a video sequence is M×S+E=305M\times S+E=305.

2 Main Results

Results on the video-based generation benchmark. In Tab. 1, we present a comprehensive evaluation of our LongVLM against state-of-the-art models on the video-based generation benchmark . Our LongVLM outperforms all other models across all the evaluation aspects. Particularly noteworthy is its significant advantage in Detail Orientation (DO) and Consistency (C), showing improvements of +0.17 and +0.65, respectively, over BT-Adapter . These results underscore superior capability of LongVLM in fine-grained video understanding and robust generation performance.

Results on zero-shot video question-answering. In Tab. 2, we compare the performance of LongVLM against various existing methods on three zero-shot video QA datasets: ANET-QA , MSRVTT-QA and MSVD-QA . Our model achieves the highest accuracy of 47.6%, 59.8%, and 70.0% on the three QA datasets, surpassing the previous SOTA approach BT-Adapter by 1.9%, 2.8% and 2.5%, respectively. Furthermore, we achieve the highest score in terms of generation quality over the three datasets.

3 Ablation Study

Effects of local feature aggregation. As discussed in Sec. 1, pooling operations or query aggregation might overlook local information in achieving fine-grained understanding in long-term videos. To this end, we introduce short-term segment-level features to retain local information and temporal structure within long-term videos. The first two rows in Tab. 3 present the effects for the design of using local features as the visual representations for videos. We compare the token merging module with local pooling operation using a pooling kernel and stride of within each short-term segment. The proposed hierarchical merging module achieves higher scores over spatial-temporal pooling operation. This could be attributed to the dynamic aggregation mechanism via the token similarity in the merging module, while averaging pooling statically aggregates visual tokens within each small 3D window. Additionally, we observe that aggregating local features for short-term segments either improves or maintains comparable performance across all evaluation metrics compared to the SOTA models which extract global semantics only, highlighting the significance of preserving local features for short-term segments in long-term video understanding.

Effects of global semantics integration. Inspired by the human visual system that using a combination of local and global information for recognizing video content , we propose to enhance visual representation by injecting global semantic features into local features. The last two rows in Tab. 3 demonstrate the effects of integrating global semantics. Compared to the first two rows, introducing global semantic features significantly enhances performance compared to models using local features only across all evaluation aspects. The notable improvements in Contextual Understanding (CU) and Consistency (C) underscore the significance of integrating global semantic information with local short-term features. Moreover, concatenating global features before local features yields better results than the opposite concatenation order. This arrangement allows each the local feature to access the global semantic information across the entire video by leveraging the causal attention mechanism in the LLM. Consequently, this enriches the contextual information of the local features and enhances the response consistency of the model.

Effects of MM. We report the model performance on VideoChatGPT benchmark and ANET-QA task on the selection of MM, i.e., M={10,20,30,40}M=\{10,20,30,40\}, keeping the same number of global semantic tokens in Tab. 4. For ANET-QA task, we also report the averaging GPU memory usage for generating each response. In general, the token length involves a trade-off between memory costs and performance. A shorter token sequence reduces computational costs for generating a single new token using LLM, thereby lowering memory costs for generating responses to individual user queries. However, it may also lead to insufficient visual information for generating accurate responses. The performance of our model is beneficial from the suitable length of visual tokens. Specifically, using M=10M=10 leads to the lowest performance in both tasks, while it still shows a comparable performance compared to the SOTA models, indicating the effectiveness of our model design. Increasing MM from 10 to 40 results in a significant improvement in terms of most evaluation aspects, while the setting of M=40M=40 leads to neglecting improvement but requires more memory cost compared to M=30M=30. Therefore, we choose M=30M=30 for our model.

Effects of EE. We evaluate the model performance on the VideoChatGPT benchmark with varying EE by selecting the [CLS] tokens from the last 1,5,10,15,20,241,5,10,15,20,24 visual encoder layers, while maintaining the same MM for local features. As depicted in Tab. 5, increasing the number of global semantic tokens from 1 to 5 improves the generation quality scores in terms of all evaluation aspects. However, increasing EE from 5 to 24 leads to degraded performance, possibly because the [CLS] tokens from earlier layers carry less semantic information for the model. Therefore, we choose E=5E=5 in our model.

4 Qualitative Results

As illustrated in Sec. 1, Fig. 1(b) demonstrates the advancement of our model in terms of fine-grained understanding in long-term videos. Despite taking the same number of video frames as input, our model excels in capturing detailed information within the videos, discerning nuances like fixing chain rather than fixing wheel. In comparison, Video-ChatGPT can describe the overall video content but may inaccurately recognize detailed information. For instance, it might identify objects such as helmets and gloves in the scene but erroneously recognize the location for these objects. This emphasize the importance of decomposing long videos into multiple short-term segments and aggregate local features to achieve fine-grained understanding in videos. The examples depicted in Fig. 4 ablate the effectiveness of integrating global semantic information into local short-term features. With global semantics integration, the model is able to recognize the actions (long jump) and objects (axe) compared to the variant that using local features only. We provide more generated examples in Fig. 3(b) and Fig. 5 from ANET-QA and Video-ChatGPT benchmark, respectively, which showcase the precise description of the video content generated by our LongVLM.

Conclusion

In this work, we have introduced LongVLM, an effective and efficient multimodal LLM designed for long-term video understanding. By extracting local features for short-term segments, we efficiently model local dependencies while preserving the temporal structure of sequential events in long-term video sequences. Through the integration of local and global information, LongVLM captures detailed information and provides consistent and accurate responses for long-term videos.

Limitations and Further work. While we introduce a novel video conversation model for fine-grained long-term video comprehension, our framework is specifically designed for video-to-text generation scenarios. Future work may include extending our framework into video-centric multimodal generation tasks and training the model on large-scale, extended-duration videos for long-context understanding.

References