VideoLLM: Modeling Video Sequence with Large Language Models

Guo Chen, Yin-Dong Zheng, Jiahao Wang, Jilan Xu, Yifei Huang, Junting Pan, Yi Wang, Yali Wang, Yu Qiao, Tong Lu, Limin Wang

Introduction

The advent of phenomenon-level language applications, such as ChatGPT , has showcased LLMs’ [61; 62; 7; 58; 64; 101; 75; 83] remarkable zero-shot capability in effectively addressing multiple NLP or vision-centric tasks. The remarkable sequence modeling and reasoning capabilities that these large language models exhibited can be traced back to their acquisition through rigorous pre-training with substantial parameters on large-scale corpora. Despite the amazing achievements in processing language sequences, understanding video sequences that record the real world’s objective laws and can be regarded as long image sequences is far from the level of present LLM.

Video sequence understanding involves various real-world applications, such as surveillance systems , autonomous vehicles , robotics , and wearable devices . Simply put, it involves AI systems in the real-time processing of visual information streams, reasoning them in the context of long-term time series, and then providing responses. The vanilla paradigm for video sequence understanding tasks relies on task-specific designs [93; 103; 11; 98; 51; 100; 39] to encode or decode video sequences, thereby achieving a promising performance but brings additional tailored cost. Compared with natural language, there is no scalable video sequence model that can be seamlessly adapted to different video sequence tasks. This is primarily attributed to the challenges associated with large-scale video self-supervision, which arise from the expensive nature of temporal-intensive visual annotation, as well as the time-consuming process of acquiring and processing extensive video data. As a result, there is a pressing demand for an efficient method that can offer fundamental modeling capabilities for tasks involving video sequence understanding.

In this work, we present a novel paradigm called VideoLLM, as shown in Figure 1, which aligns video and language sequences and harnesses LLMs’ reasoning and understanding capabilities. This paradigm enables videos to engage in reasoning about real-world events through the medium of language. Specifically, it is composed of three core components: (1) a temporal-wise unitization method to encode unit-wise data stream, (2) an appended semantic translator to transfer visual semantics to language semantics, and (3) a decoder-only LLM as a generalist video sequence reasoner for various video sequence understanding tasks. The design allows sequence tasks with different modalities (e.g. visual and text) to be seamlessly integrated, as we verified in the experiments visual-only tasks such as temporal action detection and action anticipation, etc., and visual-language tasks such as temporal grounding and highlight detection, etc. The unit-wise encoding and decoder-only reasoning enable the system to run with minimal delay, greatly meeting real-time or interactive systems’ experience requirements.

In contrast to the long-term temporal post-fusion approach proposed in , our method emphasizes learning short-term visual token representations for effectively integrating frozen LLMs. This adaptation is conducted within a well-pretrained LLM with robust sequence processing and causal reasoning abilities. Consequently, long-term video modeling can be disregarded, effectively simplifying the complexity of the system design. Compared to recent API-based or ensemble-based visual understanding applications [12; 97; 68; 54; 45], we offer an end-to-end system-level approach for video understanding by bridging visual models and LLMs, enhancing the overall efficiency of the long-term video sequence understanding pipeline. Moreover, our method achieves maximal decoupling between short-term and long-term visual modeling, enabling the flexible adoption of heterogeneous short-term visual encoding techniques while rapidly incorporating state-of-the-art LLMs.

Our contributions can be succinctly summarized as follows:

(1) We present VideoLLM, a novel framework that harnesses the sequence reasoning capabilities of pre-trained LLMs to tackle video sequence understanding tasks through the medium of language. By aligning videos with language, VideoLLM enables simultaneous reasoning about language logic and the evolution of real-world states through unified modeling.

(2) We reexamine the characteristics and challenges associated with various video sequence understanding tasks and develop a novel, plug-and-play adaptation scheme to adapt off-the-shelf visual encoders and advanced LLMs effectively. This scheme is built upon a unified adaptation principle, eliminating the need for task-specific customization.

(3) We conduct extensive experiments across four datasets, encompassing eight video sequence understanding tasks. These tasks encompass diverse settings, including data accessibility (causal or non-causal), perceptual objectives (memory or anticipation), prediction granularity (segment-level or frame-level), and modalities (vision-only or vision-language). The experiments employ a range of LLMs, such as GPT-2, T5, and OPT. Comparative analyses against task-specific tailored models demonstrate that our VideoLLM achieves state-of-the-art or comparable performance on these tasks, employing comparable or fewer trainable parameters. These results effectively establish LLM as an effective video reasoner, while validating the efficacy of our proposed VideoLLM framework for multiple video sequence understanding tasks.

Related Work

Video Sequence Understanding tasks can be categorized into two types based on the granularity of predictions: timestamp-level tasks and segment-level tasks. Timestamp-level tasks aim to predict closed-set properties at each time step or filter suitable time steps based on textual conditions. For example, [25; 87; 93; 21; 98] implement online action detection or action segmentation tasks to predict the category of each time step in a video stream. Similarly, [103; 26; 24; 67] implement action anticipation tasks to predict the action category that occurs after a certain time gap. Additionally, methods such as [39; 52] achieve text-based highlight detection. Segment-level tasks involve predicting segment boundaries in a video sequence based on closed-set categories or open text. Related tasks include moment query [50; 94; 102; 96; 99] and natural language query [100; 65; 92]. The model proposed in this paper is tested on multiple video sequence understanding tasks to verify the language models’ capability to reason about videos from different perspectives.

2 Vision Models

Vision Models, including image and video models, have recently been developed rapidly, mainly focusing on representing short-term vision information. Vision models are divided into convolution, transformer, and hybrid networks. Convolution models learn spatial [32; 28; 90; 56; 95; 84] or space-time [82; 9; 23; 77; 76; 81; 57] visual representations by aggregating neighborhood information using 2D or 3D convolution operators. With the great success of the transformer in the NLP field, the visual transformer has also been continuously developed. The visual transformer models space [18; 55; 86; 74; 5; 20] or space-time [19; 22; 6; 3; 73; 80] through an attention mechanism. Due to the data-hungry problem caused by the lack of inductive bias in the transformer network, a hybrid network [85; 46; 2; 47; 88] combining attention mechanism and convolution operator is proposed to improve performance.

3 Large Language Models

Large Language Models have emerged in recent years in natural language processing. These models usually contain billions to hundreds of billions of parameters and are trained on large text corpora [61; 62; 89; 64; 30; 13; 75]. The core architecture of the model is based on the Transformer while the objective functions range from masked language modeling [17; 53; 35], generative language modeling [61; 62; 7] and permuted language modeling . Among these works, the generative-based language models showed promising results [62; 7] on a wide range of natural language understanding benchmarks. Beginning with the representative work GPT-3 , a series of works [69; 63; 30; 101; 13; 75] scaled up the model and pre-training data and demonstrated strong few-shot and zero-shot performance. Despite the promising results on natural language tasks, the capability of the models are still less explored in multimodal domain. In this paper, we attempt to discover the long-range modeling capacity of LLMs in improving video understanding.

4 Multimodal Models

Multimodal Models aim to learn joint vision and language representation for multimodal downstream tasks. The dominant works are VLP models trained end-to-end on large-scale image/video-text pairs [60; 34; 44; 40; 79; 4; 49]. To relieve the high computation resources, modulated vision-language models adopted frozen unimodal or multimodal pre-trained encoders with learnable modules [43; 42; 1]. These models leveraged strong representation ability of large language models for alignment or generation tasks. BLIP-2 trained a lightweight Transformer to compress the visual tokens and built a bridge between vision output and language input. Flamingo injected visual features into LLM by adding intermediate cross-attention Transformer layers.

Preliminary

The current Language Model can be mainly sorted into encoder-decoder and decoder-only structures. The encoder-decoder uses bidirectional Masked Language Modeling to restore corrupted tokens in a document for textual representation learning, such as BERT and T5 . Alternatively, the decoder-only (GPT family , OPT ) uses unidirectional Language Modeling to directly maximize the likelihood of the sequence under the forward autoregressive factorization. These two training mechanisms grant the language model powerful language sequence modeling and reasoning capabilities. Model parameters and data size of Language models are continuous growth. Table 1 lists the model parameter amount and pre-training token size. These models usually adopt different network structures, training strategies, and corpora. We will explore various LLMs’ performance, advantages, and drawbacks as video sequence reasoners.

2 Tasks

VideoLLM is verified on 8 video understanding tasks across 4 datasets in Table 2. Online Action Detection, Action Segmentation, and Temporal Action Detection focus on detecting and recognizing actions and their temporal boundaries. Online Captioning generates textual descriptions of video content, while Highlight Detection identifies exciting parts and generates summaries. Action Anticipation and Long-term Anticipation predict future actions and content in advance, respectively. Moment Query quickly retrieves specific segments or events in a video. Nature Language Query localize a temporal segment through a textual question.

VideoLLM

VideoLLM is a novel online video reasoning system that aims to apply large-scale pre-trained Large Language Models to video sequence understanding tasks through parameter-efficient transfer learning. It directly borrows the sequence modeling ability of LLM to video sequence reasoning, allowing vision to flow in a natural time sequence in the form of language.

This section will overview the VideoLLM architecture, as shown in Figure 2. Specifically, VideoLLM comprises several components: Modality Encoder, Semantic Translator, decoder-only Reasoner, and simple task heads. In this framework, each short video clip is tokenized using corresponding audio and video encoders and then sequentially processed by the LLM. It is important to note that our unified LLM naturally integrates textual conditions into the framework. Furthermore, our framework allows for the easy integration of various human prompts, commands, human-computer interaction techniques, and parameter-efficient fine-tuning techniques to improve model performance and efficiency.

We adopt a temporal-wise unitization method to process unit-wise visual (or audio and other modality) information for utilizing LLMs to understand video streams comprehensively. We naturally consider integrating natural language modeling with LLMs for unified processing to achieve multimodal understanding.

2 Semantic Translator

The language model is essentially a blind who can receive language input and learn various knowledge, but it has no vision and cannot directly perceive the visual world. Therefore, we need to translate the visual semantics into language representations that the language model can interpret.

3 Decoder-only Reasoner

As detailed in Table 2, our objective is to enable our VideoLLM to accommodate a broad range of video sequence understanding tasks. However, the disparate constraints inherent to these tasks, including their respective inputs and outputs, are a potential obstacle to achieving this goal. To better understand the multifaceted nature of these tasks, we have classified them into four categories, which may exhibit some overlap, as illustrated in Figure 3. This section will discuss efficiently adapting LLMs to address different video understanding tasks.

We employ LLM with a decoder-only structure, denoted as M\mathcal{M}, as the key component of our video sequence reasoner, informed by three critical considerations. First, compelling evidence indicates that decoder-only LLMs are particularly adept at handling causal reasoning tasks for language sequences. Second, the most advanced and high-performing large language models in the current landscape are predominantly decoder-only and are subject to continuous optimization by the research community. Third, a real-world video processor should ideally be designed around a unidirectional visual data flow to maximize performance. This design philosophy aligns seamlessly with the underlying structure of decoder-only language models. Subsequently, we provide a succinct overview of our adaptation method.

Online Reasoning. Online Reasoning primarily focuses on real-time prediction of the category or caption for the most recently attended data unit, which in this paper refers to a new short-term video clip. Given a playing video stream and working memory m={sv−t+1,sv−t+2,...,svi,...,sv0}m=\{s_{v}^{-t+1},s_{v}^{-t+2},...,s_{v}^{i},...,s_{v}^{0}\}, where tt is the number of seen tokens in memory and sv0s_{v}^{0} is the latest translated token. In the training phase, mm will be fed into M\mathcal{M} to construct a causal sequence c={c−t+1,c−t+2,...,ci,...,c0}c=\{c^{-t+1},c^{-t+2},...,c^{i},...,c^{0}\} for parallel training. We use two linear layers to predict the category of each token svis_{v}^{i} and its next token svi+1s_{v}^{i+1}. Thanks to the causal structure of decoder-only LLM, we do not need to calculate the context of the entire sequence when accepting a novel token in the inference phase, compared with a bidirectional encoder. We only make sv0s_{v}^{0} cross-attend to the historical context to calculate new states c0c^{0}. Additionally, we use each cic^{i} as the hidden states for online captioning and input into an extra generative language model Mg\mathcal{M}_{g} (e.g., GPT-2 ) for autoregressive text generation.

Future Prediction. Given a sequence of seen tokens m={sv−t+1,sv−t+2,...,svi,...,sv0}m=\{s_{v}^{-t+1},s_{v}^{-t+2},...,s_{v}^{i},...,s_{v}^{0}\} as the working memory, model need predict the next NfN_{f} tokens or events. In this case, we still utilize the causal structure, supervising each seen token to learn future representations. For predicting different NfN_{f} future states, we use NfN_{f} normalization layers to separate NfN_{f} anticipation presentations a={a1,a2,...,ai,...,aNf}a=\{a^{1},a^{2},...,a^{i},...,a^{N_{f}}\}.

Memory Retrieval. Memory Retrieval often is an offline task to detect event segments in a closed category set or by a text condition. In our online system, however, the task can evaluate the model’s understanding of segment-level transitions and evolutions in the video. Given a sequence of seen tokens m={sv−t+1,sv−t+2,...,svi,...,sv0}m=\{s_{v}^{-t+1},s_{v}^{-t+2},...,s_{v}^{i},...,s_{v}^{0}\} as the working memory, to get the context of the whole video, we use the last token sv0s_{v}^{0} to predict segments in the memory. Another alternative is to concatenate a learnable token svqs_{v}^{q} or at the end of the mm to learn the memory summary. To predict at most NmN_{m} possible segments with category-closed in memory, similar to future prediction, we use NmN_{m} normalization layers to separate NmN_{m} segment-level memory presentations ms={ms1,ms2,...,msi,...,msNm}m_{s}=\{m_{s}^{1},m_{s}^{2},...,m_{s}^{i},...,m_{s}^{N_{m}}\}. Then we adopt two linear layers to predict the category and boundary of each segment. The segments are matched with ground truth through Hungarian matching algorithm for supervision. For memory retrieval based on text condition, we concatenate text presentation yty_{t} or yey_{e} at the end of mm and feed them into M\mathcal{M} together. Hence, M\mathcal{M} can generate the causal sequence conditioned on text for retrieving matched moments.

Dense Prediction. Dense Prediction can be likened to an offline reasoning task where the goal is to predict the category of each token or identify highlight tokens based on textual conditions. In this work, we treat dense prediction as an online task, which serves as a simplified implementation of online action segmentation or highlight detection. Our system uses decoder-only LLM as the default video reasoner and handles online prediction and text conditions like the aforementioned tasks. However, it is worth exploring whether a bidirectional reasoner can provide performance improvements for memory-related tasks. Therefore, we also consider a bidirectional encoder as a potential candidate for our video reasoner, which we evaluate in subsequent experiments.

In summary, our experimental objective is to assess the intrinsic capability of M\mathcal{M} in understanding video sequences. To accomplish this, we propose three fundamental adaptation principles, which have been adhered to by the aforementioned methods. Firstly, we exclusively supervise tasks by relying on the final output of M\mathcal{M}, instead of employing multi-stage supervision as demonstrated in the works of and . Secondly, we refrain from incorporating prior operators, such as convolution layers, into M\mathcal{M}. Lastly, we employ linear layers for each task to transform the hidden states generated by M\mathcal{M} into task results, thereby eschewing the utilization of intricate task-specific heads.

4 Model Training

The training process of VideoLLM involves three fine-tuning methods for training the model.

Basic Tuning. When working with a frozen language model, the optimization of VideoLLM primarily focuses on fine-tuning the semantic translation and output layers. In this scenario, the model’s performance completely relies on the capabilities of the LLM after semantic translation.

Partial Tuning. The partial tuning method involves optimizing specific parts of the LLM in addition to the basic tuning. We adopt three settings for partial tuning: optimizing all bias parameters, optimizing the first block, and optimizing the last block.

PEFT Tuning. The widely popular and effective parameter-efficient fine-tuning (PEFT) techniques in NLP, such as LoRA , Prompt Tuning , and Prefix Tuning , have also been applied to optimize VideoLLM.

Experiments

Dataset and Tasks. In order to thoroughly assess the capabilities of LLMs in video understanding, we performed experiments on four datasets, covering a total of eight tasks. The details of these tasks and datasets are presented in Table 2. The tasks were categorized into four types, as illustrated in Figure 3: Online Reasoning, Future Prediction, Memory Retrieval, and Dense Prediction. This diverse set of tasks allows for comprehensive evaluations from various perspectives, including data accessibility (causal or non-causal), perceptual objectives (memory or anticipation), and prediction granularity (segment-level or frame-level), modalities (vision-only or vision-language).

Evaluation and Metrics. Our model evaluation is conducted in accordance with previous studies [25; 103; 21; 29; 27; 39; 10]. Specifically, we measure the accuracy of online action detection and action anticipation tasks using class-mean recall@5(%) following the established standard protocol . To assess the performance of our model in the action segmentation task, we report the framewise accuracy (Acc), segmental edit distance (ED), and the segmental F1 score at overlapping thresholds of 25% denoted as F1@25. For the Long-term anticipation task, we submit our results to the EvalAI platform to evaluate the test set. Consistent with the approach employed in , we evaluate the mean Average Precision (mAP) under multiple temporal Intersection over Union (tIoU) thresholds, specifically {0.1;0.2;0.3;0.4;0.5}\{0.1;0.2;0.3;0.4;0.5\}, for the Moment Query task. In addition, we report the recall@k, where k = 1, and the IoU=m metric, where m = {0.3,0.5}\{0.3,0.5\}, for the Nature Language Query task.

Implementation Details. To ensure fairness and facilitate meaningful comparisons within the research community, we employ various visual encoders [9; 60; 23; 91; 73; 104; 88] that have been pretrained on different datasets [16; 36; 60; 27; 15] to extract visual features. This approach helps establish alignment with existing community settings and ensures equitable evaluations. Note that, the same modality encoder could share semantic translator. In this work, using different encoders and semantic translators for aligning community settings is a special case. In particular, we adopt the fundamental settings proposed in [93; 103] for the Online Action Detection and Action Anticipation tasks. We leverage the settings introduced in for the Action Segmentation task. The Online Captioning task follows the settings outlined in . Similarly, we adhere to the settings specified in for the Long-term Anticipation, Moment Query and Nature Language Query task. The Highlight Detection task builds upon the settings presented in .

2 Main Results and Analysis

Which language model performs better? Figure 4 presents the comparison results between three base-level LMs, GPT-2 , T5 Decoder , and OPT . The results are obtained through the basic tuning method. We select representative metrics for each task for intuitive comparison. From the results, we can see that different language models have different performances on different video sequence understanding tasks. Both GPT-2 and OPT are better than T5 decoder in future prediction tasks (see AA and LTA in the figure). On the contrary, OPT is significantly better than GPT-2 and T5 decoder in OAD task. For Moment Retrieval tasks, we find that GPT-2 can still gain dominance (see MQ and NLQ in the figure). It is worth noting that T5 Decoder has a great advantage over GPT-2 and OPT in dense prediction tasks (see AS and HD in the figure). For online captioning, GPT-2 attains the best performance, compared with T5 Decoder and OPT. We suppose that using GPT-2 as video sequence reasoner M\mathcal{M} better aligns the text generator Mg\mathcal{M}_{g} (also GPT-2) we used from . In general, the structure and training strategy of the language model will result in different processing capabilities for video sequences and exhibit different adept abilities. In fact, when we calculated their average scores based on the results, we found that GPT-2 and T5 decoder were basically on par, and OPT was slightly worse than GPT-2 and T5 decoder.

Which Tuning method performs better? To evaluate the influence of various tuning methods on performance, we opt OAD as the experimental object. It is a causal dense prediction task, providing a more realistic representation of performance alterations. Table 3 presents the Action Top-5 Recall achieved through the utilization of various tuning methods, along with the corresponding increase in trainable parameters compared to the basic tuning approach. We employ rr as a uniform representation of the hyperparameter for the three PEFT tuning methods, and carry out experiments using r=1/2/4/8r=1/2/4/8. As depicted in the table, employing LoRA with different rr results in a decline in performance. Conversely, the other tuning methods exhibit performance improvements of at least 0.2 points in the Action Top-5 Recall metric. Although fine-tuning the first or last block can yield performance gains, it also entails a significantly larger number of trainable parameters compared to the other methods. Remarkably, when employing prefix tuning with r=4r=4, the model achieves the best outcome, attaining an Action Top-5 Recall of 21.4, surpassing the basic tuning method by 1.3 points.

Comparison to the state-of-the-art methods. Table 4 presents the evaluation results for seven video sequence understanding tasks. It is important to note that the OC task is not included in this analysis due to the lack of comparable sequence-level methods. To thoroughly assess the effectiveness of VideoLLM, we conduct a comparative analysis with other cutting-edge methods that are specifically tailored to individual tasks. The reported results for VideoLLM represent the most favorable performance achieved from numerous combinations. To evaluate the OAD task, we reproduce the existing state-of-the-art methods [93; 103] and adopt the same evaluation metrics as the AA task. Notably, we ensure a fair comparison by excluding the data augmentation techniques employed by Testra . Our model demonstrates higher or comparable performance in both the OAD and AA tasks. Particularly, our approach achieves a higher Unseen Action Top-5 Recall, highlighting the ability of utilizing LLMs to ensure and potentially enhance generalization in unseen scenarios. For the AS task, our model outperforms the state-of-the-art method MS-TCN in terms of F1@25, edit distance, and accuracy. It is worth emphasizing that our adaptation principle solely relies on the sequence modeling capability of the LMs themself, without introducing any local prior operator or multi-stage refinement. This observation emphasizes that a language sequence-trained model can serve as a robust initialization for video sequence modeling. We also apply our adaptation principles to MS-TCN and ASFormer , with the corresponding results presented in the table. In the table, SS-TCN† refers to the deep network with a single-stage supervision mentioned in the MS-TCN paper. These results demonstrate a significant inferiority to our single-stage adaptation. Furthermore, we compare VideoLLM against state-of-the-art or baseline methods on multiple sub-tasks, namely LTA, MQ, and NLQ, of Ego4D . The evaluation conducted on the LTAv2 test set, using the EvalAI platform, shows that our model outperforms the official baseline methods. Moreover, under the constraints of the adaptation principle, our model exhibits a slight performance superiority over VSGN , which employs an anchor-based prior setting for the MQ task. In the realm of visual-language tasks, our models exhibit substantial superiority over existing state-of-the-art methods [10; 39]. This finding underscores the impressive performance of language models once the vision-to-language semantic translation is accomplished. Furthermore, in addition to the performance comparisons, we also compare the trainable parameters with these methods. The table reveals that our method necessitates approximately 2M to 15M learnable parameters across multiple tasks, with most of these parameters primarily utilized in semantic translator and task head. This substantiates the parameter efficiency of our proposed framework. In summary, these results convincingly demonstrate the adaptability of our proposed framework across a diverse range of video sequence understanding tasks, each with its own unique settings.

Scale of LLM. We also assess the scalability of utilizing LLMs as video sequence reasoners for our approach, through experimental evaluations conducted on the OAD task. Figure 5 displays the Action Top-5 Recall achieved by employing LLMs with varying scales of total parameters. In these experiments, we scale up three decoder-only LLMs, namely GPT-2, T5 Decoder, and OPT, and solely fine-tune two projectors using the basic tuning method. This ensures a comprehensive evaluation of the intrinsic capabilities possessed by these LLMs. As depicted in the figure, when utilizing language models with parameter sizes less than 2B, compelling evidence suggests that larger models yield more substantial improvements in video sequence reasoning. Among the three models, it is worth noting that OPT-1.3B yields the most favorable results, achieving a remarkable 23.4 Action Top-5 Recall. Furthermore, when considering the overall performance improvement trend observed during the scaling-up process, it becomes evident that OPT outperforms T5 Decoder, which, in turn, surpasses GPT-2. However, for larger LLMs, their performance begins to decline. One plausible explanation for this phenomenon is that the dimension-expansion projector causes the model to overfit, as the dimension of the extracted feature sequence is typically less than 2048. In conclusion, these experiments effectively demonstrate the scalability of our method to LLMs, highlighting their potential for adapting video sequence reasoning tasks.

Advanced LLM. We further scale up OPT and T5 decoder to 6.7B and utilize the latest 7B LLaMA model. The performance of T5 and OPT, as depicted in Table 5, continues to align with the declining trend observed in Figure 5. Notably, the performance of LLaMA closely approximates that of OPT.

Encoder vs. Decoder. We conducted experiments to compare the performance of bidirectional and unidirectional sequence reasoners on three tasks: AS, HD , and NLQ. For the bidirectional sequence reasoner, we employed the T5 encoder, while the unidirectional sequence reasoner utilized the T5 decoder. A comprehensive comparison of all task metrics is presented in Table 6.

As evident from the table, the bidirectional reasoner consistently outperformed the unidirectional reasoner in most cases. This discrepancy is particularly prominent in AS tasks, where the bidirectional reasoner exhibited a significantly higher level of performance compared to its unidirectional counterpart. This may be attributed to the importance of bidirectional attention in confirming temporal correlations and pre-post-action relationships within a complete event during action segmentation. In the case of visual-language tasks, HD and NLQ, the bidirectional reasoner also showcased a slight advantage over the unidirectional reasoner. However, it is worth noting that the Rank1@0.3 obtained by the OPT on the NLQ task, as depicted in Figure 4, is comparable to that achieved by the T5 Encoder (7.3 vs 7.4). This suggests that the decoder-only unidirectional reasoner holds the potential to achieve performance on par with the bidirectional reasoner.

Conclusion and Future Work

In this paper, we propose a novel video understanding framework called VideoLLM, which transfers the sequence causal reasoning abilities of large language models (LLMs) from natural language processing to video understanding. The VideoLLM framework comprises a well-designed Modality Encoder and a Semantic Translator, which convert inputs from different modalities into a unified token sequence. This sequence is then fed into a decoder-only reasoner realized by the large-scale language pretrained and parameter-frozen LLM, which possesses the ability to decode and output meaningful high-level semantics. With the help of simple task heads, the output of the LLM corresponds to various specific video understanding tasks. Extensive experiments were conducted on eight tasks from four different datasets using multiple LLMs and fine-tuning methods to evaluate the effectiveness of VideoLLM. The experimental results demonstrate that LLMs’ comprehension and reasoning abilities can be effectively applied to video understanding tasks. In our future work, we will further explore the potential of LLM. Building upon time series reasoning, we aim to incorporate serialized information about the appearance of video frames, enabling LLM to achieve a more comprehensive video understanding across the entire spatiotemporal dimension.

References