PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Lin Xu, Yilin Zhao, Daquan Zhou, Zhijie Lin, See Kiong Ng, Jiashi Feng
Introduction
Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in image comprehension when trained on large-scale image-text pairs . Analogous to the image domain, the recent video understanding models also explore a similar pipeline to fine-tune LLMs on large-scale video-text data . However, this method suffers a high cost of computing resources and video data annotations. A more pragmatic approach is to adapt the pre-trained image-domain MLLMs to video data .
An intuitive method for image MLLM adaption is to encode multiple video frames into a sequence of features and directly feed them into MLLMs, as the Large language Models(LLMs) are native for processing sequential features and shown to be capable of understanding temporal information . However, we empirically found two technical challenges when extending image MLLMs to video data in this way. First, compared to zero-shot applications, training the image MLLM on video data does not always increase the performance but introduces performance vulnerability to the change of inquiry prompts. Secondly, increasing the size of the language model component does not improve the video understanding performance. Those two observations are counter-intuitive since scaling up model sizes and exposing models to more downstream data are typically considered beneficial for model performance.
We then conducted a series of studies to investigate the root cause of these two observations. For the first one, we found it is mainly due to the limited information encoded by the image encoder. When experimenting on LLaVA with 4-frame inputs, we experimentally found that, as shown in Figure 3, some visual feature tokens have dominantly larger norms over the others during the fine-tuning process. These tokens lead to shorter text descriptions with lower quality. As demonstrated in Figure 2, the 4-frame models tend to generate shorter texts with training on more samples. We conjecture that the large-norm features have obtained global video information and thus suppress the norms of other tokens, due to the softmax calculation during the self-attention. This leads the generated description to be short. Even worse, if the prompt template changes, the learned MLLMs would completely collapse, leading to rather short descriptions or even no response. We observe that adding more video frames could mitigate the suppression of the majority of the tokens. However, this would lead to significantly larger memory consumption.
Thus, there is a trade-off between the number of frames and the computation cost. The intuitive way is to downsample the video frames. However, directly averaging the spatial and temporal dimensions as has been done in VideoChatGPT loses too much spatial information and also does not achieve optimal performance during the scaling of the training dataset. Thus, the target is to find the minimum spatial resolution of each frame that does not degrade the scaling curve. To achieve this, we adopt a pooling operation to explore the optimal settings such that it does not degrade the benefits of increasing the temporal receptive field. The impact of the pooling operation is shown in Figure 7.
For the second observed phenomenon, we believe one main reason is the poor quality of the video dataset, compared to the image dataset. Specifically, many of the video datasets are in question-answer formats and the descriptions of the videos might be short. Thus, as the model learns the temporal description from the video dataset, the description of other metrics such as the objects and the spatial relations degrades. the stronger the LLM is, the faster the output degrades. Instead of building high-quality video datasets, we choose to explore architectural and optimization algorithms to better preserve the learned information in image datasets during the learning of the temporal information on video datasets. To achieve this, we utilize the tricks of weight fusion. We set two groups of weights: one from the image pre-raining and one with video dataset fine-tuning. After training, we searched to find the optimal combination of the image-based model weights and the video-based model weights in the hope that the combined model could gather benefits from both datasets. The process is termed post-training optimization in this paper and its impacts are shown in Figure 5.
We performed a thorough initial investigation for directly applying image large multi-modality models to video tasks and found several failure modes. We then introduce an elegantly simple yet highly potent pooling strategy that systematically achieves the optimal balance between training efficiency and caption accuracy.
We introduce a post-training model merging method that could effectively reduce the forgetting phenomenon of the large language models during multi-modality fine-tuning. With this, we are able to get a large video multi-modality model with 34B LLMs without the extra creation of high-quality datasets.
We conduct extensive experiments to verify the superiority of the proposed model and achieve a new state-of-the-art across various video understanding benchmarks, especially for video captioning tasks with dense captions. With Pool-LLaVA, we do the re-captioning of the top 1M video data from Panda-70M with highly dense and accurate bilingual captions.
Related Works
Video Multi-modality Models process video input and generate responses according to user commands. Commonly, they incorporate a projection network , inter-modality attention or a modality perceiver as learnable interfaces. These interfaces are instrumental in melding the spatial-temporal dynamics of videos with large language models’ (LLMs) processing capabilities , by transforming video content into a sequence of tokens that LLMs can adeptly analyze. Parameter efficient learning schemes are adapted to reduce the computational cost. Among them, BLIP marked a significant milestone by integrating a frozen vision encoder with BLIP to enhance video processing efficiency, with only the newly added Q-Former learnable. Demonstrating remarkable zero-shot capabilities in Video Question Answering (VQA), it outperformed existing techniques of its time. Extending the innovations of its predecessors, Video-ChatGPT introduced the trailblazing approach of video instruction tuning, along with creating a dataset of high-quality instructional data. This initiative set a new standard for assessing models through video-based text generation benchmarks. VideoChat employed cross-attention mechanisms to skillfully condense video tokens, aligning user queries with the dialogue context to enhance the model’s interpretative capabilities. Building upon these advances, VideoChat2 refined the approach with a multi-stage bootstrapping technique that honed in on modality alignment and instruction tuning, amassing a robust collection of high-quality video data for fine-tuning instruction-driven tasks. VILA proposes more advanced training recipes. Further integrating modalities, Video-LLaVA leveraged a pre-aligned encoder adaptable to both images and videos, facilitating shared projections and enabling synergistic training across image and video-related tasks. CAT introduces both video and audio to futher enhance understanding.
Long videos present significant challenges due to their intrinsic high computational complexity and extensive memory requirements. Handling the entire span of a long video with video tokens poses difficulties in jointly capturing spatial details and temporal dynamics effectively. In response, Video Language Models (Video MLLMs) have adopted sophisticated temporal modeling techniques to address these challenges with greater efficiency. MovieChat implemented a novel memory-based mechanism within transformers, strategically combining similar frames to reduce both computational load and memory footprint. Chat-UniVi debuted a harmonized approach for processing images and videos, innovatively condensing spatial and temporal tokens through dynamic token merging, utilizing k-NN algorithms for improved efficiency. LLaMA-VID innovated with a dual-token approach that effectively condensed video representations by segregating context and content tokens, allowing for more efficient compression. VTimeLLM emphasize the boundaries of videos by introducing a new question answering dataset. Advancing this innovation, Vista-LLaMA introduced EDVT-Attention along with a sequential vision projector that meticulously curates visual tokens and condenses temporal tokens, progressively amalgamating them with a Q-former mechanism. To further optimize the handling of extended videos, certain models emphasized the selective processing of keyframes, thus diminishing the volume of video frames required and streamlining the overall computational demands.
Pipelined Video Understanding
Capitalizing on the Video MLLM framework, a novel approach emerged involving the use of pre-existing Video Models coupled with LLMs through a multi-stage process of video modality conversion. This method entails translating video content into textual narratives, typically through the employment of pretrained VideoLMs, before integrating with an LLM in the final phase. By encapsulating videos as text tokens, it leverages the LLMs’ adeptness at navigating textual data, thereby permitting the interpretation of temporal sequences via these crafted descriptions. VideoChat-Text adeptly converts video streams into comprehensive text descriptions, encapsulating a range of video elements. Meanwhile, LLoVi unveiled an efficient, LLM-centric framework tailored for addressing queries that span long video durations. Here, video captioning agents transcribe videos into detailed textual descriptions which the LLMs then distill to enhance long-duration video comprehension. While the aforementioned methodologies primarily translate video into text for LLM processing, LLMs are concurrently being explored for their capacity to facilitate video analysis through program generation. ViperGPT is a pioneering example, harnessing code-producing LLMs, including the likes of GPT-3 Codex . It effectively utilizes a visual module API catering to text-based queries and crafts programs that scrutinize image or video content, furnishing informed responses to those queries. Similarly, ProViQ engages an LLM to craft Python scripts that enact multi-stage procedural reasoning in the context of zero-shot video queries, processing these scripts to ascertain solutions to posed questions.
Method & Analysis
Adapting image MLLMs into the video domain can be tricky and vulnerable to the designs of model structures. In this section, we first present some challenges encountered when extending image MLLMs to video, drawing insights from our comprehensive experiments and analyses. Corresponding solutions to these challenges will be presented, forming the integral framework of PLLaVA .
where is the text inputs and r is the output texts. Nonetheless, during our efforts to train the MLLM in this scenario, we encountered two issues that hindered us from achieving optimally performance models.
The first observation is that the models trained with n-frame could be highly sensitive to prompt patterns when dealing with generation tasks. Figure 3 illustrates such a phenomenon. We divide the prompts into two categories: in-distribution (IND) and Out-of-Distribution (OOD). In the left part of the figure, when generating under the prompt pattern used in training (IND), the model can generate decent descriptions about the video despite its tendency of shorter generation length with more data samples trained. However, if we prompted the model with OOD prompts, in which we just changed the tags for the two roles in a conversation, the quality of the generated response then drastically declined. The generation has content in normal length under the model trained for 3750 steps. However, for the longer trained models, the generations are shorter for 7500 steps, and even no response for 11250 steps. This example demonstrate the vulnerability of the n-frame method.
Dominant tokens.
In view of the vulnerability of n-frame models stated above, we proceeded to analyze the variance between models at their initial stages of training and when fully trained. By visualizing the norm of vision tokens across models trained at different stages, we observed a trend towards the emergence of dominant tokens( with high norms) as training samples increased, as shown by the histograms in Figure 3. Furthermore, the twin-tower distribution is much wider when trained with more data. Therefore, we speculate there exists a plausible correlation between these dominant tokens and the degradation of generation under OOD prompt. The distribution comparisons between n-frame and the proposed PLLaVA can further validate the conjecture, which is explained in Sec. 4.4.
Data scaling failures.
2 Model Scaling Degradation
Our investigation on current video models reveals that increasing the model size does not typically result in significant improvements in performance for most models. We draw the performance of a recent work IG-VLM and our attempts in Figure 5. IG-VLM achieves almost no difference when applying 7B, 13B, and 34B models of LLaVA-Next . In our attempts of with pooling features (the first column of Figure 5), the performance of LLaVA-Next 34B is even worse than its 13B LLaVA-Next model. For IG-VLM, the input video frames are combined to a grid view image, confined by the resolution, leading to the unsatisfactory scaling ability. As for our attempts, we found a tendency of shorter generations with larger MLLMs, thus we owe the degradation to the quality of video-text data pairs, which undermines the generation ability of LLMs in MLLM models.
3 PLLaVA
Our initial attempts on n-frame and VideoChatGPT reveal the intricacies of adapting image-focused MLLMs to the video domain, encountering the data scaling problem. The former introduces a small amount of frames due to the limit of memory, whereas the latter compresses over 100 frames of information with pooling strategy. However, similar outcomes occur to both situations.
In view of the necessity of temporal information and the prohibited costs of dealing with very long video input to MLLMs, pooling is an intuitive and simple way to fulfill both of the requirements. The above two problems may stem from inadequacy of frame information and mishandling on the frame features. Therefore, in this paper, we deeply look into the pooling strategies for video features used in MLLMs.
Definition
These features are then concatenated with text input embeddings and fed into the LLM to generate responses. We also include a LoRA module to adapt the LLM to video-related generation tasks. In conclusion, the trainable weights include Multimodal Projector and LLM LoRA.
Within this framework, we investigated the impact of pooling through grid search analysis. Our findings suggest that pooling on the spatial dimension yields favorable outcomes, whereas temporal dimension pooling is associated with decreased performance. For a thorough exploration of our search process and the rationale behind this conclusion, please refer to Sec. 4.2.
4 Post Optimization
Regarding the problem of performance decline associated with scaled model size, such degradation may stem from diminished language proficiency resulting from training on low-quality video-text data samples. To mitigate this, we propose a post-training optimization approach for the parameters of the video MLLM. It involves blending the trained Language Model (LLM) on video data with the original LLM of the base image MLLM. For a pretrained MLLM with LLM parameters and given input , the output hidden states from the LoRA fine-tuned LLM can be acquired as follows:
where are a low-rank learnable parameters for adapting , and is used to scale the learned low-rank weight.
As part of our post-training optimization process, we tune the mix ratio between the original LLMs and the trained LLMs (incorporating LoRA weights) by varying the value of during inference. Our experiments indicate that lower yields significantly better generative performance.
Experiments
We leverage instructional video-to-text datasets to extend the capabilities of image MLLMs to handle video inputs. The training data are sourced from the dataset used in VideoChat2 , which embraces data for various video understanding tasks, including 27k conversation videos from VideoChat and Video-ChatGPT , 80k data of classification tasks from Kinetics and SthSthV2 , 450k captioned data from Webvid , YouCook2 , TextVR and VideoChat, 117 reasoning data from NextQA and CLEVRER and 109K annotated questioning answering data samples from Webvid, TGIF and Ego4D . In total, we use 783k instructional tuning data.
We evaluate our trained models with the following video-to-text benchmarks. First, the open-ended Video Question Answer (VideoQA) includes MSVD-QA , MSRVTT-QA , ActivityQA , and TGIF QA . Responses in these question-answering benchmarks typically consist of single-word answers. GPT-3.5 is used to evaluate the accuracy (Accuracy, with answers true/false) and quality (Score, ranging from 0 to 5) of the models’ responses. Additionally, we adopt the Video-based Generative Performance benchmark (referred to as VCG Score), introduced by VideoChatGPT . These benchmarks often involve longer answers, encompassing five aspects of video understanding: CI (Correctness of Information), DO (Detail Orientation), CU (Context Understanding), TU (Temporal Understanding), and CO (Consistency). The generation is also assessed using the GPT-3.5 model. Furthermore, we also use the multi-choice Question Answering benchmark, MVBench , comprising 20 tasks that demand nuanced temporal comprehension of videos. This benchmark does not necessitate evaluation from the GPT-3.5 model.
Models and Implementation Details
PLLaVA is constructed upon the image MLLMs, LLaVA Next models 7B, 13B, and 34B. We utilize their pre-trained weights available in the Hugging Face libraryhttps://huggingface.co/docs/transformers/en/model_doc/llava_next and integrate an average pooling to reduce feature dimensions before passing the input visual features to the LLM generation component. For the pooling layer, we uniformly sample 16 frames as input and set the target pooling shape to be , where corresponds to the input dimension of the LLMs. During training, we employ a batch size of 128 and a learning rate of 2e-5, with a cosine scheduler and a warmup ratio of 0.03. All the reported results are evaluated on models trained for 6250 steps. For evaluation, we adopt the GPT-3.5-turbo-0125 model across all the benchmarks.
2 Impact of Pooling Operation Design
Considering the unsatisfying performance of the complete pooling on temporal and spatial dimensions adopted in Video-ChatGPT and the limitation information in the straightforward n-frame method, we further explore the influence of poling strategies here.
Pooling can be done both temporally and spatially. In this part, we aim to figure out the answer to two questions: 1) which dimension is more suitable to be pooled to save the computational cost and 2) what is the largest compression ratio along that dimension. To achieve this, we plot a model curve based on the LLaVA-1.5 7B model with different temporal and spatial dimensions controlled via pooling operation. Specifically, for the spatial dimension, we picked an input video feature with shape (4,24,24,), where 4 is the frame numbers (temporal dimension), 2424 is the original spatial dimension of frame features, and is the embedding dimension of each visual token. The target spatial shapes are chosen at evenly spaced intervals between 1 and 24, resulting in a set of spatial shapes { | ]}. The MVBench and VCG Score performance of these spatial pooling shapes are shown in Figure 7(a) and 7(b). It is observed that downsampling the spatial dimension by 50% does not degrade the model performance. Further reducing the spatial dimension would lead to a significant performance drop. Considering the tradeoff between computational overhead and performance, 1212 can be a target spatial dimension.
We further experimented on the temporal dimension. Several target pooling shapes were chosen with spatial dimensions fixed as 12, including (4,12,12), (8,12,12), and (16,12,12). We study the pooling performance tendency when altering the number of input video frames, indicating the downsampling rate of poolings. For example, pooling from (64,24,24) to (4,12,12) indicates every 16 frames are fused, then the downsampling rate should be 6.25%. All of the resulting model curves are shown in Figure 7(c) and 7(d). Different from spatial pooling, the model performance is sensitive to the temporal pooling. As illustrated in these two figures, all lines achieve better performance with lower downsmapling rates. In other words, pooling along temporal dimension always downgrades the model performance.
Pooling Impact
We found that pooling over more video frames not only improves the model efficiency but also makes the model more robust to user enquires. During our experiments, we evaluated models under different training iterations with two sets of prompts. For example, we vary the role tag from ‘USER’ to ‘Human’ during evaluation and the results are as shown in Figure 3. The figure shows that the visual feature norms learned with the pooling operation show consistent distributions under different training iterations compared to the 4-frame method that shows dominant tokens. This is also reflected in the model responses where the pooling method gives consistent good text responses while the 4-frames method gives shorter and shorter text responses as the training goes longer, or even no response when out-of-distribution prompts are used. This conclusion can be further validated by Figure 2. With pooling introduced, no matter what prompt is used or how much training sampled is learned, the text generation lengths with the pooling method are consistent. We owe the stability in generation to the smoothing ability of pooling, which eliminates the influence of dominant high norm tokens. For more rigorous analysis from the perspective of mathematical proofs, we leave it for future work.
3 Quatitative Results
Table 2 demonstrates the results on VideoQA. PLLaVA 34B significantly outperforms all the existing methods on the Accuracy and Score metrics of MSVD, MSRVTT, ActivityNet, and TGIF. Compared to GPT-4V, PLLaVA 34B achieves improvement margins of 3.6, 4.9, 3.9, and 15.3 on these four benchmarks. The performance of PLLaVA with 7B and 13B model sizes also exceeds all the baselines on the Score metric. These results not only prove the capability of our model in conducting video question answering but also highlight the superiority of our pooling strategy in scaling model size.
PLLaVA also achieved a new state-of-the-art in the average VCG score. The 7B, 13B, and 34B versions have all outperformed their best counterparts of the same LLM size, with margins of 2.9%, 7.1%, and 12.6%, respectively. Notably, PLLaVA achieves superior performance on CI(correctness of information), DO(Detail Orientation), and CU(Context Understanding) compared to the previous SOTA, with 34B exceeding them by 5.8%, 6.7%, 9.2%. These results indicate that PLLaVA will be of great potential to do detailed video captioning. As for TU(temporal understanding), PLLaVA 34B exceeds its fair opponent IG-VLM LLaVA 34B by 6%. Compared with models that utilize the specialized video encoder, VideoChat2, or a more complicated frame combination method, Chat-Univ, PLLaVA still has some room for improvement by fingering the pooling strategy or incorporating a better vision encoder. CO(Consistency) measures generation consistency when the model encounters different questions that lead to similar answers. Compared to baselines except for IG-VLM, our model achieves much better consistency.
MVBench is a comprehensive video understanding benchmark, focusing on questions that require overall comprehension of multiple frames. As shown in Table 3, PLLaVA surpasses the previous SOTA VideoChat2 with a margin of 13.7% on average across 20 tasks. If we look into each aspect of MVBench, our method performs very well, concerning 17 out of 20 tasks of MVBench, which shows that our model has the superiority to understand many fine-grained details about videos accurately. However, we also noticed some aspects of our model still need to improve, such as CI(CounterFactual Inference) and OS(object shuffle). CI is used to predict what might happen if an event occurs, and OS is used to locate the final position of an object in an occlusion game. These two require strong reasoning ability and imagination to answer. VideoChat2 is pretrained with a large amount of video data with a specialized video encoder and fine-tuned with both video and image reasoning data, thus presenting better performance in these aspects.
4 Analysis
Our PLLaVA is a simple and parameter-efficient method to adapt image MLLMs into the video domain. We also provide a feasible way to scale the models to larger sizes, which we found is hard to achieve in other methods such as ChatUniv and IG-VLM . In the following, we further provide some analysis related to the explanations on pooling shapes and the influence of LoRA weight on different tasks.
In Sec.4.2, we have illustrated the impact of temporal and spatial poolings, concluding that pooling along the temporal dimension consistently results in decreased performance compared to retaining the original frame numbers. We attribute this phenomenon to the interference with token features. In image MLLMs, features are derived from images/video frames using CLiP-ViT models, which produce embedded patches for each image/video frame, resulting in a video feature with shape . Pooling changes the dimensions of (time), (height), and (weight). In contrast to pooling along the spatial dimension (local pooling on single images/frames, changing and ), pooling along the temporal dimension (changing ) risks altering the original frame features. To validate the guess, we visualize token similarities among spatial and temporal token neighbors for a video feature in Figure 8. The two subfigures reveal significantly higher similarities within spatial neighbors than temporal neighbors. This observation supports the potential distortion of original token features caused by temporal pooling. LLMs are designed for sequence understanding. Even without preprocessing on temporal information aggregation, they can model temporal relations.
Image? Video? or Both?
Post-training optimization is defined as the combination of the LLMs’ parameters of image MLLMs and learned LLMs’ LoRA weights from video samples. A suitable fusion ratio could be highly efficient in boosting model performance trained under low-quality video-text samples. Here, we discuss the influence of different choices of fusion ratio on the understanding performance. As shown in Figure 9, the x-axis represents the alpha value of LoRA. 0 indicates no LoRA weights added, and 32 means the LoRA weights are fully applied to LLMs. We observed distinct trends between MVBench and VCG Score. The former exhibits a peak around alpha 20, while the latter performs best near alpha 4. This variance can be attributed to the nature of these two benchmarks: VCG typically involves longer length generations, whereas MVBench focuses on multiple-choice question answering, placing less emphasis on language generation ability. Consequently, weights learned from video-text data samples are more tailored for MVBench tasks. In this way, a larger portion of video weights are beneficial for MVBench. Moreover, from these two figures, it’s evident that combining video and image weights leads to better performance than at the extremes of 0 and 32.
5 Case Studies
Apart from these quantitative results, we also qualitatively investigate the video understanding abilities of PLLaVA models. We have shown several caption examples in Figure 10. According to the video clips, compared to IG-VLM, PLLaVA 34B recognizes more details about videos, including the clothes worn by the main characters, the environment, and even some of the words in this video. Besides, as shown in Figure 10(b), PLLaVA can understand the video content more correctly, in which people are playing badminton rather than volleyball. These mistakes made by IG-VLM could be caused by the lowered resolution when concatenating frames into the grid view in the method design. Pooling reduces dimension after frames are encoded, thus leading to less information loss.
6 Dense Recaption
In view of the caption ability of PLLaVA , we further tested its recaption task and contributed 1K video Inter4K caption dataset. An example is shown in Figure 11. Compared to Open-Sora GPT-4 pipeline, our model captures better caption details and also highlights motion information in the video, demonstrate PLLaVA ’s potential to contribute to the video generation community.
Conclusion
In this paper, we conduct an initial investigation for extending image-language models to videos with a simple yet extremely effective method, termed PLLaVA . With the new model, it is easier to scale the training with more data and larger large language models with a more controllable strategy for over-training and performance saturation. PLLaVA ’s ability of giving detailed captions also contributes to the community development of multimodal understanding and generation.