Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Bin Lin, Yang Ye, Bin Zhu, Jiaxi Cui, Munan Ning, Peng Jin, Li Yuan

Introduction

Recently, LLMs have gained rapid popularity in the AI community, such as GPT-3.5, GPT-4 , PaLM , and BLOOM . They rely on their powerful language comprehension abilities to follow human-provided instructions and provide corresponding responses. Typically, LLMs can only respond within the text input provided by the user, which is insufficient because human interaction with the world involves multiple channels, such as visual and textual. To this end, recent works have mapped images into text-like tokens, enabling LLMs to emerge with the ability to comprehend images. Despite their effectiveness, empowering LLMs to understand videos is more challenging than image-only comprehension tasks. Nevertheless, recent work has made initial strides in enabling interactions between video and language.

However, most current LVLMs can primarily handle a single visual modality, either image-language or video-language. We compare different LVLM paradigms as shown in Fig. 1, where VideoChat and Video-LLaMA utilize a share visual encoder to handle both images and videos. However, due to the inherent differences in the media types of images and videos, it is challenging to learn a unified representation, and the performance falls significantly behind that of the specialized video expert model, Video-ChatGPT. Therefore, X-LLM and Macaw-LLM allocate a modality-specific encoder for each modality, attempting to enable a LLM to comprehend images or videos through several projection layers. But their performances are inferior to dedicated video expert models such as Video-ChatGPT . We attribute this phenomenon to the lack of alignment before projection. Because image features and video features reside in their own spaces, this poses a challenge for a LLM to learn their interactions from several poor projection layers. Some similar phenomenon such as alignment before fusion has been discussed by ALBEF and ViLT in multi-model models. More recently, ImageBind-LLM focuses on enabling the LLM to simultaneously process multiple modal inputs by pre-aligning each modality to a common feature space . Based on a large image-language model, ImageBind-LLM converts other modalities into the most similar image features by retrieving from a training-free image cached database. However, the indirect alignment approach of ImageBind-LLM may lead to performance degradation, and the LLM has no knowledge of actual video data.

In this work, we introduce Video-LLaVA, a simple but powerful baseline for the LVLM simultaneously handling both images and videos. Specifically, As shown in Fig. 1, Video-LLaVA initially aligns the representations of images and videos to a unified visual feature space. Since the visual representations are already aligned prior to projection, we employ a shared projection layer to map the unified visual representation for the LLM. To enhance computational efficiency, Video-LLaVA undergoes joint training of images and videos, achieving remarkable results with 1 training epoch.

As a result, The proposed Video-LLaVA greatly enhances the ability of the LLM to simultaneously understand both images and videos. For image understanding, Video-LLaVA surpasses advanced LVLMs such as mPLUG-owl-7B and InstructBLIP-7B in 5 image benchmarks. Additionally, utilizing 4 benchmark toolkits for a more comprehensive evaluation, Video-LLaVA-7B even outperforms IDEFICS-80B by 6.4% in MMBench. Moreover, similar trends can be observed in video understanding, where Video-LLaVA surpasses Video-ChatGPT by 5.8%, 9.9%, 18.6%, and 10.1% respectively on the MSVD, MSRVTT, TGIF, and ActivityNet video question-answering datasets. Extensive ablation experiments demonstrate that alignment before projection yields greater benefits. Additionally, joint training of images and videos can facilitate a unified visual representation in LLM comprehension.

We summarize our primary contributions as follows:

We introduce Video-LLaVA, a powerful LVLM baseline. During the training process, Video-LLaVA binds visual signals to the language feature space, unifying visual representations, and proposes a solution to align before projection. We enable an LLM to perform visual reasoning capabilities on both images and videos simultaneously.

Extensive experiments demonstrate that a unified visual representation benefits LLMs in learning to simultaneously handle both images and videos, validating the complementarity of modalities, showcasing significant superiority when compared to models specifically designed for either images or videos.

Related Work

When the well-known commercial model ChatGPT was introduced, the The AI community released open-source Large Language Models (LLMs) by instruction tuning and increasing model sizes. These include LLaMA , Vicuna , Alpaca , and more recently, LLaMA 2 . These models are tuned with instruction sets to emulate conversations between humans and AI assistants. Furthermore, InstructGPT is trained based on GPT-3 with 175 billion parameters through aligning with human preferences. However, LLMs can only interact within text. In this work, we introduce Video-LLaVA, which builds upon the powerful reasoning capabilities of LLM to extend modality interactions to images and videos.

2 Large Vision-Language Models

When extending LLMs to multi-modal, especially involving images and videos, the main approaches can be categorized into two types in Tab. 1: i) treating LLM as a scheduler, ii) treating LLM as a decoder.

LLMs as scheduler In the scheduler-based methods, various visual models are treated as plug-and-play modules. LLM schedules them according to the specific visual task requirements, like the assembly of building blocks. Some of these methods focus on images, such as VisualChatGPT and HuggingGPT , while MM-REACT and ViperGPT can also handle videos. A key characteristic of these scheduler-based LVLMs is that they do not require end-to-end training, hence eliminating the need for pre-alignment and joint training of each modality.

LLMs as decoder Regarding the approach of treating LLM as a decoder, this is our primary focus. MiniGPT-4 aligns image tokens to the input of the large language model through several linear projection layers. However, this alignment is weak and lacks feedback from human instructions. Subsequently, mPLUG-Owl adopts a two-stage training approach. In the first stage, images are aligned with language using an auto-regressive pretraining style, and the second stage involves instruction tuning through using a human instruction dataset. With the increasing scale of large language model backends, approaches such as InstructBLIP and LLaVA collecte the larger human instruction datasets to train a larger LVLMs (e.g. 13B parameters). Each answer of instruction datasets strictly follow to the given instructions. Then they undergo end-to-end training using human instruction datasets, enabling the LLM with visual reasoning capabilities. Moreover, Video-ChatGPT design a 100k video instruction dataset, successfully empowering LLMs to comprehend videos. VideoChat and Video-LLaMA achieve this by conducting joint training, allowing LLMs to simultaneously handle images and videos. Expanding LLMs to additional visual modalities typically requires pre-alignment, as seen in LLaMA-Adapter and ImageBind-LLM . They bind other modalities to the image space through ImageBind’s modality encoder. These models have demonstrated that a unified feature space is advantageous for enhancing LLM’s multi-modal reasoning capabilities. Distinguished from prior work, Video-LLaVA not only pre-aligns image and video features but also conducts joint training of images and videos, facilitating LLMs in learning multi-modal reasoning capabilities from a unified visual representation.

Video-LLaVA

Framework Overview As shown in Fig. 2, Video-LLaVA consists of LanguageBind encoders fVf_{\mathbf{V}}M to extract features from the raw visual signal (e.g. images or videos), a large language model fLf_{\mathbf{L}} such as Vicuna, visual projection layers fPf_{\mathbf{P}} and a word embedding layer fTf_{\mathbf{T}}. We initially obtain visual features using LanguageBind encoders. LanguageBind encoders are capable of mapping different modalities into the textual feature space, thereby providing us with a unified visual representation. Subsequently, the unified visual representation is encoded by shared projection layers, which is then combined with tokenized textual queries and fed into a large language model to generate corresponding responses.

United Visual Representation Our goal is to map images and videos into a shared feature space to enable the large language model to learn from a unified visual representation. We assume that the same information can be conveyed through multiple media. For example, a running dog can be expressed through language, a image or a video simultaneously. Therefore, we can compress information from different modalities into a common feature space, allowing the model to extract information from a dense feature space, facilitating modality interactions and complementarity. Hence, we chose the modality encoders from LanguageBind , which align images and videos with the textual feature space.

Alignment Before Projection Specifically, LanguageBind initializes from OpenCLIP , naturally aligning images and language in a shared feature space. Subsequently, it aligns video representations to the language space using 3 million video-text pairs from VIDAL-10M . By sharing a language feature space, the image and video representations ultimately converge into a unified visual feature space, which we refer to as emergent alignment of images and videos. Therefore, our video encoder and image encoder are initialized from the LanguageBind encoders zoo, pre-aligning the inputs for LLM and reducing the gap between representations of different visual signals. The unified visual representation is fed into LLM after passing through a shared projection layer.

2 Training Pipeline

Overall, the process of generating responses by Video-LLaVA is similar to that of a large language model (e.g. GPT series). Given a textual input XT\mathbf{X}_{\text{T}} and visual signals XV\mathbf{X}_{\text{V}}, the input signals are encoded into a sequence of tokens according to Eq. 1. By maximizing the likelihood probability in Eq. 2, the model ultimately achieves multi-modal understanding capabilities.

where LL is the length of the generated sequence XA\mathbf{X}_{\text{A}}, and θ\theta is a trainable parameter. We dynamically conduct joint training on images and videos, wherein a single batch contains both image and video samples simultaneously.

where rr represents the round number. As shown in Eq. 3, when r>1r>1 we concatenate the conversations from all previous rounds with the current instruction as the input for this round. The training objective remains the same as in the previous stage. After this stage, the model learns to generate corresponding responses based on different instructions and requests. The LLM are also involved in training at this stage.

Experiments

Model Settings We employ Vicuna-7B v1.5 as the large language model. The visual encoders are derived from LanguageBind, initialized from ViT-L/14. The text tokenizer is sourced from LLaMA, with approximately 32,000 classes. The share projection layers consist of 2 fully connected layers.

Data Details As shown in Fig. 3, for the stage of understanding pretraining, we use a subset of 558K LAION-CC-SBU image-text pairs with BLIP captions, which is sourced from CC3M and filtered by Liu et al. . The video-text pairs are derived from a subset provided by Valley , and we have access to 702k out of a total of 703k pairs, originating from WebVid . For the stage of instruction tuning, We gathered instructional datasets from two sources, including a 665k image-text instruction dataset from LLaVA v1.5 and a 100k video-text instruction dataset from Video-ChatGPT.

Training Details In the training process, we resize and crop each image, resulting in a size of 224×224 for each processed image. We uniformly sample 8 frames from each video, and each frame undergoes image pre-processing. The data in each batch is a random combination of images and videos. In the first stage, we train for one epoch with a batch size of 256, using the AdamW optimizer with a cosine learning rate schedule. In the second stage, we reduce the batch size to 128. The initial learning rate for both stages is set to 1e-3, with a warmup ratio of 0.03. Additional hyper-parameter settings can be found in the appendix.

2 Quantitative Evaluation

As shown in Tab. 2, Video-LLaVA achieves the best performance on 8/9 image understanding benchmarks, and ranks the second on the other.

Zero-shot Image Question-answering To begin with, We evaluate our approach for image understanding on five academic image question-answering benchmarks. Compared to the state-of-the-art model InstructBLIP-7B, Video-LLaVA demonstrates powerful image understanding capabilities, outperforming across all five question-answering benchmarks. Additionally, Video-LLaVA exhibits competitive results compared to several more powerful LVLMs, which are tuned based on 13B or 65B LLM, such as surpassing InstructBLIP-13B by 14.7% on VisWiz, highlighting its strong understanding ability in natural visual environments.

Evaluation under Benchmark Toolkits Additionally, we evaluate LVLMs using several benchmark toolkits for visual instruction tuning. These benchmark toolkits provide a detailed assessment of the model’s capabilities through robust evaluation metrics. Video-LLaVA outperform InstructBLIP-7B by 24.9%, 12.2%, and 5.8% on MMBench, LLaVA-Bench, and MM-Vet, respectively. It is worth noting that Video-LLaVA-7B still demonstrates advanced performance compared to larger LLM models, surpassing InstructBLIP-13B by 6.4% on MM-Vet and IDEFICS-80B by 6.4% on MMBench. These results demonstrate that Video-LLaVA exhibits a strong understanding of semantic aspects of scenes, enabling it to answer open-ended and free-form natural language questions about images.

Zero-shot Video Understanding As shown in Tab. 3, we conduct a quantitative assessment of the video question-answering capabilities of large video-language models on four datasets, including MSVD-QA , MSRVTT-QA , TGIF-QA and ActivityNet-QA . The evaluation pipeline for video understanding follows Video-ChatGPT. We report the accuracy and score, which is assessed using GPT-Assistant. Video-LLaVA consistently outperforms Video-ChatGPT in terms of question-answering accuracy, which is an advanced large video-language model. Moreover, Video-LLaVA surpasses the powerful baseline of Video-ChatGPT by 5.8%, 9.9%, 18.6%, and 10.1% on MSRVTT, MSVD, TGIF, and ActivityNet, respectively. Additionally, we conduct comparisons with the recent SOTA model, Chat-UniVi . Despite Chat-UniVi utilizing more datasets such as MIMIC-IT , Video-LLaVA still demonstrate competitive results, surpassing Chat-UniVi on MSVD, MSRVTT, and TGIF datasets. In summary, these results validate Video-LLaVA’s ability to comprehend videos and provide contextually appropriate responses based on instructions.

Object Hallucination Evaluation As shown in Tab. 4, we report evaluation results for zero-shot object hallucinations, utilizing a evaluation pipeline derived from a polling-based query method . Video-LLaVA demonstrates competitive performance across three subsets: random, popular, and adversarial. Specifically, when compared to the 7B foundation model, Video-LLaVA consistently outperforms MM-GPT across all three POPE hallucination evaluation subsets. Furthermore, when benchmarked against the larger 13B LLM, Video-LLaVA even surpasses Mini-GPT4 comprehensively. The successful performance of Video-LLaVA in object hallucination detection validates the consistency between unified visual representations and the generation of textual descriptions.

Exhibition Board In Fig. 4, we select several classic examples to explore the multi-modal understanding capabilities of Video-LLaVA. For image understanding, we compare it with GPT-4. The first two images are from GPT-4, while the last image is from LLaVA. The responses from Video-LLaVA are more comprehensive, intuitive, and logical compared to GPT-4. For example, in the first image, Video-LLaVA not only predict what is about to happen but also identify that the glove is red and the ball is blue, which GPT-4 fail to recognize. For video understanding, we do not carefully select the videos. Videos are sourced from Video-ChatGPT, which is an advanced large video-language modeL. Overall, we observe that the sentences generated by Video-LLaVA and Video-ChatGPT are very similar. However, Video-LLaVA excel at extracting key information from the videos based on the given instruction, as demonstrated by the highlighted purple text. Furthermore, leveraging a unified visual representation, we observe that Video-LLaVA demonstrates the capability to comprehend inputs that consist of both images and videos simultaneously. As depicted by the bold font in Fig. 4, it serves as compelling evidence that a LLM backend possesses robust handling abilities for both images and videos. These results demonstrate that Video-LLaVA possesses the ability to understand both images and videos, learned from a unified visual representation.

3 Ablation Results

To validate the performance degradation caused by separated visual representation, we conduct experiments to to explore the performance of the LLM learning from different visual representations. We define the use of LanguageBind image encoder as unified visual representation while the MAE encoder is separated visual representation, which is a well-known and effective image feature extractor. We only replace the image encoder with the MAE image encoder of the same scale and keep the LanguageBind video encoder. We compare the united visual representation and the separated visual representation on 13 benchmarks, including 9 image understanding benchmarks and 4 video understanding benchmarks.

For Image Understanding The unified visual representation demonstrates strong performance, surpassing the separated visual representation comprehensively across 5 image question-answering datasets and 4 benchmark toolkits in Fig. 5. Additionally, we observe a significant margin of performance improvement in the unified visual representation on the POPE, MMBench, LLaVA-Bench, and MM-Vet benchmark toolkits. This highlights that the unified visual representation not only enhances performance in image question-answering but also provides benefits in other aspects of image understanding, such as reducing object hallucination and improving OCR capabilities.

For Video Understanding Due to replacing the image encoder with the MAE encoder, the video features and image features are no longer unified during LLM’s initial learning of visual representations. In Fig. 6, compared to separated visual representation, the united visual representation significantly improves performance across 4 video question-answering datasets. Separated visual representations not only exhibit lower accuracy in question-answering, but also demonstrate a similar trend in answer scores. These results demonstrate that the unified visual representation can help the LLM further learn and understand videos.

3.2 Joint Training

This subsection aims to validate the complementarity of images and videos during joint training, which can mutually enhance the LLM’s understanding of images and videos based on a unified visual representation.

For Image Understanding As shown in Fig. 7, We find that both images and videos benefit from joint training, demonstrating mutual improvement in visual understanding. In comparison to LLaVA, we conduct evaluations of image question-answering on VisWiz, focusing on three aspects: i) unanswerable, predicting whether visual questions are unanswerable; ii) number, tasks related to numerical understanding; and iii) other, additional visual understanding tasks. Video-LLaVA outperform LLaVA in unanswerable and number tasks, indicating that joint training with videos alleviates the object hallucination in images and enhances the understanding of numerical signals in images. A similar trend is observed on the LLaVA-Bench, where video data significantly improves LLM’s performance in complex reasoning and image conversation tasks.

For Video Understanding In Tab. 5, we evaluate our model on four video question-answering datasets. Compared to Video-LLaVA∗ without image in training, the model trained with joint images and videos achieves comprehensive improvements across all four video datasets. These results demonstrate that joint training of images and videos facilitates LLM’s understanding of visual representations.

Conclusion and Future Directions

In this work, we introduce Video-LLaVA, a simple but powerful large visual-language baseline model. We propose a novel framework to address the issue of misalignment before projection, utilizing a LanguageBind encoder to pre-bind visual signals into the language feature space. To enable a LLM to comprehend both images and videos simultaneously, we conduct joint training on images and videos, allowing the LLM to learn multi-modal interactions from a unified visual representation. Extensive experiments demonstrate that joint training on images and videos mutually benefits performance. Furthermore, we validate that aligning visual representations before projection aids LLM learning. Remarkably, LLM, after learning from a unified visual representation, exhibits the remarkable ability to simultaneously engage with both images and videos, showcasing a powerful comprehension of unified visual concepts. These results collectively demonstrate the effectiveness of the Video-LLaVA training framework. As a unified visual training framework, the performance of Video-LLaVA even surpasses that of expert models designed specifically for images or videos.

Future work While Video-LLaVA exhibits strong competitiveness in both images and videos, we observe that it faces difficulty in grasping temporal relationships and spatio-temporal localization. Video-LLaVA can serve as a baseline to extend to additional visual-related modalities, such as depth and infrared images. Additionally, we could explore how to incorporate timestamp embeddings effectively, enabling large visual-language models to answer questions related to temporal relationships.

References