MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Kirolos Ataallah, Xiaoqian Shen, Eslam Abdelrahman, Essam Sleiman, Deyao Zhu, Jian Ding, Mohamed Elhoseiny
Introduction
In recent years, Large Language Models (LLMs) research has witnessed remarkable advancements, with prominent models like GPT-4 , Llama 2 , and Mistral showcasing unprecedented capabilities in processing and generating textual data. However, typical LLMs are inherently limited to text-centric tasks and do not naturally capture the multimodal nature of human interaction with the world. While some strides have been made towards integrating images into LLMs, exemplified by works such as MiniGPT and LLaVa, the incorporation of temporal information from videos remains relatively underexplored and presents significant research challenges.
Unlike static images, videos present a temporal dimension, comprising sequences of frames, essential for understanding dynamic visual content alongside textual input. In this study, we endeavor to adapt LLMs to comprehend the temporal intricacies inherent in video sequences. Previous efforts, such as Video-ChatGPT , have relied on spatial and temporal pooling techniques for fusing information across frames. However, these approaches often suffer from information loss and may not fully exploit the temporal dynamics of video data. Another example is LLaMA-VID , which attempted to address the constraints of LLMs in processing long videos by representing each frame with only two tokens, resulting in significant information loss.
In contrast, our approach leverages the concatenation of every four adjacent visual tokens, effectively reducing the token count while mitigating information loss. As depicted in Figure 2, we incorporate subtitles for each frame, allowing the representation of each frame as a combination of visual tokens extracted by the visual encoder and text tokens derived from LLM tokenizers. Such an approach enables the LLM to comprehend the video content more comprehensively, facilitating responses to both visual and textual queries.
To validate the effectiveness of our proposed methodology, we conduct thorough evaluations across multiple benchmarks. These evaluations include assessments based on the Video-ChatGPT benchmark , which evaluates aspects such as information correctness, detail orientation, contextual understanding, temporal comprehension, and consistency in video understanding. Additionally, we employ zero-shot evaluation methodologies encompassing open-ended questions and multiple-choice formats. As shown in Figure 1, the proposed MiniGPT4-Video outperforms existing state-of-the-art methods (7B) by notable margins of 4.22%, 1.13%, 20.82%, and 13.1% on the MSVD, MSRVTT, TGIF, and TVQA benchmarks, respectively.
Related work
The advancements in NLP and LLMs have inspired a wide range of methods and models known as vision language models (VLM) for understanding cross-modalities . OpenAI’s CLIP model aligns an image and language encoder on a large-scale dataset of image-text pairs with a contrastive loss, enabling it to perform multimodal-retrieval. More recent methods have harnessed the power of LLMs. For example, Flamingo leverages training on web-scraped image-text pairs to reveal the in-context learning capabilities of LVLMs. Similarly, BLIP-2 , which integrates off-the-shelf frozen pre-trained image encoders and large language models to bridge the modality gap with a Querying Transformer. There is a trend to use large-scale instruction-tuning datasets to fine-tune LVLMs. For example, LLaVA explores the instruction tuning of its vision language model on generated GPT-4 data, while MiniGPT-v2 and InstructBLIP use BLIP-2 to construct their datasets. As LLMs continue to achieve better results in image understanding, recent work has begun scaling these models to the more challenging video domain.
2 LLM-Based Video Understanding
Recently, vision-language models such as LLaVA have been extended to the video domain to process short videos 5 minutes on average or less, with similar capabilities such as visual question-answering and captioning. Video-LLaMA and VideoChat extend the BLIP-2 architecture for video embedding extraction and both employ two streams for audio and visual signals. Video-LLaMA employs a Video Q-Former and an Audio Q-Former for the two streams, while VideoChat has a video embedder and a perception toolkit for captioning, tags, etc. On the other hand, Video-ChatGPT leverages a single stream where the architecture first encodes each frame and then has a spatial and temporal pooling process that is finally mapped to an LLM with a linear layer. Video LLaVA takes advantage of the LanguageBind module to map both image and video inputs to the same embedding space. Otter proposed an instruction-tuned version of OpenFlamingo , such that it can also process multiple video frames as input.
MiniGPT4-Video
MiniGPT-v2 , has successfully translated visual features into the LLM space, enabling understanding of single images. However, extending this capability to multiple frames for video comprehension entails fine-tuning the LLM to process these frames and learn the temporal dynamics.As shown in Figure 2 Due to constraints imposed by the LLM’s context window, each video undergoes frame sub-sampling, with the number of frames (N) determined by the LLM’s context window. Subsequently, the visual frames are aligned with textual descriptions using a pre-trained model, EVA-CLIP , followed by a mapping into the large language model space using a linear layer. Similar to MiniGPT-v2 , we condense every four adjacent visual tokens in each image into a single token, thereby reducing token count per image by 75%, from 256 to 64. During training the subtitles are provided with the dataset but while inference or when there is no subtitle for the video, we utilize speech-to-text model such as whisper to generate the subtitles of the video. Frame subtitles are tokenized using the LLM tokenizer, and the visual and text tokens are concatenated for each sampled frame. Instruction tokens are appended to the end of the input sequence, and the model then outputs the answer to the question.
2 Training Pipeline
Large-scale image-text pair pretraining. In the first stage, we train a linear layer, similar as , which projects the visual feature encoded by the vision encoder (e.g. EVA-CLIP ) to the LLM’s text space with captioning loss. We leverage a combined image captioning dataset that includes images from LAION , Conceptual Captions , and SBU to align the visual feature with LLM’s input space.
Large-scale video-text pair pretraining. In the second stage, we enable the model to understand videos by taking multiple frames as input. Specifically, we sample a maximum of N frames from each video. During this stage, we use the predefined prompts in the following template: [INST]
Video question answering instruction finetuning. In this phase, we adopt the same training strategy implemented in the second stage but focus on leveraging high-quality video-question-answering datasets for instruction fine-tuning. This fine-tuning stage helps to enhance the model’s ability to interpret the input video and generate precise responses to the corresponding questions. The template is the same as the second stage with
3 Implementation Details
Throughout three training stages, we maintained a batch size of 4 and utilized the AdamW optimizer in conjunction with a cosine learning rate scheduler, setting the learning rate to 1e-4. Our visual backbone is EVA-CLIP , with the frozen weights. Notably, we trained the linear projection layer and performed efficient fine-tuning of the language model using LoRA . Specifically, we fine-tuned the and components with a rank (r) of 64 and a LoRA-alpha value equal 16. The entire model was trained with a consistent image resolution of pixels, ensuring uniformity across all stages.
Experiments
The Condensed Movies Video Captions dataset (CMD) comprises approximately 15,938 videos, each spanning one to two minutes in length. However, CMD’s captions exhibit limited quality, characterized by an average sentence length of 14 words. The Webvid dataset boasts a vast collection of two million videos. To align with CMD’s duration criteria, we refined this dataset to include videos ranging from one to two minutes in length. On the other hand, the Video Instruction Dataset offers a rich resource of 100,000 question-answer pairs distributed across 13,224 videos, distinguished by meticulous annotations. Noteworthy for its high-quality annotations, this dataset presents detailed answers to questions, averaging 57 words per sentence. Spanning diverse question types, including Video Summarization and Description-based QAs, it addresses spatial, temporal, relationship, and reasoning aspects, alongside creative or generative QAs.
Evaluation Benchmarks The Video-ChatGPT benchmark , leveraging the ActivityNet-200 dataset , is meticulously designed to evaluate video-based conversation models’ text generation capabilities across five crucial dimensions: Correctness of Information, Detail Orientation, Contextual Understanding, Temporal Understanding, and Consistency. In assessing model performance on open-ended questions, established datasets such as MSRVTT-QA , MSVD-QA , TGIF-FrameQA , and ActivityNet-QA are employed. Furthermore, for multi-choice questions, model performance is scrutinized using the TVQA dataset , which is centered around popular TV shows. The validation set comprises 15,253 QA pairs, providing a robust framework for evaluation.
2 Evaluation Metrics
Aligned with the evaluation methodology established by Video-ChatGPT , we employed GPT-3.5 turbo to juxtapose model outputs with ground truth data, subsequently computing both accuracy and a score. The accuracy metric indicates the degree of correspondence between the model’s output and the ground truth, while the score ranges from 0 to 5, signifying the level of alignment between the model output and the ground truth. A score of 0 indicates a significant deviation from the ground truth, while a score of 5 suggests close alignment. To ensure a fair and consistent comparison with the results presented in Video-ChatGPT , we adopted the same prompt for our evaluations.
3 Results
For a comprehensive evaluation of our proposed architecture, we assessed its performance across three benchmark types: Video-ChatGPT, Open-ended Questions, and Multiple-Choice Questions (MCQs). In the Video-ChatGPT benchmark, depicted in Table 1, our model is comparable with the previous methods without subtitles. When we add the subtitles as input, our model achieves the state-of-the-art in all five dimensions, which verified that our model can utilize the subtitle information to improve the video understanding. In the zero-shot evaluation of open-ended and multiple-choice question benchmarks, as illustrated in Figure 1 and Table 2, our proposed MiniGPT4-Video significantly outperforms existing state-of-the-art methods. It achieves notable margins of improvement 4.22%, 1.13%, 20.82%, and 13.1% on the MSVD, MSRVTT, TGIF, and TVQA benchmarks, respectively. The results, both with and without subtitles as shown in Table 2, further demonstrate that integrating subtitle information alongside visual cues significantly enhances performance, with accuracy rising from 33.9% to 54.21% on TVQA. While subtitles contribute substantially to performance improvements on TVQA, their inclusion doesn’t offer added value for datasets like MSVD-QA, MSRVTT-QA, TGIF-QA, and ActivityNet, where questions are exclusively vision-based.
Qualitative Results
Here in this section we show some qualitative results for our model to show the performance of it and its ability to answer different questions. For each example you can open the video link which attached to the figure description to watch the video.
Conclusion
In summary, MiniGPT4-Video offers a compelling solution for video question answering, effectively amalgamating visual and conversational comprehension within the video domain. By directly inputting both visual and textual tokens, MiniGPT4-Video empowers the Language Modeling Model (LLM) to grasp the intricate relationships between video frames, showcasing promising proficiency in understanding temporal dynamics within video content. Despite its notable achievements, MiniGPT4-Video faces a limitation imposed by the context window of the LLM. Specifically, the current version requires video lengths of 45 frames for the Llama 2 version (equivalent to less than one and a half minutes at a sampling rate of 0.5 frames per second) and 90 frames for the Mistral version (equivalent to less than three minutes). Future research endeavors will focus on extending the model’s capabilities to handle longer video sequences, thereby addressing this limitation and further enhancing its applicability and effectiveness in real-world scenarios.