MM-VID: Advancing Video Understanding with GPT-4V(ision)
Kevin Lin, Faisal Ahmed, Linjie Li, Chung-Ching Lin, Ehsan Azarnasab, Zhengyuan Yang, Jianfeng Wang, Lin Liang, Zicheng Liu, Yumao Lu, Ce Liu, Lijuan Wang
Introduction
People around the world create numerous videos on a daily basis , including user-generated live streams, video-game live streams, short clips, movies, sports broadcasts, advertising, and more. Videos serve as a versatile medium for conveying information and content through various modalities , such as text, visuals, and audio. Developing methods that can learn from diverse modalities will enable us to design cognitive machines with enhanced capabilities for analyzing uncurated real-world videos, extending beyond the confines of hand-curated datasets. However, this rich representation introduces many challenges for the study of video understanding, particularly when dealing with extended-duration videos .
Understanding long videos, especially those spanning over an hour, is a complex task that demands advanced methods capable of analyzing sequences of images and audio across multiple episodes. This challenge is compounded by the need to extract information from various sources, such as distinguishing speakers , identifying characters , and maintaining narrative coherence . Additionally, answering questions based on video evidence requires a deep comprehension of the content, context, and subtitles. When it comes to live streaming and gaming videos , there are challenges in processing dynamic environments in real-time, requiring semantic understanding, and the ability of long-term strategy planning .
Recently, substantial advances have been made with large pre-trained video models and video-language models , which have demonstrated their reasoning capabilities for video content. However, these models are usually trained on short clips (e.g., 10-second videos in Kinetics and VATEX ) or pre-defined action classes (e.g., 174 classes in Something-Something v1 ). Consequently, these models may fall short in providing a detailed comprehension of intricate videos in real world . To achieve a more comprehensive understanding of the videos we encounter in daily life, we need methods capable of addressing complex challenges. It involves not only identifying who are in the scene and what they do, but also pinpointing when and how they act, while recognizing subtle nuances and visual cues across different scenes. The aim of this work is to address these challenges and explore methods that can be applied directly to real-world video understanding. Our approach involves breaking down extended video content into coherent narratives and subsequently employing these generated stories for video analysis.
Recent advances in Large Multimodal Models (LMMs) , such as GPT-4V(ision) , have demonstrated significant breakthroughs in processing both input images and text for multimodal understanding. This has sparked interest in applying LMMs to the video domain. In this work, we present MM-Vid, a system that integrates specialized tools with GPT-4V for video understanding. Given an input video, MM-Vid performs multimodal pre-processing, including scene detection and automatic speech recognition (ASR), to collect important information in the video. The input video is then split into multiple clips according to the scene detection algorithm. Then, we employ GPT-4V, which takes the clip-level video frames as input and generates a detailed description for each video clip. Finally, GPT-4 is adopted to generate a coherent script for the full video, conditioning on the clip-level video descriptions, ASR, and video metadata if available. As shown in Figure 1, the generated script allows MM-Vid to perform a diverse set of video tasks.
Experimental results demonstrate the effectiveness of MM-Vid in different challenging scenarios. MM-Vid is able to comprehend hour-long videos through multiple modalities, and localize specific events with correct timestamps. MM-Vid also demonstrates intriguing results in an interactive environment, such as predicting the possible next steps when playing a video game or interacting with a graphical user interface (GUI) .
Related Work
Conventional Video Understanding Methods. Early work in computer vision centered on building video foundation models . These models, with different neural network architecture designs and training methods, have achieved great breakthrough at analyzing short video clips , typically lasting less than 30 seconds. However, these models are typically pre-trained with vision modality only, and thus may require specific adjustment or fine-tuning for multimodal downstream tasks.
Video-Language Models. Recent studies have made remarkable improvements in multimodal representation learning for video-and-language understanding. These advancements have been particularly evident in popular downstream tasks such as video question answering , text-video retrieval and video captioning . Building on this momentum, researchers typically embark on a pretrain-finetune paradigm: initially pre-training a video-language foundation model on large-scale video-text pairs, followed by a fine-tuning process on specific downstream datasets. However, these methods are usually trained on short video clips, often restricted to durations of around 10 seconds, posing potential challenges in comprehending longer video sequences.
Visual Instruction Tuning. Inspired by the breakthrough of Large Language Models (LLMs) , recent studies suggest using a frozen LLM combined with an image encoder and a few learnable modules for video understanding tasks. Specifically, researchers propose the visual instruction tuning , which aims to fine-tune the learnable modules and thus enable LLMs to generate textual descriptions for the video content. While promising performance is presented, these models may fall short when it comes to handling videos with extended duration. Our work aims to fill this gap, exploring methods that can be directly applied to the understanding of long videos in real-world situations.
Prompting LLMs for Video Understanding. Recently, researchers explore the LangChain system paradigm , which aims to integrate expert tools with existing LLMs to create new functionalities. For example, VLog uses BLIP2 and GRIT as dense image captioners, Whisper as ASR translator, and ChatGPT as a reasoner. By transcribing a given video to textual descriptions (e.g., document), it enables ChatGPT for video question-answering tasks. Inspired by the efficacy of these tool-using approaches , we explore integration with GPT-4V for video understanding.
Preliminary Study with GPT-4V(ision)
Recent studies show that GPT-4V can accept a range of inputs, such as textual descriptions, questions, or even visual cues like images or short video clips. GPT-4V’s inherent ability to comprehend visual inputs and generate contextually relevant text opens the door for a wide range of applications. By introducing a sequence of frames as input, GPT-4V can grasp temporal relationships and interactions, aiding in the identification and interpretation of dynamic visual content.
MM-Vid
Figure 2 shows the overview of our system pipeline. MM-Vid takes the video file as input, and outputs a script describing the video contents. The generated script enables LLMs to achieve various video understanding capabilities. MM-Vid consists of four modules: (i) Multimodal Pre-Processing, (ii) External Knowledge Collection, (iii) Clip-Level Video Description Generation, and (iv) Script Generation. We describe each module in detail below.
Multimodal Pre-Processing. Starting with an input video file, our process begins by using the established ASR tool to extract transcriptions from the video. Following this, we divide the video into several short video clips. This process involves uniform sampling of video frames, with each clip consisting of 10 frames. To enhance the overall quality of frame sampling, we use established scene detection tools like PySceneDetect to help identify crucial scene boundaries.
External Knowledge Collection. We incorporate external knowledge into our input prompts to GPT-4V. This involves gathering available information, such as metadata, title, abstract, and face photos of characters within the video. In our experiments, the metadata, title, and abstract are gathered from YouTube.
Clip-Level Video Description Generation. During our multimodal pre-processing, the input video is segmented into multiple clips. For each clip, which typically consists of 10 frames, we employ GPT-4V to generate video descriptions. By feeding the video frames along with the associated text prompt into the model, GPT-4V utilizes the input to generate detailed descriptions that capture the visual elements, actions, and events depicted in those frames.
In addition, we explore the use of visual prompting, where the character’s face photos are presented alongside the character’s name in the input to GPT-4V. Our empirical results suggest that visual prompting is helpful to enhance the quality of video descriptions, particularly for more accurate character identification. These findings align with the insights from .
Script Generation using LLM. After generating the descriptions for each video clip, we use GPT-4 to integrate these clip-level descriptions into a coherent script. This script serves as a comprehensive description of the entire video, and is used by GPT-4 for a diverse set of video understanding tasks.
MM-Vid for Streaming Inputs
Figure 3 shows the diagram of MM-Vid when applied to the context of streaming inputs. Our system operates as an agent within a dynamic environment where streaming video frames serve as the primary input. In this context, the agent continually receives streaming video frames as states, representing the ongoing visual information unfolding in the environment. These states are then processed by GPT-4V to make informed decisions and generate responses.
By continually analyzing the streaming video frames, MM-Vid plays a crucial role in transforming raw visual data into meaningful insights, making it valuable for applications such as video game play, the embodied agent, and GUI navigation.
Experiments
We implement MM-Vid based on MM-ReAct codebase. We use the Automatic Speech Recognition (ASR) tool publicly available via the Azure Cognitive Services APIs , and utilize PySceneDetect for scene detection.
2 MM-Vid Capabilities
Figures 4-9 provide illustrative examples of MM-Vid’s complete execution flow. When a user uploads a video file, MM-Vid initiates the process by first assessing the estimated video length. Subsequently, it performs multimodal pre-processing by invoking expert tools, including scene detection and ASR. Additionally, MM-Vid collects external knowledge, encompassing video metadata such as title and abstract.
Following this preliminary stage, MM-Vid proceeds to generate clip-level video descriptions for each segment of the video. Finally, it invokes GPT-4, integrating these clip-level descriptions into a coherent script. Once the script is generated, it empowers LLMs to provide a summarized understanding of the video content. That equips the system to address users’ questions with grounded answers. We discuss MM-Vid’s distinct capabilities as below.
Grounded Question-Answer (QA). The generation of a comprehensive script empowers our system with the capability of grounded QA. As shown in Figure 8, let us consider a scenario where a user poses the question, “Show me the most exciting moment in this video.” In response, MM-Vid displays a highlight, specifically featuring a home run, and provides the corresponding timestamp. When a user asks “Who are the best pitchers in this video?” MM-Vid addresses the question by referring to relevant evidence in the generated script. This grounding capability owes its success to the extensive and detailed script generation process, which documents essential timestamps and significant events within the video, enabling accurate and contextually grounded responses to user inquiries.
Multimodal Reasoning. MM-Vid considers multimodal inputs, including video frames, speech transcriptions, and external knowledge if available. In Figure 8, when a user inquires, “How did you know the sound is different?” MM-Vid explains that this information was derived from the commentator’s remarks during the game. The examples illustrate MM-Vid’s multimodal reasoning capabilities, where it integrates both visual and auditory cues to provide contextually accurate responses to user queries.
Hour-Long Video Comprehension. Figures 10-13 demonstrate MM-Vid’s capabilities in processing lengthy videos. In this example, MM-Vid effectively analyzes a documentary video spanning approximately 50 minutes in duration. For simplicity, the intermediate outputs are omitted in the figures, and only the final generated script is presented. We observe that MM-Vid is able to generate a long script with the corresponding timestamps to represent the documentary video. By leveraging this generated script as contextual information, MM-Vid is equipped to perform a range of tasks, including summarizing the lengthy video, addressing specific queries raised within the video, and indexing pivotal moments.
Multi-Video Episodic Analysis. MM-Vid’s proficiency in handling extensive video content can be expanded to encompass multiple lengthy videos, as illustrated in Figures 14-16. In these examples, we upload multiple episodes to MM-Vid, showcasing its ability to perform a variety of complex tasks. MM-Vid exhibits the capability to summarize the video series, engage in cross-episode reasoning, provide detailed descriptions of character journeys across multiple episodes, and facilitate grounded QA interactions.
Character Identification. We found that incorporating visual prompts enhances the quality of script generation, particularly with regards to character identification. In Figure 17, we illustrate this by providing MM-Vid with additional inputs consisting of characters’ face photos and their corresponding names. MM-Vid effectively utilizes these visual prompts to identify the characters depicted in the video, based on the provided face photos. As a result, the script generation process is notably improved, ensuring more accurate and contextually relevant descriptions of characters and their interactions within the video content.
Speaker Identification. Our exploration has revealed another valuable application of visual prompting in enhancing the quality of Automatic Speech Recognition (ASR). In Figures 18-19, we highlight a scenario where conventional ASR struggles to accurately recognize the number of speakers and their identities in the video. Visual prompting plays a pivotal role in enhancing ASR performance by providing contextual cues to identify individuals and attribute speech to specific speakers. This improvement ensures more precise transcriptions, enabling a more accurate representation of the dialogue and interactions within the video content.
Audio Description Generation. Audio descriptions play a crucial role in making videos accessible to individuals who are blind, have low vision, or face difficulties in visually understanding the content. These descriptions provide contextual narration of meaningful visual elements, clarify speakers, and convey the essence of visual information within a video. In our experiments, we also explore MM-Vid’s performance in audio description generation. We experiment with videos where there is limited or no speech content. In Figure 20, we showcase an example featuring a short film of Mr. Bean taking an exam, which primarily lacks speech. Without ASR inputs, MM-Vid processes the video and generates a detailed script. This shows MM-Vid’s versatility in handling various types of video content and its potential in creating inclusive and accessible multimedia content.
Self-Refinement. While the generated script offers a comprehensive understanding of video content, our experiments have unveiled occasional inaccuracies, especially in cases involving blurry or low-resolution video frames, as demonstrated in Figure 21. In this example, MM-Vid mistakenly identifies a bird as a rock due to the challenges posed by the video’s visual quality. To address such inconsistencies and elevate the overall accuracy of the generated script, we employ a self-refinement approach . This involves revising the script based on both the initially generated script and a concurrently generated video summary. Through this process, MM-Vid is able to rectify errors and inaccuracies, resulting in a more refined output.
Fast-Changing Short Videos. In Figure 22, we present an example of our experimentation with fast-changing short-form videos, such as those found on platforms like TikTok. Short videos often feature non-standard frame sizes and significantly shorter durations compared to conventional videos. Remarkably, MM-Vid excels at accurately describing the cooking recipes depicted in these short videos, despite the distinct characteristics of such content.
These examples demonstrate the versatility of MM-Vid in processing a diverse array of video content. Whether dealing with lengthy documentaries, episodic series, or short-form clips, MM-Vid adapts seamlessly to the unique attributes of each video type, consistently delivering meaningful and contextually relevant descriptions.
3 Applications to Interactive Environments
In the following section, we evaluate MM-Vid when applying to the context of streaming inputs. MM-Vid serves as an agent in an interactive environment, continually receiving streaming video frames as the inputs.
Embodied Agent. Figure 23 illustrates an example where MM-Vid is applied to an egocentric video captured by a head-mounted camera. This video, collected from Ego4D dataset , provides a brief glimpse into the wearer’s daily life within their home environment. Remarkably, MM-Vid showcases its capability in understanding such video content and assists the user in a few practical tasks. Specifically, MM-Vid helps the user locate items like the pink jacket and the laptop within the home. Additionally, it generates a list of the user’s activities within a specified time range, offering insights into the wearer’s daily routine.
Playing Video Games. Figures 24-27 demonstrate the results of applying MM-Vid to a Mario video game . In these experiments, our agent consistently receives three video frames as states and calculates the next possible control action. Remarkably, our agent displays an understanding of the specific video game dynamics and generates reasonable action controls to play the game effectively. These examples highlight MM-Vid’s ability to comprehend and navigate in an interactive gaming environment. Interested readers may find the full gameplay demonstration on our project website.
GUI Navigation. Figures 28-32 provide the demonstration of MM-Vid’s performance in the GUI navigation scenario. In this context, the agent continually receives iPhone screenshots and previous user actions as states. The agent effectively predicts the possible next steps in the user’s journey, which may include clicking on the correct shopping apps, initiating searches for items of interest, and ultimately placing an order. These results demonstrate MM-Vid’s remarkable ability to interact with graphical user interfaces, facilitating seamless and intelligent navigation through digital interfaces.
4 User Study
We explore the potential of MM-Vid for people who are blind or have low vision. Audio description (AD) provides an auditory narration integrated into the video’s soundtrack, offering important visual details that may not be discernible from the main video soundtrack. Such descriptions play a pivotal role in conveying essential visual content to those with visual impairments.
To assess the efficacy of MM-Vid in generating audio descriptions (AD), we conduct a user study. We invited 9 participants for the evaluation. 4 participants were either blind or had low vision, while the remaining 5 had normal vision. All the participants have normal hearing. For the purposes of the experiments, we segregated participants into two distinct groups: (i) Group with visual impairments, and (ii) Group with normal vision.
Our experiments utilize a curated set of videos, which are mainly suggested by the American Council of the BlindThe Audio Description Project: https://adp.acb.org/. We also collected accessibility videos from YouTubeApple Accessibility: https://www.youtube.com/watch?v=SL7YSqlEd8k. For every video used in our evaluation, participants are exposed to two versions: the first containing human-crafted AD and the second powered by MM-Vid-generated AD. Both renditions are narrated using text-to-speech (TTS) technology.
We have designed two questionnaires for the two groups, referenced in Table 1 and Table 2, respectively. Participants with visual impairments are instructed to base their evaluation exclusively on auditory cues. In contrast, those with normal vision are instructed to consider both visual and auditory elements.
The assessment adopts the standardized Likert scale for ratings. For each posed question, participants are guided to assign a score ranging from 0 to 10, with higher values indicating more favorable feedback. Furthermore, participants are urged to share feedback and remarks concerning their overall experience.
4.2 Results on the Group with Visual Impairments
We utilized 3 different videos for our evaluation, with durations of 1 minute, 1 minute 42 seconds, and 2 minutes 42 seconds, respectively. Each of the 4 participants with visual impairment was well versed with screen reader and other common accessibility tools. After listening to the audio descriptions for each video, they were asked to respond to the 4 questions outlined in Table 1. Hypotheses and Results H1: The MM-Vid-generated audio description and original video dialogues are effectively presented to the participants. Results: Using the Likert scale (0=Not Effective to 10=Most Effective) the participants rated the effectiveness of the delivery of human-crafted AD and MM-Vid-generated AD. On average, participants gave for MM-Vid-generated AD and for human-crafted AD, which shows a MM-Vid-generated AD very close to human-crafted one in terms of effective delivery (Figure 5). H2: Participants are able to follow the main story line of the video based on MM-Vid-generated audio description only. Results: Using the Likert scale (0=Not Informative to 10=Highly Informative) the participants rated the informativeness of human-crafted AD and MM-Vid-generated AD. On average, participants gave for MM-Vid-generated AD and for human-crafted AD, which shows little difference in informativeness between MM-Vid-generated AD and human-crafted one (Figure 5). H3: MM-Vid-generated AD and human-crafted AD are close in terms of voice and audio quality. Results: Using the Likert scale (0=Low Quality to 10=High Quality) the participants rated the voice and audio quality on average as for MM-Vid-generated AD and for human-crafted AD. This minimal difference between the scores indicates the close-to-human voice and audio quality of MM-Vid-generated AD (Figure 5). Discussion: The results show that the participants’ overall satisfaction of MM-Vid-generated ADs was on average around 2 points less than human-crafted ones in the Likert scale (0=Not Satisfied to 10=Highly satisfied) (Figure 5). Some of the difficulties indicated by participants while listening to MM-Vid-generated ADs were 1) occasional overlaps between AD audio and original video dialogues 2) wrong descriptions due to hallucinations of GPT-4V(ision). Regardless of the difference in overall satisfaction, all the participants agreed that MM-Vid-generated AD can provide a cost-effective and scalable solution. Thus, millions of videos that cannot afford to be professionally audio described, can be auto-processed by a tool like MM-Vid to make them accessible to the visual-impaired community.
4.3 Results on the Group with Normal Vision
For sighted individuals, we used the same set of videos as we used for individuals with visual impairments. All of our 5 participants answered to 6 questions listed in Table 2 after watching videos embedded with MM-Vid-generated AD as subtitles and audio track. Hypotheses and Results H1: The MM-Vid-generated AD is accurate and conveys essential information without overloading the listener. Results: The sighted individuals rated the clarify and accuracy of MM-Vid-generated AD as and human-curated AD as on average, using the Likert scale (0=Not Accurate to 10=Most Accurate). In terms of conciseness, the participants on average gave for the MM-Vid-generated AD and for human-curated AD based on the Likert scale (0=Not concise to 10=Most concise). These results indicate MM-Vid-generated ADs are close to human-curated ones in terms of accuracy and conciseness (Figure 6). H2: The MM-Vid-generated ADs are in sync with visual content and do not overlap with other dialogues ensuring listeners can follow the story line. Results: Participants gave on average and to human-crafted AD and MM-Vid-generated AD respectively using the Likert scale (0=Not Informative to 10=Highly Informative). Human-crafted AD and MM-Vid-generated AD received and respectively on the aspect of timing and synchronization using the Likert scale (0=Not Effective to 10=Most Effective). These indicates while listening to MM-Vid-generated ADs participants were able to follow main story line and found the audios are in sync with video content very close to that of human-crafted ADs (Figure 6). H3: The voice and audio quality of MM-Vid-generated ADs are close to human-crafted ADs. Results: The results are very similar to results on group with visual impairments. Sighted participants rated the voice and audio quality on average as for MM-Vid-generated AD and as for human-crafted AD. Therefore the voice and audio experience did not degrade much while listening to MM-Vid-generated ADs compare to human-crafted ADs (Figure 6). Discussion: The evaluations on sighted individuals helped to cross verify the hypotheses of individuals with visual impairments, that are based on audio cues only. Although the overall satisfaction points for sighted participants with MM-Vid-generated ADs was on average 1 points lower than human-generated ADs (Figure 6), the overall satisfaction points for participants who were blind was worse. This is expected because sighted individuals had access to both audio and video modalities but individuals with visual impairments did not. We also believe the reason for lower overall satisfaction, may have been the lack of practice listening to auto generated ADs. Some of the users also mentioned they have preference between pitches of voice and number of concurrent audio channels. These may add to the reason of lower overall satisfaction.
4.4 Participant Feedback
We present a collection of interview quotes from our participants who were visually impaired, in which they share their personal experiences and insights about the audio descriptions (AD) generated by MM-Vid. The participants expressed a unanimous desire to continue utilizing this AD generation service in the future, highlighting its exceptional quality (“Nearly perfect”), intricate details (“favorite was the details”), extensive applicability (“allowed me to follow anything visual”), and the profound impact it has on them (“I did not depend on someone else”). Below, we provide additional quotes for further insight.
P1: “I understand what is going on very quickly and I did not depend on someone else.” P2: “If it’s AI-generated, there are so many places it’s not available, and we need it there.” P2: “First time listening to auto-generated AD. As a user, if I am offered this AD, I would take it.” P3: “Nearly perfect. Most favorite was the details.” P3: “More information helped me follow the storyline.” P3: “It allowed me to follow anything visual. It felt natural the way AD describes how the actor interacts with the environment.” P3: “I love animal kingdom, and I watch Wild Earth safari virtual tour. I would love to have audio descriptions of Wild Earth videos and daily safaris.” P4: “I would like to have auto-generated audio description for live conferences in Microsoft Teams.” P4: “It worked best as the original audio had not much value.”
Despite the positive feedback, not all responses were favorable:
P4: “I am skeptical when it becomes subjective. Sometimes I feel they make up stories which is not good.” P4: “After listening to the human-generated AD, I figured I misunderstood parts of the original story.” P1: “It keeps referring to the same person using their names instead of pronouns.” P4: “I don’t deal well with overlapped or two parallel audios.”
Interestingly, even those participants who provided critical feedback still rated the MM-Vid-generated AD closely to human-generated AD, during the questionnaire sessions. This indicates that, similar to human-curated AD, adapting to MM-Vid-generated ADs might necessitate some practice and acclimatization over time.
Conclusion
We have presented MM-Vid, a system that synergizes with GPT-4V for advancing video understanding. MM-Vid employs GPT-4V to transcribe video content into long and detailed scripts, thereby enriching LLMs with advanced video understanding capabilities. Experimental results demonstrate the effectiveness of MM-Vid in addressing challenging tasks, including comprehension of hour-long videos, analysis across multiple episodes, identification of characters and speakers, and interaction with video games and graphical user interfaces.
Beyond the development of the MM-Vid system, we conducted an extensive user study, drawing feedback from a varied group of participants. The outcomes of this study indicated that the audio descriptions generated by MM-Vid closely mirror the quality of those crafted by humans. In our future work, we plan to explore SoM and object tracking techniques to enhance various tasks and functionalities.
We are deeply grateful to OpenAI for providing access to their exceptional tool . We are profoundly thankful to Misha Bilenko for his invaluable guidance and support. We also extend heartfelt thanks to our Microsoft colleagues for their insights, with special acknowledgment to Cenyu Zhang, Saqib Shaikh, Ailsa Leen, Jeremy Curry, Crystal Jones, Roberto Perez, Ryan Shugart, Anne Taylor for their constructive feedback.