Apollo: An Exploration of Video Understanding in Large Multimodal Models

Orr Zohar, Xiaohan Wang, Yann Dubois, Nikhil Mehta, Tong Xiao, Philippe Hansen-Estruch, Licheng Yu, Xiaofang Wang, Felix Juefei-Xu, Ning Zhang, Serena Yeung-Levy, Xide Xia

Introduction

Despite the rapid advancements in language and image-language modeling (chinchilla; brown2020language; yang2024qwen2; liu2024visual; alayrac2022flamingo; idefics3; openai2024gpt4o), the development of video Large Multimodal Models (video-LMMs) has not kept pace. Videos provide a rich, dynamic information source, capturing nuanced temporal and spatial features beyond the reach of static images. However, video-LMMs remain under-explored, hampered by unique challenges: notably higher computational demands and a broader, more complex design space compared to their image-based counterparts (video_chat; llama-vid; oryx; aria; xu2024pllava).

Many fundamental questions about video-LMM design remain unanswered: How should videos be sampled? Which vision encoders yield optimal representations? What are the best practices for resampling video tokens? Early approaches primarily extended image-LMMs directly (xu2024slowfast; image_grid; freeva; longva) or with video-specific fine-tuning (video_chat; video_llama; maaz2023video). Recent methods introduced diverse design choices, such as longer context windows (longva), multi-modality mixing (llava; li2024llava), agent workflows (videoagent), self-training (videostar), and more. Despite these efforts, the impact of these design decisions on video-LMM performance is poorly understood. This lack of systematic investigation motivates our study.

To overcome the computational challenges of training video-LMMs, we explore whether design decisions from smaller models correlate effectively with larger ones. Traditional scaling laws (chinchilla) predict model performance based on size, but apply to models trained from scratch and require training multiple models to predict performance. Scaling laws have also been observed in LMM pretraining (mm_scaling_laws; mm_scaling). Since LMMs integrate multiple pre-trained components, it’s uncertain if these laws hold. By relaxing scaling laws, our experiments reveal that design choices made with smaller LMMs transfer to larger ones, a phenomenon we term Scaling Consistency (Sec. 3).

Utilizing these insights, we conduct an extensive study across the video-LMM design space, addressing essential aspects of video-language modeling, such as video sampling and encoding methods, token resampling and integration strategies, and data compositions (Sec. 4 & LABEL:sec:training). For instance, we discover that frames-per-second video sampling significantly outperforms standard uniform sampling used in previous works (oryx; kangaroo). We also find which vision encoder combinations are the most robust and that the Perceiver Resampler (perceiver) outperforms average pooling. When studying the numerous benchmarks available, we discovered a large portion of the performance improvements are driven primarily via language modeling and, therefore, curate ApolloBench, which significantly reduces evaluation time while improving assessment quality (Sec. 2).

We conduct a systematic exploration of the video modeling design space for Large Multimodal Models, uncovering critical factors that drive performance and providing actionable insights for future research.

We identify Scaling Consistency, where design decisions effective for smaller LMMs and datasets are transferred effectively to larger ones, reducing computational costs and enabling efficient experimentation.

We address evaluation inefficiencies by curating ApolloBench, a subset of existing benchmarks that cuts evaluation time by 41×41\times while offering detailed insights into temporal reasoning and perception tasks.

We introduce Apollo, a family of LMMs that achieves state-of-the-art results across video understanding multiple benchmarks. Notably, Apollo-33B surpasses nearly all 77B models, while Apollo-77B variant is state-of-the-art among models with less than 3030B parameters.

In Sec 2, we analyze the state of video benchmarks and introduce ApolloBench. In Sec. 3, we show how one can relax traditional scaling laws for computational savings. In Sec. 4, we explore the architecture design space. In Sec. LABEL:sec:training, we investigate different training protocols and data mixtures. Finally, in Sec. LABEL:sec:apollo, we present Apollo, a state-of-the-art family of video-LMMs.

How effective are existing video question-answering benchmarks?

The rapid advancement of video Large Multimodal Models (video-LMMs) has spurred the creation of numerous video question-answering benchmarks, including Video-MME, MLVU, LongVideoBench, and others (videomme; longvideobench; mlvu; perceptiontest; li2024videovista; lvbench; temporalbench). While this proliferation enables comprehensive evaluation, it also introduces significant resource intensiveness and redundancy. For example, evaluating a 3B-parameter model on these benchmarks requires 184 A100 GPU hours. In this section, we first analyze the quality of existing benchmarks (Sec. 2.1), their redundancy (Sec. 2.2), and introduce ApolloBench (Sec. 2.3) by building on these insights.

What drives video benchmark performance is not known. As shown by goyal2017making, some image question-answering benchmarks are largely driven by text comprehension rather than image perception. mmstar further showed that data leakage in either the LLM or LMM training stage may be further contaminating evaluation in image question-answering benchmarks. To evaluate the state of video question answering benchmarks, we evaluated ten open-source LMMs on several benchmarks: Video-MME (videomme), TempCompass (tempcompass), LongVideoBench (longvideobench), MLVU (mlvu), NExTQA (nextqa), and PerceptionTest (perceptiontest)—under three different settings:

Video: Models prompted with video input using their standard video sampling. Green in Fig. 2, left.

Image: Models are provided only the center frame of each video. Red in Fig. 2, left.

Text: Models are prompted with only the original question, without any visual input. Blue in Fig. 2, left.

As illustrated in Fig. 2, left, a significant portion of existing benchmarks are answered solely through text comprehension alone (blue boxplots) or only using the center frame (red boxplots), indicating that LMMs do not rely on video perception in a large portion of existing benchmarks. We sorted the benchmarks by the difference between the Video and Text performance (light blue). A benchmark relies more and more on its video perception capabilities when this bar is high. When examining Fig. 2, left, it is apparent that as videos get longer, the reliance on video perception decreases (compare Video-MME S/M/L). To evaluate how much of the benchmarks require video input to answer the question, we also plot the difference between the Video and Image performance (yellow). Some benchmarks can almost be entirely solved using a single frame. For example, in line with buch2022revisiting, we find that NExTQA is solved using a single frame. Perception-test also behaves similarly. Finally, when studying Fig. 2, left, a high variance in the box plot is desired as this indicates more highly discriminative benchmarks. Among all the existing benchmarks, Video-MME (Short), MLVU, and TempCompass emerge as the top performers.

2 Redundancy in existing benchmarks

To evaluate the redundancy in video question answering benchmarks, we evaluated ten open-source LMMs on several benchmarks: Video-MME (videomme), TempCompass(tempcompass), LongVideoBench (longvideobench), MLVU (mlvu), NExTQA (nextqa), and PerceptionTest (perceptiontest). We then calculated the correlation of each of the benchmarks to each other, the result of which can be seen in Fig. 2, right. Our analysis revealed significant redundancy among benchmarks, as evidenced by the block-diagonal correlation matrix, where we can identify groups of benchmarks that are highly correlated.

To evaluate the effect of different question types and video durations, we also evaluated the correlations between video duration groups. We find that the performance of models on short and long videos within Video-MME (videomme) exhibits an R2=0.94R^{2}=0.94, see App. Fig. LABEL:sup:fig:eval:videomme, while in LongVideoBench, R2>0.92R^{2}>0.92 between all duration groups. To assess the effect of question format, we studied the TempCompass (tempcompass) dataset, which has different question formats (multiple-choice, yes/no, caption matching, and caption generation), and found that they are also highly correlated (R2>0.8R^{2}>0.8), indicating that varying question types do not significantly diversify the evaluation (see App. Fig. LABEL:sup:fig:eval:tempcompass).

3 Introducing ApolloBench

Motivated by these insights, we set out to curate a more effective and efficient benchmark suite called ApolloBench. We focused on multiple-choice questions to eliminate the need for external tools like ChatGPT, ensuring a consistent and cost-effective evaluation process (freeva).

We filtered out questions that could be correctly answered by more than 50% of the models with either text or image inputs, removing questions that do not require video perception (see Fig. 2, left, ApolloBench). Subsequently, we identified five broad temporal perception categories: Temporal OCR, Egocentric, Spatial, Perception, and Reasoning. Questions were then manually categorized into each one of these categories. We selected the top 400400 questions from these categories that exhibited the most discrimination between models via entropy and manually verified each one to validate the correctness of the selected questions. Evaluating on ApolloBench is 41×41\times faster while being highly correlated with existing benchmarks (see Fig. 2, right) and more influenced by video perception (Fig. 2, left). For more details, see App. Sec. LABEL:app:sec:benchmark_analysis and App. Fig. LABEL:fig:benchmark_creation_flowchart.

Scaling Consistency: How small can you go during model design?

Developing Large Multimodal Models (LMMs) poses significant computational challenges, especially when training on extensive datasets with billion-parameter models. To make the research process more efficient, it is essential to determine whether smaller LMMs and datasets can reliably inform design decisions for larger ones. Traditional scaling laws require training multiple models of varying sizes for each design decision to derive how performance scales with model size. However, in the context of LMMs, which typically utilize multiple pre-trained components (e.g., vision encoders, language models), scaling each component individually is impractical due to the lack of availability of such components and the immense computational resources required. As such, we set to relax these scaling laws and instead reason about correlation or transfer of design decisions between models of different sizes.

This section investigates the correlation between design decisions made on LMMs of different sizes. Specifically, we selected 2121 model variations encompassing various design aspects such as architecture, video sampling methods, training strategies, and data mixtures. Each variation was trained using four different Large Language Models (LLMs): Qwen22-0.50.5B, Qwen22-1.51.5B, Qwen1.51.5-44B, and Qwen22-77B (bai2023qwen; yang2024qwen2), resulting in a total of 8484 models. We then analyzed the correlation (R2R^{2}) between the performance of these models (see App. Fig. LABEL:sup:fig:scaling_consistency). Our findings reveal that design decisions on models of a critical size (∼2−4\sim 2-4B) correlate highly (R2>0.9R^{2}>0.9) with those on larger models, a phenomenon we term Scaling Consistency (see Fig. 3). For instance, the R2R^{2} between the 44B and 77B models is 0.9380.938, indicating a strong predictive relationship. Please refer to the App. Sec. LABEL:app:sec:scaling for a detailed analysis.

Scaling laws typically require training models of various sizes to study performance trends. However, due to limited availability and high computational cost, scaling laws are rarely applied to LMMs. In contrast, Scaling Consistency demonstrates that design decisions made on moderately sized models (∼2−4\sim 2-4B) and datasets transfer reliably to larger models, even across different model families. This allows researchers to make informed design choices without extensive scaling studies. Our primary goal is to show that design decisions transfer reliably, reducing computational burden and accelerating research.

In Fig. 3, left, we plot the R2R^{2} values between models of various sizes and the 7B LLM model variant. The correlation with the 77B LLM increases approximately log-linearly with the size of the smaller LLMs and generalizes between model families. This behavior is not observed with smaller models, e.g., 0.50.5B, where R2R^{2} immediately drops below 0.80.8, and no log-linear behavior can be observed. This reinforces the existence of a critical model size (∼2−4\sim 2-4B) where design decisions transfer reliably—a phenomenon we term Scaling Consistency. Scaling Consistency seems to generalize between model families, as a mix of Qwen1.5 and Qwen2 models were utilized in this study. For example, while the Qwen2-1.5B and Qwen1.5-4B model variants had similar performance, the 4B Qwen1.5-4B was still more correlated than the 1.5B model. Please refer to the App. Sec. LABEL:app:sec:scaling for a comprehensive analysis.

Impact of dataset size.

We examined the impact of dataset size on model performance by training models using the same data mixture but varying the dataset size from 7575K to 11M samples. The results are shown in Fig. 3, right, where the correlation of the 0.5/1.5/40.5/1.5/4B models trained on varying datasets sized to 77B trained on the full dataset can be seen as a function of dataset size. Focusing on the 4B LLM variant, we observed that the correlation (R2R^{2}) with larger models plateaus around ∼500K\sim 500K samples, indicating that increasing the dataset size beyond this point yields diminishing returns in terms of informing design decisions. In contrast, smaller models (e.g., 0.5B and 1.5B) exhibited less consistent behavior, with their R2R^{2} values fluctuating more across different dataset sizes. This suggests that a dataset size of approximately 500K samples is sufficient for moderately sized models (2–4 billion parameters) to reliably transfer design insights to larger models.

Exploring the video-LMM design space: what influences effective model design?

In this section, we analyze key architectural design choices shaping the performance of Large Multimodal Models (LMMs) in video-language tasks. We focus on four critical aspects: (I) Video sampling (Sec. 4.1) where we compare uniform and fps video sampling and evaluate the effect tokens and frames per second have on downstream performance. (II) Video representation (Sec. 4.2) where we explore how image and video encoders impact video representation and show which encoder and encoder pairs lead to the best performance. (III) Video token resampling (Sec. 4.3) where we test different visual token resamplers. (IV) Video token integration (Sec. LABEL:sec:arch:integration) where we examine various strategies to integrate the visual token into the text tokens.

Using Scaling Consistency, we opted to perform the following exploration using Qwen2.52.5 33B (yang2024qwen2) and trained on a dataset of 750750K samples. As demonstrated in Sec. 3, these findings exhibit a strong correlation (R2>0.9R^{2}>0.9) with results on larger models and across different model families. Unless stated otherwise, a Perceiver Resampler (perceiver) was employed, with 1616 tokens per frame at a frame rate of 22 fps. The dual encoders used were InternVideo22 (internvideo2) and SigLIP-SO400400M (siglip). When training on images, images were duplicated before being encoded by the video encoders for fully integrated encoding, as we found it to be slightly more performant with fewer parameters and complexity (see App. Sec. LABEL:sec:training:unified_split). This is in line with video_llava.

Videos can be sampled in many ways, from uniform sampling - uniformly sampling NN frames from the video (video_chat; video_llava; chat_univi; zhang2024llavanextvideo), to fps sampling - sampling at a set number of frames per second. While many recent methods have preferred fps sampling (oryx; llava), they default to uniform sampling when video durations exceed their frame sampling capacity (usually ∼64\sim 64). The maximum frame capacity is typically constrained due to the memory requirement at the vision encoder and or the LLMs context window.

Uniform frame sampling enables simplified training because the effective ‘vision batch size’ (i.e., the number of frames that need to be encoded) remains constant. However, training video-LMMs with uniform frame sampling means that the time difference between concurrent frames changes with each video, effectively setting a different ‘video speed’ in every iteration. As a result, when uniformly sampling NN frames from videos of varying lengths, the effective playback speed represented in the sampled frames changes. For a shorter video, NN uniformly sampled frames represent a slower playback (more frames per second of actual content), while for a longer video, those same NN frames represent a faster playback. This will likely hamper the LMMs’ capability to learn about the speed of objects in videos. Meanwhile, methods employing fps sampling must either limit the maximum video duration or the maximum number of frames (above which, they default to uniform frame sampling) in training or suffer from similar issues as in uniform sampling. An alternate approach is to sample ‘video clips’ of NN frames at a set fps (or duration) and, when reaching the maximum token count, space these out instead. Here, rather than uniformly spacing out the sampled video frames, the NN frames encoded by the video encoder maintain the same effective fps, and only frames of concurrent ‘clips’ are spaced out. Methods that utilize video encoders (video_chat; video_llava), where multiple frames are encoded together, should use such frame sampling as video encoders are typically trained at a constant fps (vjepa; internvideo2; languagebind; videomae).

To evaluate the effect of fps vs. uniform sampling, we trained four models that, while training, we uniformly sampled 8,16,328,16,32, or 6464 frames. To test whether performance differences are due to the different frame sampling at test or train time, we evaluated these models with uniform and fps sampling. The results of this experiment can be seen in Fig. 4, left and middle. We found that uniform frame sampling consistently underperformed compared to fps sampling, Fig. 4, left. As can be seen, this performance gap is not due to the different number of frames sampled at test time, Fig. 4, middle. Therefore, we conclude that the uniform frame sampling of videos causes this performance gap during training.

Finding 3:

2 Video representation

Training effective video encoders is challenging due to the high memory requirements for processing large video datasets and the comparatively low quality of available supervision. While early approaches predominantly used dedicated video encoders (video_chat; mvbench; video_llava), recent developments favor image encoders instead (kangaroo; llava; oryx). This shift arises because image encoders, although lacking temporal integration, still produce higher-quality representations that the LLM can readily leverage. Another possibility is that in this approach, image and video datasets can be fully integrated, possibly benefiting from image-video transfer and allowing the utilization of the much larger, more diverse, and more efficient image instruction tuning datasets (llava; longva).

Multiple studies have conducted extensive investigations into visual representation within image-LMMs (nveagle; cambrian1). idefics2 found that SigLIP outperformed even much larger encoders, such as EVA-CLIP-5B. griffon showed that input image resolution influences performance more than token count, which may have influenced idefics2’s ablation. visualtokenizer compared encoders trained with supervision and found where each is preferable. However, whether image or video encoders are preferable for video-LMMs and what influences their performance is unclear. As such, we set out to find what drives good video representation in LMMs. We trained LMMs with several image and video encoders and their combinations and evaluated how this design decision impacted the final model performance. Our study includes diverse language- and self-supervised video/image encoders:

InternVideo2 (internvideo2): trained in two stages: (1) unmasked video token reconstruction, (2) crossmodal contrastive learning aligning video with audio, speech, and text. Encodes four frames.

LanguageBind-Video v1.5 (languagebind): initialized with an OpenCLIP model, trained contrastively with a frozen text encoder. Encodes eight frames.

VideoMAE (videomae): trained through self-supervised learning by masking video patches with a reconstruction loss. Encodes sixteen frames.

V-JEPA (vjepa): trained through self-supervised learning by predicting masked spatio-temporal regions in a learned latent representation space. Encodes sixteen frames.

SigLIP-SO400M (siglip): a shape-optimized model trained using a sigmoid loss function for language-image pre-training.

LanguageBind-Image (languagebind): one of the OpenCLIP image encoders and is not further tuned.

DINOv2 (dinov2): trained using a self-supervised teacher-student framework where the teacher guides the student to produce consistent representations across different image views.

As seen in Fig. 5, left, language-supervised encoders consistently outperform self-supervised encoders, in line with observations in prior work (nveagle). In the single-encoder setups, SigLIP-SO400400M had the best performance compared to all image/video encoders, demonstrating that video encoders must be improved to replace image encoders. Video encoders outperform image encoders only on Temporal Perception, indicating that LLMs struggle with fine-grained temporal integration (e.g., estimating speed and direction of movement).

Finding 5:

3 Video token resampling