VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
Shijie Zhou, Alexander Vilesov, Xuehai He, Ziyu Wan, Shuwang Zhang, Aditya Nagachandra, Di Chang, Dongdong Chen, Xin Eric Wang, Achuta Kadambi
Introduction
Humans possess an innate ability to perceive, track, and interpret motion as well as spatial and temporal changes, enabling rich interpretations of complex dynamic events from both egocentric and allocentric perspectives . When observing an object move, we can inherently process any changes such as lateral shifts, rotational directions, and periodic or repeated actions unfolding along a specific trajectory . These sophisticated perceptual abilities are a product of our spatiotemporal cognition , and form an essential foundation that allows us to comprehend and reason about physical phenomena, object interactions, and causal relationships within our environment .
Vision language models (VLMs), which also have the potential to perceive motion and spatiotemporal changes in videos, represent a prominent class of methods aimed at emulating or surpassing human capabilities in integrated visual and linguistic reasoning . While previous work on VLMs has primarily focused on static visual understanding—through large-scale training on paired language and image data —or explored video understanding tasks such as captioning and scene understanding , we find that their strong performance in these tasks does not naturally translate into robust spatiotemporal reasoning. This limitation is particularly striking given that state-of-the-art VLMs are typically trained on datasets comprising hundreds of billions of tokens . In contrast, human infants naturally develop robust spatiotemporal cognition within the first few months of life .
Another key challenge that limits VLM performance on spatiotemporal tasks is the need to implicitly or explicitly reconstruct a four-dimensional (4D) representation—3D space + time—of dynamic scenes, and subsequently reason over this reconstruction . As illustrated in Fig. 1, the car is advancing forwards and turning to the left in its own frame of reference. However, from the camera’s perspective, its motion appears as a combination of heading to the right and receding into the distance despite the car being in the center of the frame due to camera rotation. Human observers can seamlessly disentangle these complex dynamics, accurately interpreting trajectories by synthesizing diverse visual cues including camera rotation compensation, stationary scene landmarks, prior knowledge of 3D and 4D environmental structures, and perspective projections . The inability of current VLMs to integrate spatial, temporal, and semantic cues in a human-like manner stems from their fundamentally different encoding paradigms for visual and language information. This discrepancy highlights a significant gap between human and machine spatiotemporal understanding, suggesting that future VLMs may benefit from insights in cognitive science and neuroscience to develop more advanced mechanisms for perceiving, integrating, reconstructing, and reasoning over dynamic scenes.
In order to effectively characterize and challenge the existing spatiotemporal reasoning abilities of VLMs, we introduce VLM4D, a rigorous benchmark specifically designed to probe the spatiotemporal grounding capabilities of current vision language models. Through this contribution, we aim to catalyze research that addresses the critical gap in spatiotemporal understanding and reasoning within VLMs and provide a foundational analysis highlighting key deficiencies in existing models.
We summarize our main contributions as follows:
We propose the first benchmark VLM4D explicitly designed to rigorously evaluate the spatiotemporal (4D) reasoning capabilities of Vision Language Models (VLMs).
We introduce a novel, meticulously curated dataset consisting of diverse real-world and synthetic video sequences paired with carefully crafted spatiotemporal question-answer (QA) annotations.
We analyze the critical limitations of contemporary VLMs in spatiotemporal reasoning and conduct experiments exploring potential solutions, highlighting fundamental challenges and outlining clear directions for impactful future research.
Related Work
Early methods for video understanding before the advent of large vision language models (VLMs) leveraged trajectory-based representations and separate modeling of spatial and motion cues for tasks like action and motion recognition . Recently, VLMs have evolved rapidly by fully leveraging the significant achievements of Large Language Models (LLMs) and large-scale visual instruction tuning datasets . While VLMs exhibit transformative potential for applications such as embodied AI , robotics , scene generation , and world modeling , most existing methods remain constrained to static images, focusing narrowly on spatial understanding while overlooking the dynamic temporal dimension inherent in real-world interactions. To bridge this gap, emerging research has begun exploring video modality integration, aiming to equip VLMs with spatial-temporal awareness critical for tasks like video comprehension, where both contextual details and motion dynamics are essential. For example, VideoLLM-MoD proposes to address the efficiency issue when processing long-term video by mixture-of-depths. introduces VideoRefer to enhance the finer-level (like object-level) spatial-temporal video understanding of VLMs. Grounded-VideoLLM also targets for fine-grained video understanding through incorporating an additional temporal stream. In this work, we aim to rigorously evaluate the 4D spatial-temporal reasoning capabilities of state-of-the-art VLMs, probing how and to what extent these models internalize spatial intelligence and temporal dependencies.
VLM Benchmarks
Following the development trends of VLMs, benchmarking VLMs shares the similar trajectory by first evaluating vision QA on static images , to align with models’ early focus on 2D understanding. As VLMs evolved to tackle dynamic scenarios, benchmarks expanded to evaluate general-purpose video comprehension tasks that probe temporal coherence and event understanding . Notably, MMVU further proposes a knowledge-intensive benchmark to assess the expert-level reasoning ability of current video-based large models. However, while these works assess perception and semantic understanding, they largely overlook the explicit evaluation of spatial-temporal awareness, a core capability for real-world applications requiring 4D (3D space + time) reasoning. Recent efforts like pioneer benchmarks for 3D visual-spatial intelligence but restrict evaluation to static 3D scene, neglecting the interplay of object motion and temporal dynamics intrinsic to videos. In this work, we introduce VLM4D, the first benchmark designed to holistically evaluate the 4D intelligence in VLMs, unifying spatial understanding, temporal continuity, and motion reasoning. By curating tasks that demand precise analysis of dynamic interactions (e.g., direction prediction, perspective anticipation, and motion reasoning), VLM4D exposes critical gaps in current models’ ability to internalize spatiotemporal relationships. Our work not only advances the granularity of VLM evaluation but also shares insights and potential solutions to improve the model performance.
The VLM4D Benchmark
We introduce VLM4D, the first benchmark specifically designed to test the spatiotemporal reasoning ability of VLMs. VLM4D consists of 1000 videos paired with over 1800 question-answer pairs, each carefully designed to assess both spatial and temporal understanding jointly. The majority of these videos are sourced from datasets with rich spatiotemporal characteristics, thus ensuring a diverse range of motion-related scenarios. We also augment the dataset with synthetic videos generated by a world-foundation model, Cosmos , that has been modified using techniques introduced in to obtain more accurate correspondence between motion-oriented prompts and the resulting generated video. Fig. 2 illustrates the composition of our dataset.
Unlike prior work that often relies heavily on LLMs and VLMs to generate first iterations of benchmarks and datasets followed by human quality control, we found that existing VLMs and automated methods showed significant limitations in terms of reliability and quality. This shortcoming necessitated direct human annotations that were then followed by augmentation by LLMs to ensure a high-quality benchmark. An overview of the benchmark curation pipeline is shown in Fig. 3.
Real-world videos were sourced from datasets with rich spatiotemporal characteristics that ensured diverse motion and perspective variations. For egocentric data, we relied mainly on the Ego4D dataset , while most exocentric data points were collected from the DAVIS and YouTube-VOS datasets. To minimize confounders and to focus attention of VLM abilities to only spatiotemporal reasoning, we preprocessed the videos by temporally segmenting and centering them around the most relevant action, thus resulting in videos with an average duration of - seconds. This ensures that the key event described in the question is clear and reduces ambiguities or confounders that would reduce VLM accuracy.
Synthetic Video Data Generation
The rapid advancement of video generation techniques has led to their widespread application in diverse domains like robotics and immersive entertainment , making 4D reasoning on synthetic content essential for VLMs and underscoring the need to include such videos in our benchmark. For synthetic video generation, we use Cosmos as our video generation backbone. To ensure that the generated videos align with the intended object moving directions, we incorporate input bounding boxes as additional spatial guidance. Specifically, we follow the approach introduced in modifying the diffusion forward steps to enforce object localization constraints at each timestep, ensuring consistency between the generated object direction and the user-specified trajectory. The average duration of generated synthetic videos is seconds. To maintain high-quality outputs, we perform a manual verification step after generation, filtering out low-quality videos and retaining only those that accurately match the specified directions. Once a video is generated, we use an LLM (GPT-4o) to create two types of evaluation questions: Directional questions, derived from the textual prompt used to generate the video; and False Positive (FP) questions (constituting 10%, matching the ratio in the real dataset), which query non-existent objects in the scene. Both question types follow the format: “What direction is the Object Name moving?”, where the model must select one of four possible answers: “left”, “right”, “not moving”, or “no Object Name there”. A final manual review is conducted to filter out or revise ambiguous questions, ensuring the quality of the questions and ground-truth answers.
QA Generation and Quality Control
Question-answer pairs are primarily constructed through human annotations. The question answer pairs are then supplemented with alternative answers by an LLM (GPT-4o) for multiple choice (MC) questions. To ensure high-quality annotations, we applied a three-round, multi-person cross-checked verification process, during which ambiguous videos were filtered out and vague, misleading, or incorrect QA pairs were refined to improve spatial and temporal alignment between language and visual content. Fig. 4 showcases some qualitative examples of annotations for different types of videos.
Assessing Human Performance
To establish a human performance baseline on our benchmark, we conducted an evaluation in which participants independently answered 100 randomly sampled questions from the dataset. The accuracy of human responses was then aggregated to approximate the performance of human spatiotemporal reasoning on the dataset.
2 Categorizing Spatiotemporal Performance
To systematically evaluate spatiotemporal reasoning capabilities, we first categorize videos into two primary groups: egocentric (first-person) videos and exocentric (third-person) videos. Egocentric videos are sourced from the Ego4D dataset, where scenes are captured from a head-mounted camera, thus offering dynamic video data that is inherently coupled with the individual’s actions. Exocentric videos encompass a diverse range of recorded scenes, from sports footage to everyday scenes. Beyond this categorization, we also evaluate spatiotemporal performance across four dimensions: translational movement (TM), rotational movement (RM), spatiotemporal counting (STM), and false positives (FP), with their proportions shown in Fig. 2. Translational movement assesses a model’s ability to track linear motion within scenes, while rotational movement assesses the understanding of changes in orientation and perspective shifts over time. Spatiotemporal counting extends these core motion-based tasks by requiring a more complex reasoning strategy to determine the number of objects performing a translation or rotational movement. Lastly, the false positives category evaluates the model’s critical thinking in determining whether an object or event actually occurred within the spatiotemporal context. By structuring the benchmark along these axes, we aim for a comprehensive framework for assessing spatiotemporal reasoning (Fig. 5).
black Organization Model Release Real Synthetic Overall Ego-centric Exo-centric Average Directional FP Average User Study Human Performance 99.6 99.7 99.7 95.8 100 96.2 98.8 Random Random Selection 24.4 23.2 23.6 25.5 24.7 25.4 24.1 Latest Proprietary VLMs OpenAI GPT-4o 2024-11 55.5 62.2 60.0 49.5 53.3 49.9 57.5 lightgray Google Gemini-2.5-Pro 2025-6 64.6 62.9 63.5 54.8 80.0 57.3 62.0 Anthropic Claude-Sonnet-4 2025-5 52.6 52.1 52.2 44.0 86.7 48.3 51.3 xAI Grok-2-Vision 2024-12 48.8 49.7 49.4 49.3 66.7 51.0 49.8 black Open-source Image VLMs Meta Llama-4-Maverick-17B 2025-4 52.6 54.3 53.8 53.3 51.1 53.0 53.6 Llama-4-Scout-17B 2025-4 48.6 56.2 53.7 53.3 75.6 55.5 54.1 lightgray Microsoft Phi-4-Multimodal 2025-3 41.0 35.4 37.2 37.5 11.1 34.8 36.6 Phi-3.5-Vision 2024-7 33.4 38.8 37.1 23.3 37.8 24.7 34.0 lightgray DeepSeek DeepSeek-VL2 2024-12 33.6 32.9 33.1 31.8 46.7 33.3 33.2 lightgray Shanghai AI Lab InternVL2.5-38B 2024-11 46.6 50.1 48.9 43.3 57.8 44.7 47.9 InternVL2.5-8B 2024-11 39.0 44.0 42.4 40.8 42.2 40.9 42.0 lightgray Mistral AI Pixtral-12B 2024-9 32.3 25.8 27.9 24.3 22.2 24.0 27.0 lightgray Rhymes Aria 2024-11 47.2 44.0 45.1 38.5 71.1 41.8 44.3 black Open-source Video VLMs Alibaba Qwen2.5-VL-7B 2025-1 42.3 43.7 43.3 43.5 64.4 45.6 43.8 Qwen2.5-VL-72B 2025-1 54.3 52.5 53.1 49.5 80.0 52.6 53.0 Qwen2-VL-7B 2024-8 36.1 34.7 35.2 40.5 35.6 40.0 36.3 Qwen2-VL-72B 2024-9 48.1 43.0 44.6 40.8 73.3 44.0 44.5 DAMO VideoLLama3-2B 2025-1 53.2 42.5 46.0 34.3 55.6 36.4 43.7 VideoLLama3-7B 2025-1 49.4 45.1 46.5 42.8 53.3 43.8 45.9 Shanghai AI Lab InternVideo2.5-8B 2025-1 57.2 50.5 52.7 44.3 46.7 44.5 50.7 InternVideo2-8B 2024-8 35.6 39.3 38.1 43.0 0.0 38.7 38.2 LLaVA LLaVA-One-Vision-7B 2024-9 36.8 35.6 36.0 37.8 35.6 37.5 36.3 LLaVA-NeXT-Video-34B 2024-6 29.6 31.6 30.9 24.5 55.6 27.6 30.1 black
We evaluate 23 most recently released VLMs thus covering a wide range of model sizes, architectures, and training methodologies. For closed-source VLMs, we evaluate GPT-4o , Gemini 2.5 Pro , Claude Sonnet 4 , and Grok-2-Vision . For open-source models, we include Llama 4 , DeepSeek-VL , Qwen2.5-VL , Qwen2-VL , InternVL2.5 , Aria , InternVideo 2.5 , InternVideo2 , Phi-4-multimodal , Phi-3.5-vision , Pixtral , VideoLLama3 , Llava-One-Vision , Llava-NeXT-Video . When available, we evaluate different sizes for each model, resulting in models ranging from 2 to 72 billion parameters.
Evaluation Settings
The evaluations were performed in a zero-shot setting with video or a set of sampled frames of video, followed by the prompt forming the input. For each model, we evaluate on two different inference settings. In the first setting, the model prompted to directly output (DO) the answer immediately without any reasoning, and in the second evaluation setting, the model is directed to create intermediate reasoning steps, chain-of-thought (CoT) , before inferring the final answer.
Metrics
Following prior work and given the nature of our target task, we adopt multiple-choice questions (MCQs) for evaluation, using accuracy as the primary metric. For the two inference settings described earlier, we employ LLM-as-Judge to assess the outputs of VLMs. We opt for this method instead of string or template matching, as VLMs—especially under chain of thought (CoT) prompting—often generate all possible answer choices during reasoning, with varying frequencies and slight formatting differences. In some cases, the final answer may even contradict the reasoning. To better evaluate whether the model truly understands the video, we prompt two advanced LLMs (GPT-o3 and o4-mini) to grade based on the full CoT reasoning, not just the final answer. We then perform a cross-check between their judgments and manually resolve any disagreements. The evaluation results are reported in Sec. 3.2.
2 Benchmark Results
Our comprehensive evaluation on the VLM4D benchmark, detailed in Sec. 3.2, systematically assesses the spatiotemporal reasoning capabilities of modern Vision-Language Models (VLMs). The results highlight a clear hierarchy, with proprietary models demonstrating superior performance over their open-source counterparts. Google’s Gemini-2.5-Pro emerges as the top-performing model with an overall accuracy of 62.0%, followed by OpenAI’s GPT-4o at 57.5%. Within the open-source domain, image-based models like Meta’s Llama-4-Scout-17B (54.1%) and video-supported models like Alibaba’s Qwen2.5-VL-72B (53.0%) demonstrate highly competitive results, even outperforming some proprietary counterparts. Despite these achievements, a significant performance gap persists when compared to human accuracy (98.8%), underscoring that sophisticated 4D awareness remains a formidable challenge for AI. Notably, performance varies significantly across different categories, such as real versus synthetic data and ego-centric versus exo-centric perspectives, indicating that current models lack generalized spatiotemporal understanding.
Despite significant advances in VLMs, their abilities to understand and reason about motion, spatial relationships, and temporal coherence remains fundamentally underdeveloped . Chain of thought (CoT) is widely employed as a method to improve accuracy through step-by-step reasoning. We showcase a comparison between CoT and DO in Fig. 6. Overall, there is no indication of a large advantage of CoT over all evaluated models. Upon deeper exploration of the CoT reasoning of some models, we observe that the reasoning process was primarily flawed in the following ways: irrelevant information and arriving at conclusions that are inconsistent with the reasoning process. Larger models exhibited strategies that would be similar to how a human processes spatiotemporal information, but the resulting execution falls short of human performance. This demonstrates a disconnect between its visual and linguistic knowledge. We provide examples of this behavior in the supplementary material.
2 Deficiencies in Spatiotemporal Labeling
Another avenue of exploration we undertook is to understand the richness of spatiotemporal labels in popular supervised fine-tuning (SFT) VLM datasets. Typically, video captioning occurs at the ‘scene’ level, lacking fine-grained temporal, spatial, and object-level details. We performed an extensive analysis, encompassing over 2 million samples . We performed this analysis through string-matching of spatiotemporal descriptors related to directionality, translational motion, rotation, and perspective shifts and provide the overall results in Fig. 7. We then performed a manual finegrained evaluation of the ShareGPT4Video dataset which we found had the highest density of spatiotemporal dataset. We found that from a sample of 100 labels that were detected as spatiotemporal, less than 10% of them were judged as accurate upon human evaluation. This result underscores the inadequacy of current dense captioning approaches, which frequently generate spatiotemporal descriptors without capturing precise motion dynamics. We provide more detailed analysis and explanations in the supplementary material.
To probe promising future solutions for enhancing spatiotemporal video understanding, we propose two approaches that address some of the shortcomings of current state-of-the-art VLMs: fine-tuning a VLM on data-rich in spatiotemporal actions and the other leveraging 4D reconstruction and feature fields jointly with a VLM. SFT refines the model’s abilities by training on datasets that contain temporally and spatially rich actions and interactions. By integrating structured visual representations and targeted fine-tuning, these approaches enhance video-language models’ ability to interpret motion. The second method lifts the feature space of VLMs into a temporally coherent 4D feature field, providing structured scene representations that improve motion and spatial reasoning in the stage of decoding and inference.
We evaluate on a subset split of the real dataset by randomly splitting the real-world dataset into a training and testing split (80% / 20%) and we try settings using synthetic/real/both for training. We conducted the experiments using Qwen 2VL (7B) and Qwen 2.5VL (7B) through LLama-Factory , and compared the performance before and after supervised fine-tuning in Tab. 2. The results demonstrated an improvement in accuracy in spatiotemporal reasoning, suggesting that performance gains can be obtained through targeted training. However, the addition of synthetic data does not necessarily increase performance over using real data alone, suggesting the importance of synthetic data quality.
D Feature Fields Reconstruction
Recent advances in 3D/4D feature fields reconstruction methods have significantly enhanced the vision foundation model’s performance in 3D/4D space by integrating structured latent scene representations into the model’s inference stage. Inspired by the promising results of feature lifting, we explore enhancing the InternVideo2-8B model with spatiotemporal awareness by adopting the strategy proposed in Feature4X , which constructs the VLM’s 2D feature space along the time dimension into a 4D feature field. To assess this approach, we evaluate performance on a subset of the VLM4D benchmark, specifically leveraging all 50 videos from the DAVIS 2016 dataset . Our experimental evaluation compares inference performance across three distinct input modalities: original 2D videos; rendered novel global-view videos (which provide broader contextual information in 2D format); and reconstructed global feature fields (which implicitly incorporate 4D scene-level information during reasoning). Table 3 reveals that reconstructed feature fields achieve the highest accuracy across both reasoning types. This success stems from two key advantages: the inherent structure of 4D representations and the ability of feature field inference to avoid the rendering artifacts of RGB reconstruction (global view video). However, the current approach requires per-scene optimization as a post-processing step, limiting its generalizability and making it computationally intensive.
Through the construction of the VLM4D benchmark, we evaluate the spatiotemporal reasoning capabilities of various vision language models (both open-source and proprietary). While more recently released models demonstrate improved performance over their counterparts, they remain significantly behind human proficiency. Overall, our work questions whether VLMs possess spatiotemporal reasoning abilities that are imperative to have for more sophisticated visual agents in fields ranging from robotics to interactive AI systems that require a deep understanding of dynamic visual environments. We hope to inspire future work to explore novel approaches for integrating spatiotemporal grounding, thereby enhancing their spatiotemporal reasoning capabilities and facilitating robust deployment.
A VLM4D Benchmark Statistics
Tab. A presents the breakdown of our VLM4D benchmark dataset. Additionally, Fig. A visualizes the detailed performance of VLMs across different question categories. For models that support only image input, we convert videos into multi-frame image sequences, using the maximum number of frames allowed within the model’s context window. For models that support video input, we follow their default frame rate settings, typically 1 fps.
We begin by analyzing four individual collections of datasets, which together contribute a substantial body of data for our experiments. The datasets used in this study are:
As shown in Tab. C, these datasets contain over 2 million samples in total, a robust foundation for evaluating and benchmarking the spatiotemporal validity of the highly used video instruction tuning datasets.
We show our target strings in Tab. E. In our comprehensive analysis of the ShareGPT-4o dataset (Fig. E), we observed that over 40% of the captions incorporate at least one target category. As depicted in Fig. G, our target string search highlights a significant emphasis on the directional descriptors “left” and “right.” Furthermore, the analysis of negative samples, illustrated in Fig. F, indicates that these directional terms are seldom employed in conjunction with rotational or translational actions. This observation is further substantiated by the minimal overlaps between directional descriptors and action-related terms, as shown in Fig. G. Collectively, this shows a notable gap in the dataset’s ability to capture complex spatiotemporal relationships, particularly those involving dynamic textures and nuanced motion patterns. The statistics are summarized in Tab. D.
Please refer to Fig. I - M for detailed responses from all evaluated VLMs under both chain-of-thought (CoT) and direct output (DO) prompting, based on the given example video and question.
We utilize the Feature4X framework (Fig. H) for 4D reconstruction experiments conducted on our dataset. Given an input monocular RGB video, Feature4X reconstructs the dynamic 3D scene by employing dynamic 3D Gaussians, specifically Dynamic 3D Gaussian Splatting, which represent dynamic foreground elements that deform over time. These dynamic Gaussians are guided by a 4D Motion Scaffold, a sparse graph of trajectory nodes, enabling the interpolation of dense motion trajectories and features for each Gaussian efficiently. A separate set of static 3D Gaussians represents static background elements.
Feature4X introduces a unified latent feature embedding, distilled from various foundational 2D models, which facilitates multiple downstream tasks such as segmentation, scene editing, and visual question answering (VQA). Specifically, Feature4X extracts video segment features from the InternVideo2-Chat model, a foundation model fine-tuned for video question answering.
This unified feature field is directly used by the InternVideo decoder for inference, bypassing the video encoding step entirely. This approach significantly improves inference efficiency and retains comprehensive structural information from the 4D scene representation, which surpasses the context available from the original 2D videos alone. Consequently, this method enhances downstream tasks by providing richer spatiotemporal context and semantic consistency.