Cosmos World Foundation Model Platform for Physical AI

NVIDIA, :, Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chattopadhyay, Yongxin Chen, Yin Cui, Yifan Ding, Daniel Dworakowski, Jiaojiao Fan, Michele Fenzi, Francesco Ferroni, Sanja Fidler, Dieter Fox, Songwei Ge, Yunhao Ge, Jinwei Gu, Siddharth Gururani, Ethan He, Jiahui Huang, Jacob Huffman, Pooya Jannaty, Jingyi Jin, Seung Wook Kim, Gergely Klár, Grace Lam, Shiyi Lan, Laura Leal-Taixe, Anqi Li, Zhaoshuo Li, Chen-Hsuan Lin, Tsung-Yi Lin, Huan Ling, Ming-Yu Liu, Xian Liu, Alice Luo, Qianli Ma, Hanzi Mao, Kaichun Mo, Arsalan Mousavian, Seungjun Nah, Sriharsha Niverty, David Page, Despoina Paschalidou, Zeeshan Patel, Lindsey Pavao, Morteza Ramezanali, Fitsum Reda, Xiaowei Ren, Vasanth Rao Naik Sabavat, Ed Schmerling, Stella Shi, Bartosz Stefaniak, Shitao Tang, Lyne Tchapmi, Przemek Tredak, Wei-Cheng Tseng, Jibin Varghese, Hao Wang, Haoxiang Wang, Heng Wang, Ting-Chun Wang, Fangyin Wei, Xinyue Wei, Jay Zhangjie Wu, Jiashu Xu, Wei Yang, Lin Yen-Chen, Xiaohui Zeng, Yu Zeng, Jing Zhang, Qinsheng Zhang, Yuxuan Zhang, Qingqing Zhao, Artur Zolkowski

Introduction

Physical AI is an AI system equipped with sensors and actuators: the sensors allow it to observe the world, and the actuators allow it to interact with and modify the world. It holds the promise of freeing human workers from physical tasks that are dangerous, laborious, or tedious. While several fields of AI have advanced significantly thanks to data and compute scaling in the recent decade, Physical AI only inches forward. This is largely because scaling training data for Physical AI is much more challenging, as the desired data must contain sequences of interleaved observations and actions. These actions perturb the physical world and may cause severe damage to the system and the world. This is especially true when the AI is still in its infancy when exploratory actions are essential. A World Foundation Model (WFM), a digital twin of the physical world that a Physical AI can safely interact with, has been a long-sought remedy to the data scaling problem.

In this paper, we introduce the Cosmos World Foundation Model (WFM) Platform for building Physical AI. We are mainly concerned with the visual world foundation model, where the observations are presented as videos, and the perturbations can exist in various forms. As illustrated in Fig.˜2, we present a pre-training-and-then-post-training paradigm, where we divide WFMs into pre-trained and post-trained WFMs. To build a pre-trained WFM, we leverage a large-scale video training dataset to expose the model to a diverse set of visual experiences so it can become a generalist. To build a post-trained WFM, we fine-tune the pre-trained WFM to arrive at a specialized WFM using a dataset collected from a particular Physical AI environment for the targeted, specialized Physical AI setup. Fig.˜1 shows example results from our pre-trained and post-trained WFMs.

Data determines the ceiling of an AI model. To build a high-ceiling pre-trained WFM, we develop a video data curation pipeline. We use it to locate portions of videos with rich dynamics and high visual quality that facilitate learning of physics encoded in visual content. We use the pipeline to extract about 100M clips of videos ranging from 2 to 60 seconds from a 20M hour-long video collection. For each clip, we use a visual language model (VLM) to provide a video caption per 256 frames. Video processing is computationally intensive. We leverage hardware implementations of the H.264 video encoder and decoder available in modern GPUs for decoding and transcoding. Our video data curation pipeline leverages many pre-trained image/video understanding models. These models have different throughputs. To maximize the overall throughput for generating trainable video data, we build a Ray-based orchestration pipeline (Moritz et al., 2017). The details are described in Sec.˜3.

We explore two scalable approaches for building pre-trained WFMs discussed in Sec.˜5. These approaches are transformer-based diffusion models and transformer-based autoregressive models. A diffusion model generates videos by gradually removing noise from a Gaussian noise video. An autoregressive model generates videos piece by piece, conditioned on the past generations following a preset order. Both approaches decompose a difficult video generation problem into easier sub-problems, making it more tractable. We leverage state-of-the-art transformer architectures for their scalability. In Sec.˜5.1, we present a transformer-based diffusion model design that exhibits strong world-generation capabilities. In Sec.˜5.2, we present a transformer-based autoregressive model design for world generation.

Both the transformer-based diffusion model and transformer-based autoregressive model use tokens as representations of videos, where the former uses continuous tokens in the form of vectors, and the latter uses discrete tokens in the form of integers. We note that tokenization for videos—a process that transforms videos into a set of tokens—is highly nontrivial. Video contains rich information about the visual world. However, to facilitate learning of the WFMs, we need to compress videos into sequences of compact tokens while maximally preserving the original contents in the videos as the computation complexity of world foundation model training grows with the token counts. In many ways, building a video tokenizer is similar to building a video codec. We develop an attention-based encoder-decoder architecture to learn video tokenization for both continuous and discrete tokens described in Sec.˜4.

We fine-tune the pre-trained WFMs to arrive at post-trained WFMs for various Physical AI tasks in Sec.˜6. In Sec.˜6.1, we fine-tune our pre-trained diffusion WFM to make it camera pose conditional. This post-training creates a navigable virtual world where users can explore the created world by moving the virtual viewpoint around. In Sec.˜6.2, we fine-tune our WFMs on various robotic tasks, which consist of video-action sequences. We show that by leveraging the pre-trained WFMs, we can better predict the future state of the world based on the action taken by the robot. In Sec.˜6.3, we demonstrate how the pre-trained WFMs can be fine-tuned for various autonomous driving-related tasks.

Our intended use of the developed WFMs is for Physical AI builders. To better protect the developers when using the world foundation models, we develop a powerful guardrail system that consists of a pre-Guard to block harmful inputs and a post-Guard to block harmful outputs. The details are described in Sec.˜7.

We aim to build a world foundation model platform to help Physical AI builders advance their systems. To achieve this goal, we make our pre-trained world foundation models and tokenizers available under the NVIDIA Open Model License at NVIDIA Cosmos and NVIDIA Cosmos Tokenizer respectively. The pre-training script and post-training script will be available at NVIDIA Nemo Framework together with the video data curation pipeline to help builders craft their fine-tuning datasets. While this paper makes several improvements in world foundation model design, the world foundation model problem is still far from being solved. Additional research is required to advance the state-of-the-art further.

World Foundation Model Platform

Let x0:tx_{0:t} be a sequence of visual observations of the real world from time to tt. Let ctc_{t} be the perturbation to the world. As illustrated in Fig.˜3, a WFM is a model W\mathcal{W} that predicts the future observation at time t+1t+1, x^t+1\hat{x}_{t+1}, based on the past observation x0:tx_{0:t} and the current perturbation ctc_{t}. In our case, x0:tx_{0:t} is an RGB video, while ctc_{t} is a perturbation that can take many forms. It can be an action taken by the Physical AI, a random perturbation, a text description of the perturbation, etc.

We believe a WFM is useful to Physical AI builders in many ways, including (but not limited to)

Policy evaluation. This refers to evaluating the quality of a policy model in a Physical AI system. Instead of evaluating a trained policy by deploying it to a Physical AI system operating in the real world, one could instead let the digital copy of the Physical AI system interact with the world foundation model. The WFM-based evaluation is more cost-effective and time-efficient. With the WFM, builders can deploy the policy model in unseen environments that are otherwise unavailable. WFMs can help developers rule out incapable policies quickly and focus the physical resources on a few promising ones.

Policy initialization. A policy model generates actions to be taken by the Physical AI system based on the current observations and the given task. A well-trained WFM, which models the dynamic patterns of the world based on the input perturbations, can serve as a good initialization of the policy model. This helps address the data scarcity problem in Physical AI.

Policy training. A WFM paired with a reward model can be a proxy for the physical world to provide feedback to the policy model in a reinforcement learning setup. The agent can gain proficiency in solving tasks by interacting with the WFM.

Planning or model-predictive control. A WFM can be used to simulate different future states following different action sequences taken by a Physical AI system. A cost/reward module can then be used to quantify the performance of these different action sequences based on the outcomes. The Physical AI can then execute the best action sequence based on the simulation results as a whole, as in planning algorithms or in a receding horizon manner, as in model-predictive control. The accuracy of the world model upper-bounds the performance of these decision-making strategies.

Synthetic data generation. A WFM can be used to generate synthetic data for training. It can also be fine-tuned to be conditioned on rendering metadata such as depth or semantic maps. One can use the conditional WFM for the Sim2Real use case.

While we list the possibilities, this paper does not include empirical results in applying Cosmos WFMs to them. We are eager to verify the claims in future work.

2 Current Cosmos

Fig.˜4 visualizes what is available in the Cosmos WFM platform that is included in this paper, which includes video curator, video tokenization, world foundation model pre-training, world foundation model post-training, and guardrail.

Video curator. We develop a scalable video data curation pipeline. Each video is split into individual shots without scene changes. A sequence of filtering steps is then applied to the clips to locate high-quality and dynamic information-rich subsets for training. These high-quality shots are then annotated using a VLM. We then perform semantic de-duplication to construct a diverse but compact dataset.

Video tokenization. We develop a family of video tokenizers of different compression ratios. These tokenizers are causal. The token computation for the current frames is not based on future observation. This causal design has several benefits. On the training side, it makes joint image and video training possible since a causal video tokenizer is also an image tokenizer when the input is a single image. This is important for the video model to leverage image datasets for training, which contain rich appearance information of the worlds and tend to be more diverse. On the application side, causal video tokenizers are better aligned with Physical AI systems that live in the causal world.

WFM pre-training. We explore two scalable approaches for building pre-trained world foundation models—the diffusion model and the autoregressive model. We use the transformer architecture for its scalability.

For the diffusion-based WFM, the pre-training consists of two steps: 1) Text2World generation pre-training and 2) Video2World generation pre-training. Specifically, we train the model to generate a video world based on the input text prompt. We then fine-tune it to generate a future video world based on the past video and an input text prompt, which we refer to as the Video2World generation task.

For the autoregressive-based WFM, the pre-training consists of two steps: 1) vanilla next token generation and 2) text-conditioned Video2World generation. We first train the model to generate a future video world based on the input of past video—foresight generation. We then fine-tune it to generate a future video world based on the past video and a text prompt.

The video2world generation model is a pre-trained world model that generates the future based on the current observation (the past video) and control input (prompt). For both diffusion-based and autoregressive-based WFMs, we build a family of models with different capacities and study their effectiveness on various downstream applications.

We further fine-tune our pre-trained diffusion WFM to arrive at a diffusion decoder to enhance the generation results of the autoregressive model. To better control the WFM, we also built a prompt upsampler based on a Large Language Model (LLM).

World model post-training. We show applications of the pre-trained WFMs on several downstream Physical AI applications. We fine-tune a pre-trained WFM with the camera pose as the input prompt. This allows us to navigate freely in the created world. We also demonstrate how our pre-trained WFMs might be fine-tuned for humanoid and autonomous driving tasks.

Guardrail. For safe usage of the developed world foundation models, we develop a guardrail system where harmful inputs and outputs are blocked.

Data Curation

We describe our video curation pipeline, which produces high-quality training datasets for both tokenizers and WFMs. As shown in Fig.˜5, our pipeline consists of 5 main steps: 1) splitting, 2) filtering, 3) annotation, 4) deduplication, and 5) sharding. Every step is tailored to improve the data quality and accommodate the requirements of model training. We first present our raw dataset and then describe each step in detail.

We use both proprietary video datasets and publicly available open-domain Internet videos to train our models. Our goal is to enable Physical AI developers. To this end, we curate the video training dataset to cover various Physical AI applications and target the following video categories:

Hand motion and object manipulation (16%),

These videos offer a broad coverage of different visual objects and actions. Their diversity improves the generalization of our WFMs and helps the models handle different downstream tasks. The unstructured nature of these videos and their sheer volume creates many challenges to processing them efficiently from both an algorithmic and an infrastructural perspective. The videos can be encoded with a wide variety of codecs and have different aspect ratios, resolutions, lengths, \etc. Many videos have also been post-processed or edited with different visual effects, which may induce unwanted artifacts in the generated videos and hurt the performance of the world models if not appropriately handled.

In total, we accumulate about 20M hours of raw videos with resolutions from 720p to 4k. However, a significant amount of the video data is either semantically redundant or does not contain useful information for learning the physics of the world. Hence, we design a sequence of data processing steps to find the most valuable parts of the raw videos for training. We also collect image data as joint-image-and-video training has been shown to improve the visual quality of the generated videos and accelerate the model training. Thanks to the modular design of our data curation pipeline, we can use it to process both image and video data and generate datasets for both pre-training and fine-tuning. We generate about 10810^{8} video clips for pre-training and about 10710^{7} for fine-tuning.

2 Splitting

Our videos have arbitrary lengths, and modern deep-learning models cannot directly consume very long videos. Also, many videos contain shot transitions. They can start from one scene and then transition to a different scene where the two scenes can be disconnected entirely, \eg, from two people talking in a modern kitchen in New York City to a scene of lions chasing zebra in an African savanna. It is important to segment each video based on its shot changes and generate visually consistent video clips so that the model can learn visual content transitions that are physically plausible instead of artificially edited.

Splitting aims to temporally segment raw videos of arbitrary lengths into clips without shot changes. It takes the raw videos as input and generates each shot’s start and end frame indices. Clips shorter than 2s are discarded, as they could be shot transitions or visual effects. Clips longer than 60s are further split to have a maximal length of 60s. The subsequent filtering steps can then determine whether a clip contains useful information for learning the physics of the world.

Shot boundary detection is a classical computer vision problem. Existing methods detect shot boundaries based on changes in the visual feature space, but they differ in how to learn visual features from video frames. We evaluate several algorithms for the task in Tab.˜1: PySceneDetect (Castellano, 2024), Panda70M (Chen et al., 2024), TransNetV2 (Soucek and Lokoc, 2024), and AutoShot (Zhu et al., 2023).

PySceneDetect is a popular library that detects shot changes by thresholding the temporal change of color histogram in HSV space. Note that it is also adopted by the recent MovieGen work (Polyak et al., 2024). Panda70M augments PySceneDetect with CLIP-embedding-based stitching and filtering. TransNetV2 and AutoShot, on the other hand, are neural network-based, predicting a probability of each frame being a transition frame given a 100-frame rolling input window.

It is critical to select an algorithm that can handle heavily edited videos well, as they often have complex shot changes compounded with various visual effects. This motivates us to build a dedicated benchmark to evaluate whether the method can generate clips with clean shot cuts from videos. Our benchmark (named ShotBenchShotBench is available at https://github.com/NVlabs/ShotBench.) includes existing datasets, such as RAI, BBC Planet Earth (AI Image Lab, University of Modena, 2016), ClipShots (Tang et al., 2018) and SHOT (Zhu et al., 2023). For ClipShots, we define the transition frame as the midpoint of the start and end of each shot annotation to be consistent with other datasets.

Tab.˜1 compares the different methods on ShotBench. We set the confidence threshold to 0.4 for both TransNetV2 and AutoShot. For Panda70M, we follow their implementation for splitting, excluding the filtering steps, for a fair comparison. End-to-end learning-based approaches (\eg, TransNetV2 and AutoShot) perform much better than methods using hand-crafted features or heuristic rules (\eg, PySceneDetect and Panda70M). Though TransNetV2 and AutoShot perform comparably on existing datasets, we found TransNetV2 works better on more challenging shot changes. Using an end-to-end neural network (\ie, TransNetV2) also allows us to increase the throughput of splitting by leveraging modern GPUs for acceleration without the hurdle of hybrid approaches (such as Panda70M) that use complicated logic to combine PySceneDetect and ImageBind embeddings (Girdhar et al., 2023).

2.2 Transcoding

Our videos use many different codecs with various settings, which poses challenges to data curation. We re-encode each video clip from shot detection into a consistent, high-quality mp4 format. This simplifies the subsequent data curation process. With a unified video codec, the stability and efficiency of our dataloader for model training are also greatly improved. We use the h264_nvenc codec with a high bitrate and stress test our setting using videos with fast motion and high-frequency texture to ensure no perceptible visual degradation.

We thoroughly evaluate different hardware and software configurations for transcoding to maximize the throughput in Tab.˜2. Modern GPUs provide hardware-accelerated video encoding and decoding capabilities. NVIDIA L40S has hardware accelerators for both decoding (NVDEC) and encoding (NVENC), whereas NVIDIA H100 only has NVDEC. We compensate H100 with the maximum available CPU cores (28 instead of 1) for a fair comparison with L40S in Tab.˜2. L40S has about 17% higher throughput than H100 (0.0674 \vs0.0574). For software configurations, switching from libx264 to h264_nvenc and transcoding multiple clips from the same video in batches significantly boost the throughput. We observe issues with ffmpeg fully utilizing NVDEC/NVENC accelerators, especially on multi-GPU nodes. Replacing ffmpeg with PyNvideoCodec for video stream transcoding leads to much higher accelerator utilization and the biggest throughput improvement (0.3702 \vs0.1026). We only keep ffmpeg for audio remixing and use PyNvideoCodec to better leverage the computing power in the GPUs. We achieve a ∼6.5×\sim 6.5\times increase in throughput when combining all the improvements together.

3 Filtering

The video clips produced from the splitting step are noisy, with vastly different qualities covering various topics. We design the filtering step to 1) remove video clips whose visual quality fails to meet our minimal requirements, 2) select high-quality video clips suitable for fine-tuning, and 3) tailor the data distribution for building WFMs. We achieve the above goal by doing motion filtering, visual quality filtering, text filtering, and video type filtering.

We have two main goals in motion filtering: 1) remove videos that are static or with random abrupt camera motion (usually from hand-held cameras) and 2) tag videos with different types of camera motion (\eg, pan, zoom, tilt, \etc), which can provide additional information to guide model training.

We build a lightweight classifier for motion filtering. The input to the classifier is a sequence of motion vectors or optical flow extracted from a video clip. The classifier is based on the ViT architecture and is trained with labeled videos. We experiment with motion vectors from h264 codec, the Farneback optical flow algorithm(Farnebäck, 2003), and an NVIDIA TensorRT-accelerated optical flow estimation network. We find that the classifier built on top of the NVIDIA TensorRT-accelerated optical flow estimation works the best, producing high classification accuracy for motion filtering.

3.2 Visual Quality Filtering

We consider two criteria, distortion and appearance quality, for visual quality-based filtering. First, we remove video clips with distortions, such as artifacts, noise, blur, low sharpness, overexposure, underexposure, \etc. We use a video quality assessment model trained on human-rated videos based on DOVER (Wu et al., 2023a). This gives a perceptual quality score per clip, and we use the scores to remove clips that are in the bottom 15%15\%. Second, we filter out video clips with low appearance quality. We apply an image aesthetic model (Schuhmann, 2022) on sampled frames from an input clip. We set a conservative threshold, \ie, 3.53.5, since aesthetics are less important for Physical AI.

3.3 Text Overlay Filtering

Some of our videos are post-processed to add text to include additional information for the viewer. We also find that text tends to co-occur with different visual effects. Our goal is to learn the physics of the world. It is crucial to remove videos with such excessive text. Note that we focus on text added in post-processing instead of text in the original scene from which the video is created, such as the street names in driving videos.

We train an MLP-based binary classifier to detect such videos. The input to the classifier is a video embedding extracted using InternVideo2 (Wang et al., 2025). We use a proprietary VLM to build the training set to label positive and negative videos. Our trained model achieves high prediction accuracy in the validation set.

3.4 Video Type Filtering

To adjust the training data distribution and filter out unwanted video types, we design a comprehensive taxonomy that categorizes videos based on their content type and visual style. We train a classifier to label each video clip with categories from the taxonomy. We refine our data by excluding specific video types that could lead to poor generation quality or unrealistic dynamics, such as abstract visual patterns, video game footage, animated content, \etc. We further adjust the data distribution by upsampling from categories that are more relevant to WFMs (\eg, human action, human and object interaction, \etc) and downsampling on categories that are less important (\eg, nature or landscape videos).

Given the absence of pre-existing labeled datasets matching our taxonomy, we leverage a proprietary VLM to create training and evaluation data for the classifier. For each video clip, we prompt the VLM with eight uniformly sampled frames and query for the most appropriate taxonomy label. Using the annotated data, we train an MLP classifier on the same InternVideo2 embeddings from text filtering.

4 Annotation

Text descriptions are usually paired with image and video data to provide supervision and conditions for world model training. We use a VLM to generate high-quality and consistent captions for each video clip. We configure the VLM in a way such that it focuses on the material facts and details in the videos. Using this approach to provide descriptions of videos instead of relying on Alt text also eases the burden of learning for world models as we do not need to adapt to different text styles or formats during training.

We test several SOTA methods (\ie, VFC (Ge et al., 2024), Qwen2-VL (Wang et al., 2024b), VILA (Lin et al., 2024b; Xue et al., 2024)) for caption generation on our videos, and find VILA generates more accurate descriptions based on a small-scale human evaluation. We use an internal VILA model with 13B parameters, fine-tuned for video captioning. It has an enlarged context window suitable for processing long, multi-frame contexts, with a max input and output token length of 5904 and 256, respectively. To improve the inference efficiency, we use an FP8-quantized TensorRT-LLM engine, resulting in a 10 ×\times speed-up in throughput compared to a PyTorch half-precision baseline, as shown in Tab.˜3. We prompt VILA with “Elaborate on the visual and narrative elements of the video in detail” and feed it 8 uniformly sampled frames from the input clip. The average length of captions is 559 characters or 97 words.

5 Deduplication

Given the sheer volume of our videos, there could be duplicated or near-duplicated samples in the training set. It is critical to deduplicate the data to create a more balanced and diverse data distribution. It also improves the efficiency of training and reduces the chance of memorizing specific training samples.

We adopt the approach from SemDeDup (Abbas et al., 2023) and DataComp (Gadre et al., 2024) for scalable semantic deduplication. We reuse the InternVideo2 embeddings computed during filtering and cluster the embeddings using a multi-node GPU-accelerated implementation of k-means (RAPIDS, 2023) with k=10,000k=10,000. We compute the pairwise distances within each cluster of embeddings to identify duplicates. When duplicated videos are detected, we choose the video with the highest resolution to ensure no quality is lost due to deduplication. To avoid storing the entire pairwise distance matrix in GPU memory, we calculate on-the-fly the necessary upper-triangular matrix and argmax reduction in blocks of 256. We remove about 30%30\% of training data during deduplication.

We also leverage the extracted InternVideo2 embeddings and clustering results to build a visual search engine that supports querying the whole training dataset with free-form text and videos. The search engine is useful for debugging issues in our data and understanding the gap between the pre-training dataset and downstream applications.

6 Sharding

This step aims to package the processed video clips into webdatasets that our model trainer can directly consume for training. We shard the videos based on their resolution, aspect ratio, and length to align with our training curriculum. Besides pre-training datasets, we also create fine-tuning datasets with even higher quality by leveraging the different filters described above.

7 Infrastructure

Our data processing infrastructure uses AnyScale Ray (Moritz et al., 2017) to implement a streaming pipeline system for geographically distributed clusters, addressing two key challenges in large-scale ML workflows: efficient resource utilization across homogeneous nodes and robust operation over high-latency connections to data sources. By decoupling data transfer from computation, pipelines operate efficiently with remote data storage while maintaining memory requirements that scale with pipeline complexity rather than dataset size, enabling unbounded stream processing.

Our architecture enables concurrent utilization of complementary hardware resources through parallel pipeline stages, for instance, simultaneously using network bandwidth for data ingestion, NVDEC units for video decoding, and GPUs for compute-intensive transformations. We extend the Fragmentation Gradient Descent algorithm (Weng et al., 2023) to optimize this multi-resource allocation, with our scheduler automatically scaling individual stages to maintain balanced throughput across specialized hardware accelerators.

Tokenizer

Tokenizers are fundamental building blocks of modern large-scale models. They transform raw data into more efficient representations by learning a bottle-necked latent space discovered in an unsupervised manner. Specifically, visual tokenizers map raw and redundant visual data—such as images and videos—into compact semantic tokens, making them crucial for handling high-dimensional visual data. This ability not only enables efficient training of large-scale transformer models but also democratizes their inference on limited computational resources. Fig.˜6 schematically illustrates the tokenization training pipeline where the goal is to train the encoder and decoder so that the bottleneck token representation maximally preserves visual information in the input.

Tokenizers come in two types: continuous and discrete (see Fig.˜7 for illustrations). Continuous tokenizers encode visual data into continuous latent embeddings, as in latent diffusion models like Stable Diffusion (Rombach et al., 2022) or VideoLDM (Blattmann et al., 2023b). These embeddings are suitable for models that generate data by sampling from continuous distributions. Discrete tokenizers encode visual data into discrete latent codes, mapping them into quantized indices, as seen in autoregressive transformers such as VideoPoet (Kondratyuk et al., 2024). This discrete representation is necessary for models such as GPT that are trained with the cross-entropy loss. Fig.˜7 illustrates the two types of tokens.

The success of tokenizers largely relies on their ability to deliver high compression rates without compromising their subsequent visual reconstruction quality. On one hand, high compression reduces storage and computational demands. On the other hand, excessive compression can lead to the loss of essential visual details. This trade-off presents a significant challenge in tokenizer design.

We present Cosmos Tokenizer, a suite of visual tokenizers that includes both continuous and discrete tokenizers for images and videos. Cosmos Tokenizer offers exceptional visual reconstruction quality and inference efficiency. It offers a range of compression rates to accommodate diverse computational constraints and application needs. Tab.˜4 presents a comparison of different visual tokenizers and their capabilities.

We design Cosmos Tokenizer using a lightweight and computationally efficient architecture with a temporally causal mechanism. Specifically, we employ causal temporal convolution layers and causal temporal attention layers to preserve the natural temporal order of video frames, ensuring seamless tokenization of images and videos using a single unified network architecture.

We train our tokenizers directly on high-resolution images and long-duration videos without limiting the categories or aspect ratios. Unlike existing tokenizers that focus on specific data categories and sizes, the Cosmos Tokenizer operates across various aspect ratios—including 1:1, 3:4, 4:3, 9:16, and 16:9. They are temporally length-agnostic during inference, capable of tokenizing beyond the temporal length on which it was trained.

We also evaluate our tokenizers on standard image and video benchmarking datasets, including MS-COCO 2017 (Lin et al., 2014), ImageNet-1K (Deng et al., 2009), and DAVIS (Perazzi et al., 2016). To facilitate the video tokenization study for Physical AI applications, we curate a video dataset that covers many video categories for Physical AI, ranging from fish-eye, robotics, driving, human activities, and spatial navigation. The dataset is available at github.com/NVlabs/TokenBench.

As shown in Fig.˜8, our evaluation results demonstrate Cosmos Tokenizer significantly outperforms existing tokenizers by a large margin—for instance, achieving a +4 dB PSNR improvement in reconstruction quality on DAVIS videos. It runs up to 12×12\times faster and can encode videos up to 8 seconds at 1080p and 10 seconds at 720p in one shot without running out of memory on a single NVIDIA A100 GPU with 80GB memory. A suite of pre-trained models, with spatial compression of 8×8\times and 16×16\times, and temporal compression factors of 4×4\times and 8×8\times is available at github.com/NVIDIA/Cosmos-Tokenizer.

Our architecture employs a temporally causal design, ensuring that each stage processes only current and past frames, independent of future frames. Unlike common approaches, our tokenizer operates in the wavelet space, where inputs are first processed by a 2-level wavelet transform. Specifically, the wavelet transform maps the input video x0:Tx_{0:T} in a group-wise manner to downsample the inputs by a factor of four along xx, yy, and tt. The groups are formed as: {x0,x1:4,x5:8,...,x(T−3):T}→{g0,g1,g2,...,gT/4}\{x_{0},x_{1:4},x_{5:8},...,x_{(T-3):T}\}\rightarrow\{g_{0},g_{1},g_{2},...,g_{T/4}\}. Subsequent encoder stages process the frames in a temporally causal manner as {g0,g0:1,g0:2,...}→{ξ0,ξ1,ξ2,...}\{g_{0},g_{0:1},g_{0:2},...\}\rightarrow\{\xi_{0},\xi_{1},\xi_{2},...\}. Successive encoder stages follow a similar scheme, finally outputting the tokens z0:T′z_{0:T^{\prime}}. The causal design helps adapt models built on top of the tokenizer to downstream Physical AI applications that often operate on the temporal causal setting. The wavelet transform allows us to operate on a more compact video representation that eliminates redundancies in pixel information, allowing the remaining layers to focus on more semantic compression.

Our encoder stages (post wavelet transform) are implemented using a series of residual blocks interleaved with downsampling blocks. In each block, we employ a spatio-temporal factorized 3D convolution, where we first apply a 2D convolution with a kernel size of 1×k×k1\times k\times k to capture spatial information, followed by a temporal convolution with a kernel size of k×1×1k\times 1\times 1 to capture temporal dynamics. We use left padding of k−1k-1 to ensure causality. To capture long-range dependencies, we utilize a spatio-temporal factorized causal self-attention with a global support region—for instance, 1+T′1+T^{\prime} for the last encoder block. We use the Swish activation function (Ramachandran et al., 2017) for non-linearity. We leverage Layer Normalization (LayerNorm) (Lei Ba et al., 2016) instead of Group Normalization (GroupNorm) (Wu and He, 2018), which prevents large magnitudes from appearing in specific regions of the latent space or reconstructed outputs (Karras et al., 2020; Sadat et al., 2024). The decoder mirrors the encoder, replacing the downsampling blocks with an upsampling block. Fig.˜9 depicts an overview of the overall Cosmos Tokenizer architecture.

We employ the vanilla autoencoder (AE) formulation to model the continuous tokenizer’s latent space. For discrete tokenizers, we adopt the Finite-Scalar-Quantization (FSQ) (Mentzer et al., 2023) as the latent space quantizer. The latent dimension for the continuous tokenizers is 1616, whereas for the discrete tokenizers, it is 66, which represents the number of the FSQ levels, which are (8,8,8,5,5,5)(8,8,8,5,5,5). This configuration corresponds to a vocabulary size of 64,00064{,}000.

2 Training Strategy

We employ a joint training strategy by alternating mini-batches of images and videos at a preset frequency. We only supervise the final output of our tokenizer’s decoder. We do not use auxiliary losses tapped into the latent spaces, such as commitment or KL prior losses. For example, if a VAE (Kingma, 2013) formulation were used for continuous tokenizers instead of the vanilla AE, one would need to have the KL prior loss. If a VQ-VAE (van den Oord et al., 2017) were used for discrete quantization instead of the FSQ, one would need to have the commitment loss.

We employ a two-stage training scheme. In the first stage, we optimize with the L1 loss that minimizes the pixel-wise RGB difference between the input and reconstructed video (x^0:T\hat{x}_{0:T}), given by

and the perceptual loss based on the VGG-19 features (Simonyan and Zisserman, 2014), given by,

In the second stage, we use the optical flow (OF) loss (Teed and Deng, 2020) to handle the temporal smoothness of reconstructed videos,

and the Gram-matrix (GM) loss (Gatys et al., 2016) to enhance the sharpness of reconstructed images,

Additionally, we use adversarial loss in the fine-tuning stage to further enhance reconstruction details, particularly at large compression rates.

We train the image tokenizers (denoted as CI and DI) at two compression rates: 8×88\times 8 and 16×1616\times 16. Similarly, we train the video tokenizers (denoted as CV and DV) at three compression rates: 4×8×84\times 8\times 8, 8×8×88\times 8\times 8, and 8×16×168\times 16\times 16. Here, the compression rates are expressed as H×WH\times W for images and T×H×WT\times H\times W for videos, where TT represents the temporal dimension, and HH and WW represent the spatial dimensions.

For the video tokenizers, we create two variants:

Cosmos-0.1-Tokenizer: Trained using mini-batches sampling a smaller number of video frames (49 frames for CV and 17 frames for DV).

Cosmos-1.0-Tokenizer: Trained using mini-batches sampling a larger number of video frames (121 frames for CV and 49 frames for DV).

This approach ensures flexibility in handling varying temporal and spatial resolutions for image and video data.

3 Results

We extensively evaluate our Cosmos Tokenizer suite on various image and video benchmark datasets. For the evaluation of image tokenizers, we follow prior art to evaluate MS-COCO 2017 (Lin et al., 2014) and ImageNet-1K (Deng et al., 2009). We use the MS-COCO 2017 validation subset of 5,0005{,}000 images, and ImageNet-1K validation subset of 50,00050{,}000 images as image evaluation benchmark.

TokenBench. For video tokenizer evaluation, there is not yet a standard benchmark for high-resolution and long-duration videos. To this end, we introduce a benchmark called TokenBench to cover a wide variety of domains, including robotic manipulation, driving, egocentric, and web videos, and standardize the evaluation. We resort to existing video datasets that are commonly used for various tasks, including BDD100K (Yu et al., 2020), EgoExo-4D (Grauman et al., 2024), BridgeData V2 (Walke et al., 2023), and Panda-70M (Chen et al., 2024). We randomly sample 100100 videos from each dataset and preprocess them by taking the first 1010 seconds and resizing the short size to 10801080. For Panda-70M, we manually filter out the videos with low-quality content and small motions. For EgoExo-4D, we randomly pick 100100 scenes and sample one egocentric video and one exocentric video. This results in a total of 500500 videos. Some examples of TokenBench can be found in Fig.˜10. We release TokenBench at the github.com/NVlabs/TokenBench.

In addition to TokenBench, we also evaluate our video tokenizers on the DAVIS dataset at 10801080p resolution.

Baselines and evaluation metrics. We evaluate our tokenizers at various compression rates to showcase their effectiveness for different computational needs. We compare each of these tokenizers with state-of-the-art image and video tokenizers. Tab.˜4 presents the specific SOTA tokenizers we compared against in various settings. The evaluation metrics include Peak Signal-to-Noise Ratio (PSNR), Structural Similarity (SSIM), reconstruction Fréchet Inception Distance (rFID) (Heusel et al., 2017) for images, and reconstruction Fréchet Video Distance (rFVD) (Unterthiner et al., 2019) for videos.

Quantitative results. Tabs.˜6 and 6 summarize the average quantitative metrics of continuous and discrete video tokenizers on various benchmarks. As shown in both tables, Cosmos Tokenizer achieves state-of-the-art performance in all the metrics compared to prior arts on both the DAVIS video dataset and TokenBench, with a spatial-temporal compression ratio of 4×8×84\times 8\times 8. Moreover, even with 2×2\times and 8×8\times higher compression ratios (\ie, 8×8×88\times 8\times 8 and 8×16×168\times 16\times 16), Cosmos Tokenizer still achieves better quality than prior art, showcasing an excellent compression-quality trade-off.

Tabs.˜8 and 8 summarize the average quantitative metrics of continuous and discrete image tokenizers on various image benchmarks, covering a wide range of image types. As shown, compared to prior arts, Cosmos Tokenizer consistently achieves state-of-the-art results with a compression ratio of 8×88\times 8. More importantly, at a 4×4\times larger compression ratio of 16×1616\times 16, the image quality of Cosmos Tokenizer is often comparable or even better than prior art at 8×88\times 8 compression ratio, as shown in Tabs.˜8 and 8.

These quantitative results on a variety of image and video benchmark datasets confirm that Cosmos Tokenizer is able to better represent visual content with large spatial-temporal compression.

Runtime performance. Tab.˜9 shows the number of parameters and the averaged encoding and decoding times per image or per video frame, measured on a single A100 80GB GPU. In comparison, we also list the parameters and the average speeds of prior state-of-the-art tokenizers. As shown, for both image and video tokenizers, Cosmos Tokenizer is 2×∼12×2\times\sim 12\times faster while maintaining the smallest model size compared to prior arts, showing that Cosmos Tokenizer has high efficiency for encoding and decoding visual content.

World Foundation Model Pre-training

Pre-trained WFMs are generalists that capture general knowledge of real-world physics and natural behaviors. We exploit two different scalable deep learning paradigms, diffusion models and autoregressive models, to build two families of WFMs. Both diffusion models and autoregressive models break a difficult generation problem into a sequence of easier sub-problems and have been turbo-charging the development of generative models. In the case of diffusion models, the difficult generation problem is divided into a sequence of denoising problems. In the case of autoregressive models, the difficult generation problem is divided into a sequence of next-token prediction problems. We discuss how we scale these deep learning paradigms using various parallelization techniques tailored for modern GPUs in our endeavor of building pre-trained WFMs. We train all of the WFM models reported in the paper using a cluster of 10,00010{,}000 NVIDIA H100 GPUs in a time span of three months.

In Tab.˜10, we present a map of our pre-trained WFMs and their companions. For the diffusion-based WFM family, we start by building two Text2World models of 7B and 14B, respectively, which render Cosmos-1.0-Diffusion-7B-Text2World and Cosmos-1.0-Diffusion-14B-Text2World. These models can map text prompts to videos of visual worlds. We then fine-tune the Text2World models to take additional video input, representing the current observation. The result is a Video2World model where the future video is predicted based on the current observation (input video) and the perturbation (text prompt). These diffusion models are latent diffusion models that take continuous tokens. We use Cosmos-1.0-Tokenizer-CV8x8x8 to produce the visual tokens. The training text prompts for the WFMs are produced by a VLM through video description generation. These descriptions follow a different distribution of human descriptions of videos. To mitigate the domain gap, we build Cosmos-1.0-PromptUpsampler-12B-Text2World based on the Mistral-NeMo-12B-Instruct model (Mistral and NVIDIA, 2024) to help convert human text prompts to those preferred by our diffusion-based WFMs.

For the autoregressive-based WFM family, we first build two base models that are 4B and 12B in size, respectively, to predict future videos purely based on the current video observation. We name them Cosmos-1.0-Autoregressive-4B and Cosmos-1.0-Autoregressive-12B, respectively. These are Llama3-style GPT models trained from scratch for the video prediction task and bear no language understanding. To enable autoregressive-based WFMs to utilize textual information for next token prediction, we incorporate T5 embeddings of the input text prompt into the WFMs through cross-attention layers added to the transformer blocks. These autoregressive WFMs use Cosmos-1.0-Tokenizer-DV8x16x16, which maps an input video to a few integers. The heavy compression of the tokenizer can sometimes lead to undesired distortions. To address the problem, we build a diffusion decoder (Cosmos-1.0-Diffusion-7B-Decoder-DV8x16x16ToCV8x8x8) through fine-tuning the Cosmos-1.0-Diffusion-7B-Text2World model to map discrete tokens in the DV8x16x16 space to continuous tokens in the CV8x8x8 space.

Our diffusion-based WFMs are latent diffusion models that operate within a learned latent space of a tokenizer, enabling a compact, reduced-dimensional representation of videos. This design choice offers several advantages: it reduces computational costs during both training and inference while simplifying the denoising task (Rombach et al., 2022; Hoogeboom et al., 2024). To tokenize videos into latent representations, we employ Cosmos-1.0-Tokenizer-CV8x8x8.

To train our diffusion WFMs, we adopt the approach outlined in EDM (Karras et al., 2022, 2024). The denoising score matching loss for the denoiser DθD_{\theta}, evaluated at a noise level σ\sigma, is defined as

where x0∼pdata{\mathbf{x}}_{0}\sim p_{\rm{data}} is a clean image or video sampled from the training set, {\mathbf{n}}\sim{\mathcal{N}}\big{(}\mathbf{0},\sigma^{2}{\mathbf{I}}\big{)} is i.i.d. Gaussian noise, and DθD_{\theta} is a noise-conditioned neural network tasked with denoising the corrupted sample x0+n{\mathbf{x}}_{0}+{\mathbf{n}}. We adhere to the preconditioning design introduced in EDM for parameterizing DθD_{\theta}. The overall training loss is defined as a weighted expectation of L(Dθ;σ){\mathcal{L}}(D_{\theta};\sigma) over the noise levels:

where the distribution of noise levels σ\sigma is controlled by hyperparameters PmeanP_{\text{mean}} and PstdP_{\text{std}}. σdata\sigma_{\text{data}} is the standard deviation of the training data, and the weighting function λ(σ)\lambda(\sigma) ensures equal contribution of each noise level at the beginning of the training. However, as training progresses, this balance may deteriorate. To mitigate this issue, we treat the optimization over various noise levels as a form of multi-task learning. We utilize the uncertainty-based weighting approach by introducing u(σ)u(\sigma) as a continuous uncertainty function quantifying the uncertainty for the denoising objective L(Dθ,σ){\mathcal{L}}(D_{\theta},\sigma) at noise level σ\sigma. We use a simple MLP to parameterize u(σ)u(\sigma) and minimize the overall loss L(Dθ){\mathcal{L}}(D_{\theta}) during training. Intuitively, the contribution of loss at noise level σ\sigma is weighted down if the model is uncertain about the task, \ie, if u(σ)u(\sigma) is high. At the same time, the model is penalized for this uncertainty, encouraging u(σ)u(\sigma) to be as low as possible.

Compared to recent video generative models that adopt the Gaussian flow matching formulation (Polyak et al., 2024; Kong et al., 2024), our work is derived from the diffusion score matching perspective (Ho et al., 2020; Song et al., 2020). However, as shown by Gao et al. (2024a), these frameworks are theoretically equivalent, sharing fundamental similarities in their objectives and training procedures. Our EDM-based formulation aligns with these insights, mainly differing in the choice of preconditioning designs and hyperparameters. In practice, we have not encountered any performance limitations with the EDM formulation.

1.2 Architecture

In this section, we describe the design of our denoiser network DθD_{\theta} that builds upon DiT (Peebles and Xie, 2023), which was originally designed for label-conditioned image generation. We adapt its architecture to better suit our goal of controllable video generation. We visualize the overall network design in Fig.˜11.

3D patchification. The input to our network is a latent representation of shape T×C×H×WT\times C\times H\times W for both image and video data, with images differentiated by a video with a single frame. To prepare inputs for our denoiser network, we first “patchify” the state using a linear layer and subsequently flatten it. This process involves projecting non-overlapping cubes of shape (pt,ph,pw)(p_{t},p_{h},p_{w}) into individual token inputs for the network. Consequently, after patchification, an image or video is reshaped into a one-dimensional, spatiotemporal sequence of length THW/(ptphpw)THW/(p_{t}p_{h}p_{w}). We use pt=1,ph=pw=2p_{t}=1,p_{h}=p_{w}=2 for our denoiser network.

Hybrid positional embedding with FPS-aware 3D RoPE and learnable embedding. We employ a 3D-factorized Rotary Position Embedding (RoPE) (Su et al., 2024) to allow the generation of arbitrary size, aspect ratio, and video length. Specifically, we partition the feature dimension into three approximately equal chunks, each applying RoPE with positional information along the temporal, height, and width axes, respectively. In practice, this can be implemented efficiently without splitting and concatenation in each block by concatenating frequency embeddings in their respective axes and reusing RoPE kernels optimized for Large Language Models (LLMs). To further support video synthesis with varying frame rates, we rescale temporal frequencies based on the training video’s Frames Per Second (FPS). Due to RoPE’s relative positional encoding property and our 3D factorization design, the FPS-aware design is compatible with our joint image-video training. An additional benefit of RoPE is evident during progressive training when we alter resolution or video length. By leveraging Neural Tangent Kernel (NTK)-RoPE (Peng and Quesnelle, 2023), we observe rapid model convergence, achieving reasonable performance even within 5,0005{,}000 training steps. Additionally, we find that adding an extra learnable absolute positional embedding per transformer block can further enhance the model, reduce training loss, and reduce morphing artifacts in generated videos.

Cross-attention for text conditioning. We rely on cross-attention layers in our network for incorporating linguistic information. Each transformer block consists of sequential self-attention, cross-attention, and feed-forward layers. While self-attention operates over spatiotemporal tokens, cross-attention integrates semantic context using T5-XXL (Raffel et al., 2020) embeddings as keys and values, enabling effective text conditioning.

Query-key normalization. In the early stages of training, we observe instability in the growth of attention logits, leading to a collapse of attention entropy. We follow existing literature (Dehghani et al., 2023; Wortsman et al., 2023; Esser et al., 2024) to normalize query QQ and key KK before the attention operation. We use Root Mean Square Normalization (RMSNorm) (Zhang and Sennrich, 2019) with learnable scales for all self-attention and cross-attention layers within our network.

AdaLN-LoRA. We find that DiT’s adaptive layer normalization (AdaLN) layers (Xu et al., 2019; Peebles and Xie, 2023) account for a significant portion of the model parameters while contributing negligibly to the computational complexity in terms of FLOPs. Inspired by W.A.L.T (Gupta et al., 2024a), we implement Low-Rank Adaptation (LoRA) (Hu et al., 2022) to decompose the dense linear projections in these layers into low-rank approximations. For Cosmos-1.0-Diffusion-7B, this architectural optimization achieves a 36% reduction in parameter count (from 11B to 7B parameters) while maintaining performance parity across all evaluation metrics, demonstrating the effectiveness of our parameter-efficient design.

1.3 Training Strategy

This section outlines the methodologies employed to train our models on datasets spanning multiple modalities, resolutions, aspect ratios, and conditioning inputs.

Joint image and video training. To leverage the vast abundance of high-quality, diverse image datasets in model training, we implement an alternating optimization strategy that interleaves batches of image and video data. To facilitate cross-modal knowledge transfer between image and video domains, we adopt a domain-specific normalization scheme that aligns the latent distributions using sufficient statistics estimated independently for image and video data. This approach is motivated by the observation that reducing the distributional shift between image and video latent representations improves generation quality. Furthermore, we observe non-stationary statistics across temporal and channel dimensions in video latent representations. To address this heterogeneity, we employ a normalization strategy that applies frame-wise and channel-wise standardization to video latent representations, effectively encouraging them to better approximate an isotropic Gaussian prior distribution.

Beyond cross-modality knowledge transfer, our normalization scheme provides an important theoretical benefit: scale invariance in the signal-to-noise ratio during training. Consider two zero-mean latent representations with different scales: one standardized to unit variance, and another with variance 4. When adding Gaussian noise N(0,σ2){\mathcal{N}}(0,\sigma^{2}) to achieve a desired signal-to-noise ratio for the standardized representation, we must scale the noise to N(0,4σ2){\mathcal{N}}(0,4\sigma^{2}) for the unnormalized representation to maintain the same ratio. By standardizing all latent representations, we ensure consistent signal-to-noise ratios across different scales, facilitating model adaptation even when the underlying tokenizer is updated during training.

To maintain computational efficiency, we balance image and video batch sizes to ensure comparable memory utilization across GPUs. However, we observe that the video batch denoising loss exhibits slower convergence compared to the image batch loss. We attribute this to the inherent temporal redundancy in video frames, which results in smaller gradient magnitudes for video batches. Drawing inspiration from recent advances in multi-resolution image training (Chen, 2023; Hoogeboom et al., 2023; Atzmon et al., 2024), we address this convergence discrepancy by scaling the video batch noise levels by the square root of the frame count relative to image batch noise levels.

Progressive training. We adopt a progressive training strategy, with the specifics of each stage detailed in Tab.˜12. The initial stage involves training on videos and images at a resolution of 512 pixels, using videos composed of 57 frames. Subsequently, we transition to the target resolution of 720 pixels, increasing the video length to 121 frames. After pre-training on massive data, we fine-tune the model on a high-quality subset for O(10k)\mathcal{O}(10k) iterations with a linearly decaying learning rate. Consistent with findings from Dai et al. (2023), we also find that fine-tuning can improve the quality of the generated videos.

Multi-aspect training. To accommodate content with varying aspect ratios, we organize the data into five distinct buckets corresponding to ratios of 1:1, 3:4, 4:3, 9:16, and 16:9, assigning each image or video to the bucket with the closest aspect ratio. During training, each data parallel process group samples from one bucket, allowing different buckets across different parallel process groups. We implement longest-side resizing to maximally preserve the original content information described in the prompt. For batch processing, we apply reflection padding to missing pixels and supply the padding mask to the diffusion backbone, enabling precise control during inference.

Mixed-precision training. We maintain two copies of the model weights: one in BF16 and another in FP32. During the forward and backward passes, the BF16 weights are used to improve training efficiency, resulting in gradients and activations also in BF16 format. For parameter updates, the weights are updated in FP32 to ensure numerical stability. The updated FP32 parameters are then copied and cast to BF16 for the next iteration. To further stabilize training, we scale the loss of denoising score matching in Eq.˜5 by a factor of 10. We also find that lower betas and eps coefficients in AdamW significantly reduce loss spikes. For our 14B diffusion model training, we rarely encountered loss spikes, and there were no non-recoverable loss spikes.

Text conditioning. For our Text2World models, we employ T5-XXL (Raffel et al., 2020) as the text encoder. We zero-pad T5 embeddings to maintain a fixed sequence length of 512. To enhance text-context alignment, we adopt classifier-free guidance (Ho and Salimans, 2022). Unlike prior works (Balaji et al., 2022; Saharia et al., 2022) that randomly zero out text embeddings, we omit this step due to the effectiveness of negative prompts during inference. Notably, as a text-to-image generator, our model excels in generating high-fidelity images even without guidance, a capability we attribute to the high-quality training dataset. While classifier-free guidance typically promotes mode-seeking behavior for preferred visual content, we find that careful data selection achieves a similar effect. However, for video generation, the lack of comparable high-quality data leads to suboptimal results under low guidance settings. Consequently, higher guidance values are required to produce satisfactory content in video-generation tasks.

Image and video conditioning. We extend our Text2World models to build Video2World models that support image and video conditioning by incorporating previous frame(s) into the generation process. Specifically, the conditional frame(s) are concatenated with the generated frames along the temporal dimension. To improve robustness against variations in input frame(s) during inference, we introduce augmented noise to the conditional frames during training. The sigma value for this augmented noise is sampled with Pmean=−3.0,Pstd=2.0P_{\text{mean}}=-3.0,P_{\text{std}}=2.0. Additionally, the input to the diffusion model is concatenated along the channel dimension with a binary mask that distinguishes conditional frames from generated frames. The loss function excludes contributions from the locations of conditional frames, focusing exclusively on the generated output. To improve generalization, we randomly vary the number of conditional frames during training. During inference, the model can flexibly operate with either a single conditional frame (image) or multiple previous frames as input.

1.4 Scaling Up

Here, we outline the techniques that enable efficient scaling of our diffusion WFMs. We analyze the memory requirements of our models, discuss parallelism strategies, and compare our training setup against other video diffusion models and state-of-the-art LLMs.

Memory requirements. The four major components that consume the GPU memory are:

Model parameters: 10 bytes per parameter. Our mixed precision training stores model parameters in both FP32 and BF16, alongside Exponential Moving Average (EMA) weights in FP32.

Gradients: 2 bytes per parameter. We store the gradients in BF16.

Optimizer states: 8 bytes per parameter. We use AdamW (Loshchilov and Hutter, 2019) as our optimizer and store the optimizer states (\ie, first and second moments) in FP32.

Activations: (2×number_of_layers×15×seq_len×batch_size×d_model)(2\times\text{number\_of\_layers}\times 15\times\text{seq\_len}\times\text{batch\_size}\times\text{d\_model}) bytes. We store the activations in BF16. Tab.˜13 provides details of the stored activations for major operations within the network. To optimize memory usage, we implement selective activation checkpointing (Chen et al., 2016; Korthikanti et al., 2023), recomputing activations for memory-limited layers such as normalization functions.

For instance, our 14B model (Cosmos-1.0-Diffusion-14B-Text2World) requires approximately 280 GB for model parameters, gradients, and optimizer states, alongside 310 GB for activations during high-resolution pre-training. Given the 80GB HBM3 limit of NVIDIA H100 GPUs, we employ Fully Sharded Data Parallelism (FSDP) and Context Parallelism (CP) to distribute memory demands across multiple GPUs.

Fully Sharded Data Parallelism (FSDP). FSDP improves memory efficiency by sharding model parameters, gradients, and optimizer states across devices. It gathers parameters only when needed during computation and releases them afterward. Unlike standard data parallelism, which duplicates parameters across devices, FSDP distributes parameters, gradients, and optimizer states, with each device managing only its shard. This approach minimizes memory usage to the largest temporarily unsharded parameter set alongside its shard of parameters, gradients, and optimizer states. For our implementation, we utilize a sharding factor of 32 for the 7B model and 64 for the 14B model to balance memory and communication latency.

Context Parallelism (CP). Scaling transformers for long-context settings introduces challenges with increased FLOPs and activation memory. CP addresses these challenges by distributing computation and activations across multiple GPUs. It works by splitting both the query QQ and the key-value (K,V)(K,V) along their sequence dimensions into CP_SIZE chunks, where CP_SIZE is the number of GPUs within a CP group. Each GPU processes one chunk of QQ and iteratively accumulates partial attention outputs using blocks of (K,V)(K,V) stored in the same CP group. Different implementations of CP utilize different communication primitives, including all-gather (Dubey et al., 2024), P2P (Liu et al., 2023a), and all-to-all (Jacobs et al., 2023). We employ the P2P variant from TransformerEngine (NVIDIA, 2024e), which overlaps computation and communication by transferring (K,V)(K,V) blocks between GPUs while simultaneously processing attention. When block sizes are carefully chosen, this overlap effectively hides data transfer latency. We organize CP groups within NVLink-connected GPUs and overlap CP ranks with FSDP ranks for optimal utilization. For image iterations with shorter contexts, CP is disabled to improve throughput. Cross-attention layers do not use CP due to the shorter sequence lengths of (K,V)(K,V), which results in insufficient computation to mask communication latency.

Comparison with other video generative models. Our parallelism strategy is deliberately streamlined compared to approaches outlined in HunyuanVideo (Kong et al., 2024) and MovieGen (Polyak et al., 2024), which incorporate Tensor Parallelism (TP) and its extension, Sequence Parallelism (SP). Despite excluding TP/SP, our setup achieves comparable Model FLOPs Utilization (MFU). While TP/SP remains valuable in certain scenarios, such as larger models or alternative network topologies, a detailed analysis of tradeoffs is left for future work.

Comparison with large language models. Unlike LLMs, which are typically pre-trained with shorter context lengths, long-context settings significantly increase FLOPs due to the quadratic cost of self-attention. While FLOPs for LLMs are commonly calculated as 6×seq_len×P6\times\text{seq\_len}\times P, where PP is the number of parameters (Kaplan et al., 2020), we note that this formula is inaccurate for our diffusion WFMs. We provide the forward pass FLOPs of each key operation in Tab.˜13.

1.5 Prompt Upsampler

During training, our WFMs use detailed video descriptions as input text prompts to produce high-quality videos. However, during inference, user prompts may vary in length, structure, and style, often being much shorter. To bridge this gap between training and inference text prompts, we develop a prompt upsampler to transform original input prompts into more detailed and enriched versions. It can improve the prompts by adding more details and maintaining a consistent description structure, which leads to higher quality output.

The main requirements for the prompt upsampler are:

Fidelity to the input prompts: The upsampled prompt must faithfully preserve the key elements of the original user input, including the main characters, actions or motions, key attributes, and overall intent.

Alignment with training distribution: The upsampled prompt should closely resemble the distribution of training prompts of WFMs in terms of length, language structure, and style.

Enhanced visual details: The upsampled prompt should be designed to prompt the WFMs to generate more accurate imagery.

Prompt upsampler for Text2World model. We fine-tune Mistral-NeMo-12B-Instruct (Mistral and NVIDIA, 2024) to build our prompt upsampler. To obtain paired data, that is, short prompts simulating user input and the corresponding long prompts reflecting the distribution of training prompts, we use a VLM to generate short captions based on our training long prompts and corresponding videos. This long-to-short data creation strategy is effective in (1) preserving the authentic video content and distribution from detailed training prompts of WFMs and (2) ensuring fidelity between the short and long prompts. The resulting prompt upsampler is termed Cosmos-1.0-PromptUpsampler-12B-Text2World.

Prompt upsampler for Video2World model. For the Video2World model, the input consists of video conditions and a user text prompt. To enhance the user prompt, we utilize an open-source VLM, Pixtral-12B (Agrawal et al., 2024), combined with zero-shot prompt engineering, to upsample the prompt into a detailed description that considers both the video conditions and the user prompt. We found the vanilla Pixtral-12B model works well out of the box and did not proceed to perform a similar fine-tuning described above.

1.6 Results

In Fig.˜12, we present qualitative results generated by our Cosmos-1.0-Diffusion-7B-Text2World and Cosmos-1.0-Diffusion-14B-Text2World models. Both models produce videos of high visual quality, motion dynamics, and text alignment. Compared to the 7B model, the 14B model is able to generate videos capturing more complex visual details and intricate motions.

We show generated videos from Video2World 7B and 14B models in Fig.˜13. The Video2World models support both image and video conditioning and can generate extended videos in an autoregressive manner. As demonstrated in Fig.˜13, our Video2World models produce photorealistic videos with good motion dynamics and visual fidelity. The 14B model, again, generates better videos in terms of scene richness and motion stability.

2 Autoregressive-based World Foundation Model

In autoregressive WFMs, we formulate world simulation generation as a next-token prediction task similar to language modeling. We start by converting a video into a sequence of discrete video tokens V={v1,v2,…,vn}\mathcal{V}=\{v_{1},v_{2},\dots,v_{n}\} using the Cosmos Discrete Tokenizer introduced in Sec.˜4. Then we train a Transformer decoder (Vaswani et al., 2017) to predict the next video token using past video tokens as context, similar to large language models (LLMs) (Brown et al., 2020; Jiang et al., 2023; Dubey et al., 2024). Specifically, the training objective is to minimize the following negative log-likelihood (NLL) loss:

where the conditional probability PP of the predicted next video token viv_{i} is modeled by a Transformer decoder with parameters Θ\Theta.

Our autoregressive-based WFM architecture is illustrated in Fig.˜14. We make several modifications to the standard transformer model architecture tailored for our video generation task, including adding 1) 3D-aware positional embeddings, 2) cross-attention to enable textual inputs for better control, and 3) QK-Normalization (Wortsman et al., 2023).

3D positional embeddings. Similar to our diffusion-based WFM (Sec.˜5.1.2), we incorporate two complementary positional embedding mechanisms: 3D factorized Rotary Position Embedding (RoPE) for relative positions and 3D factorized absolute positional embedding (APE) for absolute coordinates. These mechanisms work in concert to provide comprehensive spatial and temporal information throughout the network.

3D Rotary Position Embedding (RoPE). We apply 3D RoPE to our model to encode relative positional information across the temporal, height, and width dimensions. During training, we adopt a multi-stage training strategy in which the sequence length of videos increases as the training progresses. To adapt the 3D RoPE to the changing temporal duration, we use YaRN (Peng et al., 2023), a compute-efficient technique designed to extend the context window of RoPE. We apply YaRN extension only along the temporal axis as the video sequence length increases only along the temporal dimension. By utilizing YaRN, our model can extrapolate to context lengths longer than those encountered during the initial stages of training.

3D Absolute Positional Embedding (APE). In addition to 3D RoPE, we incorporate a 3D APE within each transformer block to complement the relative positional encoding. This APE encodes positional information using sinusoidal embeddings factorized across temporal, height, and width dimensions, ensuring the model is aware of absolute positions. The embedding is added directly to the input tensor at each stage, enriching the positional context for the transformer. We find combining absolute and relative positional encodings enhances model performance, reduces training loss, and minimizes morphing artifacts in generated videos. Notably, while our diffusion-based WFM (Sec.˜5.1.2) employs learnable embeddings, we adopt sinusoidal-based embeddings for APE in our autoregressive-based WFM.

Vocabulary. Tokenization is a crucial step that turns input text into a sequence of discrete tokens in large language models (LLMs). In LLMs, the vocabulary of possible tokens is determined by the LLM’s tokenizer (\eg, tiktoken introduced by OpenAI (2022)) trained on a large corpus of text with algorithms such as Byte Pair Encoding (BPE) (Gage, 1994).

For our autoregressive models, we use our Cosmos-1.0-Tokenizer-DV8x16x16 as the tokenizer. As introduced in Sec.˜4, we leverage the Finite-Scalar-Quantization (FSQ) (Mentzer et al., 2023) to quantize the 66-dimensional latent space into (8,8,8,5,5,5)(8,8,8,5,5,5) levels. This quantization leads to a vocabulary size of 8×8×8×5×5×5=64,0008\times 8\times 8\times 5\times 5\times 5=64{,}000.

Cross-attention for text conditioning. In addition to the self-attention blocks present in the transformer architecture, we add cross-attention layers to enable the model to condition on input text. Similar to diffusion-based WFM(Sec.˜5.1.2), cross-attention is applied between the features of the transformer model and text embeddings obtained from a pre-trained text encoder (T5-XXL). In our experiments, we add cross-attention blocks after every self-attention layer.

Query-key normalization. In order to enhance training stability, we incorporate Query-Key Normalization (QKNorm) (Wortsman et al., 2023). QKNorm addresses instability in attention mechanisms by normalizing the query (QQ) and key (KK) vectors before computing their dot product, thereby preventing the softmax function from saturating and ensuring more effective learning. After normalization, the dot product is scaled by a learnable parameter γ\gamma instead of the fixed 1/dk1/\sqrt{d_{k}}. This learnable scaling factor allows the model to adaptively control the magnitude of the attention scores, enhancing flexibility and expressivity.

Z-loss. To further improve training stability, we introduce a stabilization term known as the z-loss (de Brébisson and Vincent, 2016) into our training objective. The z-loss penalizes deviations of the logits from zero, effectively discouraging the model from generating excessively large logit values that could result in numerical instability or gradient explosions. The z-loss is defined as the sum of the squared logits as Lz-loss=λ⋅∑izi2\mathcal{L}_{\text{z-loss}}=\lambda\cdot\sum_{i}z_{i}^{2}. We found z-loss to be critical in maintaining gradient norms to a healthy range, especially when scaling the training to a large number of GPU nodes. Empirically, we found that the z-loss coefficient λ=3×10−4\lambda=3\times 10^{-4} strikes an optimal balance, effectively stabilizing training without adversely affecting model performance.

2.2 Scaling Up

This section describes the techniques that enable efficient scaling of our autoregressive WFMs. We briefly analyze the memory consumption of our models, discuss parallelism strategies, and compare our training setup with other autoregressive models.

Memory requirements. During training, GPU memory is mainly consumed by:

Model parameters: 6 bytes per parameter. We store the model parameters in both BF16 and FP32.

Gradients: 2 bytes per parameter. We store the gradients in BF16.

Optimizer states: 8 bytes per parameter. We store the first and second moments of AdamW (Loshchilov and Hutter, 2019) both in FP32.

Activations: Approximately (2×number_of_layers×17×seq_len×batch_size×d_model)(2\times\text{number\_of\_layers}\times 17\times\text{seq\_len}\times\text{batch\_size}\times\text{d\_model}) bytes. We refer readers to Korthikanti et al. (2023) for a detailed analysis of activation memory of state-of-the-art autoregressive models.

For instance, our 12B model (Cosmos-1.0-Autoregressive-12B) demands approximately 192 GB of memory for its parameters, gradients, and optimizer states combined. As this is beyond a single NVIDIA H100 GPU’s 80GB HBM3 capacity, we leverage tensor parallelism (TP) (Shoeybi et al., 2019) and its extension, sequence parallelism (SP) (Korthikanti et al., 2023), to distribute the memory requirements and computation across multiple GPUs.

Tensor Parallelism (TP). Tensor Parallelism (TP) (Shoeybi et al., 2019) splits the weights of linear layers along either the input or output feature dimensions, with the choice guided by the goal of minimizing inter-GPU communication. For example, in a two-layer feedforward network, the weights of the first layer are partitioned along the output feature dimension, while those of the second layer are partitioned along the input feature dimension. This arrangement allows intermediate activations to be processed locally without requiring communication between GPUs. The final outputs are then combined using all-reduce communication. By employing TP, each GPU stores only a fraction, specifically 1/TP_SIZE1\mathbin{/}\text{TP\_SIZE}, of the weights for linear layers. However, the default implementation of TP still replicates activations along the sequence dimension for operations like LayerNorm, resulting in redundancy.

Sequence Parallelism (SP). SP (Korthikanti et al., 2023) extends Tensor Parallelism by further partitioning the context along the sequence dimension. This approach is applicable to operators, such as LayerNorm and Dropout in self-attention layers, where each element in the sequence can be processed independently. With SP enabled. Each GPU stores only a fraction, specifically 1/TP_SIZE1\mathbin{/}\text{TP\_SIZE}, of the activations.

Comparison with other autoregressive models. Compared to popular LLMs, our model doesn’t leverage memory-saving attention variants such as MQA or GQA. Otherwise, our autoregressive model is deliberately designed to closely resemble the architecture of LLMs (Brown et al., 2020; Dubey et al., 2024; Jiang et al., 2023; Adler et al., 2024; Team, 2024b; Yang et al., 2024a), as this alignment offers flexibility and scalability. Experiments that leverage more parallelisms, such as context parallelism and pipeline parallelism, to further scale up the model sizes and context lengths are left for future works.

2.3 Training Strategy

We perform pre-training of our autoregressive WFMs in multiple stages.

Stage 1: In the first stage, the model is trained using the video prediction objective. Given the first frame as the input condition, the model is trained to predict future video frames. A context length of 17 frames is used for this task, \ie, the model predicts 16 future frames with the first frame as input.

Stage 1.1: This stage performs video prediction but with an increased context length of 34 frames. We use the YaRN extension on the temporal dimension to increase the context length of RoPE.

Stage 2: In stage 2 of our training, we introduce text conditioning to our model. Text embeddings are incorporated using newly initialized cross-attention layers. The model is trained with a 34-frame context. To improve text-to-video generation ability, the model is trained using joint image and video data as described in Sec.˜5.1.3. When image batches are used, we use a larger batch size as the context length for images is much smaller than that of videos.

All our models are trained with a fixed spatial resolution of 640×1024640\times 1024.

Cooling down. After pre-training, we conduct a “cooling-down” phase with high-quality data, similar to LLM training practices (Dubey et al., 2024). During this phase, we linearly decay the learning rate to while training on high-quality image-video pairs. The cooling-down phase is carried out over 30,00030{,}000 iterations.

We train two sets of autoregressive-based WFMs. We start by building two base models: one with a 4B capacity and the other with a 12B capacity. These are pure next-video token predictors that do not take text prompts as input. We then derive a Video2World version from each of the base models, where we add cross-attention layers to them to leverage text prompt inputs for next video token prediction.

Cosmos-1.0-Autoregressive-4B: a 4B transformer model for next video token prediction. This model is trained using stage 1 and stage 1.1 of the multi-stage training objective.

Cosmos-1.0-Autoregressive-5B-Video2World: a 5B transformer model derived from our Cosmos-1.0- Autoregressive-4B and trained additionally with stage 2 of the multi-stage training objective.

Cosmos-1.0-Autoregressive-12B: a 12B transformer model for next video token prediction. This model is trained using stage 1 and stage 1.1 of the multi-stage training objective.

Cosmos-1.0-Autoregressive-13B-Video2World: a 13B transformer model derived from Cosmos-1.0-Autoregressive-12B and trained additionally with stage 2 of the multi-stage training objective.

2.4 Inference Optimization Towards Real-Time Generation

Our Cosmos Autoregressive WFMs share architectural similarities with LLMs, enabling us to leverage established LLM inference optimization techniques to address the sequential decoding bottleneck. We implement a combination of key-value caching, tensor parallelism, and torch.compile, following the gpt-fasthttps://github.com/pytorch-labs/gpt-fast implementation in PyTorch (Paszke et al., 2019).

Speculative decoding. To further accelerate our autoregressive WFMs, we apply the Medusa speculative decoding framework (Cai et al., 2024). Unlike common speculative decoding approaches that require a separate draft model (Leviathan et al., 2023) or training-free methods with limited speedup (Teng et al., 2024), Medusa extends the transformer backbone with extra decoding heads to predict multiple subsequent tokens in parallel. It then verifies these speculated tokens with rejection sampling. The inference is thus accelerated by alleviating the bottleneck of one-token-at-a-time processing. We demonstrate the potential of the Medusa technique in visual autoregressive acceleration without compromising the quality of generated outputs.

In our implementation, we fine-tune our pre-trained autoregressive WFMs by introducing Medusa heads into the architecture. These heads are strategically inserted after the last transformer hidden states, where all backbone parameters and the final unembedding layer are shared across different heads. Each Medusa head is a single-layer FFN with SiLU activation and residual connection. We further merge the weight matrices of multiple Medusa heads into a unified FFN to maximize parallelism during token prediction. Note that we do not use the tree-based attention mechanism from Cai et al. (2024).

To investigate the optimal Medusa setup for our autoregressive WFMs, we conduct an in-depth study from two aspects: (1) which transformer layers to fine-tune and (2) how many Medusa heads to add. For the first problem, we compare between full fine-tuning and selective layer freezing. We observe that only fine-tuning the Medusa heads gives poor multi-token prediction, while full fine-tuning incurs quality degradation. We empirically identify that unfreezing the last two transformer layers and the final unembedding layer while keeping the backbone frozen yields the best performance. This strategy ensures our Medusa training achieves decent speculative decoding accuracy without suffering from catastrophic forgetting.

To explore the optimal number of Medusa heads, we calculate the model token throughput and forward pass count with different numbers of Medusa heads. The ablation studies are conducted on 8 ×\times H100 GPUs and evaluated on 50 unseen test videos of 640×1024640\times 1024 resolution. The results in Tab.˜15 suggest that our Medusa framework can effectively accelerate inference, with up to 2.0×2.0\times token throughput and 4.6×4.6\times less forward pass for the 4B model, and up to 3.2×3.2\times token throughput and 6.1×6.1\times less forward pass for the 5B model. We show that though more Medusa heads can reduce the number of forward passes needed to generate, it may slow down the overall token throughput. We find that 99 Medusa heads yield the best trade-off between computational efficiency and model performance.

In Tab.˜16, we show performance analysis of autoregressive WFMs with Medusa integration. This analysis was conducted on H100 GPUs and evaluated on test videos of 640×1024640\times 1024 resolution in the BF16 precision. Results show that the Medusa implementation consistently accelerates inference for both 4B and 5B models under different GPU configurations.

Low-resolution adaptation for real-time inference. We pursue real-time inference by adapting our model to a lower spatial resolution of 320×512320\times 512, which results in a lower number of tokens per video. Specifically, we first fine-tune the discrete video tokenizer (Cosmos-1.0-Tokenizer-DV8×\times16×\times16 in Sec.˜4) on 320p low-resolution videos using videos from the target Physical AI domain. Then, we fine-tune our autoregressive WFM that is pre-trained in 640×1024640\times 1024 resolution (Cosmos-1.0-Autoregressive-4B in Sec.˜5.2.3) with this low-resolution tokenizer on videos of 320×512320\times 512 resolution from the target Physical AI domain. Finally, we add the Medusa heads to the fine-tuned low-resolution autoregressive WFM.

We conducted inference benchmarking on 8 ×\times H100 GPUs using torch.compile’s “max-autotune” mode in BF16 precision, and evaluated with 10-FPS input videos from the target Physical AI domain. In Tab.˜17, we report the average token throughput and frame generation speed achieved in this setup. We observe that our model can generate 10 video frames in less than 1 second, demonstrating that we can achieve real-time video generation at 10 FPS.

2.5 Diffusion Decoder

Our Cosmos tokenizer uses a lightweight encoder-decoder architecture to perform aggressive compression, which reduces the number of tokens for our WFM training. As a result of aggressive compression, it could sometimes lead to blurriness and visible artifacts in video generation, especially in the autoregressive WFM setting, where only a few integers are used to represent a rich video through discrete tokenization. We resort to the diffusion decoder design (Ramesh et al., 2022; OpenAI, 2024a) to address the limitation. Specifically, we build a more powerful tokenizer decoder by fine-tuning Cosmos-1.0-Diffusion-7B-Text2Video in Sec.˜5.1.

Fig.˜16 illustrates how we train a diffusion decoder for our autoregressive WFMs. For each training video, we use Cosmos-1.0-Tokenizer-CV8x8x8 and Cosmos-1.0-Tokenizer-DV8x16x16 to compute a continuous token video and a corresponding discrete token video, respectively. We note that Cosmos-1.0-Tokenizer-CV8x8x8 can produce higher quality video outputs than Cosmos-1.0-Tokenizer-DV8x16x16 thanks to the more gentle continuous tokenization process and the less aggressive compression scheme (8×8×88\times 8\times 8 instead of 8×16×168\times 16\times 16).

The discrete token video is treated as the conditional input to the denoiser of the Cosmos-1.0-Diffusion-7B model. To compute the conditional input, we first embed each discrete token of the discrete token video into a 16-dimensional vector based on a learnable vocabulary embedding layer. We then upsample the embedding 2×2\times along the xx and yy directions so that the conditional input will be of the same size as the noisy input to the denoiser from the continuous token video. We concatenate the noisy continuous inputs with the conditional inputs along the channel dimension, which becomes the input to the diffusion denoiser. The first layer of the denoiser is channel-dimension expanded to accommodate the new input shape. We fine-tune the updated Cosmos-1.0-Diffusion-7B by removing the added noise. As the discrete token video is not noise-corrupted, the denoiser learns to leverage the residing information in the conditional input for denoising. The result is a higher-quality decoder for the tokenizer that decodes the discrete token by solving a reserve diffusion problem.

Fig.˜16 illustrates the inference. The output discrete token video (under 8×16×168\times 16\times 16 discrete compression) from our autoregressive WFM is decoded into a video through two steps. First, we roll out the conditional denoiser to generate a continuous token video (under 8×8×88\times 8\times 8 continuous compression) based on the autoregressive WFM output. Next, the continuous token video is decoded by Cosmos-1.0-Tokenizer-CV8x8x8 to produce the resulting RGB video.

2.6 Results

In Fig.˜17, we show qualitative results of our autoregressive WFMs using different model sizes. In the unprompted setting, comparing Cosmos-1.0-Autoregressive-4B and Cosmos-1.0-Autoregressive-12B model, we observe that the 12B model generates videos with better motion and sharper details. Similarly, in the prompted setting, comparing Cosmos-1.0-Autoregressive-5B-Video2World and Cosmos-1.0-Autoregressive-13B-Video2World reveals that the 13B model gets better motion than the 5B model.

In Fig.˜18, we show the enhancements obtained when using the diffusion decoder. The outputs of the autoregressive model are blurry mainly due to the lossy compression in our discrete tokenizer. The use of the diffusion decoder can enhance details while preserving the content.

We empirically find the outputs of the autoregressive-based Text2World WFMs do not improve with upsampled prompts from the prompt upsampler discussed in Sec.˜5.1.5. We hypothesize this is possibly due to the fact that these WFMs are pretrained with pure video generation tasks for most of the training. They are not forced hard enough to leverage text inputs.

2.7 Limitations

One notable failure case observed in the generated videos of our autoregressive WFMs is objects unexpectedly appearing from below. Fig.˜19 illustrates an example of this issue. To understand the failure rate of our models, we conduct a systematic study by creating an evaluation set of 100100 Physical AI inputs to our autoregressive WFMs. We generate videos with all our models using two input modes—image (single-frame) conditioning and video (9-frame) conditioning. For all generated videos, we manually inspect the failure cases and report the failure rate in Tab.˜18. We observe that the smaller models Cosmos-1.0-Autoregressive-4B and Cosmos-1.0-Autoregressive-5B-Video2World show a higher corruption rate in single frame conditioning, while the larger models Cosmos-1.0-Autoregressive-12B and Cosmos-1.0-Autoregressive-13B-Video2World are more robust. Generation with 99-frame video conditioning is stable for all models, with a failure rate lower than 2%2\%.

3 Evaluation

Pre-trained WFMs are generalists of visual world simulation. Their capabilities should be measured across multiple aspects. Here, we evaluate our models on two aspects. First, we evaluate the 3D consistency of the generated videos. An ideal WFM should generate video simulations from geometrically plausible 3D worlds. Second, we evaluate the physics alignment of the generated videos. We calculate how well the rendered dynamics adhere to the laws of physics. Evaluation of WFMs is a highly nontrivial task. We acknowledge that there are several other important aspects required for evaluation. We leave a more comprehensive evaluation as future work.

WFMs are designed to simulate 3D worlds through video generation, and it is essential to evaluate how well the generated videos are consistent with the 3D structure of the visual world. In addition to appearing realistic, the generated videos should maintain coherence with the physical principles of scenes through time, a key requirement for downstream Physical AI applications.

Test data and baseline model. We focus on the scenario of static scenes in order to effectively measure 3D consistency of videos with existing tools based on multi-view geometry. We curate a dataset of 500 videos randomly chosen from the test set of the RealEstate10K dataset (Zhou et al., 2018). We additionally caption the videos using a proprietary VLM to obtain text prompts that describe the videos as static scenes, so one does not need to consider scene motions for metric computation. We compare against VideoLDM (Blattmann et al., 2023b) as the baseline method.

Metrics. Generated videos are effectively 2D projections of the underlying 3D visual worlds. We design the following metrics to measure the 3D consistency of generated videos.

Geometric consistency. We evaluate the 3D consistency of our generated worlds by quantifying how the epipolar geometry constraints are satisfied, including the Sampson error (Sampson, 1982; Hartley and Zisserman, 2003) and the success rate of camera pose estimation algorithms (Schönberger and Frahm, 2016; Schönberger et al., 2016) on the generated videos.

View synthesis consistency. We evaluate the ability of world foundation models to synthesize images at interpolated novel viewpoints while maintaining coherence with the underlying 3D structure.

The Sampson error is the first-order approximation of the distance from one interest point to its corresponding epipolar line in another view. Given NN point correspondences (represented in homogeneous coordinates) {(xˉi,yˉi)}i=1N\{\left(\bar{\mathbf{x}}_{i},\bar{\mathbf{y}}_{i}\right)\}_{i=1}^{N} in a given frame pair, we define the Sampson error as

and F\mathbf{F} is the fundamental matrix estimated from the correspondences. We use the square root version of the error function to make the metric more intuitive in pixel units. We use a combination of SuperPoint (DeTone et al., 2018) and LightGlue (Lindenberger et al., 2023) to detect and match keypoint correspondences from a frame pair and estimate F\mathbf{F} using OpenCV’s 8-point RANSAC algorithm. We normalize the average error by the diagonal length of the frame with respect to a 960×540960\times 540 canvas.

We also evaluate 3D consistency of a generated video with its ability to self-synthesize novel viewpoints. Following the common practice of novel view synthesis literature (Mildenhall et al., 2020), we hold out every 8 frames as the test frames and fit a 3D Gaussian splatting model (Kerbl et al., 2023) with the rest of the training frames using the default settings from the Nerfstudio library (Tancik et al., 2023). We report the Peak Signal-to-Noise Ratio (PSNR), Structural Similarity (SSIM), and LPIPS (Zhang et al., 2018) as the metrics to quantify the quality of the synthesized test views.

Results. We present the quantitative evaluation results in Tab.˜19. The Cosmos WFMs achieve significantly better 3D consistency than our baseline model in terms of both geometric and view synthesis consistency. Not only are the interest points from Cosmos WFMs more 3D-consistent, but the camera pose estimation success rate is also notably higher, reflecting both improved overall quality and enhanced 3D consistency, even reaching the level of real-world videos. Among the cases where camera poses were successfully estimated, the synthesized held-out views demonstrate higher quality across all image synthesis metrics. These results highlight the capability of our Cosmos WFMs to generate 3D-consistent videos, establishing them as effective world simulators.

3.2 Physics Alignment

An ideal WFM should exhibit a strong understanding of the laws of physics and produce future observations that respect them. While our pre-trained WFMs exhibit a certain level of physics understanding and advance the state-of-the-art, one can still easily generate examples that do not obey the law of physics. We believe additional steps in data curation where physically implausible videos are removed are required, as well as improved model design. While we leave a strong physics-aligned WFM as future work, we are still interested in measuring how much intuitive physics naturally emerges from large-scale data-driven pre-training.

To explore this, we design a controlled benchmark dataset using a physics simulation engine, taking inspiration from (Kang et al., 2024). We generate physics-grounded simulations to test the adherence of our pre-trained WFMs to Newtonian physics and rigid body dynamics. Specifically, we use simulation to generate physically correct photorealistic videos of test scenarios specific to physical laws of interest. These reference “ground truth” videos are then compared with “predicted” videos produced by a WFM given shared context (past observations and perturbation).

Synthetic data generation. Using PhysX (NVIDIA, 2024c) and Isaac Sim (NVIDIA, 2024a), we design eight 3D scenarios aimed at evaluating different physical effects:

Free-falling object(s): objects dropping on a plane (gravity, collision, \etc)

Tilted planar slope: objects rolling down an incline (gravity, moment of inertia, \etc)

U-shaped slope: objects rolling down a U-shaped slope (potential, kinetic energy, \etc)

Stable stack: a stack of objects in equilibrium (balanced forces)

Unstable stack: a stack of objects in imbalance (gravity, collision, \etc)

Dominoes: sequence of rectangular bricks falling in sequence (transfer of momentum, collision, \etc)

Seesaw: objects on either side of a seesaw (torque, rotational inertia, \etc)

Gyroscope: a spinning top on a flat surface (angular momentum, precession, \etc)

For each scenario, we randomize the number and type of dynamic objects (varying sizes, textures, shapes), selecting from Omniverse assets (NVIDIA, 2024b), as well as the background appearance. We simulate the kinematic state of objects over time and render the output videos from 4 different static camera views. In total, we render 800 1080p videos of 100 frames in length. The objects in each simulation roll-out are positioned so that they are all visible from the first frame to avoid any existence ambiguity.

Metrics. We are interested in assessing the adherence to physical laws by comparing the simulated ground-truth video to the output directly generated by the WFM. Therefore, to produce future observations, we condition our WFMs on the first few frames (either 1 or 9 frames) of the ground truth video. When applicable, we additionally condition a WFM on a text prompt (obtained using a proprietary VLM by captioning the conditioning frames), focusing on the kinematic state of the objects being simulated in the past observations. Please refer to Fig.˜20 for some examples of simulated versus predicted scenarios. For evaluation, we use the following metrics:

Pixel-level metrics. For a pixel-level comparison, we compute the Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM) to compare a predicted frame from the WFM rollout with the reference frame from the ground truth video.

Feature-level metrics. For a slightly higher-level semantic comparison, we calculate DreamSim similarity scores (Fu et al., 2023), a feature similarity metric, between the predicted and reference frames.

Object-level metrics. Finally, since we care most about how objects of interest are impacted by the ongoing physical phenomenon, we use tracking to compute object-level metrics that eliminate confounders (background changes, visual quality, \etc). Since the test conditions are synthetically generated, we have access to the ground-truth instance segmentation masks of the dynamic objects in the scenes. Using SAMURAI (Yang et al., 2024b), we propagate the ground-truth instance masks in the first frame through the rest of the predicted video frames to extract tracks, allowing us to quantify object-level metrics. We compute the intersection-over-union (IoU) between ground truth and predicted object masks for each frame and object of interest.

We average these metrics across frames in a video, across videos in the evaluation set, and across four random seeds for rollouts. PSNR and SSIM are computed on all frames, excluding the ones used for conditioning.

Results. Quantitative results on physical alignment are outlined in Tab.˜20. Based on quantitative and qualitative results, we make the following observations. Unsurprisingly, the models are able to better predict the overall object kinematics with more frames as conditioning input (which allows us to better infer 1st and 2nd order quantities such as speed and acceleration).

From the table, we also find that our diffusion WFMs perform better in pixel-level prediction than our autoregressive WFMs on the 9-frame conditional setting. This correlates with our visual observation that the diffusion-based WFMs render videos with higher visual quality. We also note that our results do not suggest that the larger model performs better on our physics alignment. While we observe larger models render videos with higher visual quality, all the WFMs equally struggle with physics adherence and require better data curation and model design.

More generally, we observe that the rigid-body simulations described above already test the limits of our WFMs, serving as valuable tools for identifying specific failure cases. These range from low-level issues like object impermanence (spontaneous appearance and disappearance of objects) and deformation (shape changes) to more complex problems such as implausible kinematics, violation of gravity, \etc. We believe such structured simulations offer a useful methodology to test physics alignment. We, therefore, intend to improve them over time by incorporating more complex scenarios, enhancing photorealism to bridge the sim-to-real gap (since WFM pre-training data consists of real videos), and refining our evaluation metrics for a more comprehensive assessment of physical understanding.

Post-trained World Foundation Model

In this section, we demonstrate how our Cosmos WFMs can be fine-tuned to support diverse Physical AI applications. We include examples from post-training our WFM with camera control to achieve 3D navigable visual world generation, post-training our WFM with action control on two different robotic setups for two different robotic manipulation tasks, and post-training our WFM with multi-view support for training autonomous driving agents.

Tab.˜21 provides a list of the discussed post-trained WFMs in different subsections of this section. We also list the conditional inputs to highlight the operation mode. Note that for each model, we add “-Sample“ to emphasize our goal is to provide sample applications of our pre-trained WFMs. Those models are by no means a complete system or a production model for any real-world applications. The developer would need to fine-tune the WFMs on their custom datasets for their Physical AI setups for their target applications.

Through camera pose conditioning, we integrate camera control into Cosmos-1.0-Diffusion-7B-Video2World, making it an effective 3D world simulator. We term the result post-trained WFM as Cosmos-1.0-Diffusion-7B-Video2World-Sample-CameraCond. We focus on generating 3D worlds from a single reference input image, leveraging camera control to produce temporally coherent and 3D-consistent video simulations from the specified camera trajectories, where changes in perspective align with the underlying 3D structure of the scene.

We use DL3DV-10K (Ling et al., 2024), a large-scale video dataset of static scenes, for this task. As a preprocessing step, we chunk all videos into clips with 256 frames. To obtain camera pose annotations densely for all frames within a clip, we run structure-from-motion on the chunked clips using GLOMAP (Pan et al., 2025). We set the camera pose of the first frame to be the identity transform and compute the relative camera poses for all subsequent frames. We also use a proprietary VLM to caption the videos to obtain text prompts that describe the videos as static scenes.

1.2 Fine-tuning

We add camera control conditioning by concatenating the sampled latent embeddings with Plücker embeddings (Sitzmann et al., 2021), which has the same spatial dimensions as the latent embeddings. Specifically, given the camera pose, we compute the Plücker coordinates via

where c\mathbf{c} is the camera center location and d\mathbf{d} is the unit ray direction of each latent pixel (where the latent embedding is treated as a downsampled image). All the camera poses are relative with respect to the initial frame. The Cosmos-1.0-Tokenizer-CV8x8x8 used by Cosmos-1.0-Diffusion-7B-Video2World models has a temporal compression rate of 8×8\times, and thus for every 8 frames, we use the Plücker embedding at the 4th frame to concatenate with the corresponding latent representation.

We resized the input frames of our training videos to 704×1252704\times 1252 and padded them to 704×1280704\times 1280 with reflection. We sample 57 frames during training. The training objective and other hyper-parameters are the same as the base Diffusion WFM training (Sec.˜5.1.3).

1.3 Evaluation

We assume a single reference image of the world is given and generate the future rollout as a video from the input image. We compare against CamCo (Xu et al., 2024), the state-of-the-art model for camera-controllable video generation under this setup. For a fair comparison, we use the CamCo model that was also fine-tuned on the DL3DV-10K (Ling et al., 2024) training set. As our post-trained WFM generates 57 frames and CamCo can only generate 14 frames, we compare the same 57-frame trajectories where we temporally downsample by 4×4\times for CamCo. The video resolution from CamCo is limited to 256×256256\times 256. We additionally maximally center-crop the input image and test frames for evaluation.

For the test data, we use the same 500 samples from the RealEstate10K (Zhou et al., 2018) test set previously described in Sec.˜5.3.1. We use the initial frame as the reference image and camera trajectories provided by the dataset as the camera control input, which we additionally rescale such that the distance between two ends of trajectories is normalized to 1.

Metrics. Following Xu et al. (2024), we evaluate the camera controllability of the post-trained world model in two aspects: video generation quality and 3D consistency. For video quality, we use the Fréchet Inception Distance (FID) (Heusel et al., 2017) and the Fréchet Video Distance (FVD) (Unterthiner et al., 2019) to assess the qualities at the frame and video levels, respectively. We use the same test data as the reference videos to compute the metrics (note that they are not used for pixel-level comparisons).

For 3D consistency, we evaluate via the ability of structure-from-motion (Schönberger and Frahm, 2016; Schönberger et al., 2016; Pan et al., 2025) libraries to re-estimate the camera poses, and we compare the results against the input camera control trajectories. Given NN frames in the video, we quantify the camera trajectory error into two terms: the average rotation error ϵrot\epsilon_{\text{rot}} and translation error ϵtrans\epsilon_{\text{trans}}, defined respectively as

where Ri\mathbf{R}_{i} and ti\mathbf{t}_{i} are the input rotation and translation of the ii-th frame (serving as ground truth), and R^i\mathbf{\hat{R}}_{i} and t^i\mathbf{\hat{t}}_{i} are the re-estimated quantities. To account for ambiguities from camera pose estimation results up to a similarity transformation, we follow Lin et al. (2021) and run Procrustes analysis on the predicted camera trajectories to align against the ground truth.

Comparisons. We present the results in Tab.˜22. First, our post-trained WFM can generate realistic and coherent 3D worlds. This is evidenced by the lower FID/FVD scores (higher visual quality) and the higher camera pose estimation success rate. Cosmos-1.0-Diffusion-7B-Video2World-Sample-CameraCond demonstrates better camera control, as the camera trajectory re-estimation is significantly closer to the original control input.

We also provide visual comparisons in Fig.˜21. While CamCo struggles to generate content beyond the input image, Cosmos-1.0-Diffusion-7B-Video2World-Sample-CameraCond effectively generates visuals that adhere to the structure of a 3D world. Note that both models were post-trained on DL3DV-10K and evaluated on the RealEstate10K dataset, which introduces a significant distribution shift between training and testing. The Cosmos model successfully overcomes this distribution shift while also demonstrating its capability to generalize to unseen input camera trajectories.

Qualitative results. Fig.˜22 shows our results from joystick-like control input on the camera, including moving forward, moving backward, rotating left, and rotating right. This demonstrates the use case where one can navigate the simulated world using a joystick to control the model in generating future video frames. A Physical AI agent could also use such control to predict the future of the world under different scenarios.

To show the diversity of the generation, we show generation results from the same input image and camera control with different random seeds in Fig.˜23. Cosmos-1.0-Diffusion-7B-Video2World-Sample-CameraCond is able to generate different worlds while still maintaining 3D spatial and temporal coherence in the videos. This could be used to simulate different possible futures given the current states.

2 Post-training WFM for Robotic Manipulation

A world model has the potential to serve as a powerful planner and simulator for robotic manipulation. Here, we demonstrate how we fine-tune our pre-trained WFMs for two tasks: (1) instruction-based video prediction and (2) action-based next-frame generation. For instruction-based video prediction, the input is the current video frame of a robot as well as a text instruction, and the output is a predicted video of the robot following the instruction. For action-based next-frame prediction, the input is the current video frame of a robot as well as an action vector between the current and next frame, and the output is the predicted next frame showing the result of the robot performing the specified action. Given a sequence of actions, the model can be run autoregressively to predict a video of the robot executing the given actions.

We curate two datasets for the two tasks described above. For instruction-based video prediction, we created an internal dataset called the Cosmos-1X dataset. It comprises approximately 200 hours of egocentric videos captured by EVE, a humanoid robot from 1x.Tech (Technologies, 2024) performing a variety of tasks, including navigation, folding clothes, cleaning tables, picking up objects, \etc. From the raw videos, we selected approximately 12,00012{,}000 episodes ranging from 1 to 9 seconds. Each episode is labeled with a one-sentence instruction, which is later upsampled with a proprietary VLM. The videos are captured at 30 FPS with a resolution of 512×512512\times 512.

For action-based next-frame generation, we used a public dataset called Bridge (Ebert et al., 2022), with the same configuration as a prior work (Zhu et al., 2024) for comparison. The Bridge dataset includes approximately 20,00020{,}000 episodes of third-person views of a robot arm performing different tasks in a kitchen environment, with videos of 320×256320\times 256 resolution captured at 5 FPS. For each video frame, the corresponding action is defined as a 7-dimensional vector in the gripper coordinate space (Δx,Δy,Δz,Δθr,Δθp,Δθy,Δ\mboxGripper)(\Delta x,\Delta y,\Delta z,\Delta\theta_{r},\Delta\theta_{p},\Delta\theta_{y},\Delta\mbox{Gripper}) as in OpenVLA (Kim et al., 2024).

2.2 Fine-tuning

We fine-tune both our Cosmos-1.0-Diffusion-7B-Video2World (Sec.˜5.1) and Cosmos-1.0-Autoregressive-5B-Video2World (Sec.˜5.2) for instruction-based video prediction and action-based next-frame prediction tasks.

For instruction-based video prediction, we build two models based on the base WFMs. The first is called Cosmos-1.0-Diffusion-7B-Video2World-Sample-Instruction, and the second is called Cosmos-1.0-Autoregressive-5B-Video2World-Sample-Instruction. We compute the T5 embedding of the instruction, which is added to the finetuing of the base model via cross-attention.

For action-based next-frame prediction, we also build two models based on the base WFMs. The first one is called Cosmos-1.0-Diffusion-7B-Video2World-Sample-ActionCond, and the second one is called Cosmos-1.0-Autoregressive-5B-Video2World-Sample-ActionCond.

Since action is a new modality not encountered during pre-training, we introduce additional modules inside our models for conditioning. For Cosmos-1.0-Autoregress-5B-Video2World-Sample-ActionCond, we add an action embedder MLP to project the action vector into a tensor, which is then incorporated into the model via cross-attention. For Cosmos-1.0-Diffusion-7B-Video2World-Sample-ActionCond, we also add an action embedder MLP to predict the action into a tensor but instead, incorporate it into the model by adding it to the timestamp embedding of the DiT modules.

2.3 Evaluation

For instruction-based video prediction, we fine-tune VideoLDM (Blattmann et al., 2023b) on the Cosmos-1X dataset and obtained VideoLDM-Instruction as a baseline for comparison. To evaluate the video generation performance of the models, we define the following dimensions:

Instruction following: Is the generated video aligned with the input language instruction?

Object permanence: Do objects present in the scene remain throughout the generated video?

Verity: Does the generated video faithfully represent the real world without unexpected imaginary objects?

Overall: Is the generated video reasonable for the robot to plan accordingly?

Human evaluators are tasked to observe a pair of anonymous videos generated by different models but with the same language instruction and compare them along the dimensions listed above. A group of ten human evaluators performed the evaluation over 23 test episodes. The statistical results are summarized in Fig.˜24.

As shown, we find that both Cosmos-1.0-Diffusion-7B-Video2World-Sample-Instruction and Cosmos-1.0-Autoregressive-5B-Video2World-Sample-Instruction perform better than VideoLDM-Instruction along the four evaluation dimensions. Cosmos-1.0-Diffusion-7B-Video2World-Sample-Instruction achieved 78.3%78.3\% overall preference compared to 13.0%13.0\% for VideoLDM-Instruction. Cosmos-1.0-Autoregressive-5B-Video2World-Sample-Instruction has also achieved better performance than diffusion-based VideoLDM-Instruction. Some predicted video frames for both fine-tuned WFMs are presented in Fig.˜25, which shows the quality of the predicted videos.

For action-based next frame prediction, we fine-tuned our models on the Bridge dataset. As a baseline, we fine-tune IRASim (Zhu et al., 2024) to derive an action-based next-frame prediction model IRASim-Action. We perform the next-frame prediction autoregressively to generate videos. To evaluate video generation quality, we compare the generated videos against ground truth videos over 100 episodes randomly selected from the official Bridge test set.

The computed metrics are summarized in Tab.˜23, including PSNR, SSIM, Latent L2 (Zhu et al., 2024), and FVD. As shown, both Cosmos-1.0-Autoregressive-5B-Video2World-Sample-ActionCond and Cosmos-1.0-Diffusion-7B-Video2World-Sample-ActionCond models outperform the baseline model (IRASim-Action). Some predicted video frames are presented in Fig.˜26, which shows the quality of the predicted videos compared to the ground truth.

3 Post-training WFM for Autonomous Driving

A world model for in-the-wild driving scenes has the potential to serve as a powerful simulation engine for training autonomous driving agents. As most autonomous vehicles are equipped with multiple cameras viewing different directions, an ideal world model for an autonomous vehicle should also be a multi-view one, preferably matching the precise setup of the sensors in the target vehicle. Here, we demonstrate how we fine-tune our pre-trained WFM to create a multi-view world model for autonomous driving tasks.

We curate an internal dataset called the Real Driving Scene (RDS) dataset. It comprises approximately 3.6 million 20-second surround-view video clips (equivalent to approximately 20,00020{,}000 hours of data) captured using an NVIDIA internal driving platform. Each clip is recorded from six camera views: front, left, right, rear, rear-left, and rear-right. In addition, the dataset includes ego-motion information that we use to construct the trajectory data. We use the recorded timestamps of the front camera video to synchronize the frames of all other views.

This dataset was selected from a large labeled data corpus to match a target distribution of data attributes. The specific attribute tags include:

Contender vehicle density (\eg, none, low, medium, high)

Weather (\eg, clear, raining, snowing, fog)

Ego vehicle speed (\eg, standing, low, local, highway speeds)

Ego vehicle behavior (\eg, high, medium, low curvature trajectories and accelerations)

Road type/population density (based on OpenStreetMap definitions: rural, residential, urban).

Additionally, the dataset was augmented through a second data-mining run to ensure a minimum number of clips containing rare road structures (\eg, tollbooths, bridges, tunnels, speed bumps, \etc). Finally, videos from each camera view are captioned separately, starting with a template text string: “The video is captured from a camera mounted on a car. The camera is facing forward|left|right|backward|rear-left|rear-right.”

3.2 Fine-tuning

We fine-tune our Cosmos-1.0-Diffusion-7B-Text2World (Sec.˜5.1) into a multiple-view world model using the RDS dataset. To ensure consistent video generation across multiple views, we slightly modify the architectural design described in Sec.˜5.1 and fine-tune the WFM to generate videos from all six cameras simultaneously.

We build three multi-view world models, summarized in Tab.˜21. The first one is called Cosmos-1.0-Diffusion-7B-Text2World-Sample-MultiView, which is a multi-view world model that can generate six camera views based on a text prompt input. The second one is called Cosmos-1.0-Diffusion-7B-Text2World-Sample-MultiView-TrajectoryCond. This model is built on top of Cosmos-1.0-Diffusion-7B-Text2World-Sample-AV-MultiView and takes an additional trajectory input as the conditional input signal. The final model, Cosmos-1.0-Diffusion-7B-Video2World-Sample-MultiView, is fine-tuned from the Diffusion-7B-Video2World-Sample-MultiView model to support video-based conditioning. It achieves this by incorporating previous frames into the generation process. Cosmos-1.0-Diffusion-7B-Video2World-Sample-MultiView can take the video output from Cosmos-1.0-Diffusion-7B-Text2World-Sample-MultiView and generate its extension. All three models output 6 views of 57 frames of video at a resolution of 848×480848\times 480.

View-independent positional embedding and view embedding. Instead of extending the FPS-aware 3D RoPE Positional Embedding to include an additional view dimension, we opt to use the same positional embedding described in Sec.˜5.1 independently to each view. To represent view differences, we modify the denoising function DθD_{\theta} to take an additional view embedding as input. That is, the camera view information is supplied through global view embeddings instead of positional embedding.

View-dependent cross-attention. In our multi-view setting, each of the six views of the same scene would have a different video description. While we treat all six views as a whole as the state of the diffusion process and perform self-attention among all the elements in the six views for denoising, we find it beneficial to employ view-dependent cross-attention for textual inputs. Specifically, the cross-attention operation for each view only attends to the textual description for the specific view. Note that each view has a different video description in our dataset. With the view embedding and view-dependent cross-attention, we derive Cosmos-1.0-Diffusion-7B-Text2World-Sample-MultiView from fine-tuning Cosmos-1.0-Diffusion-7B-Text2World.

Trajectory control condition. Optionally, in addition to the text condition, we fine-tune the model to produce videos that conform to the given future trajectory paths to enable more precise control of the agent. This enables the generation of unique driving scenarios that adhere to both driving trajectories recorded by real-world data and driving environments specified by the input text descriptions. The fine-tuned model is Cosmos-1.0-Diffusion-7B-Text2World-Sample-MultiView-TrajectoryCond.

We define a trajectory as a sequence of 64 points in the 3D space, representing a sequence of translations of the agent from the initial position (0,0,0)(0,0,0) to the final destination, with each point separated by a 0.1-second interval. We compute the embedding of the trajectory input and make the result a conditional input to the denoiser of the fine-tuned Cosmos-1.0-Diffusion-7B-Video2World model. We note that it is possible to achieve more fine-grained control signals by giving a per-interval action vector, following prior works (Kim et al., 2020, 2021; Hu et al., 2023) or as in the robotic manipulation task (Sec.˜6.2). We leave such extensions for future work.

3.3 Evaluation

We first present text-conditioned qualitative results in Fig.˜27. Using Cosmos-1.0-Diffusion-7B-Text2World-Sample-MultiView, we generate a 57-frame video with six views, which is then extended to 201 frames using the Cosmos-1.0-Diffusion-7B-Video2World-Sample-MultiView model. In Fig.˜28, we demonstrate how the pre-trained world model enhances generalization, enabling the generation of rare or out-of-domain scenes from the RDS dataset, such as driving on a river. Lastly, Fig.˜29 showcases the results from Cosmos-1.0-Diffusion-7B-Text2World-Sample-MultiView-TrajectoryCond, where the ego car accurately follows the input trajectory.

For quantitative results, as a baseline, we followed the same fine-tuning recipe to fine-tune VideoLDM (Blattmann et al., 2023b) to derive a multi-view world model called VideoLDM-MultiView. We use a set of evaluation metrics measuring video generation quality, multi-view consistency, and trajectory following accuracy. To evaluate video generation quality, we use 1000 samples to compute the scores. For consistency-related metrics, to better understand different models’ behaviors under different scenarios, we categorize the ground-truth trajectories into four types: moving forward, turning left, turning right, and others (including static or complex movements). For each category, we gathered 200 samples and their corresponding prompts and conditions, totaling 800 samples. Below, we provide detailed descriptions of the metrics and the results.

Generation quality. We utilize Fréchet Inception Distance (FID) (Heusel et al., 2017) and Fréchet Video Distance (FVD) (Unterthiner et al., 2019) to measure the quality of the generated videos relative to the real ones. We first calculate a score per view by extracting 16 frames from each video. We then report the average score across all views per method. As shown in Tab.˜24, we find both Cosmos-1.0-Diffusion-7B-Text2World-Sample-MultiView and Cosmos-1.0-Diffusion-7B-Text2World-Sample-MultiView-TrajectoryCond significantly outperform VideoLDM-MultiView on both metrics, demonstrating the superior quality of our pre-trained 7B Diffusion-based WFM over the VideoLDM-MultiView baseline.

Multi-view consistency. We use an extended version of the Sampson error (Sampson, 1982; Hartley and Zisserman, 2003) formulated in Sec.˜5.3.1 to quantify the geometry consistency of the generated multi-view videos. As the ground-truth videos in our RDS dataset share similar fisheye camera intrinsic parameters, we use the median calibration to undistort the keypoints to a regular pinhole camera with a uniform size of 960×540960\times 540 and 120 degrees of horizontal FoV. Under this setting, two metrics are computed for the generated multi-view video:

Temporal Sampson Error (TSE) measures whether the content generated for each camera is consistent over time. It is the median Sampson error of adjacent frames for each of the views.

Cross-view Sampson Error (CSE) measures whether multi-view consistency is preserved over time. It is the Sampson error across different generated views averaged in time. The fundamental matrix used in CSE is estimated using the keypoints accumulated across all temporal frames.

As shown in Tab.˜24, we find that both Cosmos-1.0-Diffusion-7B-Text2World-Sample-MultiView and Cosmos-1.0-Diffusion-7B-Text2World-Sample-MultiView-TrajectoryCond render better multi-view geometry consistency over VideoLDM-MultiView. The overall geometric plausibility of the generated videos is much better for a world model fine-tuned from our WFM. We also note that, with the trajectory control condition, such consistency is further improved thanks to the explicit 3D guidance as Cosmos-1.0-Diffusion-7B-Text2World-Sample-MultiView-TrajectoryCond is ranked the best.

Trajectory consistency: Trajectory Agreement Error (TAE). We design a robust multi-view camera pose estimation pipeline similar to the one used in Liang et al. (2024) based on the formulation of Teed and Deng (2021). Such a pose estimation pipeline features an online dynamic mask generation module and a highly efficient dense bundle adjustment module, reaching a robust and real-time performance for estimating multi-view camera poses. We use this pipeline to estimate the camera poses of the front camera, separately using two multi-view camera configurations that consider “front + left-front” cameras and “front + right-front” cameras. We then calculate their trajectory errors to show their agreement, reflecting the consistency of the multi-view generation. Specifically, we compute the Absolute Trajectory Error (ATE) and Relative Pose Error for both the translational component (RPE-t) and the rotational component (RPE-R). We normalize the length of the trajectories to 1.0 for fair comparison and exclude the cases with minor camera movements (\eg, a car stopped at a red light).

As shown in Tab.˜25, the results echo the findings from the multi-view geometry consistency, where the trajectory consistency of the world models fine-tuned from the Cosmos WFM is much better than that from the VideoLDM-MultiView. We note that the post-trained Cosmos world models have trajectory consistency that is close to real-world videos.

Trajectory consistency: Trajectory Following Error (TFE). Furthermore, for the model where we have trajectory control conditions being fed into the model, we use the same camera pose estimation pipeline used above to compute the poses of the front camera using multi-view information and compare the predicted trajectory with respect to the ground truth trajectory condition. This measures how well the model follows the given trajectory path. As shown in Tab.˜25, the trajectory error estimated using the generated videos from our Cosmos post-trained world models is only <<7cm less precise than the ground-truth oracle. Such a considerably slight margin shows that our model is able to accurately follow the given trajectory path, which is crucial for training autonomous driving agents.

Objects tracking consistency. Finally, we applied object detection and tracking using YOLOv11x (Khanam and Hussain, 2024) on the generated 8-second videos. Human annotators were tasked with identifying instances where the tracking algorithm misinterpreted physically impossible scenarios, such as two distinct objects (\eg, a person and a car) merging incorrectly into a single tracked entity. To evaluate this, we provided annotators with a random sample of 20 generated videos containing 157 objects. Remarkably, none of the 157 objects exhibited any physically impossible scenarios, demonstrating the physical consistency and object permanence of our generated driving videos.

Guardrails

For the safe use of our WFMs, we develop a comprehensive guardrail system. It consists of two stages: the pre-Guard stage and the post-Guard stage. The pre-Guard stage leverages Aegis (Ghosh et al., 2024) and a keyword list to block harmful prompts. The post-Guard stage blocks harmful visual outputs using a video content safety classifier and a face blur filter. The pipeline is illustrated in Fig.˜30.

Our pre-Guard is a text-domain guardrail comprising an LLM-based guardrail for semantically complex prompts and a simple blocklist-based checker for explicitly unsafe keywords.

The blocklist heuristic will act as the first line of defense to mitigate the risk of generating unsafe content. This is designed to block explicitly harmful generations by doing a keyword search on the prompt against a hard-coded blocklist of a large corpus of explicit and objectionable words. Input words are lemmatized using WordNetLemmatizer, a tool that uses a lexical database of the English language (Miller, 1995) to extract the root word from its variants. For example, the root word of “abacii” is “abacus”. These lemmatized words are then compared to the words in the hard-coded blocklist, and the entire prompt is rejected if any profanity is found. We use a comprehensive set of keywords to maximally protect our users.

1.2 Aegis Guardrail

As the second line of defense, we use Aegis-AI-Content-Safety-LlamaGuard-LLM-Defensive-1.0 (Ghosh et al., 2024), which is a fine-tuned version of Llama-Guard (Inan et al., 2023) trained on NVIDIA’s Aegis Content Safety Dataset covering NVIDIA’s broad taxonomy of 13 critical safety risk categories. There are two versions of AEGIS 1.0, the defensive version and the permissive version. The defensive version adopts a tighter permission boundary than the permissive version. Cosmos uses the defensive version of Aegis to block potentially harmful user prompts that attempt to generate harmful content. If the input prompt is categorized as unsafe by this prompt filter, the video is not generated, and an error message is displayed.

For using Aegis as a prompt filter, we classify the prompt as unsafe if it falls into the following categories: violence, sexual, criminal planning, weapons, substance abuse, suicide, child sexual abuse material, hatred, harassment, threat, and profanity. Any prompt that does not fall into the above categories is considered safe from the prompt-filtering standpoint.

2 Post-Guard

Our post-Guard is a vision-domain guardrail comprising a video content safety filter and a face blur filter for the generated output.

The Video Content Safety Filter is a frame-level multi-class classifier trained on our video dataset and generation results. Among the classes, some are considered safe, while others are unsafe. A major challenge in training the classifier is in balancing false positives, where safe content is mistakenly flagged as unsafe, and false negatives, where unsafe content is wrongly classified as safe. To minimize classification errors, we carefully balanced the data during training.

We collect three kinds of ground truth annotated data. First, we sample a large set of videos from our dataset, extract frames, and determine its class using a VLM. Next, we generate synthetic videos with our WFMs using a set of prompts to ensure coverage of corner cases and least-represented content categories. Finally, human annotators provide the “gold standard” labels for a portion of our dataset, adding a vital layer of validation and helping us continuously refine the accuracy of our classifier. We extract the SigLIP (Zhai et al., 2023) embedding for each video frame and train a simple MLP classifier on the embeddings.

During inference, we generate a SigLIP embedding for every frame and then apply the classifier. The entire video is flagged as unsafe if any frame is classified as unsafe.

2.2 Face Blur Filter

We use RetinaFace (Deng et al., 2020), a state-of-the-art face detection model, to identify facial regions with high confidence scores. For any detected face region larger than 20×2020\times 20 pixels, we apply pixelation to obscure the regions while preserving the overall scene composition for Physical AI applications.

3 Red Team Effort

We employ a dedicated red team to actively probe the system using both standard and adversarial examples that are collected in an internal attack prompt dataset. These video outputs are annotated by a team of expert annotators, who were specially trained for our task, to classify the generated video on a scale of 1-5 on multiple categories of harm related to the taxonomy in Sec.˜7.1.2. These annotations also specify the start and end-frames where the unsafe content is detected, thereby generating high-quality annotations. The red team also probed each guardrail component independently with targeted examples to identify weaknesses and improve performance in edge cases. As of the date of publication, the red team has tested and annotated over 10,00010{,}000 distinct prompt-video pairs that were carefully crafted to cover a broad range of unsafe content.

Related Work

World models. The concept of “world models” originated from the seminal work of Ha and Schmidhuber (2018), which proposed learning a representation of the real world using neural network models to predict future states given current states and inputs. An accurate representation of the physical world model enables not only reliable prediction of future states but also informed decision-making. This concept of modeling the physical world is not new; traditional automation and robotics industries have long employed mathematical models based on physics laws and system identification in planning and control algorithms (Murray et al., 2017). However, these system-specific models, typically confined to low-dimensional state spaces, restrict generalization and knowledge transfer across different systems, limiting model reuse when applied to new tasks or environments. Recent advances in deep learning, particularly generative AI, have made it possible to learn world models directly from visual observations.

Modern world model pipelines can be categorized based on their backbone architecture. Most works (Hafner et al., 2019, 2021; Kim et al., 2020, 2021; Hafner et al., 2023; Hansen et al., 2024), including the original paper (Ha and Schmidhuber, 2018) by Ha and Schmidhuber, employ a recurrent neural network to model system state evolution in a latent space learned via an autoencoder. More recent trends view world models as generative models in visual observation space, often in the form of conditional video generative models (\eg, action-to-video, text-to-video). These models can be either autoregressive (Yang et al., 2023; Micheli et al., 2023; Robine et al., 2023; Bruce et al., 2024; Liu et al., 2024b) or diffusion-based (Valevski et al., 2024; Alonso et al., 2024; Ding et al., 2024), as considered in this work. Another promising approach is generative simulation (Nasiriany et al., 2024; Hua et al., 2024), which combines generative AI and physical simulators to model the real world.

A well-trained world model can be applied in various ways, including verification (Hu et al., 2023), planning-based model predictive control (Hansen et al., 2024; Bar et al., 2024), and model-based reinforcement learning (Yang et al., 2023; Robine et al., 2023; Alonso et al., 2024; Zhang et al., 2024; Ding et al., 2024). The effectiveness of world models has been demonstrated in domains such as computer games (Hafner et al., 2021; Kim et al., 2020; Bruce et al., 2024; Valevski et al., 2024; Alonso et al., 2024), real-world robots (Wu et al., 2023b; Yang et al., 2023), and autonomous driving (Kim et al., 2021; Blattmann et al., 2023b; Hu et al., 2023; Zhao et al., 2024a). We envision that foundational world models will have transformative impacts on these industries.

Video generative models. The field of video generative models has undergone rapid development in recent years. From the initial models that produced short, low-resolution videos, the field has evolved significantly, with video generative models now at the forefront of generative AI research (Ho et al., 2022; Huang et al., 2024). Recent years have seen the emergence of impressive video generative models, such as Sora, Dream Machine, Gen 3 and Kling, capable of producing realistic, high-resolution videos (Luma, 2024; OpenAI, 2024b; KuaiShou, 2024; Runway, 2024). These advancements have been made in just a few years since the release of the first video generative model.

Most existing work on video generative models focuses on text-to-video tasks, which generate videos based on text prompt inputs (Yang et al., 2024d; Ma et al., 2024; Lin et al., 2024a; Blattmann et al., 2023a; Ge et al., 2023; Girdhar et al., 2024). These models enable users to create impressive videos using carefully designed text prompts. Other popular tasks include Image-to-Video that generates videos starting from a given image frame (Blattmann et al., 2023a; Wang et al., 2024c; Ren et al., 2024; Wang et al., 2021b; Mallya et al., 2022; Gururani et al., 2023), Video-to-Video that generates new videos given a reference video (Wang et al., 2018, 2019; Mallya et al., 2020; Ku et al., 2024; Liu et al., 2024a) and Action-to-Video that generates videos based on actions driven by the development of world models and embodied AI (Bruce et al., 2024; Valevski et al., 2024; Alonso et al., 2024; Tulyakov et al., 2018).

The majority of video generative models adopt the diffusion model framework (Blattmann et al., 2023a; Lin et al., 2024a; Ge et al., 2023; Ma et al., 2024; Yang et al., 2024d) to gradually transform noise into video sequences. Autoregressive models have also been employed for video generation, offering the advantage of handling video and other modalities in a unified manner (Kondratyuk et al., 2024; Deng et al., 2024; Liu et al., 2024c). While autoregressive models have shown promise, diffusion-based video models still excel in terms of visual quality. Our goal is to help Physical AI developers advance their applications. We believe that the diffusion-based and autoregressive-based models both have their pros and cons. Diffusion-based models could render videos with better visual quality. Autoregressive-based models can better leverage all sorts of techniques developed by the LLM community. We build both diffusion-based (Cosmos-Diffusion) and autoregressive-based (Cosmos-Autoregressive) WFMs and make them available to the Physical AI builders.

Video generation with camera control. 3D-consistent video generation traces back to early works in view synthesis and 3D reconstruction, where the community sought to create 3D-consistent videos using neural rendering applied to various 3D representations (Zhou et al., 2018; Mildenhall et al., 2020; Wang et al., 2021a; Li et al., 2023; Kerbl et al., 2023). Within this line of research, single-image 3D view synthesis (Tucker and Snavely, 2020; Wiles et al., 2020; Yu et al., 2021; Lin et al., 2023b; Charatan et al., 2024) is particularly challenging, typically requiring the learning of a strong 3D prior model from multi-view image datasets. As such 3D prior models often suffer to scale well, learning-based view synthesis has also been explored through a purely data-driven approach using scalable Transformer architectures (Vaswani et al., 2017; Dosovitskiy et al., 2021). This bypasses the need for explicit 3D prior knowledge (Rombach et al., 2021; Sajjadi et al., 2022): instead of relying on neural rendering applied to 3D representations, novel views are synthesized directly by neural networks conditioned on camera inputs (Tatarchenko et al., 2016). This paradigm has been successfully scaled up with diffusion models (Liu et al., 2023b), finding broad applications in 3D asset generation (Poole et al., 2023; Lin et al., 2023a; Shi et al., 2023; Qian et al., 2024; Li et al., 2024; NVIDIA, 2024d). Recent advances in video generation quality suggest the potential for achieving full 3D consistency through the scaling of training video data (Brooks et al., 2024). Camera controllability on such models has since become an active area of investigation (He et al., 2024a; Wang et al., 2024g; Xu et al., 2024) for its great potential applications to robotics and autonomous navigation.

Generative models for robotic control. Recent advances in deep generative models have sparked significant interest in their application to robotic control. Several approaches have emerged, with one line of work directly employing diffusion models as visuomotor policies, demonstrating substantial improvements in imitation learning in various robotic tasks (Chi et al., 2023; Wang et al., 2024f; Prasad et al., 2024; Ke et al., 2024). Two other threads more related to this work are the use of pre-trained image and video generation models as motion planners and the use of image and video data for generative pre-training. The generative motion planning approach (Ko et al., 2024; Zhou et al., 2024b; Finn and Levine, 2017; Black et al., 2023; Du et al., 2024) aims to enhance generalization to unseen environments by generating intermediate visual sub-goals rather than explicit action sequences. This visual representation strategy proves more robust, as image and video sub-goals can generalize across diverse environmental setups, unlike action sequences that are typically environment- and task-specific. The generative pre-training approach (Gupta et al., 2024b; Cheang et al., 2024; He et al., 2024b) leverages large-scale image and video datasets for pre-training. While Gupta et al. (2024b) extract and utilize features from pre-trained text-to-image diffusion models to guide subsequent policy learning, Cheang et al. (2024) and He et al. (2024b) use a two-stage framework: first pre-training the model to predict future frames, then fine-tuning it to jointly predict both actions and future frames.

Generative models for autonomous driving. Video generative models have the potential to revolutionize autonomous driving simulation by enabling the generation of realistic driving videos conditioned on diverse input modalities, such as text, images, trajectories, 3D data, or maps (Kim et al., 2021; Blattmann et al., 2023b; Gao et al., 2024d; Wang et al., 2023a, 2024e; Lu et al., 2025; Jia et al., 2023; Yang et al., 2024c; Hu et al., 2023; Gao et al., 2024c, b). Despite their potential, existing approaches have been limited by constraints in data scale (Wang et al., 2023a, 2024e; Lu et al., 2025; Jia et al., 2023; Gao et al., 2024c, b), resolution (Yang et al., 2024c; Hu et al., 2023), and the number of camera views (Blattmann et al., 2023b; Gao et al., 2024d), restricting their effectiveness as comprehensive driving world simulators. To overcome these limitations, we leverage the capabilities of a powerful pre-trained WFM to develop a flexible and scalable driving simulator. Our model achieves high resolution, elevated frame rates, and multi-view consistency.

Tokenizer. There has been a fairly long history of learning latent features that reproduce the input visual data (Kingma, 2013; van den Oord et al., 2017; Hinton et al., 1995; He et al., 2022). Recently, such models, also known as tokenizers, have been widely incorporated as essential components to improve the efficiency of training large-scale generative models (Rombach et al., 2022; Esser et al., 2021).

Continuous visual tokenizers, often including Autoencoder (AE) and Variational Autoencoder (VAE), compress visual data into a continuous latent space where diffusion-based models can be efficiently trained on (Song et al., 2020; Lipman et al., 2022; Ho et al., 2020). At inference time, the generated latents are decoded back to RGB space with the tokenizer decoder. Various diffusion models have been trained in such a way for image (Rombach et al., 2022; Ramesh et al., 2022; Betker et al., 2023; Dai et al., 2023; Podell et al., 2024; FLUX, 2024; Gafni et al., 2022) and video generation (Zeng et al., 2024; Blattmann et al., 2023b, a; Ge et al., 2023; Brooks et al., 2024; Wang et al., 2023b; Yu et al., 2023b; An et al., 2023; Girdhar et al., 2024).

Discrete visual tokenizers additionally involve a quantizer (van den Oord et al., 2017; Zhao et al., 2024b; Mentzer et al., 2023; Yu et al., 2024a; Lee et al., 2022; Yu et al., 2024b) that further discretizes the continuous latents into a discrete space, allowing for easy integration into large language models (LLMs) and vision language models (VLMs) alongside other modalities, such as text and audio. Thus, discrete tokenizers are deployed in various visual understanding (Wu et al., 2024; Team, 2024a; Sun et al., 2024b; Wang et al., 2024d) as well as image (Esser et al., 2021; Ramesh et al., 2021; Yu et al., 2022; Chang et al., 2022; Sun et al., 2024a) and video generation tasks (Yan et al., 2021; Villegas et al., 2023; Hong et al., 2023; Wu et al., 2022; Ge et al., 2022; Yu et al., 2023a; Luo et al., 2024; Kondratyuk et al., 2024).

Cosmos tokenizers are extensively built based on previous studies, \eg, FSQ (Mentzer et al., 2023) and causal architecture (Yu et al., 2023a), with the goal of creating a suite of efficient and high-quality tokenizers.

Conclusions and Discussions

The Cosmos World Foundation Models mark a significant step towards building general-purpose simulators for the physical world. This work outlines our comprehensive approach, including the data curation pipeline, the design of continuous and discrete tokenizers, the architecture of diffusion and autoregressive world foundation models, and the fine-tuning process for diverse downstream Physical AI tasks. Notably, we demonstrate the adaptability of our pre-trained world models to critical applications, including 3D world navigation, robotic manipulation, and autonomous vehicle systems, which demand both 3D consistency and action controllability.

Limitations. Despite the progress, the development of world foundation models is still in the early stages. Current models, including ours, fall short as reliable simulators of the physical world. We observe that our models still suffer from issues, including the lack of object permanence, inaccuracies in contact-rich dynamics, and inconsistency in instruction following. Additionally, the realism of the generated videos does not always reflect adherence to fundamental physical principles, such as gravity, light interactions, and fluid dynamics.

Evaluation presents another significant challenge. Defining robust rubrics for humans to evaluate physical fidelity is hard as such assessments are often influenced by personal biases, backgrounds, and other subjective factors. Moreover, these evaluations may not align positively with metrics used in downstream Physical AI tasks. In order to address these challenges, promising directions include the development of automated evaluators powered by multi-modal LLMs and leveraging existing physical simulators to enable reproducible and interactive evaluation, thereby reducing dependence on human evaluation.

Autoregressive \vsDiffusion WFMs. Our evaluation results in 3D consistency (Sec.˜5.3.1) and video generation for robotics (Sec.˜6.2) indicate that diffusion-based WFMs currently deliver better generation quality. Through fine-tuning, diffusion-based WFMs are able to incorporate diverse control signals, including camera pose, end-effector positions, or autonomous vehicle trajectories, and generate outputs of novel formats like multi-view videos. However, autoregressive-based WFMs possess significant untapped potential. They could (1) leverage pre-trained weights from large language models (LLMs) to inherit extensive world knowledge and (2) enable faster generation through the use of advanced inference optimization techniques designed for causal attention. If these capabilities are fully realized, autoregressive WFMs may become particularly well-suited for applications requiring interactive control or real-time processing, such as planning and simulation in robotics. Importantly, the boundary between diffusion and autoregressive models is not rigid. Recent advancements have shown that diffusion transformers with bidirectional attention can be distilled into student transformers with causal attention, enabling support for key-value caching during inference (Yin et al., 2024). Similarly, autoregressive models can incorporate locally bidirectional attention to generate images via diffusion heads (Zhou et al., 2024a). Exploring these hybrid approaches and their trade-offs remains an active and promising area of research. We plan to investigate these formulations further and provide a comprehensive analysis in future work.

Appendix A Contributors and Acknowledgements

Data Curation Jacob Huffman, Francesco Ferroni, Alice Luo, Niket Agarwal, Hao Wang, Jing Zhang, David Page, Vasanth Rao Naik Sabavat, Sriharsha Niverty, Erik Barker, Lindsey Pavao, Stella Shi, Prithvijit Chattopadhyay, Shitao Tang, Yin Cui, Yunhao Ge, Qianli Ma, Yifan Ding, Seungjun Nah, Siddharth Gururani, Jiashu Xu, Grace Lam, Tiffany Cai, Jibin Varghese, Pooya Jannaty, Jay Zhangjie Wu, Yuxuan Zhang, Huan Ling, Hanzi Mao, Heng Wang

Tokenizer Jinwei Gu, Xian Liu, Songwei Ge, Ting-Chun Wang, Haoxiang Wang, Fitsum Reda

Diffusion-based World Foundation Model Pre-training Qinsheng Zhang, Lin Yen-Chen, Xiaohui Zeng, Huan Ling, Shitao Tang, Maciej Bala, Ting-Chun Wang, Yu Zeng, Seungjun Nah, Qianli Ma, Hanzi Mao

Autoregressive-based World Foundation Model Pre-training Haoxiang Wang, Yifan Ding, Xian Liu, Jiaojiao Fan, Xiaohui Zeng, Yogesh Balaji

Prompt Upsampler Yunhao Ge, Haoxiang Wang, Jiashu Xu, Yin Cui

Diffusion Decoder Huan Ling, Jiaojiao Fan, Fitsum Reda, Yogesh Balaji, Hanzi Mao, Qinsheng Zhang

3D Consistency Pre-training Evaluation Jiahui Huang, Chen-Hsuan Lin

Physics Alignment Pre-training Evaluation Francesco Ferroni, Prithvijit Chattopadhyay, Xinyue Wei, Qianli Ma, Gergely Klár, Chen-Hsuan Lin

Camera Control Post-training Evaluation Xiaohui Zeng, Tsung-Yi Lin, Jingyi Jin, Chen-Hsuan Lin

Robotics Post-training Evaluation Lin Yen-Chen, Wei-Cheng Tseng, Yunhao Ge, Xian Liu, Shitao Tang, Fangyin Wei, Lyne Tchapmi, Yu Zeng, Qingqing Zhao, Yin Cui, Zhaoshuo Li, Jinwei Gu

Autonomous Driving Post-training Evaluation Seung Wook Kim, Jay Zhangjie Wu, Jiahui Huang, Francesco Ferroni, Michele Fenzi, Daniel Dworakowski, Despoina Paschalidou, Ed Schmerling, Shiyi Lan, Laura Leal-Taixe, Sanja Fidler, Huan Ling

Guardrail Jibin Varghese, Arslan Ali, Grace Lam, Pooya Jannaty

A.2 Contributors

Anqi Li, Arsalan Mousavian, Artur Zolkowski, Bartosz Stefaniak, Dieter Fox, Ethan He, Kaichun Mo, Morteza Ramezanali, Przemek Tredak, Wei Yang, Xiaowei Ren, Yongxin Chen, Zeeshan Patel

A.3 Acknowledgments

We thank 1X Technologies for generously providing humanoid robot data and offering invaluable support for the post-training for robotic manipulation in this technical report.

We thank Aarti Basant, Akan Huang, Alex Qi, Alexis Bjorlin, Amanda Moran, Amol Fasale, Ankit Patel, Arash Vahdat, Aryaman Gupta, Ashna Khetan, Ashwath Aithal, Bor-Yiing Su, Bryan Catanzaro, Charles Hsu, Chris Pruett, Christopher Horvath, Clark Doan, Coulten Holt, Dane Aconfora, Deepak Narayanan, Dennis Chang, Dheeraj Kapur, Dong Ahn, Ebrar Erdem, Elmar Haussmann, Gandhi Vaithilingam, Henry Estela, Henry Vera, Herb Woodruff, Imad El Hanafi, Jashojit Mukherjee, Jason Sewall, Jensen Huang, John Dickinson, John Dickinson, Jonah Alben, Jonah Philion, Josh Abbott, Jun Gao, Kumar Anik, Lee Ditiangkin, Luke Alonso, Madison Huang, Marek Dabek, Mark Arnold, Max Ehrlich, Michele Ferretti, Misbah Mubarak, Misha Smelyanskiy, Mohamed Fawzy, Mohammad Harrim, Mohammad Shoeybi, Omkar Mehta, Pallab Bhattacharya, Paniz Karbasi, Pasha Shamis, Raju Wagwani, Rick Izzo, Robert Hero, Sharon Clay, Songyan Tang, Sophia Huang, Sridhar Bhuvanapalli, TJ Galda, Thomas Volk, Tobias Lasser, Vaibhav Ranglani, Vijay Anand Korthikanti, Yazdan Aghaghiri, Yugi Guvvala, and Zekun Hao for their feedback and engineering support.

We thank Iain Cunningham, Jim Fan, Marco Pavone, Meredith Price, Nikki Pope, Scott Reed, and Yuke Zhu for their feedback on the early draft of this technical report.

References