WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens

Xiaofeng Wang, Zheng Zhu, Guan Huang, Boyuan Wang, Xinze Chen, Jiwen Lu

Introduction

The next significant leap in artificial intelligence is expected to come from systems that possess a profound understanding of the dynamic visual world. At the core of this advancement are world models, crucial for comprehending and predicting the dynamic nature of our world. World models hold great promise for learning motion and physics in the general world, which is essential for video generation.

The early exploration of world models primarily focus on gaming scenarios, which proposes a generative neural network model capable of learning compressed representations of spatial and temporal dynamics within game environments. Subsequent research in the Dreamer series further validated the efficacy of world models across diverse gaming scenarios. Considering its structured nature and paramount importance, autonomous driving has become a forefront domain for the practical application of world models. Various approaches are introduced to explore the efficacy of world models in autonomous driving scenarios. Furthermore, DayDreamer has extended the application of world models to encompass real-world robotic environments, However, current world models are predominantly confined to gaming, robotics, and autonomous driving, lacking the capability to capture the motion and physics of the general world. Besides, relevant research in world models mainly relies on Recurrent Neural Networks (RNNs) and diffusion-based methods to model visual dynamics. While these approaches have yielded some success in video generation, they encounter challenges in effectively capturing the motion and physics in general world scenes.

In this paper, we introduce WorldDreamer, which pioneers the construction of general world models for video generation. Drawing inspiration from the successes of large language models (LLMs) , we predict the masked visual tokens to effectively model the intricate dynamics of motion and physics embedded in visual signals. Specifically, WorldDreamer involves encoding images into discrete tokens using VQGAN . We then randomly mask a portion of these tokens and utilize the unmasked tokens to predict the masked ones, a process integral to capturing the underlying motion and physics in visual data. WorldDreamer is constructed on the Transformer architecture . Regarding the spatial-temporal priority inherent in video signals, we propose the Spatial Temporal Patchwise Transformer (STPT), which enables attention to focus on localized patches within a temporal-spatial window, facilitating the learning of visual signal dynamics and accelerating the convergence of the training process. Additionally, WorldDreamer integrates language and action signals through cross-attention, to construct multi-modal prompts for interaction within world model. Notably, compared to diffusion-based methods, WorldDreamer capitalizes on the reuse of LLM infrastructure and benefits from optimizations developed over years for LLMs, including model scaling learning recipes. Besides, WorldDreamer exhibits a remarkable speed advantage, parallel decoding videos with just a few iterations, which is ∼\sim3×\times faster than diffusion-based methods . Therefore, WorldDreamer holds great promise for constructing a general world model from visual signals.

The main contributions of this paper can be summarized as follows: (1) We introduce WorldDreamer, the first general world model for video generation, which learns general world motion and physics. (2) We propose the Spatial Temporal Patchwise Transformer (STPT), which enhances the focus of attention on localized patches within a temporal-spatial window. This facilitates easier learning of visual signal dynamics and expedites the training process. (3) We conduct extensive experiments to verify that WorldDreamer excels in generating videos across different scenarios, including natural scenes and driving environments. WorldDreamer showcases versatility in executing tasks such as text-to-video conversion, image-to-video synthesis, video editing, and action-to-video generation (see Fig. 1).

Related Work

Currently, state-of-the-art video generation models are primarily classified into two categories: Transformer-based methods and diffusion-based methods.

Transformer-based methods. The Transformer-based video generation methods are derived from the general family of LLMs . Typically, these methods employ autoregressive prediction of the next token or parallel decoding of masked tokens to generate videos. Drawing inspiration from image generation techniques , VideoGPT integrates VQVAE with Transformer-based token prediction, enabling it to autoregressively predict visual tokens for video generation. Furthermore, GAIA-1 integrates various modalities, including text descriptions, images, and driving actions, resulting in the generation of autonomous driving scenario videos. Unlike these autoregressive methods, some Transformer-based approaches , draw inspiration from , accelerating video generation through parallel decoding. In addition to these methods, VideoPoet adopts video tokenizer and generates exceptionally high-quality videos based on parallel decoding. The incorporation of Transformer models into video language models showcases their formidable zero-shot capability in handling various tasks during pretraining. Therefore, employing Transformer-based mask image models as the foundation for general world models emerges as a promising avenue.

Diffusion based methods. Compared to Transformer-based models, there has been extensive research employing diffusion-based models for video generation. VideoLDM introduces a temporal dimension to the latent space of the 2D diffusion model and fine-tuned it using videos, effectively transforming the image generator into a video generator and enabling high-resolution video synthesis. Similarly, LVDM explores lightweight video diffusion models, making use of a low-dimensional 3D latent space. Make-A-Video also employs a pre-trained text-to-image model, eliminating the need for large-scale video training. Moreover, in the Imagen Video , a cascading video diffusion model is built upon the pretrained 2D diffusion model . DiffT and W.A.L.T improve the video generation by utilizing a Transformer-based Diffusion network. Recently, Emu Video and PixelDance propose a two-step factorization approach for text-to-video generation, wherein the process is initially decomposed into text-to-image conversion, followed by image-to-video synthesis. This methodology capitalizes on the effectiveness of contemporary text-to-image models, strategically directing the focus of the video diffusion model training toward the learning of motion dynamics. However, diffusion-based methods have difficulty integrating multiple modalities within a single model. Furthermore, these diffusion-based approaches struggle to produce results that accurately capture dynamics and motion.

2 World Models

World models play a pivotal role in comprehending and predicting the dynamic nature of our environment, holding immense potential for acquiring insights into motion and physics on a global scale. Initially, the exploration of world model focuses primarily on gaming scenarios, presenting a generative neural network model capable of learning condensed representations of spatial and temporal dynamics within game environments. Subsequent research within the Dreamer series affirmed the effectiveness of world models across a diverse array of gaming scenarios. Given its structured nature and critical significance, the domain of autonomous driving has emerged as a forefront application area for world models. Numerous approaches have been introduced to assess the efficacy of world models in autonomous driving scenarios. Additionally, DayDreamer has expanded the scope of world models to encompass real-world robotic environments. However, it is noteworthy that current world models primarily operate within the realms of gaming, robotics, and autonomous driving, lacking the capability to comprehensively capture the motion and physics of the general world.

WorldDreamer

The overall framework of WorldDreamer is depicted in Fig. 2. The initial phase involves encoding visual signals (i.e., images and videos) into discrete tokens using a visual tokenizer. These tokens undergo a carefully devised masking strategy before being processed by STPT. Meanwhile, textual and action signals are separately encoded into embeddings, which serve as multimodal prompts. STPT engages in the pivotal task of predicting the masked visual tokens, which are then decoded by visual decoders, facilitating video generation and editing in multiple contexts.

To train WorldDreamer, we construct triplets of Visual-Text-Action data, where the training supervision solely involves predicting masked visual tokens without any additional supervision signals. WorldDreamer also supports training without text or action data, which not only reduces the difficulty of data collection but also enables WorldDreamer to learn unconditional or single-condition video generation. At inference time, WorldDreamer can accomplish various video generation and video editing tasks: (1) For image-to-video, only a single image input is needed, considering the remaining frames as masked. WorldDreamer can also predict the future frames based on both single image condition and text condition. (2) For video stylization, a video segment can be input, with a random masking of certain pixels. WorldDreamer can alter the video style, such as creating an autumn-themed effect, based on both the input language. (3) For text-to-video, providing language input allows WorldDreamer to predict the corresponding video, assuming that all visual tokens are masked. (4) For video inpainting, a video segment can be input, with a manually masked region of interest. WorldDreamer can fill in the masked portion based on the input language and unmasked visual signals. (5) For action-to-video, inputting the initial frame of a driving scene along with future driving commands allows WorldDreamer to predict future frames.

The subsequent subsections elaborate on the model architecture and the masking strategy.

2 Model Architecture

Preliminery WorldDreamer utilizes VQGAN to tokenize visual signals:

where I∈RN×H×W×3I\in\mathcal{R}^{N\times H\times W\times 3} are NN frames of visual inputs. VQGAN Fv\mathcal{F}_{v} downsamples the resolution by 16×\times, which produces visual tokens TV∈RN×h×wT_{\text{V}}\in\mathcal{R}^{N\times h\times w} (h=H4,w=W4h=\frac{H}{4},w=\frac{W}{4}). The VQGAN has a vocabulary size of 8192 and is trained with billions of images . For text inputs, we employ the pretrained T5 to map them into high-dimensional embeddings ET∈RK×CTE_{\text{T}}\in\mathcal{R}^{K\times C_{\text{T}}}, where KK is the sequence length and CTC_{\text{T}} is the embedding channel. To be compatible with the feature learning in STPT, the text embeddings are repeated for NN frames, and the embedding channel is mapped into CVC_{\text{V}}. Furthermore, Multi Layer Perception (MLP) is utilized to encode action inputs, which generates action embeddings EA∈RN×CVE_{\text{A}}\in\mathcal{R^{N\times C_{\text{V}}}}. The text embeddings and action embeddings are concatenated, producing the multimodal prompt embeddings EM∈RN×(K+1)×CVE_{\text{M}}\in\mathcal{R}^{N\times(K+1)\times C_{\text{V}}}. Note that either text or action embedding can be empty, enabling unconditional learning.

During training, and the optimizing objective is to predict the masked visual token conditioned on unmasked tokens and multimodal prompts:

where T^V\hat{T}_{\text{V}} are masked visual tokens, and T~V\widetilde{T}_{\text{V}} are unmasked visual tokens.

STPT STPT leverages the foundation of U-ViT while strategically enhancing its architecture to better capture the intricacies of spatial-temporal dynamics in video data. Specifically, STPT confines the attention mechanism within spatial-temporal patches. Additionally, to seamlessly incorporate multimodal information, spatial-wise cross-attention is employed to integrate multimodal embeddings. For input tokens T^V\hat{T}_{\text{V}}, STPT transforms them into visual embeddings EV∈RN×h×w×CVE_{\text{V}}\in\mathcal{R}^{N\times h\times w\times C_{\text{V}}} by referencing a learnable codebook. The size of this codebook is set to 8193, exceeding the codebook size of VQGAN by 1, which enables compatibility with masked tokens. In each layer of STPT, as illustrated in Fig. 3, the visual embeddings are first processed through a 3D convolutional network. Then these embeddings are spatially partitioned into several patches EP∈RN×h/s×w/s×CVE_{\text{P}}\in\mathcal{R}^{N\times h/s\times w/s\times C_{\text{V}}}, where we empirically set patch stride ss as 2. Subsequently, each patch embeddings are flattened for spatial-temporal patchwise self-attention:

where Gs\mathcal{G}_{s} is the flatten operation that maps the embedding dimension to RNhw/s2×CV\mathcal{R}^{Nhw/s^{2}\times C_{\text{V}}}, and Gs−1\mathcal{G}_{s}^{-1} is the reverse operation. Fs\mathcal{F}_{\text{s}} is the standard self-attention. These patches are then concatenated and reshaped back to their original dimensions. In the following, the spatial-wise cross attention is applied, which facilitates feature interaction between visual embeddings and multimodal embeddings:

where Fc\mathcal{F}_{\text{c}} is the cross-attention operation that regards the frame number as batch size. After being processed through LL layers of STPT, the feature dimensionality of EVE_{\text{V}} is mapped to the codebook size of VQGAN. This enables the utilization of softmax to calculate the probability of each token, facilitating the prediction of masked visual tokens. Finally, cross-entropy loss is employed to optimize the proposed STPT:

where PSTPT(T~V,EM)\mathcal{P}_{\text{STPT}}(\widetilde{T}_{\text{V}},E_{\text{M}}) are visual token probabilities predicted by the STPT.

Notably, the proposed STPT can be trained jointly with videos and images. For image inputs, we simply replace the attention weight of Fs\mathcal{F}_{\text{s}} as a diagonal matrix . Simultaneous training on both video and image datasets offers a substantial augmentation of the training samples, enabling more efficient utilization of extensive image datasets. Besides, the joint training strategy has significantly enhanced WorldDreamer’s capability to comprehend temporal and spatial aspects within visual signals.

3 Mask Strategy

Mask strategy is crucial for training WorldDreamer, following , we train WorldDreamer utilizing a dynamic masking rate based on cosine scheduling. Specifically, we sample a random mask rate r∈r\in in each iteration, and totally 2hwπ(1−r2)−12\frac{2hw}{\pi}(1-r^{2})^{\frac{-1}{2}} tokens are masked in each frame. Note that we employ the same token mask across different frames. This decision is grounded in the similarity of visual signals between adjacent frames. Using different token masks could potentially lead to information leakage during the learning process. In comparison to an autoregressive mask scheduler, the dynamic mask schedule employed in our approach is crucial for parallel sampling at inference time, which enables the prediction of multiple output tokens in a single forward pass. This strategy capitalizes on the assumption of a Markovian property, where many tokens become conditionally independent given other tokens . The inference process also follows a cosine mask schedule, selecting a fixed fraction of the highest-confidence masked tokens for prediction at each step. Subsequently, these tokens are unmasked for the remaining steps, effectively reducing the set of masked tokens. As shown in Fig. 4, diffusion-based methods usually require ∼\sim30 steps to reduce noise, and autoregressive methods need ∼\sim200 steps to iteratively predict the next token. In contrast, WorldDreamer, parallel predicts masked tokens in about 10 steps, presenting a 3×∼20×3\times\sim 20\times acceleration compared to diffusion-based or autoregressive methods.

Experiment

We employ a diverse set of images and videos to train WorldDreamer, enhancing its understanding of visual dynamics. The specific data utilized in this training includes:

Deduplicated LAION-2B The original LAION dataset presented challenges such as data duplication and discrepancies between textual descriptions and accompanying images. We follow to address these issues. Specifically, we opted to utilize the deduplicated LAION-2B dataset for training WorldDreamer. This refined dataset excludes images with a watermark probability exceeding 50% or an NSFW probability surpassing 45%. The deduplicated LAION dataset was made available by , following the methodology introduced in .

WebVid-10M WebVid-10M comprises approximately 10 million short videos, each lasting an average of 18 seconds and primarily presented in the resolution of 336×596336\times 596. Each video is paired with associated text correlated with the visual content. A challenge posed by WebVid-10M is the presence of watermarks on all videos, resulting in the watermark being visible in all generated video content. Therefore, we opted to further refine WorldDreamer leveraging high-quality self-collected video-text pairs.

Self-collected video-text pairs We obtain publicly available video data from the internet and apply the procedure detailed in to preprocess the obtained videos. Specifically, we use PySceneDetect to detect the moments of scene switching and obtain video clips of a single continuous scene. Then, we filtered out clips with slow motion by calculating optical flow. Consequently, 500K high-quality video clips are obtained for training. For video caption, we extract the 10th, 50th, and 90th percentile frames of the video as keyframes. These key frames are processed by Gemini to generate captions for each keyframe. Additionally, Gemini is instructed to aggregate these individual image captions into an overall caption for the entire video. Regarding that highly descriptive captions enhance the training of generative models , we prompt Gemini to generate captions with as much detail as possible. The detailed captions allow WorldDreamer to learn more fine-grained text-visual correspondence.

NuScenes NuScenes is a popular dataset for autonomous driving, which comprises a total of 700 training videos and 150 validation videos. Each video includes approximately 20 seconds at a frame rate of 12Hz. WorldDreamer utilizes the front-view videos in the training set, with a frame interval of 6 frames. In total, there are approximately 28K driving scene videos for training. For video caption, we prompt Gemini to generate a detailed description of each frame, including weather, time of the day, road structure, and important traffic elements. Then Gemini is instructed to aggregate these image captions into an overall caption for each video. Furthermore, we extract the yaw angle and velocity of the ego-car as the action metadata.

2 Implementation Details

Train details WorldDreamer is first trained on a combination of WebVid and LAION datasets. For WebVid videos, we extract 16 frames as a training sample. For the LAION dataset, 16 independent images are selected as a training sample. Each sample is resized and cropped to an input resolution of 256×256256\times 256. WorldDreamer is trained over 2M iterations with a batch size of 64. The training process involves the optimization with AdamW and a learning rate of 5×10−55\times 10^{-5}, weight decay 0.01. To enhance the training and extend the data scope, WorldDreamer is further finetuned on self-collected datasets and nuScenes data, where all (1B) parameters of STPT can be trained. During the finetuning stage, the input resolution is 192×320192\times 320, and each sample has 24 frames. WorldDreamer is finetuned over 20K iterations with a batch size of 32, and the learning rate is 1×10−51\times 10^{-5}.

Inference details At inference time, Classifier-Free Guidance (CFG) is utilized to enhance the generation quality. Specifically, we randomly eliminate multimodal embeddings for 10% of training samples. During inference, we calculate a conditional logit cc and an unconditional logit uu for each masked token. The final logits gg are then derived by adjusting away from the unconditional logits by a factor of β\beta, referred to as the guidance scale:

For the predicted visual tokens, we employ the pretrained VQGAN decoder to directly output the video. Notably, WorldDreamer can generate a video consisting of 24 frames at a resolution of 192×320192\times 320, which takes only 3 seconds on a single A800.

3 Visualizations

We have conducted comprehensive visual experiments to demonstrate that WorldDreamer has acquired a profound understanding of the general visual dynamics of the general world. Through detailed visualizations and results, we present compelling evidence showcasing Worlddreamer’s ability to achieve video generation and video editing across diverse scenarios.

Image to Video WorldDreamer excels in high-fidelity image-to-video generation across various scenarios. As illustrated in Fig. 5, based on the initial image input, Worlddreamer has the capability to generate high-quality, cinematic landscape videos. The resulting videos exhibit seamless frame-to-frame motion, akin to the smooth camera movements seen in real films. Moreover, these videos adhere meticulously to the constraints imposed by the original image, ensuring a remarkable consistency in frame composition. It generates subsequent frames adhering to the constraints of the initial image, ensuring remarkable frame consistency.

Text to Video Fig. 6 demonstrates WorldDreamer’s remarkable proficiency in generating videos from text across various stylistic paradigms. The produced videos seamlessly align with the input language, where the language serves as a powerful control mechanism for shaping the content, style, and camera motion of the videos. This highlights WorldDreamer’s effectiveness in translating textual descriptions into visually faithful video content.

Video Inpainting As depicted in Fig. 7, WorldDreamer exhibits an exceptional ability for high-quality video inpainting. By providing a mask outlining the specific area of interest and a text prompt specifying desired modifications, WorldDreamer intricately alters the original video, yielding remarkably realistic results in the inpainting process.

Video Stylization Fig. 8 shows that WorldDreamer excels in delivering high-quality video stylization. By supplying a randomly generated visual token mask and a style prompt indicating desired modifications, WorldDreamer convincingly transforms the original video, achieving a genuinely realistic outcome in the stylization process.

Action to Video WorldDreamer shows the ability to generate videos based on actions in the context of autonomous driving. As shown in Fig. 9, given identical initial frames and different driving actions, WorldDreamer can produce distinct future videos corresponding to different driving actions (e.g., controlling the car to make a left-turn or a right-turn).

Conclusion

In conclusion, WorldDreamer marks a notable advancement in world modeling for video generation. Unlike traditional models constrained to specific scenarios, WorldDreamer capture the complexity of general world dynamic environments. WorldDreamer frames world modeling as a visual token prediction challenge, fostering a comprehensive comprehension of general world physics and motions, which significantly enhances the capabilities of video generation. In experiments, WorldDreamer shows exceptional performance across scenarios like natural scenes and driving environments, showcasing its adaptability in tasks such as text-to-video conversion, image-to-video synthesis, and video editing.

References