Pandora: Towards General World Model with Natural Language Actions and Video States
Jiannan Xiang, Guangyi Liu, Yi Gu, Qiyue Gao, Yuting Ning, Yuheng Zha, Zeyu Feng, Tianhua Tao, Shibo Hao, Yemin Shi, Zhengzhong Liu, Eric P. Xing, Zhiting Hu
Introduction
A world model (WM) is an abstract representation that an intelligent system uses to understand and simulate the real world. The model encompasses various aspects of the environment, including physical laws, spatiotemporal knowledge, objects, scenes, agents, and their dynamic interactions. In particular, it allows to predict the future states of the world in response to different actions. Building a general world model, therefore, can serve for interactive content creation, such as generating realistic virtual scenes for video games and movies, developing immersive experiences in virtual and augmented reality, and creating dynamic simulations for training and educational purposes. Perhaps of even more significance is that a general WM provides a foundation for robust, grounded reasoning in AI systems, enabling them to anticipate complex environments and plan actions, such as robots navigating disaster scenes safely. WMs also hold the potential to power long-horizon reasoning that improves decision making in fields like logistics and healthcare, by simulating various scenarios and outcomes and identifying the most effective solutions.
Current large language models (LLMs) are adept at generating human language and are used as surrogates for world models in certain reasoning tasks . However, language alone is a fundamentally insufficient and inefficient modality for describing various aspects of the world, such as intuitive physics (e.g., predicting fluid flow based on its viscosity) . Moreover, LLMs lack a robust understanding of physical and temporal dynamics in the real world, relying on patterns in textual data without comprehending the underlying realities they describe . On the other hand, contemporary video generation models can produce high-quality video content from given initial frames or text prompts . While these models can animate consistent sequences to visualize diverse scenes, they miss the complex interactive nature of the real world, lacking the ability for causal control and intervention with arbitrary actions during simulations. Recent work has also developed interactive world models at scale, such as GAIA-1 for auto-driving, UniSim for robotic manipulation, and Genie for 2D games. These models are typically specific to certain domains, permitting limited sets of actions and/or states.
This work presents \scalerel*X, a step towards a general world model that simulates world states across various domains by generating videos and allows real-time control through arbitrary actions expressed in natural language. \scalerel*X is an autoregressive model that sequentially processes actions (free text) and previous states (videos) as inputs and generates next states (videos) as outputs (Figure 2). \scalerel*X introduces a staged training strategy akin to the successful recipe of training LLMs , including: (1) large-scale pretraining with massive video and text data, respectively, to learn domain-general understanding of the world and production of consistent video simulations; and (2) instruction tuning with high-quality text-video sequential data to learn any-time text controllability during video generation.
Crucially, the pretraining stage allows for the separate training of text and video modeling. We thus can simply reuse existing pretrained LLMs and (text-to-)video generation models that have already achieved domain generality and video consistency in their own pretraining. We only need to stitch and align the language and video models together with necessary additional modules and lightweight tuning as described in §2. More specifically, in this work, we use the Vicuna-7B-v1.5 language model and the DynamiCrafter text-to-video model as the backbone. Using larger, more sophisticated pretrained models (such as GPT-4 and Sora) is expected to yield stronger performance. For the instruction tuning stage, we craft a large diverse set of action-state sequential data, by re-captioning general-domain videos and synthesizing with various simulators for robots, in-/out-door activities, driving, 2D games, and more. Similar to instruction tuning of LLMs that boosts their instructability in general unseen domains, tuning on the curated data boosts the world model’s real-time controllability that generalizes to broad unseen states and actions.
We illustrate extensive outputs generated by \scalerel*X across various domains in §3. The model demonstrates a range of desirable properties not exhibited by previous models. The results also indicate great potential for further enhancement with larger-scale training in the future.
The model simulates video states across broad domains: \scalerel*X is capable of generating videos across a wide range of general domains, such as indoor/outdoor, natural/urban, human/robot, 2D/3D, and other scenarios. This domain generality is primarily due to the large-scale video pretraining (inherited from the pretrained video model).
The model permits on-the-fly control with free-text actions: \scalerel*X accepts natural language actions as inputs during video generation to direct future world states. This differs crucially from previous text-to-video models which allow text prompts only at the beginning of the video. The on-the-fly control fulfills the promise of the world model to support interactive content generation and enhance robust reasoning and planning. The capability is enabled by the autoregressive architecture of the model (which permits text inputs at any time), the pretrained LLM backbone (which understands any text expressions), and the instruction tuning stage (which substantially enhances the effectiveness of control).
Action controllability transfers across domains: As above, instruction tuning with high-quality data allows the model to learn effective action control and transfer to different unseen domains. We demonstrate that actions learned from a specific domain apply seamlessly to states in diverse new domains.
Autoregressive model backbone enables longer videos: Existing video generation models based on diffusion architectures typically produce videos of a fixed length (e.g., 2 seconds). By integrating the pretrained video model with the LLM autoregressive backbone, \scalerel*X is capable of extending the video duration indefinitely in an autoregressive manner. Together with the additional training (e.g., instruction tuning), we show \scalerel*X can generate longer videos (e.g., 8 seconds) of higher quality.
Methods
*X is an autoregressive world model. Given the previous states of the world, e.g., images or video clips, and a natural language action, it predicts the next state of the world, which is also a video clip. Specifically, it formulates a state transition distribution:
where and are the state and action at time step , respectively. Each state is a single or a sequence of video frames, and each action is a sequence of text tokens. At the first time step, the state is one single image, and the states at the following steps are video clips.
Figure 2 gives an overview of the model architecture. The two core components of of \scalerel*X include the autoregressive backbone, which stems from a pretrained LLM, and the video generator, which is initialized with a pretrained video model. To stitch the two components together, other necessary components are added, including a vision encoder, and two adapters connecting the vision encoder to the LLM backbone, and the LLM backbone to the video generator, respectively.
At each time step , the autoregressive backbone accepts three sets of embedding vectors as inputs: (1) the first is the sequence of visual embeddings, by the vision encoder followed by the adapter (a Q-Former ), that encodes the previous world state ; (2) the second is the token embeddings of the text words in action ; and (3) the third is a sequence of learnable embedding vectors (a.k.a. query embeddings). The length and positions of the query embeddings correspond exactly to those of the output embeddings by the autoregressive backbone to be fed to the video generator. Intuitively, the query embeddings stimulate the model to start generating videos . The autoregressive backbone then generates a sequence of output embeddings. The adapter, which is a Q-Former, accepts the output embeddings and produces a new sequence of embeddings. Finally, the video generator takes the embeddings and generates the video clip outputs . To improve the consistency of the new video clip with the preceding video clip , the video generator additionally takes the last four frames of as input (or the single image of as input if is the initial state ). In addition, the video generator will take an FPS number to control the motion level of the video. The number of frames generated in each video clip depends on the specific pretrained video model used for initializing the video generator. As described below, we used the DynamiCrafter which generates 16 frames.
2 Staged Training
A general world model needs to achieve consistency, controllability, and generality—it needs to generate consistent videos to describe the world state accurately, allow on-the-fly control by accepting natural language actions at any time during video generation, and perform the above well across all diverse domains (with different scenes and actions).
To this end, direct training of the world model requires massive high-quality (video , text , video , ) sequences as training data, which is hard to obtain in practice. We instead devise a two-stage training strategy consisting of pretraining and instruction tuning.
The pretraining stage aims to acquire a few key capabilities, including (1) consistent general video generation of the video generator, (2) general text understanding of the autoregressive backbone to process actions, and (3) alignment of the representation spaces between the two components. The first two capabilities can be learned separately by training the video generator and the autoregressive backbone individually, or even by just plugging in existing pretrained video models and LLMs that already possess these capabilities during their own pretraining. The reuse of separately pretrained video and language models significantly reduces the training costs of the world model.
In the instruction tuning stage, we train the model on a curated video dataset with high-quality instructions (actions) that focus on the dynamics of the videos. This training is aimed at enhancing the model’s ability to follow natural language instructions and accurately predict subsequent video states based on these directions.
We describe more details of the two training stages in the next sections, respectively.
The pretraining stage aims to achieve the core capabilities of consistency and generality as described above. This is similar to the process of building an LLM where large-scale pretraining enables the LLM to generate consistent/fluent text in general domains.
General understanding of natural language can be achieved by massive training on text data, and generation of consistent general videos can be achieved by massive training on video data. Both of these have been done on existing pretrained LLMs and video generation models. We thus can directly reuse these models.
Specifically, we use Chat-Univi , which is a Vicuna-7B-v1.5 LLM equipped with a vision encoder, as our base LLM (autoregressive backbone), and DynamiCrafter as our base video generation model. DynamiCrafter is a diffusion model pretrained to generate a video given an image and a text prompt.
In the pretraining stage, we additionally want to align the pretrained LLM backbone and video model, so that the output embeddings from the LLM can be passed along to the video model as input for video generation. We used a video caption dataset WebVid-10M for the alignment training. For each (video, caption), we feed the first frame of the video and the caption into the LLM+vision encoder and get the output embeddings from the LLM. Meanwhile, we feed the caption into the text encoder of the video generation model and get the caption embeddings. We aim to match the two embeddings, so that the output embeddings from LLM can be understood by the vidoe generation model (just as how it understands the embeddings from its text encoder). Specifically, we minimize the L2 loss between the two sets of embeddings, and trains the parameters of the adapter between the LLM and video generator, as well as the query embeddings. Both the pretrained LLM and video generation model are fixed at this stage.
2.2 Instruction Tuning for Real-Time Controllability
This stage aims to gain real-time controllability by training the model on high-quality instruction tuning data. We construct such a dataset, which contains captions to precisely describe the dynamics of different clips in each video. With the data, we finetune the model by minimizing a diffusion loss on the videos given the instructions. In this stage, both the video generator and query embeddings are finetuned, while other components are fixed.
Below we describe the creation of the instruction tuning data in more details. An overview of the collected data is summarized in Table 1. The data come from both public corpus and simulators with careful data processing.
To make the dataset general, we use a large-scale video dataset, Panda-70M . We first filter the dataset by aesthetic score evaluation, optical flow magnitude assessment, cut detection, static video detection, and clip length filtering. Different from previous text-to-video models, our model emphasizes the controllability of natural language actions towards the next state. Therefore, we do re-captioning of the videos to get better captions that focus on the dynamics of the videos. we prompt GPT-4 Turbo to generate captions describing the dynamics of four frames sampled from each video clip. This process yields a total of 500k video-text pairs. Besides Panda-70M, we also collect video-action pairs from existing action-annotated datasets, including Something-Something V2 , BridgeData V2 , and EPIC-KITCHENS . This includes 260k examples.
To provide our model with more diverse and accurate training experience, we use simulation environments to collect video-action pairs. CARLA is a simulation platform for autonomous driving. It supports flexible modifications to the environment at runtime, making it suitable for simulating unexpected actions, such as Change the weather to Sunset or Add a car to the front. We sampled 75k video-action pairs from Carla. MP3D and StreetLearn are indoor and urban panorama scans. We built simulation environments to render these 3D scans. Turning actions such as turn right for 60 degrees can be constructed by gradually changing camera poses and collecting corresponding image projections. Besides, we prompt GPT-4 Turbo to generate scene descriptions, so that the instructions include both turning actions and the final scene descriptions. We got 70k data from MP3D and 146k data from StreetLearn, respectively. HM3D is a 3D environment dataset of real-world indoor scenes. We used Habitat-Lab to render these indoor scenes and collect data by sampling trajectories randomly. We created 152k data from it. Finally, we used Coinrun for collecting 2D game simulation data, resulting in 30k data.
Qualitative Results
We show qualitative results that demonstrate the core capabilities of \scalerel*X as a world simulator. Readers are encouraged to refer to https://world-model.ai for live video examples. We aim to report more quantitative results in the future.
*X is a general world model capable of generating videos across a broad range of domains. It permits on-the-fly control with free-text actions, i.e., it can accept text action control anytime during the video generation and predict future world states accordingly. We show the generation results of indoor/outdoor videos in Figure 3, robot/human videos in Figure 4, and 2D/3D game videos in Figure 5. In Figure 6, we also show videos that correctly demonstrate basic physical phenomena, demonstrating the model’s understanding of real-world physical concepts.
2 Action Controllability Transfer
Although some actions and their corresponding motion patterns only appear in some of the simulation data, we found that \scalerel*X can transfer the action controllability to different unseen domains. As shown in Figure 7 and Figure 8, \scalerel*X transfers 2D game ability from Coinrun and 3D simulator ability from HM3D to other unseen domains, respectively.
3 Autoregressively Generating Longer Videos
With the autoregressive backbone, \scalerel*X is capable of generating longer videos of higher quality in an autoregressive manner. \scalerel*X is trained on videos with up to 5 seconds (40 frames), but it is able to generate longer videos. We show the results of generating 8-second (64-frame) videos in Figure 9.
4 Limitations
*X can struggle to generate videos with high quality and good controllability. Figure 10 shows failure cases about semantics understanding, motion control, and video consistency.
When conducting small-scale exploratory experiments, we found that the data quality, i.e., the precision of the dynamics descriptions, has great influence on the model performance. In the domains where high-quality simulation data exists, the model easily gains great controllability. But in the domains of public video datasets, where captions generated by GPT-4 Turbo are noisy, the model does not show good performance. However, when we increased the training compute, controllability across general domains emergents on the model. We show a result comparison between the models trained with small-scale and large-scale training compute in Figure 11. We hypothesize it is because increasing data size can mitigate some of the noises in the data. The results indicate the great potential of building a stronger general world model by larger-scale training.
Related Works
World models simulate the future state of the world based on its previous states and given actions . Previous world models in AI systems are usually designed for specific domains. For example, in robotics domain, world models are usually used for model-based reinforcement learning in specific simulators . In robotics domain, world models are capable of predicting future image or video states across diverse robotics environments. These predictive capabilities are important for robots to understand the environments, make informed decisions, and execute tasks accurately. Besides the robotics domain, world models are also widely used in autonomous driving , where they mainly focus on path planning and real-time decision-making, which is pivotal in enabling vehicles to navigate complex environments safely and efficiently. There are also world models for 2D games . For example, Genie is a generative model capable of simulating an interactive 2D game given an image. In this work, we make a step towards building a more general world model that simulates any-domain states given any-text actions at any time.
Video generation models aim to synthesize realistic videos given text prompts or initial frames. Recent successes in diffusion models have paved the way for their application in the video generation domain . For example, additional modules are introduced into the existing image diffusion models to facilitate video generation capabilities. However, the length of generated videos is limited due to the non-autoregressive nature. Consequently, the Diffusion Transformer (DiT) has been proposed to allow for autoregressive generation, and Sora has further scaled it up, achieving remarkable success in generating long, high-quality video. Furthermore, as the strong understanding and generation ability of LLMs, have explored the usage of LLMs in vision generation domain. Additionally, incorporate LLMs for video generation to enhance the semantic understanding. Previous models are designed to generate scenes from input descriptions, yet they frequently lack the ability to control actions or predict real-world states. On the contrary, \scalerel*X is a hybrid autoregressive-diffusion model, thus it is capable of on-the-fly control over video generation.
Conclusion
We presented \scalerel*X as a step towards building a general world model. The model is able to simulate world states by generating videos across different domains, and control the video on the fly with natural language actions. \scalerel*X introduces a staged training recipe that allows to reuse and integrate existing pretrained language and video models. We believe larger-scale training with larger backbone models (e.g., GPT-4 and Sora) will lead to further improvement in terms of domain generality, video consistency, and action controllability. We are also excited about extending the model by incorporating other modalities, such as audio, to better measure and simulate the world.