DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving
Xiaofeng Wang, Zheng Zhu, Guan Huang, Xinze Chen, Jiagang Zhu, Jiwen Lu
Introduction
Spurred by insights from AGI (Artificial General Intelligence) and the principles of embodied AI, a profound transformation in autonomous driving is underway. Autonomous vehicles rely on sophisticated systems that engage with and comprehend the real driving world. At the heart of this evolution is the integration of world models . World models hold great promise for generating diverse and realistic driving videos, encompassing even long-tail scenarios, which can be utilized to train various driving perception approaches. Furthermore, the predictive capabilities in world models facilitate end-to-end driving, ushering in a new era of autonomous driving experiences.
Deriving latent dynamics of world models from visual signals was initially introduced in video prediction . By extrapolating from observed visual sequences, video prediction methods can infer future states of the environment, effectively modeling how objects and entities within a scene will evolve over time. However, modeling the intricate driving scenarios in pixel space is challenging due to the large sampling space . To alleviate this problem, recent research endeavors have sought innovative strategies to enhance sampling efficiency. ISO-Dream explicitly disentangles visual dynamics into controllable and uncontrollable states. MILE strategically incorporates world modeling within the Bird’s Eye View (BEV) semantic segmentation space, complementing world modeling with imitation learning. SEM2 further extends the Dreamer framework into BEV segmentation maps, utilizing Reinforce Learning (RL) for training. Despite the progress witnessed in world models, a critical limitation in relevant research lies in its predominant focus on simulation environments.
In this paper, we propose DriveDreamer, which pioneers the construction of comprehensive world models from real driving videos and human driver behaviors. Considering the intricate nature of modeling real-world driving scenes, we introduce the Autonomous-driving Diffusion Model (Auto-DM), which empowers the ability to create a comprehensive representation of the complex driving environment. We propose a two-stage training pipeline. In the first stage, we train Auto-DM by incorporating traffic structural information as intermediate conditions, which significantly enhances sampling efficiency. Consequently, Auto-DM exhibits remarkable capabilities in comprehending real-world driving scenes, particularly concerning the dynamic foreground objects and the static background. In the second-stage training, we establish the world model through video prediction. Specifically, driving actions are employed to iteratively update future traffic structural conditions, which enables DriveDreamer to anticipate variations in the driving environment based on different driving strategies. Moreover, DriveDreamer extends its predictive prowess to foresee forthcoming driving policies, drawing from historical observations and Auto-DM features. Thus creating a executable, and predictable driving world model.
The main contributions of this paper can be summarized as follows: (1) We introduce DriveDreamer, which is the first world model derived from real-world driving scenarios. DriveDreamer can jointly enable the generation of high-quality driving videos and reasonable driving policies. (2) To enhance the comprehension of real-world driving scenes and expedite the world model convergence, we introduce the Autonomous-driving Diffusion Model and a two-stage training pipeline. The first-stage training enables the comprehension of traffic structural information, and the second-stage video prediction training empowers the predictive capacity. (3) DriveDreamer can controllably generate driving scene videos that are highly aligned with traffic constraints (see Fig. 1), enhancing the training of driving perception methods (e.g., 3D detection). Besides, DriveDreamer can generate future driving policies based on historical observations and Auto-DM features. Notably, DriveDreamer achieves promising planning results in open-loop assessments on the nuScenes dataset.
Related Work
Diffusion models represent a family of probabilistic generative models that progressively introduce noise to data and subsequently learn to reverse this process for the purpose of generating samples . These models have recently garnered significant attention due to their exceptional performance in various applications, setting new benchmarks in image synthesis , video generation , and 3D content generation . To enhance the controllable generation capability, ControlNet , GLIGEN , T2I-Adapter and Composer have been introduced to utilize various control inputs, including depth maps, segmentation maps, canny edges, and sketches. Concurrently, BEVControl , MagicDrive and DrivingDiffuson incorporate layout conditions to enhance image generation. The fundamental essence of diffusion-based generative models lies in their capacity to comprehend and understand the intricacies of the world. Harnessing the power of these diffusion models, DriveDreamer seeks to comprehend the complex realm of autonomous-driving scenarios.
2 Video Generation
Video generation and video prediction are effective approaches to understanding the visual world. In the realm of video generation, several standard architectures have been employed, including Variational Autoencoders (VAEs) , auto-regressive models , flow-based models , and Generative Adversarial Networks (GANs) . Recently, the burgeoning diffusion models have also been extended to the domain of video generation. Video diffusion models exhibit higher-quality video generation capabilities, producing realistic frames and transitions between frames while offering enhanced controllability. They accommodate various input control conditions such as text, canny, sketch, semantic maps, and depth maps.
Video prediction models represent a specialized form of video generation models, sharing numerous similarities. In particular, video prediction involves anticipating future video changes based on historical video observations . DriveGAN establishes associations between driving actions and pixels, predicting future driving videos by specifying future driving policies. In contrast, DriveDreamer incorporates structured traffic conditions, text prompts, and driving actions as inputs, empowering precise, realistic video and action generation that are faithfully aligned with real-world driving scenarios.
3 World Models
World models have been extensively explored in model-based imitation learning, demonstrating remarkable success in various applications . These approaches typically leverage Variational Autoencoders (VAE) and Long Short-Term Memory (LSTM) to model transition dynamics and rendering functionality. World methods target at establishing dynamic models of environments, enabling agents to be predictive of the future. This aspect is of paramount importance in autonomous driving, where precise predictions about the future are essential for safe maneuvering. However, constructing world models in autonomous driving presents unique challenges, primarily due to the high sample complexity inherent in real-world driving tasks . To address these problems, ISO-Dream introduces an explicit disentanglement of visual dynamics into controllable and uncontrollable states. MILE strategically incorporates world modeling within the BEV semantic segmentation space, enhancing world modeling through imitation learning. SEM2 extends the Dreamer framework into BEV segmentation maps, employing reinforcement learning for training. Despite the progress witnessed in world models, a critical limitation in relevant research lies in its predominant focus on simulation environments. The transition to real-world driving scenarios remains an under-explored frontier.
DriveDreamer
The overall framework of DriveDreamer is depicted in Fig 3. The framework begins with an initial reference frame and its corresponding road structural information (i.e., HDMap and 3D box ). Within this context, DriveDreamer leverages the proposed ActionFormer to predict forthcoming road structural features in the latent space. These predicted features serve as conditions and are provided to Auto-DM, which generates future driving videos. Simultaneously, the utilization of text prompts allows for dynamic adjustments to the driving scenario style (e.g., weather and time of the day). Moreover, DriveDreamer incorporates historical action information and the multi-scale latent features extracted from Auto-DM, which are combined to generate reasonable future driving actions. In essence, DriveDreamer offers a comprehensive framework that seamlessly integrates multi-modal inputs to generate future driving videos and driving policies, thereby advancing the capabilities of autonomous-driving systems.
Regarding the extensive search space of establishing world models in real-world driving scenarios, we introduce a two-stage training strategy for DriveDreamer. This strategy is designed to significantly enhance sampling efficiency and expedite model convergence. The two-stage training is illustrated in Fig. 2. There are two steps in the first-stage training. Step 1 involves utilizing the single-frame structured condition, which guides DriveDreamer to generate driving scene image, facilitating its comprehension of structural traffic constraints. Step 2 extends its understanding into video generation. The second-stage training enables DriveDreamer to interact with the environment and predict future states effectively. This phase takes an initial frame image along with its corresponding structured information as input. Simultaneously, sequential driving actions are provided, with the model expected to generate future driving videos and future driving actions. In the following sections, we delve into the specifics of the model architecture and training pipelines.
Auto-DM. In DriveDreamer, we introduce Auto-DM, to model and comprehend driving scenarios from real-world driving videos. It is noted that comprehending driving scenes solely from pixel space presents challenges due to extensive search space in real-world driving scenarios. To mitigate this, we explicitly incorporate structured traffic information as conditional inputs.
where is MLP layers, is CLIP embed box categories features, is Fourier embedding , and is the concatenation operation. Then gated self-attention is leveraged to integrate position embeddings with visual signals from the original UNet features :
where is a learnable parameter, is self-attention, and is the token selection operation that considers visual tokens only .
To further empower Auto-DM with comprehension of driving dynamics, we introduce temporal attention layers to enhance frame coherence in the generated videos:
where we first reshape the visual signal from to . The shape transformation facilitates the frame-wise self-attention layers to learn inter-frame dynamics. denotes temporal position embeddings that are encoded by sinusoidal function . Finally, we restore the visual signal to its original dimensions, thus ensuring the feature integrity. Notably, the same architecture can be extended to generate multi-view images (see Fig. 6), where the solely attends to neighbor views. Additionally, a stack of frame-wise attention and view-wise attention contributes to multi-view video generation (see supplement for more details).
Furthermore, cross-attention layers are utilized to facilitate feature interactions between text inputs and visual signals, empowering text descriptions to influence driving scene attributes such as weather and time of day. In the next, we will elaborate on the first-stage training pipeline, which involves two steps.
Step 1 training. The Auto-DM incorporates input solely from a single frame of structured traffic conditions, coupled with supervision from a single-frame image. For structured traffic conditions, HDMaps and 3D boxes are obtained either from human annotations or pertained perception methods (e.g., LAV , BEVerse , UniAD ). Then three-channel HDMaps (lane boundary, lane divider, and pedestrian crossing) and eight-corner 3D boxes are projected onto the image plane to generate corresponding conditions. Notably, during step 1 training, temporal attention layers are omitted, which enables the network to focus exclusively on learning the traffic structural constraints, expediting the convergence of the training process.
Step 2 training. The Auto-DM incorporates input from multiple frames of structured traffic conditions and is supervised using driving videos. In contrast to step 1, learning from videos allows Auto-DM to gain a deeper understanding of the intricate motion transitions in driving scenarios. Building upon the pretrained models established in step 1, step 2 incorporates temporal attention layers into the model architecture. These additional parameters enable the Auto-DM to focus on the temporal dynamics present in the input data, further enhancing its ability to capture and interpret the nuanced temporal aspects of driving scenes.
In step 1 and step 2 training, the proposed Auto-DM is trained using the same noise schedule as the underlying image model . Specifically, the forward process gradually adds noise to the latent feature , resulting in the noisy latent feature . Then we train to predict the noise we added, and the trainable parameters are optimized via:
where denotes the trainable parameters involved in the gated self-attention, temporal attention, and cross-attention layers, and time step is uniformly sampled from .
2 Second-stage Training
Based on the first-stage training, DriveDreamer has obtained comprehension of the structured traffic information. However, the desired world model should also be predictive of the future and can interact with the environment. Therefore, we embark on the second phase of our approach. In this phase, we leverage the video prediction task to establish the driving world model. Specifically, the video prediction task entails providing an initial observation , as well as driving actions , with the desired outcome being the future driving videos , and future driving actions .
ActionFormer. Recall that the trained Auto-DM can generate driving videos based on sequential structured information . However, in the video prediction task, future traffic structural conditions beyond the present timestamp is unavailable. To address this challenge, we introduce the ActionFormer, which leverages driving actions to iteratively predict future structural conditions. The overall architecture of ActionFormer is in Fig. 5. Firstly the initial structural conditions are encoded and flattened into 1D latent space. The latent features are concatenated and aggregated by self-attention and MLP layers to generate the hidden state . Subsequently, cross-attention layers are utilized to construct associations between hidden states and driving actions. Then latent variable is parameterized as:
where are layers to learn Gaussian parameters. To predict future hidden states, we employ Gated Recurrent Units (GRUs) to iteratively make updates:
These hidden states are concatenated with action features and are decoded into future traffic structural conditions. It’s noted that the Actionformer forecasts future traffic conditions at the feature level, which mitigates noise interference at the pixel level, resulting in more robust predictions. Besides the traffic structural conditions generated by Actionformer and the text prompt condition, we process the reference image condition similar to . Based on the above conditions, we extend Auto-DM to jointly generate future driving videos and driving actions . We formalize this process as a generative probabilistic model, where the joint probability can be factorized as:
Considering updating hidden states is a deterministic process (Eq. 6), only latent variables are needed to be inferred to maximize the marginal likelihood of observation . Therefore, variational distribution is introduced to conduct variational inference:
where . Similar to , the variational lower bound can be derived as:
Note that the posterior and prior matching is not included, as we empirically find the simplified variational lower bound produces similar plausible results. In Eq. 10, the video prediction and action prediction parts can be modeled by Gaussian distributions and Laplace distribution . Therefore, we employ mean-squared error and loss to optimize the video prediction training. , are learnable layers involved in ActionFormer, Auto-DM, video decoder (i.e., VAE decoder) and action decoder. For action prediction details, we first pool multi-scale UNet features from Auto-DM. The pooled features are concatenated with historical action features, which are then decoded by MLP layers to generate future driving actions.
Based on the two-stage training, DriveDreamer has acquired a comprehensive understanding of the driving world, encompassing the structural constraints of traffic, predictions of future driving states, and interaction with the established world model.
Experiment
Dataset. The training data is sourced from the real-world driving dataset nuScenes , comprising a total of 700 training videos and 150 validation videos. Each video includes 20 seconds of footage captured by six surround-view cameras. The videos have a frame rate of 12Hz, resulting in 1M video frames available for training. During the first-stage training, we utilize the nuScenes-devkit to acquire HDMap annotations (lane boundary, lane divider, and pedestrian crossing) corresponding to 12Hz frames, which are then projected onto the image plane. Considering the nuScenes dataset only provides 2Hz 3D bounding box annotations, we supplement this with 12Hz bounding box annotations from . In the second-stage training, we employ the yaw angle and velocity of the ego-car as the driving action inputs. Besides, we extract scene description information (e.g., weather and time) from the nuScenes annotation, which serves as text conditions.
Training. The proposed Auto-DM is built upon Stable Diffusion v1.4 , whose original parameters are frozen. In step 1 of first-stage training, our model is trained for 40 epochs with a batch size of 16. In step 2, Auto-DM is trained for 10 epochs with a batch size of 1, with video frame length , and spatial size of . During second-stage video prediction training, our model predicts 16 frame driving videos and 16 future driving actions , and the model is trained for 10 epochs on a batch size of 1. All the experiments are conducted on A800 GPUs, and we use the AdamW optimizer with a learning rate .
Evaluation. We conducted a comprehensive evaluation of the proposed DriveDreamer, employing both qualitative and quantitative assessments. We utilized frame-wise Fréchet Inception Distance (FID) and Fréchet Video Distance (FVD) to evaluate the generation quality, where the evaluated image is resized to . Besides, to verify the generated images enhance the training of driving perception methods, DriveDreamer is evaluated through 3D object detection, with FCOS3D and BEVFusion as baseline methods. Furthermore, we test the performance of driving policy generation. Following the settings in , we evaluate output driving trajectories for future 3 seconds.
2 Controllable Driving Video Generation
The proposed DriveDreamer exhibits a profound comprehension of driving scenarios, capable of controllably generating diverse driving videos. In this subsection, we first demonstrate that, based on first-stage training, DriveDreamer can generate diverse driving videos under structured traffic conditions. Besides, we verify that the generated images can enhance the training of driving perception methods. Furthermore, DriveDreamer showcases its versatility by responding to different input actions, allowing for the control of the vehicle’s trajectory and consequently generating diverse driving videos.
As shown in Fig. 1 and Fig 6, DriveDreamer exhibits proficiency in producing images and videos that adhere meticulously to structured traffic conditions (more visualizations are in supplement). Significantly, we can also manipulate the text prompt to induce variations in the generated videos, encompassing changes in weather and time of day. To further validate the generation quality, we extract 4K traffic conditions (from the nuScenes training set) to generate driving images. The generated images are combined with real images for training the 3D detection task. Results in Tab. 1 indicate that training with our synthetic data significantly enhances the performance of 3D detection. Specifically, compared with training without synthetic data, the mAP metrics of FCOS3D and BEVFusion are improved by 0.7 and 3.0.
In addition to the utilization of structured traffic conditions for generating driving videos, DriveDreamer exhibits the capability to diversify the generated driving videos by adapting to different driving actions. As depicted in Fig. 1 (more visualizations are in supplement), starting from an initial frame paired with its corresponding structural information, DriveDreamer can generate distinct videos based on various driving actions, such as videos depicting left and right turns. In summary, DriveDreamer excels in producing a wide spectrum of driving scene videos, characterized by both high controllability and diversity. Thus, DriveDreamer holds promise for training autonomous-driving systems across a wide range of tasks, encompassing even corner cases and long-tail scenarios.
In the quantitative experiment, we extract ego-car driving actions from the nuScenes validation set as conditions to generate driving videos. For comparison, we train DriveGAN on the nuScenes dataset, employing the same training settings as those used for Drivedreamer. Besides, we train Drivedreamer without ActionFormer as a baseline (specifically, the action features are directly concatenated with the zero-padded structured traffic conditions). The results are presented in Tab. 2, where we evaluate the quality of generated videos. Notably, our approach without first-stage training achieves superior FID and FVD scores compared to DriveGAN. This observation underscores the effectiveness of leveraging a powerful diffusion model in visually comprehending driving scenarios. Furthermore, our findings reveal that Drivedreamer after first-stage training, exhibits an improved understanding of the structured information within driving scenes, resulting in higher-quality video generation. Lastly, we observe that the proposed ActionFormer effectively leverages the traffic structural information knowledge acquired during the first-stage training. Compared to the concatenation baseline approach, the ActionFormer iteratively updates future structured information based on input actions, which further enhances the quality of generated videos.
3 Driving Action Generation
In addition to its capacity for generating highly controllable driving videos, DriveDreamer demonstrates the ability to predict reasonable driving actions. As depicted in Fig. 1, provided with an initial frame condition and past driving actions, DriveDreamer can generate future driving actions that align with real-world scenarios . Furthermore, we conduct a quantitative assessment of the prediction accuracy. Specifically, MLP layers are utilized to encode past driving action information. Additionally, multi-scale UNet features are pooled as visual cues. The two modality features are then concatenated to learn future driving trajectories (more implementation details are in supplement). The results of open-loop evaluation on the nuScenes dataset are presented in Tab. 3. Remarkably, the average trajectory error of DriveDreamer is merely 0.29m, surpassing the performance of the multi-modality method VAD . In addition, DriveDreamer relatively decreases the average collision rate reported in by 21%, confirming that the visual features learned by DriveDreamer contribute to end-to-end autonomous driving, thereby enhancing driving safety.
Discussion and Conclusion
DriveDreamer represents a significant advancement in the field of world modeling, particularly in the context of autonomous driving. By focusing on real-world driving scenarios and harnessing the power of the diffusion model, DriveDreamer has demonstrated its ability to comprehend complex environments, generate high-quality driving videos, and formulate realistic driving policies. While prior research primarily concentrated on gaming or simulated environments, DriveDreamer extends the boundaries of world modeling to encompass the intricacies of actual driving conditions. DriveDreamer paves the way for future research in autonomous driving, emphasizing the importance of real-world representation for more accurate modeling and decision-making in this critical domain.
References
Implementation Details
Condition encoders. In DriveDreamer, diverse encoders are employed to embed different condition inputs, including the reference image, HDMap, 3D box, and action. The detailed architectures of these encoders are listed in Table 4. For spatially aligned conditions, such as the reference image and HDMap , a stack of 2D convolution layers is utilized to perform downsampling, ensuring the final output dimensions align with those of the diffusion noise. For unstructured conditions like the 3D box and action , Multilayer Perceptron (MLP) layers are employed for encoding features.
Multi-view generation. The framework of DriveDreamer can be easily extended to multi-view image/video generation. The model architecture comparison between video generation, multi-view image generation and multi-view video generation are shown in Fig. 7. For multi-view image generation, the model framework is the same as that of video generation, except that the frame-vise attention layers are replaced with view-wise attention layers. Besides, the view-wise attention layers construct associations solely between adjacent views. For multi-view video generation, view-wise attention layers and frame-wise attention layers are stacked to process diffusion latent features, which results in view-consistent and frame-consistent videos (see Fig. 8).
Action prediction architecture. For action prediction, the multi-modal features are first concatenated:
where is the average pooling operation, are multi-scale UNet features, and is the encoded driving action (i.e., velocity and yaw angle) features. Then we use MLP layers to learn future driving actions. For trajectory prediction evaluation, following , is additionally extracted from high-level command, accelerate and past trajectories, and we use the same action feature encoder of .
Synthetic data training. We leverage data generated by DriveDreamer to augment the training of 3D detection tasks. Specifically, DriveDreamer is fine-tuned with higher-resolution images, where the training data is from nuScenes . Consequently, DriveDreamer can generate high-fidelity images with a resolution of . Then the generated images are resized to the original resolution of , which can be utilized to train various off-the-shelf 3D detectors. During the training process, we randomly select 4000 samples (3D boxes and HDMap) from the nuScenes training set, which are employed to generate multi-view images. These synthetic data are mixed with the original training set to train 3D detectors. In the experiment, we train each baseline (i.e., FCOS3D and BEVFusion ) for 12 epochs. The results presented in Tab. 1 demonstrate that our approach significantly improves the performance of downstream tasks.
Visualizations
As shown in Fig. 9, DriveDreamer exhibits significant proficiency in producing a diverse range of driving scene videos that adhere meticulously to structured traffic conditions, comprising elements such as HDMaps and 3D boxes. Significantly, we can also manipulate the text prompt to induce variations in the generated videos, encompassing changes in weather and time of day. This heightened adaptability contributes substantially to the multifaceted nature of the generated video outputs. In addition to the utilization of structured traffic conditions for generating driving videos, DriveDreamer exhibits the capability to diversify the generated driving videos by adapting to different driving actions. As depicted in Fig. 10, starting from an initial frame paired with its corresponding structural information, DriveDreamer can generate distinct videos based on various driving actions, such as videos depicting left and right turns. Apart from its capacity for generating highly controllable driving videos, DriveDreamer demonstrates the ability to predict reasonable driving actions. As depicted in Fig. 11, provided with an initial frame condition and past driving actions, DriveDreamer can generate future driving actions that align with real-world scenarios. Comparative analysis of the generated actions against corresponding ground truth videos reveals that DriveDreamer consistently predicts sensible driving actions, even in complex situations such as intersections, obeying traffic lights, and executing turns.