MagicDrive: Street View Generation with Diverse 3D Geometry Control

Ruiyuan Gao, Kai Chen, Enze Xie, Lanqing Hong, Zhenguo Li, Dit-Yan Yeung, Qiang Xu

Introduction

The high costs associated with data collection and annotation often impede the effective training of deep learning models. Fortunately, cutting-edge generative models have illustrated that synthetic data can notably boost performance across various tasks, such as object detection (Chen et al., 2023c) and semantic segmentation (Wu et al., 2023b). Yet, the prevailing methodologies are largely tailored to 2D contexts, primarily relying on 2D bounding boxes (Lin et al., 2014; Han et al., 2021) or segmentation maps (Zhou et al., 2019) as layout conditions (Chen et al., 2023c; Li et al., 2023b).

In autonomous driving applications, a thorough grasp of the 3D environment is essential. This demands reliable techniques for tasks like Bird’s-Eye View (BEV) map segmentation (Zhou & Krähenbühl, 2022; Ji et al., 2023) and 3D object detection (Chen et al., 2020; Huang et al., 2021; Liu et al., 2023a; Ge et al., 2023). A genuine 3D geometry representation is crucial for capturing intricate details from 3D annotations, such as road elevations, object heights, and their occlusion patterns, as shown in Figure 2. Consequently, generating multi-camera street-view images according to 3D annotations becomes vital to boost downstream perception tasks.

For street-view data synthesis, two pivotal criteria are realism and controllability. Realism requires that the quality of the synthetic data should align with that of real data; and in a given scene, views from varying camera perspectives should remain consistent with one another (Mildenhall et al., 2020). On the other hand, controllability emphasizes the precision in generating street-view images that adhere to provided conditions: the BEV map, 3D object bounding boxes, and camera poses for views. Beyond these core requirements, effective data augmentation should also grant the flexibility to tweak finer scenario attributes, such as prevailing weather conditions or the time of day. Existing solutions like BEVGen (Swerdlow et al., 2023) approach street view generation by encapsulating all semantics within BEV. Conversely, BEVControl (Yang et al., 2023a) starts by projecting 3D coordinates to image views, subsequently using 2D geometric guidance. However, both methods compromise certain geometric dimensions—height is lost in BEVGen and depth in BEVControl.

The rise of diffusion models has significantly pushed the boundaries of controllable image generation quality. Specifically, ControlNet (Zhang et al., 2023a) proposes a flexible framework to incorporate 2D spatial controls based on pre-trained Text-to-Image (T2I) diffusion models (Rombach et al., 2022). However, 3D conditions are distinct from pixel-level conditions or text. The challenge of seamlessly integrating them with multi-camera view consistency in street view synthesis remains.

In this paper, we introduce MagicDrive, a novel framework dedicated to street-view synthesis with diverse 3D geometry controls In this paper, our 3D geometry controls contain control from road maps, 3D object boxes, and camera poses. We do not consider others like the exact shape of objects or background contents. . For realism, we harness the power of pre-trained stable diffusion (Rombach et al., 2022), further fine-tuning it for street view generation. One distinctive component of our framework is the cross-view attention module. This simple yet effective component provides multi-view consistency through interactions between adjacent views. In contrast to previous methods, MagicDrive proposes a separate design for objects and road map encoding to improve controllability with 3D data. More specifically, given the sequence-like, variable-length nature of 3D bounding boxes, we employ cross-attention akin to text embeddings for their encoding. Besides, we propose that an addictive encoder branch like ControlNet (Zhang et al., 2023a) can encode maps in BEV and is capable of view transformation. Therefore, our design achieves geometric controls without resorting to any explicit geometric transformations or imposing geometric constraints on multi-camera consistency. Finally, MagicDrive factors in textual descriptions, offering attribute control such as weather conditions and time of day.

Our MagicDrive framework, despite its simplicity, excels in generating strikingly realistic images that align with road maps, 3D bounding boxes, and varied camera perspectives. Besides, the images produced can enhance the training for both 3D object detection and BEV segmentation tasks. Furthermore, MagicDrive offers comprehensive geometric controls at the scene, background, and foreground levels. This flexibility makes it possible to craft previously unseen street views suitable for simulation purposes. We summarize the main contributions of this work as:

The introduction of MagicDrive, an innovative framework that generates multi-perspective camera views conditioned on BEV and 3D data tailored for autonomous driving.

The development of simple yet potent strategies to manage 3D geometric data, effectively addressing the challenges of multi-camera view consistency in street view generation.

Through rigorous experiments, we demonstrate that MagicDrive outperforms prior street view generation techniques, notably for the multi-dimensional controllability. Additionally, our results reveal that synthetic data delivers considerable improvements in 3D perception tasks.

Related Work

Diffusion Models for Conditional Generation. Diffusion models (Ho et al., 2020; Song et al., 2020; Zheng et al., 2023) generate images by learning a progressive denoising process from the Gaussian noise distribution to the image distribution. These models have proven exceptional across diverse tasks, such as text-to-image synthesis (Rombach et al., 2022; Nichol et al., 2022; Yang et al., 2023b), inpainting (Wang et al., 2023a), and instructional image editing (Zhang et al., 2023b; Brooks et al., 2023), due to their adaptability and competence in managing various form of controls (Zhang et al., 2023a; Li et al., 2023b) and multiple conditions (Liu et al., 2022a; Gao et al., 2023). Besides, data synthesized from geometric annotations can aid downstream tasks such as 2D object detection (Chen et al., 2023c; Wu et al., 2023b). Thus, this paper explores the potential of T2I diffusion models in generating street-view images and benefiting downstream 3D perception models.

Street View Generation. Numerous street view generation models condition on 2D layouts, such as 2D bounding boxes (Li et al., 2023b) and semantic segmentation (Wang et al., 2022). These methods leverage 2D layout information corresponding directly to image scale, whereas the 3D information does not possess this property, thereby rendering such methods unsuitable for leveraging 3D information for generation. For street view synthesis with 3D geometry, BEVGen (Swerdlow et al., 2023) is the first to explore. It utilizes a BEV map as a condition for both roads and vehicles. However, the omission of height information limits its application in 3D object detection. BEVControl (Yang et al., 2023a) amends the loss of object’s height by the height-lifting process. Similarly, Wang et al. (2023b) also projects 3D boxes to camera views to guide generation. However, the projection from 3D to 2D results in the loss of essential 3D geometric information, like depth and occlusion. In this paper, we propose to encode bounding boxes and road maps separately for more nuanced control and integrate scene descriptions, offering enhanced control over the generation of street views.

Multi-camera Image Generation of a 3D scene fundamentally requires viewpoint consistency. Several studies have addressed this issue within the context of indoor scenes. For instance, MVDiffusion (Tang et al., 2023) employs panoramic images and a cross-view attention module to maintain global consistency, while Tseng et al. (2023) leverage epipolar geometry as a constraining prior. These approaches, however, primarily rely on the continuity of image views, a condition not always met in street views due to limited camera overlap and different camera configurations (e.g., exposure, intrinsic). Our MagicDrive introduces extra cross-view attention modules to UNet, which significantly enhances consistency across multi-camera views.

Preliminary

Conditional Diffusion Models. Diffusion models (Ho et al., 2020; Song et al., 2020) generate data (x0\bm{x}_{0}) by iteratively denoising a random Gaussian noise (xT\bm{x}_{T}) for TT steps. Typically, to learn the denoising process, the network is trained to predict the noise by minimizing the mean-square error:

where ϵθ\bm{\epsilon}_{\theta} is the network to train, with parameters θ\theta, c\bm{c} is optional conditions, which is used for the conditional generation, t∈[0,T]t\in[0,T] is the time-step, ϵ∈N(0,I)\bm{\epsilon}\in\mathcal{N}(0,I) is the additive Gaussian noise, and αˉt\bar{\alpha}_{t} is a scalar parameter. Latent diffusion models (LDM) (Rombach et al., 2022) is a special kind of diffusion model, where they utilize a pre-trained Vector Quantized Variational AutoEncoder (VQ-VAE) (Esser et al., 2021) and perform diffusion process in the latent space. Given the VQ-VAE encoder as z=E(x)z=\mathcal{E}(x), one can rewrite ϵθ(⋅)\bm{\epsilon}_{\theta}(\cdot) in Equation 1 as ϵθ(αˉtE(x0)+1−αˉtϵ,t,c)\bm{\epsilon}_{\theta}(\sqrt{\bar{\alpha}_{t}}\mathcal{E}(\bm{x}_{0})+\sqrt{1-\bar{\alpha}_{t}}\bm{\epsilon},t,\bm{c}) for LDM. Besides, LDM considers text describing the image as condition cc.

Street View Generation with 3D Information

The overview of MagicDrive is depicted in Figure 3. Operating on the LDM pipeline, MagicDrive generates street-view images conditioned on both scene annotations (S\mathbf{S}) and the camera pose (P\mathbf{P}) for each view. Given the 3D geometric information in scene annotations, projecting all to a BEV map, akin to BEVGen (Swerdlow et al., 2023) or BEVControl (Yang et al., 2023a), doesn’t ensure precise guidance for street view generation, as exemplified in Figure 2. Consequently, MagicDrive categorizes conditions into three levels: scene (text and camera pose), foreground (3D bounding boxes), and background (road map); and integrates them separately via cross-attention and an additive encoder branch, detailed in Section 4.1. Additionally, maintaining consistency across different cameras is crucial for synthesizing street views. Thus, we introduce a simple yet effective cross-view attention module in Section 4.2. Lastly, we elucidate our training strategies in Section 4.3, emphasizing Classifier-Free Guidance (CFG) in integrating various conditions.

As illustrated in Figure 3, two strategies are employed for information injection into the UNet of diffusion models: cross-attention and additive encoder branch. Given that the attention mechanism (Vaswani et al., 2017) is tailored for sequential data, cross-attention is apt for managing variable length inputs like text tokens and bounding boxes. Conversely, for grid-like data, such as road maps, the additive encoder branch is effective in information injection (Zhang et al., 2023a). Therefore, MagicDrive employs distinct encoding modules for various conditions.

Ideally, the model learns the geometric relationship between bounding boxes and camera pose through training. However, the distribution of the number of visible boxes to different views is long-tailed. Thus, we bootstrap learning by filtering visible objects to each view (viv_{i}), i.e., fvizf_{viz} in Equation 6. Besides, we also add invisible boxes for augmentation (more details in Section 4.3).

Road Map Encoding. The road map has a 2D-grid format. While Zhang et al. (2023a) shows the addictive encoder can incorporate this kind of data for 2D guidance, the inherent perspective differences between the road map’s BEV and the camera’s First-Person View (FPV) create discrepancies. BEVControl (Yang et al., 2023a) employs a back-projection to transform from BEV to FPV but complicates the situation with an ill-posed problem. In MagicDrive, we propose that explicit view transformation is unnecessary, as sufficient 3D cues (e.g., height from object boxes and camera pose) allow the addictive encoder to accomplish view transformation. Specifically, we integrate scene-level and 3D bounding box embeddings into the map encoder (see Figure 3). Scene-level embeddings provide camera poses, and box embeddings offer road elevation cues. Additionally, incorporating text descriptions facilitates the generation of roads under varying conditions (e.g., weather and time of day). Thus, the map encoder can synergize with other conditions for generation.

2 Cross-view Attention Module

In multi-camera view generation, it is crucial that image synthesis remains consistent across different perspectives. To maintain consistency, we introduce a cross-view attention module (Figure 4). Given the sparse arrangement of cameras in driving contexts, each cross-view attention allows the target view to access information from its immediate left and right views, as in Equation 7; here, tt, ll, and rr are the target, left, and right view respectively. Then, the target view aggregates such information with skip connection, as in Equation 8, where hv\bm{h}^{v} indicates the hidden state of the target view.

We inject cross-view attention after the cross-attention module in the UNet and apply zero-initialization (Zhang et al., 2023a) to bootstrap the optimization. The efficacy of the cross-view attention module is demonstrated in Figure 4 right, Figure 5, and Figure 6. The multilayered structure of UNet enables aggregating information from long-range views after several stacked blocks. Therefore, using cross-view attention on adjacent views is enough for multi-view consistency, further evidenced by the ablation study in Appendix C.

3 Model Training

Classifier-free Guidance reinforces the impact of conditional guidance (Ho & Salimans, 2021; Rombach et al., 2022). For effective CFG, models need to discard conditions during training occasionally. Given the unique nature of each condition, applying a drop strategy is complex for multiple conditions. Therefore, our MagicDrive simplifies this for four conditions by concurrently dropping scene-level conditions (camera pose and text embeddings) at a rate of γs\gamma^{s}. For boxes and maps, which have semantic representations for null (i.e., padding token in boxes and in maps) in their encoding, we maintain them throughout training. At inference, we utilize null for all conditions, enabling meaningful amplification to guide generation.

Training Objective and Augmentation. With all the conditions injected as inputs, we adapt the training objective described in Section 3 to the multi-condition scenario, as in Equation 9.

Besides, we emphasize two essential strategies when training our MagicDrive. First, to counteract our filtering of visible boxes, we randomly add 10% invisible boxes as an augmentation, enhancing the model’s geometric transformation capabilities. Second, to leverage cross-view attention, which facilitates information sharing across multiple views, we apply unique noises to different views in each training step, preventing trivial solutions to Equation 9 (e.g., outputting the shared component across different views). Identical random noise is reserved exclusively for inference.

Experiments

Dataset and Baselines. We employ the nuScenes dataset (Caesar et al., 2020), a prevalent dataset in BEV segmentation and detection for driving, as the testing ground for MagicDrive. We adhere to the official configuration, utilizing 700 street-view scenes for training and 150 for validation. Our baselines are BEVGen (Swerdlow et al., 2023) and BEVControl (Yang et al., 2023a), both recent propositions for street view generation. Our method considers 10 object classes and 8 road classes, surpassing the baseline models in diversity. Appendix B holds additional details.

Evaluation Metrics. We evaluate both realism and controllability for street view generation. Realism is mainly measured using Fréchet Inception Distance (FID), reflecting image synthesis quality. For controllability, MagicDrive is evaluated through two perception tasks: BEV segmentation and 3D object detection, with CVT (Zhou & Krähenbühl, 2022) and BEVFusion (Liu et al., 2023a) as perception models, respectively. Both of them are renowned for their performance in each task. Firstly, we generate images aligned with the validation set annotations and use perception models pre-trained with real data to assess image quality and control accuracy. Then, data is generated based on the training set to examine the support for training perception models as data augmentation.

Model Setup. Our MagicDrive utilizes pre-trained weights from Stable Diffusion v1.5, training only newly added parameters. Per Zhang et al. (2023a), a trainable UNet encoder is created for EmapE_{map}. New parameters, except for the zero-init module and the class token, are randomly initialized. We adopt two resolutions to reconcile discrepancies in perception tasks and baselines: 224×\times400 (0.25×\times down-sample) following BEVGen and for CVT model support, and a higher 272×\times736 (0.5×\times down-sample) for BEVFusion support. Unless stated otherwise, images are sampled using the UniPC (Zhao et al., 2023) scheduler for 20 steps with CFG at 2.02.0.

2 Main Results

Realism and Controllability Validation. We assess MagicDrive’s capability to create realistic street-view images with the annotations from the nuScenes validation set. As shown by Table 1, MagicDrive outperforms others in image quality, yielding notably lower FID scores. Regarding controllability, assessed via BEV segmentation tasks, MagicDrive equals or exceeds baseline results at 224×\times400 resolution due to the distinct encoding design that enhances vehicle generation precision. At 272×\times736 resolution, our encoding strategy advancements enhance vehicle mIoU performance. Cropping large areas negatively impacts road mIoU on CVT. However, our bounding box encoding efficacy is backed by BEVFusion’s results in 3D object detection.

Training Support for BEV Segmentation and 3D Object Detection. MagicDrive can produce augmented data with accurate annotation controls, enhancing the training for perception tasks. For BEV segmentation, we augment an equal number of images as in the original dataset, ensuring consistent training iterations and batch sizes for fair comparisons to the baseline. As shown in Table 3, MagicDrive significantly enhances CVT in both settings, outperforming BEVGen, which only marginally improves vehicle segmentation. For 3D object detection, we train BEVFusion models with MagicDrive ’s synthetic data as augmentation. To optimize data augmentation, we randomly exclude 50% of bounding boxes in each generated scene. Table 2 shows the advantageous impact of MagicDrive ’s data in both CAM-only (C) and CAM+LiDAR (C+L) settings. It’s crucial to note that in CAM+LiDAR settings, BEVFusion utilizes both modalities for object detection, requiring more precise image generation due to LiDAR data incorporation. Nevertheless, MagicDrive’s synthetic data integrates seamlessly with LiDAR inputs, highlighting the data’s high fidelity.

3 Qualitative Evaluation

Comparison with Baselines. We assessed MagicDrive against two baselines, BEVGen and BEVControl, synthesizing multi-camera views for the same validation scenes (the comparison with BEVGen is in the Appendix D). Figure 5 illustrates that MagicDrive generates images markedly superior in quality to BEVControl, particularly excelling in accurate object positioning and maintaining consistency in street views for backgrounds and objects. Such performance primarily stems from MagicDrive ’s bounding box encoder and its cross-view attention module.

Multi-level Controls. The design of MagicDrive introduces multi-level controls to street-view generation through separation encoding. This section demonstrates the capabilities of MagicDrive by exploring three control signal levels: scene level (time of day and weather), background level (BEV map alterations and conditional views), and foreground level (object orientation and deletion). As illustrated in Figure 1, Figure 6, and Appendix E, MagicDrive adeptly accommodates alterations at each level, maintaining multi-camera consistency and high realism in generation.

4 Extension to Video Generation

We demonstrate the extensibility of MagicDrive to video generation by fine-tuning it on nuScenes videos. This involves modifying self-attention to ST-Attn (Wu et al., 2023a), adding a temporal attention module to each transformer block (Figure 7 left), and tuning the model on 7-frame clips with only first and last frames having bounding boxes. We sample initial noise independently for each frame using the UniPC (Zhao et al., 2023) sampler for 20 steps and illustrate an example in Figure 7 right.

Furthermore, by utilizing the interpolated annotations from ASAP (Wang et al., 2023c) like DriveDreamer (Wang et al., 2023b), MagicDrive can be extended to 16-frame video generation at 12Hz trained on Nvidia V100 GPUs. More results (e.g., video visualization) can be found on our website.

Ablation Study

Bounding Box Encoding. MagicDrive utilizes separate encoders for bounding boxes and road maps. To demonstrate the efficacy, we train a ControlNet (Zhang et al., 2023a) that takes the BEV map with both road and object semantics as a condition (like BEVGen), denoted as “w/o EboxE_{box}” in Table 4. Objects in BEV maps are relatively small, which require separate EboxE_{box} for accurate vehicle annotations, as shown by the vehicle mIoU performance gap. Applying visible object filter fvizf_{viz} significantly improves both road and vehicle mIoU by reducing the optimization burden. A MagicDrive variant incorporating EboxE_{box} with BEV of road and object semantics didn’t enhance performance, emphasizing the importance of integrating diverse information through different strategies.

Effect of Classifier-free Guidance. We focus on the two most crucial conditions, i.e. object boxes and road maps, and analyze how CFG affects the performance of generation. We change CFG from 1.5 to 4.0 and plot the change of validation results from CVT in Figure 8. Firstly, by increasing CFG scale, FID degrades due to notable changes in contrast and sharpness, as seen in previous studies (Chen et al., 2023c). Secondly, retaining the same map for both conditional and unconditional inference eliminates CFG’s effect on the map condition. As shown by blue lines of Figure 8, increasing CFG scale results in the highest vehicle mIoU at CFG=2.5, but the road mIoU keeps decreasing. Thirdly, with M={0}M=\{0\} for unconditional inference in CFG, road mIoU significantly increases. However, it slightly degrades the guidance on vehicle generation. As mentioned in Section 4.3, CFG complexity increases with more conditions. Despite simplifying training, various CFG choices exist during inference. We leave the in-depth investigation for this case as future work.

Conclusion

This paper presents MagicDrive, a novel framework to encode multiple geometric controls for high-quality multi-camera street view generation. With the separation encoding design, MagicDrive fully utilizes geometric information from 3D annotations and realizes accurate semantic control for street views. Besides, the proposed cross-view attention module is simple yet effective in guaranteeing consistency across multi-camera views. As evidenced by experiments, the generations from MagicDrive show high realism and fidelity to 3D annotations. Multiple controls equipped MagicDrive with improved generalizability for the generation of novel street views. Meanwhile, MagicDrive can be used for data augmentation, facilitating the training for perception models on both BEV segmentation and 3D object detection tasks.

Limitation and Future Work. We show failure cases from MagicDrive in Figure 9. Although MagicDrive can generate night views, they are not as dark as real images (as in Figure 9a). This may be due to that diffusion models are hard to generate too dark images (Guttenberg, 2023). Figure 9b shows that MagicDrive cannot generate unseen weathers for nuScenes. Future work may focus on how to improve the cross-domain generalization ability of street view generation.

Acknowledgement. This work is supported in part by General Research Fund (GRF) of Hong Kong Research Grants Council (RGC) under Grant No. 14203521 and in part by the Research Matching Grant Scheme under Grant No. 8601109. We gratefully acknowledge the support of MindSpore, CANN (Compute Architecture for Neural Networks) and Ascend AI Processor used for this research.

References

APPENDIX

Appendix A Object Filtering

In Equation 6, we employ fvizf_{viz} for object filtering to facilitate bootstrap learning. We show more details of fvizf_{viz} here. Refer to Figure 10 for illustration. For the sake of simplicity, each camera’s Field Of View (FOV) is not considered. Objects are defined as visible if any corner of their bounding boxes is located in front of the camera (i.e., zvi>0z^{v_{i}}>0) within each camera’s coordinate system. The application of fvizf_{viz} significantly lightens the workload of the bounding box encoder, evidence for which can be found in Section 6.

Appendix B More Experimental Details

Semantic Classes for Generation. To support most perception models on nuScenes, we try to include semantics commonly used in most settings (Huang et al., 2021; Zhou & Krähenbühl, 2022; Liu et al., 2023a; Ge et al., 2023). Specifically, for objects, ten categories include car, bus, truck, trailer, motorcycle, bicycle, construction vehicle, pedestrian, barrier, and traffic cone. For the road map, eight categories include drivable area, pedestrian crossing, walkway, stop line, car parking area, road divider, lane divider, and roadblock.

Optimization. We train all newly added parameters using AdamW (Loshchilov & Hutter, 2019) optimizer and a constant learning rate at 8e−58e^{-5} and batch size 24 (total 144 images for 6 views) with a linear warm-up of 3000 iterations, and set γs=0.2\gamma^{s}=0.2.

Appendix C Ablation on Number of Attending Views

In Table 5, we demonstrate the impact of varying the number of attended views on evaluation results. Attending to a single view yields superior FID results; the reduced influx of information from neighboring views simplifies optimization for that view. However, this approach compromises mIoU, also reflecting less consistent generation, as depicted in Figure 11. Conversely, incorporating all views deteriorates performance across all metrics, potentially due to excessive information causing interference in cross-attention. Since each view has an intersection with both left and right views, attending to one view cannot guarantee consistency, especially for foreground objects, while attending to more views requires more computation. Thus, we opt for 2 attended views in our main paper, striking a balance between consistency and computational efficiency.

Appendix D Qualitative Comparison with BEVGen

Figure 12 illustrates that MagicDrive generates images with higher quality compared to BEVGen (Swerdlow et al., 2023), particularly excelling in objects. Such enhancement can be attributed to MagicDrive’s utilization of the diffusion model and the adoption of a customized condition injection strategy.

Appendix E More results with control from different conditions

Figure 13 shows scene level control (time of day) and background level control (BEV map alterations). MagicDrive can effectively reflect these changes in control conditions through the generated camera views.

Appendix F More Experiments with 3D Object Detection

In Table 6, we show additional experimental results on training 3D object detection models using synthetic data produced by MagicDrive. Given that BEVFusion utilizes a lightweight backbone (i.e., Swin-T (Liu et al., 2021)), model performance appears to plateau with training through 1-2×\times epochs (2×\times: 20 for CAM-Only and 6 for CAM+LiDAR). Reducing epochs can mitigate this saturation, allowing more varied data to enhance the model’s perceptual capacity in both settings. This improvement is evident even when epochs for 3D object detection are further reduced to 0.5×\times. Our MagicDrive accurately augments street-view images with the annotations. Future works may focus on annotation sampling and construction strategies for synthetic data augmentation.

Appendix G More Discussion

Note that MagicDrive-generated street views can currently only perform as augmented samples to train with real data, and it is exciting to train detectors solely with generated data, which will be explored in the future. More flexible usage of the generated street views beyond data augmentation, especially incorporation with generative pre-training (Chen et al., 2023a; Zhili et al., 2023), contrastive learning (Chen et al., 2021; Liu et al., 2022b) and the large language models (LLMs) (Chen et al., 2023b; Gou et al., 2023), is an appealing future research direction. It is also interesting to utilize the geometric controls in different circumstances beyond 3D scenarios (e.g., multi-object tracking (Li et al., 2023a) and concept removal (Liu et al., 2023b)).

Appendix H Detailed Analysis on 3D Object Detection with Synthetic Data

We provide per-class AP for 3D object detection from the nuScenes validation set using BEVFusion in Table 7. From the results, we observe that, firstly, the improvements for large objects are significant, for example, buses, trailers, and construction vehicles. Secondly, objects with less diverse appearances, such as traffic cones and barriers, show more improvement, especially compared to trucks. Thirdly, we note that the improvement is marginal for cars, while significant for pedestrians, motorcycles, and bicycles. This may be because the baseline already performs well for cars. For pedestrians, motorcycles, and bicycles, even though distant objects from the ego car are generated less faithfully, MagicDrive can synthesize high-quality objects near the ego car, as shown in Figure 17-18. Therefore, more accurate detection of objects near the ego car contributes to improvements for these classes. Overall, mAP improvement comes with promotion in all classes’ AP, indicating MagicDrive can indeed help the training of perception models.

Appendix I More results for BEV segmentation

BEVFusion is also capable of BEV segmentation and considers most of the classes we used in the BEV map condition. Due to the lack of baselines, we present the results in Table 8 to facilitate comparison for future works. As can be seen, the 272×\times736 resolution does not outperform the 224×\times400 resolution. This is consistent with the results from CVT in Table 1 on the Road segment. Such results confirm that better map controls rely on maintaining the original aspect ratio for generation training (i.e., avoiding cropping on each side).

Appendix J Generalization of Camera Parameters

To improve generalization ability, MagicDrive encodes raw camera intrinsic and extrinsic parameters for different perspectives. However, the generalization ability is somewhat limited due to nuScenes fixing camera poses for different scenes. Nevertheless, we attempt to exchange the intrinsic and extrinsic parameters between the three front cameras and three back cameras. The comparison is shown in Figure 14. Since the positions of the nuScenes cameras are not symmetrical from front to back, and the back camera has a 120∘ FOV compared to the 70∘ FOV of the other cameras, clear differences between front and back views can be observed for the same 3D coordinates.

Appendix K More Generation Results

We show some corner-case (Li et al., 2022) generations in Figure 15, and more generations in Figure 16-Figure 18.