DisCoScene: Spatially Disentangled Generative Radiance Fields for Controllable 3D-aware Scene Synthesis
Yinghao Xu, Menglei Chai, Zifan Shi, Sida Peng, Ivan Skorokhodov, Aliaksandr Siarohin, Ceyuan Yang, Yujun Shen, Hsin-Ying Lee, Bolei Zhou, Sergey Tulyakov
Introduction
3D-consistent image synthesis from single-view 2D data has become a trendy topic in generative modeling. Recent approaches like GRAF and Pi-GAN introduce 3D inductive bias by taking neural radiance fields as the underlying representation, gaining the capability of geometry modeling and explicit camera control. Despite their success in synthesizing individual objects (e.g., faces, cats, cars), they struggle on scene images that contain multiple objects with non-trivial layouts and complex backgrounds. The varying quantity and large diversity of objects, along with the intricate spatial arrangement and mutual occlusions, bring enormous challenges, which exceed the capacity of the object-level generative models .
Recent efforts have been made towards 3D-aware scene synthesis. Despite the encouraging progress, there are still fundamental drawbacks. For example, Generative Scene Networks (GSN) achieve large-scale scene synthesis by representing the scene as a grid of local radiance fields and training on 2D observations from continuous camera paths. However, object-level editing is not feasible due to spatial entanglement and the lack of explicit object definition. On the contrary, GIRAFFE explicitly composites object-centric radiance fields to support object-level control. Yet, it works poorly on challenging datasets containing multiple objects and complex backgrounds due to the absence of proper spatial priors.
To achieve high-quality and controllable scene synthesis, the scene representation stands out as one critical design focus. A well-structured scene representation can scale up the generation capability and tackle the aforementioned challenges. Imagine, given an empty apartment and a furniture catalog, what does it take for a person to arrange the space? Would people prefer to walk around and throw things here and there, or instead figure out an overall layout and then attend to each location for the detailed selection? Obviously, a layout describing the arrangement of each furniture in the space substantially eases the scene composition process . From this vantage point, here comes our primary motivation — an abstract object-oriented scene representation, namely a layout prior, could facilitate learning from challenging 2D data as a lightweight supervision signal during training and allow user interaction during inference. More specifically, to make such a prior easy to obtain and generalizable across different scenes, we define it as a set of object bounding boxes without semantic annotation, which describes the spatial composition of objects in the scene and supports intuitive object-level editing.
In this work, we present DisCoScene, a novel 3D-aware generative model for complex scenes. Our method allows for high-quality scene synthesis on challenging datasets and flexible user control of both the camera and scene objects. Driven by the aforementioned layout prior, our model spatially disentangles the scene into compositable radiance fields which are shared in the same object-centric generative model. To make the best use of the prior as a lightweight supervision during training, we propose global-local discrimination which attends to both the whole scene and individual objects to enforce spatial disentanglement between objects and against the background. Once the model is trained, users can generate and edit a scene by explicitly controlling the camera and the layout of objects’ bounding boxes. In addition, we develop an efficient rendering pipeline tailored for the spatially-disentangled radiance fields, which significantly accelerates object rendering and scene composition for both training and inference stages.
Our method is evaluated on diverse datasets, including both indoor and outdoor scenes. Qualitative and quantitative results demonstrate that, compared to existing baselines, our method achieves state-of-the-art performance in terms of both generation quality and editing capability. Tab. 1 compares DisCoScene with relevant works. it is worth noting that, to the best of our knowledge, DisCoScene stands as the first method that achieves high-quality 3D-aware generation on challenging datasets like Waymo , while enabling interactive object manipulation.
Related Work
3D-aware Image Synthesis. Generative Adversarial Networks (GANs) have achieved remarkable success in 2D image synthesis , and have recently been extended to 3D-aware image generation. VON and HoloGAN introduce voxel representations to the generator and use neural rendering to project 3D voxels into 2D space. Then, GRAF and Pi-GAN propose to use implicit functions to learn NeRF from single-view image collections, resulting in better multi-view consistency compared to voxel-based methods. GOF , ShadeGAN , and GRAM introduce occupancy field, albedo field and radiance surface instead of radiance field to learn better 3D shapes. However, high-resolution image synthesis with direct volumetric rendering is usually expensive. Many works resort to convolutional upsamplers to improve the rendering resolution and quality with lower computation overhead. While some other works adopt patch-based sampling and sparse-voxel to speed up training and inference. Note that most of these methods are restricted to well-aligned objects and fail on more complex, multi-object scene imagery. Our work instead naturally handles multi-object scenes with spatial disentangled object-level radiance fields, which can be scaled to very challenging real-world scene datasets.
Scene Generation. Scene generation has been a long-standing task. Early works like attempt to model a complex scene by trying to generate it. Recently, with the successes in generative models, scene generation has been advanced significantly. Among them, one popular line is to resort to the setups of image-to-image translation from given conditions, i.e., semantic masks , object-attribute graph . Although able to synthesize photorealistic scene images, they struggle to manipulate the objects in 3D space due to the lack of 3D understanding. Some works reuse the knowledge from 2D GAN models to achieve scene manipulation like the camera pose. But they suffer from poor multi-view consistency due to inadequate geometry modeling. Another active line of work explores adding 3D inductive biases to the scene representation. BlockGAN and GIRAFFE introduce compositional voxels and radiance fields to encode the object structures, but their object control can only be performed at simple diagnostic scenes. DepthGAN introduces depth as a 3D prior but is hard to achieve manipulation and multi-view consistency. GSN proposes to represent a scene with a grid of local radiance fields. However, since this local radiance field does not properly link to the object semantics, individual objects cannot be manipulated with versatile user control. Our work proposes to use an abstract layout prior to spatially disentangle the whole scene into object-centric radiance fields, which enables 3D-aware image synthesis on challenging real-world imagery like Waymo .
Method
The overall framework is illustrated in Fig. 1. We employ layout as an explicit prior to disentangle objects in our approach (Sec. 3.1). Based on the layout prior, we introduce our spatially disentangled radiance fields (Sec. 3.2) and an efficient rendering pipeline (Sec. 3.3) to achieve controllable 3D-aware scene generation. We also describe our global-local discrimination, which makes training on challenging datasets possible (Sec. 3.4). Finally, we discuss our model’s training and inference details on 2D image collections (Sec. 3.5).
There exist many representations of a scene, including the popular choice of scene graph , where objects and their relations are denoted as nodes and edges. Although graph can describe a scene in rich details, its structure is hard to process and the annotation is laborious to obtain in our case. Therefore, we opt to represent the scene layout in a much-simplified manner – a set of bounding boxes without category annotation, where counts objects in the scene. Concretely, each bounding box is defined with parameters, including rotation , translation , and scale .
where comprises Euler angles, which are easier to convert into rotation matrix . Using this notation, the bounding box can be transformed from a canonical bounding box , i.e., a unit cube at the coordinate origin:
2 Spatially Disentangled Radiance Fields
Since we use the layout as an internal representation, it naturally disentangles the whole scene into several objects. We can leverage multiple individual generative NeRFs to model different objects, but it can easily lead to an overwhelmingly large number of models and poor training efficiency. To alleviate this issue, we propose to infer generative object radiance field in the canonical space, to allow weight sharing among objects:
Spatial Condition. Although object bounding boxes are used as a prior, their latents are still randomly sampled regardless of their spatial configuration, leading to illogical arrangements. To synthesize scene images and infer object radiance fields with proper semantics, we adopt the location and scale of each object as a condition for the generator to encode more consistent intrinsic properties, i.e., shape and category. To this end, we simply modify Eq. 4 by concatenating the latent code with the Fourier features of object location and scale:
Therefore, semantic clues are injected into the layout in an unsupervised manner, without explicit category annotation.
3 Efficient Rendering Pipeline
As aforementioned, we use spatial-disentangled radiance fields to represent scenes. However, naïve point sampling solutions can lead to prohibitive computational overhead when rendering multiple radiance fields. Considering the independence of objects’ radiance fields, we can achieve much more efficient rendering by only focusing on the valid points within the bounding boxes.
Background Point Sampling. We adopt different background sampling strategies depending on the dataset. In general, we do fixed depth sampling for bounded backgrounds in indoor scenes and inherit the inverse parametrization of NeRF++ for complex and unbounded outdoor scenes, which uniformly samples background points in an inverse depth range. More details can be founded in the Supplementary Materials.
For any ray that does not intersect with boxes, its color and density are set to and , respectively. So that the foreground object map can be formulated as:
Since the background points are sampled at a fixed depth, we can directly adopt Eq. 6 to evaluate background points in the global space without sorting. And the background map can also be obtained by volume rendering similar to Eq. 7. Finally, and are alpha-blended into the final image with alpha extracted from Eq. 9:
Although our rendering pipeline efficiently composites multiple radiance fields, it still suffers from slow performance when rendering high-resolution images. To mitigate this issue, we render a high-dimensional feature map instead of a -channel color in a smaller resolution, followed by a StyleGAN2-like architecture that upsamples the feature map to the target resolution.
4 Local & Global Discrimination
5 Training and Inference
where is the softplus function, and and are the extracted object patches of synthesized image and real image , respectively. stands for the loss weight of the object discriminator. The last two terms in Eq. 13 are the gradient penalty regularizers of both discriminators, with and denoting their weights.
Inference. Besides high-quality scene generation, our method naturally supports object editing by manipulating the layout prior as shown in Fig. 1. Various applications are shown in Sec. 4.3. In particular, ray marching at a small resolution () may cause aliasing especially when moving the objects. We adopt supersampling anti-aliasing (SSAA) to perform ray marching at a temporary higher resolution () and downsample the feature map to the original resolution before the upsampler. This strategy is used only for object synthesis, and we do not change the background resolution during inference.
Experiments
Datasets. We evaluate DisCoScene on three multi-object scene datasets, including Clevr , 3D-Front , and Waymo . Clevr is a diagnostic multi-object dataset. We use the official script to render scenes with 2 and random primitives. Our Clevr dataset consists of samples in resolution. 3D-Front is an indoor scene dataset, containing a collection of houses with rooms. We obtain bedrooms after filtering out rooms with uncommon arrangements or unnatural sizes and use BlenderProc to render images per room from random camera positions, resulting in a total of images. Waymo is a large-scale autonomous driving dataset with video sequences of outdoor scenes. Six images are provided for each frame, and we only keep the front view. We also apply heuristic rules to filter out small and noisy cars and collect a subset of images. Because the width is always larger than height on Waymo, we adopt the black padding to make images square, similar with StyleGAN2 . More details about data preprocessing and rendering are included in Supplementary Materials.
Baselines. We compare with both 2D and 3D GANs. For 2D, we compare with StyleGAN2 on image quality. As for 3D, we compare with EpiGRAF , VolumeGAN , and EG-3D on object generation, and GIRAFFE , GSN on scene generation. We use the baseline models either released along with their papers or official implementations to train on our data.We fail to train GSN on Clevr and Waymo with the official implementation, hence we do not report the quantitative results.
2 Main Results
Qualitative Comparison. Fig. 2 presents the synthesized images in a resolution of of our method and baselines on all the datasets. We compare our method on explicit camera control and object editing with baselines.
GSN and EG3D, with a single radiance field, can manipulate the global camera of the synthesized images. GSN highly depends on the training camera trajectories. Thus in our setting where the camera positions are randomly sampled, it suffers from the poor image quality and multi-view consistency. As for EG3D, although it converges on the datasets, the object fidelity are lower than our method. On Clevr with a narrow camera distribution, the results of EG3D are inconsistent. In the first example, the color of the cylinder changes from gray to green across different views. Meanwhile, our method learns better 3D structure of the objects and achieves better camera control. On the challenging Waymo dataset, it is difficult to encode huge street scenes within a single generator, thus we train GIRAFFE and our DisCoScene in the camera space to evaluate object editing. GIRAFFE struggles to generate realistic results and, while manipulating objects, their geometry and appearance are not preserved well. Our approach is capable of handling these complicated scenarios with good variations. Wherever the object is placed and regardless of how the rotation is carried out, the synthesized objects are substantially better and more consistent than GIRAFFE. It demonstrates the effectiveness of our spatially disentangle radiance fields built upon the layout prior.
Quantitative Comparison. Tab. 2 reports the quantitative metrics on the quality of results, including FID and KID . All metrics are calculated between generated samples and all real images. DisCoScene consistently outperforms baselines with significant improvement on all datasets. Besides, training cost in V100 days and testing cost in ms/image (on a single V100 over samples) are also included to reflect the efficiency of our model. Note that the inference cost of 3D-aware models is evaluated on generating radiance fields rather than images. In such a case, EG3D and EpiGRAF are not fast as excepted due to the heavy computation on tri-planes. With comparable training and testing cost, it even achieves similar level of image quality with state-of-the-art 2D GAN baselines, e.g., StyleGAN2 , while allowing for explicit camera control and object editing that are otherwise challenging.
3 Controllable Scene Generation
The layout prior in our model enables versatile user controls of scene objects. In what follows, we evaluate the flexibility and effectiveness of our model through various 3D manipulation applications in different datasets. Examples are shown in Fig. 3 and more results can be found in Supplementary Materials.
Rearranging Objects. We can transform bounding boxes to rearrange (rotation and translation) the objects in the scenes without affecting their appearance. Transforming shapes in Clevr, furniture in 3D-Front, and cars in Waymo all show consistent results. In particular, rotating symmetric shapes (i.e., spheres and cylinders) in Clevr shows little changes, suggesting desired multi-view consistency. Our model can properly handle mutual occlusion. Take the blue cube from Clevr as example (-st row of Fig. 3), our model can produce new occlusions between it and the grey cylinder and generate high-quality renderings.
Removing and Cloning Objects. Users can update the layout by removing or cloning bounding boxes. Our method seamlessly removes objects with the background inpainted realistically, even without training on any pure background, including the challenging dataset of Waymo (-rd row of Fig. 3). Object cloning is also naturally supported, by copying and pasting a box to a new location in the layout.
Restyling Objects. Although appearance and shape are not explicitly modeled by the latent code, we can reuse the encoded hierarchical knowledge to perform object restyling. Like , we arbitrarily sample latent codes and perform style-mixing on different layers to achieve independent control over appearance and shape. Fig. 3 presents the restyling results on certain objects, i.e., the front cylinder in Clevr, the bed in 3D-Front, and the left car in Waymo.
Camera Movement. Explicit camera control is also permitted. Even for Clevr that is trained on very limited camera ranges, we can rotate the camera up to an extreme side view. Our model also produces consistent results when rotating the camera on 3D-Front (-nd row of Fig. 3).
4 Ablation Study
We ablate main components of our approach to better understand their individual contributions. In addition to the FID score that measures the quality of the entire image, we also provide another metric FID to measure the quality of individual objects. Specifically, we use the projected 2D boxes to crop objects from the synthesized images and then perform FID evaluation against the ones from real images.
Spatial Condition. To analyze how spatial condition (S-Cond) affects the quality of generation, we compare results with models trained with and without S-Cond on 3D-Front (LABEL:fig:spatial-cond). For example, our full model consistently infers beds at the center of rooms, while the baseline predicts random items like tables or nightstands that rarely appear in the middle of bedrooms. These results demonstrate that spatial condition can assist the generator with appropriate semantics from simple layout priors. Note that this correlation between spatial configurations and object semantics is automatically emerged without any supervision. We also numerically compare the image quality on these two models in Tab. 4, which shows that S-Cond also achieves better image quality at both scene- and object-level, because more proper semantics are more in line with the native distribution of real images.
Supersampling Anti-Aliasing. We adopt a simple super-sampling (SSAA) strategy to reduce edge aliasing by sampling more points during inference (Sec. 3.5). Thanks to our efficient object point sampling, doubling the resolution of foreground points keeps a similar inference speed ( ms/image), comparable with original speed ( ms/image). Results with different sampling points are shown in LABEL:fig:anti-aliasing. Taking the right boundary of the cabinet as an example (see the zoom-in insets for better visualization), when the cabinet is moved, SSAA achieves more consistent boundary compared with the jaggy one in the baseline.
Neural Renderer for Shadow. We adopt the StyleGAN2-like neural renderer to boost the rendering efficiency (Sec. 3.3). Besides the low computational cost, the added capacity of the neural renderer also brings better implicit modeling of realistic lighting effects such as shadowing. Therefore, without handling the shadowing effect in our rendering pipeline, our model can still synthesize high-quality shaodws on datasets such as Clevr (LABEL:fig:upsampler). This is because the large receptive field brought by convolutions and upsampler blocks make the neural renderer be aware of the object locations and progressively add shadows to the low resolution features rendered from radiance fields.
Discussion and Conclusion
Real Image Editing. Fig. 6 shows that it is possible to embed a real image into the latent space of our pretrained model using pivotal tuning inversion (PTI) . Besides reconstruction, all object manipulation operations are supported to edit the image. As one of the very first steps towards 3D scene editing from a single image, we believe that our method proves a promising venue and can inspire future research efforts along this direction.
Limitations and Future Work. Our model requires the abstract layout prior as the input. For in-the-wild datasets, we need monocular 3D object detector to infer pseudo layouts. While existing approaches attempt to learn the layout in an end-to-end manner, they struggle to generalize to complex scenes consisting of multiple objects. So it would be interesting to explore 3D layout estimation for complex scenes and combine with our approach end-to-end. Also, although our work shows significant improvement over existing 3D-aware scene generators, it is still challenging to learn on the street scenes in the global space due to the limited model capacity. Large-scale NeRFs might be one potention solutions.
Conclusion. This work presents DisCoScene, a method for controllable 3D-aware scene synthesis on challenging datasets. By taking spatially disentangled radiance fields as the representation based on a very abstract layout prior, our method is able to generate high-fidelity scene images and allows for versatile object-level editing.
Acknowledgements. We thank Jiatao Gu, Willi Menapace, Jian Ren, Panos Achlioptas, Tai Wang, and Zian Wang for fruitful discussions and comments about this work.
References
A1. Implementation Details of DisCoScene
Background with NeRF++ . The outdoor datasets, i.e. Waymo, have unbounded backgrounds. It is insufficient to model the whole scene in the image within a fixed bounding box. Therefore, we inherit the inverse parametrization of NeRF++ to model the background in Waymo:
where . The background points are uniformly sampled in an inverse depth range of where denotes the starting depth of the background.
Constant Latents for Upsampler. We adopt similar architecture and parameters of the synthesis network from StyleGAN2 as the upsampler for the rendered 2D feature map. Note that since our model handles multiple radiance fields, different spatial locations of the convolution feature maps should be modulated by different codes belonging to specific objects, making it costly to upsample the feature map. Thus we disable the spatial-aware modulation by setting as a constant tensor with value , which significantly reduces the computation overhead.
A2. Implementation Details of Baselines
Because of the wildly divergent data distribution, the training parameters vary greatly on different datasets. Tab. 5 and Tab. 6 list the detailed training configurations of different datasets for each baseline. FOV, Rangedepth, and Steps denote the field of view, the depth range, and the number of sampling steps along a camera ray, respectively. Rangeh and Rangev denote the horizontal and vertical angle ranges of the camera pose . Sample_Dist denotes the sampling scheme of the camera pose. We only use Gaussian or uniform sampling in our experiments. is the loss weight of the gradient penalty.
VolumeGAN . We use the official implementation of VolumeGAN.https://github.com/genforce/volumegan We train VolumeGAN with images. The coordinates range of feature volume is adjustable for different datasets. We adopt the training configuration in Tab. 5 to train VolumeGAN models.
EpiGRAF We use the official implementation of EpiGRAF.https://github.com/universome/epigraf We inherit the patch-wise training scheme to train EpiGRAF with the same data and camera parameters at the target resolution shown in Tab. 5.
EG3D . We use the official implementation of EG3D. https://github.com/NVlabs/eg3d Different from VolumeGAN and EpiGRAF, EG3D renders the whole radiance field within a bounding box, so we inherit the larger camera radius than the ones of EpiGRAF and VolumeGAN for training. Since the original EG3D requires pose annotations for training, we add a pose sampler in it to enable the training on all three datasets as the global annotations are not always available. We adjust the loss weight of gradient penalty on different datasets to achieve the best performance. Hyperparameters used for training are available in Tab. 6
GSN . We use the official GSN implementation.https://github.com/apple/ml-gsn. GSN highly dependents on input camera sequences and we find it very difficult to converge at a narrow camera distribution, i.e. Waymo and Clevr. On 3D-Front, we set the length of camera sequence to , and it can converge to some extent. We don’t leverage depth supervision for a fair comparison with our method.
GIRAFFE . We use the official implementation of GIRAFFE.https://github.com/autonomousvision/giraffe The number of boxes for training follows the configuration of ours on each dataset. The bounding box generator of GIRAFFE is tuned specifically for each dataset for a fair comparison.
A3. Data Preparation
Clevr . We use the official script to render scenes with Cube, Cylinder, and Sphere primitives. The camera position is jittered in a small scale. And the dataset is rendered in a resolution with samples.
3D-Front . We use BlenderProc to render images per room in 3D-Front. We move the center of each room to the coordinate origin and then sample the camera positions on the upper sphere between 2 to 3 where is the diagonal length of room.
Waymo . We only keep the front view of Waymo for the model training. However, there exist lots of occluded and noisy cars in Waymo, we design several heuristic rules to filter it. Specifically, we require the camera depth of car is less than 40 and the area of cars is larger than 40000 pixels in original image size (). We then adopt the black padding to make images square and then resize it in to resolution.
A5. Efficiency of Rendering Pipeline
Naïve point sampling solutions where the density and color of spatial points are inferred with multiple object radiance fields can lead to prohibitive computational overhead. Therefore we propose an efficient rendering pipeline by only focusing on the valid points within the bounding boxes. Tab. 7 presents the training cost in V100 days and testing cost in ms/image (on a single V100 over samples) with our efficient rendering and Naïve rendering. Our rendering pipeline can handle multiple objects efficiently, with nearly 1.5 and 2 times faster training and inference speed, respectively.
A6. Additional Results
We include a demo video, which shows more results of various 3D manipulation applications. From the video, we can see that our method both achieves good generation quality and enables precise object control. We also include comparisons with the state-of-the-art methods, i.e., EG3D and GIRRAFE , in the video.