RenderDiffusion: Image Diffusion for 3D Reconstruction, Inpainting and Generation
Titas Anciukevičius, Zexiang Xu, Matthew Fisher, Paul Henderson, Hakan Bilen, Niloy J. Mitra, Paul Guerrero
Introduction
Image diffusion models now achieve state-of-the-art performance on both generation and inference tasks. Compared to alternative approaches (e.g. GANs and VAEs), they are able to model complex datasets more faithfully, particularly for long-tailed distributions, by explicitly maximizing likelihood of the training data. Many exciting applications have emerged in only the last few months, including text-to-image generation , inpainting , object insertion , and personalization .
However, in 3D generation and understanding, their success has so far been limited, both in terms of quality and diversity of the results. Some methods have successfully applied diffusion models directly to point cloud or voxel data , or optimized a NeRF using a pre-trained diffusion model . This limited success in 3D is due to two problems: first, an explicit 3D representation (e.g., voxels) leads to significant memory demands and affects convergence speed; and more importantly, a setup that requires access to explicit 3D supervision is problematic as 3D model repositories contain orders of magnitude fewer data compared to image counterparts—a particular problem for large diffusion models which tend to be more data-hungry than GANs or VAEs.
In this work, we present RenderDiffusion – the first diffusion method for 3D content that is trained using only 2D images. Like previous diffusion models, we train our model to denoise 2D images. Our key insight is to incorporate a latent 3D representation into the denoiser. This creates an inductive bias that allows us to recover 3D objects while training only to denoise in 2D, without explicit 3D supervision. This latent 3D structure consists of a triplane representation that is created from the noisy image by an encoder, and a volumetric renderer that renders the 3D representation back into a (denoised) 2D image. With the triplane representation, we avoid the cubic memory growth for volumetric data, and by working directly on 2D images, we avoid the need for 3D supervision. Compared to latent diffusion models that work on a pre-trained latent space , working directly on 2D images also allows us to obtain sharper generation and inference results. Note that RenderDiffusion does assume that we have the intrinsic and extrinsic camera parameters available at training time.
We evaluate RenderDiffusion on in-the-wild (FFHQ, AFHQ) and synthetic (CLEVR, ShapeNet) datasets and show that it generates plausible and diverse 3D-consistent scenes (see Figure 1). Furthermore, we demonstrate that it successfully performs challenging inference tasks such as monocular 3D reconstruction and inpainting 3D scenes from masked 2D images, without specific training for those tasks. We show improved reconstruction accuracy over a state-of-the-art method in monocular 3D reconstruction that was also trained with only monocular supervision.
In short, our key contribution is a denoising architecture with an explicit latent 3D representation, which enables us to build the first 3D-aware diffusion model that can be trained purely from 2D images.
Related Work
Generative models. To achieve high-quality image synthesis, diverse generative models have been proposed, including GANs , VAEs , independent component estimation , and autoregressive models . Recently, diffusion models have achieved state-of-the-art results on image generation and many other image synthesis tasks . Uniquely, diffusion models can avoid mode collapse (a common challenge for GANs), achieve better density estimation than other likelihood-based methods, and lead to high sample quality when generating images. We aim to extend such powerful image diffusion models from 2D image synthesis to 3D content generation and inference.
While many generative models (like GANs) have been extended for 3D generation tasks , applying diffusion models on 3D scenes is still relatively an unexplored area. A few recent works build 3D diffusion models using point-, voxel-, or SDF-based representations , relying on 3D (geometry) supervision. Except for the concurrent works , these methods focus on shape generation only and do not model surface color or texture – which are important for rendering the resulting shapes. Instead, we combine diffusion models with advanced neural field representations, leading to complete 3D content generation. Unlike implicit models , our model generates both shape and appearance, allowing for realistic image synthesis under arbitrary viewpoints.
Neural field representations. There has been exponential progress in the computer vision community on representing 3D scenes as neural fields , allowing for high-fidelity rendering in various reconstruction and image synthesis tasks . We utilize the recent tri-plane based representation in our diffusion model, allowing for compact and efficient 3D modeling with volume rendering.
Recent NeRF-based methods have developed generalizable networks trained across scenes for efficient few-shot 3D reconstruction. While our model is trained only on single-view images, our method can achieve high inference quality comparable to such a method like PixelNeRF that requires multi-view data during training. Several concurrent approaches use 2D diffusion models as priors for this task ; unlike ours they cannot synthesise new images or scenes a priori.
Neural field representations have also been extended to building 3D generative models , where most methods are based on GANs . Two concurrent works – DreamFusion (extending DreamFields ) and Latent-NeRF – leverage pre-trained 2D diffusion models as priors to drive NeRF optimization, achieving promising text-driven 3D generation.
We instead seek to build a 3D diffusion model and achieve direct object generation via sampling. Another concurrent work, GAUDI , introduces a 3D generative model that first learns a triplane-based latent space using multi-view data, and then builds a diffusion model over this latent space. In contrast, our approach only requires single-view 2D images and enables end-to-end 3D generation from image diffusion without pre-training any 3D latent space. DiffDreamer casts 3D scene generation as repeated inpainting of RGB-D images rendered from a moving camera; this inpainting uses a 2D diffusion model. However this method cannot generate scenes a priori.
Method
Our method builds on the successful training and generation setup of 2D image diffusion models, which are trained to denoise input images that have various amounts of added noise . At test time, novel images are generated by applying the model in multiple steps to progressively recover an image starting from pure noise samples. We keep this training and generation setup, but modify the architecture of the main denoiser to encode the noisy input image into a 3D representation of the scene that is volumetrically rendered to obtain the denoised output image. This introduces an inductive bias that favors 3D scene consistency, and allows us to render the 3D representation from novel viewpoints. Figure 2 shows an overview of our architecture. In the following, we first briefly review 2D image diffusion models (Section 3.1), then describe the novel architectural changes we introduce to obtain a 3D-aware denoiser (Section 3.2).
Diffusion models generate an image by moving a starting image progressively closer to the data distribution through multiple denoising steps .
To train the model, noisy images are created by repeatedly adding Gaussian noise starting from a training image :
where is a variance schedule that increases from to and controls how much noise is added in each step. We use a cosine schedule in our experiments. To avoid unnecessary iterations, we can directly obtain from in a single step using the closed form:
Reverse process.
The reverse process aims at reversing the steps of the forward process by finding the posterior distribution for the less noisy image given the more noisy image :
Note that is unknown (it is the image we want to generate), so we cannot directly compute this distribution, instead we train a denoiser with parameters to approximate it. Typically only the mean needs to be approximated by the denoiser, as the variance does not depend on the unknown image . Please see Ho et al. for a derivation.
We could directly train a denoiser to predict the mean , however, Ho et al. show that a denoiser can be trained more stably and efficiently by directly predicting the total noise that was added to the original image in Eq. 2. We follow Ho et al., but since our denoiser will also be tasked with reconstructing a 3D version of the scene shown in as intermediate representation, we train to predict instead of the noise :
where denotes the training loss. Once trained, at generation time, the model can then approximate the mean of the posterior as:
This approximate posterior is sampled in each generation step to progressively get the less noisy image from the more noisy image .
2 3D-Aware Denoiser
where denotes the concatenated parameters and of the encoder and renderer. The output image is a denoised version of the input image, thus it has to be rendered from the same viewpoint. We assume the viewpoint of the input image to be available. Note that the noise is applied directly to the source/rendered images.
A triplane representation factorizes a full 3D feature grid into three 2D feature maps placed along the three (canonical) coordinate planes, giving a significantly more compact representation . Each feature map has a resolution of , where is the number of feature channels. The feature for any given 3D point is then obtained by projecting the point to each coordinate plane, interpolating each feature map bilinearly, and summing up the three results to get a single feature vector of size . We denote this process of bilinear sampling from as .
Triplane encoder.
The triplane encoder transforms an input image of size into a triplane representation of size , where . We use the U-Net architecture commonly employed in diffusion models as a basis, but append additional layers (without skip connections) to output feature maps in the size of the triplanes. More architectural details are given in the supplementary material.
Triplane renderer.
The triplane renderer performs volume rendering using the triplane features and outputs an image of size . At each 3D sample point along rays cast from the image, we obtain density and color with an MLP as . The final color for a pixel is produced by integrating colors and densities along a ray using the same explicit volume rendering approach as MipNeRF . We use the two-pass importance sampling approach of NeRF ; the first pass uses stratified sampling to place samples along each ray, and the second pass importance-samples the results of the first pass.
3 Score-Distillation Regularization
4 3D Reconstruction
Unlike existing 2D diffusion models, we can use RenderDiffusion to reconstruct 3D scenes from 2D images. To reconstruct the scene shown in an input image , we pass it through the forward process for steps, and then denoise it in the reverse process using our learned denoiser . In the final denoising step, the triplanes encode a 3D scene that can be rendered from novel viewpoints. The choice of introduces an interesting control that is not available in existing 3D reconstruction methods. It allows us to trade off between reconstruction fidelity and generalization to out-of-distribution input images: At , no noise is added to the input image and the 3D reconstruction reproduces the scene shown in the input as accurately as possible; however, out-of-distribution images cannot be handled. With larger values for , input images that are increasingly out-of-distribution can be handled, as the denoiser can move the input images towards the learned distribution. This comes at the cost of reduced reconstruction fidelity, as the added noise removes some detail from the input image, which the denoiser fills in with generated content.
Experiments
We evaluate RenderDiffusion on three tasks: monocular 3D reconstruction, unconditional generation, and 3D-aware inpainting.
For training and evaluation we use real-world human face dataset (FFHQ), a cat face dataset (AFHQv2) as well as generated datasets of scenes from CLEVR and ShapeNet . We adopt FFHQ and AFHQv2 from EG3D , which uses off-the-shelf estimator to extract approximate extrinsics and augments dataset with horizontal image flips. We generate a variant of the CLEVR dataset which we call the CLEVR1 dataset, where each scene contains a single object standing on the plane at the origin. We randomize the objects in different scenes to have different colors, shapes, and sizes and generate scenes, where are used for training and the rest for testing. To evaluate our method on more complex shapes, we use objects from three categories of the ShapeNet dataset. ShapeNet contains man-made objects of various categories that are represented as textured meshes. We use shapes from the car, plane, and chair categories, each placed on a ground-plane. We use a total of objects from each category: for training and for testing. To render the scenes, we sample viewpoints uniformly on a hemisphere centered at the origin. Viewing angles that are too shallow (below ) are re-sampled. of these viewpoints are used for training, and are reserved for testing. We use Blender to render each of the objects from each of the viewpoints.
1 Monocular 3D Reconstruction
We evaluate 3D reconstruction on test scenes from each of our three ShapeNet categories. Since these images are drawn from the same distribution as the training data, we do not add noise, i.e. we set . Reconstruction is performed on the non-noisy image with one iteration of our denoiser . Note that our method does not require the camera viewpoint of the image as input.
We compare to two state-of-the-art methods that are also trained without 3D supervision. EG3D is a generative model that uses triplanes as its 3D representation and, like our method, trains with only single images as supervision. As it does not have an encoder, we perform GAN inversion to obtain a 3D reconstruction . Specifically, we optimize the latent vector and latent noise maps of the StyleGAN2 decoder and super-resolution blocks in EG3D to match the input image when rendered from the ground truth viewpoint. We optimize EG3D for 1000 steps for each of random initializations, minimizing the difference in VGG16 features between generated and target images, then pick the best result. As a second baseline we use PixelNeRF , a recent approach for novel view synthesis from sparse views. Unlike our method, this is trained with multi-view supervision, giving it a significant advantage over both ours and EG3D. We therefore treat this baseline as a reference that we do not expect to outperform. Both EG3D and PixelNeRF were carefully tuned and re-trained on our datasets.
Metrics.
We evaluate the 3D reconstruction performance by comparing all held-out test set views of the reconstructed scenes to the ground truth renders using two metrics: PSNR and SSIM . We average the results over all test images. Evaluating in image-space from multiple views has the advantage that it measures the accuracy not only of the shape, but also of the (possibly view-dependent) color of the surface at every point.
Results.
Qualitative results from our method and the baselines on ShapeNet are shown in Figure 4 and additional results including the CLEVR dataset, are shown in the supp. material. We see that EG3D usually predicts shapes that appear realistic from all angles. However, these often differ somewhat in shape or color from the input image, sometimes drastically (e.g. the 4th and 5th cars). In contrast, PixelNeRF produces faithful reconstructions of its input views. However, when rotated away from an input view the images are often blurry or lack details. Our RenderDiffusion achieves a balance between the two, preserving the identity of shapes, but yielding sharp and plausible details in back-facing regions. Quantitative results are shown in Table 1. In agreement with the qualitative results, all methods perform better on the simpler CLEVR1 dataset than on the three ShapeNet classes. RenderDiffusion out-performs EG3D across all datasets, achieving an average PSNR on ShapeNet of 26.1, versus 24.1 for EG3D. PixelNeRF, which is not directly comparable since it receives multi-view supervision during training, achieves higher performance still, with an average PSNR on ShapeNet of 28.9. Note that on GeForce GTX 1080 Ti our method performs reconstruction in a single pass in under seconds per scene compared to 3 min. for EG3D inversion. On FFHQ and AFHQ (Fig. 3, top four rows), our method accurately reconstructs input faces and cats, predicting a plausible depth map and renderings from novel views.
Reconstruction on out-of-distribution images.
Using a 3D-aware denoiser allows us to reconstruct a 3D scene from noisy images, where information that is lost to the noise is filled in with generated content. By adding more noise, we can generalize to input images that are increasingly out-of-distribution, at the cost of reconstruction fidelity. In Figure 6, we show 3D reconstructions from photos that have significantly different backgrounds and materials than the images seen at training time. We see that results with added noise () generalize better than results without added noise (), at the cost of less accurate shapes and poses of the reconstructed models. In addition, in the supplementary material, we show how results vary with differing amounts of noise .
2 Unconditional Generation
We show results for unconditional generation; for additional quantitative results, please refer to the supplemental.
We compare against EG3D , which is the most similar existing work to ours, since it too uses a triplane representation for the 3D scene, and a similar rendering approach. We also compare with the older methods pi-GAN and GIRAFFE . All three baselines use adversarial training rather than denoising diffusion.
Results.
Qualitative examples from both methods are shown in Fig. 5. We see that scenes generated by both models appear realistic from all viewpoints, and are 3D-consistent. This is particularly notable for RenderDiffusion, since our method performs the denoising process from just a single viewpoint (top rows of Fig. 5). We observe somewhat higher diversity of both color and shape among the samples from our model than those from EG3D. However, both methods are able to generate complex structures (e.g. slatted chair backs), and synthesise physically-plausible shadows cast onto the ground-plane. Quantitatively (Tab. 2) our method performs better than EG3D and GIRAFFE w.r.t. dataset coverage (see supp. for a definition and additional results) on two of the three ShapeNet classes, while EG3D is best for the other datasets. On FFHQ and AFHQ, our model again produces plausible, 3D-consistent samples (Fig. 3), though with some artifacts visible at larger azimuth angles.
3 3D-Aware Inpainting
We also apply our trained model to the task of inpainting masked 2D regions of an image while simultaneously reconstructing the 3D shape it shows. We follow an approach similar to RePaint , but using our 3D denoiser instead of their 2D architecture. Specifically, we condition the denoising iterations on the known regions of the image, by setting in known regions to the noised target pixels, while sampling the unknown regions as usual based on . Thus, the model performs 3D-aware inpainting, finding a latent 3D structure that is consistent with the observed part of the image, and also plausible in the masked part.
Qualitative results are shown in Figure S8. For each masked image, we show scnes resulting from two different noise samples. Even though our model was not trained explicitly on this task, we can see that it generates diverse and plausible inpainted regions that match the observed part of the image. In the supplementary material, we give quantitative results on this task, by sampling multiple completions for each input, and measuring how close they can get to the ground-truth.
Conclusion
We have presented RenderDiffusion, the first 3D diffusion model that can be trained purely from posed 2D images, without requiring any explicit 3D supervision. Our denoising architecture incorporates a triplane rendering model which enforces a strong inductive bias and produces 3D-consistent generations. RenderDiffusion can be used to infer a 3D scene from an image, for 3D editing using 2D inpainting, and for 3D scene generation. We have shown competitive performance on sampling and inference tasks, in terms of both quality and diversity of results.
Our method currently has several limitations. First, our generated images still lag behind GANs in some cases, probably due to more blurry outputs; however ours is the first 3D-aware diffusion model trained with 2D images, and there are engineering techniques that could enhance the quality, like upsampling models. Second, we have introduced a score distillation regularization to prevent learning trivial geometry on FFHQ/AFHQ, which results in a loss of fidelity in the generated 3D models; we expect this to be less of an issue as score distillation methods improve. Third, we require the training images to include camera extrinsics, and our triplanes are positioned in a global coordinate system, which restricts generalization across object placements. This limitation can be addressed by utilizing off-the-shelf pose estimation, or rendering everything in the camera-view. Finally, we would like to support object editing and material editing, enabling a more expressive 3D-aware 2D image-editing workflow.
Acknowledgements
We would like to thank Yilun Du, Christopher K. I. Williams, Noam Aigerman, Julien Philip and Valentin Deschaintre for valuable discussions. HB was supported by the EPSRC Visual AI grant EP/T028572/1; NM was partially supported by the UCL AI Centre.
References
S7 Overview of Supplementary
In this supplementary material, we provide additional architecture details (Section S8), additional results for unconditional generation (Section S9), additional results for 3D-aware inpainting (Section S10), an additional experiment where generate multiple ShapeNet categories with a single model (Section S11), and additional results for reconstruction (Section S12).
S8 Architecture Details
Here we summarise the architecture of the denoiser network. Code, training configurations and datasets are publicly available at https://github.com/Anciukevicius/RenderDiffusion.
The triplane encoder transforms the input image of size into a triplane representation of size . We choose and for our experiments on ShapeNet, as we found that the increased triplane resolution improves the quality of our results, and , for CLEVR1. Similar to other 2D diffusion models , we use a UNet architecture for the triplane encoder. The UNet consists of 8 down and up blocks . Each block consists of 2 ResNet blocks that additionaly take a timestep embedding, and linear attention. If the triplane has larger resolution than the input image, we append additional up blocks to the UNet that upsample the image to the triplane resolution. These have the same architecture as the other UNet blocks, except that they do not use skip connections, as there is no down block in the UNet with the corresponding resolution.
Triplane renderer.
To render triplanes, we mostly follow EG3D ; however, we use explicit volumetric rendering that samples points along the ray and queries a 2-layer fully-connected neural network to output color and a density . The network takes as input 32-dimensional sum-pooled interpolations of triplane features. Unlike EG3D, we also use a positional embedding of the 3D sample position as input to the network, which allows the network to represent parts of the ground plane that extend beyond the triplanes with a single constant feature.
S9 Additional Unconditional Generation Results
We evaluate the distributions of generated scenes quantitatively using four metrics. is the Fréchet Inception Distance (FID) computed between training views of the generated scenes and training views of all training set scenes. is the FID computed between test views of the generated scenes and test views of all scenes from the test set. The coverage metric (cov.) measures how well the generated distribution covers the data distribution. It is defined as the fraction of training set images with neighborhoods that contain at least one generated sample, with a neighborhood defined based on the 3-nearest neighbors. Similarly, the density metric (dens.) measures how close generated samples are to the data distribution, by calculating the average number of real samples whose neighborhoods contain each generated sample. Neighborhoods are defined in the feature space of a VGG-16 network (last hidden layer) that was applied to all training set views of a generated scene. The latter two metrics are similar to the improved recall and precision metrics of , but avoid certain pathological behaviors .
Results on the synthetic datasets are presented in Tab. S3. We see that EG3D performs well on the FID metric, with ours second for CLEVR and pi-GAN second for ShapeNet. Our approach tends to perform better on the coverage and density metrics, while pi-GAN is particularly poor on these. To interpret our quantitative performance relative to EG3D, we refer to the qualitative results shown in Figure 5 of the main paper, where we see that our shapes are slightly more blurry than EG3D, but exhibit more variety and similar shape quality, apart from the blurriness. We hypothesize that the blurriness introduces a bias that the FID metric is highly sensitive to (as the blurriness may affect the feature average that the FID is based on). Coverage and density are less sensitive to the blurriness, as they don’t rely on an average over all samples and instead provide a more detailed comparison of the sample distributions by working with sample neighborhoods. This interpretation of the quantitative results suggests that, leaving aside the blurriness, our method generates samples that better cover the data distribution, at a comparable sample quality. This is in line with current understanding of the differences between diffusion models like RenderDiffusion, and GANs like EG3D. Note that our models were not fully converged at the time of measuring these results, and we observed that the blurriness gradually decreases over the course of the training, making it likely that the blurriness can be reduced with additional training. Tab. S4 shows quantitative results on the real datasets FFHQ (photos of human faces) and AFHQ (photos of cat faces); see also the qualitative results in the main paper.
Additional qualitative results
The supplementary video shows additional uncurated (not cherry-picked) qualitative results for unconditional generation, shown from a camera that rotates around the object. Fig. S10 shows uncurated generated samples from GIRAFFE and pi-GAN, on the three ShapeNet classes.
S10 Additional Inpainting Results
To quantitatively measure how well our generative model inpaints 3D scenes, we treat inpainting as a 3D reconstruction task with occlusions, where the mask is the occluder. Similar to the unoccluded case, we compare renders of the reconstructed scene from test set viewpoints to ground truth renders using PSNR and SSIM as metrics. Since there are often many plausible inpaintings (i.e. the task is ambiguous), we sample different inpaintings with our model and select the best matching one. For CLEVR we choose as the mask often hides majority of the object (increasing the degree of ambiguity), while for ShapeNet we choose . This gives as an indication if the distribution of generated scenes for a given masked input image includes the ground truth scene. To choose the masked-out region of each image, we use a square with width and height equal to 40% of the image resolution (e.g. for an image of size the mask will be of size ). The mask is placed uniformly at random within a square region of side length of the image size, itself centered in the image. This ensures the mask always covers part of the foreground object, not just the background. Illustration of masked inputs and diversity in RenderDiffusion predictions is shown in Fig. S8.
Quantitative results on this task are given in Tab. S5. We compare the reconstruction performance with and without masked input. We can see that in most cases, the performance for the two cases is comparable, indicating that a scene resembling the ground truth is contained in the output distribution. We show additional qualitative results with multiple seeds in the supplementary video.
S11 Multi-Category Generation Results
To further demonstrate that our model can represent complex, multi-modal distributions, we perform an additional experiment where a single model is trained jointly on multiple ShapeNet categories. Specifically, we train RenderDiffusion on the union of the chair and plane categories, otherwise using the same architecture and training protocol as described in the main paper.
In Fig. S9, we show qualitative results from this model. We see that RenderDiffusion has successfully captured both modes of the dataset, sampling plausible chairs and airplanes. As in the single-category experiments in the main paper, the samples are 3D-consistent, exhibit plausible depth-maps, and look realistic from novel test viewpoints.
S12 Additional Reconstruction Results
In Figure S12, we show addditional reconstruction results for CLEVR1 and ShapeNet chair datasets. In Figure S11, we show reconstruction from out-of-distribution images with different amounts of added noise, ranging from no noise at to noise steps at . Adding larger amounts of noise results in reconstructions that are more generic and increasingly diverge from the input image, as the generative model fills in details covered by the noise, but also show increasingly higher-quality shapes.