HoloDiffusion: Training a 3D Diffusion Model using 2D Images
Animesh Karnewar, Andrea Vedaldi, David Novotny, Niloy Mitra
Introduction
Diffusion models have rapidly emerged as formidable generative models for images, replacing others (e.g., VAEs, GANs) for a range of applications, including image colorization , image editing , and image synthesis . These models explicitly optimize the likelihood of the training samples, can be trained on millions if not billions of images, and have been shown to capture the underlying model distribution better than previous alternatives.
A natural next step is to bring diffusion models to 3D data. Compared to 2D images, 3D models facilitate direct manipulation of the generated content, result in perfect view consistency across different cameras, and allow object placement using direct handles. However, learning 3D diffusion models is hindered by the lack of a sufficient volume of 3D data for training. A further question is the choice of representation for the 3D data itself (e.g., voxels, point clouds, meshes, occupancy grids, etc.). Researchers have proposed 3D-aware diffusion models for point clouds , volumetric shape data using wavelet features and novel view synthesis . They have also proposed to distill a pretrained 2D diffusion model to generate neural radiance fields of 3D objects . However, a diffusion-based 3D generator model trained using only 2D image for supervision is not available yet.
In this paper, we contribute HoloDiffusion, the first unconditional 3D diffusion model that can be trained with only real posed 2D images. By posed, we mean different views of the same object with known cameras, for example, obtained by means of structure from motion .
We make two main technical contributions: (i) We propose a new 3D model that uses a hybrid explicit-implicit feature grid. The grid can be rendered to produce images from any desired viewpoint and, since the features are defined in 3D space, the rendered images are consistent across different viewpoints. Compared to utilizing an explicit density grid, the feature representation allows for a lower resolution grid. The latter leads to an easier estimation of the probability density due to a smaller number of variables. Furthermore, the resolution of the grid can be decoupled from the resolution of the rendered images. (ii) We design a new diffusion method that can learn a distribution over such 3D feature grids while only using 2D images for supervision. Specifically, we first generate intermediate 3D-aware features conditioned only on the input posed images. Then, following the standard diffusion model learning, we add noise to this intermediate representation and train a denoising 3D UNet to remove the noise. We apply the denoising loss as photometric error between the rendered images and the Ground-Truth training images. The key advantage of this approach is that it enables training of the 3D diffusion model from 2D images, which are abundant, sidestepping the difficult problem of procuring a huge dataset of 3D models for training.
We train and evaluate our method on the Co3Dv2 dataset where HoloDiffusion outperforms existing alternatives both qualitatively and quantitatively.
Related Work
Neural rendering is a class of algorithms that partially or entirely use neural networks to approximate the light transport equation.
The 2D versions of neural rendering include variants of pix2pix , deferred neural rendering , and their follow-up works. The common theme in all these methods is that a post-processor neural network (usually a CNN) maps neural feature images into photorealistic RGB images.
The 3D versions of neural rendering have recently been popularized by NeRF , which uses a Multi Layer Perceptron (MLP) to model the parameters of the 3D scene (radiance and occupancy) and a physically-based rendering procedure (Emission-Absorption raymarching). NeRF solves the inverse rendering problem where, given many 2D images of a scene, the aim is to recover its 3D shape and appearance. The success of NeRF gave rise to many follow-up works . While NeRF uses MLPs to represent the occupancy and radiance of the underlying scene, different representations were explored in .
Few-view reconstruction.
In many cases, dense image supervision is unavailable, and one has to condition the reconstruction on a small number of scene views instead. Since 3D reconstruction from few-views is ambiguous, recent methods aid the reconstruction with 3D priors learned by observing many images of an object category. Works like CMR , C3DM , and UMR learn to predict the parameters of a mesh by observing images of individual examples of the object category. DOVE also predicts meshes, but additionally leverages stronger constraints provided by videos of the deformable objects.
Others aimed at learning to fit NeRF given only a small number of views; they do so by sampling image features at the 2D projections of the 3D ray samples. Later works have improved this formulation by using transformers to process the sampled features. Finally, ViewFormer drops the rendering model and learns a fully implicit transformer-based new-view synthesizer. Recently, BANMo reconstructed deformable objects with a signed distance function.
2 3D Generative Models
Early 3D generative models leveraged adversarial learning as the primary form of supervision. PlatonicGAN learns to generate colored 3D shapes from an unstructured corpus of images belonging to the same category, by rendering a voxel grid from random viewpoints so that an adversarial discriminator cannot distinguish between the renders and natural images sampled from a large uncurated database. PrGAN differs from PlatonicGAN by focusing only on rendering untextured 3D shapes. To deal with the large memory footprint of voxels, HoloGAN adjusts PlatonicGAN by rendering a low-resolution 2D feature image of a coarse voxel grid, followed by a 2D convolutional decoder mapping the feature render to the final RGB image. The results are, however, not consistent with camera motion.
Inspired by the success of NeRF , GRAF also trains in a data setting similar to PlatonicGAN but, differently from PlatonicGAN, represents each generated scene with a neural radiance field. The GRAF pipeline was subsequently improved by PiGAN , leveraging a SIREN-based architecture. Similar to HoloGAN, StyleNerf first renders a radiance feature field followed by a convolutional super-resolution network. EG3D further improves the pipeline by initially decoding a randomly sampled latent vector to a tri-plane representation followed by a NeRF-style rendering of a radiance field supported by the tri-plane. EpiGRAF further improves upon the triplane-based 3D generation. GAUDI also uses the tri-plane while building upon DeVries et. al. which used a single plane representing the floor map of the indoor room scenes being generated.
Besides radiance fields, other shape representations have also been explored. While VoxGRAF replaces the radiance field of GRAF with a sparse voxel grid, StyleSDF employs signed distance fields, and Neural Volumes propose a novel trilinearly-warped voxel grid. Wu et al. differentiably render meshes and aid the adversarial learning with a set of constraints exploiting symmetry properties of the reconstructed categories. Recently, GET3D differentiably converts an initial tri-plane representation to colored mesh, which is finally rendered.
The aforementioned approaches are trained solely by observing uncurated category-centric image collections without the need for any explicit 3D supervision in form of the ground truth 3D shape or camera pose. However, since the rendering function is non-smooth under camera motion, these methods can either successfully reconstruct image databases with a very small variation in camera poses (e.g., fronto-parallel scenes such as portrait photos of cat or human faces) or datasets with a well-defined distribution of camera intrinsics end extrinsics (e.g., synthetic datasets). We tackle the pose estimation problem by leveraging a dataset of category-centric videos each containing multiple views of the same object. Observing each object from a moving vantage point allows for estimating accurate scene-consistent camera poses that provide strong constraints.
While most approaches focus on generating shapes of isolated instances of object categories, GIRAFFE and BlockGAN extend GRAF and HoloGAN to reconstruct compositions of objects and their background. Alternative approaches focus on text-conditioned 3D shape generation , or learn a generative model by observing a single self-similar scene.
D diffusion models.
Diffusion models for 3D shape learning have been explored only very recently. Luo et al. use full 3D supervision to learn a generative diffusion model of point clouds. In a concurrent effort, Watson et al. learns a new-view synthesis function which, conditioned on a posed image of a scene, generates a new view of the scene from a specified target viewpoint. Differently from us, do not employ an explicit image formation model which may lead to geometrical inconsistencies between generated viewpoints.
HoloDiffusion
We start by discussing the necessary background and notation on diffusion models in Sec. 3.1, and then we introduce our method in Sec. 3.2, Sec. 3.3, and Sec. 3.4.
Given i.i.d. samples from (an unknown) data distribution , the task of generative modeling is to find the parameters of a parametric model that best approximate . Diffusion models are a class of likelihood-based models centered on the idea of defining a forward diffusion (noising) process , for . The noising process converts the data samples into pure noise, i.e., . The model then learns the reverse process , which iteratively converts the noise samples back into data samples starting from the purely Gaussian sample .
The Denoising Diffusion Probabilistic Model (DDPM) , in particular, defines the noising transitions using a Gaussian distribution, setting
Samples can be easily drawn from this distribution by using the reparameterization trick:
One similarly defines the reverse denoising step using a Gaussian distribution:
where, the is the denoising network with learned parameters . The sequence defines the noise schedule as:
We use a linear time schedule with steps.
The denoising must be applied iteratively for sampling the target distribution . However, for training the model we can draw samples directly from as:
It is common to use the network to predict the noise component instead of the signal component in eq. 2; which has the interpretation of modeling the score of the marginal distribution up to a scaled constant . Instead, we employ the “-formulation” from eq. 3, which has recently been explored in the context of diffusion model distillation and modeling of text-conditioned videos using diffusion . The reasoning behind this design choice will become apparent later.
Training the “-formulation” of a diffusion model comprises minimizing the following loss:
encouraging to denoise sample to predict the clean sample .
Sampling.
Once the denoising network is trained, sampling can be done by first starting with pure noise, i.e., , and then iteratively refining times using the network , which terminates with a sample from target data distribution :
2 Learning 3D Categories by Watching Videos
Our goal is to train a generative model where is a representation of the shape and appearance of a 3D object; furthermore, we aim to learn this distribution using only the 2D training videos .
D feature grids.
Next, we discuss how to build a diffusion model for the distribution of feature grids. One might attempt to directly apply the methodology of Sec. 3.1, setting , but this does not work because we have no access to ground-truth feature grids for training; instead, these 3D models must be inferred from the available 2D videos while training. We solve this problem in the next section.
3 Bootstrapped Latent Diffusion Model
In this section, we show how to learn the distribution of feature grids from the training videos alone. In what follows, we use the symbol as a shorthand for .
The training videos provide RGB images and their corresponding camera poses , but no sample feature grids from the target distribution . As a consequence, we also have no access to the noised samples required to evaluate the denoising objective eq. 6 and thus learn a diffusion model.
To solve this issue, we introduce the BLDM (Bootstrapped Latent Diffusion Model). BLDM can learn the denoiser-cum-generator given samples from an auxiliary distribution , which is closely related but not identical to the target distribution .
As shown in fig. 2, our idea is to obtain the auxiliary samples as a (learnable) function of the corresponding training videos . To this end, we use a design strongly inspired by Warp-Conditioned-Embedding (WCE) , which demonstrated compelling performance for learning 3D object categories. Specifically, given a training video containing frames , we generate a grid of auxiliary features by projecting the 3D coordinate of the each grid element to every video frame , sampling corresponding 2D image features, and aggregating those into a single -dimensional descriptor per grid element. The 2D image features are extracted by a trainable encoder (we use the ResNet-32 encoder ) . This process is detailed in the supplementary material.
Auxiliary denoising diffusion objective.
The standard denoising diffusion loss eq. 6 is unavailable in our case because the data samples are unavailable. Instead, we leverage the “-formulation” of diffusion to employ an alternative diffusion objective which does not require knowledge of . Specifically, we replace eq. 6 with a photometric loss
which compares the rendering of the denoising of the noised auxiliary grid to the (known) image with pose . Equation 8 can be computed because the image and camera parameters are known and is derived from the auxiliary sample , whose computation is given in the previous section.
Train/test denoising discrepancy.
Our denoiser takes as input a sample from the noised auxiliary distribution instead of the noised target distribution . While this allows to learn the denoising model by minimizing eq. 8, it prevents us from drawing samples from the model at test time. This is because, during training, learns to denoise the auxiliary samples (obtained through fusing image features into a voxel-grid), but at test time we need instead to draw target samples as specified by eq. 7 per sampling step. We address this problem by using a bootstrapping technique that we describe next.
Two-pass diffusion bootstrapping.
In order to remove the discrepancy between the training and testing sample distributions for the denoiser , we first use the latter to obtain ‘clean’ voxel grids from the training videos during an initial denoising phase, and then apply a diffusion process to those, finetuning as a result.
Our bootstrapping procedure rests on the assumption that once is minimized, the denoisings of the auxiliary grids follow the clean data distribution , i.e., for the optimal denoiser parameters that minimize . Simply put, the denoiser learns to denoise both the diffusion noise and the noise resulting from imperfect reconstructions. Note that our assumption is reasonable since recent single-scene neural rendering methods have demonstrated successful recovery of high-quality 3D shapes solely by optimizing the photometric loss via differentiable rendering.
Given that is now capable of generating clean data samples, we can expose it to the noised version of the clean samples by executing a second denoising pass in a recurrent manner. To this end, we define the bootstrapped photometric loss :
with denoting the diffusion of input grid at time . Intuitively, eq. 9 evaluates the photometric error between the ground truth image and the rendering of the doubly-denoised grid .
4 Implementation Details
HoloDiffusion training finds the optimal model parameters by minimizing the sum of the photometric and the bootstrapped photometric losses using the Adam optimizer with an initial learning rate (decaying ten-fold whenever the total loss plateaus) until convergence is reached.
In each training iteration, we randomly sample 10 source views from a randomly selected training video to form the grid of auxiliary features . The auxiliary features are noised to form and later denoised with . Afterwards is noised and denoised again during the two-pass bootstrap procedure. To avoid two rendering passes in each training iteration (one for and the second for ), we randomly choose to optimize with 50-50 probability in each iteration as a lazy regularization. The photometric losses compare renders of the denoised voxel grid to 3 different target views (different from the source views).
The differentiable rendering function from eqs. 8 and 9 uses Emission-Absorption (EA) ray marching as follows. First, given the knowledge of the camera parameters , a ray is emitted from each pixel of the rendered image . We sample 3D points on each ray at regular intervals . For each point , we sample the corresponding voxel grid feature , where stands for trilinear interpolation. The feature is then decoded by an MLP as with parameters to obtain the density and the RGB color of each 3D point. The MLP is designed so that the color depends on the ray direction while the density does not, similar to NeRF . Finally, EA ray marching renders the ’s pixel color as a weighted combination of the sampled colors. The weights are defined as where .
Experiments
In this section, we evaluate our method. First we perform the quantitative evaluation and then follow it by visualizing samples for assessing the quality of generations.
For our experiments, we use CO3Dv2 , which is currently the largest available dataset of fly-around real-life videos of object categories. The dataset contains videos of different object categories and each video makes a complete circle around the object, showing all sides of it. Furthermore, camera poses and object foreground masks are provided with the dataset (they were obtained by the authors by running off-the-shelf Structure-from-Motion and instance segmentation software, respectively).
We consider the four categories Apple, Hydrant, TeddyBear and Donut for our experiments. For each of the categories we train a single model on the 500 “train” videos (i.e. approx. frames in total) with the highest camera cloud quality score, as defined in the CO3Dv2 annotations, in order to ensure clean ground-truth camera pose information. We note that all trainings were done on -to- V100 32GB GPUs for 2 weeks.
We consider the prior works pi-GAN , EG3D , and GET3D as baselines for comparison. Pi-GAN generates radiance fields represented by MLPs and is trained using an adversarial objective. Similar to our setting, they only use 2D image supervision for training. EG3D uses the feature triplane, decoded by an MLP as the underlying representation, while needing both the images and the camera poses as input to the training procedure. GET3D is another GAN-based baseline, which also requires the images and camera poses for training. Apart from this, GET3D also requires the fg/bg masks for training; which we supply in form of the masks available in CO3Dv2. Since GET3D applies a Deformable Marching Tetrahedra step in the pipeline, the samples generated by them are in the form of textured meshes.
Quantitative evaluation.
We report Frechet Inception Distance (FID) , and Kernel Inception Distance (KID) for assessing the generative quality of our results. As shown in Table 1, our HoloDiffusion produces better scores than EG3D and GET3D. Although pi-GAN gets better scores than ours on some categories, we note that the 3D-agnostic training procedure of pi-GAN cannot recover the proper 3D structure of the unaligned shapes of CO3Dv2. Thus, without the 3D-view consistency, the 3D neural fields (MLPs) produced by pi-GAN essentially mimic a 2D image GAN.
Qualitative evaluation.
Figure 4 depicts random samples generated from all the methods under comparison. HoloDiffusion produces the most appealing, consistent and realistic samples among all. Figure 3 further analyzes the viewpoint consistency of pi-GAN compared to ours. It is evident that, although individual views of pi-GAN samples look realistic, their appearance is inconsistent with the change of viewpoint. Please refer to the project webpage for more examples and videos of the generated samples.
Conclusion
We have presented HoloDiffusion, an unconditional 3D-consistent generative diffusion model that can be trained using only posed-image supervision. At the core of our method is a learnable rendering module that is trained in conjunction with the diffusion denoiser, which operates directly in the feature space. Furthermore, we use a pretrained feature encoder to decouple the cubic volumetric memory complexity from the final image rendering resolution. We demonstrate that the method can be trained on raw posed image sets, even in the few-image setting, striking a good balance between quality and diversity of results.
At present, our method requires access to camera information at training time. One possibility is to jointly train a viewpoint estimator to pose the input images, but the challenge may be to train this module from scratch as the input view distribution is unlikely to be uniform . An obvious next challenge would be to test the setup for conditional generation, either based on images (i.e., single view reconstruction task) or using text guidance. Beyond generation, we would also like to support editing the generated representations, both in terms of shape and appearance, and compose them together towards scene generation. Finally, we want to explore multi-class training where diffusion models, unlike their GAN counterparts, are known to excel without suffering from mode collapse.
Societal Impact
Our method primarily contributes towards the generative modeling of 3D real-captured assets. Thus as is the case with 2D generative models, ours is also prone to misuse of generated synthetic media. In the context of synthetically generated images, our method could potentially be used to make fake 3D view-consistent GIFs or videos. Since we only train our models on the virtually harmless Co3D (Common objects in 3D) dataset, our released models could not be directly used to infer potentially malicious samples.
As diffusion models can be prone to memorizing the training data in limited data settings , our models can also be used to recover the original training samples. Analyzing the severity and the extent to which our models suffer from this, is an interesting future direction for exploration.
Acknowledgements
Animesh and Niloy were partially funded by the European Union’s Horizon 2020 research and innovation programme under the Marie Skłodowska-Curie grant agreement No. 956585. This research has been partly supported by MetaAI and the UCL AI Centre. Finally, Animesh thanks Alexia Jolicoeur-Martineau for the helpful and insightful guidance on diffusion models.
References
Appendix A Views2Voxel-grid Unprojection Mechanism
Given a training video containing frames , we generate a grid of auxiliary features by using the following procedure. We first project the 3D coordinate of each grid element to every video frame and sample corresponding 2D image features. The 2D image features are obtained using a ResNet32 encoder . We use bilinear interpolation for sampling continuous values and use zero-features for projected points that lie outside the Image. Thus, we obtain features (corresponding to each frame in the video) for each grid element of the voxel-grid. We accumulate these features using the Accumulator MLP . The accumulator takes as input , where denotes concatenation and corresponds to the viewing direction corresponding to the camera center of frame, and outputs . Finally, we compute the feature at each of the voxel grid centers as a weighted sum of the newly mapped features:
Appendix B Implementation Details
In this section, we provide more details related to implementing our proposed method.
Our proposed pipeline (Fig 2. of main paper) contains three neural components: The Encoder, Diffusion UNet and Renderer. The Encoder network is a ResNet32 model . For the main diffusion network, we use a 3D variant of the UNet used by Dhariwal and Nichol . The model comprises residual blocks containing downsampling, upsampling, and self-attention blocks (with additive residual connections).
B.2 Renderer
In order to decode the generated voxel-grid of features into density and radiance fields, we use a NeRF-like MLP (Multi-layer perceptron). The MLP contains 4 layers of 256 hidden units with a skip-connection on the 3rd hidden layer. The skip connection concatenates the input features with the intermediate hidden layer features. Similar to NeRF, and for the reasons described in Zhang et al. , we also input the view-directions at a latter layer in the MLP. The input features are not encoded, but we apply sinusoidal encodings to the input viewing directions with max frequency level . The activation functions used are: LeakyReLU for the hidden layers, Softplus for the density output head, and the Sigmoid for the radiance output head. All trainable weights are initialized using the Xavier uniform initialization . Figure I shows the detailed architecture of the RenderMLP.
B.3 Training Details
We train the full HoloDiffusion pipeline for epochs over the dataset containing the object-centric videos. During training, we randomly sample 11 source views for unprojecting into the initial voxel-grid, and 1 target (reserved) novel view for computing loss. The latter enforces 3D structure in the generated samples. We use distance between the rendered views and the G.T. views as the photometric-consistency loss. In terms of hardware, we train all our models on 4-8 32GB-V100 GPUs, with a batch-size equal to the number of GPUs in use, i.e., each GPU processes one voxel-grid during training. We use Adam optimizer with a learning rate () of and default values of , , and for all the trainable networks during training.
B.4 Diffusion Details
We use the DDPM diffusion-formulation for our bootstrap-latent-diffusion module as described in section 4.2 of the main paper. We use the default time-steps and the default schedule in our experiments: wherein we set . Rest of the values are obtained by linearly interpolating between the and . Finally, to improve the input conditioning of our diffusion module, we apply to the voxel features to constrain their values in the range of , as proposed in Karras et al. . This allows us to apply clipping during sampling.