HoloDiffusion: Training a 3D Diffusion Model using 2D Images

Animesh Karnewar, Andrea Vedaldi, David Novotny, Niloy Mitra

Introduction

Diffusion models have rapidly emerged as formidable generative models for images, replacing others (e.g., VAEs, GANs) for a range of applications, including image colorization , image editing , and image synthesis . These models explicitly optimize the likelihood of the training samples, can be trained on millions if not billions of images, and have been shown to capture the underlying model distribution better than previous alternatives.

A natural next step is to bring diffusion models to 3D data. Compared to 2D images, 3D models facilitate direct manipulation of the generated content, result in perfect view consistency across different cameras, and allow object placement using direct handles. However, learning 3D diffusion models is hindered by the lack of a sufficient volume of 3D data for training. A further question is the choice of representation for the 3D data itself (e.g., voxels, point clouds, meshes, occupancy grids, etc.). Researchers have proposed 3D-aware diffusion models for point clouds , volumetric shape data using wavelet features and novel view synthesis . They have also proposed to distill a pretrained 2D diffusion model to generate neural radiance fields of 3D objects . However, a diffusion-based 3D generator model trained using only 2D image for supervision is not available yet.

In this paper, we contribute HoloDiffusion, the first unconditional 3D diffusion model that can be trained with only real posed 2D images. By posed, we mean different views of the same object with known cameras, for example, obtained by means of structure from motion .

We make two main technical contributions: (i) We propose a new 3D model that uses a hybrid explicit-implicit feature grid. The grid can be rendered to produce images from any desired viewpoint and, since the features are defined in 3D space, the rendered images are consistent across different viewpoints. Compared to utilizing an explicit density grid, the feature representation allows for a lower resolution grid. The latter leads to an easier estimation of the probability density due to a smaller number of variables. Furthermore, the resolution of the grid can be decoupled from the resolution of the rendered images. (ii) We design a new diffusion method that can learn a distribution over such 3D feature grids while only using 2D images for supervision. Specifically, we first generate intermediate 3D-aware features conditioned only on the input posed images. Then, following the standard diffusion model learning, we add noise to this intermediate representation and train a denoising 3D UNet to remove the noise. We apply the denoising loss as photometric error between the rendered images and the Ground-Truth training images. The key advantage of this approach is that it enables training of the 3D diffusion model from 2D images, which are abundant, sidestepping the difficult problem of procuring a huge dataset of 3D models for training.

We train and evaluate our method on the Co3Dv2 dataset where HoloDiffusion outperforms existing alternatives both qualitatively and quantitatively.

Related Work

Neural rendering is a class of algorithms that partially or entirely use neural networks to approximate the light transport equation.

The 2D versions of neural rendering include variants of pix2pix , deferred neural rendering , and their follow-up works. The common theme in all these methods is that a post-processor neural network (usually a CNN) maps neural feature images into photorealistic RGB images.

The 3D versions of neural rendering have recently been popularized by NeRF , which uses a Multi Layer Perceptron (MLP) to model the parameters of the 3D scene (radiance and occupancy) and a physically-based rendering procedure (Emission-Absorption raymarching). NeRF solves the inverse rendering problem where, given many 2D images of a scene, the aim is to recover its 3D shape and appearance. The success of NeRF gave rise to many follow-up works . While NeRF uses MLPs to represent the occupancy and radiance of the underlying scene, different representations were explored in .

Few-view reconstruction.

In many cases, dense image supervision is unavailable, and one has to condition the reconstruction on a small number of scene views instead. Since 3D reconstruction from few-views is ambiguous, recent methods aid the reconstruction with 3D priors learned by observing many images of an object category. Works like CMR , C3DM , and UMR learn to predict the parameters of a mesh by observing images of individual examples of the object category. DOVE also predicts meshes, but additionally leverages stronger constraints provided by videos of the deformable objects.

Others aimed at learning to fit NeRF given only a small number of views; they do so by sampling image features at the 2D projections of the 3D ray samples. Later works have improved this formulation by using transformers to process the sampled features. Finally, ViewFormer drops the rendering model and learns a fully implicit transformer-based new-view synthesizer. Recently, BANMo reconstructed deformable objects with a signed distance function.

2 3D Generative Models

Early 3D generative models leveraged adversarial learning as the primary form of supervision. PlatonicGAN learns to generate colored 3D shapes from an unstructured corpus of images belonging to the same category, by rendering a voxel grid from random viewpoints so that an adversarial discriminator cannot distinguish between the renders and natural images sampled from a large uncurated database. PrGAN differs from PlatonicGAN by focusing only on rendering untextured 3D shapes. To deal with the large memory footprint of voxels, HoloGAN adjusts PlatonicGAN by rendering a low-resolution 2D feature image of a coarse voxel grid, followed by a 2D convolutional decoder mapping the feature render to the final RGB image. The results are, however, not consistent with camera motion.

Inspired by the success of NeRF , GRAF also trains in a data setting similar to PlatonicGAN but, differently from PlatonicGAN, represents each generated scene with a neural radiance field. The GRAF pipeline was subsequently improved by PiGAN , leveraging a SIREN-based architecture. Similar to HoloGAN, StyleNerf first renders a radiance feature field followed by a convolutional super-resolution network. EG3D further improves the pipeline by initially decoding a randomly sampled latent vector to a tri-plane representation followed by a NeRF-style rendering of a radiance field supported by the tri-plane. EpiGRAF further improves upon the triplane-based 3D generation. GAUDI also uses the tri-plane while building upon DeVries et. al. which used a single plane representing the floor map of the indoor room scenes being generated.

Besides radiance fields, other shape representations have also been explored. While VoxGRAF replaces the radiance field of GRAF with a sparse voxel grid, StyleSDF employs signed distance fields, and Neural Volumes propose a novel trilinearly-warped voxel grid. Wu et al. differentiably render meshes and aid the adversarial learning with a set of constraints exploiting symmetry properties of the reconstructed categories. Recently, GET3D differentiably converts an initial tri-plane representation to colored mesh, which is finally rendered.

The aforementioned approaches are trained solely by observing uncurated category-centric image collections without the need for any explicit 3D supervision in form of the ground truth 3D shape or camera pose. However, since the rendering function is non-smooth under camera motion, these methods can either successfully reconstruct image databases with a very small variation in camera poses (e.g., fronto-parallel scenes such as portrait photos of cat or human faces) or datasets with a well-defined distribution of camera intrinsics end extrinsics (e.g., synthetic datasets). We tackle the pose estimation problem by leveraging a dataset of category-centric videos each containing multiple views of the same object. Observing each object from a moving vantage point allows for estimating accurate scene-consistent camera poses that provide strong constraints.

While most approaches focus on generating shapes of isolated instances of object categories, GIRAFFE and BlockGAN extend GRAF and HoloGAN to reconstruct compositions of objects and their background. Alternative approaches focus on text-conditioned 3D shape generation , or learn a generative model by observing a single self-similar scene.

D diffusion models.

Diffusion models for 3D shape learning have been explored only very recently. Luo et al. use full 3D supervision to learn a generative diffusion model of point clouds. In a concurrent effort, Watson et al. learns a new-view synthesis function which, conditioned on a posed image of a scene, generates a new view of the scene from a specified target viewpoint. Differently from us, do not employ an explicit image formation model which may lead to geometrical inconsistencies between generated viewpoints.

HoloDiffusion

We start by discussing the necessary background and notation on diffusion models in Sec. 3.1, and then we introduce our method in Sec. 3.2, Sec. 3.3, and Sec. 3.4.

Given NN i.i.d. samples {xi}i=1N\{x^{i}\}_{i=1}^{N} from (an unknown) data distribution p(x)p(x), the task of generative modeling is to find the parameters θ\theta of a parametric model pθ(x)p_{\theta}(x) that best approximate p(x)p(x). Diffusion models are a class of likelihood-based models centered on the idea of defining a forward diffusion (noising) process q(xt∣xt−1)q(x_{t}|x_{t-1}), for t∈[0,T]t\in[0,T]. The noising process converts the data samples into pure noise, i.e., q(xT)≈q(xT∣xT−1)=N(0,I)q(x_{T})\approx q(x_{T}|x_{T-1})=\mathcal{N}(0,I). The model then learns the reverse process p(xt−1∣xt)p(x_{t-1}|x_{t}), which iteratively converts the noise samples back into data samples starting from the purely Gaussian sample xTx_{T}.

The Denoising Diffusion Probabilistic Model (DDPM) , in particular, defines the noising transitions using a Gaussian distribution, setting

Samples can be easily drawn from this distribution by using the reparameterization trick:

One similarly defines the reverse denoising step using a Gaussian distribution:

where, the Dθ\mathcal{D}_{\theta} is the denoising network with learned parameters θ\theta. The sequence αt\alpha_{t} defines the noise schedule as:

We use a linear time schedule with T=1000T=1000 steps.

The denoising must be applied iteratively for sampling the target distribution p(x)p(x). However, for training the model we can draw samples xtx_{t} directly from q(xt∣x0)q(x_{t}|x_{0}) as:

It is common to use the network Dθ(xt,t)\mathcal{D}_{\theta}(x_{t},t) to predict the noise component ϵ\epsilon instead of the signal component xt−1x_{t-1} in eq. 2; which has the interpretation of modeling the score of the marginal distribution q(xt)q(x_{t}) up to a scaled constant . Instead, we employ the “x0x_{0}-formulation” from eq. 3, which has recently been explored in the context of diffusion model distillation and modeling of text-conditioned videos using diffusion . The reasoning behind this design choice will become apparent later.

Training the “x0x_{0}-formulation” of a diffusion model Dθ\mathcal{D}_{\theta} comprises minimizing the following loss:

encouraging Dθ\mathcal{D}_{\theta} to denoise sample xt∼N(αˉtx0,(1−αˉt)I)x_{t}\sim\mathcal{N}(\sqrt{\bar{\alpha}_{t}}x_{0},(1-\bar{\alpha}_{t})I) to predict the clean sample x0x_{0}.

Sampling.

Once the denoising network Dθ\mathcal{D}_{\theta} is trained, sampling can be done by first starting with pure noise, i.e., xT∼N(0,I)x_{T}\sim\mathcal{N}(0,I), and then iteratively refining TT times using the network Dθ\mathcal{D}_{\theta}, which terminates with a sample from target data distribution x0∼q(x0)=p(x)x_{0}\sim q(x_{0})=p(x):

2 Learning 3D Categories by Watching Videos

Our goal is to train a generative model p(V)p(V) where VV is a representation of the shape and appearance of a 3D object; furthermore, we aim to learn this distribution using only the 2D training videos {si}i=1N\{s^{i}\}_{i=1}^{N}.

D feature grids.

Next, we discuss how to build a diffusion model for the distribution p(V)p(V) of feature grids. One might attempt to directly apply the methodology of Sec. 3.1, setting x=Vx=V, but this does not work because we have no access to ground-truth feature grids VV for training; instead, these 3D models must be inferred from the available 2D videos while training. We solve this problem in the next section.

3 Bootstrapped Latent Diffusion Model

In this section, we show how to learn the distribution p(V)p(V) of feature grids from the training videos ss alone. In what follows, we use the symbol V\mathcal{V} as a shorthand for p(V)p(V).

The training videos provide RGB images II and their corresponding camera poses PP, but no sample feature grids VV from the target distribution V\mathcal{V}. As a consequence, we also have no access to the noised samples Vt∼N(αˉtV0,(1−αˉt)I)V_{t}\sim\mathcal{N}(\sqrt{\bar{\alpha}_{t}}V_{0},(1-\bar{\alpha}_{t})I) required to evaluate the denoising objective eq. 6 and thus learn a diffusion model.

To solve this issue, we introduce the BLDM (Bootstrapped Latent Diffusion Model). BLDM can learn the denoiser-cum-generator Dθ\mathcal{D}_{\theta} given samples Vˉ∼Vˉ\bar{V}\sim\mathcal{\bar{V}} from an auxiliary distribution Vˉ\mathcal{\bar{V}}, which is closely related but not identical to the target distribution V\mathcal{V}.

As shown in fig. 2, our idea is to obtain the auxiliary samples Vˉ\bar{V} as a (learnable) function of the corresponding training videos ss. To this end, we use a design strongly inspired by Warp-Conditioned-Embedding (WCE) , which demonstrated compelling performance for learning 3D object categories. Specifically, given a training video ss containing frames IjI_{j}, we generate a grid Vˉ∈dV×S×S×S\bar{V}\in{}^{d^{V}\times S\times S\times S} of auxiliary features Vˉ:mno∈dV\bar{V}_{:mno}\in^{d_{V}} by projecting the 3D coordinate xmnoV\mathbf{x}^{V}_{mno} of the each grid element (m,n,o)(m,n,o) to every video frame IjI_{j}, sampling corresponding 2D image features, and aggregating those into a single dVd_{V}-dimensional descriptor per grid element. The 2D image features are extracted by a trainable encoder (we use the ResNet-32 encoder ) EE. This process is detailed in the supplementary material.

Auxiliary denoising diffusion objective.

The standard denoising diffusion loss eq. 6 is unavailable in our case because the data samples VV are unavailable. Instead, we leverage the “x0x_{0}-formulation” of diffusion to employ an alternative diffusion objective which does not require knowledge of VV. Specifically, we replace eq. 6 with a photometric loss

which compares the rendering rζ(Dθ(Vˉt,t),Pj)r_{\zeta}(\mathcal{D}_{\theta}(\bar{V}_{t},t),P_{j}) of the denoising Dθ(Vˉt,t)\mathcal{D}_{\theta}(\bar{V}_{t},t) of the noised auxiliary grid Vˉt\bar{V}_{t} to the (known) image IjI_{j} with pose PjP_{j}. Equation 8 can be computed because the image IjI_{j} and camera parameters PjP_{j} are known and Vˉt\bar{V}_{t} is derived from the auxiliary sample Vˉ\bar{V}, whose computation is given in the previous section.

Train/test denoising discrepancy.

Our denoiser Dθ\mathcal{D}_{\theta} takes as input a sample Vˉt\bar{V}_{t} from the noised auxiliary distribution Vˉt\bar{\mathcal{V}}_{t} instead of the noised target distribution Vt\mathcal{V}_{t}. While this allows to learn the denoising model by minimizing eq. 8, it prevents us from drawing samples from the model at test time. This is because, during training, Dθ\mathcal{D}_{\theta} learns to denoise the auxiliary samples Vˉ∈Vˉ\bar{V}\in\mathcal{\bar{V}} (obtained through fusing image features into a voxel-grid), but at test time we need instead to draw target samples V∈VV\in\mathcal{V} as specified by eq. 7 per sampling step. We address this problem by using a bootstrapping technique that we describe next.

Two-pass diffusion bootstrapping.

In order to remove the discrepancy between the training and testing sample distributions for the denoiser Dθ\mathcal{D}_{\theta}, we first use the latter to obtain ‘clean’ voxel grids from the training videos during an initial denoising phase, and then apply a diffusion process to those, finetuning Dθ\mathcal{D}_{\theta} as a result.

Our bootstrapping procedure rests on the assumption that once Lphoto\mathcal{L}_{\text{photo}} is minimized, the denoisings Dθ(Vˉt,t)\mathcal{D}_{\theta}(\bar{V}_{t},t) of the auxiliary grids Vˉ∼Vˉ\bar{V}\sim\mathcal{\bar{V}} follow the clean data distribution V\mathcal{V}, i.e., Dθ⋆(Vˉt,t)∼V\mathcal{D}_{\theta^{\star}}(\bar{V}_{t},t)\sim\mathcal{V} for the optimal denoiser parameters θ⋆\theta^{\star} that minimize Lphoto\mathcal{L}_{\text{photo}}. Simply put, the denoiser Dθ\mathcal{D}_{\theta} learns to denoise both the diffusion noise and the noise resulting from imperfect reconstructions. Note that our assumption Dθ⋆(Vˉt,t)∼V\mathcal{D}_{\theta^{\star}}(\bar{V}_{t},t)\sim\mathcal{V} is reasonable since recent single-scene neural rendering methods have demonstrated successful recovery of high-quality 3D shapes solely by optimizing the photometric loss via differentiable rendering.

Given that Dθ⋆\mathcal{D}_{\theta^{\star}} is now capable of generating clean data samples, we can expose it to the noised version of the clean samples VV by executing a second denoising pass in a recurrent manner. To this end, we define the bootstrapped photometric loss Lphoto′\mathcal{L}^{\prime}_{\text{photo}}:

with ϵt′(Z)∼N(αˉt′Z,(1−αˉt′)I)\epsilon_{t^{\prime}}(Z)\sim\mathcal{N}(\sqrt{\bar{\alpha}_{t^{\prime}}}Z,(1-\bar{\alpha}_{t^{\prime}})I) denoting the diffusion of input grid ZZ at time t′t^{\prime}. Intuitively, eq. 9 evaluates the photometric error between the ground truth image II and the rendering of the doubly-denoised grid Dθ(ϵt′(Dθ(Vˉt,t),t′))\mathcal{D}_{\theta}(\epsilon_{t^{\prime}}(\mathcal{D}_{\theta}(\bar{V}_{t},t),t^{\prime})).

4 Implementation Details

HoloDiffusion training finds the optimal model parameters θ,ζ\theta,\zeta by minimizing the sum of the photometric and the bootstrapped photometric losses Lphoto+Lphoto′\mathcal{L}_{\text{photo}}+\mathcal{L}^{\prime}_{\text{photo}} using the Adam optimizer with an initial learning rate 5⋅10−55\cdot 10^{-5} (decaying ten-fold whenever the total loss plateaus) until convergence is reached.

In each training iteration, we randomly sample 10 source views {Ij}\{I_{j}\} from a randomly selected training video sis^{i} to form the grid of auxiliary features Vˉ\bar{V}. The auxiliary features are noised to form Vˉt\bar{V}_{t} and later denoised with Dθ(Vˉt)\mathcal{D}_{\theta}(\bar{V}_{t}). Afterwards Dθ(Vˉt)\mathcal{D}_{\theta}(\bar{V}_{t}) is noised and denoised again during the two-pass bootstrap procedure. To avoid two rendering passes in each training iteration (one for Lphoto\mathcal{L}_{\text{photo}} and the second for Lphoto′\mathcal{L}^{\prime}_{\text{photo}}), we randomly choose to optimize Lphoto′\mathcal{L}^{\prime}_{\text{photo}} with 50-50 probability in each iteration as a lazy regularization. The photometric losses compare renders rζ(⋅,Pj)r_{\zeta}(\cdot,P_{j}) of the denoised voxel grid to 3 different target views (different from the source views).

The differentiable rendering function rζ(V,Pj)r_{\zeta}(V,P_{j}) from eqs. 8 and 9 uses Emission-Absorption (EA) ray marching as follows. First, given the knowledge of the camera parameters PjP_{j}, a ray ru∈S2\mathbf{r}_{\mathbf{u}}\in\mathcal{S}^{2} is emitted from each pixel u∈{0,…,H−1}×{0,…,W−1}\mathbf{u}\in\{0,\dots,H-1\}\times\{0,\dots,W-1\} of the rendered image I^j∈3×H×W\hat{I}_{j}\in{}^{3\times H\times W}. We sample NSN_{S} 3D points (pi)i=1NS(\mathbf{p}_{i})_{i=1}^{N_{S}} on each ray at regular intervals Δ∈ℜ\Delta\in\real. For each point pi\mathbf{p}_{i}, we sample the corresponding voxel grid feature V[pi]∈dVV[\mathbf{p}_{i}]\in{}^{d^{V}}, where V[⋅]V[\cdot] stands for trilinear interpolation. The feature V[pi]V[\mathbf{p}_{i}] is then decoded by an MLP as fζ(V[pi],ru):=(σi,ci)f_{\zeta}(V[\mathbf{p}_{i}],\mathbf{r}_{\mathbf{u}}):=(\sigma_{i},\mathbf{c}_{i}) with parameters ζ\zeta to obtain the density σi∈\sigma_{i}\in and the RGB color ci∈3\mathbf{c}_{i}\in^{3} of each 3D point. The MLP ff is designed so that the color c\mathbf{c} depends on the ray direction ru\mathbf{r}_{\mathbf{u}} while the density σ\sigma does not, similar to NeRF . Finally, EA ray marching renders the ru\mathbf{r}_{\mathbf{u}}’s pixel color cru=∑i=1NSw(pi)ci\mathbf{c}_{\mathbf{r}_{\mathbf{u}}}=\sum_{i=1}^{N_{S}}w(\mathbf{p}_{i})\mathbf{c}_{i} as a weighted combination of the sampled colors. The weights are defined as w(pi)=Ti−Ti+1w(\mathbf{p}_{i})=T_{i}-T_{i+1} where Ti=e−∑1i−1σiΔT_{i}=e^{-\sum_{1}^{i-1}\sigma_{i}\Delta}.

Experiments

In this section, we evaluate our method. First we perform the quantitative evaluation and then follow it by visualizing samples for assessing the quality of generations.

For our experiments, we use CO3Dv2 , which is currently the largest available dataset of fly-around real-life videos of object categories. The dataset contains videos of different object categories and each video makes a complete circle around the object, showing all sides of it. Furthermore, camera poses and object foreground masks are provided with the dataset (they were obtained by the authors by running off-the-shelf Structure-from-Motion and instance segmentation software, respectively).

We consider the four categories Apple, Hydrant, TeddyBear and Donut for our experiments. For each of the categories we train a single model on the 500 “train” videos (i.e. approx. 500×100500\times 100 frames in total) with the highest camera cloud quality score, as defined in the CO3Dv2 annotations, in order to ensure clean ground-truth camera pose information. We note that all trainings were done on 22-to-88 V100 32GB GPUs for 2 weeks.

We consider the prior works pi-GAN , EG3D , and GET3D as baselines for comparison. Pi-GAN generates radiance fields represented by MLPs and is trained using an adversarial objective. Similar to our setting, they only use 2D image supervision for training. EG3D uses the feature triplane, decoded by an MLP as the underlying representation, while needing both the images and the camera poses as input to the training procedure. GET3D is another GAN-based baseline, which also requires the images and camera poses for training. Apart from this, GET3D also requires the fg/bg masks for training; which we supply in form of the masks available in CO3Dv2. Since GET3D applies a Deformable Marching Tetrahedra step in the pipeline, the samples generated by them are in the form of textured meshes.

Quantitative evaluation.

We report Frechet Inception Distance (FID) , and Kernel Inception Distance (KID) for assessing the generative quality of our results. As shown in Table 1, our HoloDiffusion produces better scores than EG3D and GET3D. Although pi-GAN gets better scores than ours on some categories, we note that the 3D-agnostic training procedure of pi-GAN cannot recover the proper 3D structure of the unaligned shapes of CO3Dv2. Thus, without the 3D-view consistency, the 3D neural fields (MLPs) produced by pi-GAN essentially mimic a 2D image GAN.

Qualitative evaluation.

Figure 4 depicts random samples generated from all the methods under comparison. HoloDiffusion produces the most appealing, consistent and realistic samples among all. Figure 3 further analyzes the viewpoint consistency of pi-GAN compared to ours. It is evident that, although individual views of pi-GAN samples look realistic, their appearance is inconsistent with the change of viewpoint. Please refer to the project webpage for more examples and videos of the generated samples.

Conclusion

We have presented HoloDiffusion, an unconditional 3D-consistent generative diffusion model that can be trained using only posed-image supervision. At the core of our method is a learnable rendering module that is trained in conjunction with the diffusion denoiser, which operates directly in the feature space. Furthermore, we use a pretrained feature encoder to decouple the cubic volumetric memory complexity from the final image rendering resolution. We demonstrate that the method can be trained on raw posed image sets, even in the few-image setting, striking a good balance between quality and diversity of results.

At present, our method requires access to camera information at training time. One possibility is to jointly train a viewpoint estimator to pose the input images, but the challenge may be to train this module from scratch as the input view distribution is unlikely to be uniform . An obvious next challenge would be to test the setup for conditional generation, either based on images (i.e., single view reconstruction task) or using text guidance. Beyond generation, we would also like to support editing the generated representations, both in terms of shape and appearance, and compose them together towards scene generation. Finally, we want to explore multi-class training where diffusion models, unlike their GAN counterparts, are known to excel without suffering from mode collapse.

Societal Impact

Our method primarily contributes towards the generative modeling of 3D real-captured assets. Thus as is the case with 2D generative models, ours is also prone to misuse of generated synthetic media. In the context of synthetically generated images, our method could potentially be used to make fake 3D view-consistent GIFs or videos. Since we only train our models on the virtually harmless Co3D (Common objects in 3D) dataset, our released models could not be directly used to infer potentially malicious samples.

As diffusion models can be prone to memorizing the training data in limited data settings , our models can also be used to recover the original training samples. Analyzing the severity and the extent to which our models suffer from this, is an interesting future direction for exploration.

Acknowledgements

Animesh and Niloy were partially funded by the European Union’s Horizon 2020 research and innovation programme under the Marie Skłodowska-Curie grant agreement No. 956585. This research has been partly supported by MetaAI and the UCL AI Centre. Finally, Animesh thanks Alexia Jolicoeur-Martineau for the helpful and insightful guidance on diffusion models.

References

Appendix A Views2Voxel-grid Unprojection Mechanism

Given a training video ss containing frames IjI_{j}, we generate a grid Vˉ∈dV×S×S×S\bar{V}\in{}^{d^{V}\times S\times S\times S} of auxiliary features Vˉ:mno∈dV\bar{V}_{:mno}\in^{d_{V}} by using the following procedure. We first project the 3D coordinate xmnoV\mathbf{x}^{V}_{mno} of each grid element (m,n,o)(m,n,o) to every video frame IjI_{j} and sample corresponding 2D image features. The 2D image features fmnojf_{mno}^{j} are obtained using a ResNet32 encoder E(Ij)E(I_{j}). We use bilinear interpolation for sampling continuous values and use zero-features for projected points that lie outside the Image. Thus, we obtain NframesN_{\text{frames}} features (corresponding to each frame in the video) for each grid element of the voxel-grid. We accumulate these features using the Accumulator MLP Aacc\mathcal{A}_{acc}. The accumulator Aacc\mathcal{A}_{acc} takes as input [fmnoj;vj][f_{mno}^{j};v^{j}], where [;][;] denotes concatenation and vjv^{j} corresponds to the viewing direction corresponding to the camera center of jthj^{\text{th}} frame, and outputs [σmnoj;f′mnoj][\sigma^{j}_{mno};{f^{\prime}}_{mno}^{j}]. Finally, we compute the feature at each of the voxel grid centers as a weighted sum of the newly mapped features:

Appendix B Implementation Details

In this section, we provide more details related to implementing our proposed method.

Our proposed pipeline (Fig 2. of main paper) contains three neural components: The Encoder, Diffusion UNet and Renderer. The Encoder network is a ResNet32 model . For the main diffusion network, we use a 3D variant of the UNet used by Dhariwal and Nichol . The model comprises residual blocks containing downsampling, upsampling, and self-attention blocks (with additive residual connections).

B.2 Renderer

In order to decode the generated voxel-grid of features into density and radiance fields, we use a NeRF-like MLP (Multi-layer perceptron). The MLP contains 4 layers of 256 hidden units with a skip-connection on the 3rd hidden layer. The skip connection concatenates the input features with the intermediate hidden layer features. Similar to NeRF, and for the reasons described in Zhang et al. , we also input the view-directions at a latter layer in the MLP. The input features are not encoded, but we apply sinusoidal encodings to the input viewing directions with max frequency level L=4L=4. The activation functions used are: LeakyReLU for the hidden layers, Softplus for the density output head, and the Sigmoid for the radiance output head. All trainable weights are initialized using the Xavier uniform initialization . Figure I shows the detailed architecture of the RenderMLP.

B.3 Training Details

We train the full HoloDiffusion pipeline for 10001000 epochs over the dataset containing the object-centric videos. During training, we randomly sample 11 source views for unprojecting into the initial voxel-grid, and 1 target (reserved) novel view for computing loss. The latter enforces 3D structure in the generated samples. We use L2L2 distance between the rendered views and the G.T. views as the photometric-consistency loss. In terms of hardware, we train all our models on 4-8 32GB-V100 GPUs, with a batch-size equal to the number of GPUs in use, i.e., each GPU processes one voxel-grid during training. We use Adam optimizer with a learning rate (α\alpha) of 0.000050.00005 and default values of β1\beta_{1}, β2\beta_{2}, and ϵ\epsilon for all the trainable networks during training.

B.4 Diffusion Details

We use the DDPM diffusion-formulation for our bootstrap-latent-diffusion module as described in section 4.2 of the main paper. We use the default t=1000t=1000 time-steps and the default βt\beta_{t} schedule in our experiments: wherein we set β0=0.0001;β999=0.02\beta_{0}=0.0001;\beta_{999}=0.02. Rest of the βt\beta_{t} values are obtained by linearly interpolating between the β0\beta_{0} and β999\beta_{999}. Finally, to improve the input conditioning of our diffusion module, we apply tanhtanh to the voxel features to constrain their values in the range of , as proposed in Karras et al. . This allows us to apply clipping during sampling.