3D Neural Field Generation using Triplane Diffusion

J. Ryan Shue, Eric Ryan Chan, Ryan Po, Zachary Ankner, Jiajun Wu, Gordon Wetzstein

Introduction

Diffusion models have seen rapid progress, setting state-of-the-art (SOTA) performance across a variety of image generation tasks. While most diffusion methods model 2D images, recent work has attempted to develop denoising methods for 3D shape generation. These 3D diffusion methods operate on discrete point clouds and, while successful, exhibit limited quality and resolution.

In contrast to 2D diffusion, which directly leverages the image as the target for the diffusion process, it is not directly obvious how to construct such 2D targets in the case of 3D diffusion. Interestingly, recent work on 3D-aware generative adversarial networks (GANs) (see Sec. 2 for an overview) has demonstrated impressive results for 3D shape generation using 2D generators. We build upon this idea of learning to generate triplane representations that encode 3D scenes or radiance fields as a set of axis-aligned 2D feature planes. The structure of a triplane is analogous to that of a 2D image and can be used as part of a 3D generative method that leverages conventional 2D generator architectures.

Inspired by recent efforts in designing efficient 3D GAN architectures, we introduce a neural field-based diffusion framework for 3D representation learning. Our approach follows a two-step process. In the first step, a training set of 3D scenes is factored into a set of per-scene triplane features and a single, shared feature decoder. In the second step, a 2D diffusion model is trained on these triplanes. The trained diffusion model can then be used at inference time to generate novel and diverse 3D scenes. By interpreting triplanes as multi-channel 2D images and thus decoupling generation from rendering, we can leverage current (and likely future) SOTA 2D diffusion model backbones nearly out of the box. Fig. 1 illustrates how a single object is generated with our framework (top), and how two generated objects—even with different topologies—can be interpolated (bottom).

We introduce a generative framework for diffusion on 3D scenes that utilizes 2D diffusion model backbones and has a built-in 3D inductive bias.

We show that our approach is capable of generating both high-fidelity and diverse 3D scenes that outperform state-of-the-art 3D GANs.

Related Work

Implicit neural representations, or neural fields, hold the SOTA for 3D scene representation . They either solely learn geometry or use posed images to jointly optimize geometry and appearance . Neural fields represent scenes as continuous functions, allowing them to scale well with scene complexity compared to their discrete counterparts . Initial methods used a single, large multilayer perceptron (MLP) to represent entire scenes , but reconstruction with this approach can be computationally inefficient because training such a representation requires thousands of forward passes through the large model per scene. Recent years have shown a trend towards locally conditioned representations, which either learn local functions or locally modulate a shared function with a hybrid explicit–implicit representation . These methods use small MLPs, which are efficient during inference and significantly better at capturing local scene details. We adopt the expressive hybrid triplane representation introduced by Chan et al. . Triplanes are efficient, scaling with the surface area rather than volume, and naturally integrate with expressive, fine-tuned 2D generator architectures. We modify the triplane representation for compatibility with our denoising framework.

Generative synthesis in 2D and 3D.

Some of the most popular generative models include GANs , autoregressive models , score matching models , and denoising diffusion probabilistic models (DDPMs) . DDPMs are arguably the SOTA approach for synthesizing high-quality and diverse 2D images . Moreover, GANs can be difficult to train and suffer from issues like mode collapse whereas diffusion models train stably and have been shown to better capture the full training distribution.

In 3D, however, GANs still outperform alternative generative approaches . Some of the most successful 3D GANs use an expressive 2D generator backbone (e.g., StyleGAN2 ) to synthesize triplane representations which are then decoded with a small, efficient MLP . Because the decoder is small and must generalize across many local latents, these methods assign most of their expressiveness to the powerful backbone. In addition, these methods treat the triplane as a multi-channel image, allowing the generator backbone to be used almost out of the box.

Current 3D diffusion models are still very limited. They either denoise a single latent or do not utilize neural fields at all, opting for a discrete point-cloud-based approach. For example, concurrently developed single-latent approaches generate a global latent for conditioning the neural field, relying on a 3D decoder to transform the scene representation from 1D to 3D without directly performing 3D diffusion. As a result, the diffusion model does not actually operate in 3D, losing this important inductive bias and generating blurry results. Point-cloud-based approaches , on the other hand, give the diffusion model explicit 3D control over the shape, but limit its resolution and scalability due to the coarse discrete representation. While showing promise, both 1D-to-3D and point cloud diffusion approaches require specific architectures that cannot easily leverage recent advances in 2D diffusion models.

In our work, we propose to directly generate triplanes with out-of-the-box SOTA 2D diffusion models, granting the diffusion model near-complete control over the generated neural field. Key to our approach is our treatment of well-fit triplanes in a shared latent space as ground truth data for training our diffusion model. We show that the latent space of these triplanes is grounded spatially in local detail, giving the diffusion model a critical inductive bias for 3D generation. Our approach gives rise to an expressive 3D diffusion model.

Triplane Diffusion Framework

Here, we explain the architecture of our neural field diffusion (NFD) model for 3D shapes. In Section 3.1, we explain how we can represent the occupancy field of a single object using a triplane. In Section 3.2, we describe how we can extend this framework to represent an entire dataset of 3D objects. In Section 3.3, we describe the regularization techniques that we found necessary to achieve optimal results. Finally, Sections 3.4 and 3.5 illustrate training and sampling from our model. For an overview of the pipeline at inference, see Figure 3.

The feature planes and mlp can be jointly optimized to represent the occupancy field of a shape.

2 Representing a Class of Objects with Triplanes

We aim to convert our dataset of shapes into a dataset of triplanes so that we can train a diffusion model on these learned feature planes. However, because the MLP and feature planes are typically jointly learned, we cannot simply train a triplane for each object of the dataset individually. If we did, the MLP’s corresponding to each object in our dataset would fail to generalize to triplanes generated by our diffusion model. Therefore, instead of training triplanes for each object in isolation, we jointly optimize the feature planes for many objects simultaneously, along with a decoder that is shared across all objects. This joint optimization results in a dataset of optimized feature planes and an MLP capable of interpreting any triplane from the dataset distribution. Thus, at inference, we can use this MLP to decode feature planes generated by our model.

In practice, during training, we are given a dataset of II objects, and we preprocess the coordinates and ground-truth occupancy values of JJ points per object. Typically, J=10MJ=10\text{M}, where 5M points are sampled uniformly throughout the volume and 5M points are sampled near the object surface. Our naive training objective is a simple L2L2 loss between predicted occupancy values \scnf(i)(xj(i))\textrm{\sc{nf}}^{(i)}(\mathbf{x}^{(i)}_{j}) and ground-truth occupancy values \scoj(i)\textrm{\sc o}^{(i)}_{j} for each point, where xj(i)\mathbf{x}^{(i)}_{j} denotes the jjth point from the iith scene:

During training, we optimize Equation 2 for a shared MLP parameterized by ϕ\phi, as well as the feature planes corresponding to every object in our dataset:

3 Regularizing Triplanes for Effective Generalization

Following the procedure outlined in the previous section, we can learn a dataset of triplane features and a shared triplane decoder; we can then train a diffusion model on these triplane features and sample novel shapes at inference. Unfortunately, the result of this naive training procedure is a generative model for triplanes that produces shapes with significant artifacts.

We find it necessary to regularize the triplane features during optimization to simplify the data manifold that the diffusion model must learn. Therefore, we include total variation (TV) regularization terms with weight λ1\lambda_{1} in the loss function to ensure that the feature planes of each training scene do not contain spurious high-frequency information. This strategy makes the distribution of triplane features more similar to the manifold of natural images (see supplement), which we found necessary to robustly train a diffusion model on them (see Sec. 4).

While the trained feature values are unbounded, our DDPM backbone requires training inputs with values in the range . We address this by normalizing the feature planes before training, but this process is sensitive to outliers. As a result, we include an L2 regularization term on the triplane features with weight λ2\lambda_{2} to discourage outlying values.

We also include an explicit density regularization (EDR) term. Due to our ground-truth occupancy data being concentrated on the surface of the shapes, there is often insufficient data to learn a smooth outside-of-shape volume. Our EDR term combats this issue by sampling a set of random points from the volume, offsetting the points by a random vector ω\boldsymbol{\omega}, feeding both sets through the MLP, and calculating the mean squared error. Notationally, this term can be represented as EDR(\scnf(x),ω)=∥\scnf(x)−\scnf(x+ω)∥22\textrm{EDR}\left(\textrm{\sc{nf}}\left(\mathbf{x}\right),\boldsymbol{\omega}\right)=\|\textrm{\sc{nf}}\left(\mathbf{x}\right)-\textrm{\sc{nf}}\left(\mathbf{x}+\boldsymbol{\omega}\right)\|_{2}^{2}. We find this term necessary to remove floating artifacts in the volume (see Sec. 4)

Our training objective, with added regularization terms, is as follows:

4 Training a Diffusion Model for Triplane Features

The forward or diffusion processes is a Markov chain that gradually adds Gaussian noise to the triplane features, according to a variance schedule β1,β2,…,βT\beta_{1},\beta_{2},\ldots,\beta_{T}

The goal of training a diffusion model is to learn the reverse process. For this purpose, a function approximator ϵθ\boldsymbol{\epsilon}_{\theta} is needed that predicts the noise ϵ∼N(0,I)\epsilon\sim\mathcal{N}\left(\mathbf{0},\mathbf{I}\right) from its noisy input. Typically, this function approximator is implemented as a variant of a convolutional neural network defined by its parameters θ\theta. Following , we train our triplane diffusion model by optimizing the simplified variant of the variational bound on negative log-likelihood:

where tt is sampled uniformly between 1 and TT.

5 Sampling Novel 3D Shapes

The unconditional generation of shapes at inference is a two-stage process that involves sampling a triplane from the trained diffusion model and then querying the neural field.

where ϵ∼N(0,I)\boldsymbol{\epsilon}\sim\mathcal{N}(\mathbf{0},\mathbf{I}) for all but the very last step (i.e., t=1t=1), at which ϵ=0\boldsymbol{\epsilon}=0 and σt2=βt\sigma_{t}^{2}=\beta_{t}.

We use the marching cubes algorithm to extract meshes from the resulting neural fields. Note that our framework is largely agnostic to the diffusion backbone used; we choose to use ADM , a 2D state-of-the-art diffusion model.

Source code and pre-trained models will be made available.

Experiments

To compare NFD against existing 3D generative methods, we train our model on three object categories from the ShapeNet dataset individually. Consistent with previous work , we choose the categories: cars, chairs and airplanes. Each mesh is normalized to lie within 3^{3} and then passed through watertighting. The generation of ground truth triplanes then works as follows: we precompute the occupancies of 10M points per object, where 5M points are distributed uniformly at random in the volume, and 5M points are sampled within a 0.01 distance from the mesh surface. We then train an MLP jointly with as many triplanes as we can fit in the GPU memory of a single A6000 GPU. In our case, we initially train on the first 500 objects in the dataset. After this initial joint optimization, we freeze the shared MLP and use it to optimize the triplanes of the remaining objects in the dataset. All triplanes beyond the first 500 are optimized individually with the same shared MLP; thus, the training of these triplanes can be effectively parallelized.

Evaluation metrics.

As in , we choose to evaluate our model using an adapted version of Fréchet inception distance (FID) that utilizes rendered shading images of our generated meshes. Shading-image FID overcomes limitations of other mesh-based evaluation metrics such as the light-field-descriptor (LFD) by taking human perception into consideration. Zheng et al. provide a detailed discussion of the various evaluation metrics for 3D generative models. Following the method , shading images of each shape are rendered from 20 distinct views; FID is then compared across each view and averaged to obtain a final score:

where gg and rr represent the generated and training datasets, while μi,\Upsigmai\mu^{i},\Upsigma^{i} represent the mean and covariance matrices for shading images rendered from the ithi^{\text{th}} view, respectively.

Along with FID, we also report precision and recall scores using the method proposed by Sajjadi et al. . While FID correlates well with perceived image quality, the one-dimensional nature of the metric prevents it from identifying different failure modes. Sajjadi et al. aim to disentangle FID into separate metrics known as precision and recall, where the former correlates to the quality of the generated images and the latter represents the diversity of the generative model.

Baselines.

We compare our method against state-of-the-art point-based and neural-field-based 3D generative models, namely PVD and SDF-StyleGAN . For evaluation, we use the pre-trained models for both methods on the three ShapeNet categories listed above. Note that PVD is inherently a point-based generative method and therefore does not output a triangle mesh needed for shading image rendering. To circumvent this, we choose to convert generated point clouds to triangle meshes using the ball-pivoting algorithm .

Results.

We provide qualitative results, comparing samples generated by our method to samples generated by baselines, in Figure 4. Our method generates a diverse and finely detailed collection of objects. Objects produced by our method contain sharp edges and features that we would expect to be difficult to accurately reconstruct—note that delicate features, such as the suspension of cars, the slats in chairs, and armaments of planes, are faithfully generated. Perhaps more importantly, samples generated by our model are diverse—our model successfully synthesizes many different types of cars, chairs, and planes, including reproductions of several varieties that we would expect to be rare in the training dataset.

In comparison, while PVD also produces a wide variety of shapes, it is limited by its nature to generating only coarse object shapes. Furthermore, because PVD produces a fixed-size point cloud with only 2048 points, it cannot synthesize fine elements.

SDF-StyleGAN creates high-fidelity shapes, accurately reproducing many details, such as airplane engines and chair legs. However, our method is more capable of capturing very fine features. Note that while SDF-StyleGAN smooths over the division between tire and wheel well when generating cars, our method faithfully portrays this gap. Similarly, our method synthesizes the tails and engines of airplanes, and the legs and planks of chairs, with noticeably better definition. Our method also apparently generates a greater diversity of objects than SDF-StyleGAN. While SDF-StyleGAN capably generates varieties of each ShapeNet class, our method reproduces the same classes with greater variation. This is expected, as a noted advantage of diffusion models over GANs is better mode coverage.

We provide quantitative results in Table 1. The metrics tell a similar story to the qualitative results. Quantitatively, NFD outperforms all baselines in FID, precision, and recall for each ShapeNet category. FID is a standard one-number metric for evaluating generative models, and our performance under this evaluation indicates the generally better quality of object renderings. Precision evaluates the renderings’ fidelity, and recall evaluates their diversity. Outperforming baselines in both precision and recall suggest that our model produces higher fidelity of shapes and a more diverse distribution of shapes. This is consistent with the qualitative results in Figure 4, where our method produced sharper and more complex objects while also covering more modes.

Semantically meaningful interpolation.

Figure 5 shows latent space interpolation between pairs of generated neural fields. As shown in prior work , smooth interpolation in the latent space of diffusion models can be achieved by interpolation between noise tensors before they are iteratively denoised by the model. As in their method, we sample from our trained model using a deterministic DDIM, and we use spherical interpolation so that the intermediate latent noise retains the same distribution. Our method is capable of smooth latent space interpolation in the generated triplanes and their corresponding neural fields.

1 Ablation Studies

We validate the design of our framework by ablating components of our regularization strategies using the cars dataset.

As discussed by Park et al. , the precision of the ground truth decoded meshes is limited by the finite number of point samples guiding the training of the decision boundaries. Because we rely on a limited number of pre-computed coordinate–occupancy pairs to train our triplanes, it is easy to overfit to this limited training set. Even when optimizing a single triplane in isolation (i.e., without learning a generative model), this overfitting manifests in “floater” artifacts in the optimized neural field. Figure 6 shows an example where we fit a single triplane with and without density regularization. Without density regularization, the learned occupancy field contains significant artifacts; with density regularization, the learned occupancy field captures a clean object.

Triplane regularization.

Regularization of the triplanes is essential for training a well-behaved diffusion model. Figure 7 compares generated samples produced by our entire framework, with and without regularization terms. If we train only with Equation 2, i.e., without regularization terms, we can optimize a dataset of triplane features and train a diffusion model to generate samples. However, while the surfaces of the optimized shapes will appear real, the triplane features themselves will have many high-frequency artifacts, and these convoluted feature images are a difficult manifold for even a powerful diffusion model to learn. Consequently, generated triplane features produced by a trained diffusion model decode into shapes with significant artifacts. We note that these artifacts are present only in generated samples; shapes directly factored from the ground-truth shapes are artifact-free, even without regularization.

Training with Equation 4 introduces TV, L2, and density regularizing factors. Triplanes learned with these regularization terms are noticeably smoother, with frequency distributions that more closely align with those found in natural images (see supplement). As we would expect, a diffusion model more readily learns the manifold of regularized triplane features. Samples produced by a diffusion model trained on these regularized shapes decode into convincing and artifact-free shapes.

Discussion

In summary, we introduce a 3D-aware diffusion model that uses a 2D diffusion backbone to generate triplane feature maps, which are assembled into 3D neural fields. Our approach improves the quality and diversity of generated objects over existing 3D-aware generative models by a large margin.

Similarly to other generative methods, training a diffusion model is slow and computationally demanding. Diffusion models, including ours, are also slow to evaluate, whereas GANs, for example, can be evaluated in real-time once trained. Luckily, our method will benefit from improvements to 2D diffusion models in this research area. Slow sampling at inference could be addressed by more efficient samplers and potentially enable real-time synthesis.

Future Work.

We have demonstrated an effective way to generate occupancy fields, but in principle, our approach can be extended to generating any type of neural field that can be represented by a triplane. In particular, triplanes have already been shown to be excellent representations for radiance fields, so it seems natural to extend our diffusion approach to generating NeRFs. While we demonstrate successful results for unconditional generation, conditioning our generative model on text, images, or other input would be an exciting avenue for future work.

Ethical Considerations.

Generative models, including ours, could be extended to generate DeepFakes. These pose a societal threat, and we do not condone using our work to generate fake images or videos of any person intending to spread misinformation or tarnish their reputation.

Conclusion.

3D-aware object synthesis has many exciting applications in vision and graphics. With our work, which is among the first to connect powerful 2D diffusion models and 3D object synthesis, we take a significant step towards utilizing emerging diffusion models for this goal.

Acknowledgements

We thank Vincent Sitzmann for valuable discussions. This project was in part supported by Samsung, the Stanford Institute for Human-Centered AI (HAI), the Stanford Center for Integrated Facility Engineering (CIFE), NSF RI #2211258, Autodesk, and a PECASE from the ARO.

References