Single-Stage Diffusion NeRF: A Unified Approach to 3D Generation and Reconstruction
Hansheng Chen, Jiatao Gu, Anpei Chen, Wei Tian, Zhuowen Tu, Lingjie Liu, Hao Su
Introduction
Synthesizing 3D visual contents has gained significant attention in computer vision and graphics, thanks to advances in neural rendering and generative models. Although numerous methods have emerged to handle individual tasks, such as single-/multi-view 3D reconstruction and 3D content generation, it remains a major challenge to develop a comprehensive framework that bridges the state of the art of multiple tasks. For instance, neural radiance fields (NeRF) have shown impressive results in novel view synthesis by solving the inverse rendering problem via per-scene fitting, which is suitable for dense-view inputs but difficult to generalize to sparse observations. In contrast, many sparse-view 3D reconstruction methods rely on feed-forward image-to-3D encoders, but they are unable to handle ambiguity in the occluded region and generate crisp images. Regarding unconditional generation, 3D-aware generative adversarial networks (GAN) are partially limited in their usage of single-image discriminators, which cannot reason cross-view relationships to effectively learn from multi-view data.
In this paper, we propose a unified approach to various 3D tasks (Fig. 1) by developing a holistic model that learns generalizable 3D priors from multi-view images. Inspired by the success of 2D diffusion models , we present the Single-Stage Diffusion NeRF (SSDNeRF), which models the generative prior of scene latent codes with a 3D latent diffusion model (LDM).
While similar LDMs have been applied in 2D and 3D generation in previous work , they typically require two-stage training, where the first stage pretrains the variational auto-encoders (VAE) or auto-decoders without diffusion models. In the case of diffusion NeRFs, however, we argue that two-stage training induces noisy patterns and artifacts in the latent code due to the uncertain nature of inverse rendering, particularly when training from sparse-view data, which prevents the diffusion model from learning a clean latent manifold effectively. To address this issue, we introduce a novel single-stage training paradigm that enables end-to-end learning of diffusion and NeRF weights (§ 4.1). This approach blends the generative and the rendering biases coherently for improved performance overall and allows for training on sparse-view data. Additionally, we show that the learned 3D priors of unconditional diffusion models can be exploited for flexible test-time scene sampling from arbitrary observations (§ 4.2).
We evaluate SSDNeRF on multiple datasets of categorical single-object scenes, demonstrating strong performance overall. Our approach represents a significant step towards a unified framework for various 3D tasks.
To summarize, our main contributions are as follows:
We introduce SSDNeRF, a unified approach to all-round performance in unconditional 3D generation and image-based reconstruction;
We propose a novel single-stage training paradigm that jointly learns NeRF reconstruction and diffusion model from multi-view images of a large number of objects. Notably, this enables training on as sparse as three views per scene, which is previously infeasible;
A guidance-finetuning sampling scheme is developed to exploit the learned diffusion priors for 3D reconstruction from arbitrary number of views at test time.
Related Work
The generative adversarial framework has been successfully adapted for 3D generation by integrating projection-based rendering into the generator. A variety of 3D representations have been explored previously, including point clouds, cuboids, spheres and voxels in early works, the more recent radiance fields and feature fields with volume renderer, and differentiable surface with mesh renderer. The above methods are all trained with 2D image discriminators that are unable to reason cross-view relationships, making them heavily dependent on model bias for 3D consistency. As a result, multi-view data cannot be effectively exploited to learn complex and diverse geometries. 3D GANs are mostly applied in unconditional generation. Although 3D completion from images is possible through GAN inversion , faithfulness is not guaranteed due to limited latent expressiveness, as shown in .
View-Conditioned Regression and Generation
Sparse-view 3D reconstruction can be tackled by regressing novel views from input images. Various architectures have been proposed to encode images into volume features, which can be projected to supervised target views through volume rendering. However, they cannot reason ambiguity and generate diverse and meaningful contents, which often leads to blurry results. In contrast, image-conditioned generative models are better at synthesizing distinct contents. 3DiM proposes to generate novel views from a view-conditioned image diffusion model, but the model lacks 3D consistency bias. distill priors of image-conditioned 2D diffusion models into NeRFs to enforce 3D constraints. These methods are parallel to our track as they model cross-view relationships in the image space, while our model is inherently 3D.
Auto-Decoders and Diffusion NeRF
NeRF’s per-scene fitting scheme can be generalized to multi-scene fitting by sharing part of the parameters across all scenes, leaving the rest as individual scene codes . Therefore, multi-scene NeRFs can be trained as auto-decoders , where the code bank and shared decoder weights are jointly learned. With proper architectures, scene codes can be treated as latents with Gaussian priors, allowing 3D completion and even generation . However, like 3D GANs, the latents are not expressive enough for faithful reconstruction of detailed objects. improve upon vanilla auto-decoders with latent diffusion priors. DiffRF leverages the diffusion prior to perform 3D completion. These methods train the auto-decoders and diffusion models in two separate stages, which is subject to the limitations in § 3.2.
Background
With this objective, the model is trained as an auto-decoder , where the scene codes can be interpreted as the latent codes, and the plenoptic function can be regarded as a decoder in the form of , assuming independent Gaussians.
An auto-decoder with trained weights can perform unconditional generation by decoding latent codes drawn from a Gaussian prior . However, to ensure continuity in generation, a low-dimensional latent space and a complex decoder is required, which adds to the difficulty in optimizing the latent code to faithfully reconstruct any given views.
2 Latent Diffusion Models
Latent diffusion models (LDM) learn a prior distribution in the latent space with parameters , which enables the usage of more expressive latent representations, such as 2D grids for images . For neural field generation, previous work adopts a two-stage training scheme, where the auto-decoder is trained first to obtain the per-scene latent , which is then treated as real data to train the LDM. The LDM injects Gaussian perturbation into the code , yielding a noisy code at diffusion time step , under empirical noise schedule functions . A denoising network with trainable weights is then tasked with removing the noise from to predict a denoised code . The network is typically trained with a simplified L2 denoising loss:
where , is an empirical time dependent weighting function, and formulates the time-conditioned denoising network.
With trained weights , one can sample from the diffusion prior using a variety of solvers (e.g., DDIM ) that recursively denoise , starting from random Gaussian noise , until reaching the denoised state . Moreover, the sampling process can be guided by the gradients of the rendering loss against known observations, allowing 3D reconstruction from images at test time .
Limitations of Two-Stage Training for 3D Tasks
While LDMs with 2D image VAEs are typically trained in two stages , training LDMs with NeRF auto-decoders poses an unprecedented challenge. An expressive latent code is underdetermined when obtained via rendering-based optimization, leading to noisy patterns that distract denoising networks (top-left of Fig. 2). Additionally, reconstructing NeRFs from sparse views without a learned prior is exceptionally difficult (bottom-left of Fig. 2), limiting training to dense-views settings.
Proposed Method
To build a holistic model that unifies 3D generation and reconstruction, we propose SSDNeRF, a framework that conjoins the expressive triplane NeRF auto-decoder with a triplane latent diffusion model. Fig. 3 provides an overview of the model. In the following subsections, we elaborate on how training and testing are performed in detail.
Single-stage training constrains scene codes with both terms in the loss function, allowing the learned prior to complete the parts unseen to rendering. This is particularly beneficial to training on sparse-view data, where the expressive triplane codes are severely underdetermined.
Comparison to Two-Stage Generative Neural Fields
2 Image-Guided Sampling and Finetuning
To achieve generalizable test-time NeRF reconstruction that covers a wide spectrum from single-view to dense observations, we propose performing image-guided sampling and then finetuning the sampled codes considering both the diffusion prior and rendering likelihood.
Following the reconstruction-guided sampling method by Ho et al. , we compute the approximated rendering gradients w.r.t. a noisy code , defined as:
where \mathopen{}\mathclose{{}\left(\alpha^{(t)}/\sigma^{(t)}}\right)^{2\omega} is an additional weighting factor based on signal-to-noise ratio (SNR), with hyperparameter chosen to be 0.5 or 0.25 in our work. The guidance gradients are then combined with unconditional score prediction, expressed as a correction to the denoising output :
We observe that the reconstruction guidance alone cannot strictly enforce rendering constraints towards faithful reconstruction. To address this issue, we reuse the training objective in Eq. (4) to finetune the sampled scene code , while freezing the diffusion and decoder parameters:
While finetuning with rendering loss is common in view-conditioned NeRF regression methods , our finetuning approach differs in the use of diffusion prior loss on the 3D scene code, which significantly enhances generalization to novel views, as demonstrated in § 5.3.
3 Implementation Details
This subsection briefly describes some important technical details. More details can be found in the supplementary.
Denoising Parameterization and Weighting
The denoising model is implemented as a U-Net as in DDPM , with a total of 122M parameters. Its input and output are noisy and denoised triplane features, respectively, with channels of all three planes stacked together. For the prediction format, we adopt the -parameterization in , such that . Regarding the weighting function in the diffusion loss in Eq. (2), LSGM employs two different mechanisms for optimizing latents and diffusion weights , respectively, which we find unstable with NeRF auto-decoders. Instead, we observe that the SNR-based weighting w^{(t)}=\mathopen{}\mathclose{{}\left(\alpha^{(t)}/\sigma^{(t)}}\right)^{2\omega} used in Eq. (5) works well with our models.
Experiments
We conduct experiments on the ShapeNet SRN and Amazon Berkeley Objects (ABO) Tables datasets for benchmarking with previous work. The SRN dataset provides single-object scenes in two categories, i.e., Cars and Chairs, with a train/test split of 2458/703 for Cars and 4612/1317 for Chairs. Each train scene has 50 random views from a sphere and each test scene has 251 spiral views from the upper hemisphere. The ABO Tables dataset provides a train/test split of 1520/156 table scenes, where each scene has 91 views from the upper hemisphere. For both datasets, we use the provided renderings (resized to 128×128) with ground truth poses for training and testing.
2 Unconditional Generation
In this section, we conduct evaluations for unconditional generation using the SRN Cars and ABO Tables dataset. The Cars dataset poses a challenge in generating sharp and intricate textures, whereas the Tables dataset comprises of diverse geometries with realistic materials. Models are trained on all images of the training set for 1M iterations.
For SRN Cars, following Functa , we sample 704 scenes from the diffusion model, and render each scene using the fixed 251 camera poses from the test set. For ABO Tables, following DiffRF , we sample 1000 scenes and render each scene with 10 random cameras. We adopt standard generation metrics including Fréchet Inception Distance (FID) and Kernel Inception Distance (KID) . The metrics’ reference sets are all images in the test set for SRN Cars and all images in the entire dataset for ABO Tables, respectively.
Comparison to the State of the Art
As shown in Table 7, on SRN Cars, SSDNeRF (1-stage) outperforms EG3D in KID (a more suitable measure for small datasets) by a clear margin. Meanwhile, its FID is drastically better than Functa, which uses an LDM but with low dimensional latent codes. On ABO Tables, SSDNeRF shows significantly better performance than EG3D and DiffRF.
Single- vs. Two-stage
On SRN Cars, we compare the proposed single-stage training against two-stage training with tuned TV regularization using the same model architecture. The results in Table 7 indicate substantial advantage of single-stage training (KID/10−3 3.47 vs. 6.38).
Qualitative Results
As shown in Fig. 4, SSDNeRF generates more regular geometries than the slightly skewed and distorted shapes by EG3D . Compared to DiffRF , our method produces sharp details and reflective materials, thanks to our more expressive model with latents of higher spatial resolution and view-dependent NeRF decoder.
3 Sparse-View NeRF Reconstruction
This section presents experiments on 3D reconstruction from sparse-view images of unseen objects in SRN Cars and Chairs test sets. The Cars dataset presents the challenge of recovering distinct textures, while the Chairs dataset requires accurate reconstruction of diverse shapes. Models are trained on all images of the training set for 80K iterations, as we find that longer schedule leads to decaying performance in reconstructing unseen objects. This behaviour is in accordance with the interpolation results in § 5.5.
We use the evaluation protocol and metrics in PixelNeRF . Given input images sampled from each test scene, we obtain the triplane scene code via guidance-finetuning and evaluate novel view synthesis quality with respect to the unseen images. The image quality metrics include average peak signal-to-noise-ratio (PSNR), structural similarity (SSIM) , and Learned Perceptual Image Patch Similarity (LPIPS) . In addition, we evaluate the FID between all synthesized images and ground truth images as in 3DiM .
Comparison to the State of the Art
Table 2 compares SSDNeRF against previous approaches in single-view and two-view reconstruction settings. Overall, SSDNeRF reaches the best LPIPS of all tasks, indicating the best perceptual fidelity. In contrast, 3DiM generates high quality images (best FID) but with the lowest fidelity to the ground truth (lowest PSNR); CodeNeRF reports the best PSNR on single-view Cars, but its limited expressiveness leads to blurry outputs (Fig. 5) and less competitive LPIPS and FID; VisionNeRF achieves a balanced performance on all single-view metrics, but may struggle to generate textural details on the unseen side of cars (e.g., the other side of the ambulance in Fig. 5). Moreover, SSDNeRF exhibits a clear advantage in two-view reconstruction, achieving the best performance on all relevant metrics.
Single- vs. Two-stage
As demonstrated in Table 3, the model trained in a single stage (A0) outperforms the same architecture trained in two stages with TV regularization (A1) in all metrics of single-view reconstruction.
Ablation Studies on Test-Time Finetuning
As shown in Table 3, we evaluate the effectiveness of test-time finetuning and the contribution of the learned diffusion prior with two ablation experiments: (A2) removing the diffusion loss during finetuning and using only the rendering loss, and (A3) omitting the finetuning process entirely. The results indicate that finetuning with single-view rendering loss provides only marginal improvements over guided sampling (A2 vs. A3), while the learned diffusion prior significantly boosts the LPIPS and FID scores (A0 vs. A2), highlighting its importance in recovering sharp and distinct contents. Moreover, the qualitative results in Fig. 5 reveal that views with higher overlap to the input view benefit the most from finetuning, meeting our expectation that finetuning helps faithfully reconstruct the exact observations.
Sparse-to-Dense Reconstruction
To validate that SSDNeRF seamlessly bridges sparse- and dense-view NeRF reconstruction, we evaluate its novel view synthesis performance with the number of input views varying from 1 to 32. We compare our model to the triplane NeRF baseline trained as an auto-decoder with optional TV regularization instead of diffusion prior. Meanwhile, we also evaluate CodeNeRF , an auto-decoder with 256-d latent codes. The results in Fig. 6 show that SSDNeRF excels in all settings, especially in 1 to 4 views. In contrast, CodeNeRF is outperformed by vanilla triplane NeRF with more views.
4 Training SSDNeRF on Sparse-View Dataset
In this section, we train SSDNeRF on a sparse-view subset of the full SRN Cars training set, in which a fixed set of only three views are randomly picked from each scene. Note that a reasonable decline in performance compared to dense-view training is expected as the whole training dataset has been reduced to 6% of its original size.
We adopt a training trick that resets the triplane codes to their mean value halfway through training. This helps to prevent the model from getting stuck in a local minimum that overfits geometric artifacts. We also double the length of the training schedule accordingly. The model achieves a decent FID of 19.04±1.10 and a KID/10−3 of 8.28±0.60. Results are visualized in Fig. 7.
Single-View Reconstruction
We adopt the same training strategy as in § 5.3. With our guidance-finetuning approach, the model achieves an LPIPS score of 0.106, even outperforming most of the previous methods in Table 2 that use the full training set.
Comparison to TV Regularization
Fig. 8 (b) shows the RGB images and geometries represented by the scene latent codes learned from three views during training. By comparison, vanilla triplane auto-decoder with TV regularization (Fig. 8 (a)) often fails to reconstruct a scene from sparse views, leading to severe geometric artifacts. As a result, previously it has been infeasible to train two-stage models with expressive latents on sparse-view data.
5 NeRF Interpolation
Following DDIM , we can sample two initial values , interpolate them using spherical linear interpolation , and then use the deterministic solver to obtain interpolated samples. However, as noted by , standard Gaussian diffusion models often result in non-smooth interpolation. In SSDNeRF (with results shown in Fig. 9), we find that the model (a) trained with early stopping for sparse-view reconstruction produces reasonably smooth transitions between samples, while the model (b) trained with a longer schedule for unconditional generation produces distinct yet discontinuous samples. This suggests that early stopping preserves a smoother prior, leading to better generalization for sparse-view reconstruction.
Conclusion
In this paper, we propose SSDNeRF, which combines the diffusion model and NeRF representation through a novel single-stage training paradigm with an end-to-end justifiable loss. Notably, it overcomes the limitations in previous work where implicit neural fields must be obtained from dense observations first, before training the diffusion models to learn their manifold. With strong performance on multiple benchmarks, SSDNeRF demonstrates a significant advancement towards a unified framework for general 3D content manipulation.
Currently, our method relies on ground truth camera parameters during both training and testing. Future work may explore transform-invariant models. Additionally, the diffusion prior can become discontinuous with prolonged training, which affects generalization. Although early stopping is temporarily used, a better network design or a larger training dataset may be able to address this problem fundamentally.
Acknowledgements
We thank Norman Müller for sharing the baseline results on ABO Tables. Hansheng Chen and Wei Tian acknowledge the funding by the National Natural Science Foundation of China (No. 52002285), the Shanghai Science and Technology Commission (No. 21ZR1467400), the original research project of Tongji University (No. 22120220593), the National Key R&D Program of China (No. 2021YFB2501104), and the Natural Science Foundation of Chongqing (No. 2023NSCQ-MSX4511).
References
A Details on Batch-Wise Rendering Loss
B Implementation and Hyperparameters
We implement our models using PyTorch and MMGeneration toolkit . Our NeRF renderer is based on a public codebase torch-ngp , which employs a density-based grid pruning strategy for efficient real-time rendering.
B.2 Hyperparameters
The major difference between unconditional- and reconstruction-purposed models is the training schedule, where reconstruction-purposed training stops early at 80K iterations, as mentioned in the main paper. Other differences lie in the U-Net dropout rate and latent learning rate, which may have marginal effects on the reconstruction performance.
Regarding the Langevin correction step in the form of with step size and independent noise , we observe that this technique is more effective in reconstructing Chairs than Cars. Therefore, to reduce inference time, Langevin correction is not used for SRN Cars dataset. Our intuition is that Chairs dataset exhibits higher variety in geometry, and Langevin correction helps better explore the latent space by injecting random noising during sampling.
B.3 Training and Inference Time
We train all our models using two RTX 3090 GPUs, each processing a batch of 8 scenes. On average, a single outer training step takes around 0.5 sec, 80K iterations take around 11 hours, and 1M iterations cost around 6 days.
C Additional Model Details
In the interest of reproducibility, this section provides additional details about the models used in our experiments. These techniques were not discussed in the main paper, because they are not essential components of the proposed method, and they seem to have negligible effect on the overall results (Table 5). Nevertheless, we have included them in our implementation to maintain consistency with an earlier version of our codebase where they were found to be useful at one stage.
In an earlier version of our implementation of the diffusion model, we use the prediction format as in DDPM instead of the current format proposed by . To stabilize denoising-based sampling process, the format requires clipping the denoised prediction at each step, which is suitable for bounded data. This motivated us to bound the latent code element-wise via an additional Tanh layer.
Because our final models have switched to the prediction format, Tanh mapping may not be an essential component of SSDNeRF, as indicated in Table 5.
C.2 Additional L2 Regularization
L2 latent regularization in auto-decoder training originates from the assumed Gaussian latent prior . In two-stage diffusion NeRF or occupancy field models, L2 regularization helps control the norm of the latent codes and discourage outlying values with respect to the clipping during sampling. During single-stage training and test-time finetuning, we also keep this regularization term in the actual loss function:
D Experiment Details and Additional Results
Table 5 presents more details on the experiment settings, testing hyperparameters, and evaluation results of sparse-to-dense reconstruction on SRN Cars dataset.
Overall, we find that more iterations and higher learning rate are required when finetuning on more input views, but the learning rate should not exceed the upper bound of 0.08 for stability, and a maximum of 200 outer loop iterations (totaling 1600 inner loop iterations) are sufficient for dense-view settings.
D.2 Single-View Reconstruction from Real Images
In this subsection, we provide addition experiments on single-view NeRF reconstruction from real images, using the model trained on the synthetic SRN Cars dataset. This demonstrates the generalization capability of SSDNeRF under substantial domain gap.
We extract images of vehicles from the KITTI 3D object detection dataset , which provides annotated 3D bounding boxes of objects in the camera view. We use the provided ground truth bounding box dimensions and poses to align the objects in the same world coordinate system as in SRN Cars dataset. In addition, we leverage the segmentation masks annotated by Heylen et al. to remove the background. All images are cropped and resized to 128×128. In real applications, one could also use a monocular 3D object detector and an instance segmentation model to obtain these inputs.
Testing Hyperparameters
Qualitative Results and Failure Case
D.3 Addition Qualitative Examples
We show randomly sampled scenes generated by SSDNeRF in Figure 12, Figure 13, and Figure 14. For single-view reconstruction, we compare the novel views predicted by SSDNeRF to those predicted by CodeNeRF and VisionNeRF in Figure 15 and Figure 16.