NeRDi: Single-View NeRF Synthesis with Language-Guided Diffusion as General Image Priors

Congyue Deng, Chiyu "Max'' Jiang, Charles R. Qi, Xinchen Yan, Yin Zhou, Leonidas Guibas, Dragomir Anguelov

Introduction

Novel view synthesis is a long-existing problem in computer vision and computer graphics. Recent progresses in neural rendering such as NeRFs have made huge strides in novel view synthesis. Given a set of multi-view images with known camera poses, NeRFs represent a static 3D scene as a radiance field parametrized by a neural network, which enables rendering at novel views with the learned network. A line of work has been focusing on reducing the required inputs to NeRF reconstructions, ranging from dense inputs with calibrated camera poses to sparse images with noisy or without camera poses . Yet the problem of NeRF synthesis from one single view remains challenging due to its ill-posed nature, as the one-to-one correspondence from a 2D image to a 3D scene does not exist. Most existing works formulate this as a reconstruction problem and tackle it by training a network to predict the NeRF parameters from the input image . But they require matched multiview images with calibrated camera poses as supervision, which is inaccessible in many cases such as images from the Internet or captured by non-expert users with mobile devices. Recent attempts have been focused on relaxing this constraint by using unsupervised training with novel-view adversarial losses and self-consistency . But they still require the test cases to follow the training distribution which limits their generalizability. There is also work that aggregates priors learned on synthetic multi-view datasets and transfers them to in-the-wild images using data distillation. But they are missing fine details with poor generalizability to unseen categories.

Despite the difficulty of 2D-to-3D mapping for computers, it is actually not a difficult task for human beings. Humans gain knowledge of the 3D world through daily observations and form a common sense of how things should look like and should not look like. Given a specific image, they can quickly narrow down their prior knowledge to the visual input. This makes humans good at solving ill-posed perception problems like single-view 3D reconstruction. Inspired by this, we propose a single-image NeRF synthesis framework without 3D supervision by leveraging large-scale diffusion-based 2D image generation model (Figure 1). Given an input image, we optimize for a NeRF by minimizing an image distribution loss for arbitrary-view renderings with the diffusion model conditioned on the input image. An unconstrained image diffusion is the ‘general prior’ which is inclusive but also vague. To narrow down the prior knowledge and relate it to the input image, we design a two-section semantic feature as the conditioning input to the diffusion model. The first section is the image caption which carries the overall semantics; the second is a text embedding extracted from the input image with textual inversion , which captures additional visual cues. These two sections of language guidance facilitate our realistic NeRF synthesis with semantic and visual coherence between different views. In addition, we introduce a geometric loss based on the estimated depth of the input view for regularizing the underlying 3D structure. Learned with all the guidance and constraints, our model is able to leverage the general image prior and perform zero-shot NeRF synthesis on single image inputs. Experimental results show that we can generate high quality novel views from diverse in-the-wild images. To summarize, our key contributions are:

We formulate single-view reconstruction as a conditioned 3D generation problem and propose a single-image NeRF synthesis framework without 3D supervision, using 2D priors from diffusion models trained on large image datasets.

We design a two-section semantic guidance to narrow down the general prior knowledge conditioned on the input image, enforcing synthesized novel views to be semantically and visually coherent.

We introduce a geometric regularization term on estimated depth maps with 3D uncertainties.

We validate our zero-shot novel view synthesis results on the DTU MVS dataset, achieving higher quality than supervised baselines. We also demonstrate our capability of generating novel-view renderings with high visual quality on in-the-wild images.

Related Work

The recently proliferating NeRF representation has shown great success in novel view synthesis, which is a long-existing task in computer graphics and vision. Combining differentiable rendering with neural network scene parametrizations, NeRF is able to recover the underlying 3D scene from a collection of posed images and render it at novel views realistically. A number of follow-up works have been focusing on relaxing NeRF inputs to less informative data such as unposed images or sparse views . As less data gives rise to a more complex optimization landscape, a variety of regularization losses have been studied, for example: RegNeRF regularizes the geometry and appearance of patches, DDP and DS-NeRF regularize the depth maps, DietNeRF enforces semantic consistency between views by minimizing a CLIP feature loss, and GNeRF adopts a patch-based adversarial loss. Another line of work learns NeRF-based novel-view prediction for few- or single-image inputs by pre-training a scene prior on a large dataset of 3D scenes containing dense views . With additional self-supervision techniques such as equivariance or cycle-consistency , the learning of scene priors can be done simply from sparse- or single-view data, or even purely from unposed image collections with an image adversarial loss . These two lines of works both have their specialties and constraints: the first is generalizable to any scene configurations, but is also less competitive in the more challenging scenarios such as single-image novel view synthesis with high quality requirements; the second, on the other hand, has strong ability of inferring unseen novel views from very limited inputs, but is also restricted to certain scene categories modeled by their scene priors learned from the training data. In our work, we leverage a diffusion-based image prior for NeRF synthesis that is general enough for modeling variations of in-the-wild images while having the adaptivity to each specific input image.

Diffusion-based generative models

Denoising diffusion probabilistic models , or score-based generative models , have recently caught a surge of interests due to their simple designs and excellent performances across a variety of computer vision tasks such as image generation , completion , and editing . In visual content creation, language-guided image diffusion models such as DALL-E2 , Imagen and Stable Diffusion have shown great success in generating photo realistic images with strong semantic correlation to the given text-prompt inputs. In additional to the success of 2D image diffusion models, more recent works have also extend diffusion models to 3D content generation. generate 3D pointclouds with point diffusions. 3DiM shows uncertainty-aware novel view synthesis with image diffusions conditioned on input views and poses, but it does not have guaranteed multiview consistency as no underlying 3D representation is adopted. More related to ours are DreamFusion and GAUDI that also generate NeRFs with diffusions: generates NeRFs under language guidance by optimizing for their renderings at randomly sampled views with a 2D image diffusion model ; trains a diffusion model on the latent space of NeRF scenes, but the learned scene distribution is limited to a set of indoor 3D scenes and does not generalize to in-the-wild images. Similar to , we also leverage 2D image diffusions to optimize for the NeRF renderings at novel views, but instead of unconstrained NeRF generation with user-specified language inputs, we study how to faithfully capture the the features of single-view image inputs and use it to constrain the novel-view image distributions.

Method

An overview of our method is shown in Figure 2. Given an input image x0{\mathbf{x}}_{0}, we would like to learn a NeRF representation Fω:(x,y,z)→(c,σ)F_{\omega}:(x,y,z)\to({\mathbf{c}},\sigma) as its 3D reconstructionHere we use a Lambertian NeRF without view direction inputs for enforcing stronger multiview consistency.. The NeRF holds the rendering equation that, for any camera view with pose P{\mathbf{P}}, one can sample camera rays r(t)=o+td{\mathbf{r}}(t)={\mathbf{o}}+t{\mathbf{d}} and render the image x{\mathbf{x}} at this view with

where T(t)=exp⁡(−∫tntσ(s)ds)T(t)=\exp\left(-\int_{t_{n}}^{t}\sigma(s){\textnormal{d}}s\right). For more details, please refer to Mildenhall et al. . For simplicity, we denote this whole rendering equation by x=f(P,ω){\mathbf{x}}=f({\mathbf{P}},\omega) which means NeRF ff renders image x{\mathbf{x}} at camera pose P{\mathbf{P}} with parameters ω\omega. Instead of predicting the NeRF parameters ω\omega from x0{\mathbf{x}}_{0} in a forward pass, we formulate this as a conditioned 3D generation problem

where we optimize the NeRF to follow a 3D scene distribution conditioned on that its rendering f(P0,ω)f({\mathbf{P}}_{0},\omega) at a given view P0{\mathbf{P}}_{0} should be the input image x0{\mathbf{x}}_{0}

Directly learning the 3D scene distribution prior requires large 3D datasets, which is less straightforward to acquire and restricts its application to unseen scene categories. To enable better generalizability to in-the-wild scenarios, we instead leverage 2D image priors and reformulate the objective into

Here, s{\mathbf{s}} is an additional semantic guidance term that we apply to further restrict the prior image distribution to fit the generation context. In contrast to DreamFusion which also utilizes language-guided image diffusion model as 2D image priors for sampled views, our main contribution stands in our approach for further constraining the identity of the generated 3D volume to be consistent with the inputs.

We cover more details on this novel-view distribution loss in Sec. 3.1. We utilize natural language descriptions of the scene as the semantic guidance s{\mathbf{s}}. More details on this will be discussed in Sec. 3.2. In addition, as the image diffusion model only operates on the rendered rgb colors, we further apply a geometric regularization with a depth map estimated at the input view to facilitate the NeRF optimization (Sec. 3.3)

Denoising Diffusion Probabilistic Models (DDPM) are a type of generative models that learn a distribution over training data samples. Recently, there are many advances in language guided image synthesis with diffusion models. We build our method upon the recent Latent Diffusion Model (LDM) for its high quality and efficiency in image generation. It adopts a pre-trained image auto-encoder with an encoder E(x)=z{\mathcal{E}}({\mathbf{x}})={\mathbf{z}} mapping images x{\mathbf{x}} into latent codes s{\mathbf{s}} and a decoder D(E(x))=x{\mathcal{D}}({\mathcal{E}}({\mathbf{x}}))={\mathbf{x}} recovering the images. The diffusion process is then trained in the latent space by minimizing the objective

where tt is a diffusion time scale, ϵ∼N(0,1){\epsilon}\sim{\mathcal{N}}(0,1) is a random noise sample, zt{\mathbf{z}}_{t} is the latent code z{\mathbf{z}} noised to time tt with ϵ{\epsilon}, and ϵθ{\epsilon}_{\theta} is the denoising network with parameters θ\theta to regress the noise ϵ{\epsilon}. The diffusion model also takes a conditioning input s{\mathbf{s}} which is encoded as cθ(s)c_{\theta}({\mathbf{s}}) and serves as guidance in the denoising process. For text-to-image generation models such as the LDM, cθc_{\theta} is a pre-trained large language model that encodes the conditional text s{\mathbf{s}}.

In a pre-trained diffusion model, the network parameters θ\theta are fixed, and we can instead optimize for the input image x{\mathbf{x}} with the same objective which transforms x{\mathbf{x}} to follow the image distribution priors conditioned on s{\mathbf{s}}. Let x=f(P,ω){\mathbf{x}}=f({\mathbf{P}},\omega) be our NeRF rendering at arbitrarily sampled view P{\mathbf{P}}, we can back propagate gradients to the NeRF parameters ω\omega and thus get a stochastic gradient descent on ω\omega.

2 Semantics-Conditioned Image Priors

We argue that the prior distribution over all in-the-wild images is not specific enough to guide the novel view synthesis from an arbitrary image. We thus introduce a well-designed guidance s{\mathbf{s}} that narrows down the generic prior over natural images to a prior of images related to the input image x0{\mathbf{x}}_{0}. Here we choose text as the guidance, which is flexible for describing arbitrary input images. Text-to-image diffusion models such as LDM utilize a pre-trained large language model as the language encoder to learn a conditional distribution over images conditioned on language. This serves as a natural gateway for us to utilize language as a means to restrict the image prior space.

The most straightforward way of getting a text prompt from the input image is to use an image captioning or classification network S{\mathcal{S}} trained on (image, text) datasets and predict a text s0=S(x0){\mathbf{s}}_{0}={\mathcal{S}}({\mathbf{x}}_{0}). However, while text description can summarize the semantics of the image, it leaves a huge space of ambiguities, making it hard to include all the visual details in the image especially with limited prompt length. In Figure 3 top row, we show the images generated with the caption “a collection of products” from the input image on the left. While their semantics are highly accurate with respect to the language description, the generated images have very high variances in their visual patterns and low correlations to the input image.

Textual inversion , on the other hand, optimizes for the text embedding of one or few images from a text-based image diffusion model. With the LDM Equation 5, we can optimize for the text embedding s∗{\mathbf{s}}_{*} for the input image x0{\mathbf{x}}_{0} by

In Figure 3 middle row, images generated with textual inversion are shown. The colors and visual cues of the input image are well captured (orange-colored elements, food, and even the brand logos). However, the semantics at the macro level is sometimes wrong (second column is a person playing sports). One reason is that, different from the multi-image scenarios where textual inversion can discover the common contents of these images, it is unclear for one single image what the key features are that the text embedding should focus on.

To reflect both semantic and visual characteristics of the input image in the novel view synthesis task, we combine these two methods by concatenating their text embeddings to form a joint feature s=[s0,s∗]{\mathbf{s}}=[{\mathbf{s}}_{0},{\mathbf{s}}_{*}] and use it as the guidance in the diffusion process in Equation 5. Figure 3 bottom row shows the images generated with this joint feature, with balanced semantics and visual cues.

3 Geometric Regularization

While image diffusion shapes the appearance of the NeRF, multiview consistency is difficult to enforce as the underlying 3D geometry can be different even with the same image rendering , making the gradient back-propagation (from the image diffusion to the NeRF parameters ω\omega) highly non-controllable. To this end, we further incorporate a geometric regularization term on the input view depth to alleviate this issue. We adopt the Dense Prediction Transformer (DPT) model trained on 1.4 million images for zero-shot monocular depth estimation and apply it to the input image x0{\mathbf{x}}_{0} to estimate a depth map d0,est{\mathbf{d}}_{0,\text{est}}. We use this estimated depth to regularize the depth

rendered by the NeRF at input view P0{\mathbf{P}}_{0}. Due to the ambiguities of the estimated depth (including scales, shifts, camera intrinsics) and estimation error (Figure 4), we cannot back project pixels with depth to 3D and compute the regularization directly. Instead, we maximize the Pearson correlation between the estimated depth map and the NeRF-rendered depth

which measures if the rendered depth distribution and the noisy estimated depth distribution are linearly correlated.

Experiments

Now we demonstrate our efficacy in synthesizing realistic NeRFs with single-view inputs. Section 4.1 presents a quantitative comparison between our method and the state-of-the-art single-view NeRF reconstruction methods on a synthetic dataset. Section 4.2 shows a qualitative comparison as well as more synthesis results of our method on in-the-wild images.

We evaluate our method on the DTU MVS dataset with 15 test scenes as specified in . For each input image, we use GPT-2 to generate a caption. We manually correct the obvious mistakes made by GPT-2 while trying our best to avoid introducing additional details. The scenes and their captions are listed in the supplementary material.

Implementation details

For the NeRF model, we implement the multi-resolution grid sampler as described in . For the diffusion model, we employ the text-guided diffusion model from which was pre-trained on the LAION-400M dataset . While operates on 512×512512\times 512 images, NeRF’s volumetric rendering at this resolution would incur an extensive computational burden. Thus, at the randomly sampled novel views, we render 128×128128\times 128 images and resize them to 512×512512\times 512 before feeding them to the encoder of . At the input view, we render at the same resolution as the input image to compute the image reconstruction and depth correlation losses.

Baselines

We compare with two state-of-the-art single-view NeRF reconstruction algorithms, PixelNeRF and its fine-tuned model with CLIP feature consistency loss as proposed by DietNeRF , both of which trained on the training set data from the DTU MVS dataset. To gain better convergence, we use the predictions from as an initialization for our 3D scene optimization. But our method is directly applied to the test scenes without any additional fine-tuning on the DTU training set.

Results

Table 1 shows the quantitative comparison between our method and the baselines. Following the convention, we report the standard image quality metrics PSNR and SSIM . Our PSNR and SSIM are slightly lower than pixelNeRF which directly learns the scene distributions from the DTU training set and are on par with DietPixelNeRF which enforces semantic consistency between views. However, we emphasize that these two metrics are less indicative in our scenario as they are local pixel-aligned similarity metrics between the synthesized novel views and the ground truth images but uncertainties naturally exist in single-view 3D inference. The middle column of the first scene in Figure 5 shows an example of such uncertainty. The height of the tallest snack bag in the input image cannot be inferred as its top extrudes beyond the camera view. The width of the toy pig in the left column of the third scene is another example which cannot be inferred from the input side view. In both cases our method guesses its novel view (bottom row) in a reasonable sense but different from the ground truth (top row). In addition, we also measure novel views with LPIPS , which is a perceptual metric computing the Mean Squared Error (MSE) between normalized features from all layers of a pre-trained VGG encoder . Our method shows a significant improvement on this metric compared to the baselines as the diffusion model helps to improve image qualities while the language guidance maintains the multi-view semantic consistency.

Figure 5 shows a qualitative comparison between our method and the baselines. With the scene initialization from , our method removes the noises and blurriness, synthesizing high quality novel views.

2 Images in the Wild

Figure 6 shows a qualitative comparison between our method and existing state-of-the-art single-image to 3D synthesis methods for in-the-wild images . Input images are adopted from the Google Scanned Objects dataset with their category labels (‘bag’ and ‘hat’) as captions. Similar to ours, DietNeRF uses an input-view constrained NeRF optimization technique where they minimize the CLIP feature between arbitrary view renderings. While CLIP features enforce consistent appearances, they fail to capture the global semantics of the object. SS3D is a forward-prediction model for 3D geometries that transfers the priors learned on synthetic datasets to in-the-wild images with knowledge distillation. While it generates more structured global geometries, it fails to capture the fine geometric details of the input image. The geometries of the hats in the bottom rows are also incorrect, with only the silhouette shape preserved but the structure of ‘hat’ shape missing.

More results

Figure LABEL:fig:results_internet shows our results on images of objects from the internet. The text prompts are words or phrases used to search for the images. The backgrounds are masked out using an off-the-shelf dichotomous image segmentation network from . For each input, we show 3 different novel views that are distant from the input view. Figure 7(b) shows our results on images with more complex contents and backgrounds from the COCO dataset which contains (image, caption) pairs. Within camera views close to the input, our model is still able to generate realistic renderings. But it can hardly generalize to distant views due to the limited capacity of the NeRF scene box.

3 Ablation Studies

We conduct ablation studies to show the efficacy of our two-section semantic guidance and geometric regularization.

Figure 8(a) shows the ablation of the two text embeddings s0{\mathbf{s}}_{0} from image captions and s∗{\mathbf{s}}_{*} from textual inversion. Without the captions s0{\mathbf{s}}_{0}, the model fails to learn the overall semantics and cannot generate a meaningful object. While both the full model and the caption-only one (without textual inversion) successfully generate backack novel views, the results without textual inversion s∗{\mathbf{s}}_{*} have more blurriness and noises. A zoom-in comparison is shown in Figure 8(b).

Figure 8(c) shows another comparison of models with and without textual inversion s∗{\mathbf{s}}_{*} on the can example from Figure 7(b) left. In the object regions visible to the input view, the full model better recovers the fine details (the white letters on the lateral); and in the invisible regions, the full model completes the appearances with coherent styles of the input (red and white textures at the back of the can), while the model without textual inversion does not have such appearance coherency. The model with textual inversion can even synthesize the pull tab at the top (second column of the zoom-in views) by inferring from the input side view that this is a can containing drinks.

Geometric regularization

Figure 9 shows an ablation on the geometric regularization term. Both image renderings and depth maps are visualized. The full model is able to synthesize realistic novel views with coherent 3D geometry. The model without the regularization on the input view depth can still generate realistic appearances at novel views with the diffusion model, but the underlying 3D geometry is erroneous and multi-view consistency is not enforced. As a sanity check, we also visualize the results with only the depth loss but without the diffusion model. The model is unable to generate a realistic NeRF due to the 3D ambiguities of monocular depth as stated in Section 3.3.

Conclusions

In this paper, we propose a novel framework for zero-shot single-view NeRF synthesis for images in the wild without 3D supervision. We leverage the general image priors in 2D diffusion models and apply them to the 3D NeRF generation conditioned on the input image. To efficiently use these priors in synthesizing consistent views, we design a two-section language guidance as conditioning inputs to the diffusion model which unifies the semantic and visual features of the input image. To our knowledge, we are the first to combine semantic and visual features in the text embedding space and apply it to novel view synthesis. In addition, we introduce a geometric regularization term while addressing the 3D ambiguity of monocular-estimated depth maps. Our experimental results show that, with well-designed guidance and constraints, one can leverage general image priors to specific image-to-3D, enabling us to build generalizable and adaptable reconstruction frameworks.

As our method relies on multiple large pre-trained image models , any biases in these models will affect our synthesis results. Figure 10(a) shows an example where the image diffusion model can generate two shoes even the text prompt is “a single shoe”, resulting in our synthesized NeRF showing the features of multiple shoes. Our method is also less robust to highly deformable instances, as our language guidance focuses on semantics and styles but lacks a global description of physical states and dynamics. Figure 10(b) shows such a failure case. Renderings from each independent view are visually plausible but represent different states of the same instances.

Besides, while formulation-wise the optimization is applicable to any scenes, it is more suitable for object-centric images as it takes the underlying assumption that the scene has exactly the same semantics from any view, which is not true for large scenes with complex configurations due to view changes and occlusions. The text embedding learned from textual inversion is of the dimension of a single-world embedding, limiting its expressiveness in representing the subtleties complex contents.

Appendix A Additional Results

Figure 12 shows our additional results and comparisons for images in the wild. The results are presented in 4 groups, each group containing 3 objects from similar classes but with different content details and appearances. We use this to test the capability of each method in capturing the overall semantics and visual feature variations from input images.

For a fair comparison, DietNeRF is also optimized with the estimated depth map from the input image. While DietNeRF is able to maintain appearance consistency between different views, it fails to capture the overall geometry of the objects, especially when the object has complex geometric structures (such as the chairs in the 1st group, and the baskets in the 3rd group). In the 4th group (the skirts), our generated textures form the unseen back regions are also closer to the input image than DietNeRF.

Our method also addresses the naturally existing ambiguity in novel-view inference, especially for the occluded regions in the input view. For example, in the 3rd group in Figure 12, the unseen spaces of the baskets are filled with different fruits/flowers/vegetables, instead of duplicating the input views as DietNeRF . As a feature or as an inductive bias, such synthesis results are also affected by the 2D distribution from the image diffusion model. For example, Figure 11 shows the image generation results by with text prompt ‘a pumpkin’. Half of them are Jack-o’-lanterns. This makes our synthesized pumpkin also having the Jack-o’-lantern face at its back (the 3rd row of the 2nd group).

Comparison to SS3D [45]

As a geometry-based method, SS3D captures better global geometries than DietNeRF even without the depth regularization, especially on the object classes covered by ShapeNet where the

References