Putting NeRF on a Diet: Semantically Consistent Few-Shot View Synthesis

Ajay Jain, Matthew Tancik, Pieter Abbeel

Introduction

In the novel view synthesis problem, we seek to rerender a scene from arbitrary viewpoint given a set of sparsely sampled viewpoints. View synthesis is a challenging problem that requires some degree of 3D reconstruction in addition to high-frequency texture synthesis. Recently, great progress has been made on high-quality view synthesis when many observations are available. A popular approach is to use Neural Radiance Fields (NeRF) to estimate a continuous neural scene representation from image observations. During training on a particular scene, the representation is rendered from observed viewpoints using volumetric ray casting to compute a reconstruction loss. At test time, NeRF can be rendered from novel viewpoints by the same procedure. While conceptually very simple, NeRF can learn high-frequency view-dependent scene appearances and accurate geometries that allow for high-quality rendering.

Still, NeRF is estimated per-scene, and cannot benefit from prior knowledge acquired from other images and objects. Because of the lack of prior knowledge, NeRF requires a large number of input views to reconstruct a given scene at high-quality. Given 8 views, Figure 2B shows that novel views rendered with the full NeRF model contain many artifacts because the optimization finds a degenerate solution that is only accurate at observed poses. We find that the core issue is that prior 3D reconstruction systems based on rendering losses are only supervised at known poses, so they overfit when few poses are observed. Regularizing NeRF by simplifying the architecture avoids the worst artifacts, but comes at the cost of fine-grained detail.

Further, prior knowledge is needed when the scene reconstruction problem is underdetermined. 3D reconstruction systems struggle when regions of an object are never observed. This is particularly problematic when rendering an object at significantly different poses. When rendering a scene with an extreme baseline change, unobserved regions during training become visible. A view synthesis system should generate plausible missing details to fill in the gaps. Even a regularized NeRF learns poor extrapolations to unseen regions due to its lack of prior knowledge (Figure 2D).

Recent work trained NeRF on multi-view datasets of similar scenes to bias reconstructions of novel scenes. Unfortunately, these models often produce blurry images due to uncertainty, or are restricted to a single object category such as ShapeNet classes as it is challenging to capture large, diverse, multi-view data.

In this work, we exploit the consistency principle that “a bulldozer is a bulldozer from any perspective”: objects share high-level semantic properties between their views. Image recognition models learn to extract many such high-level semantic features including object identity. We transfer prior knowledge from pre-trained image encoders learned on highly diverse 2D single-view image data to the view synthesis problem. In the single-view setting, such encoders are frequently trained on millions of realistic images like ImageNet . CLIP is a recent multi-modal encoder that is trained to match images with captions in a massive web scrape containing 400M images . Due to the diversity of its data, CLIP showed promising zero- and few-shot transfer performance to image recognition tasks. We find that CLIP and ImageNet models also contain prior knowledge useful for novel view synthesis.

We propose DietNeRF, a neural scene representation based on NeRF that can be estimated from only a few photos, and can generate views with unobserved regions. In addition to minimizing NeRF’s mean squared error losses at known poses in pixel-space, DietNeRF penalizes a semantic consistency loss. This loss matches the final activations of CLIP’s Vision Transformer between ground-truth images and rendered images at different poses, allowing us to supervise the radiance field from arbitrary poses. In experiments, we show that DietNeRF learns realistic reconstructions of objects with as few as 8 views without simplifying the underlying volumetric representation, and can even produce reasonable reconstructions of completely occluded regions. To generate novel views with as few as 1 observation, we fine-tune pixelNeRF , a generalizable scene representation, and improve perceptual quality.

Background on Neural Radiance Fields

A plenoptic function, or light field, is a five-dimensional function that describes the light radiating from every point in every direction in a volume such as a bounded scene. While explicitly storing or estimating the plenoptic function at high resolution is impractical due to the dimensionality of the input, Neural Radiance Fields parameterize the function with a continuous neural network such as a multi-layer perceptron (MLP). A Neural Radiance Field (NeRF) model is a five-dimensional function fθ(x,d)=(c,σ)f_{\theta}(\mathbf{x},\mathbf{d})=(\mathbf{c},\sigma) of spatial position x=(x,y,z)\mathbf{x}=(x,y,z) and viewing direction (θ,ϕ)(\theta,\phi), expressed as a 3D unit vector d\mathbf{d}. NeRF predicts the RGB color c\mathbf{c} and differential volume density σ\sigma from these inputs. To encourage view-consistency, the volume density only depends on x\mathbf{x}, while the color also depends on viewing direction d\mathbf{d} to capture viewpoint dependent effects like specular reflections. Images are rendered from a virtual camera at any position by integrating color along rays cast from the observer according to volume rendering :

where the ray originating at the camera origin o\mathbf{o} follows path r(t)=o+td\mathbf{r}(t)=\mathbf{o}+t\mathbf{d}, and the transmittance T(t)=exp⁡(−∫tntfσ(r(s))ds){T(t)=\exp\left(-\int_{t_{n}}^{t_{f}}\sigma(\mathbf{r}(s))ds\right)} weights the radiance by the probability that the ray travels from the image plane at tnt_{n} to tt unobstructed. To approximate the integral, NeRF employs a hierarchical sampling algorithm to select function evaluation points near object surfaces along each ray. NeRF separately estimates two MLPs, a coarse network and a fine network, and uses the coarse network to guide sampling along the ray for more accurately estimating (1). The networks are trained from scratch on each scene given tens to hundreds of photos from various perspectives. Given observed multi-view training images {Ii}\{I_{i}\} of a scene, NeRF uses COLMAP SfM to estimate camera extrinsics (rotations and origins) {pi}\{\mathbf{p}_{i}\}, creating a posed dataset D={(Ii,pi)}\mathcal{D}=\{(I_{i},\mathbf{p}_{i})\}.

NeRF Struggles at Few-Shot View Synthesis

View synthesis is a challenging problem when a scene is only sparsely observed. Systems like NeRF that train on individual scenes especially struggle without prior knowledge acquired from similar scenes. We find that NeRF fails at few-shot novel view synthesis in several settings.

NeRF overfits to training views Conceptually, NeRF is trained by mimicking the image-formation process at observed poses. The radiance field can be estimated repeatedly sampling a training image and pose (I,pi)(I,\mathbf{p}_{i}), rendering an image I^pi\hat{I}_{\mathbf{p}_{i}} from the same pose by volume integration (1), then minimizing the mean-squared error (MSE) between the images, which should align pixel-wise:

In practice, NeRF samples a smaller batch of rays across all training images to avoid the computational expense of rendering full images during training. Given subsampled rays R\mathcal{R} cast from the training cameras, NeRF minimizes:

With many training views, LMSE\mathcal{L}_{\text{MSE}} provides training signal to fθf_{\theta} densely in the volume and does not overfit to individual training views. Instead, the MLP recovers accurate textures and occupancy that allow interpolations to new views (Figure 2A). Radiance fields with sinusoidal positional embeddings are quite effective at learning high-frequency functions , which helps the MLP represent fine details.

Unfortunately, this high-frequency representational capacity allows NeRF to overfit to each input view when only a few are available. LMSE\mathcal{L}_{\text{MSE}} can be minimized by packing the reconstruction I^p\hat{I}_{\mathbf{p}} of training view (I,p)(I,\mathbf{p}) close to the camera. Fundamentally, the plenoptic function representation suffers from a near-field ambiguity where distant cameras each observe significant regions of space that no other camera observes. In this case, the optimal scene representation is underdetermined. Degenerate solutions can also exploit the view-dependence of the radiance field. Figure 2B shows novel views from the same NeRF trained on 8 views. While a rendered view from a pose near a training image has reasonable textures, it is skewed incorrectly and has cloudy artifacts from incorrect geometry. As the geometry is not estimated correctly, a distant view contains almost none of the correct information. High-opacity regions block the camera. Without supervision from any nearby camera, opacity is sensitive to random initialization.

Regularization fixes geometry, but hurts fine-detail High-frequency artifacts such as spurious opacity and rapidly varying colors can be avoided in some cases by regularizing NeRF. We simplify the NeRF architecture by removing hierarchical sampling and learning only a single MLP, and reducing the maximum frequency positional embedding in the input layer. This biases NeRF toward lower frequency solutions, such as placing content in the center of the scene farther from the training cameras. We also can address some few-shot optimization challenges by lowering the learning rate to improve initial convergence, and manually restarting training if renderings are degenerate. Figure 2C shows that these regularizers successfully allow NeRF to recover plausible object geometry. However, high-frequency, fine details are lost compared to 2A.

No prior knowledge, no generalization to unseen views As NeRF is estimated from scratch per-scene, it has no prior knowledge about natural objects such as common symmetries and object parts. In Figure 2D, we show that NeRF trained with 14 views of the right half of a Lego vehicle generalizes poorly to its left side. We regularized NeRF to remove high-opacity regions that originally blocked the left side entirely. Even so, the essential challenge is that NeRF receives no supervisory signal from LMSE\mathcal{L}_{\text{MSE}} to the unobserved regions, and instead relies on the inductive bias of the MLP for any inpainting. We would like to introduce prior knowledge that allows NeRF to exploit bilateral symmetry for plausible completions.

Semantically Consistent Radiance Fields

Motivated by these challenges, we introduce the DietNeRF scene representation. DietNeRF uses prior knowledge from a pre-trained image encoder to guide the NeRF optimization process in the few-shot setting.

DietNeRF supervises fθf_{\theta} at arbitrary camera poses during training with a semantic loss. While pixel-wise comparison between ground-truth observed images and rendered images with LMSE\mathcal{L}_{\text{MSE}} is only useful when the rendered image is aligned with the observed pose, humans are easily able to detect whether two images are views of the same object from semantic cues. We can in general compare a representation of images captured from different viewpoints:

If ϕ(x)=x\phi(x)=x, Eq. (4) reduces to Lfull\mathcal{L}_{\text{full}} up to a scaling factor. However, the identity mapping is view-dependent. We need a representation that is similar across views of the same object and captures important high-level semantic properties like object class. We evaluate the utility of two sources of supervision for representation learning. First, we experiment with the recent CLIP model pre-trained for multi-modal language and vision reasoning with contrastive learning . We then evaluate visual classifiers pre-trained on labeled ImageNet images . In both cases, we use similar Vision Transformer (ViT) architectures.

A Vision Transformer is appealing because its performance scales very well to large amounts of 2D data. Training on a large variety of images allows the network to encounter multiple views of an object class over the course of training without explicit multi-view data capture. It also allows us to transfer the visual encoder to diverse objects of interest in graphics applications, unlike prior class-specific reconstruction work that relies on homogeneous datasets . ViT extracts features from non-overlapping image patches in its first layer, then aggregates increasingly abstract representations with Transformer blocks based on global self-attention to produce a single, global embedding vector. ViT outperformed CNN encoders in our early experiments.

In practice, CLIP produces normalized image embeddings. When ϕ(⋅)\phi(\cdot) is a unit vector, Eq. (4) simplifies to cosine similarity up to a constant and a scaling factor that can be absorbed into the loss weight λ\lambda:

We refer to LSC\mathcal{L}_{\text{SC}} (5) as a semantic consistency loss because it measures the similarity of high-level semantic features between observed and rendered views. In principle, semantic consistency is a very general loss that can be applied to any 3D reconstruction system based on differentiable rendering.

2 Interpreting representations across views

The pre-trained CLIP model that we use is trained on hundreds of millions of images with captions of varying detail. Image captions provide rich supervision for image representations. On one hand, short captions express semantically sparse learning signal as a flexible way to express labels . For example, the caption “A photo of hotdogs” describes Fig. 2A. Language also provides semantically dense learning signal by describing object properties, relationships and appearances such as the caption “Two hotdogs on a plate with ketchup and mustard”. To be predictive of such captions, an image representation must capture some high-level semantics that are stable across viewpoints. Concurrently, found that CLIP representations capture visual attributes of images like art style and colors, as well as high-level semantic attributes including object tags and categories, facial expressions, typography, geography and brands.

In Figure 3, we measure the pairwise cosine similarity between CLIP representations of views circling an object. We find that pairs of views have highly similar CLIP representations, even for diametrically opposing cameras. This suggests that large, diverse single-view datasets can induce useful representations for multi-view applications.

3 Pose sampling distribution

We augment the NeRF training loop with LSC\mathcal{L}_{\text{SC}} minimization. Each iteration, we compute LSC\mathcal{L}_{\text{SC}} between a random training image sampled from the observation dataset I∼DI\sim\mathcal{D} and rendered image I^p\hat{I}_{\mathbf{p}} from random pose p∼π\mathbf{p}\sim\pi. For bounded scenes like NeRF’s Realistic Synthetic scenes where we are interested in 360∘ view synthesis, we define the pose sampling distribution π\pi to be a uniform distribution over the upper hemisphere, with radius sampled uniformly in a bounded range. For unbounded forward-facing scenes or scenes where a pose sampling distribution is difficult to define, we interpolate between three randomly sampled known poses p1,p2,p3∼D\mathbf{p}_{1},\mathbf{p}_{2},\mathbf{p}_{3}\sim\mathcal{D} with pairwise interpolation weights α1,α2∼U(0,1)\alpha_{1},\alpha_{2}\sim\mathcal{U}(0,1).

4 Improving efficiency and quality

Volume rendering is computationally intensive. Computing a pixel’s color evaluates NeRF’s MLP fθf_{\theta} at many points along a ray. To improve the efficiency of DietNeRF during training, we render images for semantic consistency at low resolution, requiring only 15-20% of the rays as a full resolution training image. Rays are sampled on a strided grid across the full extent of the image plane, ensuring that objects are mostly visible in each rendering. We found that sampling poses from a continuous distribution was helpful to avoid aliasing artifacts when training at a low resolution.

In experiments, we found that LSC\mathcal{L}_{\text{SC}} converges faster than LMSE\mathcal{L}_{\text{MSE}} for many scenes. We hypothesize that the semantic consistency loss encourages DietNeRF to recover plausible scene geometry early in training, but is less helpful for reconstructing fine-grained details due to the relatively low dimensionality of the ViT representation ϕ(⋅)\phi(\cdot). We exploit the rapid convergence of LSC\mathcal{L}_{\text{SC}} by only minimizing LSC\mathcal{L}_{\text{SC}} every kk iterations. DietNeRF is robust to the choice of kk, but a value between 10 and 16 worked well in our experiments. StyleGAN2 used a similar strategy for efficiency, referring to periodic application of a loss as lazy regularization.

As backpropagation through rendering is memory intensive with reverse-mode automatic differentiation, we render images for LSC\mathcal{L}_{\text{SC}} with mixed precision computation and evaluate ϕ(⋅)\phi(\cdot) at half-precision. We delete intermediate MLP activations during rendering and rematerialize them during the backward pass . All experiments use a single 16 GB NVIDIA V100 or 11 GB 2080 Ti GPU.

Since LSC\mathcal{L}_{\text{SC}} converges before LMSE\mathcal{L}_{\text{MSE}}, we found it helpful to fine-tune DietNeRF with LMSE\mathcal{L}_{\text{MSE}} alone for 20-70k iterations to refine details. Alg. 1 details our overall training process.

Experiments

In experiments, we evaluate the quality of novel views synthesized by DietNeRF and baselines for both synthetically rendered objects and real photos of multi-object scenes. (1) We evaluate training from scratch on a specific scene with 8 views §5.1. (2) We show that DietNeRF improves perceptual quality of view synthesis from only a single real photo §5.2. (3) We find that DietNeRF can reconstruct regions that are never observed §5.3, and finally (4) run ablations §6.

Datasets The Realistic Synthetic benchmark of includes detailed multi-view renderings of 8 realistic objects with view-dependent light transport effects. We also benchmark on the DTU multi-view stereo (MVS) dataset used by pixelNeRF . DTU is a challenging dataset that includes sparsely sampled real photos of physical objects.

Low-level full reference metrics Past work evaluates novel view quality with respect to ground-truth from the same pose with Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM) . PSNR expresses mean-squared error in log space. However, SSIM often disagrees with human judgements of similarity .

Perceptual metrics Deep CNN activations mirror aspects of human perception. NeRF measures perceptual image quality using LPIPS , which computes MSE between normalized features from all layers of a pre-trained VGG encoder . Generative models also measure sample quality with feature space distances. The Fréchet Inception Distance (FID) computes the Fréchet distance between Gaussian estimates of penultimate Inception v3 features for real and fake images. However, FID is a biased metric at low sample sizes. We adopt the conceptually similar Kernel Inception Distance (KID), which measures the MMD between Inception features and has an unbiased estimator . All metrics use a different architecture and data than our CLIP ViT encoder.

NeRF’s Realistic Synthetic dataset includes 8 detailed synthetic objects with 100 renderings from virtual cameras arranged randomly on a hemisphere pointed inward. To test few-shot performance, we randomly sample a training subset of 8 images from each scene. Table 1 shows results. The original NeRF model achieves much poorer quantitative quality with 8 images than with the full 100 image dataset. Neural Volumes performs better as it tightly constrains the size of the scene’s bounding box and explicitly regularizes its scene representation using a penalty on spatial gradients of voxel opacity and a Beta prior on image opacity. This avoids the worst artifacts, but reconstructions are still low-quality. Simplifying NeRF and tuning it for each individual scene also regularizes the representation and helps convergence (+5.1 PSNR over the full NeRF). The best performance is achieved by regularizing with DietNeRF’s LSC\mathcal{L}_{\text{SC}} loss. Additionally, fine-tuning with LMSE\mathcal{L}_{\text{MSE}} even further improves quality, for a total improvement of +8.5 PSNR, -0.2 LPIPS, and -156 FID over NeRF. This shows that semantic consistency is a valuable prior for high-quality few-shot view synthesis. Figure 4 visualizes results.

2 Single-view synthesis by fine-tuning

NeRF only uses observations during training, not inference, and uses no auxiliary data. Accurate 3D reconstruction from a single view is not possible purely from LMSE\mathcal{L}_{\text{MSE}}, so NeRF performs poorly in the single-view setting (Table 2).

To perform single- or few-shot view synthesis, pixelNeRF learns a ResNet-34 encoder and a feature-conditioned neural radiance field on a multi-view dataset of similar scenes. The encoder learns priors that generalize to new single-view scenes. Table 2 shows that pixelNeRF significantly outperforms NeRF given a single photo of a held-out scene. However, novel views are blurry and unrealistic (Figure 5). We propose to fine-tune pixelNeRF on a single scene using LMSE\mathcal{L}_{\text{MSE}} alone or using both LMSE\mathcal{L}_{\text{MSE}} and LSC\mathcal{L}_{\text{SC}}. Fine-tuning per-scene with MSE improves local image quality metrics, but only slightly helps perceptual metrics. Figure 6 shows that pixel-space MSE fine-tuning from one view mostly only improves quality for that view.

We refer to fine-tuning with both losses for a short period as DietPixelNeRF. Qualitatively, DietPixelNeRF has significantly sharper novel views (Fig. 5, 6). DietPixelNeRF outperforms baselines on perceptual LPIPS, FID, and KID metrics (Tab. 2). For the very challenging single-view setting, ground-truth novel views will contain content that is completely occluded in the input. Because of uncertainty, blurry renderings will outperform sharp but incorrect renderings on average error metrics like MSE and PSNR. Arguably, perceptual quality and sharpness are better metrics than pixel error for graphics applications like photo editing and virtual reality as plausibility is emphasized.

3 Reconstructing unobserved regions

We evaluate whether DietNeRF produces plausible completions when the reconstruction problem is underdetermined. For training, we sample 14 nearby views of the right side of the Realistic Synthetic Lego scene (Fig. 7, right). Narrow baseline multi-view capture rigs are less costly than 360∘ captures, and support unbounded scenes. However, narrow-baseline observations suffer from occlusions: the left side of the Lego bulldozer is unobserved. NeRF fails to reconstruct this side of the scene, while our Simplified NeRF learns unrealistic deformations and incorrect colors (Fig. 7, left). Remarkably, DietNeRF learns quantitatively (Tab. 3) and qualitatively more accurate colors in the missing regions, suggesting the value of semantic image priors for sparse reconstruction problems. We exclude FID and KID since a single scene has too few samples for an accurate estimate.

Ablations

Choosing an image encoder Table 4 shows quality metrics with different semantic encoder architectures and pre-training datasets. We evaluate on the Lego scene with 8 views. Large ViT models (ViT L) do not improve results over the base ViT B. Fixing the architecture, CLIP offers a +1.8 PSNR improvement over an ImageNet model, suggesting that data diversity and language supervision is helpful for 3D tasks. Still, both induce useful representations that transfer to view synthesis.

Varying LMSE\mathcal{L}_{\text{MSE}} fine-tuning duration Fine-tuning DietNeRF with LMSE\mathcal{L}_{\text{MSE}} can improve quality by better reconstructing fine-details. In Table 5, we vary the number of iterations of fine-tuning for the Realistic Synthetic scenes with 8 views. Fine-tuning for up to 50k iterations is helpful, but reduces performance with longer optimization. It is possible that the model starts overfitting to the 8 input views.

Related work

Few-shot radiance fields Several works condition NeRF on latent codes describing scene geometry or appearance rather than estimating NeRF per scene . An image encoder and radiance field decoder are learned on a multi-view dataset of similar objects or scenes ahead of time. At test time, on a new scene, novel viewpoints are rendered using the decoder conditioned on encodings of a few observed images. GRAF renders patches of the scene every iteration to supervise the network with a discriminator . Concurrent to our work, IBRNet also fine-tunes a latent-conditioned radiance field on a specific scene using NeRF’s reconstruction loss, but needed at least 50 views. Rather than generalizing between scenes through a shared encoder and decoder, meta-learn radiance field weights that can be adapted to a specific scene in a few gradient steps. Meta-learning improves performance in the few-view setting. Similarly, a signed distance field can be meta-learned for shape representation problems . Much literature studies single-view reconstruction with other, explicit 3D representations. Notable recent examples include voxel , mesh and point-cloud approaches.

Novel view synthesis, image-based rendering Neural Volumes proposes a VAE encoder-decoder architecture to predict a volumetric representation of a scene from posed image observations. NV uses priors as auxiliary objectives like DietNeRF, but penalizes opacity based on geometric intuitions rather than RGB image semantics. TBNs learn an autoencoder with a 3-dimensional latent that can be rotated to render new perspectives for a single-category. SRNs fit a continuous representation to a scene and also generalize to novel single-category objects if trained on a large multi-view dataset. It can be extended to predict per-point semantic segmentation maps . Local Light Field Fusion estimates and blends multiple MPI representations for each scene. Free View Synthesis uses geometric approaches to improve view synthesis in unbounded in-the-wild scenes. NeRF++ also improves unbounded scenes using multiple NeRF models and changing NeRF’s parameterization.

Semantic representation learning Representation learning with deep supervised and unsupervised approaches has a long history . Without labels, generative models can learn useful representations for recognition , but self-supervised models like CPC tend to be more parameter efficient. Contrastive methods including CLIP learn visual representations by matching similar pairs of items, such as captions and images , augmentated variants of an image , or video patches across frames .

Conclusions

Our results suggest that single-view 2D representations transfer effectively to challenging, underconstrained 3D reconstruction problems such as volumetric novel view synthesis. While pre-trained image encoder representations have certainly been transferred to 3D vision applications in the past by fine-tuning, the recent emergence of visual models trained on enormous 100M+ image datasets like CLIP have enabled surprisingly effective few-shot transfer. We exploited this transferrable prior knowledge to solve optimization issues as well as to cope with partial observability in the NeRF family of scene representations, offering notable improvements in perceptual quality. In the future, we believe “diet-friendly” few-shot transfer will play a greater role in a wide range of 3D applications.

Acknowledgements

This material is based upon work supported by the National Science Foundation Graduate Research Fellowship under grant number DGE-1752814 and by Berkeley Deep Drive. We would like to thank Paras Jain, Aditi Jain, Alexei Efros, Angjoo Kanazawa, Aravind Srinivas, Deepak Pathak and Alex Yu for helpful feedback and discussions.

References

Appendix A Experimental details

For most few-view Realistic Synthetic experiments, we randomly subsample 8 of the available 100 training renders. Views are not manually selected. However, to compare the ability of NeRF and DietNeRF to extrapolate to unseen regions, we manually selected 14 of the 100 views mostly showing the right side of the Lego scene. For DTU experiments where we fine-tune pixelNeRF , we use the same source view as . This viewpoint was manually selected and is shared across all 15 scenes.

Simplified NeRF baseline

The published version of NeRF can be unstable to train with 8 views, often converging to a degenerate solution. We found that NeRF is sensitive to MLP parameter initialization, as well as hyperparameters that control the complexity of the learned scene representation. For a fair comparison, we tuned the Simplified NeRF baseline on each Realistic Synthetic scene by modifying hyperparameters until object geometry converged. Table 6 shows the resulting hyperparameter settings for initial learning rate prior to decay, whether the MLP fθf_{\theta} is viewpoint dependent, number of samples per ray queried from the fine and coarse networks, and the maximum frequency sinusoidal encoding of spatial position (x,y,z)(x,y,z). The fine and coarse networks are used in for hierarchical sampling. ✗ denotes that we do not use the fine network.

Implementation

Our implementation is based on a PyTorch port of NeRF’s original Tensorflow code. We re-train and evaluate NeRF using this code. For memory efficiency, we use 400×\times400 images of the scenes as in rather than full-resolution 800×\times800 images. NV is trained with full-resolution 800×800800\times 800 views. NV renderings are downsampled with a 2x2 box filter to 400×400400\times 400 to compute metrics. We train all NeRF, Simplified NeRF and DietNeRF models with the Adam optimizer for 200k iterations.

Metrics

Our PSNR, SSIM, and LPIPS metrics use the same implementation as based on the scikit-image Python package . For the DTU dataset, excluded some poses from the validation set as ground truth photographs had excessive shadows due to the physical capture setup. We use the same subset of validation views.

For both Realistic Synthetic and DTU scenes, we also included FID and KID perceptual image quality metrics. While PSNR, SSIM and LPIPS are measured between pairs of pixel-aligned images, FID and KID are measured between two sets of image samples. These metrics compare the distribution of image features computed on one set of images to those computed on another set. As distributions are compared rather than individual images, a sufficiently large sample size is needed. For the Realistic Synthetic dataset, we compute the FID and KID between all 3200 ground-truth images (across train, validation and testing splits and across scenes), and 200 rendered test images at the same resolution (25 test views per scene). Aggregating across scenes allows us to have a larger sample size. Due to the setup of the Neural Volumes code, we use additional samples for rendered images for that baseline. For the DTU dataset, we compute FID and KID between 720 rendered images (48 per scene across 15 validation scenes, excluding the viewpoint of the source image provided to pixelNeRF) and 6076 ground-truth images (49 images including the source viewpoint across 124 training and validation scenes). FID and KID metrics are computed using the torch-fidelity Python package .

Appendix B Per-scene metrics

In Figure 8, we compare the cosine similarity of two views with the distance between their camera origins for each pair of scenes in the Realistic Synthetic dataset. When sampling both views from the same scene, views have high cosine similarity (diagonal). For 6 of the 8 scenes, there is some dependence on the relative poses of the camera views, though similarity is high across all camera distances. For views sampled from different scenes, similarity is low (cosine similarity around 0.5).

Quality metrics

Table 7 shows PSNR, SSIM and LPIPS metrics on a per-scene basis for the Realistic Synthetic dataset. FID and KID metrics are excluded as they need a larger sample size. We bold the best method on each scene, and underline the second-best method. Across all scenes in the few-shot setting, DietNeRF or DietNeRF fine-tuned for 50k iterations with LMSE\mathcal{L}_{\text{MSE}} performs best or second-best.

Appendix C Qualitative results and ground-truth

In this section, we provide additional qualitative results. Figure 9 shows the ground-truth training views used for 8-shot Realistic Synthetic experiments. These views are sampled at random from the training set of . Random sampling models challenges with real-world data capture such as uneven view sampling. It may be possible to improve results if views are carefully selected.

In Figure 10, we provide additional renderings of Realistic Synthetic scenes from testing poses for baseline methods and DietNeRF. Neural Volumes generally converges to recover coarse object geometry, but has wispy artifacts and distortions. On the Ship scene, Neural Volumes only recovers very low-frequency detail. Simplified NeRF suffers from occluders that are not visible from the 8 training poses. DietNeRF has the highest quality reconstructions without these distortions or occluders, but does miss some high-frequency detail. An interesting artifact is the leakage of green coloration to the back of the chair.

Finally, in Figure 11, we show renderings from pixelNeRF and DietPixelNeRF on all DTU dataset validation scenes not included in the main paper. Starting from the same checkpoint, pixelNeRF is fine-tuned using LMSE\mathcal{L}_{\text{MSE}} for 20k iterations, whereas DietPixelNeRF is fine-tuned using LMSE+LSC\mathcal{L}_{\text{MSE}}+\mathcal{L}_{\text{SC}} for 20k iterations. DietPixelNeRF has sharper renderings. On scenes with rectangular objects like bricks and boxes, DietPixelNeRF performs especially well. However, the method struggles to preserve accurate geometry in some cases. Note that the problem is under-determined as only a single view is observed per scene.

Appendix D Adversarial approaches

While NeRF is only supervised from observed poses, conceptually, a GAN uses a discriminator to compute a realism loss between real and generated images that need not align pixel-wise. Patch GAN discriminators were introduced for image translation problems and can be useful for high-resolution image generation . SinGAN trains multiscale patch discriminators on a single image, comparable to our single-scene few-view setting. In early experiments, we trained patch-wise discriminators per-scene to supervise fθf_{\theta} from novel poses in addition to LSC\mathcal{L}_{\text{SC}}. However, an auxiliary adversarial loss led to artifacts on Realistic Synthetic scenes, both in isolation and in combination with our semantic consistency loss.