pixelNeRF: Neural Radiance Fields from One or Few Images

Alex Yu, Vickie Ye, Matthew Tancik, Angjoo Kanazawa

Introduction

We study the problem of synthesizing novel views of a scene from a sparse set of input views. This long-standing problem has recently seen progress due to advances in differentiable neural rendering . Across these approaches, a 3D scene is represented with a neural network, which can then be rendered into 2D views. Notably, the recent method neural radiance fields (NeRF) has shown impressive performance on novel view synthesis of a specific scene by implicitly encoding volumetric density and color through a neural network. While NeRF can render photorealistic novel views, it is often impractical as it requires a large number of posed images and a lengthy per-scene optimization.

In this paper, we address these shortcomings by proposing pixelNeRF, a learning framework that enables predicting NeRFs from one or several images in a feed-forward manner. Unlike the original NeRF network, which does not make use of any image features, pixelNeRF takes spatial image features aligned to each pixel as an input. This image conditioning allows the framework to be trained on a set of multi-view images, where it can learn scene priors to perform view synthesis from one or few input views. In contrast, NeRF is unable to generalize and performs poorly when few input images are available, as shown in Fig. 1.

Specifically, we condition NeRF on input images by first computing a fully convolutional image feature grid from the input image. Then for each query spatial point x\mathbf{x} and viewing direction d\mathbf{d} of interest in the view coordinate frame, we sample the corresponding image feature via projection and bilinear interpolation. The query specification is sent along with the image features to the NeRF network that outputs density and color, where the spatial image features are fed to each layer as a residual. When more than one image is available, the inputs are first encoded into a latent representation in each camera’s coordinate frame, which are then pooled in an intermediate layer prior to predicting the color and density. The model is supervised with a reconstruction loss between a ground truth image and a view rendered using conventional volume rendering techniques. This framework is illustrated in Fig. 2.

PixelNeRF has many desirable properties for few-view novel-view synthesis. First, pixelNeRF can be trained on a dataset of multi-view images without additional supervision such as ground truth 3D shape or object masks. Second, pixelNeRF predicts a NeRF representation in the camera coordinate system of the input image instead of a canonical coordinate frame. This is not only integral for generalization to unseen scenes and object categories , but also for flexibility, since no clear canonical coordinate system exists on scenes with multiple objects or real scenes. Third, it is fully convolutional, allowing it to preserve the spatial alignment between the image and the output 3D representation. Lastly, pixelNeRF can incorporate a variable number of posed input views at test time without requiring any test-time optimization.

We conduct an extensive series of experiments on synthetic and real image datasets to evaluate the efficacy of our framework, going beyond the usual set of ShapeNet experiments to demonstrate its flexibility. Our experiments show that pixelNeRF can generate novel views from a single image input for both category-specific and category-agnostic settings, even in the case of unseen object categories. Further, we test the flexibility of our framework, both with a new multi-object benchmark for ShapeNet, where pixelNeRF outperforms prior approaches, and with simulation-to-real transfer demonstration on real car images. Lastly, we test capabilities of pixelNeRF on real images using the DTU dataset , where despite being trained on under 100 scenes, it can generate plausible novel views of a real scene from three posed input views.

Related Work

Novel View Synthesis. The long-standing problem of novel view synthesis entails constructing new views of a scene from a set of input views. Early work achieved photorealistic results but required densely captured views of the scene . Recent work has made rapid progress toward photorealism for both wider ranges of novel views and sparser sets of input views, by using 3D representations based on neural networks . However, because these approaches fit a single model to each scene, they require many input views and substantial optimization time per scene.

There are methods that can predict novel view from few input views or even single images by learning shared priors across scenes. Methods in the tradition of use depth-guided image interpolation . More recently, the problem of predicting novel views from a single image has been explored . However, these methods employ 2.5D representations, and are therefore limited in the range of camera motions they can synthesize. In this work we infer a 3D volumetric NeRF representation, which allows novel view synthesis from larger baselines.

Sitzmann et al. introduces a representation based on a continuous 3D feature space to learn a prior across scene instances. However, using the learned prior at test time requires further optimization with known absolute camera poses. In contrast, our approach is completely feed-forward and only requires relative camera poses. We offer extensive comparisons with this approach to demonstrate the advantages our design affords. Lastly, note that concurrent work adds image features to NeRF. A key difference is that we operate in view rather than canonical space, which makes our approach applicable in more general settings. Moreover, we extensively demonstrate our method’s performance in few-shot view synthesis, while GRF shows very limited quantitative results for this task.

Learning-based 3D reconstruction. Advances in deep learning have led to rapid progress in single-view or multi-view 3D reconstruction. Many approaches propose learning frameworks with various 3D representations that require ground-truth 3D models for supervision. Multi-view supervision is less restrictive and more ecologically plausible. However, many of these methods require object masks; in contrast, pixelNeRF can be trained from images alone, allowing it to be applied to scenes of two objects without modification.

Most single-view 3D reconstruction methods condition neural 3D representations on input images. The majority employs global image features , which, while memory efficient, cannot preserve details that are present in the image and often lead to retrieval-like results. Spatially-aligned local image features have been shown to achieve detailed reconstructions from a single view . However, both of these methods require 3D supervision. Our method is inspired by these approaches, but only requires multi-view supervision.

Within existing methods, the types of scenes that can be reconstructed are limited, particularly so for object-centric approaches (e.g. ). CoReNet reconstructs scenes with multiple objects via a voxel grid with offsets, but it requires 3D supervision including the identity and placement of objects. In comparison, we formulate a scene-level learning framework that can in principle be trained to scenes of arbitrary structure.

Viewer-centric 3D reconstruction For the 3D learning task, prediction can be done either in a viewer-centered coordinate system, i.e. view space, or in an object-centered coordinate system, i.e. canonical space. Most existing methods predict in canonical space, where all objects of a semantic category are aligned to a consistent orientation. While this makes learning spatial regularities easier, using a canonical space inhibits prediction performance on unseen object categories and scenes with more than one object, where there is no pre-defined or well-defined canonical pose. PixelNeRF operates in view-space, which has been shown to allow better reconstruction of unseen object categories in , and discourages the memorization of the training set . We summarize key aspects of our approach relative to prior work in Table 1.

Background: NeRF

The volumetric radiance field can then be rendered into a 2D image via

The rendered pixel value for camera ray r\mathbf{r} can then be compared against the corresponding ground truth pixel value, C(r)\mathbf{C}(\mathbf{r}), for all the camera rays of the target view with pose P\mathbf{P}. The NeRF rendering loss is thus given by

where R(P)\mathcal{R}(\mathbf{P}) is the set of all camera rays of target pose P\mathbf{P}.

Limitations While NeRF achieves state of the art novel view synthesis results, it is an optimization-based approach using geometric consistency as the sole signal, similar to classical multiview stereo methods . As such each scene must be optimized individually, with no knowledge shared between scenes. Not only is this time-consuming, but in the limit of single or extremely sparse views, it is unable to make use of any prior knowledge of the world to accelerate reconstruction or for shape completion.

Image-conditioned NeRF

To overcome the NeRF representation’s inability to share knowledge between scenes, we propose an architecture to condition a NeRF on spatial image features. Our model is comprised of two components: a fully-convolutional image encoder EE, which encodes the input image into a pixel-aligned feature grid, and a NeRF network ff which outputs color and density, given a spatial location and its corresponding encoded feature. We choose to model the spatial query in the input view’s camera space, rather than a canonical space, for the reasons discussed in §\S 2. We validate this design choice in our experiments on unseen object categories (§\S 5.2) and complex unseen scenes (§\S 5.3). The model is trained with the volume rendering method and loss described in §\S 3.

In the following, we first present our model for the single view case. We then show how this formulation can be easily extended to incorporate multiple input images.

We now describe our approach to render novel views from one input image. We fix our coordinate system as the view space of the input image and specify positions and camera rays in this coordinate system.

Given a input image I\mathbf{I} of a scene, we first extract a feature volume W=E(I)\mathbf{W}=E(\mathbf{I}). Then, for a point on a camera ray x\mathbf{x}, we retrieve the corresponding image feature by projecting x\mathbf{x} onto the image plane to the image coordinates π(x)\pi(\mathbf{x}) using known intrinsics, then bilinearly interpolating between the pixelwise features to extract the feature vector W(π(x))\mathbf{W}(\pi(\mathbf{x})). The image features are then passed into the NeRF network, along with the position and view direction (both in the input view coordinate system), as

where γ(⋅)\gamma(\cdot) is a positional encoding on x\mathbf{x} with 66 exponentially increasing frequencies introduced in the original NeRF . The image feature is incorporated as a residual at each layer; see §\S 5 for more information. We show our pipeline schematically in Fig. 2.

In the few-shot view synthesis task, the query view direction is a useful signal for determining the importance of a particular image feature in the NeRF network. If the query view direction is similar to the input view orientation, the model can rely more directly on the input; if it is dissimilar, the model must leverage the learned prior. Moreover, in the multi-view case, view directions could serve as a signal for the relevance and positioning of different views. For this reason, we input the view directions at the beginning of the NeRF network.

2 Incorporating Multiple Views

Multiple views provide additional information about the scene and resolve 3D geometric ambiguities inherent to the single-view case. We extend our model to allow for an arbitrary number of views at test time, which distinguishes our method from existing approaches that are designed to only use single input view at test time. Moreover, our formulation is independent of the choice of world space and the order of input views.

In the case that we have multiple input views of the scene, we assume only that the relative camera poses are known. For purposes of explanation, an arbitrary world coordinate system can be fixed for the scene. We denote the iith input image as I(i)\mathbf{I}^{(i)} and its associated camera transform from the world space to its view space as P(i)=[R(i)t(i)]\mathbf{P}^{(i)}=\begin{bmatrix}\mathbf{R}^{(i)}&\mathbf{t}^{(i)}\end{bmatrix}.

For a new target camera ray, we transform a query point x\mathbf{x}, with view direction d\mathbf{d}, into the coordinate system of each input view ii with the world to camera transform as

To obtain the output density and color, we process the coordinates and corresponding features in each view coordinate frame independently and aggregate across the views within the NeRF network. For ease of explanation, we denote the initial layers of the NeRF network as f1f_{1}, which process inputs in each input view space separately, and the final layers as f2f_{2}, which process the aggregated views.

We encode each input image into feature volume W(i)=E(I(i))\mathbf{W}^{(i)}=E(\mathbf{I}^{(i)}). For the view-space point x(i)\mathbf{x}^{(i)}, we extract the corresponding image feature from the feature volume W(i)\mathbf{W}^{(i)} at the projected image coordinate π(x(i))\pi(\mathbf{x}^{(i)}). We then pass these inputs into f1f_{1} to obtain intermediate vectors:

The intermediate V(i)\mathbf{V}^{(i)} are then aggregated with the average pooling operator ψ\psi and passed into a the final layers, denoted as f2f_{2}, to obtain the predicted density and color:

In the single-view special case, this simplifies to Equation 3 with f=f2∘f1f=f_{2}\circ f_{1}, by considering the view space as the world space. An illustration is provided in the supplemental.

Experiments

We extensively demonstrate our approach in three experimental categories: 1) existing ShapeNet benchmarks for category-specific and category-agnostic view synthesis, 2) ShapeNet scenes with unseen categories and multiple objects, both of which require geometric priors instead of recognition, as well as domain transfer to real car photos and 3) real scenes from the DTU MVS dataset .

Baselines For ShapeNet benchmarks, we compare quantitatively and qualitatively to SRN and DVR , the current state-of-the-art in few-shot novel-view synthesis and 2D-supervised single-view reconstruction respectively. We use the 2D multiview-supervised variant of DVR. In the category-agnostic setting (§\S 5.1.2), we also include grayscale rendering of SoftRas results. Color inference is not supported by the public SoftRas code. In the experiments with multiple ShapeNet objects, we compare with SRN, which can also model entire scenes.

For the experiment on the DTU dataset, we compare to NeRF trained on sparse views. Because NeRF is a test-time optimization method, we train a separate model for each scene in the test set.

Metrics We report the standard image quality metrics PSNR and SSIM for all evaluations. We also include LPIPS , which more accurately reflects human perception, in all evaluations except in the category-specific setup (§\S 5.1.1). In this setting, we exactly follow the protocol of SRN to remain comparable to prior works , for which source code is unavailable.

Implementation Details For the image encoder EE, to capture both local and global information effectively, we extract a feature pyramid from the image. We use a ResNet34 backbone pretrained on ImageNet for our experiments. Features are extracted prior to the first 44 pooling layers, upsampled using bilinear interpolation, and concatenated to form latent vectors of size 512512 aligned to each pixel.

To incorporate a point’s corresponding image feature into the NeRF network ff, we choose a ResNet architecture with a residual modulation rather than simply concatenating the feature vector with the point’s position and view direction. Specifically, we feed the encoded position and view direction through the network and add the image feature as a residual at the beginning of each ResNet block. We train an independent linear layer for each block residual, in a similar manner as AdaIn and SPADE , a method previously used with success in . Please refer to the supplemental for additional details.

We first evaluate our approach on category-specific and category-agnostic view synthesis tasks on ShapeNet.

We perform one-shot and two-shot view synthesis on the “chair” and “car” classes of ShapeNet, using the protocol and dataset introduced in . The dataset contains 6591 chairs and 3514 cars with a predefined split across object instances. All images have resolution 128×128128\times 128.

A single model is trained for each object class with 50 random views per object instance, randomly sampling either one or two of the training views to encode. For testing, We use 251 novel views on an Archimedean spiral for each object in the test set of object instances, fixing 1-2 informative views as input. We report our performance in comparison with state-of-the-art baselines in Table 2, and show selected qualitative results in Fig. 4. We also include the quantitative results of baselines TCO and dGQN reported in where applicable, and the values available in the recent works ENR and GRF in this setting.

PixelNeRF achieves noticeably superior results despite solving a problem significantly harder than SRN because we: 1) use feed-forward prediction, without test-time optimization, 2) do not use ground-truth absolute camera poses at test-time, 3) use view instead of canonical space.

Ablations. In Table 3, we show the benefit of using local features and view directions in our model for this category-specific setting. Conditioning the NeRF network on pixel-aligned local features instead of a global code (−-Local vs Full) improves performance significantly, for both single and two-view settings. Providing view directions (−-Dirs vs Full) also provides a significant boost. For these ablations, we follow an abbreviated evaluation protocol on ShapeNet chairs, using 25 novel views on the Archimedean spiral.

1.2 Category-agnostic Object Prior

While we found appreciable improvements over baselines in the simplest category-specific benchmark, our method is by no means constrained to it. We show in Table 4 and Fig. 5 that our approach offers a much greater advantage in the category-agnostic setting of , where we train a single model to the 1313 largest categories of ShapeNet. Please see the supplemental for randomly sampled results.

We follow community standards for 2D-supervised methods on multiple ShapeNet categories and use the renderings and splits from Kato et al. , which provide 24 fixed elevation views of 64×6464\times 64 resolution for each object instance. During both training and evaluation, a random view is selected as the input view for each object and shared across all baselines. The remaining 2323 views are used as target views for computing metrics (see §\S 5).

2 Pushing the Boundaries of ShapeNet

Taking a step towards reconstruction in less controlled capture scenarios, we perform experiments on ShapeNet data in three more challenging setups: 1) unseen object categories, 2) multiple-object scenes, and 3) simulation-to-real transfer on car images. In these settings, successful reconstruction requires geometric priors; recognition or retrieval alone is not sufficient.

Generalization to novel categories. We first aim to reconstruct ShapeNet categories which were not seen in training. Unlike the more standard category-agnostic task described in the previous section, such generalization is impossible with semantic information alone. The results in Table 5 and Fig. 6 suggest our method learns intrinsic geometric and appearance priors which are fairly effective even for objects quite distinct from those seen during training.

We loosely follow the protocol used for zero-shot cross-category reconstruction from [55, yan2017perspective]. Note that our baselines do not evaluate in this setting, and we adapt them for the sake of comparison. We train on the airplane, car, and chair categories and test on 10 categories unseen during training, continuing to use the Kato et al. renderings described in §\S 5.1.2.

Multiple-object scenes. We further perform few-shot 360°360\degree reconstruction for scenes with multiple randomly placed and oriented ShapeNet chairs. In this setting, the network cannot rely solely on semantic cues for correct object placement and completion. The priors learned by the network must be applicable in an arbitrary coordinate system. We show in Fig. 7 and Table 5 that our formulation allows us to perform well on these simple scenes without additional design modifications. In contrast, SRN models scenes in a canonical space and struggles on held-out scenes.

We generate training images composed with 20 views randomly sampled on the hemisphere and render test images composed of a held out test set of chair instances, with 50 views sampled on an Archimedean spiral. During training, we randomly encode two input views; at test-time, we fix two informative views across the compared methods. In the supplemental, we provide example images from our dataset as well as additional quantitative results and qualitative comparisons with varying numbers of input views.

Sim2Real on Cars. We also explore the performance of pixelNeRF on real images from the Stanford cars dataset . We directly apply car model from §\S 5.1.1 without any fine-tuning. As seen in Fig. 8, the network trained on synthetic data effectively infers shape and texture of the real cars, suggesting our model can transfer beyond the synthetic domain.

Synthesizing the 360°360\degree background from a single view is nontrivial and out of the scope for this work. For this demonstration, the off-the-shelf PointRend segmentation model is used to remove the background.

3 Scene Prior on Real Images

Finally, we demonstrate that our method is applicable for few-shot wide baseline novel-view synthesis on real scenes in the DTU MVS dataset . Learning a prior for view synthesis on this dataset poses significant challenges: not only does it consist of more complex scenes, without clear semantic similarities across scenes, it also contains inconsistent backgrounds and lighting between scenes. Moreover, under 100 scenes are available for training. We found that the standard data split introduced in MVSNet contains overlap between scenes of the training and test sets. Therefore, for our purposes, we use a different split of 88 training scenes and 15 test scenes, in which there are no shared or highly similar scenes between the two sets. Images are down-sampled to a resolution of 400×300400\times 300.

We train one model across all training scenes by encoding 33 random views of a scene. During test time, we choose a set of fixed informative input views shared across all instances. We show in Fig. 9 that our method can perform view synthesis on the held-out test scenes. We further quantitatively compare the performance of our feed-forward model with NeRF optimized to the same set of input views in Fig. 10. Note that training each of 60 NeRFs took 14 hours; in contrast, pixelNeRF is applied to new scenes immediately without any test-time optimization.

Discussion

We have presented pixelNeRF, a framework to learn a scene prior for reconstructing NeRFs from one or a few images. Through extensive experiments, we have established that our approach can be successfully applied in a variety of settings. We addressed some shortcomings of NeRF, but there are challenges yet to be explored: 1) Like NeRF, our rendering time is slow, and in fact, our runtime increases linearly when given more input views. Further, some methods (e.g. ) can recover a mesh from the image enabling fast rendering and manipulation afterwards, while NeRF-based representations cannot be converted to meshes very reliably. Improving NeRF’s efficiency is an important research question that can enable real-time applications. 2) As in the vanilla NeRF, we manually tune ray sampling bounds tn,tft_{n},t_{f} and a scale for the positional encoding. Making NeRF-related methods scale-invariant is a crucial challenge. 3) While we have demonstrated our method on real data from the DTU dataset, we acknowledge that this dataset was captured under controlled settings and has matching camera poses across all scenes with limited viewpoints. Ultimately, our approach is bottlenecked by the availability of large-scale wide baseline multi-view datasets, limiting the applicability to datasets such as ShapeNet and DTU. Learning a general prior for 360°360\degree scenes in-the-wild is an exciting direction for future work.

Acknowledgements

We thank Shubham Goel and Hang Gao for comments on the text. We also thank Emilien Dupont and Vincent Sitzmann for helpful discussions.

References

Appendix

A Additional Results

In this section, we provide additional qualitative and quantitative results for several key experiments. The reader is encouraged to refer to the video and website for a richer, animated presentation of qualitative results.

We show randomly sampled results for the category-agnostic setting (§\S 5.1.2) in Fig. 11, Fig. 12, and Fig. 13. Specifically, we sample 66 uniformly random objects for each of the 13 largest ShapeNet categories and show comparisons to the baselines as in the main paper. Two random views are selected from the 2424 available views to be source and target views respectively.

A.2 Generalization to novel categories

In Table 6 we show a detailed breakdown of metrics by category on unseen categories, as promised in the main paper.

A.3 Two-object Scenes

We show samples from our rendered dataset in Fig. 14. An analysis of performance as more views become available is in Table 7, for our method when compared with SRN. We also show randomly sampled results of scenes when given two input views in Figure 15. We train our model using two random views, and give the model either one, two, or three fixed informative views during inference.

A.4 DTU

In Fig. 16, we show quantitative results for each scene as well as renderings of of all test scenes not shown in the main paper.

In Table 8 we provide means and standard deviations of metrics for our method and NeRF on the DTU test set, with 1, 3, 6, 9 views. The PSNR here was plotted in Fig. 10 of the main paper

B Reproducibility

Here we describe implementation details in the interest of reproducibility. A general remark is that due to the high compute cost, we did not spend significant effort to tune the architecture or training procedure, and it is possible that variations can perform better, or that smaller models may suffice.

As briefly discussed in the main paper, we use a ResNet34 backbone and extract a feature pyramid by taking the feature maps prior to the first pooling operation and after the first ResNet 33 layers. For a H×WH\times W image, the feature maps have shapes

These are upsampled bilinearly to H/2×W/2H/2\times W/2 and concatenated into a volume of size 512×H/2×W/2512\times H/2\times W/2. For a 64×6464\times 64 image, to avoid losing too much resolution, we skip the first pooling layer, so that the image resolutions are at 1/2,1/2,1/4,1/81/2,1/2,1/4,1/8 of the input rather than 1/2,1/4,1/8,1/161/2,1/4,1/8,1/16. We use ImageNet pretrained weights provided through PyTorch.

We employ a fully-connected ResNet architecture with 55 ResNet blocks and width 512512, similar to that in . To enable arbitrary number of views as input, we aggregate across the source-views after block 33 using an average-pooling operation. This architecture is illustrated in Fig. 18. We remark that due to computational cost, we did not tune this architecture very much in practice.

To improve the sampling efficiency, in practice, we also use coarse and fine NeRF networks fc,fff_{c},f_{f} as in the vanilla NeRF , both of which share an identical architecture described above. Note that the encoder EE is not duplicated.

More precisely, we use 64 stratified uniform and 16 importance samples, and additionally take 16 fine samples with a normal distribution (SD 0.01) around the expected ray termination (i.e. depth) from the coarse model, to further promote denser sampling near the surface.

NeRF rendering hyperparameters We use positional encoding γ\gamma from NeRF for the spatial coordinates, with exponentially increasing frequencies:

Note that we do not apply the encoding to the view directions. In all experiments, we set L=6L=6. We also concatenate the input coordinates along the encoding as in the NeRF implementation. ω\omega is a scaling factor, set (rather arbitrarily) to 1.51.5 for the single-category, category-agnostic ShapeNet experiments as well as the DTU experiment, and to 2.02.0 for the multi-object experiment. While the exponent base can be tuned, in practice we left it at 22 as in NeRF.

The sampling bounds were set manually for each dataset. They were [1.25,2.75][1.25,2.75] for ShapeNet chairs, [0.8,1.8][0.8,1.8] for ShapeNet cars, [1.2,4.0][1.2,4.0] for Kato et al. renderings (category agnostic, novel category), [4.0,9.0][4.0,9.0] for our rendered 22-object dataset, and [0.1,5.0][0.1,5.0] for input.

We use a white background color in NeRF to match the ShapeNet renderings, except in the DTU setup where a black background is used.

We implement all models using the PyTorch framework.

B.2 Experimental Details

We first provide general details about the metrics and training procedure common to all experiments, then present more specific details for each experimental setting in subsections.

We use PSNR and SSIM from the scikit-image package as in SRN , whereas LPIPS is computed with the code provided by the LPIPS authors after normalizing the pixel values to the $$ range. We use the VGG network version of LPIPS following NeRF .

For all experiments, we take the learning rate to be 10−410^{-4}. We use a batch size of 44 instances and 128128 rays per instances.

B.2.1 Single-category ShapeNet

We train for 400000 iterations, which took roughly 66 days on a single Titan RTX. For efficiency, we sample rays from within a tight bounding box around the object for the first 300000 iterations, after which we remove the bounding box to avoid background artifacts. Further, we use 22 input views for the first 300000 iterations and after that, we randomly choose to take either 11 or 22 views as input to encourage the model to work with either 11 or 22 views.

SRN’s evaluation protocol is followed: in the 1-view case, we use view 6464 as input, and in the 2-view case, we use views 6464 and 128128.

For SRN , we use the pretrained chair model from the public GitHub repository. Note that SRN requires a test-time training step (latent inversion) to generate result images; we apply latent inversion for 170000 iterations for both the 1-view and 2-view cases for chairs.

Recall that, due to a camera sampling bug, we use an updated car dataset provided by the SRN author. Thus, we follow instructions in the Github to train a model on the new dataset; we train for 400000 iterations and apply latent inversion for 100000 iterations for each of the 1-view and 2-view cases. Note the quantitative results we report are slightly lower than that in in the single-view case, but substantially higher than in the original SRN paper, which used the bugged renderings. For the remaining baselines, we only report numbers from the relevant papers on the same task.

B.2.2 Category-agnostic ShapeNet

We train our model for 800000 iterations on the entire training set, where rays are sampled from within a tight bounding box for the first 400000 iterations. This took about 6 days on an RTX 2080Ti.

As discussed in the main paper, we evaluate on the test split from as provided by DVR . To ensure fairness, we sampled a random input view to encode for each object and use this view for all baselines as well.

For DVR , we use the pretrained 2D multiview-supervised model from the public GitHub and the provided rendering code (in render.py). For SoftRas , we similarly use the pretrained ShapeNet model from the public GitHub repo and obtain images using their renderer library.

Since SRN did not originally evaluate in this setting, we train a model for this category-agnostic setting using the public code. We train for 1 million iterations and perform latent inversion for 260000 iterations, taking about 14 days on a Titan RTX in total.

B.2.3 Generalization to Novel Categories

We train our model for 680000 iterations across all instances of 33 categories: airplane, car, and chair. Rays are sampled from within a tight bounding box for the first 400000 iterations. This took about 5 days on a GTX 1080Ti.

Since there are more than 25000 objects in the 10 remaining categories, it would be computationally prohibitive to evaluate on all of them. For our purposes, we sample 25% of the objects from each category for testing, using the remaining for validation. The protocol otherwise remains the same as in the category-agnostic model (§\S B.2.2).

We train SRN and DVR as in §\S B.2.2. For DVR we turned off the use of visual hull depth for sampling, since this information was not provided for all instances of the dataset shipped with DVR.

B.2.4 Two-object Scenes

We generate more complex synthetic scenes consisting of two ShapeNet chairs. We subdivide ShapeNet chairs into 2715 training instances and 1101 test instances. We generate scenes by randomly placing instances within each split around the origin, rotated randomly about each object’s z-axis, and render 128×128128\times 128 resolution images. Per instance, we render 20 training views sampled binned uniform on the hemisphere, and 50 testing views sampled on an Archimedean spiral, similar to the SRN protocol.

We compare with SRN as our baseline on this task, using the publicly available code. We note that SRN performs prediction in a canonical object-centric coordinate system, and used this version for this task. We train a model for evaluation on this two-object dataset using one, two, and three input views. We first train the model for 1 million iterations. Then for each number of input views, we fix the set of input views per instance and perform latent inversion for 150,000 iterations.

B.2.5 Sim2Real on Real Car Images

We use car images from the Stanford Cars dataset . PointRend is applied to the images to obtain foreground masks and bounding boxes. After removing the background with this mask, the image is then translated and rescaled so that the center of the bounding box is at the center of the image and the shorter side of the bounding box is 1/4 of the image side length, 128. This normalization heuristic is motivated by the observation that the shorter side roughly corresponds to the height or width of the car, which is a more constant quantity than the length.

For evaluation, we set the camera pose to identity and use the same sampling strategy and bounds as at train time for the single-category cars model.

B.2.6 Real Images on DTU

A single model is trained on the 88 training scenes. We use exposure level 33 only. Note that while there are several views per scene with incorrect exposure throughout the DTU dataset, we did not remove them for training purposes. At each training step, a random color jitter augmentation is applied equally to all views of each object.

While we solve a very different task from MVSNet which predicts depth maps from short-baseline views and is 2.5D supervised, we considered using the MVSNet DTU split to conform to standards for training on DTU. However, we found that the the split contained effective overlap across the train/val/test sets, making it a poor benchmark of cross-scene generalization, as shown in Fig. 17. For our purposes, we created a different split to avoid this issue: we use scans 8, 21, 30, 31, 34, 38, 40, 41, 45, 55, 63, 82, 103, 110, 114 for testing and all other scans except 1, 2, 7, 25, 26, 27, 29, 39, 51, 54, 56, 57, 58, 73, 83, 111, 112, 113, 115, 116, 117 for training.

We downsampled all DTU images to 400×300400\times 300 and adjusted the world scale of all scans by a factor or 1/3001/300.

We separately evaluate using 1, 3, 6, 9 informative input views and calculate image metrics with the remaining views. Specifically, we selected views 25, 22, 28, 40, 44,48, 0, 8, 13 for input, taking a prefix of this list when less than 99 views are used. We exclude the views with bad exposure (they are 3, 4, 5, 6, 7, 16, 17, 18, 19, 20, 21, 36, 37, 38, 39) for testing.

We train a total of 60 NeRFs for comparison, one for each scene and number of input views, using the original NeRF TensorFlow code. Each NeRF is trained for 400000 iterations with ray batch size 128, which takes about 14 hours on an RTX Titan, to ensure convergence.

We found that NeRF did not converge in 5 cases when given 6 or 9 views, including in the case of smurf (scan 82) shown in the video. This is possibly due to exposure variation in the DTU dataset. For these scenes, we initialized the model to the trained weights for the same scene with 33 less views and train for about 200000200000 additional iterations to get reasonable results.