FWD: Real-time Novel View Synthesis with Forward Warping and Depth

Ang Cao, Chris Rockwell, Justin Johnson

Introduction

Given several posed images, novel view synthesis (NVS) aims to generate photorealistic images depicting the scene from unseen viewpoints. This long-standing task has applications in graphics, VR/AR, bringing life to still images. It requires a deep visual understanding of geometry and semantics, making it appealing to test visual understanding.

Early work on NVS focused on image-based rendering (IBR), where models generate target views from a set of input images. Light field or proxy geometry (like mesh surfaces) are typically constructed from posed inputs, and target views are synthesized by resampling or blending warped inputs. Requiring dense input images, these methods are limited by 3D reconstruction quality, and can perform poorly with sparse input images.

Recently, Neural Radiance Fields (NeRF) have become the leading methods for NVS, using MLPs to represent the 5D radiance field of the scene implicitly. The color and density of each sampling point are queried from the network and aggregated by volumetric rendering to get the pixel color. With dense sampling points and differentiable renderer, explicit geometry isn’t needed, and densities optimized for synthesis quality are learned. Despite impressive results, they are not generalizable, requiring MLP fitting for each scene with dense inputs. Also, they are extremely slow because of tremendous MLP query times for a single image.

Generalizable NeRF variants like PixelNeRF , IBRNet and MVSNeRF emerge very recently, synthesizing novel views of unseen scenes without per-scene optimization by modeling an MLP conditioned on sparse inputs. However, they still query the MLP millions of times, leading to slow speeds. Albeit the progress of accelerating NeRF with per-scene optimization , fast and generalizable NeRF variants are still under-explored.

In this paper, we target a generalizable NVS method with sparse inputs, refraining dense view collections. Both real-time speed and high-quality synthesis are expected, allowing interactive applications. Classical IBR methods are fast but require dense input views for good results. Generalizable NeRF variants show excellent quality without per-scene optimization but require intense computations, leading to slow speeds. Our method, termed FWD, achieves this target by Forward Warping features based on Depths.

Our key insight is that explicitly representing the depth of each input pixel allows us to apply forward warping to each input view using a differentiable point cloud renderer. This avoids the expensive volumetric sampling used in NeRF-like methods, enabling real-time speed while maintaining high image quality. This idea is deeply inspired by the success of SynSin , which employs a differentiable point cloud renderer for single image NVS. Our paper extends SynSin to multiple inputs settings and explores effective and efficient methods to fuse multi-view information.

Like prior NVS methods, our approach can be trained with RGB data only, but it can be progressively enhanced if noisy sensor depth data is available during training or inference. Depth sensors are becoming more prevalent in consumer devices such as the iPhone 13 Pro and the LG G8 ThinQ, making RGB-D data more accessible than ever. For this reason, we believe that methods making use of RGB-D will become increasingly useful over time.

Our method estimates depths for each input view to build a point cloud of latent features, then synthesizes novel views via a point cloud renderer. To alleviate the inconsistencies between observations from various viewpoints, we introduce a view-dependent feature MLP into point clouds to model view-dependent effects. We also propose a novel Transformer-based fusion module to effectively combine features from multiple inputs. A refinement module is employed to inpaint missing regions and further improve synthesis quality. The whole model is trained end-to-end to minimize photometric and perceptual losses, learning depth and features optimized for synthesis quality.

Our design possesses several advantages compared with existing methods. First, it gives both high-quality and high-speed synthesis. Using explicit point clouds enables real-time rendering. In the meanwhile, differentiable renderer and end-to-end training empower high-quality synthesis results. Also, compared to NeRF-like methods, which cannot synthesize whole images during training because of intensive computations, our method could easily utilize perceptual loss and refinement module, which noticeably improves the visual quality of synthesis. Moreover, our model can seamlessly integrate sensor depths to further improve synthesis quality. Experimental results support these analyses.

We evaluate our method on the ShapeNet and DTU datasets, comparing it with representative NeRF-variants and IBR methods. It outperforms existing methods, considering speed and quality jointly: compared to IBR methods we improve both speed and quality; compared to recent NeRF-based methods we achieve competitive quality at realtime speeds (130-1000×\times speedup). A user study demonstrates that our method gives the most perceptually pleasing results among all methods. The code is available at https://github.com/Caoang327/fwd_code.

Related Work

Novel view synthesis is a long-standing problem in computer vision, allowing for the generation of novel views given several scene images. A variety of 3D representations (both implicit and explicit) have been used for NVS, including depth and multi-plane images , voxels , meshes , point clouds and neural scene representations . In this work, we use point clouds as our 3D representations for computational and memory efficiency.

Image-based Rendering. IBR synthesizes novel views from a set of reference images by weighted blending . They generally estimate proxy geometry from dense captured images for synthesis. For instance, Riegler et al. uses multi-view stereo to produce scene mesh surface and warps source view images to target views based on proxy geometry. Despite promising results in some cases, they are essentially limited by the quality of 3D reconstructions, where dense inputs (tens to hundreds) with large overlap and reasonable baselines are necessary for decent results. These methods estimate geometry as an intermediate task not directly optimized for image quality. In contrast we input sparse views and learn depth jointly to optimize for synthesis quality.

Neural Scene Representations. Recent work uses implicit scene representations for view synthesis . Given many views, neural radiance fields (NeRF) show impressive results , but require expensive per-scene optimization. Recent methods generalize NeRF without per-scene optimization by learning a shared prior, with sparse inputs. However these methods require expensive ray sampling and therefore are very slow. In contrast, we achieve significant speedups using explicit representations. Some concurrent work accelerates NeRF by reformulating the computation , using precomputation , or adding view dependence to explicit 3D representations ; unlike ours, these all require dense input views and per-scene optimization.

Utilizing RGB-D in NVS. The growing availability of annotated depth maps facilitates depth utilization in NVS , which serves as extra supervision or input to networks. Our method utilizes explicit depths as 3D representations, allowing using sensor depths as additional inputs for better quality. Given the increasing popularity of depth sensors, integrating sensor depths is a promising direction for real-world applications.

Depth has been used in neural scene representations for speedups , spaser inputs and dynamic scenes. However, these works still require per-scene optimization. Utilizing RGB-D inputs to accelerate generalizable NeRF like is still an open problem.

Differentiable Rendering and Refinement. We use advances in differentiable rendering to learn 3D end-to-end. Learned geometries rely heavily on rendering and refinement to quickly synthesize realistic results. Refinement has improved dramatically owing to generative modeling and rendering frameworks . Instead of aggregating information across viewpoints before rendering , we render viewpoints separately and fuse using a Transformer , enabling attention across input views.

Method

Given a sparse set of input images {Ii}i=1N\{I_{i}\}_{i=1}^{N} and corresponding camera poses {Ri,Ti}\{R_{i},T_{i}\}, our goal is to synthesize a novel view with camera pose {Rt,Tt}\{R_{t},T_{t}\} fast and effectively. The depths {Disen}\{D^{sen}_{i}\} of IiI_{i} captured from sensors are optionally available, which are generally incomplete and noisy.

The insight of our method is that using explicit depths and forward warping enables real-time rendering speed and tremendous accelerations. Meanwhile, to alleviate quality degradations caused by inaccurate depth estimations, a differentiable renderer and well-designed fusion & refinement modules are employed, encouraging the model to learn geometry and features optimized for synthesis quality.

As illustrated in Figure 2, with estimated depths, input view IiI_{i} is converted to a 3D point cloud Pi\mathcal{P}_{i} containing geometries and view-dependent semantics of the view. A differentiable neural point cloud renderer π\pi is used to project point clouds to target viewpoints. Rather than directly aggregating point clouds across views before rendering, we propose a Transformer-based module TT fusing rendered results at target view. Finally, a refinement module RR is employed to generate final outputs. The whole model is trained end-to-end with photometric and perceptual loss.

We use point clouds to represent scenes due to their efficiency, compact memory usage, and scalability to complex scenes. For input view IiI_{i}, point cloud Pi\mathcal{P}_{i} is constructed by estimating depth DiD_{i} and feature vectors Fi′F^{\prime}_{i} for each pixel in the input image, then projecting the feature vectors into 3D space using known camera intrinsics. The depth DiD_{i} is estimated by a depth network dd; features Fi′F^{\prime}_{i} are computed by a spatial feature encoder ff and view-dependent MLP ψ\psi.

Spatial Feature Encoder ff. Scene semantics of input view IiI_{i} are mapped to per-pixel feature vectors FiF_{i} by spatial feature encoder ff. Each feature vector in FiF_{i} is 61-dimensions and is concatenated with RGB channels, which is 64 dimensions in total. ff is built on BigGAN architecture .

Depth Network dd. Estimating depth from a single image has scaling/shifting ambiguity, losing valuable multi-view cues and leading to inconsistent estimations across views. Applying multi-view stereo algorithms (MVS) solely on sparse inputs is challenging because of limited overlap and huge baselines between input views, leading to inaccurate and low-confidence estimations. Therefore, we employ a hybrid design cascading a U-Net after the MVS module. The U-Net takes image IiI_{i} and estimated depths from the MVS module as inputs, refining depths with multi-view stereo cues and image cues. PatchmatchNet is utilized as the MVS module, which is fast and lightweight.

Depth Estimation with sensor depths. As stated, U-Net receives an initial depth estimation from the MVS module and outputs a refined depth used to build the point cloud. If sensor depth DisenD^{sen}_{i} is available, it is directly input to the U-Net as the initial depth estimations. In this setting, U-Net servers as completion and refinement module taking DisenD^{sen}_{i} and IiI_{i} as inputs, since DisenD^{sen}_{i} is usually noisy and incomplete. During training, loss LsL_{s} is employed to encourage the U-Net output to match the sensor depth.

where MiM_{i} is a binary mask indicating valid sensor depths.

View-Dependent Feature MLP ψ\psi. The appearance of the point could vary across views because of lighting and view direction, causing inconsistency between multiple views. Therefore, we propose to insert view direction changes into scene semantics to model this view-dependent effects. An MLP ψ\psi is designed to compute view-dependent features Fi′F^{\prime}_{i} by taking FiF_{i} and relative view changes Δv\Delta v from input to target view as inputs. For each point in the cloud, Δv\Delta v is calculated based on normalized view directions viv_{i} and vtv_{t}, from the point to camera centers of input view ii and target view tt. The relative view direction change is calculated as:

and the view-dependent feature Fi′F^{\prime}_{i} is:

where δ\delta is a two-layer MLP mapping Δv\Delta v to a 32-dimensions vector and ψ\psi is also a two-layer MLP.

2 Point Cloud Renderer

We use the differentiable renderer design of , which splats 3D points to the image plane and gets pixel values by blending point features. The blending weights are computed based on z-buffer depths and distances between pixel and point centers. It is implemented using Pytorch3D .

This fully differentiable renderer allows our model to be trained end-to-end, where photometric and perceptual loss gradients can be propagated to points’ position and features. In this way, the model learns to estimate depths and features optimized for synthesis quality, leading to superior quality. We show the effectiveness of it in experiments.

3 Fusion and Refinement

Unlike SynSin using a single image for NVS, fusing multi-view inputs is required in our method. A naïve fusion transforms each point cloud to target view and aggregates them into a large one for rendering. Despite high efficiency, it is vulnerable to inaccurate depths since points with wrong depths may occlude points from other views, leading to degraded results. Methods like PointNet may be feasible to apply on the aggregated point cloud for refinement, but they are not efficient with significant point numbers.

Instead, we render each point cloud individually at target viewpoints and fuse rendered results by a fusion Transformer TT. A refinement module RR is used to inpaint missing regions, decode feature maps and improve synthesis quality.

Rendered results may lose geometry cues for fusion when rendered from 3D to 2D. For instance, depths may reveal occlusion relationships across views, and relative view changes from input to target views relate to each input’s importance for fusion. Therefore, we also explored to use geometry features as position encoding while not helpful.

4 Training and Implementation Details

Our model is trained end-to-end with photometric Ll2\mathcal{L}_{l_{2}} and perceptual Lc\mathcal{L}_{c} losses between generated and ground-truth target images. The whole loss function is:

where λl2=5.0,λc=1.0\lambda_{l_{2}}=5.0,\lambda_{c}=1.0. The model is trained end-to-end on 4 2080Ti GPUs for 3 days, using Adam with learning rate 10−410^{-4} and β1=0.9,β2=0.999\beta_{1}{=}0.9,\beta_{2}{=}0.999. When sensors depths are available as inputs, Ls\mathcal{L}_{s} is used with λs=5.0\lambda_{s}=5.0.

Experiments

The goal of our paper is real-time and generalizable novel view synthesis with sparse inputs, which can optionally use sensor depths. To this end, our experiments aim to identify the speed and quality at which our method can synthesize novel images and explore the advantage of explicit depths. We evaluate our methods on ShapeNet and DTU datasets, comparing results with the SOTA methods and alternative approaches. Experiments take place with held-out test scenes and no per-scene optimization. We conduct ablations to validate the effectiveness of designs.

Metrics. We conduct A/B test to measure the visual quality, in which workers select the image most similar to the ground truth from competing methods. Automatic image quality metrics including PSNR, SSIM and LPIPS are also reported, and we find LPIPS best reflects the image quality as perceived by humans. Frames per second (FPS) during rendering is measured on the same platform (single 2080Ti GPU with 4 CPU cores). All evaluations are conducted using the same protocol (same inputs and outputs).

Model Variants. Three models are evaluated with various accessibility to depths for training and test, as defined in Table 1. FWD utilizes a pretrained PatchmatchNet as the MVS module for depth estimations, which is also updated during end-to-end training with photometric and perceptual loss. FWD-U learns depth estimations in an Unsupervised manner, sharing the same model and settings as FWD while PatchmatchNet is randomly initialized without any pretraining. FWD-D takes sensor depths as additional inputs during both training and inference. It doesn’t use any MVS module since sensor depths provide abundant geometry cues. For pretraining PatchmatchNet, we train it following typical MVS settings and using the same data splitting as NVS.

We first evaluate our model for category-agnostic synthesis task on ShapeNet . Following the setting of , we train and evaluate a single model on 13 ShapeNet categories. Each instance contains 24 fixed views of 64 ×\times 64 resolution. During training, one random view is selected as input and the rests are served as target views. For testing, we synthesize all the other views from a fixed informative view. The model is finetuned with two random input views for 2-view experiments. We find that U-Net is sufficient for good results on this dataset without the MVS module.

Qualitative comparisons to PixelNeRF are shown in Figure 4, where FWD-U gets noticeably superior results. Our synthesized results are more realistic and closely matching to target views, while PixelNeRF’s results tend to be blurry. We observe the same trend in the DTU benchmark and evaluate the visual quality quantitatively there.

We show quantitative results in Table 2, adding SRN and DVR as other baselines. Our method outperforms others significantly for LPIPS, indicating a much better perceptual quality, as corroborated by qualitative results. PixelNeRF has a slightly better PSNR while its results are blurry. Most importantly, FWD-U runs at a speed of over 300 FPS, which is 300×\times faster than PixelNeRF.

2 DTU MVS Benchmarks

We also evaluate our model on DTU MVS dataset , which is a real scene dataset consisting of 103 scenes. Each scene contains one or multiple objects placed on a table, while images and incomplete depths are collected by the camera and structured light scanner mounted on an industrial robot arm. Corresponding camera poses are provided.

As stated in , this dataset is challenging since it consists of complex real scenes without apparent semantic similarities across scenes. Also, images are taken under varying lighting conditions with distinct color inconsistencies between views. Moreover, with only under 100 scenes available for training, it is prone to overfitting in training.

We follow the same training and evaluation pipelines as PixelNeRF for all methods to give a fair comparison. The data consists of 88 training and 15 test scenes, between which there are no shared or highly similar scenes. Images are down-sampled to a resolution of 300 ×\times 400. For training, three input views are randomly sampled, with the rest as target views. For inference, we choose three fixed informative input views and synthesize other views of the scene.

Baselines. We evaluate a set of representatives of generalizable NeRF and IBR methods in two different scenarios: with RGB or RGB-D available as inputs during inference.

PixelNeRF , IBRNet and MVSNeRF are the SOTA generalizable NeRF variants, taking RGB as inputs. We use the official PixelNeRF model trained on DTU MVS and carefully retrain IBRNet and MVSNeRF with the same 3-input-view settings. PixelNeRF-DS is also included as reported in , which is PixelNeRF supervised with depths. Please note that our settings are very different from evaluations used in original papers of IBRNet and MVSNeRF.

A series of IBR methods are also evaluated. Since COLMAP fails to give reasonable outputs with sparse input images, methods using COLMAP like FVS , DeepBlending cannot estimate scene geometry in this setting. For these methods, we use depths captured by sensors as estimated depths, which should give upper-bound performance of these methods. To better cope with missing regions, we add our refinement model to DeepBlending and retrain it on DTU dataset, termed Blending-R.

For fairness, we evaluate all methods using the same protocol, distinct from some of their original settings. Although we try our best to adopt these methods, our reported results may still not perfectly reflect their true capacity.

Qualitative Results. Synthesis results are shown in Figure 5, where high-quality and geometrically correct novel views are synthesized in real-time (over 35 FPS) under significant viewpoint changes. Our refinement module faithfully inpaints invisible regions; also, synthesized images have good shadows, light reflections, and varying appearances across views, showing the efficacy of view-dependent MLP. With sensor depths, results can be further improved.

We show comparisons to baselines in Figure 6. Our methods provide noticeably better results than baselines across different depth settings. For models without depths in test, IBRNet and PixelNeRF give blurry results in areas of high detail such as the buildings in the top row, while our FWD-U and FWD give more realistic and sharper images. With sensor depths in test, baseline Blending-R produces more cogent outputs, but still struggles to distinguish objects from the background, such as in the middle row, while FWD-D gives faithfully synthesis and clear boundaries.

Quantitative Results. We evaluate synthesis quality quantitatively by user study following a standard A/B paradigm. Workers choose the closest to a ground truth image between competing methods, and are monitored using a qualifier and sentinel examples. All views in the test set (690 in total) are evaluated, and each view is judged by three workers.

In Figure 7, user study results support qualitative observations. Among all baselines with and without test depths , users choose our method as more closely matching ground truth images than others most of the time. FWD-U is selected over PixelNeRF in 65.6% of examples, and 77.8% compared to IBRNet. Also, over 90% workers prefer FWD-D to FWD , showing advantage of using sensor depths.

We show automated view synthesis metrics and speeds in Table 3. Across all depth availability settings, our method is competitive with the SOTA baselines while significantly faster. FWD-D runs in real-time and gives substantially better image quality than others. FWD has competitive metrics to PixelNeRF-DS while 1000×\times faster. Notably, NeRF variants such as PixelNeRF, IBRNet, MVSNeRF, and PixelNeRF-DS are at least two orders of magnitude slower.

The exception to highly competitive performance is weaker PSNR and SSIM of our unsupervised FWD-U against PixelNeRF and IBRNet. However, FWD-U has noticeably better perceptual quality with the best LPIPS, and human raters prefer it to other methods in A/B tests. The visual quality in figure 6 also illustrates the disparity between comparisons using PSNR and LPIPS. Meanwhile, FWD-U is above 1000×1000\times faster than PixelNeRF and above 100×100\times faster than IBRNet. Depth estimations, rendering and CNN would introduce tiny pixel shiftings, which harm the PSNR of our method. NeRF-like methods are trained to optimize L2 loss for each pixel independently, leading to blur results.

Among all methods without test depths, FWD has the best results. Although it uses a pretrained MVS module, we think this comparison is still reasonable since pretrained depth module is easy to get. Also, training depths can be easily calculated from training images since they are dense.

Baseline comparisons also show that IBR methods are fast, but do not give images that are competitive with our method. Our method outperforms them in both perceptual quality and standard metrics, showing the efficacy of proposed methods. Note that Blending+R doesn’t support variable number inputs and our refinement module improves its results significantly. We also compare FWD-U with SynSin which only receives a single input image, showing the benefits of using multi-view inputs in NVS.

3 Ablations and Analysis

We evaluate the effectiveness of our designs and study depth in more detail through ablation experiments.

Effects of Fusion Transformer. We design a model without Transformer, which concatenates point clouds across views into a bigger one for later rendering and refinement. Its results in FWD-U settings are shown in Figure 8. The ablated version is vulnerable to inaccurate depths learned in unsupervised manner and synthesizes “ghost objects” since points with bad depths occlude other views’ points.

We repeat the same ablation in FWD-D settings, shown in Table 4, which settings give much better depth estimations with sensor depths. The ablated model has notably worse results for all metrics, indicating that the proposed method is not only powerful to tackle inaccurate depth estimations, but also fuse semantic features effectively.

Effects of View Dependent MLP. For ablation, we remove the view-dependent feature MLP and report its results in Table 4. Removing this module reduces model’s ability to produce view-dependent appearances, leading to worse performance for all metrics. We show more results in Supp.

Depth Analysis and Ablations. We visualize depths in Figure 9. Estimating depths from sparse inputs is challenging and gives less accurate results because of the huge baselines between inputs. We show estimated depths from PatchmatchNet here, filtered based on the confidence scores. Therefore, refinement is essential in our design to propagate multi-view geometry cues to the whole image. Our end-to-end model learns it by synthesis losses.

We ablate the depth network in Table 5 and report depth error δ3cm\delta_{3cm}, which is the percentage of estimated depths within 3 cm of sensor depths. MVS module is critical (row 2) to give geometrically consistent depths. U-Net further refines depths and improves the synthesis quality (row 3). PatchmatchNet has its own shallow refinement layer, already giving decent refinements. Learning unsupervised MVS and NVS jointly from scratch is challenging (row 4), and training depth network without supervision first may give a good initialization for further jointly training.

Conclusion

We propose a real-time and generalizable method for NVS with sparse inputs by using explicit depths. This method inherits the core idea of SynSin while extending it to multi-view input settings, which is more challenging. Our experiments show that estimating depths can give impressive results with a real-time speed, outperforming existing methods. Moreover, the proposed method could utilize sensor depths seamlessly and improve synthesis quality significantly. With the increasing availability of mobile depth sensors, we believe our method has exciting real-world 3D applications. We acknowledge there could be the potential for the technology to be used for negative purposes by nefarious actors, like synthesizing fake images for cheating.

There are also challenges and limitations yet to be explored. 1) Although using explicit depths gives tremendous speedups, it potentially inflicts depth reliance on our model. We designed a hybrid depth regressor to improve the quality of depth by combining MVS and single image depth estimations. We also employed an effective fusion and refinement module to reduce the degrades caused by inaccurate depths. Despite these designs, the depth estimator may still work poorly in some challenging settings (like very wide camera baselines), and it would influence the synthesis results. Exploring other depth estimation methods like MiDaS could be an interesting direction for future work.

2) The potential capacity of our method is not fully explored. Like SynSin , our model (depth/feature network and refinement module especially) is suitable and beneficial from large-scale training data, while the DTU MVS dataset is not big enough and easy to overfit during training. Evaluating our method on large-scale datasets like Hypersim would potentially reveal more advantages of our model, which dataset is very challenging for NeRF-like methods.

3) Although our method gives more visually appealing results, our PSNR and SSIM are lower than NeRF-like methods. We hypothesize that our refinement module is not perfectly trained to decode RGB colors from feature vectors because of limited training data. Also, tiny misalignments caused during the rendering process may also harm the PSNR, although it is not perceptually visible.

Acknowledgement. Toyota Research Institute provided funds to support this work. We thank Dandan Shan, Hao Ouyang, Jiaxin Xie, Linyi Jin, Shengyi Qian for helpful discussions.

References

Model Architectures and Training details

As stated in the paper, our model consists of spatial feature network ff, depth network dd, view-dependent feature MLP ψ\psi, neural point cloud renderer π\pi, fusion transformer TT and refinement module RR. We show architecture details.

Spatial Feature Network ff. The spatial feature network ff contains 8 ResNet blocks, with output channels of 32, 32, 32, 64, 64, 64,64, 61 and no downsampling. Each ResNet block utilizes 3 ×\times 3 /stride 1/padding 1 convolution followed by instance norm and ReLU. Similar to , spectral normalization is utilized for each convolution for stable training. The input images are padded by 0 to fit network.

Depth Network dd. We utilize a classical U-Net for depth refinement, consisting 8 downsampling blocks and 8 upsampling blocks. We pad constant zero to make the input have feasible shape. For PatchMatchNet to estimate the initial depths, we follow the original pipeline, in which we downsample the input images into 4 scales and predict the depths in a coarse-to-fine manner.

View-dependent feature MLP ψ\psi. The relative view direction change vector is first passed through a two-layer MLP without 16 and 32 output features perspectively to get a 32-dims feature embedding. Then this 32 dims feature embedding is concatenated with the original 64 dims feature vector and passed through another two-layer MLP with output 64 and 64 output features. We use ReLU as activation function following MLP and no normalization layers.

Neural point cloud renderer π\pi. The point cloud renderer is implemented in Pytorch3D , which takes a point cloud and pose PP and project it to a 150 ×\times 200 feature map, where each pixel has 64-dims feature. The renderer fills zero with 64-dims for pixels which are invisible.

The blending weight α\alpha of the 3D point xx for pixel ll is

where ss is the Euclidean distance between point xx’s splatted center and pixel ll; rr is the radius of point xx in NDC coordinate. We set r=1.5r=1.5 pixels in our experiments. To render the value of each pixel, we employ alpha-compositing to blend all feasible points. The rendered feature FlF_{l} of pixel ll is:

where FiF_{i}, αi\alpha_{i} is the feature and alpha of 3D point ii. The rendering is based on points’ depths and we blend the Top K points with the nearest depths. We use K=16K=16 in our paper, meaning we blend at most 16 points to get the results of one pixel.

We compute DlD_{l}, the depth of pixel ll using the blending weights by:

Fusion transformer TT. The Key and Value are feature vectors from multiple views, with the shape as NView×N×C\times N\times C, where NView is the number of input views (3 in our experiments), NN is the batch size, which is H×WH\times W, and CC is the feature dimension, which is 64. The query is a learnable token with the shape as 1×C1\times C, and expanded to N batch. The transformer is a multi-head attention with 4 heads, projecting key into 16 dims embedding and output 64 dims vectors.

Refinement module RR. We use a ResNet decoder as our refinement module, which consists 8 ResNet blocks, in which we downsample at the 3rd layers and upsample at 6th and 7th layers. The output feature dims are 64, 128, 256, 256, 128, 128, 128, 3.

Two-stage training for FWD-U. We find that a two-stage training scheme would slightly improve the performance of FWD-U model ( 0.4dB). In stage one, we first train the depth networks by constructing RGB point clouds for each input view and projecting them to the target view via differentiable point cloud rendering. We directly aggregate all the point clouds from various input views and get the rendered RGB images at target views. The depth network is trained by photometric loss between rendered views and target images. This step works as a simple unsupervised MVS scheme and gives a proper initialization of the depth network. We then train the whole model with the initialized depth network in stage two. This two-stage training scheme’s intuition is that jointly training depth network and other components from scratch are unstable, and our first training stage could give a initialization for depth network.

License Discussions

ShapeNet Dataset . We conduct our experiments on the ShapeNet dataset and we cite the paper as required by the author. More specifically, we download the data from NMR, which is hosted by DVR authors.

DTU MVS Dataset . We conduct our experiments on the DTU MVS dataset. This dataset doesn’t include any license. On the other hand, we cite the paper which is required by the paper.

2 Methods

PixelNeRF . We evaluate the official code of PixelNeRF for comparison. The code is hosted in the github page https://github.com/sxyu/pixel-nerf, which uses the BSD 2-Clause ”Simplified” License.

IBRNet . We use the official code of IBRNet for comparison, which is hosted at https://github.com/googleinterns/IBRNet. This project has the Apache License 2.0.

MVSNeRF . We use the official code of MVSNeRF for comparison, which is hosted at https://github.com/apchenstu/mvsnerf with MIT License.

SynSin . The code of SynSin is hosted at https://github.com/facebookresearch/synsin with Copyright (c) 2020, Facebook All rights reserved.

Stable View Synthesis (SVS) . We use the code hosted at https://github.com/isl-org/StableViewSynthesis for evaluation, with MIT License and Copyright (c) 2021 Intel ISL (Intel Intelligent Systems Lab).

DeepBlending . We implement it according to the code at https://github.com/Phog/DeepBlending. This work is under Apache License.

PixelNeRF-DS. We use the number reported in .

Pytorch3D. We use the code from Pytorch3D [ravi2020pytorch3d]: https://github.com/facebookresearch/pytorch3d for our differentiable renderer. The code is licensed with BSD 3-Clause License.

User Study Details

As detailed in the paper, we conduct user study to evaluate perceptual quality of synthesized images. We provide more information about user study here.

We employ the standard A/B test paradigm as our user study format, which asks workers to select the closest result to a ground truth image between two competing results. Results of method A and B and ground truth target view are available during test. All views in the test set (690 views in total) are evaluated and each view is judged by 3 workers.

All tests are conducted using thehive.ai, a website similar to Amazon Mechanical Turk. Workers were given instructions and examples of tests, and then they were given a test to identify whether they understand the task. Only workers passing the test were allowed to conduct A/B test. Three images of the same view are shown in the test, where the first image is the ground truth target views and the rest two images are synthesized images. Workers are asked to select “left” or “right” image to indicate their preference. Results of method A and B are randomly placed for every test for fairness. We show the instructions and examples below.

Additional Results

Again, please see attached videos for comparison between ours and other methods.

We show times spent on each component of FWD-D model during a single forward pass in Table 6. As shown in the Table, only 30 percent and 12 percent of times are spent on the Rendering process and Fusion process, indicating that our renderer and Transformer are highly-efficient.

Moreover, we compare the time distributions of our method with PixelNeRF and IBRNet in Table 7.

2 Other ablations.

3 View number ablations.

4 Synthesis Results.

We show more synthesis results in the following. For comparison, we show results of the same scene and views. We first show the results of FWD-U in Figure 11, FWD in Figure 12 and FWD-D in Figure 13. We also show baseline results: PixelNeRF in Figure 14, IBRNet in Figure 15, MVSNeRF in Figure 16, FVS in Figure 17 and Blending-R in Figure 18. Again, please see attached videos for comparison between ours and other methods.