Worldsheet: Wrapping the World in a 3D Sheet for View Synthesis from a Single Image

Ronghang Hu, Nikhila Ravi, Alexander C. Berg, Deepak Pathak

Introduction

A 2D image is the projection of an underlying 3D world, but as humans, we have no trouble in understanding this structure and imagining how an image will look from other views. Consider the train shown in Figure 1, we can seamlessly predict other views from a single image based on the abstractions we have learned from past experience of seeing several trains, or similar shaped objects from different views. Enabling machines with such an ability to reason about 3D from a single image will bring trillions of still photos to life, with wide applications in virtual reality, animation, image editing, and robotics.

The goal of synthesizing novel views from 2D images has been pursued for decades, from early efforts relying completely on multi-view geometry , to more recent learning based approaches . Over the years, significant progress has been made in this direction. However, despite impressive photorealistic output renderings, most of these previous approaches require multiple images or ground-truth depth at test time, which severely hinders their practicality. To compensate for the lack of multiple views or 3D models at test time, methods for single-image 3D rely on statistical learning from data. This line of work can be traced back to classic works of Hoiem et al. , followed by Saxena et al. , that obtain ‘qualitative 3D’ from a single image by fitting a collection of planes onto the image.

An ideal approach to general-purpose view synthesis should not only rely on a single image at test time, but also learn from easy-to-collect supervision signal during training. In the deep learning era, there is growing interest in end-to-end methods with intermediate 3D representations supervised by multiple images and no explicit 3D information during training. However, they are mostly applied to objects , and are either category-specific, restricted to synthetic scenes, or both. Recent works address these issues by training with multiple views of real-world scenes, relying on point cloud or multiplane images as intermediate representations. However, multiplane images only perform well with relatively small viewpoint changes as each plane is at a constant depth; for point clouds, one needs to represent each point in a scene individually, making it inefficient to scale to high-resolution data or large viewpoint changes. In contrast, meshes can provide a sparser scene representation, e.g., two triangular mesh faces can theoretically represent the entire flat surface of a wall, making it ideal for single-image view synthesis. However, mesh recovery from single images has been studied mostly for object images and in a category-specific manner and not for scenes.

In this paper, we present an end-to-end approach for novel view synthesis from a single image of a scene via an intermediate mesh representation. Unlike mesh reconstruction for objects of specific categories, generating meshes for a scene is challenging as there is no notion of mean or canonical shape to start from, or silhouette from segmentation for supervision. We circumvent this problem by wrapping a deformable mesh sheet over the 3D world – much like wrapping a 2D tinfoil onto a 3D pan before baking! We name this shrink-wrapped mesh Worldsheet, a term borrowed from physics for the 2D manifold of high-dimensional strings. After generating this Worldsheet for a given view, novel views are obtained by moving the camera in 3D space (Figure 2), which allows us to train from just two views of a scene using only rendering losses without any 3D or depth supervision.

To train our model end-to-end, both reconstruction of the mesh texture from input view and rendering from a novel camera view need to be differentiable. The latter is easily handled thanks to recent differentiable mesh renderers . To address the former, we propose a differentiable texture sampler over projected 2D views, enabling gradient computation of the reconstructed texture map over the 3D mesh geometry. Furthermore, to better handle occlusions and depth discontinuities, we propose a simple extension by stacking multiple layers of Worldsheets onto the scene.

In summary, Worldsheet generates novel views by learning to predict scene geometry from a single image. Although 3D mesh reconstruction via differentiable rendering is common for objects, to our best knowledge, this is the first work to show mesh recovery for scenes just from multi-view supervision. Our model consistently outperforms prior state-of-the-art by a significant margin on three benchmark datasets (Matterport , Replica , and RealEstate10K ), and is applicable to very high-resolution images in-the-wild as shown in Figure 1.

Related work

Novel view synthesis from multiple images. Traditional novel view synthesis methods use multiple input views at test time , and are often based on different representations. Among recent works, Waechter et al. build scene meshes with diffuse appearance. StereoMag proposes multiplane images (MPIs) from a stereo image pair as a layered scene representation. NPBG captures the scene as a point cloud with neural descriptors. NeRF proposes a neural radiance field representation for scene appearance, and is followed by many extensions (see for a summary). NSVF adopts sparse voxel octrees as scene representations. FVS and SVS blend multiple source images based on a geometric scaffold. Yoon et al. combine depth from both single and multiple views to generate novel views of dynamic scenes. Access to multiple input views greatly simplifies the task, allowing the scene geometry to be recovered via multi-view stereo .

Novel view synthesis from a single image. In early works, Debevec et al. recover 3D scene models and Horry et al. fit a regular mesh to generate novel views. Liebowitz et al. and Criminsi et al. generate meshes via projective geometry constraints but these methods came at the expense of manual editing. Hoiem et al. generate automatic 3D pop-up by fitting vertical and ground planes onto the 2D image, unlike our mesh representation. More recently in , layered depth images are used for single image view synthesis based on a pre-trained depth estimator. In , online videos are used to train a scale-invariant MPI representation for view synthesis. SynSin synthesizes novel views from a single image with a feature point cloud. In contrast, we learn to construct scene meshes instead of point clouds and directly map image texture instead of feature vectors to generate novel views from large viewpoint changes.

Differentiable mesh rendering. Recent work on differentiable mesh renders allow learning 3D structures through synthesis. NMR and SoftRas reconstruct the 3D object shape as a mesh by rendering it, comparing it with the input image, and back-propagating losses to refine the mesh geometry. CMR , CSM and U-CMR build category-specific object meshes from images by deforming from a mean or template category shape through silhouette (and keypoints in ) supervision. Our method is aligned with the analysis-by-synthesis paradigm above. However, unlike most previous works that apply differentiable mesh rendering to objects, we learn the 3D geometry of scenes through the rendering losses on the novel view. Moreover, instead of predicting a texture flow as in , we propose to analytically sample the mesh texture from the input view with a differentiable texture sampler. Unlike , our differentiable texture sampler considers multiple mesh faces in the z-buffer (soft rasterization instead of only the closest one), and assumes perspective (instead of orthographic) camera projection.

Worldsheet: Rendering the World in a Sheet

In this work, we propose Worldsheet to synthesize novel views from a single image, as shown in Figure 2. Our model build a 3D scene mesh MM by warping a lattice grid (i.e. a “sheet”) onto the scene geometry, and is trained with only 2D rendering losses without any 3D or depth supervision.

From the input view image IinI_{in} of size Wim×HimW_{im}\times H_{im}, we build a scene mesh by warping a Wm×HmW_{m}\times H_{m} lattice grid (i.e. a sheet) onto the scene, as shown in Figure 2. We first extract a Wm×HmW_{m}\times H_{m} visual feature map {qw,h}\{q_{w,h}\} from IinI_{in} with a convolutional neural network. Each qw,hq_{w,h} is a feature vector at spatial location (w,h)(w,h) on the Wm×HmW_{m}\times H_{m} network output. In our implementation, we use ResNet-50 (pretrained on ImageNet) with dilation to output features {qw,h}\{q_{w,h}\}.

From each qw,hq_{w,h} on the feature map, we predict the grid offset Δx^w,h\Delta\hat{x}_{w,h} and Δy^w,h\Delta\hat{y}_{w,h} to decide how much the vertex (w,h)(w,h) on the grid should move away from its anchor positions within the image plane (we output Δx^w,h\Delta\hat{x}_{w,h} and Δy^w,h\Delta\hat{y}_{w,h} in NDC space between −1-1 to 11). We also predict how far each vertex is from the camera, i.e. its depth zw,hz_{w,h}. These values are predicted using learned mappings as

where division by (Wm−1)(W_{m}-1) and (Hm−1)(H_{m}-1) ensures that the vertices can only move within a certain range. g(⋅)g(\cdot) is a scalar nonlinear function to scale the network prediction into depth values. We use g(ψ)=αg/(σ(ψ)+ϵg)+βgg(\psi)=\alpha_{g}/(\sigma(\psi)+\epsilon_{g})+\beta_{g} in our implementation, where σ(⋅)\sigma(\cdot) is the sigmoid function and αg\alpha_{g}, βg\beta_{g} and ϵg\epsilon_{g} are fixed hyper-parameters.

Building the 3D scene mesh. We first build the mesh vertices {Vw,h}\{V_{w,h}\} from the grid offset and depth as

for w=1,⋯ ,Wmw=1,\cdots,W_{m} and h=1,⋯ ,Hmh=1,\cdots,H_{m}. Here θF\theta_{F} is the camera field-of-view, and x^w,h\hat{x}_{w,h} and y^w,h\hat{y}_{w,h} are anchor positions on the grid equally spaced from −1-1 to 11.

Then, we connect the mesh vertices {Vw,h}\{V_{w,h}\} along the edges on the grid to form mesh faces {F}\{F\} as shown in Figure 2 and obtain a 3D mesh M=({Vw,h},{F})M=(\{V_{w,h}\},\{F\}). A vertex in the mesh is connected to its 4 or 8 neighbours on the grid.

To encourage the mesh surface to be smooth unless it needs to bend to fit the scene geometry, we apply a Laplacian term Lm=∑w,h∥∑(wˉ,hˉ)∈N(w,h)(Vwˉ,hˉ−Vw,h)∥1L_{m}=\sum_{w,h}\left\|\sum_{(\bar{w},\bar{h})\in N(w,h)}\left(V_{\bar{w},\bar{h}}-V_{w,h}\right)\right\|_{1} on the mesh vertices, where N(w,h)N(w,h) are the adjacent vertices to (w,h)(w,h). In addition, we also apply an L2 regularization term Lg=∑w,h(Δx^w,h2+Δy^w,h2)L_{g}=\sum_{w,h}\left(\Delta\hat{x}_{w,h}^{2}+\Delta\hat{y}_{w,h}^{2}\right) to the grid offset.

2 Differentiable texture sampler

To render the input scene in another camera pose for novel view synthesis, we need to project image texture from the input view to the target view in a differentiable manner. While existing renderers can render an image from a scene mesh based on its texture map, they cannot directly transform image pixels in screen space between two different camera poses. In our model, we accomplish differentiable projection between two views by first reconstructing the scene mesh’s texture map from the input view (which involves inverting the texture-map-to-image perspective transform in a differentiable manner) so that it can be later rendered with the scene mesh in novel views using existing mesh renderers.

While a few approaches build a mesh texture map with a learned texture flow on objects, it is hard to apply the same to scenes, which do not have canonical shapes. Here, we take an alternative route and propose a differentiable texture sampler. We analytically sample the mesh texture T^\hat{T} as a UV texture map from the input view IinI_{in}, where gradients ∂T^/∂V\partial\hat{T}/\partial V and ∂T^/∂Iin\partial\hat{T}/\partial I_{in} over the vertex coordinates and the input image respectively can be computed.

To implement this texture sampler, we project the mesh faces onto the image plane to build a buffer (sorted in ascending z-order) containing the z values and 2D euclidean distance of points on the closest KK mesh faces whose projection overlaps image pixel pi,jp_{i,j} as in PyTorch3D . Then, we splat the RGB pixel intensities from the image IinI_{in} onto the UV texture map T^\hat{T}. Specifically, we first compute the weight wi,jkw_{i,j}^{k} denoting the contribution of the kk-th face color on pixel pi,jp_{i,j} based on the softmax blending formulation in . We then decompose the input image IinI_{in} into KK images IinkI^{k}_{in}, where Iink(i,j)=Iin⋅wi,jkI^{k}_{in}(i,j)=I_{in}\cdot w_{i,j}^{k}, and splat the RGB pixels from each IinkI^{k}_{in} to a texture map layer T^k\hat{T}^{k} as

In summary, the image pixels are splatted onto the texture space via each rasterized mesh face, and blended together to obtain the final texture map. The entire process is differentiable with respect to both IinI_{in} and the mesh vertex coordinates {V}\{V\}, as one can analytically compute ∂T^k/∂Iink\partial\hat{T}^{k}/\partial I^{k}_{in}, ∂T^k/∂fk\partial\hat{T}^{k}/\partial f^{k}, ∂fk/∂V\partial f^{k}/\partial V, and ∂wi,jk/∂V\partial w_{i,j}^{k}/\partial V.

3 Learning scene geometry by view synthesis

To synthesize a novel view, we project the mesh vertex coordinates {V}\{V\} from the input camera pose θin\theta_{in} to {Vtgt}\{V^{tgt}\} in the camera coordinate space of the target viewpoint θtgt\theta_{tgt}. Then, we render the mesh Mtgt=({Vtgt},{F})M^{tgt}=(\{V^{tgt}\},\{F\}) in the target camera pose along with its texture map T^\hat{T} to output a 2D image IoutI_{out} of size Wim×HimW_{im}\times H_{im} as the target view:

We use the differentiable mesh renderer in so that we can compute the gradients ∂Iout/∂Vtgt\partial I_{out}/\partial V^{tgt} and ∂Iout/∂T^\partial I_{out}/\partial\hat{T}. Through mesh rendering, we also obtain a foreground mask FoutF_{out} with the same size as IoutI_{out}, indicating which pixels in the rendered image IoutI_{out} are covered by the mesh and which pixels are from background color, as shown by the grey area in Figure 3 (b).

Our model is supervised with paired input and target views of a scene (along with their camera poses). We use a pixel L1 loss Loutrgb=∥Iout−Itgt∥1/(Wim⋅Him)L^{rgb}_{out}=\|I_{out}-I_{tgt}\|_{1}/(W_{im}\cdot H_{im}) and a perceptual loss Loutpc=P(Iout,Itgt)L^{pc}_{out}=P(I_{out},I_{tgt}), where ItgtI_{tgt} is the ground-truth target view image. The model then needs just a single image at test time.

4 Inpainting and image refinement

The target view image consists of two parts: things that can be directly seen from the input view IinI_{in}, and things that need to be imagined based on our prior knowledge of the visual world, as illustrated in Figure 3. As our mesh warping and rendering procedure in Sec. 3.1, 3.2 and 3.3 builds a pixel-to-pixel correspondence between the input and the target view, it only renders pixels that are visible from the input view. To obtain a plausible imagination of the invisible image regions, we apply an inpainting network GG on the rendered mesh IoutI_{out} to fill the missing regions and output a new image Ipaint=G(Iout)I_{paint}=G(I_{out}) as the final target view.

We build our inpainting network based on the generator in pix2pixHD , which translates a 4-channel input (the rendered image IoutI_{out} and its foreground mask FoutF_{out}) into a 3-channel output image IpaintI_{paint}. Our inpainting network outputs an entire image – it not only fills the invisible regions but also refines the image details in the visible regions. We apply the same RGB pixel L1 loss LpaintrgbL^{rgb}_{paint} and perceptual loss LpaintpcL^{pc}_{paint} as in Sec. 3.3 on the inpainting output IpaintI_{paint}.

Training. We train our model using the Adam optimizer with a weighted combination of losses as L=λ1Loutrgb+λ2Loutpc+λ3Lpaintrgb+λ4Lpaintpc+λ5Lg+λ6LmL=\lambda_{1}L^{rgb}_{out}+\lambda_{2}L^{pc}_{out}+\lambda_{3}L^{rgb}_{paint}+\lambda_{4}L^{pc}_{paint}+\lambda_{5}L_{g}+\lambda_{6}L_{m} with λ1=λ3=8\lambda_{1}=\lambda_{3}=8, λ2=λ4=2\lambda_{2}=\lambda_{4}=2, λ5=0.2\lambda_{5}=0.2, and λ6=10−4\lambda_{6}=10^{-4}. Our model is trained for a total of 50000 iterations with batch size 6464 and 10−410^{-4} learning rate.

We use a grid mesh with size Wm×Hm=33×33W_{m}\times H_{m}=33\times 33 (and also 65×6565\times 65 in Sec. 4.2). Following SynSin , we use Wim×Him=256×256W_{im}\times H_{im}=256\times 256 as the input and output image size. Our mesh implementation is based on PyTorch3D .

5 Extension: multi-layered Worldsheets

Although shrink-wrapping a single mesh sheet onto images works well on a wide range of scenes, one limitation is that it assumes that the foreground objects are connected to the background by mesh faces, which sometimes causes artifacts near object boundaries or depth discontinuities.

We propose an extension to address this limitation: predicting and warping multiple layers of Worldsheet onto the scene, where each sheet has a transparency channel in its texture map, loosely inspired by layered-depth images . This allows some layers to fit the foreground object and others to capture the background. Specifically, we predict grid offset and depth for each mesh sheet from the feature map {qw,h}\{q_{w,h}\} following Eqn. 1 to 3 with separate parameters. We also predict an Him×WimH_{im}\times W_{im} alpha map for each sheet using a deconvolution layer on {qw,h}\{q_{w,h}\}, which is then projected to the transparency channel in the UV texture map of the associated sheet. Finally, the multiple mesh sheets are rendered in the novel view using alpha compositing . The whole model can be trained end-to-end under the same supervision. In Sec. 4.4, we find that, qualitatively, this extension leads to better handling of occlusions and parallax effect than a single mesh sheet.

Experiments

We evaluate our model on three datasets: Matterport , Replica , and RealEstate10K , following the experimental setup and details from . We then provide analysis on in-the-wild images and multi-layered sheets.

We first train and evaluate our approach on the Matterport dataset , which contains 3D scans of homes. We load the Matterport dataset in the Habitat simulator , following the same training, validation, and test splits as in SynSin . During training, we supervise our model with paired 2D images of the input and the target views. We empirically find that it works slightly better to first train the scene mesh predictor (Sec. 3.1) and then freeze the scene mesh to further train the inpainting network (Sec. 3.4), rather than training both components jointly from scratch.

Metrics. Following SynSin , we evaluate the predicted novel view images IpaintI_{paint} using three metrics: Peak Signal-to-Noise Ratio (PSNR; higher is better), Structural Similarity (SSIM; higher is better), and Perceptual Similarity distance (Perc Sim; lower is better). The Perc Sim metric is based on the convolutional feature distance between the prediction and the ground-truth, which is shown to be highly correlated with human judgement . Since only a part of the target view image can be seen from the input image as illustrated in Figure 3, we separately evaluate these metrics on visible regions (Vis, which can be seen from the input view), invisible regions (InVis, which cannot be seen and must be imagined), and the entire image (Both). Note that the visible region masks are obtained from the ground-truth scene geometry and camera frustum (available from the Habitat simulator) instead of predicted by our mesh, and are the same as in SynSin’s evaluation.

Baselines. We compare our method to several previous approaches: Im2Im is an image-to-image translation method which predicts an appearance flow to warp an input view to the target view based on an input camera transformation. Tatarchenko et al. is similar to Im2Im, but directly predicts the target view image instead of an appearance flow. Vox w/ UNet and Vox w/ ResNet are two variants of the deep voxel representation with different encoder-decoder architectures based on UNet, or ResNet as implemented in . SynSin projects a dense feature point cloud (extracted from every image pixel) to the target camera pose and applies a refinement network on the point cloud projection to output the target view image. We also evaluate the prediction of our model before inpainting (i.e. directly using the mesh rendering output IoutI_{out} as the target view) to analyze how well our method performs with texture sampling and mesh rendering alone.

Results. The results are shown in Table 1. Even without inpainting, the mesh rendering output IoutI_{out} from our method already outperforms previous approaches by a large margin under all the three metrics on the visible regions. With the help of an inpainting network, our final output IpaintI_{paint} has significantly higher performance than previous work on both invisible and visible regions, achieving a new state-of-the-art performance on this dataset. Figure 4 shows view synthesis examples from our method and previous work on the Matterport dataset, where our method can paint things such as doorframe or sofa at more precise locations.

Generalization to the Replica dataset. Following , we also evaluate how well our model generalizes to another scene dataset, Replica , which contains high-quality laser scans of both homes and offices. We take our model trained on the Matterport dataset and directly evaluate on the Replica dataset without re-training. The results are shown in Table 1, where all methods are trained and evaluated under the same setting. It can be seen that our method achieves noticeably better generalization to this dataset and outperforms previous approaches by a large margin. Figure 5 shows view synthesis examples on the Replica dataset.

Generalization to larger viewpoint changes. We further analyze how well our and previous approaches generalize to larger camera pose changes beyond their training data. In this analysis, we sample new input-target view pairs on the test scenes with 2×2\times larger camera angle changes than in the training data, and directly evaluate all approaches on these new viewpoints without retraining. The results are shown in Table 2, where our method largely outperforms other approaches under all metrics. Figure 4 (second row) shows an example under 2×2\times larger camera angle change.

2 Evaluation on RealEstate10K

The RealEstate10K dataset consists of both indoor and outdoor scenes extracted from YouTube videos of houses. The input view and the target view are different video frames within a time range, with camera poses estimated using structure-from-motion.

On this dataset, we follow the experimental setup in SynSin and use the same training, validation, and test data. In addition to using a 33×3333\times 33 mesh, we also train our model with a higher resolution Wm×Hm=65×65W_{m}\times H_{m}=65\times 65 mesh, which is initialized from a trained 33×3333\times 33 mesh model with a new transposed convolution layer to upsample the feature map {qw,h}\{q_{w,h}\} in Sec. 3.1 to 65×6565\times 65 spatial dimensions.

We compare our method to several previous approaches. In addition to the baselines in Sec. 4.1, we also compared to three additional approaches. 3DView is a system similar to the Facebook 3D Photo based on layered depth images and is also a baseline in . Single-View MPI and StereoMag both use multiplane images (MPIs), where Single-View MPI builds MPIs from a single input image while StereoMag relies on a stereo pair using images from two different views as input at test time. Except for StereoMag, all other methods use a single view at test time.

Results. We follow the evaluation protocol of on RealEstate10K, with resultsTo compare with StereoMag that uses two input views, in the evaluation protocol of SynSin on RealEstate10K, the best metrics of two separate predictions based on each view were reported for single-view methods. We follow this evaluation protocol for consistency with on RealEstate10K in Table 3 and 4. We also report averaged metrics over all predictions in supplemental, where the trends are consistent. shown in Table 3. It can be seen that our method achieves the highest performance, outperforming previous approaches by a noticeable margin. Besides, a higher resolution 65×6565\times 65 mesh gives a further performance boost. Figure 6 shows predicted novel views on this dataset. In addition, we visualize the pixel-wise squared error map on the prediction from our method and SynSin in Figure 7, where our method paints objects at more precise locations compared to SynSin, resulting in higher PSNR and better quality.

Ablations. Since our model relies on deforming a mesh sheet, we first analyze the impact of geometric regularization on the mesh deformation. In Table 4, line 2 to 4 vary the weight of the mesh Laplacian regularization term LmL_{m} from its default value 10−410^{-4}. Comparing these variants to line 1, a higher regularization (10−310^{-3}, line 4) restricts the model’s capacity to precisely fit the scene geometry and hence hurts the performance. Meanwhile, there is only a smaller drop when decreasing this regularization weight to 10−510^{-5} or even zero, suggesting that our differentiable rendering pipeline provides robust wrapping of the mesh onto the scene.

We further study the impact of mesh resolution Wm×HmW_{m}\times H_{m} in line 5 to 8. As expected, higher mesh resolution allows fitting more fine-grained scene details and gives higher view synthesis performance, with the final 65×6565\times 65 mesh giving the best performance. In addition, we find that with enough mesh resolution (such as 65×6565\times 65), one can restrict the grid offset in Eqn. 1 and 2 to zero and only use the predicted depth in Eqn. 3 to deform the mesh, which gives only −0.13-0.13 PSNR drop (however, the grid offset makes a larger difference in lower-resolution meshes such as Figure 2).

3 Analysis: testing the limits of wrapping sheets

So far, we have shown that the idea of wrapping a mesh sheet onto an image achieves strong performance across all benchmarks. But one might wonder, how good is a planar sheet prior for novel view synthesis in any arbitrary images? To test the limits of wrapping a mesh sheet, we test it over a large variety of images including outdoor scenes, outdoor objects, indoor scenes, indoor objects, and even artistic paintings. We analyze our underlying mesh data structure by pretraining depth to fill in the zz values, and examine how well it generates novel views in-the-wild. As shown in Figure 1 (top row), although missing a few details (such as tree branches), a scene mesh sheet captures the geometric structures sufficient enough to render high-resolution (960×\times960) photorealistic novel views even from very large viewpoint changes. Please see videos at worldsheet.github.io for animation of continuously generated views. This result confirms that our mesh sheet data structure and warping procedure, despite being simple, are flexible enough to handle the variety of the visual world.

4 Analysis: multi-layered Worldsheets

As described in Sec. 3.5, to better handle sharp depth discontinuities, we explore the extension to stack multiple mesh sheet layers, so that foreground objects and the background can be placed on different layers. Interestingly, we observe that this extension does not make a noticeable improvement in RealEstate10K evaluation metrics compared to a single sheet (26.62 vs. 26.74 in PSNR) suggesting that single Worldsheet is sufficient enough for view synthesis application. However, qualitatively speaking, we notice that it allows better handling of occlusions and parallax effect in our model under large viewpoint changes, as shown in Figure 8. Please see supplemental for more details.

Conclusion

In this work, we propose Worldsheet, which synthesizes novel views from a single image by shrink-wrapping the scene with a grid mesh. Our approach jointly learns the scene geometry and generates novel views through differentiable texture sampling and mesh rendering, supervised with only 2D images of the input and the target views. The approach is category-agnostic and end-to-end trainable, resulting in state-of-the-art performance on single-image view synthesis across three datasets by a large margin. Acknowledgments. We are grateful to Alyosha Efros, Angjoo Kanazawa, Shubham Goel, Devi Parikh, Ross Girshick, Georgia Gkioxari, Justin Johnson, Brian Okorn and other colleagues at FAIR and CMU for fruitful discussions. This work was supported in part by DARPA Machine Common Sense grant (associated with D. Pathak and not associated with Facebook Inc).

References

Appendix A Continuous and large viewpoint changes

Our approach allows synthesizing continuous novel views by smoothly moving to a new camera pose that is largely different from the input. We kindly request the readers to view the videos at worldsheet.github.io to better understand the performance of our method. In these videos, we compare our synthesized novel views (from a single image) to SynSin on the RealEstate10K dataset with simulated large view-point changes (the first frame contains the input view and the rest of the frames are synthesized). From the videos, it can be seen that our model can generate novel views with much larger camera translation and rotation than in the training data, while SynSin often suffers from severe artifacts in these cases, likely because its refinement network does not generalize well to a sparser point cloud (resulting from large camera zoom-in or rotation).

We also show in these videos continuously synthesized novel views on high resolution (960×960960\times 960) images over a wide range of scenes (the first frame is the input view), as described in our analysis in Sec. 4.3 in the main paper.

Appendix B Ablation study: using depth supervision

In our experiments in the paper, we show that our model can be trained using only two views of a scene without 3D or depth supervision. In this section, we further analyze our approach by training it with depth supervision on the Matterport dataset, where the ground-truth depth can be obtained from the Habitat simulator.

In this analysis, we modify the differentiable mesh renderer to render RGB-D images from our mesh, and apply an L1 loss between the ground-truth and the rendered depth as additional supervision. We also compare with the performance of SynSin with depth supervision (reported in ). The results are shown in Table B.1. It can be seen that our model without depth supervision (the default setting; line 7) works almost equally as well as its counterpart using depth supervision (line 8) on the Matterport dataset and generalizes better to the Replica dataset. In addition, it outperforms SynSin under both supervision settings (lines 5-6).

Appendix C Additional analyses on RealEstate10K

As described in Sec. 4.2 in the main paper, we follow the evaluation protocol of SynSin on the RealEstate10K dataset. To enable comparison with StereoMag that uses two input views on this dataset, in , the best metrics of two views were reported for single-view methods. At test time, for each target view, this involves making two separate predictions based on two different input views respectively, and then selecting the best metrics between the two predictions as the score for this target view. Note that this evaluation protocol is only applied to the RealEstate10K dataset (in Table 3 and 4 in the main paper) and is not applied to Matterport or Replica.

In this section, we further evaluate by taking the average metrics over all predictions to measure how well the model does on average from a single input view and to be consistent with our evaluation on Matterport and Replica in Table 1 and 2 in the main paper. Apart from metrics over the entire image, we would also like to analyze how well each model does on rendering regions seen in the input view vs. invisible regions (where things must be imagined). However, we cannot compute the exact visibility map as there are no ground-truth geometry annotations in the RealEstate10K dataset. To get an approximation, we evaluate on the central Wim2×Him2=128×128\frac{W_{im}}{2}\times\frac{H_{im}}{2}=128\times 128 crop of the target image (Center, which is nearly always visible) and the rest of the image (Peripheral, containing most of the invisible regions).

The results of these analyses are shown in Table C.1. It can be seen that our method achieves the highest performance on both the center regions (which are mostly visible) and the peripheral regions, outperforms the previous single-view based approaches by a large margin.

Our model uses a simple ResNet-50 backbone with an output stride of 8 pixels to extract image features (see Sec. 3.1 in the main paper), while SynSin adopts a U-Net backbone that has a higher output feature resolution same as the input image (i.e. output feature stride is 1 pixel). To further study the impact of different backbone architectures, we train a variant of SynSin by replacing its U-Net backbone with the same ResNet-50 backbone pretrained on ImageNet (and upsampling its output feature map to stride 1 with a deconvolution layer) to be consistent with our model, shown in line 4 in Table C.1. Comparing it with line 3 or 6, it can be seen that this variant of SynSin with ResNet-50 backbone performs worse than the default SynSin architecture, or our model. This suggests that SynSin requires a high-resolution feature output to build a per-pixel point-cloud for view synthesis, while our model is able to work with a lower resolution (a larger stride) in the backbone.

Appendix D Details on differentiable texture sampler

Our differentiable texture sampler (Sec. 3.2 in the main paper) splats image pixels onto the texture map through each face. This splatting procedure involves three main steps: forward-mapping, normalization, and hole filling.

In the implementation above, gradients can be taken over ff through the bilinear weights.

However, forward mapping alone will lead to incorrect pixel intensity (e.g. imagine down-scaling an image to half its width and height by forward-mapping – each pixel in the low-resolution image will receive assignment from 4 pixels and become 4×4\times brighter). Hence, a second normalization step is applied:

where IoneI_{one} is an Wim×HimW_{im}\times H_{im} image with all ones as its pixel intensity. T^norm\hat{T}_{norm} contains the normalized splatting result. A threshold 10−410^{-4} is applied to avoid division by zero (which could happen due to holes described below).

Filling holes with Gaussian filtering. It is well known that the bilinear forward mapping above often leads to holes in the output (e.g. imagine up-scaling an image to a much larger size – there will be gaps in the output image as some pixels will not receive assignments). To minimize hole occurrence, in our UV texture map we assign the UV coordinates of each mesh vertex with an equally-spaced Wm×HmW_{m}\times H_{m} lattice grid, and use the same image size as the texture map size (Wuv×Huv=Wim×HimW_{uv}\times H_{uv}=W_{im}\times H_{im}). This ensures that most texels on the texture map receive assignments in forward mapping, so that holes rarely occur in T^norm\hat{T}_{norm}. However, to address corner cases, we further apply a Gaussian filter to fill the holes in T^norm\hat{T}_{norm} (where W^sum\hat{W}_{sum} is zero as no assignment is received from forward-mapping):

where FgF_{g} is a discrete 2D Gaussian kernel for image filtering (we use kernel size 7 and standard deviation 2 for FgF_{g} in our implementation). Here M^\hat{M} is a binary mask indicating which pixels have received assignments in forward mapping (i.e. 1 means valid and 0 means holes), T^g\hat{T}_{g} is the Gaussian-blurred version of T^norm\hat{T}_{norm} (where the division ensures the correct pixel intensity; otherwise it will be darker due to holes in T^norm\hat{T}_{norm}) and is used to fill only the holes in T^norm\hat{T}_{norm}. We use T^\hat{T} in Eqn. D.6 as the final splatting output.

Perspective correctness. A main purpose of our differentiable texture sampler is to build perspective-correct novel views during texture reconstruction. We note that the alternative solution of directly using the input image as a texture map by putting vertex uvuv texture coordinates in the input screen space for mesh rendering breaks perspective correctness, as shown in Figure D.1 (d). For perspective-correct novel views, one needs to invert the texture-map-to-image perspective transform when building UV texture maps from the image, which we implement in our texture sampler shown in Figure D.1 (c).

Appendix E Details on multi-layered Worldsheets

In our proposed Worldsheet model, we build a scene mesh by warping a planar sheet onto the scene. This model is capable of handling moderate occlusion and generating plausible novel views by deforming the mesh along object boundaries and refining the predicted novel view with an inpainting network. However, we also acknowledge that artifacts can sometimes occur in occluded regions or object boundaries when the disparity is very large between the input view and the novel view. This is partly because our current approach of deforming a mesh onto the scene does not capture all the fine-grained geometric details (such as the flower boundary as shown in the last two failure cases in the supplemental videos. We believe there is room for improvement in this direction (e.g. via adaptive resolution), which we are interested in exploring in future work.

In Sec. 3.5 in the main paper, we propose a simple extension with multi-layered Worldsheets. The main purpose of this extension is to separate objects or scene structures at different depth levels into different mesh layers, instead of placing them all on a single sheet. In this extension, the 3D mesh geometry of each layer provides the geometric support for the scene components, while the transparency channel in the RGBA texture maps allows segmentation between different components. For example, to represent a sofa object, a mesh layer can be wrapped onto a larger 3D surface region covering the sofa surface, with its texture map containing the sofa texture over the object region while being transparent on the surrounding regions.

Specifically, we predict and warp a total of LL mesh sheets (i.e. LL layers) onto the scene for view synthesis. For each layer l=1,⋯ ,Ll=1,\cdots,L, we predict its grid offset (Δx^w,h(l),Δy^w,h(l))\left(\Delta\hat{x}_{w,h}^{(l)},\Delta\hat{y}_{w,h}^{(l)}\right) and its depth zw,h(l)z_{w,h}^{(l)} from the convolutional feature map {qw,h}\{q_{w,h}\} similar to Sec. 3.1, and also predict a pixel-wise transparency map α(l)\alpha^{(l)} of size Him×WimH_{im}\times W_{im} in the screen space of the input view as follows.

For each layer l=1,⋯ ,Ll=1,\cdots,L, we construct a corresponding 3D mesh sheet M(l)M^{(l)} following the procedure in Sec. 3.1 and also build its UV texture map T^(l)\hat{T}^{(l)} consisting of RGBA channels by concatenating the predicted transparency values αi,j(l)\alpha_{i,j}^{(l)} with the input image and splatting them onto the mesh texture space using our differentiable texture sampler in Sec. 3.2. Finally, we render all the mesh faces from all LL layers M(1),⋯ ,M(L)M^{(1)},\cdots,M^{(L)} in the novel view along with their RGBA UV texture maps T^(1),⋯ ,T^(L)\hat{T}^{(1)},\cdots,\hat{T}^{(L)} through alpha compositing. The whole model can be trained end-to-end under the same supervision using only 2D rendering losses.

We use a total of L=3L=3 layers in our analyses. Figure E.1 visualizes this extension on multi-layered Worldsheets. It can be seen that through end-to-end training, the model learns to place scene structures at different depth levels onto different mesh layers and separate foreground objects (e.g. kitchen counter, sofa, or table) from their background. We qualitatively find that it better handles occlusions and parallax effect under large viewpoint changes, as shown in Sec. 4.4 in the main paper.

Appendix F Hyper-parameters in our model

In our nonlinear function g(⋅)g(\cdot) to scale the network prediction into depth values (in Eqn. 3 in the main paper), we use different output scales based on the depth range in each dataset. On Matterport and Replica, we use

On RealEstate10K (which has larger depth range), we double the output depth scale and use

However, we find that the performance of our model is quite insensitive to the hyper-parameters in g(⋅)g(\cdot).

In our differentiable texture sampler and the mesh renderer, we mostly follow the hyper-parameters in PyTorch3D . We use K=10K=10 faces per pixel and 1e-8 blur radius in mesh rasterization, 1e-4 sigma and 1e-4 gamma in softmax RGB blending, and background color filled with the mean RGB intensity on each dataset. On Matterport and Replica, the input views have 90-degree field-of-view. On RealEstate10K, we multiply the actual camera intrinsic matrix of each frame into its camera extrinsic RR and TT matrices, so that we can still use the same intrinsics and 90-degree field-of-view in the renderer. On high resolution images in the wild (Sec. 4.3 in the main paper), we assume 45-degree field-of-view.

We choose our mesh size WmW_{m} and HmH_{m} based on the image size. In our experiments on Matterport and Replica (Sec. 4.1 in the main paper), we use 256×256256\times 256 input image resolution following SynSin , and use pixel stride 8 on the lattice grid sheet (from which our mesh is built), resulting in Wm=Hm=1+256/8=33W_{m}=H_{m}=1+256/8=33. On the RealEstate10K dataset (Sec. 4.2 in the main paper), we additionally experiment with pixel stride 4 on the grid sheet, giving Wm=Hm=1+256/4=65W_{m}=H_{m}=1+256/4=65. In our analysis on high resolution images in the wild (resized to have the image long side equal to 960960 and padded to 960×960960\times 960 square size for ease of rendering in PyTorch3D; Sec. 4.3 in the main paper), we use Wm×Hm=129×129W_{m}\times H_{m}=129\times 129 mesh on the actual image regions (not including the padding regions).

Appendix G More visualized examples

Figure G.1 shows the depth maps from our scene mesh, where most of the scene structure is captured, giving coherent novel view projections.

Figure G.2 shows additional visualization and error map comparisons between our approach and SynSin on the RealEstate10K dataset (similar to Figure 7 in the main paper), where our method paints things in the novel view at more precise locations with lower error and higher PSNR.