Neural Rerendering in the Wild

Moustafa Meshry, Dan B Goldman, Sameh Khamis, Hugues Hoppe, Rohit Pandey, Noah Snavely, Ricardo Martin-Brualla

Introduction

Imagine spending a day sightseeing in Rome in a fully realistic interactive experience without ever stepping on a plane. One could visit the Pantheon in the morning, enjoy the sunset overlooking the Colosseum, and fight through the crowds to admire the Trevi Fountain at night time. Realizing this goal involves capturing the complete appearance space of a scene, i.e., recording a scene under all possible lighting conditions and transient states in which the scene might be observed—be it crowded, rainy, snowy, sunrise, spotlit, etc.—and then being able to summon up any viewpoint of the scene under any such condition. We call this ambitious vision total scene capture. It is extremely challenging due to the sheer diversity of appearance—scenes can look dramatically different under night illumination, during special events, or in extreme weather.

In this paper, we focus on capturing tourist landmarks around the world using publicly available community photos as the sole input, i.e., photos in the wild. Recent advances in 3D reconstruction can generate impressive 3D models from such photo collections , but the renderings produced from the resulting point clouds or meshes lack the realism and diversity of real-world images. Alternatively, one could use webcam footage to record a scene at regular intervals but without viewpoint diversity, or use specialized acquisition (e.g., Google Street View, aerial, or satellite images) to snapshot the environment over a short time window but without appearance diversity. In contrast, community photos offer an abundant (but challenging) sampling of appearances of a scene over many years.

Our approach to total scene capture has two main components: (1) creating a factored representation of the input images, which separates viewpoint, appearance conditions, and transient objects such as pedestrians, and (2) rendering realistic images from this factored representation. Unlike recent approaches that extract implicit disentangled representations of viewpoint and content , we employ state-of-the-art reconstruction methods to create an explicit intermediate 3D representation, in the form of a dense but noisy point cloud, and use this 3D representation as a “scaffolding” to predict images.

An explicit 3D representation lets us cast the rendering problem as a multimodal image translation . The input is a deferred-shading framebuffer in which each rendered pixel stores albedo, depth, and other attributes, and the outputs are realistic views under different appearances. We train the model by generating paired datasets, using the recovered viewpoint parameters of each input image to render a deep buffer of the scene from the same view, i.e., with pixelwise alignment. Our model effectively learns to take an approximate initial scene rendering and rerender a realistic image. This is similar to recent neural rerendering frameworks but using uncontrolled internet images rather than carefully captured footage.

We explore a novel strategy to train the multimodal image translation model. Rather than jointly estimating an embedding space for the appearance together with the rendering network , our system performs staged training of both. First, an appearance encoding network is pretrained using a proxy style-based loss , an efficient way to capture the style of an image. Then, the rerendering network is trained with fixed appearance embeddings from the pretrained encoder. Finally, both the appearance encoding and rerendering networks are jointly finetuned. This simple yet effective strategy lets us train simpler networks on large datasets. We demonstrate experimentally how a model trained in this fashion better captures scene appearance.

Our system is a first step towards addressing total scene capture and focuses primarily on the static parts of scenes. Transient objects (e.g., pedestrians and cars) are handled by conditioning the rerendering network on the expected semantic labeling of the output image, so that the network can learn to ignore these objects rather than trying to hallucinate their locations. This semantic labeling is also effective at discarding small or thin scene features (e.g., lampposts) whose geometry cannot be robustly reconstructed, yet are easily identified using image segmentation methods. Conditioning our network on a semantic mask also enables the rendering of scenes free of people if desired. Code will be available at https://bit.ly/2UzYlWj.

A first step towards total scene capture, i.e., recording and rerendering a scene under any appearance from in-the-wild photo collections.

A factorization of input images into viewpoint, appearance, and semantic labeling, conditioned on an approximate 3D scene proxy, from which we can rerender realistic views under varying appearance.

A more effective method to learn the appearance latent space by pretraining the appearance embedding network using a proxy loss.

Compelling results including view and appearance interpolation on five large datasets, and direct comparisons to previous methods .

Related work

Traditional methods for scene reconstruction first generate a sparse reconstruction using large-scale structure-from-motion , then perform Multi-View Stereo (MVS) or variational optimization to reconstruct dense scene models. However, most such techniques assume a single appearance, or else simply recover an average appearance of the scene. We build upon these techniques, using dense point clouds recovered from MVS as proxy geometry for neural rerendering.

In image-based rendering , input images are used to generate new viewpoints by warping input pixels into the outputs using proxy geometry. Recently, Hedman et al. introduce a neural network to compute blending weights for view-dependent texture mapping that reduces artifacts in poorly reconstructed regions. However, image-based rendering generally assumes the captured scene has static appearance, so it is not well-suited to our problem setup in which the appearance varies across images.

Neural scene rendering applies deep neural networks to learn a latent scene representation that allows generation of novel views, but is limited to simple synthetic geometry.

Appearance modeling

A given scene can have dramatically different appearances at different times of day, in different weather conditions, and can also change over the years. Garg et al. observe that for a given viewpoint, the dimensionality of scene appearance as captured by internet photos is relatively low, with the exception of outliers like transient objects. One can recover illumination models for a photo collection by estimating albedo using cloudy images , retrieving the sun’s location through timestamps and geolocation , estimating coherent albedos across the collection , or assuming a fixed viewpoint . However, these methods assume simple lighting models that do not apply to nighttime scene appearance. Radenovic et al. recover independent day and night reconstructions, but do not enable smooth appearance interpolations between the two.

Laffont et al. assign transient attributes like “fall” or “sunny” to each image, and learn a database of patches that allows for editing such attributes. Other works require direct supervision from lighting models estimated using 360-degree images , or ground truth object geometry . In contrast, we use a data-driven implicit representation of appearance that is learned from the input image distribution and does not require direct supervision.

Deep image synthesis

The seminal work of pix2pix trains a deep neural network to translate an image from one domain, such as a semantic labeling, into another domain, such as a realistic image, using paired training data. Image-to-image (I2I) translation has since been applied to many tasks . Several works propose improvements to stabilize training and allow for high-quality image synthesis . Others extend the I2I framework to unpaired settings where images from two domains are not in correspondence , multimodal outputs where an input image can map to multiple images , or unpaired datasets with multimodal outputs where an image in one domain is converted to another domain while preserving the content .

Image translation techniques can be used to rerender scenes in a more realistic domain, to enable facial expression synthesis , to fix artifacts in captured 3D performances , or to add viewpoint-dependent effects . In our paper, we demonstrate an approach for training a neural rerendering framework in the wild, i.e., with uncontrolled data instead of captures under constant lighting conditions. We cast this as a multimodal image synthesis problem, where a given viewpoint can be rendered under multiple appearances using a latent appearance vector, and with editable semantics by conditioning the output on the desired semantic labeling of the output.

Total scene capture

We define the problem of total scene capture as creating a generative model for all images of a given scene. We would like such a model to:

encode the 3D structure of the scene, enabling rendering from an arbitrary viewpoint,

capture all possible appearances of the scene, e.g., all lighting and weather conditions, and allow rendering the scene under any of them, and

understand the location and appearance of transient objects in the scene, e.g., pedestrians and cars, and allow for reproducing or omitting them.

Although these goals are ambitious, we show that one can create such a generative model given sufficient images of a scene, such as those obtained for popular tourist landmarks.

We first describe a neural rerendering framework that we adapt from previous work in controlled capture settings to the more challenging setting of unstructured photo collections (Section 3.1). We extend this model to enable appearance capture and multimodal generation of renderings under different appearances (Section 3.2). We further extend the model to handle transient objects in the training data by conditioning its inputs on a semantic labeling of the ground truth images (Section 3.3).

We adapt recent neural rerendering frameworks to work with unstructured photo collections. Given a large internet photo collection {Ii}\{I_{i}\} of a scene, we first generate a proxy 3D reconstruction using COLMAP , which applies Structure-from-Motion (SfM) and Multi-View Stereo (MVS) to create a dense colored point cloud.

An alternative to a point cloud is to generate a textured mesh . Although meshes generate more complete renderings, they tend to also contain pieces of misregistered floating geometry which can occlude large regions of the scene . As we show later, our neural rerendering framework can produce highly realistic images given only point-based renderings as input.

Given the proxy 3D reconstruction, we generate an aligned dataset of rendered images and real images by rendering the 3D point cloud from the viewpoint viv_{i} of each input image IiI_{i}, where viv_{i} consists of camera intrinsics and extrinsics recovered via SfM. We generate a deferred-shading deep buffer BiB_{i} for each image , which may contain per-pixel albedo, normal, depth and any other derivative information. In our case, we only use albedo and depth and render the point cloud by using point splatting with a z-buffer with a radius of 1 pixel.

However, the image-to-image translation paradigm used in is not appropriate for our use case, as it assumes a one-to-one mapping between inputs and outputs. A scene observed from a particular viewpoint can look very different depending on weather, lighting conditions, color balance, post processing filters, etc. In addition, a one-to-one mapping fails to explain transient objects in the scene, such as pedestrians or cars, whose location and individual appearance is impossible to predict from the static scene geometry alone. Interestingly, if one trains a sufficiently large neural network on this simple task on a dataset, the network learns to (1) associate viewpoint with appearance via memorization and (2) hallucinate the location of transient objects, as shown in Figure 2.

2 Appearance modeling

To capture the one-to-many relationship between input viewpoints (represented by their deep buffers BiB_{i}) and output images IiI_{i} under different appearances, we cast the rerendering task as multimodal image translation . In such a formulation, the goal is to learn a latent appearance vector ziaz^{a}_{i} that captures variations in the output domain IiI_{i} that cannot be inferred from the input domain BiB_{i}. We compute the latent appearance vector as zia=Ea(Ii,Bi)z^{a}_{i}=E^{a}(I_{i},B_{i}) where EaE^{a} is an appearance encoder that takes as input both the output image IiI_{i} and the deep buffer BiB_{i}. We argue that having the appearance encoder EaE^{a} observe the input BiB_{i} allows it to learn more complex appearance models by correlating the lighting in IiI_{i} with scene geometry in BiB_{i}. Finally, a rerendering network RR generates a scene rendering conditioned on both viewpoint BiB_{i} and the latent appearance vector zaz^{a}. Figure 3 shows an overview of the overall process.

To train the appearance encoder EaE^{a} and rendering network RR, we first adopted elements from recent methods in multimodal synthesis to find a combination that is most effective in our scenario. However, this combination still has shortcomings as it is unable to model infrequent appearances well. For instance, it does not reliably capture night appearances for scenes in our datasets. We hypothesize that the appearance encoder (which is jointly trained with the rendering network) is not expressive enough to capture the large variability in the data.

To improve the model expressiveness, our approach is to stabilize the joint training of RR and EaE^{a} by pretraining the appearance network EaE^{a} independently on a proxy task. We then employ a staged training approach in which the rendering network RR is first trained using fixed appearance embeddings, and finally we jointly fine-tune both networks. This staged training regime allows for a simpler model that captures more complex appearances.

We present our baseline approach, which adapts state-of-the-art multimodal synthesis techniques, and then our staged training strategy, which pretrains the appearance encoder on a proxy task.

Our baseline uses BicycleGAN with two main adaptations. First, our appearance encoder also takes as input the buffer BiB_{i}, as described above. Second, we add a cross-cycle consistency loss similar to to encourage appearance transfer across viewpoints. Let z1a=Ea(I1,B1)z^{a}_{1}=E^{a}(I_{1},B_{1}) be the captured appearance of an input image I1I_{1}. We apply a reconstruction loss between image I1I_{1} and cross-cycle reconstruction I^1=R(B1,z^1a)\hat{I}_{1}=R(B_{1},\hat{z}^{a}_{1}), where z^1a\hat{z}^{a}_{1} is computed through a cross-cycle with a second image (I2,B2)(I_{2},B_{2}), i.e. z^1a=Ea(R(B2,z1a)),B2)\hat{z}^{a}_{1}=E^{a}(R(B_{2},z^{a}_{1})),B_{2}). We also apply a GAN loss on the intermediate appearance transfer output R(B2,z1a)R(B_{2},z^{a}_{1}) as in .

Staged appearance training

The key to our staged training approach is the appearance pretraining stage, where we pretrain the appearance encoder EaE^{a} independently on a proxy task. We then train the rendering network RR while fixing the weights of EaE^{a}, allowing RR to find the correlations between output images and the embedding produced by the proxy task. Finally, we fine-tune both EaE^{a} and RR jointly.

This staged approach simplifies and stabilizes the training of RR, enabling training of a simpler network with fewer regularization terms. In particular, we remove the cycle and cross-cycle consistency losses, the latent vector reconstruction loss, and the KL-divergence loss, leaving only a direct reconstruction loss and a GAN loss. We show experimentally in Section 4 that this approach results in better appearance capture and rerenderings than the baseline model.

Appearance pretraining

To pretrain the appearance encoder EaE^{a}, we choose a proxy task that optimizes an embedding of the input images into the appearance latent space using a suitable distance metric between input images. This training encourages embeddings such that if two images are close under the distance metric, then their appearance embeddings should also be close in the appearance latent space. Ideally the distance metric we choose should ignore the content or viewpoint of IiI_{i} and BiB_{i}, as our goal is to encode a latent space that is independent of viewpoint. Experimentally we find that the style loss employed in neural style-transfer work has such a property; it largely ignores content and focuses on more abstract properties.

To train the embedding, we use a triplet loss, where for each image IiI_{i}, we find the set of kk closest and furthest neighbor images given by the style loss, from which we can sample a positive sample IpI_{p} and negative sample InI_{n}, respectively. The loss is then:

where gijg^{j}_{i} is the Gram matrix of activations at the jthj^{th} layer of a VGG network of image IiI_{i}, and α\alpha is a separation margin.

3 Semantic conditioning

To account for transient objects in the scene, we condition the rerendering network on a semantic labeling SiS_{i} of image IiI_{i} that depicts the location of transient objects such as pedestrians. Specifically, we concatenate the semantic labeling SiS_{i} to the deep buffer BiB_{i} wherever the deep buffer was previously used. This discourages the network from encoding variations caused by the location of transient objects in the appearance vector, or associating such transient objects with specific viewpoints, as shown in Figure 2.

A separate benefit of semantic labeling is that it allows the rerendering network to reason about static objects in the scene not captured in the 3D reconstruction, such as lampposts in San Marco Square. This prevents the network from haphazardly introducing such objects, and instead lets them appear where they are detected in the semantic labeling, which is a significantly simpler task. In addition, by adding the segmentation labeling to the deep buffer, we allow the appearance encoder to reason about semantic categories like sky or ground when computing the appearance latent vector.

We compute “ground truth” semantic segmentations on the input images IiI_{i} using DeepLab trained on ADE20K . ADE20K contains 150 classes, which we map to a 3-channel color image. We find that the quality of the semantic labeling is poor on the landmarks themselves, as they contain unique buildings and features, but is reasonable on transient objects.

Using semantic conditioning, the rerendering network takes as input a semantic labeling of the scene. In order to rerender virtual camera paths, we need to synthesize semantic labelings for each frame in the virtual camera path. To do so, we train a separate semantic labeling network that takes as input the deep buffer BiB_{i}, instead of the output image IiI_{i}, and estimates a “plausible” semantic labeling S^i\hat{S}_{i} for that viewpoint given the rendered deep buffer BiB_{i}. For simplicity, we train a network with the same architecture as the rendering network (minus the injected appearance vector) on samples (Bi,Si)(B_{i},S_{i}) from the aligned dataset, and we modify the semantic labelings of the ground truth images SiS_{i} and mask out the loss on pixels labeled as transient as defined by a curated list of transient object categories in ADE20K.

Evaluation

Here we provide an extensive evaluation of our system. Please also refer to the supplementary video to best appreciate the quality of the results, available in the project website: https://bit.ly/2UzYlWj.

Datasets

We evaluate our method on five datasets reconstructed with COLMAP from public images, summarized in Table 1. A separate model is trained for each dataset. We create aligned datasets by rendering the reconstructed point clouds with a minimum dimension of 600 pixels, and throw away sparse renderings (>>85%\% empty pixels), and small images (<<450 pixels across). We randomly select a validation set of 100 images per dataset.

Ablative study

We perform an ablative study of our system and compare the proposed methods in Figure 4. The results of the image-to-image translation baseline method contain additional blurry artifacts near the ground because it hallucinates the locations of pedestrians. Using semantic conditioning, the results improve slightly in those regions. Finally, encoding the appearance of the input photo allows the network to match the appearance. The staged training recovers a closer appearance in San Marco and Pantheon datasets (two bottom rows). However, in Sacre Coeur (top row), the smallest dataset, the baseline appearance model is able to better capture the general appearance of the image, although the staged training model reproduces the directionality of the lighting with more fidelity.

Reconstruction metrics

We report image reconstruction errors in the validation set using several metrics: perceptual loss , L1L_{1} loss, and PSNR. We use the ground truth semantic mask from the source image, and we extract the appearance latent vector using the appearance encoder. Staged training of the appearance fares better than the baseline for all but the smallest dataset (Sacre Coeur), where the staged training overfits to the training data and is unable to generalize. The baseline method assumes a prior distribution of the latent space and is less prone to overfitting at the cost of poorer modeling of appearance.

Appearance interpolation

The rerendering network allows for interpolating the appearance of two images by interpolating their latent appearance vectors. Figure 5 depicts two examples, showing that the staged training approach is able to generate more complex appearance changes, although its generated interpolations lack realism when transitioning between day and night. In the following, we only show results for the staged training model.

Appearance transfer

Figure 6 demonstrates how our full model can transfer the appearance of a given photo to others. It shows realistic renderings of the Trevi fountain from five different viewpoints under four different appearances obtained from other photos. Note the sunny highlights and the spotlit night illumination appearance of the statues. However, these details can flicker when synthesizing a smooth camera path or smoothly interpolating the appearance in the latent space, as seen in the supplementary video.

Image interpolation

Figure 7 shows sets of two images and frames of smooth image interpolations between them, where both viewpoint and appearance transition smoothly between them. Note how the illumination of the scene can transition smoothly from night to day. The quality of the results is best appreciated in the supplementary video.

Semantic consistency

Figure 8 shows the output of the staged training model with ground truth and predicted segmentation masks. Using the predicted masks, the network produces similar results on the building and renders a scene free of people. Note however how the network depicts pedestrians as black, ghostly figures when they appear in the segmentation mask.

Comparison to 3D reconstruction methods

We evaluated our technique against the one of Shan et al. on the Colosseum, which contains 3K images, 10M color vertices and 48M triangles and was generated from Flickr, Google Street View, and aerial images. Their 3D representation is a dense vertex-colored mesh, where the albedo and vertex normals are jointly recovered together with a simple 8-dimensional lighting model (diffuse, plus directional lighting) for each image in the photo collection.

Figure 9 compares both methods and the original ground truth image. Their method suffers from floating white geometry on the top edge of the Colosseum, and has less detail, although it recovers the lighting better than our method, thanks to its explicit lighting reasoning. Note that both models are accessing the test image to compute lighting coefficients and appearance latent vectors, with dimension 8 in both cases, and that we use the predicted segmentation labelings from BiB_{i}.

We ran a randomized user study on 20 random sets of output images that do not contain close-ups of people or cars, and were not in our training set. For each viewpoint, 200 participants chose “which image looks most real?” between an output of their system and ours (without seeing the original). Respondents preferred images generated by our system a 69.9%69.9\% of the time, with our technique being preferred on all but one of the images. We show the 20 random sets of the user study in the supplementary material.

Discussion

Our system’s limitations are significantly different from those of traditional 3D reconstruction pipelines:

Our model relies heavily on the segmentation mask to synthesize parts of the image not modeled in the proxy geometry, like the ground or sky regions. Thus our results are very sensitive to errors in the segmentation network, like in the sky region in Figure 10(c) or an appearing “ghost pole” artifact in San Marco (frame 40 of bottom row in Figure 7, best seen in video). Jointly training the neural rerenderer together with the segmentation network could reduce such artifacts.

Neural artifacts

Neural networks are known to produce screendoor patterns and other intriguing artifacts . We observe such artifacts in repeated structures, like the patterns on the floor of San Marco, which in our renderings are misaligned as if hand-painted. Similarly, the inscription above the Trevi fountain is reproduced with a distorted font (see Figure 10(a)).

Incomplete reconstructions

Sometimes an image contains partially reconstructed parts of the 3D model, creating large holes in the rendered BiB_{i}. This forces the network to hallucinate the incomplete regions, generally leading to blurry outputs (see Figure 10(b)).

Temporal artifacts

When smoothly varying the viewpoint, sometimes the appearance of the scene can flicker considerably, especially under complex appearance, such as when the sun hits the Trevi Fountain, creating complex highlights and cast shadows. Please see the supplementary video for an example.

In summary, we present a first attempt at solving the total scene capture problem. Using unstructured internet photos, we can train a neural rerendering network that is able to produce highly realistic scenes under different illumination conditions. We propose a novel staged training approach that better captures the appearance of the scene as seen in internet photos. Finally, we evaluate our system on five challenging datasets and against state-of-the-art 3D reconstruction methods.

Acknowledgements: We thank Gregory Blascovich for his help in conducting the user study, and Johannes Schönberger and True Price for their help generating datasets.

Appendix A Supplementary Results

Figure 12 shows additional results of diverse appearances modeled by our proposed staged training method on the San Marco dataset. As in Figure 6, it shows realistic renderings of five different scenes/viewpoints under four different appearances obtained from other photos.

Qualitative comparison.

We evaluate our technique against Shan et al. on the Colosseum. In Section 4, we report the result of a user study run on 20 randomly selected sets of output images that do not contain close-ups of people or cars, and were not in our training set. Figures 13, 14 show a side-by-side comparison of all 20 images used in the user study.

Quantitative evaluation with learned segmentations

To quantitatively evaluate rerendering using estimated segmentation masks, we generate semantic labelings for the validation set, as described in Section 3.3, and recompute the quantitative metrics, as in Table 1, for our proposed method. Note that estimated semantic maps will not perfectly match those of the ground truth validation images. For example, ground truth semantic maps could contain the segmentation of transient objects, like people or trees. So, it is not fair to compare reconstructions based on estimated segmentation maps to the ground truth validation images. While results in Table 2 show some performance drop as expected, we still get a reasonable performance compared to that in Table 1. In fact, we still perform better than the BicycleGAN baseline on the Trevi, Pantheon and Dubrobnik datasets, even though the BicycleGAN baseline uses ground truth segmentation masks.

Appendix B Implementation Details

We use different networks for the staged training and the baseline mode. We obtain best results for each model with different networks. Below, we provide an overview of the different architectures used in the staged training and the baseline models. Code will be available at https://bit.ly/2UzYlWj.

Our rerendering network is a symmetric encoder-decoder with skip connections. The generator is adopted from without using progressive growing. Specifically, we extend the GAN architecture in to a conditional GAN setting. The encoder/decoder operates at a 256×256256\times 256 resolution, with 6 downsampling/upsampling blocks. Each block has a downsampling/upsampling layer followed by two single-strided 3×33\times 3 conv layers with a leaky ReLu ( α=0.2\alpha=0.2) and pixel-norm layers. We add skip connections between the encoder and decoder by concatenating feature maps at the beginning of each decoder block. We use 64 feature maps at the first encoder and double the size of feature maps after each downsampling layer until it reaches size 512.

B.2 Appearance encoder architecture

B.3 Baseline architecture

We use a faithful Tensorflow implementation of the encoder-decoder network and appearance encoder in using their PyTorch released code as a guideline. We adapt their training pipeline to the single-domain supervised setup as described in Section 3.2 in our paper.

B.4 Aligned datasets

Figure 11 shows sample frames from aligned datasets we generate as described in Section 3.1.

B.5 Latent space visualization

Figure 15 visualizes the latent space learned by the appearance encoder, EaE^{a}, after appearance pretraining and finetuning in our staged training, as well as training EaE^{a} with the BicycleGAN baseline. The embedding learned during the appearance pretraining stage shows meaningful clusters, but has lower quality than the one learned after finetuning, which is comparable to the one of the BicycleGAN baseline.

References