Simple and Effective Synthesis of Indoor 3D Scenes

Jing Yu Koh, Harsh Agrawal, Dhruv Batra, Richard Tucker, Austin Waters, Honglak Lee, Yinfei Yang, Jason Baldridge, Peter Anderson

Introduction

We synthesize immersive 3D indoor scenes from one or more context images captured along a trajectory. Our aim is to generate high-resolution images and videos from novel viewpoints, including viewpoints that extrapolate far beyond the context image(s).

Solving this problem would make photos and videos interactive and immersive, with applications not only to content creation but also robotics and embodied AI. For example, models that can predict around corners could be used by navigation agents as world models (Ha and Schmidhuber 2018) for model-based planning in novel environments (Koh et al. 2021b; Finn and Levine 2017). Such models could also be used to train agents in interactive environments synthesized from static images or video.

Previous approaches attempting this under large viewpoint changes (Wiles et al. 2020; Koh et al. 2021b; Rockwell, Fouhey, and Johnson 2021) typically operate on point clouds, which are accumulated from the available context images. The use of point clouds naturally incorporates camera projective geometry into the model and helps maintain the 3D consistency of the scene (Mallya et al. 2020). To generate novel views, these approaches reproject the point-cloud relative to the target camera pose into an RGB-D guidance image (see Fig. 1). These guidance images are sparse because of missing context and completing them requires extensive inpainting and outpainting. In prior work, Pathdreamer (Koh et al. 2021b) achieves this by assuming the availability of semantic segmentations, and combining a stochastic depth and semantic segmentation (structure) generator with an RGB image generator infused with semantic, depth and guidance information via Multi-SPADE spatially adaptive normalization layers (Park et al. 2019; Mallya et al. 2020). On the other hand, PixelSynth (Rockwell, Fouhey, and Johnson 2021) creates guidance images using a differentiable point cloud renderer, generates a support set of additional views using PixelCNN++ (Salimans et al. 2017) operating on the latent space of a VQ-VAE (Razavi, van den Oord, and Vinyals 2019), combines and refines these images using a GAN (Goodfellow et al. 2014) similar to SynSin (Wiles et al. 2020), and then repeats this process over many samples using a combination of discriminator loss and the entropy of a scene classifier to select the best output.

We propose a simple alternative: an image-to-image GAN that maps directly from guidance images to high-resolution photorealistic RGB-D (see Fig. 1). Compared to Pathdreamer, our simple model forgoes the stochastic structure generator, spatially adaptive normalization layers, dependence on semantic segmentations, and multi-step training. By dropping the dependence of semantic segmentation inputs, we unlock training on a much broader range of data, such as more commonly available RGB-D datasets (Xia et al. 2018; Li and Snavely 2018; Nathan Silberman and Fergus 2012), and video data such as the RealEstate10K dataset (Zhou et al. 2018b) from YouTube. We eschew many components of PixelSynth: differentiable rendering, support set generation using PixelCNN++and a VQ-VAE, and the multiple sampling and re-ranking procedure.

Perhaps surprisingly, with random masking of the guidance images during training, plus other architectural changes supported by thorough ablation studies, our lightweight approach outperforms prior work. In human evaluations of image quality, our model is preferred to Pathdreamer in 60% of comparisons and preferred to PixelSynth in 77%. Our FID scores on 360° panoramic images from Matterport3D (Chang et al. 2017) improve over Pathdreamer’s by 27.9% relatively (from 27.2 to 19.6) when predicting single step viewpoint changes (an average of 2.2m), and from 65.8 to 58.0 when predicting over longer trajectories containing a sequence of 6 novel viewpoints. On RealEstate10K (Zhou et al. 2018b) – a collection of real estate video walkthroughs from YouTube – our FID scores outperform PixelSynth, improving from 25.5 to 23.5 (and from 23.6 to 21.5 for indoor images). Our model is capable of composing compelling 3D synthetic environments from a single image, which can be used to produce realistic video renderings.Video results: https://youtu.be/4fVG0vg7yXI.

Motivated by the strong results on image generation, we also show the usefulness of our model for data augmentation in embodied AI. We focus on the vision-and-language navigation (VLN) task, which requires an agent to follow natural language navigation instructions in previously unseen photorealistic environments. Training data for the task consists of instruction-trajectory demonstrations, where each trajectory is defined by a sequence of high-resolution 360° panoramas. Using our model, we spatially perturb the location of the training panoramas by synthesizing views up to 1.5m away and augment the training dataset. This reduces overfitting to the incidental details of these trajectories, and improves the success in unseen environments by an additional 1% on its own, or 1.5% when combined with renders of spatially-perturbed images from the Habitat (Savva et al. 2019) simulator – achieving state-of-the-art performance on the R2R test set (Tab. 4). Our code is publicly releasedhttps://github.com/google-research/se3ds to facilitate generative trajectory augmentation and applications to downstream tasks.

Related Work

Synthesizing novel views from sets of 2D images has been studied extensively from the lens of multi-view geometry (Debevec, Taylor, and Malik 1996; Avidan and Shashua 1997; Zitnick et al. 2004), and more recently with deep learning based approaches (Flynn et al. 2016; Kar, Häne, and Malik 2017; Henzler et al. 2018; Flynn et al. 2019; Srinivasan et al. 2019; Zhou et al. 2018b; Mildenhall et al. 2019). Various methods using explicit 3D scene representations have been proposed, including point cloud representations (Wiles et al. 2020), layered depth images (Dhamo et al. 2019), and mesh representations (Shih et al. 2020). More recently, Neural Radiance Fields (NeRF) (Mildenhall et al. 2020) models have achieved impressive results on novel view synthesis. NeRFs optimize an underlying continuous volumetric scene function, learning an implicit 3D scene representation from multiple images. Follow up papers have extended the NeRF formulation to large-scale environments (Tancik et al. 2022), unconstrained photo collections (Martin-Brualla et al. 2021), learning scene priors (Yu et al. 2020a), using multi-scale representations (Barron et al. 2021), and reducing the amount of input images required (Jain, Tancik, and Abbeel 2021). However, at present NeRF models primarily focus on scene representation rather than generalization to unseen environments, and are unable to perform realistic synthesis of previously unseen 3D scenes at high resolution (which is our focus).

While early work in novel view synthesis focused on smaller view changes and settings with multiple input images, recent methods tackle single-image novel view synthesis (Tucker and Snavely 2020; Hu et al. 2021; Shih et al. 2020) and large viewpoint changes (Wiles et al. 2020; Koh et al. 2021b; Rockwell, Fouhey, and Johnson 2021; Rombach, Esser, and Ommer 2021) and long-term future prediction for indoor scenes (Ren and Wang 2022). The primary challenge of this task is being able to handle both inpainting of missing pixels, as well as outpainting of large regions of the image from limited context, while maintaining consistency with the existing scene. The closest works to ours are PixelSynth (Rockwell, Fouhey, and Johnson 2021) and Pathdreamer (Koh et al. 2021b) (see introduction).

Point Cloud Rendering.

A number of previous works explore rendering novel views from point clouds. The Gibson simulator (Xia et al. 2018) combined point cloud rendering with a neural net ‘filler’ to fix artifacts and produce more realistic images, although the neural net was trained with perceptual loss (Johnson, Alahi, and Fei-Fei 2016) rather than GAN based losses. Several papers (Yu et al. 2020b; Fu, Hu, and Guo 2020) propose methods to inpaint object point clouds with new points in order to fill holes, which would allow more realistic images can be synthesized. Song et al. (2020) proposes a model which takes a colored 3D point cloud of a scene as input, and synthesizes a photo-realistic image from a novel viewpoint. Whereas, Cortinhal, Kurnaz, and Aksoy (2021) relies only on the semantics of a scene to synthesize a panoramic color image from a given full 3D LiDAR point cloud. These methods often consider smaller point clouds, which limits their efficiency and viability for high resolution imagery. For example, (Song et al. 2020) trains on point clouds of size 4096×64096\times 6, while the accumulated point cloud for our method uses RGB-D images of size 1024×5121024\times 512 (at least 64×64\times larger).

Generative Data Augmentation.

Generative models are good alternatives to standard data augmentation approaches. They can synthesize artificial samples that match the distribution and characteristics of an underlying dataset, augmenting the training set with additional novel examples (Antoniou, Storkey, and Edwards 2017; Sandfort et al. 2019). We examine whether novel view synthesis can be used to create new training trajectories for a navigation agent, by spatially-perturbing camera viewpoints post hoc.

In imitation learning, expert demonstrations must be augmented to minimize the differences between the state distribution seen in training, and those induced by the agent during inference, which otherwise causes compounding errors (Ross and Bagnell 2010). In practice this usually means augmenting the dataset with examples of recoveries from error, which may be rare in expert demonstrations. For example, robots and self-driving vehicles (Codevilla et al. 2018; Bojarski et al. 2016) can be instrumented to record from three cameras simultaneously (one facing forward and two shifted to the left and right). This allows recordings from the shifted cameras, as well as intermediate synthetically reprojected views, to be added to the training set with adjusted control signals to simulate recovery from drift (Codevilla et al. 2018; Bojarski et al. 2016). In effect, we propose a flexible, hardware-free alternative to the multi-camera setup. We investigate this in the context of indoor vision-and-language navigation (VLN) on the R2R dataset (Anderson et al. 2018b). To the best of our knowledge, we are the first to show benefits from data augmentation using novel view synthesis models in a photorealistic setting.

Approach

We aim to synthesize high-resolution images and videos from novel viewpoints in buildings, conditioning on one or more RGB and depth (RGB-D) observations as context. Specifically, given context consisting of a sequence of RGB-D image observations and their associated camera poses (I1:t−1,P1:t−1)(I_{1:t-1},P_{1:t-1}), our goal is to generate realistic RGB-D images for one or more target camera poses [Pt,Pt+1,⋯ ,PT][P_{t},P_{t+1},\cdots,P_{T}]. Target poses may require extrapolating far beyond the context images (e.g., predicting around corners), requiring the model to generate and in-fill potentially large regions of missing information – even entire rooms.

Our model requires depth values and camera poses for the context images (I1:t−1,P1:t−1)(I_{1:t-1},P_{1:t-1}) to create an accumulated point-cloud. Given a new pose PtP_{t}, we reproject our accumulated point cloud, similar to (Mallya et al. 2020; Liu et al. 2021; Koh et al. 2021b) into an image from that viewpoint. We call that image a guidance image because it will be used to guide the inpainting and outpainting process later. In our experiments, we report results using both ground-truth and estimated depth values and camera poses, depending on the dataset. When predicting over multiple steps [Pt,Pt+1,⋯ ,PT][P_{t},P_{t+1},\cdots,P_{T}], predictions are accumulated in the point cloud to maintain 3D consistency. We do not make assumptions about the input image format, which enables our model to run on equirectangular panoramas from Matterport3D (Chang et al. 2017) as well as perspective images from RealEstate10K (Zhou et al. 2018b). To accommodate both formats, we simply use the appropriate camera models (while keeping model architecture the same).

Model Architecture.

We propose a simple, single-stage, end-to-end trainable model to convert a guidance image directly into a high-resolution photorealistic RGB-D output. Fig. 1 provides an overview of this process. Our model uses an encoder-decoder CNN architecture inspired by RedNet (Jiang et al. 2018). We use ResNet-101 (He et al. 2016) as the encoder, and a ‘mirror image’ of the ResNet-101 as the decoder, replacing convolutions with transposed convolutions for upsampling. A single encoder is used for both the RGB and depth inputs in the guidance, but separate decoders are used for predicting RGB and depth outputs. Following RedNet, skip connections are introduced between the encoder and decoders to preserve spatial information.

Inpainting or outpainting large image regions that lack input guidance is a major challenge, e.g. predicting around corners or under large viewpoint changes. To overcome this, we replace all convolutions in the encoder with partial convolutions (Liu et al. 2018). Partial convolutions only convolve valid regions in the input with the convolution kernels, so the output features are minimally affected by missing pixels in the guidance images. Further, we increase the effective receptive field size by adding four additional convolutional layers to the output of the encoder helping propagate non-local information. We found this to be essential, without which the model does not learn anything meaningful.

We also experimented with the SVG (Denton and Fergus 2018) structure generator proposed in Pathdreamer (Koh et al. 2021b). This module learns a prior distribution to model stochasticity in the predictions. During training, the prior distribution is learnt by minimizing KL-divergence from the posterior distribution derived from ground truth data. At inference time, different outputs are generated by sampling from the prior distribution. We found that generation quality was worse when this module was included (as measured by FID and visual inspection), despite improvement in image diversity (more details in the Ablation Studies section). Hence, we omit the SVG module in our final model for the sake of simplicity and performance. However, we stress that this module can be easily incorporated to enable a trade-off between generation quality and diversity.

Loss Functions.

The encoder-decoder model GG is optimized to minimize a joint loss consisting of an L1L_{1} loss between the predicted depth d^t\hat{d}_{t} and the ground-truth depth dtd_{t}, and an adversarial loss on depth and RGB images. The discriminator DD used for adversarial training is based on PatchGAN (Isola et al. 2017), and applied to the 4-channel RGB-D image (either generated [I^t,d^t][\hat{I}_{t},\hat{d}_{t}] or ground-truth [It,dt][I_{t},d_{t}]). Note that while the use of separate decoders could allow for greater divergence between the RGB and depth predictions, this is mitigated by the discriminator which helps to enforce consistency between RGB and depth, which is essential for multi-step predictions. The training losses for the generator and discriminator are:

Notably, we avoid the VGG perceptual loss commonly used in prior work on conditional image synthesis (Park et al. 2019; Koh et al. 2021b; Mallya et al. 2020) as we find that it is unnecessary for strong performance. We justify this through careful ablation experiments.

In all experiments we train the model end-to-end from scratch, and we randomly mask up to 75% of the input guidance image for data augmentation. Masking has been shown to be an effective method of pre-training visual representations (He et al. 2021; Dosovitskiy et al. 2020). We similarly find that this masking strategy significantly improves the generation quality of our model in unseen environments.

View Synthesis Experiments

We conduct experiments on two datasets of diverse indoor environments: Matterport3D (Chang et al. 2017), which contains 3D meshes of 90 buildings reconstructed from 11K high-resolution RGB-D panoramas (panos), and RealEstate10K (Zhou et al. 2018b), a collection of up to 10,000 YouTube video walkthroughs of real estate properties. Few prior works attempt view synthesis from a single image under large viewpoint changes. We compare to PixelSynth (Rockwell, Fouhey, and Johnson 2021), which builds on and outperforms SynSin (Wiles et al. 2020), and Pathdreamer (Koh et al. 2021b). We report automated Fréchet Inception Distance (FID) (Heusel et al. 2017) scores (lower is better) and pairwise human evaluations of image quality.

Both PixelSynth and Pathdreamer report results on Matterport3D. However, as the two papers are concurrent work, the evaluation procedures are not comparable. PixelSynth is trained and evaluated on 256×256256\times 256 perspective renders from the reconstructed 3D meshes. These renders are not photorealistic and often contain large regions of missing/black pixels due to poor reconstruction. Pathdreamer is trained and evaluated on the high-res 1024×5121024\times 512 360° panoramic images, and evaluated over multiple steps of prediction. Therefore, on the Matterport3D dataset we follow the more challenging Pathdreamer evaluation procedure.

Following Pathdreamer, we train our model using 1024×5121024\times 512 equirectangular RGB-D images and ground-truth pose information. Unlike Pathdreamer, our model does not require ground-truth semantic segmentations as input, and we train for only single-step prediction, i.e., predicting the pano at an adjacent viewpoint using one context pano as input. Evaluations are based on

Val-Seen and Val-Unseen splits of the Room-to-Room (R2R) dataset (Anderson et al. 2018b), which are comprised of sequences of adjacent panoramas (∼\sim2.2m apart). Given the first RGB-D pano in the path and its (x,y,z)(x,y,z) pose as context, the model must generate panos for the remainder of the path given only their poses (up to a maximum of 6 steps).

FID Scores.

In Tab. 1 we report FID scores for the generated RGB images over 1–6 prediction steps (representing trajectory rollouts of 2–13m), using 1–3 panos as context. We evaluate in two settings – novel viewpoints in environments seen during training (Val-Seen) and previously unseen (Val-Unseen) environments. Since the top and bottom 12.5% of each Matterport3D pano is blurred, we crop these areas before calculating FID and re-state the previously reported results from Pathdreamer for fair comparison. On Val-Unseen, based on FID score, our model outperforms Pathdreamer in all settings. When using a single context image for 1-step prediction, we improve FID scores from 27.6 to 19.6. Similarly, when using 3 panos as context, FID scores improve from 32.0 to 22.3. We notice similar gains for multi-step (1-6 steps) prediction. On Val-seen, we improve FID scores slightly over Pathdreamer but perform worse on multi-step prediction in Val-Seen environments. We attribute this to Pathdreamer’s recurrent training regime for their Structure Generator, under which the model is trained over multiple prediction steps while accumulating its own outputs as additional context. In initial experiments we found that this improved results in the training environments, but did not generalize to unseen environments, perhaps because the accumulated context from the model predictions doesn’t reconcile with the target image used in the L1L_{1} depth reconstruction loss. In unseen environments (which is our focus), our model performs better on FID score at every step. As shown in Fig. 2, despite using a far simpler model than Pathdreamer that does not require ground-truth semantic segmentation inputs, our model produces a clearer image of the wooden cabinet at 1.4m, the wooden ceiling rafters at 5.7m, and the passageway at 7.5m away.

Human evaluations.

We perform human evaluations of 1,000 image pairs from our model and Pathdreamer. Each pair is evaluated by 5 different human evaluators, similar to Koh et al. (2021a). Since human evaluators may be unfamiliar with the distortion characteristics of 360° equirectangular images, for each pair evaluators are shown a random perspective projection from the generated images with 102° horizontal field-of-view and 16:9 aspect ratio. Images generated by our model are preferred over Pathdreamer images in 60.4% of cases (see Fig. 3). For images with unanimous preference (preferred by 5/5 raters), our model is greatly preferred: 30.7% compared to 12.3% for Pathdreamer.

Ablation Studies

To validate the design choices detailed earlier, we perform an extensive ablation study summarized in Tab. 2.

We find that the adversarial loss and L1L_{1} depth reconstruction loss are both essential to strong performance: dropping the L1L_{1} loss (row 2) reduces generation quality for 1 step ahead (FID@1 increases from 19.6 to 22.7), as well as over longer horizons (FID@{1-6} increases from 58.0 to 64.8). When we include the VGG perceptual loss (Johnson, Alahi, and Fei-Fei 2016) (row 3), we find that FID@1 is improved (19.6 to 18.2), but FID@{1-6} deteoriates (58.0 to 66.5). This is likely because visual quality is improved at the cost of RGB and depth consistency, which is essential for rollouts over multiple prediction steps. When KLD loss for stochastic noise conditioning (Denton and Fergus 2018; Koh et al. 2021b) is included (row 4), generation quality worsens slightly for both 1 step and 1-6 steps ahead (FID@1 from 19.6 to 22.6, and FID@{1-6} from 58.0 to 65.2). While modeling stochasticity can be beneficial, we leave this out to maximize generation quality. Notably, even with the inclusion of the KLD loss and noise conditioning our results still outperform Pathdreamer, particularly in 1 step predictions (FID@1 of 22.6 vs. 27.2).

Separate decoders.

As described previously, we use a shared encoder with two separate decoders for RGB and depth. Using a single shared decoder (row 5) degrades generation quality for 1 step (FID@1 deteriorates from 19.6 to 24.1), although longer rollouts see slight improvement, improving FID@{1-6} from 58.0 to 56.9. This is likely due to marginally improved consistency between the RGB and depth outputs, which is essential for multistep prediction.

Random masking and conv layers.

We insert 4 additional 3×33\times 3 convolutional layers between the encoder and the decoders to increase the model’s receptive field and better propagate global image information. This is combined with random masking of upto 75% guidance pixels during training. These changes significantly improve generation quality both in one-step predictions (FID@1 improves from 24.2 in row 6 to 19.6) and over multiple steps (FID@{1-6} improves from 68.0 to 58.0).

Ground-truth depth for GAN loss.

During training, the adversarial GAN (Goodfellow et al. 2014) hinge loss (Lim and Ye 2017) is computed on the generated RGB-D image. We experimented with replacing the generated depth channel with the ground-truth depth channel, to explore whether this would help enforce RGB and depth consistency. Our findings suggest that it does not, with FID@1 degrading from 19.6 to 24.3 and FID@{1-6} from 58.0 to 66.1.

RealEstate10K Experiments

The RealEstate10K (Zhou et al. 2018b) dataset consists of a collection of real estate walkthrough videos. In its raw form, the dataset lacks depth images and camera poses, which are required to create point clouds and re-project guidance images. The PixelSynth (Rockwell, Fouhey, and Johnson 2021) and SynSin (Wiles et al. 2020) models, which we compare to, include a depth estimation module which is trained on the RE10K dataset using reprojection losses. We use MiDaS (Ranftl, Bochkovskiy, and Koltun 2021), a pretrained transformer-based monocular depth estimation model, and we do not finetune on RE10K. See Appendix for more details. As with the Matterport3D experiments, we train our model for single-step prediction, using a perspective camera projection for the guidance images rather than the equirectangular camera model used in the Matterport3D experiments. Following PixelSynth, we select image pairs for training with a camera rotation of 20°−60°20\degree-60\degree that are estimated to be ≤1m\leq 1m apart, and train and evaluate with an image resolution of 256×256256\times 256.

Evaluation.

Given an RGB context image and a target camera pose, the model must generate an RGB image for the target pose. We evaluate using the same 3,600 context-target image pairs and camera poses as PixelSynth. Since the scaling of the camera poses is arbitrary, during evaluation we scale the MiDaS depth predictions to match the PixelSynth depth predictions as closely as possible, so that the depth predictions and camera poses have consistent scaling. Qualitatively, the guidance images used by each model are extremely similar. We verify that none of the videos contributing evaluation images are in our training set.

FID Scores.

As reported in Tab. 3, our model outperforms existing methods, achieving an FID score of 23.5 compared to 25.5 and 34.7 achieved by PixelSynth (Rockwell, Fouhey, and Johnson 2021) and SynSin (Wiles et al. 2020) respectively. Notably, PixelSynth results were achieved by generating 50 sample target images for each input example, and ranking them according to a combination of discriminator loss and the entropy of a scene classifier trained on MIT Places 365 (Zhou et al. 2018a) to select the best. Our results represent a single prediction from our model. Since we excluded outdoor scenes from our training data, we also report results on a subset of only indoor images (3,122 of the 3,600 images, based on manual inspection). On this subset, we achieve an FID score of 21.5 vs. 23.6 from PixelSynth.

Human Evaluations.

Similar to Matterport3D, we perform human evaluations on RE10K. As shown in Fig. 5, human annotators significantly prefer images generated by our model over PixelSynth, rating them as more realistic in 77.3% of cases. In addition, for images which achieve unanimous preference (selected by 5/5 raters), our model is massively preferred, spanning 40.5% of images compared to just 3.5% for PixelSynth. Qualitative comparisons in Fig. 4 show that our model completes some of the scenes by imagining adjacent rooms (Row 1, Row 2), while keeping wall and carpet colors consistent (Row 1), and introducing new elements like lamps and a wall painting (Row 3).

Trajectory Augmentation in VLN

Motivated by our strong generation results, we investigate the usefulness of our model for synthetic trajectory augmentation, i.e., augmenting the training data for a vision-based navigation robot by spatially perturbing training trajectories post hoc. We focus on vision-and-language navigation (VLN) using the Room-to-Room (R2R) dataset (Anderson et al. 2018b). This task requires an agent to follow natural language navigation instructions in unseen photorealistic indoor environments, by navigating between locations where high-res 360° Matterport3D panos have been captured. Training data consists of 14K instruction-trajectory pairs, where each trajectory is defined by a sequence of panos that are adjacent in a navigation graph. Constrained by the location of captured images, VLN agents tend to overfit to the incidental details of these trajectories (Zhang, Tan, and Bansal 2020), which contributes to the large performance drop in unseen environments (refer Tab. 5). We hypothesize that spatially perturbing the location of the captured images could reduce overfitting and improve generalization.

We base our experiments on the VLN↻\circlearrowrightBERT agent (Hong et al. 2021), an image-text cross-modal transformer with a recurrent state that is updated over time as the agent moves. The agent is trained using a mixture of imitation learning and A2C (Mnih et al. 2016). To make the baseline as strong as possible, we first upgrade the image representation used by the model from ResNet-152 (He et al. 2016) trained on Places365 (Zhou et al. 2018a) to MURAL-large (Jain et al. 2021), an EfficientNet-B7 architecture (Tan and Le 2019) trained on 1.8B web image-text pairs. As illustrated in Tab. 4, this improves the agent’s Success Rate (SR ↑\uparrow) on the R2R Val-Unseen split from 62.2% to 67.3%, which beats even the recent state-of-the-art HAMT model (Chen et al. 2021). SR is defined as the proportion of trajectories ending within 3m of the end of the target location. We also report Navigation Error (NE ↓\downarrow), the average distance in meters between the agent’s final position and the target, and Success rate weighted by the normalized inverse of the Path Length (SPL ↑\uparrow) (Anderson et al. 2018a).

Trajectory Augmentation

To implement synthetic trajectory augmentation, we continually re-sample pano locations while training the VLN agent. We perturb the agent’s position with a translation sampled uniformly at random from (−1.5m,1.5m)(-1.5\text{m},1.5\text{m}) for directions parallel to the ground plane and from (−0.1m,0.1m)(-0.1\text{m},0.1\text{m}) for height. To avoid perturbing the pano location through a wall or inside an object, we reject any perturbation that exceeds the depth returned in that direction. The pano is generated at 1024×5121024\times 512 resolution by our model trained on Matterport3D, using the nearest two ground-truth RGB-D panos as context. We compare to two alternatives: (1) data augmentation using carefully tuned SimCLR (Chen et al. 2020) operations like random cropping, color distortion and Gaussian blur operations, and (2) rendering the spatially-perturbed pano from the textured mesh using the Habitat (Savva et al. 2019) simulator.

Results

As shown in Tab. 4, trajectory augmentation using our model improves the agent’s SR and SPL on Val-Unseen by 1% (Row 5 vs Row 2), reducing the gap between Val-Seen and Val-Unseen from 8% to 4%. In contrast, Pathdreamer augmentations (Row 3), SimCLR image-based data augmentation (Row 4), or textured mesh based renders using Habitat (Row 5) produce no improvements over the upgraded model (Row 2). Trajectory augmentation using a combination of Habitat and our model (Row 7) produced the largest improvement in SR (+1.5%), virtually closing the gap between Val-Seen and Val-Unseen performance, and on the unseen test set this model outperforms all published prior work (Tab. 5). The R2R dataset (Anderson et al. 2018b) is a mature benchmark at this time, and such gains do not come easily. Unlike Pathdreamer (Koh et al. 2021b), which requires integration at inference time, our visual augmentation procedure is completely training based, and this can be applied to train virtually any off-the-shelf VLN agent.

Conclusion

Synthesizing immersive 3D environments from limited images is a challenging task brought into focus by several recent works. These approaches are highly complex, with many separately trained stages and components. We propose a simple alternative: an image-to-image GAN trained with random input masking combined with other architecture changes. Perhaps surprisingly, our approach outperforms prior work in human evaluations and on FID, and its useful for generative data augmentation as well. We achieve state-of-the-art results on the R2R dataset by spatially perturbing the training images with our model, improving generalization to unseen environments.

Acknowledgements

We would like to thank Noah Snavely and many others for insightful discussions during the development of this paper. We thank Chris Rockwell for helping to set up PixelSynth models. We also thank the Google ML Data Operations team for collecting human evaluations on our generated images.

References

Appendix A Appendix

Appendix B Implementation Details

All models are implemented in TensorFlow 2.0. We set loss weights λGAN=1.0\lambda_{\text{GAN}}=1.0 and λD=100.0\lambda_{\text{D}}=100.0. Spectral normalization is used for all convolutional layers in the generator and discriminator. We train using the Adam optimizer with parameters β1=0.5\beta_{1}=0.5 and β2=0.999\beta_{2}=0.999. The learning rates for the generator and discriminator are set to 1e−41e^{-4} and 4e−44e^{-4} respectively. The discriminator is trained for two training steps for each training step of the generator. For evaluation, we apply exponential moving average of generator weights with decay of 0.9990.999. All models are trained for 300K steps with a batch size of 128, and early stopping is applied based on FID scores on the validation dataset.

Evaluation Details.

For Matterport3D, we follow Pathdreamer to compute the FID scoreWe used https://github.com/mseitzer/pytorch-fid for computing results. using 10,000 random samples for each prediction sequence step. Random horizontal roll and flips are applied to obtain 10,000 samples for evaluation.

For RealEstate10K, we follow PixelSynth in computing the FID score over their evaluation set of 3,600 examples. As our training data does not contain outdoor images, we extract a subset of only indoor images (3,122 out of 3,600) based on an image classifier. We classifier frames as indoors if they are tagged with any of the ‘room’, ‘floor’, ‘kitchen’, ‘table’, ‘sofa’, ‘bed’ categories. We manually inspected images which had uncertain predictions. We ran the same evaluation on this subset of 3,122 images for both our model and PixelSynth.

Appendix C Architectural Details

The detailed generator and discriminator architectures for our model can be found in Tab. 6(a) and Tab. 6(b) respectively. The same architecture is used for both Matterport3D and RealEstate10K, with the exception of the input size H×WH\times W being changed to the appropriate input dimensions (1024×5121024\times 512 and 256×256256\times 256 respectively). All convolutions used in the encoder ResNet-101 are partial convolutions.

Appendix D RealEstate10K Experiment details

Appendix E COLMAP sampling details

Now, we describe the details of our RealEstate10K training dataset. We obtain camera trajectories and depth maps for Youtube videos listed by running SLAM and bundle adjustment on the video frames. RealEstate10K breaks down the YouTube video into a collection of smaller sub-sequences. Each sub-sequence consists of connected frames that were tracked by ORB-SLAM (Mur-Artal, Montiel, and Tardós 2015). Out of the original dataset, only 6,8656,865 videos are available because the rest have been taken down since the release. We ran the original pipeline from RealEstate10K but with longer sequence length (up to 1,000 frames) which results in a total of 108,275108,275 sub-sequences containing over 1919 million frames. Videos in the RealEstate10K dataset often contain walkthroughs of outdoor surroundings as well. Since we are only interested in synthesizing indoor scenes, we filter sub-sequences which contain outdoor frames from our dataset. We filter these by running a multi-label classifier which tags each frame of the sub-sequence with relevant semantic categories. We remove those sub-sequences in which no frame is tagged with any of the ‘room’, ‘floor’, ‘kitchen’, ‘table’, ‘sofa’, ‘bed’ categories. After filtering, we are left with 83,559 sub-sequences containing over 1515 million frames. We obtain camera trajectories corresponding to each sub-sequence by running COLMAP (Schönberger and Frahm 2016), a structure-from-motion pipeline. COLMAP was able to successfully build dense reconstruction for 67,834 out of the 83,559 sub-sequences.

To obtain per-frame depth maps, we use a standard multi-view stereo system in COLMAP. The depth estimates from MVS are often too noisy due to camera-blur, reflections, poor lightning, and other factors (Li et al. 2021). Following procedure outlined in Li et al. (Li et al. 2021), we first filter outlier depth values using the depth refinement method outlined in (Li and Snavely 2018). Further, we remove errorneous depth values by considering the consistency between MVS depth and depth obtained from motion parallax between adjacent pair of frames. Depth estimates from MVS has another limitation. These depth estimates are often extremely sparse. This severely limit the density of point-cloud constructed from context images, and leads to sparse guidance images for the model. To overcome this, we generate dense depth predictions from MiDaS (Ranftl, Bochkovskiy, and Koltun 2021). MiDaS is a state-of-the-art transformer based monocular depth estimation model which outputs a dense depth map from a single RGB image. Of course, the scale of the depth predictions from MiDaS for different frames are inconsistent and not aligned with the camera pose. Thus they cannot be used to create a unified point-cloud. We tackle this by a simple fix. We scale the depth predictions from MiDaS to match the COLMAP MVS depth outputs for valid pixels using the least-squares method. We select only those frames for training for which the depth map produced by MVS covers at least 10% of the image.

Appendix F Human Evaluations

We perform human evaluations of 1,000 image pairs for both the Matterport3D and RealEstate10K datasets. For Matterport3D, we do head-to-head comparisons between our model and Pathdreamer, and for RealEstate10K, we compare our model against PixelSynth. Each pair is evaluated by 5 independent human evaluators, and the user interface is shown in Fig. 7. The two models are anonymized and shown in random order to minimize biases.

Evaluation with Guidance Image

In addition to the evaluation described in our main paper and in the Human Evaluations section, we peform a secondary evaluation for RealEstate10K similar to the process in PixelSynth. In this setup, we show the input guidance image to the evaluators, and task them to select the image that best matches the partial input (see Fig. 8). Under this procedure, we expect the evaluators to focus even more on the regions of the image requiring inpainting. Under this setting, our model significantly outperforms PixelSynth, with 88.0% of examples being preferred by human evaluators (compared to 12% for PixelSynth). For examples with unanimous preference (5/5 raters prefer it), our model is overwhelming preferred: 48.4% compared to 2.0% for PixelSynth.

Appendix G Qualitative Evaluation

In this section, we present additional qualitative results generated by our model. We show cherry-picked examples (Fig. 9) and random examples (Fig. 10) on the Matterport3D dataset. Similarly, for the RealEstate10K dataset we show more cherry-picked examples (Fig. 12) and random examples (Fig. 13). We also provide more side-by-side comparisons of our model against Pathdreamer in Fig. 11.

Our model is also capable of generating compelling, high resolution video by subsampling trajectories with small viewpoint changes. We provide videos displaying generation results for unseen environments in Matterport3D, and for rotations and camera trajectories from RealEstate10K. We refer readers to the videoVideo results: https://youtu.be/4fVG0vg7yXI for these results.