Pathdreamer: A World Model for Indoor Navigation

Jing Yu Koh, Honglak Lee, Yinfei Yang, Jason Baldridge, Peter Anderson

Introduction

World models , or models of environments , are an appealing way to represent an agent’s knowledge about its surroundings. An agent with a world model can predict its future by ‘imagining’ the consequences of a series of proposed actions. This capability can be used for sampling-based planning , learning policies directly from the model (i.e., learning in a dream) , and for counterfactual reasoning . Model-based approaches such as these also typically improve the sample efficiency of deep reinforcement learning . However, world models that generate high-dimensional visual observations (i.e., images) have typically been restricted to relatively simple environments, such as Atari games and tabletops .

Our goal is to develop a generic visual world model for agents navigating in indoor environments. Specifically, given one or more previous observations and a proposed navigation action sequence, we aim to generate plausible high-resolution visual observations for viewpoints that have not been visited, and do so in buildings not seen during training. Beyond applications in video editing and content creation, solving this problem would unlock model-based methods for many embodied AI tasks, including navigating to objects , instruction-guided navigation and dialog-guided navigation . For example, an agent asked to find a certain type of object in a novel building, e.g. ‘find a chair’, could perform mental simulations using the world model to identify navigation trajectories that are most likely to include chair observations – without moving.

Building such a model is challenging. It requires synthesizing completions of partially visible objects, using as few as one previous observation. This is akin to novel view synthesis from a single image , but with potentially unbounded viewpoint changes. There is also the related but considerably more extreme challenge of predicting around corners. For example, as shown in Fig. 1, any future navigation trajectory passing the entrance of an unseen room requires the model to plausibly imagine the entire contents of that room (we dub this the room reveal problem). This requires generalizing from the visual, spatial and semantic structure of previously explored environments—which in our case are photo-realistic 3D captures of real indoor spaces in the Matterport3D dataset . A third problem is temporal consistency: predictions of unseen building regions should ideally be stochastic (capturing the full distribution of possible outcomes), but revisited regions should be rendered in a consistent manner to previous observations.

Towards this goal, we introduce Pathdreamer. Given one or more visual observations (consisting of RGB, depth and semantic segmentation for panoramas) of an indoor scene, Pathdreamer synthesizes high-resolution visual observations (RGB, depth and semantic segmentations) along a specified trajectory through future viewpoints, using a hierarchical two-stage approach. Pathdreamer’s first stage, Structure Generator, generates depth and semantic segmentations. Inspired by video prediction , these outputs are conditioned on a latent noise tensor capturing the stochastic information about the next observation (such as the layout of an unseen room) that cannot be predicted deterministically. The second stage’s Image Generator renders the depth and semantic segmentations as realistic RGB images using modified Multi-SPADE blocks . To maintain long-term consistency in the generated observations, both stages use back-projected 3D point cloud representations which are re-projected into image space for context .

Pathdreamer can generate plausible views for unseen indoor scenes under large viewpoint changes (see Figure 1), while also addressing the room reveal problem – in this case correctly hypothesizing that the unseen room revealed at position 2 will most likely resemble a kitchen. Empirically, using the Matterport3D dataset and 360∘360^{\circ} observations, we evaluate both stages of our model against prior work and reasonable baselines and ablations. We find that the hierarchical structure of the model is essential for predicting over large viewpoint changes, that maintaining both RGB and semantic context is required, and that prediction quality degrades gradually when we evaluate with trajectory rollouts of up to 13m (with viewpoints 2.25m apart on average).

Encouraged by these results, we investigate whether Pathdreamer’s RGB predictions can improve performance on a downstream task: Vision-and-Language Navigation (VLN), using the R2R dataset . VLN requires agents to interpret and execute natural language navigation instructions in a photorealistic 3D environment. A robust finding from previous research is that performance improves dramatically when agents can look ahead at unobserved parts of the environment while following an instruction . We show that replacing look-ahead observations with Pathdreamer predictions maintains around half of the gains, a finding we expect to have significant implications for VLN research. In summary, our main contributions include:

Proposing the study of visual world models for generic indoor environments and defining evaluation protocols and baselines for future work.

Pathdreamer, a stochastic hierarchical visual world model combining multiple, independent threads of previous work on video prediction , semantic image synthesis and video-to-video synthesis .

Extensive experiments characterizing the performance of Pathdreamer and demonstrating improved results on the downstream VLN task .

Related Work

Our work is closely related to the task of video prediction, which aims to predict the future frames of a video sequence. While some video prediction methods predict RGB video frames directly , many others use hierarchical models to first predict an intermediate representation (such as semantic segmentation) , which improves the fidelity of long-term predictions . Several approaches have also incorporated 3D point cloud representations, using projective camera geometry to explicitly infer aspects of the next frame . Inspired by this work, we adopt and combine both the hierarchical two-stage approach and 3D point cloud representations. Further, since our interest is in action-conditional world models, we provide a trajectory of future viewpoints to the model rather than assuming a constant frame rate and modeling camera motion implicitly, which is more typical in video generation .

Action-Conditional Video Prediction

Conditional video prediction to improve agent reasoning and planning has been explored in several tasks. This includes video prediction for Atari games conditioned on control inputs and 3D game environments like Doom . In robotics, action-conditional video prediction has been investigated for object pushing in tabletop settings to improve generalization to novel objects . This work has been restricted to simple environments and low-resolution images, such as 64×\times64 images of objects in a wooden box. To the best of our knowledge, we are the first to investigate action-conditional video prediction in building-scale environments with high-resolution (1024×\times512) images.

World Models and Navigation Priors

World models are an appealing way to summarize and distill knowledge about complex, high-dimensional environments. However, world models can differ in their outputs. While Pathdreamer predicts visual observations, there is also a vast literature on world models that predict compact latent representations of future states or other task-specific measurements or rewards . This includes recent work attempting to learn statistical regularities and other priors for indoor navigation—for example, by mining spatial co-occurrences from real estate video tours , learning to predict top-down belief maps over room characteristics , or learning to reconstruct house floor plans using audio and visual cues from a short video sequence . In contrast to these approaches, we focus on explicitly predicting visual observations (i.e., pixels) which are generic, human-interpretable, and apply to a wide variety of downstream tasks and applications. Further, recent work identifies a close correlation between image prediction accuracy and downstream task performance in model-based RL .

Embodied Navigation Agents

High-quality 3D environment datasets such as Matterport3D , StreetLearn , Gibson and Replica have triggered intense interest in developing embodied agents that act in realistic human environments . Tasks of interest include ObjectNav (navigating to an instance of a particular kind of object), and Vision-and-Language Navigation (VLN) , in which agents must navigate according to natural language instructions. Variations of VLN include indoor navigation , street-level navigation , vision-and-dialog navigation , VLN in continuous environments , and more. Notwithstanding considerable exploration of pretraining strategies , data augmentation approaches , agent architectures and loss functions , existing work in this space considers only model-free approaches. Our aim is to unlock model-based approaches to these tasks, using a visual world model to encode prior commonsense knowledge about human environments and thereby relieve the burden on the agent to learn these regularities. Underscoring the potential of this direction, we note that using the ground-truth environment for planning with beam search typically improves VLN success rates on the R2R dataset by 17-19% .

Novel View Synthesis

We position our work with respect to novel view synthesis . Methods for representing 3D scenes include point cloud representations , layered depth images , and mesh representations . Recently, neural radiance fields (NeRF) achieved impressive results by capturing volume density and color implicitly with a neural network. NeRF models synthesize very high quality 3D scenes, but a significant drawback for our purposes is that they require a large number of input views to render a single scene (e.g., 20–62 images per scene in ). More importantly, these models are typically trained to represent a single scene, and do not yet generalize well to unseen environments. In contrast, our problem demands generalization to unseen environments, using as little as one previous observation.

Pathdreamer

Pathdreamer is a world model that generates high-resolution visual observations from a trajectory of future viewpoints in buildings it has never observed. The input to Pathdreamer is a sequence of previous observations consisting of RGB images I1:t−1I_{1:t-1}, semantic segmentation images s1:t−1s_{1:t-1}, and depth images d1:t−1d_{1:t-1} (where the depth and segmentations could be ground-truth or estimates from a model). We assume that a corresponding sequence of camera poses T1:t−1T_{1:t-1} is available from an odometry system, and that the camera intrinsics are known or estimated. Our goal is to generate realistic RGB, semantic segmentation and depth images for a trajectory of future poses Tt,Tt+1,…,TTT_{t},T_{t+1},\dots,T_{T}, which may be provided up front or iteratively by some agent interacting with the returned observations. Note that we generate depth and segmentation because these modalities are useful in many downstream tasks. We assume that the future trajectory may traverse unseen areas of the environment, requiring the model to not only in-fill minor object dis-occlusions, but also to imagine entire room reveals (Figure 1).

Figure 2 shows our proposed hierarchical two-stage model for addressing this challenge. It uses a latent noise tensor ztz_{t} to capture the stochastic information about the next observation (e.g. the layout of an unseen room) that cannot be predicted deterministically. Given a sampled noise tensor ztz_{t}, the first stage (Structure Generator) generates a new depth image d^t\hat{d}_{t} and segmentation image s^t\hat{s}_{t} to provide a plausible high-level semantic representation of the scene, using as context the previous semantic and depth images s1:t−1s_{1:t-1}, d1:t−1d_{1:t-1}. In the second stage (Image Generator), the predicted semantic and depth images s^t\hat{s}_{t}, d^t\hat{d}_{t} are rendered into a realistic RGB image I^t\hat{I}_{t} using previous RGB images I1:t−1I_{1:t-1} as context. In each stage, context is provided by accumulating previous observations as a 3D point cloud which is re-projected into 2D using TtT_{t}.

Pathdreamer’s first stage is the Structure Generator, a stochastic encoder-decoder network for generating diverse, plausible segmentation and depth images. Like , to provide the previous observation context, we first back-project the previous segmentations s1:t−1s_{1:t-1} into a unified 3D semantic point cloud using the depth images d1:t−1d_{1:t-1} and camera poses T1:t−1T_{1:t-1}. We then re-project this point cloud back into pixel space using TtT_{t} to create sparse segmentation and depth guidance images st′s^{{}^{\prime}}_{t}, dt′d^{{}^{\prime}}_{t} which reflect the current pose.

To generate the noise tensor ztz_{t}, we take inspiration from SVG and learn a conditional prior noise distribution pψ(zt∣st′,dt′)p_{\psi}(z_{t}|s^{{}^{\prime}}_{t},d^{{}^{\prime}}_{t}). Intuitively, there are many possible scenes that may be generated for an unseen building region. We would like ztz_{t} to carry the stochastic information about the next observation that the deterministic encoder cannot capture, and we would like for the decoder to make good use of that information. During training, we encourage the first outcome by using a KL-divergence loss to force the prior distribution pψ(zt∣st′,dt′)p_{\psi}(z_{t}|s^{{}^{\prime}}_{t},d^{{}^{\prime}}_{t}) to be close to the posterior distribution ϕ(zt∣st,dt){\phi}(z_{t}|s_{t},d_{t}) which is conditioned on the ground-truth segmentation and depth images. We encourage the second outcome by providing the decoder with sampled ztz_{t} values from the posterior distribution qϕq_{\phi} (conditioned on the ground-truth outputs) during training. During inference, the latent noise ztz_{t} is sampled from the prior distribution pψp_{\psi} and the posterior distribution qϕq_{\phi} is not used. Both distributions are modeled using 3-layer CNNs that take their input from the encoder and output two channels representing μ\mu and σ\sigma to parameterize a multivariate Gaussian distribution N(μ,σ)\mathcal{N}(\mu,\sigma). As shown in Figure 3, the noise is useful in encoding diverse, plausible representations of unseen regions.

Overall, the Structure Generator is trained to minimize a joint loss consisting of a cross-entropy loss LceL_{\text{ce}} for semantic predictions, a mean absolute error term for depth predictions, and the KL-divergence term for the noise tensor:

where λce\lambda_{\text{ce}}, λd\lambda_{\text{d}}, and λKL\lambda_{\text{KL}} are weights determined by a grid search. We set these to 1, 100, and 0.5 respectively.

2 Image Generator: RGB

The Image Generator is an image-to-image translation GAN that converts the semantic and depth predictions s^t\hat{s}_{t}, d^t\hat{d}_{t} from the first stage into a realistic RGB image I^t\hat{I}_{t}. Our model architecture is based on SPADE blocks that use spatially-adaptive normalization layers to insert context into multiple layers of the network. As with our Structure Generator, we maintain an accumulating 3D point cloud containing all previous image observations. This provides a sparse RGB guidance image It′I^{\prime}_{t} when re-projected. Similar to Multi-SPADE , we insert two SPADE normalization layers into each residual block: one conditioned on the concatenated semantic and depth inputs [s^t,d^t][\hat{s}_{t},\hat{d}_{t}], and one conditioned on the RGB guidance image It′I^{\prime}_{t}. The sparsity of the RGB guidance image is handled by applying partial convolutions . In total Image Generator consists of 7 Multi-SPADE blocks, preceded by a single convolution block.

Following SPADE , the model is trained with the GAN hinge loss, feature matching loss , and perceptual loss from a pretrained VGG-19 model. During training, the generator is provided with the ground-truth segmentation image sts_{t} and ground-truth depth image dtd_{t}. Our discriminator architecture is based on PatchGAN , and takes as input the concatenation of the ground-truth image ItI_{t} or generated image I^t\hat{I}_{t}, the ground-truth depth image dtd_{t} and the ground-truth semantic image sts_{t}. The losses for the generator GG and the discriminator DD are:

where xt=(st,dt,It′)x_{t}=(s_{t},d_{t},I^{\prime}_{t}) denotes the complete set of inputs to the generator, ϕ(i)\phi^{(i)} denotes the output of the ithi^{\text{th}} layer of the pretrained VGG-19 model, D(i)D^{(i)} denotes the output of the discriminator’s ii-th layer (conditioning inputs st,dts_{t},d_{t} to the discriminator have been dropped to save space). Like the Structure Generator, the Image Generator is not pretrained.

3 Training and Inference

For training and evaluation we use Matterport3D , a dataset of 10.8k RGB-D images from 90 building-scale indoor environments. For each environment, Matterport3D also includes a textured 3D mesh which is annotated with 40 semantic classes of objects and building components. To align with downstream VLN tasks, in all experiments the RGB, depth and semantic images are 360∘360^{\circ} panoramas in equirectangular format.

Trajectories

To train Pathdreamer, we sampled 400k trajectories from the Matterport3D training environments. To define feasible trajectories, we used the navigation graphs from the Room-to-Room (R2R) dataset , in which nodes correspond to panoramic image locations, and edges define navigable state transitions. For each trajectory 5–8 panoramas were sampled, choosing the starting node and the edge transitions uniformly at random. On average the viewpoints in these trajectories are 2m apart. Training with relatively large viewpoint changes is desirable, since the model learns to synthesize observations with large viewpoint changes in a single step (without the need to incur the computational cost of generating intervening frames). However, this does not preclude Pathdreamer from generating smooth video outputs at high frame ratesSee https://youtu.be/StklIENGqs0 for our video generation results..

Training

The first and second stages of the model are trained separately. For the Image Generator, we use the Matterport3D RGB panoramas as training targets at 1024×\times512 resolution. We use the Habitat simulator to render ground-truth depth and semantic training inputs and stitch these into equirectangular panoramas. We perform data augmentation by randomly cropping and horizontally rolling the RGB panoramas, which we found essential due to the limited number of panoramas available.

To train the Structure Generator, we again used Habitat to render depth and semantic images. Since this stage does not require aligned RGB images for training, in this case we performed data augmentation by perturbing the viewpoint coordinates with a random Gaussian noise vector drawn from N(0,0.2m)\mathcal{N}(0,0.2\text{m}) independently along each 3D axis. The Structure Generator was trained with equirectangular panoramas at 512×\times256 resolution.

Inference

To avoid heading discontinuities during inference, we use circular padding on the image x-axis for both the Structure Generator and the Image Generator. The 512×\times256 resolution semantic and depth outputs of the Structure Generator are upsampled to 1024×\times512 using nearest neighbor interpolation before they are passed to the Image Generator. In quantitative experiments, we set the Structure Generator noise tensor ztz_{t} to the mean of the prior.

Experiments

For evaluation we use the paths from the Val-Seen and Val-Unseen splits of the R2R dataset . Val-Seen contains 340 trajectories from environments in the Matterport3D training split. Val-Unseen contains 783 trajectories in Matterport3D environments not seen in training. Since R2R trajectories contain 5-7 panoramas and at least 1 previous observation is given as context, we report evaluations over 1–6 steps, representing predictions over trajectory rollouts of around 2–13m (panoramas are 2.25m apart on average). See Figure 4 for an example rollout over 8.6m. We characterize the performance of Pathdreamer in comparison to baselines, ablations and in the context of the downstream task of Vision-and-Language Navigation (VLN).

A key feature of our approach is the ability to generate semantic segmentation and depth outputs, in addition to RGB. We evaluate the generated semantic segmentation images using mean Intersection-Over-Union (mIOU) and report results for:

Nearest Neighbor: A baseline without any learned components, using nearest-neighbor interpolation to fill holes in the projected semantic guidance image st′s^{\prime}_{t}.

Ours (Teacher Forcing): Structure Generator trained using the ground truth semantic and depth images as the previous observation at every time step.

Ours (Recurrent): Structure Generator trained while feeding back its own semantic and depth predictions as previous observations for the next step prediction. This reduces train-test mismatch and may allow the model to compensate for errors when doing longer roll-outs.

We also tried training the hierarchical convolutional LSTM from , but found that it frequently collapsed to a single class prediction. We attribute this to the large viewpoint changes and heavy occlusion in the training sequences; we believe this can be more effectively modeled with point cloud geometry than with a geometry-unaware LSTM.

As illustrated in Table 1, Pathdreamer performs far better than the Nearest Neighbor baseline regardless of the number of steps in the rollout or the number of previous observations used as context. As expected, performance in seen environments is higher than unseen. Perhaps surprisingly, in Figure 5(a) we show that Recurrent training improves results during longer rollouts in the training environments (Val-Seen), but this does not improve results on Val-Unseen, perhaps indicating that the error compensation learned by the Image Generator does not easily generalize.

In addition to accurate predictions, we also want generated results to be diverse. Figure 3 shows that our model can generate diverse semantic scenes by interpolating the noise tensor ztz_{t}, and that the RGB outputs closely reflect the generated semantic image. This allows us to generate multiple plausible alternatives for the same navigation trajectory.

RGB Generation

To evaluate the quality of RGB panoramas generated by the Image Generator, we compute the Fréchet Inception Distance (FID) between generated and real images for each step in the paths. We report results using the semantic images generated by the Structure Generator as inputs (i.e., our full model). To quantify the potential for uplift with better Structure Generators, we also report results using ground truth semantic segmentations as input. We compare to two ablated versions of our model:

No Semantics: The semantic and depth inputs sts_{t}, dtd_{t} are removed from the Multi-SPADE blocks.

SPADE: An ablation of the RGB inputs to the model, comprising the previous RGB image It−1I_{t-1} and the re-projected RGB guidance image It′I^{\prime}_{t}. The semantic image sts_{t} replaces It−1I_{t-1} as input to the model and the It′I^{\prime}_{t} input layers are removed from the Multi-SPADE blocks, making this effectively the SPADE model .

As shown in Table 2, SPADE performs the best in Val-Seen, indicating that the model has the capacity to memorize the training environments. In this case, RGB inputs are not necessary. However, our model performs noticeably better in Val-Unseen, highlighting the importance of maintaining RGB context in unseen environments (which is our focus). Performance degrades significantly in the No Semantics setting in both Val-Seen and Val-Unseen. We observed that without semantic inputs, the model is unable to generate meaningful images over longer horizons, which validates our two-stage hierarchical approach. These results are reflected in the FID scores, as well as qualitatively (Figure 6); Image Generator’s outputs are significantly crisper, especially over longer horizons. Due to the benefit of guidance images, the Image Generator’s textures are also generally better matched with the unseen environment, while SPADE tends to wash out textures, usually creating images of a standard style. Figure 5(b) plots performance for every setting step-by-step. FID of the Image Generator improves substantially when using ground truth semantics, particularly for longer rollouts, highlighting the potential to benefit from improvements to the Structure Generator.

2 VLN Results

Finally, we evaluate whether predictions from Pathdreamer can improve performance on a downstream visual navigation task. We focus on Vision-and-Language Navigation (VLN) using the R2R dataset . Because reaching the navigation goal requires successfully grounding natural language instructions to visual observations, this provides a challenging task-based assessment of prediction quality.

In our inference setting, at each step while moving through the environment we use a baseline VLN agent based on to generate a large number of possible future trajectories using beam search. We then rank these alternative trajectories using an instruction-trajectory compatibility model to assess which trajectory best matches the instruction. The agent then executes the first action from the top-ranked trajectory before repeating the process. We consider three different planning horizons, with future trajectories containing 1, 2 or 3 forward steps.

The instruction-trajectory compatibility model is a dual-encoder that separately encodes textual instructions and trajectories (encoded using visual observations and path geometry) into a shared latent space. To improve performance on incomplete paths, we introduce truncated paths into the original contrastive training scheme proposed in . The compatibility model is trained using only ground truth observations. However, during inference, RGB observations for future steps are drawn from three different sources:

Ground truth: RGB observations from the actual environment, i.e., look-ahead observations.

Pathdreamer: RGB predictions from our model.

Repeated pano: A simple baseline in which the most recent RGB observation is repeated in future steps.

Blank pano: A simple baseline in which blank images are provided as future observations.

Note that in all cases the geometry of the future trajectories is determined by the ground truth R2R navigation graphs. In Table 3, we report Val-Unseen results for this experiment using standard metrics for VLN: navigation error (NE), success rate (SR), shortest path length (SPL), normalized Dynamic Time Warping (nDTW) , and success weighted by normalized Dynamic Time Warping (sDTW) .

Consistent with prior work , we find that looking ahead using ground truth visual observations provides a robust performance boost, e.g., success rate increases from 44.6% with 1 planning step (top panel) to 59.3% with 3 planning steps (bottom panel). At the other extreme, the Repeated pano baseline is weak, with a success rate of just 35.7% with 1 planning step (top row). The Blank pano baseline is similar, with a success rate of 35.9%. This is not surprising: these baselines deny the compatibility model any useful visual representation of the next action, which is crucial to performance . However, increasing the planning horizon does improve performance even for the Repeated/Blank pano baselines, since the compatibility model is able to compare the geometry of alternative future trajectories. Finally, we observe that using Pathdreamer’s visual observations closes about half the gap between the Repeated pano baseline and the ground truth observations, e.g., 50.4% success with Pathdreamer vs. 40.6% and 59.3% respectively for the others. We conclude that using Pathdreamer as a visual world model can improve performance on downstream tasks, although existing agents still rely on using a navigation graph to define the feasible action space at each step. Pathdreamer is complementary to current SOTA model-based approaches, and a combination would likely lead to further boosts in VLN performance, which is worth investigating in future work.

Conclusion

Pathdreamer is a stochastic hierarchical visual world model that can synthesize realistic and diverse 360∘360^{\circ} panoramic images for unseen trajectories in real buildings. As a visual world model, Pathdreamer also shows strong promise in improving performance on downstream tasks, which we show with VLN. Most notably, we show that Pathdreamer captures around half the benefit of looking ahead at actual observations from the environment. The efficacy of Pathdreamer in the VLN task may be attributed to its ability to model fundamental constraints in the real world, and thus relieve agents from having to learn the geometry and visual and semantic structure of buildings. Applying Pathdreamer to other embodied navigation tasks such as Object-Nav , VLN-CE and street-level navigation are natural directions for future work.

References

Appendix A Qualitative Results

We present additional qualitative results generated by the Pathdreamer model, on the Val-Unseen split of the R2R dataset. These environments are not seen by the model during training, and act as a test of the generalization ability of Pathdreamer. Figure 7 presents cherry-picked examples of Pathdreamer generated sequences and Figure 8 presents randomly chosen examples.

A.2 Val-Seen Results

The Val-Seen split contains novel navigation trajectories for environments seen in the training split. As described in the main paper, the Structure Generator and Image Generator models perform significantly better on Val-Seen, since they can memorize the training environments with less generalization required. We observe that generation results maintain high fidelity even at large distances from the location of the input observation. Figure 9 presents cherry-picked examples and Figure 10 presents random examples.

A.3 Noise Interpolation

As described in the main paper, the Structure Generator is able to generate alternative, diverse, room reveals for a given scene. In areas without guidance pixels (i.e., in areas of the environment not seen from a previous view), interpolating over different noise vectors produces diverse and plausible outcomes. Figure 11 presents cherry-picked examples of Pathdreamer sequences when conditioned on different noise vectors. Figure 12 contains additional examples of randomly selected sequences conditioned on random noise vectors.

Appendix B Generated Videos

In addition to image generation, Pathdreamer is capable of generating continuous video sequences, simply by sequencing generated images with small viewpoint changes. We provide videos displaying Pathdreamer generated results for several unseen environments. We also compare Pathdreamer’s rendering quality using multiple observations to the output of the Habitat simulator using mesh-based rendering, illustrating a favorable comparison in quality. We refer readers to the YouTube linkhttps://youtu.be/StklIENGqs0 for the video results.

Appendix C Implementation Details

We trained this model with a batch size of 64, over 50 epochs. For the first 30 epochs the model is trained with teacher forcing, i.e., the previous observation that is used as input at each step is the ground-truth previous observation. During the teacher forcing stage, the number of ground truth context frames used for a path of length LL is decayed uniformly from L−1L-1 to 11. After 30 epochs, we switch to the recurrent setting in which the model’s previous prediction is used as the previous observation at each step during the rollout (similar to our setup at inference time). We use the Adam optimizer with parameters β1=0.9\beta_{1}=0.9 and β2=0.999\beta_{2}=0.999. We apply a learning rate which starts at 1e−41e^{-4} and warms up to 2e−42e^{-4} uniformly over 10 epochs.

Image Generator

This model was trained with a batch size of 128 over 500 epochs. We use the Adam optimizer with parameters β1=0.5\beta_{1}=0.5 and β2=0.999\beta_{2}=0.999, and a learning rate of 2e−42e^{-4} for both the generator and the discriminator. During training, the discriminator is trained for 2 steps for each generator training step. Following standard practice, at inference time we applied an exponential moving average (EMA) to the generator weights with 0.999 decay. For the choice of loss weights, we set λGAN=1\lambda_{\text{GAN}}=1, λVGG=0.07\lambda_{\text{VGG}}=0.07, and λFM=0\lambda_{\text{FM}}=0. We found experimentally that excluding the feature matching loss speeds up training throughput and did not have a significant effect on the results.

Evaluation Details

We compute the FID scoreWe used https://github.com/mseitzer/pytorch-fid for computing results. using 10,000 random samples for each prediction sequence step. As the R2R validation sets contain 783 and 340 sequences for Val-Unseen and Val-Seen respectively, we perform data augmentation with random horizontal roll and flips to acquire 10,000 samples.

We run evaluation every 2000 training steps, and select the best checkpoint on the Val-Seen and Val-Unseen set for reporting results for the respective split. We note that training time is generally significantly longer on Val-Seen as compared to Val-Unseen. Due to the benefit of overfitting on the training set for Val-Seen, performance continues to improve on even as generalization performance on Val-Unseen deteriorates.