Look Outside the Room: Synthesizing A Consistent Long-Term 3D Scene Video from A Single Image

Xuanchi Ren, Xiaolong Wang

Introduction

Single-image view synthesis has attracted a lot of attention in computer vision and computer graphics. It brings a photo to life by extrapolating beyond the input pixels and generating new pixels following the geometric structure of the scene. At the same time, the generated pixels need to be semantically coherent with the existing pixels. Current view synthesis methods which learn 3D geometric representation have shown encouraging results in generating high-quality novel views . However, these approaches can only generate views within a limited range of camera motion. For example, it will be very challenging for current approaches to synthesize what is outside the door of the room shown in the first row of Figure 1.

When synthesizing images with large camera view changes, we would also expect the generated images to be consistent. That is, when we are synthesizing with a path walking towards the door in a room, we hope that the surroundings of the path should not change all the time and reveal a single underlying world. To this end, we propose to solve the problem extended based on view synthesis: Given a single image of the 3D scene and a long-term camera trajectory as inputs, synthesize a consistent video as the output. For example, given a single input image of a room (first row of Figure 1), we synthesize the video on walking towards the door, going through the door, and navigating into a hallway with a painting on the wall. Solving such a task not only has wide applications in content generation and editing but also helps build a differentiable simulator for model-based planning and control in robotics.

To solve this problem, we seek help from autoregressive models , which have shown tremendous success in extrapolating the contents beyond the input image. For example, Rombach et al. proposes to use an autoregressive Transformer to implicitly perform large geometric transformation for view synthesis. To handle the uncertainty with a large transformation, the model is trained under a probabilistic framework which allows for sampling different novel views with the same camera. While generating realistic novel views even given a large transformation, it also leads to inconsistent and diverse outputs along a given trajectory due to the probabilistic sampling.

In this paper, to synthesize consistent long-term videos, we propose to leverage the autoregressive Transformer for sequential modeling in time with locality constraints. Instead of learning the autoregressive model between only two views of the scene , our work leverages the continuity in videos and perform sequential modeling with multiple video frames. Given a sequence of input images {x1,x2,...,xt−1}\{{x}_{1},{x}_{2},...,{x}_{t-1}\} and the previous camera trajectory {C2,C3,...,Ct−1}\{{C}_{2},{C}_{3},...,{C}_{t-1}\} and the camera for the future frame Ct{C}_{t}, we provide a probabilistic framework to predict the future frame via sampling from p(xt∣x1,C2,x2,C3,...,xt−1,Ct)p(x_{t}|x_{1},C_{2},x_{2},C_{3},...,x_{t-1},C_{t}). By conditioning multiple frames during sampling, it ensures the consistency between generated views and historical views. When inference with our Transformer model, we can start with a single input image and gradually increase the inputs using the predicted frames and previous frames.

However, it is very challenging to learn such a sequential model with the autoregressive Transformer, which uses self-attention to model a large number of relations between every two patches across space and time in the input video. To facilitate training, our key insight is that not every relational pair is equally important, and we can incorporate a locality constraint to guide the model to concentrate on the critical dependencies. Such locality constraints are introduced by the cameras. Intuitively, given a relative camera between two frames, we can roughly locate where the overlapping pixels are and where are the new pixels to synthesize. To incorporate this knowledge, we compute a bias using an MLP, which takes the relative camera as inputs, namely Camera-Aware Bias. We add this bias to the affinity matrix while performing the self-attention operation. In this way, each patch will have a stronger bias on depending on or attending to relevant patches connected by the camera. Empirically, we find the Camera-Aware Bias not only makes the optimization much easier but also plays a vital role in enforcing the consistency between frames during generation.

We perform our experiments on multiple datasets, including the RealEstate10K and Matterport3D , which mainly focus on 3D indoor scenes. Our model is able to synthesize new views with large camera motion, and generate a long-term video given a single image input as visualized in Figure 1. Our method not only outperforms state-of-the-art approaches on standard view synthesis metrics, but also achieves a significantly better gain when evaluating in terms of long-range future frames. We highlight our main contributions as follows:

A novel Transformer model on synthesizing a consistent long-term video given a single image and a trajectory as inputs.

A novel locality constraint using Camera-Aware Bias, which facilitates optimization during learning and enforces the consistency between generated frames.

State-of-the-art performance in view synthesis. Our method outperforms baselines by a large margin on the long-term frames.

Related Work

Novel View Synthesis. View synthesis has been a long-studied problem in computer vision and graphics. When synthesizing with multiple input views, 3D structural representations are often leveraged such as classical multi-view geometry , deep voxel representations , and neural radiance fields . Recently, researchers have also proposed to perform single-image view synthesis to bring a static photo to life . For example, Wiles et al. propose to perform view synthesis using 3D point clouds as intermediate representations. While these approaches work well with small camera changes, they cannot outpaint pixels far from the given view. To perform view synthesis with large camera changes, Rombach et al. propose a Transformer based autoregressive model. While this approach can synthesize diverse and realistic results, it cannot synthesize consistent views along a trajectory. To seek a balance, Rockwell et al. propose to leverage both 3D representation and the autoregressive models to achieve consistent view synthesis in indoor scenes with large camera changes. However, they are not able to generate a long-term future outside the door of the given room like our approach does. Our work is highly inspired with the idea of building “Infinite Images” in . Instead of performing explicit matching through a large-scale dataset, we synthesize the novel scene by sampling with a Transformer model.

Video Synthesis. Learning to synthesize a video provides an important manner to capture the dynamics of the world. Researchers have studied synthesizing videos from a random noise vector , predicting the future frames based on one or multiple previous frames , and translating one video from one modality to another . However, most video synthesis approaches do not consider the underline 3D geometry of the scene when predicting the pixels. Our work is mostly related to , which proposes an approach to synthesize a long-term video of outdoor nature environments given a single image and a trajectory as inputs. Different from them, we focus on 3D indoor scenes, which requires more structural reasoning when performing outpainting.

Image Extrapolation and Outpainting. Image outpainting synthesizes pixels beyond current input images in 2D. Specifically, our work is related to the autoregressive models , which perform outpainting the next pixels in a sequential manner. However, learning to predict pixels one by one introduces a large complexity in training and inference. Recently, Razavi et al. propose a novel representation with Vector Quantized Variational AutoEncoder (VQ-VAE), which performs autoregressive modeling in latent space instead of pixel space. This largely reduces the complexity in sequential modeling, and it enables Generative Adversarial Networks for synthesizing high-resolution images with Transformers. Our work is highly inspired by these works. Besides forwarding only image tokens to Transformers, we also add cameras as tokens in sequential modeling similar to .

Transformers. With the success of Transformer in language-modeling , it is also recently introduced into multiple recognition tasks in computer vision .Besides recognition tasks, it has also been widely used together with autoregressive models for image and video generations . However, it is still very challenging to optimize the self-attention module in Transformer when modeling a long sequence of visual tokens. In this paper, we propose to introduce a novel camera-aware bias as a locality constraint for better sequential modeling.

Method

We propose a Transformer based autoregressive model to encode and synthesize videos in a sequential manner. We will first introduce our network architecture, and our novel locality constraints using camera-aware bias for self-attention as shown in Figure 2. Then we discuss the detailed training procedure.

Given a single input image x1x_{1} together with a sequence of desirable camera transformations {C2,C3,...,CT}\{{C}_{2},{C}_{3},...,{C}_{T}\}, we hope to synthesize a sequence of images {x2,...,xT}\{{x}_{2},...,{x}_{T}\} with unconstrained length, ensuring high-quality and perceptual consistency without any 3D information.

Inspired by the success of sequential modeling in reinforcement learning , we propose to leverage a sequential of previous frames and cameras to synthesize future novel views. To achieve this, we need to accumulate the likelihood of generating {xt}t=2T\{x_{t}\}^{T}_{t=2} autoregressively as,

where τ∈[1,T]\tau\in[1,T] indicates timestep, and i∈[1,HW]i\in[1,HW] indicates the index inside a flattened image coordinate. Based on this, we can sample xtx_{t} from from the distribution:

However, different from the simple case that models only two adjacent views , sequential modeling poses two problems: (i) Self-attention alone does not ensure that the relationship between every two patches across space and time are properly modeled, given a large number of input patch tokens increase the optimization difficulty; (ii) More careful designs should be taken into account to ensure a consistent long-term synthesis. For the first problem, we propose a Camera-Aware Bias in self-attention as a locality constraint (Sec. 3.3). For the second problem, we propose several key techniques for both training and inference (Sec. 3.2 & Sec. 3.4).

2 Network Architecture

where yl,ky_{l,k} and zl,kIz^{I}_{l,k} are the kk-th tokens of yly_{l} and zlIz^{I}_{l}.

Image Decoder DD. Given the nearest indexes zlIz^{I}_{l}, we can decode it back to a high-fidelity image using the pretrained VQ-GAN decoder. zlIz^{I}_{l} is first embedded by the codebook B\mathcal{B}:

Camera Encoder ECE^{C}. For the camera model, we follow previous work to assume it as a pinhole one, such that a desired geometric transformation between two images can be determined by the intrinsic camera matrix KK, a rotation matrix RR, and a translation matrix tt.

where the camera parameters inside ClC_{l} are flattened and concatenated to shape M×1M\times 1 and ECE^{C} is a linear layer mapping from RR to RdeR^{d_{e}}.

Transformer T\mathcal{T}. Given the encoded images embeddings {zl}l=1L\{z_{l}\}_{l=1}^{L} and camera embeddings {Cle}l=2L\{C^{e}_{l}\}_{l=2}^{L}, we use a transformer to model the conditional probability in Eq. 1 in the latent space.

Decoupled Positional Embedding (P.E.). To deal with the spatial-temporal relationship, we propose a decoupled positional embedding. The tokens of the image are first calculated with the consideration of spatial information:

Now, the transformer T\mathcal{T} can be trained in an autoregressive way, denoted as:

where CE(.)CE(.) calculates the cross-entropy between the probabilities and given labels, and zIz^{I} is the corresponding indexes in the codebook B\mathcal{B}.

3 Camera-Aware Bias in Transformer

Self-attention in Transformer captures global dependency, which is a desirable property for novel view synthesis. However, since only self-attention and MLP are applied in the Transformer, there is a lack of inductive bias on 3D . When facing thousands of tokens, including information interaction between tokens across spatial and time, it is hard to capture the significant dependencies (e.g., whether two patches should be perceptually consistent) without any constraints and inductive bias.

An intuitive way to introduce 3D-aware inductive bias is to inject 3D convolutions. In ConvNets, 3D convolution serves as a 3D-aware inductive bias with the constraint on locality in both spatial and time . Thus, it may be beneficial to inject 3D convolutions into Transformers to introduce 3D-aware inductive bias. However, the motion between two adjacent views can be so large that the overlapping pixels in geometric transformation are not in the local window, which cannot be modeled by one time of convolution operation. Our key insight to solve this problem is that there is a clear relationship between frames in the video, such that the correspondence between frame xix_{i} and frame xjx_{j} is determined by relative camera transformation (K,Ri→j,ti→j)(K,R_{i\rightarrow j},t_{i\rightarrow j}). We can incorporate such spatial-temporal dependency between pixels as a 3D-aware inductive bias in Transformer. Inspired by the exploration on relative position bias in computing affinity matrix in self-attention based on image coordinate , we model the observed relationship as a novel Camera-Aware Bias in self-attention block, as shown in Figure 2 (b).

where softmax(.){\rm softmax(.)} here also takes similarity for cameras and other frames into account. For the similarity between frames and cameras, we do not apply any bias. Note that our design is applicable to causal self-attention by setting j<ij<i. By adding the Camera-Aware Bias, each patch will have a stronger bias depending on relevant patches connected by the camera, which serves as a 3D-aware inductive bias.

4 Training and Inference Details

We then introduce several key techniques for training and inference in our method.

Overlapping Iterative Modeling. In our task, we target generating long-term 3D scene video with unconstrained length TT. However, it is never possible to set the length of the training sequence LL to infinity. Thus, we choose an iterative modeling strategy. Given a single image x1x_{1}, we first generate x2,...xLx_{2},...x_{L} in an autoregressive manner. Then, instead of only using xLx_{L}, we aggregate information from x2,...xLx_{2},...x_{L} to generate xL+1x_{L+1} and so on. This overlapping iterative modeling allows us to inference for unconstrained length and maintains perceptual consistency. As we show in Sec. 4.4, this strategy is sufficient for a consistent long-term 3D scene video even with a small LL.

Error Accumulation. As pointed in , a key challenge in generating long sequences is dealing with the accumulation of errors. Even a tiny perturbation in each iteration can eventually lead to predictions outside the distribution and thus undesirable results. For an autoregressive Transformer, though we still need teacher forcing in training, we can partially simulate the error accumulation process during inference. We can first sample the predicted novel views from the predicted logits with the image decoder DD and then finetune the model with its own predicted outputs, which improves the visual quality for long-term synthesis as shown in Sec. 4.4.

Beam Search. During inference, we need to sample next frame xtx_{t} from Eq. 2. Considering the consistency, we need to choose xtx_{t} with the most likely sequences of tokens. However, decoding the most likely output sequence is exponential in the length of the output sequence, and thus it is intractable . We find that greedily take the most likely next step as the sequence leads to unnatural artifacts. Thus, we adopt a beam search strategy . Starting with the kk most likely codes in the VQ codebook as the first step in the sequence, we expand the top kk possible next steps instead of all possible in original algorithms for faster speed. Then we keep the kk most likely ones and repeat. In this way, we find a more optimal sample than a greedy search.

Experiments

In this section, we provide an empirical evaluation of our method. We demonstrate the power of our approach with an autoregressive Transformer on the view synthesis task.

Datasets. We follow the common protocol to evaluate our method on Matterport3D and RealEstate10K . Matterport3D consists of 3D models of scanned and reconstructed building-scale scenes, of which 6161 are for training and 1818 are for testing. To generate long-term episodes, we use an embodied agent in Habitat from one point in the scene to another point. In total, we render 60006000 videos for training and 500500 videos for testing. RealEstate10K is a collection of videos of footages of real estates (both indoor and outdoor). We follow to use 10,00010,000 videos for training and 5,0005,000 videos for testing.

Baselines. We compare our approach with three state-of-the-art single-image novel view synthesis work: SynSin , PixelSynth The current implementation of PixelSynth only supports 1010 discrete directions. We compare against it in Sec. 4.3 following their setting. and GeoGPT . SynSin and PixelSynth utilize point cloud as a geometric representation. We also adopt an improved version of SynSin, named SynSin-6x, provided by , which trained on larger view change. GeoGPT is a geometry-free method with probabilistic modeling between two adjacent views. We provide comparisons with additional baselines including in appendix Sec. B.1. Specifically, for comparing to the Infinite Nature method proposed by Liu et al. , as we are focusing on different settings and the code of is not publicly available, we compare to an alternative approximation as suggested by .

Implementation Details. For preprocessing, we resize all images into a resolution of H×W=256×256H\times W=256\times 256. For our experiments on both Matterport3D and RealEstate10K, we adopted the VQ-GAN from pretrained on RealEstate10K. The number of entries in the codebook B\mathcal{B} is 1638416384. For the Transformer, we adopt a GPT-like architecture with a stack of 3232 transformer blocks containing casual self-attention modules. During training, the training video clip consists L=3L=3 frames, which will be discussed in Sec. 4.4. The encoded image is of shape h×w=16×16h\times w=16\times 16 and the camera embedding is of length M=30M=30, which lead the total sequence length N=828N=828. We train our Transformer using a batch size of 1616 for 200K200K iterations with an AdamW optimizer (with β1=0.9\beta_{1}=0.9, β2=0.95\beta_{2}=0.95). We set the initial learning rate to 1.5×10−41.5\times 10^{-4} and apply a cosine-decay learning rate schedule towards zero. For beam search, we set k=3k=3. We defer more details to the supplementary material.

2 Evaluation on Short-Term View Synthesis

We evaluate our method against the baselines on short-term view synthesis in the considered range of previous novel view synthesis methods. In this setting, we adopt the standard metrics in view synthesis task: PSNR and LPIPS . PSNR measures pixel-wise differences between two images, and LPIPS measures the perceptual similarity in deep feature space. As pointed by , PSNR and LPIPS also measure consistency for a unimodal task, such as the short-term view synthesis. For both datasets, we randomly select test sequences with an input frame and 55 subsequent ground-truth frames.

Table 1 shows the quantitative results for our method. Without an intermediate geometry, our method can still outperform the methods with explicit geometric modeling in terms of short-term view synthesis. Moreover, our method also outperforms the geometry-free baseline, GeoGPT, by a large margin since this method does not ensure consistency.

3 Evaluation on Long-Term View Synthes

We then evaluate our method on the long-term view synthesis task. Prior work points out that PSNR and LPIPS are poor metrics for scene extrapolation tasks cause there are multiple possibilities for the output. Thus, for image quality, we follow to use FID , which is a distribution-level similarity measurement between generated images and real images. For consistency, we follow to conduct user study on Amazon Mechanical Turk following the A/B test protocol . Each user is presented with a video generated by our method and a baseline simultaneously during the user study. Then the user needs to choose a more consistent one. For both datasets, we randomly select test sequences with an input frame and 2020 subsequent GT frames with significant camera motion, of which each covers an extended range of footage. To compare with PixelSynth, we randomly sample an input frame and several outpainting directions to form a test sequence.

We report the quantitative comparisons in Table 2. For both image quality and consistency, our method is significantly better than other baselines, including geometry-based and geometry-free ones. This is consistent with qualitative results, as shown in Figure 3 and Figure 5. On Matterport3D, the gap is even more prominent due to the view angle changes being more significant. Notably, geometry-free methods achieve better image quality on long-term view synthesis. SynSin-6x performance is still not good, indicating that training previous methods on larger camera changes helps but does not account for the main issue. In addition, we follow past work to report PSNR and LPIPS in Table 3, which are poor measures for extrapolation tasks . For example, though SynSin-6x usually produces entirely gray results, as shown in Figure 5, its PSNR is good.

4 Ablation Study

We report some ablations of our method in terms of long-term view synthesis on the Matterport3D dataset.

Camera-Aware Bias. As shown in Figure 6, the Camera-Aware Bias improves the image quality and the consistency between frames. Table 4 also confirms this observation, indicating that bringing locality into autoregressive Transformer is critical, especially for consistency.

Decoupled positional embedding. We replace our decoupled positional embedding (P.E.) with a vanilla learnable positional embedding. As shown in Table 4, both image quality and consistency drop.

Error accumulation. As shown in Table 4, finetuning the model by stimulating error accumulation benefits the long-term view synthesis.

Length of video clips. We compare our default length of video clips with variants that modify the length during training. As shown in Table 5, the consistency improves significantly when the length increase from 22 to 33. When the length further increases to 55, the consistency remains nearly unchanged. For the image quality, there is a significant drop when we expand the length to 55. We hypothesize that the numbers of tokens are too large that the Transformer is difficult to optimize. Considering the computation resource and performance, we set the length of video clips to 33.

Discussion

Conclusion. We propose an autoregressive Transformer based model to solve novel view synthesis, especially when synthesizing long-term future in indoor 3D scenes. This method leverages a locality constraint based on the input cameras in self-attention to ensure consistency among generated frames. Our method can get superior performance in novel view synthesis compared to the state-of-the-art approaches. To conclude, we take a further step to explore the capabilities of geometry-free methods and manage to synthesize consistent high-fidelity 3D scenes.

Limitations and Future Work. Nevertheless, there are challenges remain. First, the current inference speed of the autoregressive models is slightly slower than vanilla models (details in appendix Sec. B.2). Further advancements in the autoregressive model still call for need. Second, current metrics like PSNR and LPIPS are not perfect to evaluate long-term view synthesis. New metrics for this task deserve more attention.

References

Appendix A Additional View Synthesis Results

Figure 7 and 8 provide additional long-term 3D scene videos synthesized by our methods. Our method is able to synthesize consistent novel views with large camera transformations while maintaining high fidelity.

A.2 Qualitative Comparison with Baselines

Figure 9 provide additional comparison with previous methods, including SynSin , SynSin-6x , GeoGPT and Appearance Flow . The details of the baselines are introduced in Sec. D. Our method is able to generate more consistent and clear.

A.3 Additional Visual Ablation Study

Figure 6 provides additional visual ablation study to validate the effectiveness of beam search strategy.

Appendix B Additional Experiment

Infinite Nature . Our paper focuses on indoor scenes while Infinite Nature proposed by Liu et al. focuses on nature scenes and the training code is currently not available online. Our problem is also more challenging given more structural constraints in indoor scenes. As an approximation, following the suggestion by Rockwell et al. , we compare to a method applying SynSin in a sequential manner, namely SynSin-Sequential. We report the FID results on Matterport3D in Table 6. We achieve significant improvements on image quality.

Video Antoencoder . We also compare to Video Autoencoder proposed by Lai et al. on Matterport3D. As shown in Table 6, it performs worse than our method. However, it is worthy to note that Video Autoencoder does not require camera ground-truths during training, which is a more challenging setting.

B.2 Time consumption.

We measure the average time to generate a frame during inference, as shown in Table 7.

Appendix C More Implementation Details

We provide more implementation details of our method.

Transformer. We follow GPT-2 architecture to implement our Transformer. We set the hidden dimension ded_{e} to 10241024, set the number of attention heads to 1616, and use a two-layer MLP with hidden size of 40964096 inside each transformer block. For an autoregressive Transformer, we adopt the teacher-force strategy with autoregressive masks during training to enable parallel computing.

VQ-GAN. We adopt the architecture and training strategy from https://github.com/CompVis/taming-transformers for our VQ-GAN part. And we use a downsampling factor of 1616, such that an image of resolution 256×256256\times 256 is encoded to 16×1616\times 16 tokens.

Appendix D Details of Baselines

SynSin. SynSin utilizes a point cloud as an intermediate geometric representation. We also consider a baseline, SynSin-6x, which is a version of SynSin trained on much larger view changes. However, these two baselines can only perform inpainting and can not generalize to large view changes. We adopt the official implementationhttps://github.com/facebookresearch/synsin.

PixelSynth . Based on SynSin, PixelSynth proposes to perform outpainting with the help of VQ-VAE2 and auto-regressive model . However, though it can perform outpainting, it still can not apply to the long-term view synthesis as our method does. For the implementation, we adopt the official onehttps://github.com/crockwell/pixelsynth.

GeoGPT . GeoGPT is a geometry-free method, which models two adjacent views as a probabilistic model. However, GeoGPT can not ensure consistency and does not explore the locality constraint in the autoregressive Transformer. For the implementation, we adopt the official onehttps://github.com/CompVis/geometry-free-view-synthesis.

Appearance Flow. Besides the baselines used in the main paper, we also compare our method with Appearance Flow, which is also a geometry-free baseline. Appearance Flow predicts a flow field that warps the original image into a novel view. However, this method can not work well on large camera changes since there are large missing areas after warping. We adopt the implementation provided by SynSin.