Geometry-Free View Synthesis: Transformers and no 3D Priors
Robin Rombach, Patrick Esser, Björn Ommer
Introduction
Imagine looking through an open doorway. Most of the room on the other side is invisible. Nevertheless, we can estimate how the room likely looks. The few visible features enable an informed guess about the height of the ceiling, the position of walls and lighting etc. Given this limited information, we can then imagine several plausible realizations of the room on the other side. This 3D geometric reasoning and the ability to predict what the world will look like before we move is critical to orient ourselves in a world with three spatial dimensions. Therefore, we address the problem of novel view synthesis (NVS) based on a single initial image and a desired change in viewpoint. In particular, we aim at specifically modeling large camera transformations, e.g. rotating the camera by and looking at previously unseen scenery. As this is an underdetermined problem, we present a probabilistic generative model that learns the distribution of possible target images and synthesizes them at high fidelity. Solving this task has the potential to transform the passive experience of viewing images into an interactive, 3D exploration of the depicted scene. This requires an approach that both understands the geometry of the scene and, when rendering novel views of an input, considers their semantic relationships to the visible content.
Interpolation vs. Extrapolation Recently, impressive synthesis results have been obtained with geometry-focused approaches in the multi-view setting , where not just a single but a large number of images or a video of a scene are available such that the task is closer to a view interpolation than a synthesis of genuinely novel views. In contrast, if only a single image is available, the synthesis of novel views is always an extrapolation task. Solving this task is appealing because it allows a 3D exploration of a scene starting from only a single picture.
While existing approaches for single-view synthesis make small camera transformations, such as a rotation by a few degrees, possible, we aim at expanding the possible camera changes to include large transformations. The latter necessitates a probabilistic framework: Especially when applying large transformation, the problem is underdetermined because there are many possible target images which are consistent with the source image and camera pose. This task cannot be solved with a reconstruction objective alone, as it will either lead to averaging, and hence blurry synthesis results, or, when combined with an adversarial objective, cause a significant mode-dropping when modeling the target distribution. To remedy these issues, we propose to model this task with a powerful, autoregressive transformer, trained to maximize the likelihood of the target data. Explicit vs. Implicit Geometry The success of transformers is often attributed to the fact that they enforce less inductive biases compared to convolutional neural networks (CNNs), which are biased towards local context. Relying mainly on CNNs, this locality-bias required previous approaches for NVS to explicitly model the overall geometric transformation, thereby enforcing yet another inductive bias regarding the three dimensional structure. In contrast, by modeling interactions between far-flung regions of source and target images, transformers have the potential to learn to represent the required geometric transformation implicitly without requiring such hand engineered operations. This raises the question whether it is at all necessary to explicitly include such biases in a transformer model. To address this question, we perform several experiments with varying degrees of inductive bias and find that our autoregressively trained transformer model is indeed capable of learning this transformation completely without built-in priors and can even learn to predict depth in an unsupervised fashion. To summarize our contributions, we (i) propose to learn a probabilistic model for single view synthesis that properly takes into account the uncertainties inherent in the task and show that this leads to significant benefits over previous state-of-the-art approaches when modeling large camera transformations; see Fig. Geometry-Free View Synthesis: Transformers and no 3D Priors. We (ii) also analyze the need for explicit 3D inductive biases in transformer architectures for the task of NVS with large viewpoint changes and find that transformers make it obsolete to explicitly code 3D transformations into the model and instead can learn the required transformation implicitly themselves. We also (iii) find that the benefits of providing them geometric information in the form of explicit depth maps are relatively small, and investigate the ability to recover an explicit depth representation from the layers of a transformer which has learned to represent the geometric transformation implicitly and without any depth supervision.
Related Work
Novel View Synthesis (NVS) We can identify three seminal works which illustrate different levels of reliance on geometry to synthesize novel views. describes an approach which requires no geometric model, but requires a large number of structured input views. describes a similar approach but shows that unstructured input views suffice if geometric information in the form of a coarse volumetric estimate is employed. can work with a sparse set of views but requires an accurate photogrammetric model. Subsequent work also analyzed the commonalities and trade-offs of these approaches . Ideally, an approach could synthesize novel views from a single image without having to rely on accurate geometric models of the scene and early works on deep learning for NVS explored the possibility to directly predict novel views or their appearance flows with convolutional neural networks (CNNs). However, results of these methods were limited to simple or synthetic data and subsequent works combined geometric approaches with CNNs.
Among these deep learning approaches that explicitly model geometry, we can distinguish between approaches relying on a proxy geometry to perform a warping into the target view, and approaches predicting a 3D representation that can subsequently be rendered in novel views. For the proxy geometry, relies on point clouds obtained from structure from motion (SfM) and multi-view stereo (MVS) . To perform the warping, use plane-sweep volumes, estimates depth at novel views and a depth probability volume. post-process MVS results to a global mesh and relies on per-view meshes . Other approaches learn 3D features per scene, which are associated with a point cloud or UV maps , and decoded to the target image using a CNN. However, all of these approaches rely on multi-view inputs to obtain an estimate for the proxy geometry.
Approaches which predict 3D representations mainly utilize layered representations such as layered depth images (LDIs) , multi-plane images (MPIs) and variants thereof . While this allows an efficient rendering of novel views from the obtained representations, their layered nature limits the range of novel views that can be synthesized with them. Another emerging approach represents a five dimensional light field directly with a multi-layer-perceptron (MLP), but still requires a large number of input views to correctly learn this MLP.
In the case of NVS from a single view, SfM approaches cannot be used to estimate proxy geometries and early works relied on human interaction to obtain a scene model . uses a large scale, scene-specific light field dataset to learn CNNs which predict light fields from a single image. assumes that scenes can be represented by a fixed set of planar surfaces. To handle more general scenes, most methods rely on monocular depth estimation to predict warps or LDIs . directly predicts an MPI, and a mesh. To handle disocclusions, most of these methods rely on adversarial losses, inspired by generative adversarial networks (GANs) , to perform inpainting in these regions. However, the quality of these approaches quickly degrades for larger viewpoint changes because they do not model the uncertainty of the task. While adversarial losses can remedy an averaging effect over multiple possible realizations to some degree, our empirical results show the advantages of properly modeling the probabilistic nature of NVS from a single image.
Self-Attention and Transformers The transformer is a sequence-to-sequence model that models interactions between learned representations of sequence elements by the so-called attention mechanism . Importantly, this mechanism does not introduce locality biases such as those present in e.g. CNNs, as the importance and interactions of sequence elements are weighed regardless of their relative positioning. We build our autogressive transformer from the GPT-2 architecture , i.e. multiple blocks of multihead self-attention, layer norm and position-wise MLP.
Generative Two Stage Approaches Our approach is based on work in conditional generative modeling combined with neural discrete representation learning (VQVAE) . The latter aims to learn discrete, compressed representations through either vector quantization or soft relaxation of the discrete assignment . This paradigm provides a suitable space to train autoregressive (AR) likelihood models on the latent representations and has been utilized to train generative models for hierarchical, class-conditional image synthesis , text-controlled image synthesis and music generation , and continuous analogues using VAEs or normalizing flows exist. Recently, demonstrated that adversarial training of the VQVAE improves compression while retaining high-fidelity reconstructions, subsequently enabling efficient training of an AR transformer model on the learned latent space (yielding a so-called VQGAN). We directly build on this work and use VQGANs to represent both source and target views and, when needed, depth maps. Concurrent to our work, develop an approach to NVS which uses a VQVAE and PixelCNN++ to outpaint large viewpoint changes.
Approach
To render a given image experienceable in a 3D manner, we allow the specification of arbitrary new viewpoints, including in particular large camera transformations . As a result we expect multiple plausible realizations for the novel view, which are all consistent with the input, since this problem is highly underdetermined. Consequently, we follow a probabilistic approach and sample novel views from the distribution
To solve this task, a model must explicitly or implicitly learn the 3D relationship between both images and . In contrast to most previous work that tries to solve this task with CNNs and therefore oftentimes includes an explicit 3D transformation, we want to use the expressive transformer architecture and investigate to what extent the explicit specification of such a 3D model is necessary at all.
Sec. 3.1 describes how to train a transformer model in the latent space of a VQGAN. Next, Sec. 3.2 shows how inductive biases can be build into the transformer and describes all bias-variants that we analyze. Finally, Sec. 3.3 presents our approach to extract geometric information from a transformer where no 3D bias has been explicitly specified.
Learning the distribution in Eq. (1) requires a model which can capture long-range interactions between source and target view to implicitly represent geometric transformations. Transformer architectures naturally meet these requirements, since they are not confined to short-range relations such as CNNs with their convolutional kernels and exhibit state-of-the-art performance . Since likelihood-based models have been shown to spend too much capacity on short-range interactions of pixels when modeling images directly in pixel space, we follow and employ a two-stage training. The first stage performs adversarially guided discrete representation learning (VQGAN), obtaining an abstract latent space that has proved to be well-suited for efficiently training generative transformers .
where denotes the length of the conditioning sequence. By using different functions various inductive biases can be incorporated into the architecture as described in Sec. 3.2. The transformer then processes the concatenated sequence to learn the distribution of plausible novel views conditioned on and ,
Hence, to train an autoregressive transformer by next-token prediction we maximize the log-likelihood of the data, leading to the training objective
2 Encoding Inductive Biases
Besides achieving high-quality NVS, we aim to investigate to what extent transformers depend on a 3D inductive bias. To this end, we compare approaches where a geometric transformation is built explicitly into the conditioning function , and approaches where no such transformation is used. In the latter case, the transformer itself must learn the required relationship between source and target view. If successful, the transformation will be described implicitly by the transformer.
Geometric Image Warping We first describe how an explicit geometric transformation results from the 3D relation of source and target images. For this, pixels of the source image are back-projected to three dimensional coordinates, which can then be re-projected into the target view. We assume a pinhole camera model, such that the projection of 3D points to homoegenous pixel coordinates is determined through the intrinsic camera matrix . The transformation between source and target coordinates is given by a rigid motion, consisting of a rotation and a translation . Together, these parameters specify the desired control over the novel view to be generated, i.e. .
To project pixels back to 3D coordinates, we require information about their depth , since this information has been discarded by their projection onto the camera plane. Since we assume access to only a single source view, we require a monocular depth estimate. Following by previous works , we use MiDaS in all of our experiments which require monocular depth information.
Because the target pixels obtained from the flow are not necessarily integer valued, we follow and implement by bilinearly splatting features across the four closest target pixels. When multiple source pixels map to the same target pixels, we use their relative depth to give points closer to the camera more weight—a soft variant of z-buffering.
In the simplest case, we can now describe the difference between explicit and implicit approaches in the way that they receive information about the source image and the desired target view. Here, explicit approaches receive source information warped using the camera parameters, whereas implicit approaches receive the original source image and the camera parameters themselves, i.e.
Thus, in explicit approaches we enforce an inductive bias on the 3D relationship between source and target by making this relationship explicit, while implicit approaches have to learn it on their own. Next, we introduce a number of different variants for each, which are summarized in Fig. 2.
(I) Our first explicit variant, expl.-img, warps the source image and encodes it in the same way as the target image:
(II) Inspired by previous works we include a expl.-feat variant which first encodes the original source image, and subsequently applies the warping on top of these features. We again use the VQGAN encoder to obtain
Implicit Geometric Transformations Next, we describe implicit variants that we use to analyze if transformers—with their ability to attend to all positions equally well—require an explicit geometric transformation built into the model. We use the same notation as for the explicit variants.
Compared to the other variants, this sequence is roughly times longer, resulting in twice the computational costs.
(VI) Implicit approaches offer an intriguing possibility: Because they do not need an explicit estimate of the depth to perform the warping operation , they hold the potential to solve the task without such a depth estimate. Thus, impl.-nodepth uses only camera parameters and source image—the bare minimum according to our task description.
(VII) Finally, we analyze if explicit and implicit approaches offer complementary strengths. Thus, we add a hybrid variant whose conditioning function is the sum of the ’s of expl.-emb in Eq. (3.2) and impl.-depth in Eq. (13).
3 Depth Readout
To investigate the ability to learn an implicit model of the geometric relationship between different views, we propose to extract an explicit estimate of depth from a trained model. To do so, we use linear probing , which is commonly used to investigate the feature quality of unsupervised approaches. More specifically, we assume a transformer model consisting of layers and of type impl.-nodepth, which is conditioned on source frame and transformation parameters only. Next, we specify a certain layer (where denotes the input) and extract its latent representation , corresponding to the positions of the provided source frame . We then train a position-wise linear classifier to predict the discrete, latent representation of the depth-encoder (see Sec. 3.2) via a cross-entropy objective from . Note that both the weights of the transformer and the VQGANs remain fixed.
Experiments
First, Sec. 4.1 integrates the different explicit and implicit inductive biases into the transformer to judge if such geometric biases are needed at all. Following up, Sec. 4.2 compares implicit variants to previous work and evaluates both the visual quality and fidelity of synthesized novel views. Finally, we evaluate the ability of the least biased variant, impl.-nodepth, to implicitly represent scene geometry, observing that they indeed capture such 3D information.
To investigates if transformers need (or benefit from) an explicit warping between source and target view we first compare how well the different variants from Sec. 3.2 (see also Fig. 2) can learn a probabilistic model for NVS. We then evaluate both the quality and fidelity of their samples.
To prepare, we first train VQGANs on frames of the RealEstate10K and ACID datasets, whose preparation is described in the supplementary. We then train the various transformer variants on the latent space of the respective first stage models. Note that this procedure ensures comparability of different settings within a given dataset, as the space in which the likelihood is measured remains fixed.
Comparing Density Estimation Quality A basic measure for the performance of probabilistic models is the likelihood assigned to validation data. Hence, we begin our evaluation of the different variants by comparing their (minimal) negative log-likelihood (NLL) on RealEstate and ACID. Based on the results in Tab. 1, we can identify three groups with significant performance differences on ACID: The implicit variants impl.-catdepth, impl.-depth, and impl.-nodepth and hybrid achieve the best performance, which indicates an advantage over the purely explicit variants. Adding an explicit warping as in the hybrid model does not help significantly.
Moreover, expl.-feat is unfavorable, possibly due to the features remaining fixed while training the transformer. The learnable features which are warped in variant expl.-emb obtain a lower NLL and thereby confirm the former hypothesis. Still there are no improvements of warped features over warped pixels as in variant expl.-img.
The results on RealEstate look similar but in this case the implicit variant without depth, impl.-nodepth, performs a bit worse than expl.-img. Presumably, accurate depth information obtained from a supervised, monocular depth estimation model are much more beneficial in the indoor setting of RealEstate compared to the outdoor setting of ACID.
Visualizing Entropy of Predictions The NLL measures the ability of the transformer to predict target views. The entropy of the predicted distribution over the codebook entries for each position captures the prediction uncertainty of the model. See Fig. 4 for a visualization of variant impl.-nodepth. The model is more confident in its predictions for regions which are visible in the source image. This indicates that it is indeed able to relate source and target via their geometry instead of simply predicting an arbitrary novel view.
Measuring Image Quality and Fidelity Since NLL does not necessarily reflect the visual quality of the images , we evaluate the latter also directly. Comparing predictions with ground-truth helps to judge how well the model respects the geometry. However, for large camera movements, large parts of the target image are not visible in the source view. Thus, we must also evaluate the quality of the content imagined by the model, which might be fairly different from that of the ground-truth, since the latter is just one of many possible realizations of the real-world.
To evaluate the image quality without a direct comparison to the ground-truth, we report FID scores . To evaluate the fidelity to the ground-truth, we report the low-level similarity metrics SSIM and PSNR, and the high-level similarity metric PSIM , which better represents human assessments of visual similarity. Tab. 1 contains the results for RealEstate10K and ACID. In general, they reflect the findings from the NLL values: Image quality and fidelity of implicit variants with access to depth are superior to explicit variants. The implicit variant without depth (impl.-nodepth) consistently achieves the same good FID scores as the implicit variants with depth (impl.-catdepth & impl.-depth), but cannot achieve quite the same level of performance in terms of reconstruction fidelity. However, it is on par with the explicit variants, albeit requiring no depth supervision.
2 Comparison to Previous Approaches
Next, we compare our best performing variants impl.-depth and impl.-nodepth to previous approaches for NVS: 3DPhoto , SynSin and InfNat . 3DPhoto has been trained on MSCOCO to work on arbitrary scenes, whereas SynSin and InfNat have been trained on RealEstate and ACID, respectively.
To assess the effect of formulating the problem probabilistically, we introduce another baseline to compare probabilistic and deterministic models with otherwise equal architectures. Specifically, we use the same VQGAN architecture as described in Sec. 3.1. However, it is not trained as an autoencoder, but instead the encoder receives the warped source image , and the decoder predicts the target image . This model, denote by expl.-det, represents an explicit and deterministic baseline. Finally, we include the warped source image itself as a baseline denoted by MiDaS .
Utilizing the probabilistic nature of our model, we analyze how close we can get to a particular target image with a fixed amount of samples. Tab. 2 and 3 report the reconstruction metrics with 32 samples per target. The probabilistic variants consistently achieve the best values for the similarity metrics PSIM, SSIM and PSNR on RealEstate, and are always among the best three on ACID, where expl.-det achieves the best PSIM values and the second best PSNR values. We show the reconstruction metrics on RealEstate as a function of the number of samples in Fig. 3. With just four samples, the performance of impl.-depth is better than all other approaches except for the SSIM values of 3DPhoto , which are overtaken by impl.-depth with 16 samples, and do not saturate with 32 samples, which demonstrates the advantages of a probabilistic formulation of NVS.
These results should be considered along with the competitive FID scores in Tab. 2 and 3 (where the implicit variants always constitute the best and second best value) and the qualitative results in Fig. 5 and 6, underlining the high quality of our synthesized views. It is striking that IS assigns the best scores to 3DPhoto and MiDaS , which contain large and plain regions of gray color in regions where the source image does not provide information about the content. Where the monocular depth estimation is accurate, 3DPhoto shows good results but it can only inpaint small areas. SynSin and InfNat can fill larger areas but, for large camera motions, their results become blurry and a similar observation holds for expl.-det. The probabilistic variants impl.-depth and impl.-nodepth consistently produce plausible results which are largely consistent with the source image, although small details sometimes differ. This shows that only the probabilistic variants are able to synthesize high quality images for large camera changes.
3 Probing for Geometry
Based on the experiments in Sec. 4.1 and Sec. 4.2, which showed that the unbiased variant impl.-nodepth is mostly on-par with the others, we investigate the question whether this model is able to develop an implicit 3D “understanding” without explicit 3D supervision. To do so, we perform linear probing experiments as described in Sec. 3.3.
Fig. 7 plots the negative cross-entropy loss and the negative PSIM reconstruction error of the recovered depth maps against the layer depth of the transformer model. Both metrics are consistent and quickly increase when probing deeper representations of the transformer model. Furthermore, both curves exhibit a peak for (i.e. after the third self-attention block) and then slowly decrease with increasing layer depth. The depth maps obtained from this linear map resemble the corresponding true depth maps qualitatively well as shown in Fig. 7. This figure demonstrates that a linear estimate of depth only becomes possible through the representation learned by the transformer () but not by the representation of the VQGAN encoder (). We hypothesize that, in order to map an input view onto a target view, the transformer indeed develops an implicit 3D representation of the scene to solve its training task.
Discussion
We have introduced a probabilistic approach based on transformers for novel view synthesis from a single source image with strong changes in viewpoint. Comparing various explicit and implicit 3D inductive biases for the transformer showed that explicitly using a 3D transformation in the architecture does not help their performance significantly. However, removing inductive biases also comes at a price.
Without priors on camera movements or warping layers, the architecture must be able to take relationships between arbitrary positions into account, which requires a compressed representation. In our experiments, compression artifacts dominate the error for small viewpoint changes. Avoiding them increases computational costs (see Sec. E). Synthesizing two views from the same image generally results in two incompatible realizations. However, we can run our approach iteratively. When synthesizing continuous trajectories, sampling still leads to flickering but this can be alleviated with deterministic sampling (see Sec. A).
To conclude, our approach is not a final solution to novel view synthesis, but an important step towards synthesizing large camera changes and understanding the need for 3D priors. Our results demonstrate significant improvements over existing approaches, and even with no depth information as input our model learns to infer depth within its internal representations. Future works should explore how to combine these capabilities and insights with improved performance at synthesizing stable high-resolution trajectories.
Geometry-Free View Synthesis Transformers and no 3D Priors
In this supplementary, we provide additional results obtained with our models in Sec. A. Sec. B summarizes models, architectures and hyperparameters that were used in the main paper. After describing details on the training and test data in Sec. C and on the uncertainty evaluations via the entropy in Sec. D, Sec. E concludes the supplementary material with a brief discussion of the compression artifacts introduced by the usage of the VQGAN as the compression model.
Appendix A Additional Results
Fig. 9 shows a preview of the videos available at https://git.io/JOnwn, which demonstrate an interface for interactive 3D exploration of images. Starting from a single image, a user can use keyboard and mouse to move the camera freely in 3D. To provide orientation, we warp the starting image to the current view using a monocular depth estimate (corresponding to the MiDaS baseline in Sec. 4.2). This enables a positioning of the camera with real-time preview of the novel view. Once a desired camera position has been reached, the spacebar can be pressed to autoregressively sample a novel view with our transformer model.
In the videos available at https://git.io/JOnwn, we use camera trajectories from the test sets of RealEstate10K and ACID, respectively. The samples are produced by our impl.-depth model, and for an additional visual comparison, we also include results obtained with the same methods that we compared to in Sec. 4.2.
Small Viewpoint Changes & Continuous Trajectories
For very small viewpoint changes, distortions due to compression dominate the error of our approach (left of Fig. 10, where the x-axis uses PSIM between source and target view as a proxy for the difficulty/magnitude of viewpoint change). Still, our approach outperforms previous approaches when considering the average over small, medium and large viewpoint changes (solid lines at the left) and the gap quickly increases when considering more difficult examples (solid lines, right). See also Sec. E on the trade-off between distortion caused by compression and computational efficiency. Our approach can also be applied to small viewpoint changes and the generation of continuous, consistent trajectories; see Fig. 11. Note that the ability of previous approaches to synthesize small viewpoint changes well does not enable high-quality synthesis of even moderately long trajectories.
Transformer Variants Over the Course of Training
Fig. 12 reports the negative log-likelihood (NLL) over the course of training on RealEstate and ACID, respectively. The models overfit to the training split of ACID early which makes training on ACID much quicker and thus allows us to perform multiple training runs of each variant with different initializations. This enables an estimate of the significance of the results by computing the mean and standard deviation over three runs (solid line and shaded area in Fig. 12, respectively).
Additional Qualitative Results
For convenience, we also include additional qualitative results directly in this supplementary. Fig. 15 and 17 show additional qualitative comparisons on RealEstate10K and ACID, as in Fig. 5 and 6 of the main paper. Fig. 16 and 18 demonstrate the diversity and consistency of samples by showing them along with their pixel-wise standard deviation. Fig. 19 contains results from the depth-probing experiment of Sec. 4.3, and Fig. 13 from the entropy visualization of Sec. 4.1.
Appendix B Architectures & Hyperparameters
In contrast to the global attention operation, the MLP is applied position-wise.
Note that non-conditioning elements, i.e. the last elements, are masked autoregressively . For all experiments, we use an embedding dimensionality , transformer blocks, 16 attention heads, two-layer MLPs with hidden dimensionalities of and a codebook of size . This setting results in a transformer with parameters. We train the model using the AdamW optimizer (with , ) and apply weight decay of 0.01 on non-embedding parameters. We train for steps, where we first linearly increase the learning rate from to during the first steps, and then apply a cosine-decay learning rate schedule towards zero.
VQGAN
Other models
For monocular depth estimates, we use MiDaS v2.1see https://github.com/intel-isl/MiDaS. We use the official implementations and pretrained models for the comparison with 3DPhoto see https://github.com/vt-vl-lab/3d-photo-inpainting, SynSin see https://github.com/facebookresearch/synsin/ and InfNat see https://github.com/google-research/google-research/tree/master/infinite_nature.
Appendix C Training and Testing Data
Training our conditional generative model requires examples consisting of . Such training pairs can be obtained via SfM applied to image sequences, which provides poses for each frame with respect to an arbitrary world coordinate system. For two frames from the sequence, the relative transformation is then given by and . However, the scale of the camera translations obtained by SfM is also arbitrary, and without access to the full sequence, underspecified.
To train the model and to meaningfully compute reconstruction errors for the evaluation, we must resolve this ambiguity. To do this, we also triangulate a sparse set of points for each sequence using COLMAP . We then compute a monocular depth estimate for each image using MiDaS and compute the optimal affine scaling to align this depth estimate with the scale of the camera pose. Finally, we normalize depth and camera translation by the minimum depth estimate.
All qualitative and quantitative results are obtained on a subset of the test splits of RealEstate10K and ACID , consisting of 564 source-target pairs, which have been selected to contain medium-forward, large-forward, medium-backward and large-backward camera motions in equal parts. We will make this split publicly available along with our code.
Since our stated goal is to model large camera transformations, our evaluation focuses on this ability and thus differs from the evaluation of SynSin, which is biased to small changes; see Tab. 4: With the SynSin evaluation (small cam-) we reproduce the officially reported numbers and our choice to evaluate at the original aspect ratio (at ) has minor effects. The last three rows show the deterioration of metrics if we remove biases to small changes : We (i) evaluate all test pairs, not just the better of two views (w/o best of 2), (ii) remove of test pairs that contain no camera change at all (w/o src tgt) and (iii) add of larger viewpoint changes (w/ large). The main paper reports results at medium and large viewpoint changes, but we include an additional analysis with small changes in Fig. 10.
Appendix D Details on Entropy Evaluation
As discussed in Sec. 4.1, the relationship between a source view and a target view can be quantified via the entropy of the probability distribution that the transformer assigns to a target view , given a source frame , camera transformation and conditioning function . More specifically, we first encode target, camera and source via the encoder and the conditioning function (see Sec. 3.1), i.e. and . Next, for each element in the sequence , the (trained) transformer assigns a probability conditioned on the source and camera:
Reshaping to the latent dimensionality and bicubic upsampling to the input’s size then produces the visualizations of transformer entropy as in Fig. 4 and Fig. 13. Note that this approach quantifies the transformers uncertainty/surprise from a single example only and does not need to be evaluated on multiple examples.
Appendix E Faithful Reconstructions/Compression Artifacts
Efficient training of the transformer models is enabled by the strong compression achieved with the VQGAN, which to some degree introduces artifacts but allows to trade compute requirements for reconstruction quality. Larger discrete codes improve fidelity (see Fig. 14, Tab. 5) but a 4 larger code leads to approximately 16 larger costs when training the transformer (-complexity of attention). Reducing such artifacts is thus a matter of scaling up hardware or training time. Additionally, it would also increase the time required to sample a novel view, which is currently seconds.