Monocular Dynamic View Synthesis: A Reality Check

Hang Gao, Ruilong Li, Shubham Tulsiani, Bryan Russell, Angjoo Kanazawa

Introduction

Dynamic scenes are ubiquitous in our everyday lives – people moving around, cats purring, and trees swaying in the wind. The ability to capture 3D dynamic sequences in a “casual” manner, particularly through monocular videos taken by a smartphone in an uncontrolled environment, will be a cornerstone in scaling up 3D content creation, performance capture, and augmented reality. Work partially done as part of HG’s internship at Adobe.

Recent works have shown promising results in dynamic view synthesis (DVS) from a monocular video . However, upon close inspection, we found that there is a discrepancy between the problem statement and the experimental protocol employed. As illustrated in Figure 1, the input data to these algorithms either contain frames that “teleport” between multiple camera viewpoints at consecutive time steps, which is impractical to capture from a single camera, or depict quasi-static scenes, which do not represent real-life dynamics.

In this paper, we provide a systematic means of characterizing the aforementioned discrepancy and propose a better set of practices for model fitting and evaluation. Concretely, we introduce effective multi-view factors (EMFs) to quantify the amount of multi-view signal in a monocular sequence based on the relative camera-scene motion. With EMFs, we show that the current experimental protocols operate under an effectively multi-view regime. For example, our analysis reveals that the aforementioned practice of camera teleportation makes the existing capture setup akin to an Olympic runner taking a video of a moving scene without introducing any motion blur.

The reason behind the existing experimental protocol is that monocular DVS is a challenging problem that is also hard to evaluate. Unlike static novel-view synthesis where one may simply evaluate on held-out views of the captured scene, in the dynamic case, since the scene changes over time, evaluation requires another camera that observes the scene from a different viewpoint at the same time. However, this means that the test views often contain regions that were never observed in the input sequence. Camera teleportation, i.e., constructing a temporal sequence by alternating samples from different cameras, addresses this issue at the expense of introducing multi-view cues, which are unavailable in the practical single-camera capture.

We propose two sets of metrics to overcome this challenge without the use of camera teleportation. The first metric enables evaluating only on pixels that were seen in the input sequence by computing the co-visibility of every test pixel. The proposed co-visibility mask can be used to compute masked image metrics (PSNR, SSIM and LPIPS ). While the masked image metrics measure the quality of rendering, they do not directly measure the quality of the inferred scene deformation. Thus, we also propose a second metric that evaluates the quality of established point correspondences by the percentage of correctly transferred keypoints (PCK-T) . The correspondences may be evaluated between the input and test frames or even within the input frames, which enable evaluation on sequences that are captured with only a single camera.

Related work

Traditional NR-SfM tackles the task of dynamic 3D inference by fitting parametric 3D morphable models , or fusing non-parametric depth scans of generic dynamic scenes . All of these approaches aim to recover accurate surface geometry at each time step and their performance is measured with ground truth 3D geometry or 2D correspondences with PCK when such ground truth is not available. In this paper, we analyze recent dynamic view synthesis methods whose goal is to generate a photo-realistic novel view. Due to their goal, these methods do not focus on evaluation against ground truth 3D geometry, but we take inspiration from prior NR-SfM works to evaluate the quality of the inferred 3D dynamic representation based on correspondences. We also draw inspiration from previous NR-SfM work that analyzed camera/object speed and 3D reconstruction quality .

Monocular dynamic neural radiance fields (dynamic NeRFs).

Dynamic NeRFs reconstruct moving scenes from multi-view inputs or given pre-defined deformation template . In contrast, there is a series of recent works that seek to synthesize high-quality novel views of generic dynamic scenes given a monocular video . These works can be classified into two categories: a deformed scene is directly modeled as a time-varying NeRF in the world space or as a NeRF in canonical space with a time-dependent deformation . The evaluation protocol in these works inherit from the original static-scene NeRF that quantify the rendering quality of held-out viewpoints using image metrics, e.g., PSNR. However, in dynamic scenes, PSNR from an unseen camera view may not be meaningful since the novel view may include regions that were never seen in the training view (unless the method can infer unseen regions using learning based approaches). Existing approaches resolve this issue by incorporating views from multiple cameras during training, which we show results in an effectively multi-view setup. We introduce metrics to measure the difficulties of an input sequence, a monocular dataset with new evaluation protocol and metrics, which show that existing methods have a large room for improvement.

Effective multi-view in a monocular video

We consider the problem of dynamic view synthesis (DVS) from a monocular video. A monocular dynamic capture consists of a single camera observing a moving scene. The lack of simultaneous multi-view in the monocular video makes this problem more challenging compared to the multi-view setting, such as reconstructing moving people from multiple cameras .

Contrary to the conventional perception that the effect of multi-view is binary for a capture (single versus multiple cameras), we show that it can be characterized on a continuous spectrum. Our insight is that a monocular sequence contains effective multi-view cues when the camera moves much faster than the scene, though technically the underlying scene is observed only once at each time step.

Although a monocular video only sees the scene from one viewpoint at a time, depending on the capture method, it can still contains cues that are effectively similar to those captured by a multi-view camera rig, which we call as effective multi-view. As shown in Figure 2, when the scene moves significantly slower than the camera (to the far right end of the axis), the same scene is observed from multiple views, resulting in multi-view capture. In this case, DVS reduces to a well-constrained multi-view stereo problem at each time step. Consider another case where the camera moves significantly faster compared to the scene so that it observes roughly the same scene from different viewpoints. As the camera motion approaches the infinity this again reduces the monocular capture to a multi-view setup. We therefore propose to characterize the amount of multi-view cues by the relative camera-scene motion.

2 Quantifying effective multi-view in a monocular video

For practicality, we propose two metrics, referred to as effective multi-view factors (EMFs). The first metric, full EMF Ω\Omega is defined as the relative ratio between the motion magnitude of the camera to the scene, which in theory characterizes the effective multi-view perfectly, but in practice can be expensive and challenging to compute. The second metric, angular EMF ω\omega is defined as the camera angular velocity around the scene look-at point, which only considers the camera motion; while approximate, it is easy to compute and characterizes object-centric captures well.

where the denominator xt+1−xt\mathbf{x}_{t+1}-\mathbf{x}_{t} denotes the 3D scene flow and the the numerator ot+1−ot\mathbf{o}_{t+1}-\mathbf{o}_{t} denotes the 3D camera motion, both over one time step forward. The 3D scene flow can be estimated via the 2D dense optical flow field and the metric depth map when available, or monocular depth map from off-the-shelf approaches in the general case. Please see the Appendix for more details. Note that Ω\Omega in theory captures the effective multi-view factor for any sequence. However, in practice, 3D scene flow estimation is an actively studied problem and may suffer from noisy or costly predictions.

Angular EMF ω𝜔\omega: camera angular velocity.

We introduce a second metric ω\omega that is easy to compute in practice. We make an additional assumption that the capture has a single look-at point in world space, which often holds true, particularly for captures involving a single centered subject. Specifically, given a look-at point a\mathbf{a} by triangulating the optical-axes of all cameras (as per ) and the frame rate NN, the camera angular velocity ω\omega is computed as a scaled expectation,

Note that even though ω\omega only considers the camera motion, it is indicative of effective multi-view in the majority of existing captures, which we describe in Section 4.1.

For both Ω\Omega and ω\omega, the larger the value, the more multi-view cue the sequence contains. For future works introducing new input sequences, we recommend always reporting angular EMF for its simplicity and reporting full EMF when possible. Next we inspect the existing experimentation practices under the lens of effective multi-view.

Towards better experimentation practice

In this section, we reflect on the existing datasets and find that they operate under the effective multi-view regime, with either teleporting camera motion or quasi-static scene motion. The reason behind the existing protocol is that monocular DVS is challenging from both the modeling and evaluation perspective. While the former challenge is well known, the latter is less studied, as we expand below. To overcome the existing challenge in the evaluation and enable future research to experiment with casually captured monocular video, we propose a better toolkit, including two new metrics and a new dataset of complex motion in everyday lives.

Visualizing the actual training data shown in Figure 1 reveals that existing datasets feature non-practical captures of either (1) teleporting/fast camera motion or (2) quasi-static/slow scene motion. The former is not representative of practical captures from a hand-held camera, e.g., a smartphone, while the latter is not representative of moving objects in daily life. Note that, out of the 2323 multi-camera sequences that these prior works used for quantitative evaluation, 2222 have teleporting camera motion, and 11 has quasi-static scene motion – the Curls sequence shown at the 5th5^{\text{th}} column in Figure 1. All 1313 single-camera sequences from HyperNeRF used for qualitative evaluation have quasi-static scene motion.

The four datasets also share a similar data protocol for generating effective multi-view input sequences from the original multi-camera rig capture. In Figure 4, we illustrate the teleporting protocol used in Nerfies and HyperNeRF as a canonical example. They sample alternating frames from two physical cameras (left and right in this case) mounted on a rig to create the training data. NSFF samples alternating frames from 2424 cameras based on the data released from Yoon et al. . D-NeRF experiments on synthetic dynamic scenes where cameras are randomly placed on a fixed hemisphere at every time step, in effect teleporting between 100100-200200 cameras. We encourage you to visit our project page to view the input videos from these datasets.

Existing works adopt effective multi-view capture for two reasons. First, it makes monocular DVS more tractable. Second, it enables evaluating novel view on the full image, without worrying about the visibility of each test pixel, as all camera views were visible during training. We show this effect in Figure 5. When trained with camera teleportation (3rd3^{\text{rd}} column), the model can generate a high-quality full image from the test view. However, when trained without camera teleportation (4th4^{\text{th}} column), the model struggles to hallucinate unseen pixels since NeRFs are not designed to predict completely unseen portions of the scene, unless they are specifically trained for generalization . Next, we propose new metrics that enable evaluation without using camera teleportation. Note that when the model is trained without camera teleportation, the rendering quality also degrades, which we also evaluate.

2 Our proposed metrics

While the existing setup allows evaluating on the full rendered image from the test view, the performance under such evaluation protocol, particularly with teleportation, confounds the efficacy of the proposed approaches and the multi-view signal present in the input sequence. To evaluate with an actual monocular setup, we propose two new metrics that evaluate only on seen pixels and measure the correspondence accuracy of the predicted deformation.

Existing works evaluate DVS models with image metrics on the full image, e.g., PSNR, SSIM and LPIPS , following novel-view synthesis evaluation on static scenes . However, in dynamic scenes, particularly for monocular capture with multi-camera validation, the test view contains regions that may not have been observed at all by the training camera. To circumvent this issue without resorting to camera teleportation, for each pixel in the test image, we propose co-visibility masking, which tests how many times a test pixel has been observed in the training images. Specifically, we use optical flow to compute correspondences between every test image and the training images, and only keep test pixels that have enough correspondences in the training images via thresholding. This results in a mask, illustrated in Figure 5, which we use to confine the image metrics. We follow the common practice from the image generation literature and adopt masked metrics, mPSNR and mLPIPS . Note that NSFF adopts similar metrics but for evaluating the rendering quality on foreground versus background regions. We additionally report mSSIM by partial convolution , which only considers seen regions during its computation. More details are in the Appendix. Using masked image metrics, we quantify the performance gap in rendering when a model is trained with or without multi-view cues in Section 5.1.

Percentage of correctly transferred keypoints (PCK-T).

Correspondences lie at the heart of traditional non-rigid reconstruction , which is overlooked in the current image-based evaluation. We propose to evaluate 2D correspondences across training frames with the percentage of correctly transferred keypoints (PCK-T) , which directly evaluates the quality of the inferred deformation. Specifically, we sparsely annotate 2D keypoints across input frames to ensure that each keypoint is fully observed during training. For correspondence readout from existing methods, we use either root finding or scene flow chaining. Please see the Appendix for details on our keypoint annotation, correspondence readout, and metric computation. As shown in Figure 6, evaluating correspondences reveal that high quality image rendering does not necessarily result in accurate correspondences, which indicates issues in the underlying surface, due to the ambiguous nature of the problem.

3 Proposed iPhone dataset

Existing datasets can be rectified by removing camera teleportation and evaluated using the proposed metrics, as we do in Section 5.1. However, even after removing camera teleportation, the existing datasets are still not representative of practical in-the-wild capture. First, the existing datasets are limited in motion diversity. Second, the evaluation baseline in existing datasets is small, which can hide issues in incorrect deformation and resulting geometry. For these reasons, we propose a new dataset called the iPhone dataset shown in Figure 7. In contrast to existing datasets with repetitive object motion, we collect 1414 sequences featuring non-repetitive motion, from various categories such as generic objects, humans, and pets. We deploy three cameras for multi-camera capture – one hand-held moving camera for training and two static cameras of large baseline for evaluation. Furthermore, our iPhone dataset comes with metric depth from the lidar sensors, which we use to provide ground-truth depth for supervision. In Section 5.2, we show that depth supervision, together with other regularizations, is beneficial for training DVS models. Please see the Appendix for details on our multi-camera capture setup, data processing, and more visualizations.

Reality check: re-evaluating the state of the art

In this section, we conduct a series of empirical studies to disentangle the recent progress in dynamic view synthesis (DVS) given a monocular video from effective multi-view in the training data. We evaluate current state-of-the-art methods when the effective multi-view factor (EMF) is low.

We consider the following state-of-the-art approaches for our empirical studies: NSFF , Nerfies and HyperNeRF . We choose them as canonical examples for other approaches , discussed in Section 2. We also evaluate time-conditioned NeRF (T-NeRF) as a common baseline . Unlike the state-of-the-art methods, it is not possible to extract correspondences from a T-NeRF. A summary of these methods can be found in the Appendix.

Datasets.

Masked image and correspondence metrics.

Following Section 4.2, we evaluate co-visibility masked image metrics and the correspondence metric. We report masked image metrics: mPSNR, mSSIM , and mLPIPS . We visualize the rendering results with the co-visibility mask. For the correspondence metric, we report the percentage of correctly transferred keypoints (PCK-T) with threshold ratio α=0.05\alpha=0.05. Additional visualizations of full image rendering and inferred correspondences can be found in the Appendix.

Implementation details.

We consolidate Nerfies and HyperNeRF in one codebase using JAX . Compared to the original official code releases, our implementation aligns all training and evaluation details between models and allows correspondence readout. Our implementation reproduces the quantitative results in the original papers. We implement T-NeRF in the same codebase. For NSFF , we tried both the official code base and a public third-party re-implementation , where the former fails to converge on our proposed iPhone dataset while the latter works well. We thus report results using the third-party re-implementation. However, note that both the original and the third-party implementation represent the dynamic scene in normalized device coordinates (NDC). As NDC is designed for forward-facing but not considered inward-facing scenes, layered artifacts may appear due to its log-scale sampling rate in the world space, as shown in Figure 9. More details about aligning the training procedure and remaining differences are provided in the Appendix. Code, pretrained models, and data are available on the project page.

1 Reality check on the Nerfies-HyperNeRF dataset

We first study the impact of effective multi-view on the Nerfies-HyperNeRF dataset. In this experiment, we rectify the effective multi-view sequences by only using the left camera during training as opposed to both the left and right cameras, illustrated in Figure 4. We denote the original setting as “teleporting” and the rectified sequences as “non-teleporting”. We train all approaches under these two settings with the same held-out validation frames and same set of co-visibility masks computed from common training frames. In Figure 8 (Top), all methods perform better across all metrics when trained under the teleporting setting compared to the non-teleporting one, with the exception of PCK-T for NSFF. We conjecture that this is because that NSFF has additional optical flow supervision, which is more accurate without camera teleportation. In Figure 8 (Bottom), we show qualitative results using Nerfies (we include visualizations of the other methods in the Appendix). Without effective multi-view, Nerfies fails at modeling physically plausible shape for broom and wires. Our results show that the effective multi-view in the existing experimental protocol inflates the synthesis quality of prior methods, and that truly monocular captures are more challenging.

Benchmark results without camera teleportation.

In Table 2 and Figure 9, we report the quantitative and qualitative results under the non-teleporting setting. Note that our implementation of the T-NeRF baseline performs the best among all four evaluated models in terms of mPSNR and mSSIM. In Figure 9, we confirm this result since T-NeRF renders high-quality novel view for both sequences. HyperNeRF produces the most photorealistic renderings, measured by mLPIPS. However it also produces distorted artifacts that do not align well with the ground truth (e.g., the incorrect shape in the Chicken sequence).

2 Reality check on the proposed iPhone dataset

We find that existing methods perform poorly out-of-the-box on the proposed iPhone dataset with more diverse and complex real-life motions. In Figure 10 (Bottom), we demonstrate this finding with HyperNeRF for it achieves the highest mLPIPS metric on the Nerfies-HyperNeRF dataset. Shown in the 3rd3^{\text{rd}} column, HyperNeRF produces visually implausible results with ghosting effects. Thus we explored incorporating additional regularizations from recent advances in neural rendering. Concretely, we consider the following: (+B) random background compositing ; (+D) a depth loss on the ray matching distance ; and (+S) a sparsity regularization for scene surface . In Figure 10 (Top), we show quantitative results from the ablation. In Figure 10 (Bottom), we show visualizations of the impact of each regularization. Adding additional regularizations consistently boosts model performance. While we find the random background compositing regularizations particularly helpful, extra depth supervision and surface regularization further improve the quality, e.g., the fan region of the paper windmill.

Benchmarked results.

Discussion and recommendation for future works

In this work, we expose issues in the common practice and establish systematic means to calibrate performance metrics of existing and future works, in the spirit of papers like . We provide initial attempts toward characterizing the difficulty of a monocular video for dynamic view synthesis (DVS) in terms of effective multi-view factors (EMFs). In practice, there are other challenging factors such as variable appearance, lighting condition, motion complexity and more. We leave their characterization for future works. We recommend future works to visualize the input sequences and report EMFs when demonstrating the results. We also recommend future works to evaluate the correspondence accuracy and strive for establishing better correspondences for DVS.

We would like to thank Zhengqi Li and Keunhong Park for valuable feedback and discussions; Matthew Tancik and Ethan Weber for proofreading. We are also grateful to our pets: Sriracha, Haru, and Mochi, for being good during capture. This project is generously supported in part by the CONIX Research Center, sponsored by DARPA, as well as the BDD and BAIR sponsors.

References

Appendix A Outline

In this Appendix, we describe in detail the following:

Computation for effective multi-view factors (EMFs) in Section B.

Computation for co-visibility mask and masked image metrics in Section C.

Summary of existing works and correspondence readout in Section D.

Summary of the capture setup and data processing for our iPhone dataset in Section E.

Summary of the implementation details and remain differences in Section F.

Additional results on the impact of effective multi-view in Section G.

Additional results on per-sequence performance breakdown in Section H.

Additional results on novel-view synthesis in Section I.

Additional results on inferred correspondence in Section J.

For better demonstration, we strongly recommend visiting our project page for videos of the capture and result visualizations.

Appendix B Computation for effective multi-view factors (EMFs)

To quantify the amount of effective multi-view in a sequence by the camera and scene motion magnitude, we propose two metrics as effective multi-view factors (EMFs), i.e., the Full EMF Ω\Omega and the angular EMF ω\omega. Note that we design our metrics to be scale-agnostic such that we can compare them across different sequences of different world scales.

We are interested in the relative scale of the camera motion compared to the object. Recall that we define Ω\Omega as the expected ratio over all visible pixels over time,

The numerator is trivially computable given the camera information. We thus focus on the denominator, i.e., the foreground 3D scene flow xt+1−xt\mathbf{x}_{t+1}-\mathbf{x}_{t}.

We estimate 3D scene flow by combining the known cameras, dense 2D optical flow, and per-frame depth maps. We estimate the 2D optical flow using RAFT . When metric depth is not available, e.g., on previous datasets , we use DPT for monocular depth estimation. Additionally, we need a foreground mask for the object, which we obtain through a video segmentation network . For each pixel location ut\mathbf{u}_{t} at time tt in the foreground mask, we can compute its 3D position xt\mathbf{x}_{t} by back-projection with the depth ztz_{t}. We then get the 2D pixel correspondence at time t+1t+1 by simply following the 2D optical flow ut+1=ut+ft→t+1(ut)\mathbf{u}_{t+1}=\mathbf{u}_{t}+\mathbf{f}_{t\rightarrow t+1}(\mathbf{u}_{t}), where ft→t+1\mathbf{f}_{t\rightarrow t+1} is a bilinearly interpolated forward flow map. After back-projection, we obtain the corresponding 3D point position xt+1\mathbf{x}_{t+1} at frame t+1t+1. In practice, extra care is needed for handling the unknown depth scale from model prediction and occlusion, discussed next.

Note that the sparse 3D points from COLMAP are all located on the static background. When projecting the sparse 3D points onto the image, some points might be occluded by the moving objects in the foreground. We handle occluded points by fitting aa and bb using RANSAC , which ignores outliers and is robust in practice.

Handling occlusions.

We identify occlusions using a forward-backward consistency check following the method of Brox et al. . We briefly summarize their method here.

Concretely, we identify an occlusion by chaining the forward flow ft→t+1\mathbf{f}_{t\rightarrow t+1} and backward flow ft+1→t\mathbf{f}_{t+1\rightarrow t} and thresholding based on warp consistency. For those pixels that have inconsistent forward and backward optical flows, defined by regions where chained forward and backward flows result in non-zero flow values, satisfying the following inequality:

The occluded pixels, along with the background pixels not belonging to the foreground mask, are excluded from the 3D scene flow computation.

Discussion.

In practice, we find that the Ω\Omega metric relies on the model estimation quality, in particular, the monocular depth prediction. We therefore design a second metric by measuring camera angular speed ω\omega. With some practical assumptions, it circumvents Ω\Omega’s limitation and does not rely on any external model estimates.

B.2 Angular EMF ω𝜔\omega: Camera angular velocity

We propose to measure camera angular speed ω\omega given the camera parameters, frame rate NN and a single 3D look-at point a\mathbf{a} obtained by triangulating all cameras, following Nerfies . Recall that ω\omega is computed as the scaled expectation,

When computing this metric, we assume that (1) the object moves at roughly constant speed, (2) the camera always fixates on the object, and (3) the distance between the camera and the object remains approximately the same over time. All sequences from existing works as well as ours meet these assumptions, except those from the NSFF and NV-DYN datasets. In their case, the cameras are always facing forward, breaking the assumption (2). However, we find that even though cameras are not fixated on the object since they are static, we can still compute the look-at point a\mathbf{a} by considering the center of mass of the foreground visible surfaces in 3D. Both datasets provide accurate foreground segmentations and MVS depth, which we use to identify and back-project foreground pixels into 3D space. The final look-at point is computed as the average foreground points over all frames.

Note that existing works only provide extracted frames from each sequence without specifying the frame rate. We identify the frame rates by re-assembling the original video using different FPS candidates and hand-picking the one that results in the most natural object and camera motion, which are verified by the original authors . We document per-sequence FPS for future reference in Table 4.

Appendix C Computation for co-visibility mask and masked image metrics

Code for both the co-visibility mask and masked image metrics are made publicly available on our project page. In this section, we provide further details for their computation processes.

In dynamic scenes, particularly for monocular capture with multi-camera validation, the test view contains regions that may not have been observed at all by the training camera. To circumvent this issue without resorting to camera teleportation, for each pixel in the test image, we propose “co-visibility” masking, which tests how many times a test pixel has been observed in the training images.

We visualize the computation process of the co-visibility mask in Figure 12. Concretely, for each (a) test frame, we first check its (b) occlusion in each training frame by the forward-backward flow consistency check according to Equation 4. We use RAFT for optical flow estimation between each test frame and each training frame. Note that we visualize occlusion as both a binary mask and its overlay on the test image. For occlusion mask visualization, the black color indicates pixels with no correspondence in the training views. We then compute the (c) co-visibility heatmap by simply summing up all test view occlusion masks. This co-visibility heatmap stores the number of times that each pixel is seen in training views. For visualization purpose, we normalize the heatmap by the number of the training frames NN. Finally, we apply a threshold β\beta to the heatmap and obtain a (d) binary co-visibility mask, which we also visualize with (e) its overlay on the test image. We adopt a conservative strategy and set β=max⁡(5,0.1⋅N)\beta=\max(5,0.1\cdot N), meaning that we deem a pixel “seen” during training and valid for evaluation when it is seen in 55 or 10%10\% of training frames, whichever is larger. This strategy ensures high recall in the masking result, i.e., the final co-visible regions are adequately seen during training when the flow estimation is noisy. For example, as shown in the third row of Figure 12 (b), the test view occlusions are inaccurate and miss the red cover of the chicken toy when it is visible in both frames. However, since the red cover is adequately seen over the whole sequence, it is still included in the final co-visibility mask.

C.2 Masked image metrics

In this work, we propose to only evaluate on regions that are adequately seen during training by co-visibility masking. We employ three masked image metrics, namely mPSNR, mSSIM and mLPIPS, which extend from their original definition, which we discuss next.

PSNR is originally defined as per-pixel mean squared error (MSE) in the log scale (with a constant negative multiplier). We compute mPSNR by simply taking the average of per-pixel PSNR scores over the masked region.

SSIM [9] →→\rightarrow mSSIM.

Comparing to PSNR, SSIM is defined on the patch level: it considers the structural similarity within each patch. In practice, it is usually implemented as convolutions where kernels are defined by the pixels in each patch. We take inspiration from Liu et al. and follow exactly their partial convolution implementation for this operation, where only the masked pixels are accounted for the final result.

LPIPS [10] →→\rightarrow mLPIPS [4, 42, 43].

LPIPS is also defined on patch level. Given two images, it computes their similarity distance in the feature space across different spatial resolution using a pretrained AlexNet model . The final similarity score is the average over all distance maps. To compute mLPIPS, we follow the previous works and first apply the co-visibility mask on the input images by zeroing out the unseen regions. Given the output distance maps at each spatial resolution, we then apply the same mask with downsampling and compute the masked average distance score. It should be noted that the pretrained AlexNet has a receptive field of 1952195^{2}. Thus when the co-visibility mask is small (most of the pixels are not seen during training), this metric can be artificially low due to the zeroing operation.

Appendix D Correspondence readout from existing works

In this section, we first review the formulation of the existing works and then describe the computation to read out correspondence from these models.

A neural radiance field (NeRF) represents a static scene as a continuous volumetric field FF that transforms a point’s position x\mathbf{x} and auxiliary variables w\mathbf{w} (e.g., view direction, latent appearance vector) to color c\mathbf{c} and density σ\sigma,

Here we briefly review representative approaches that extend NeRFs to dynamic scenes.

Similarly to traditional non-rigid reconstruction methods that explains non-rigid scenes with a static canonical space and a per-frame deformation model , Nerfies capture a non-rigid scene with one canonical NeRF FF and a per-time step view-to-canonical deformation Wt→cW_{t\rightarrow c} that takes a point x\mathbf{x} with a time-conditioned latent vector φt\bm{\varphi}_{t} to a canonical point xc\mathbf{x}_{c},

At each time step the resulting volumetric field is Ft=F∘Wt→cF_{t}=F\circ W_{t\rightarrow c}. HyperNeRF addresses topological change on top of Nerfies by outputting a two-dimensional “ambient” coordinate w\mathbf{w} encoding the topological change in addition to the canonical point xc\mathbf{x}_{c},

These two output variables are passed to the (topologically varying) canonical space mapping FF.

Time-conditioned NeRF and NSFF . Another way to handle non-rigid scenes is to directly map space-time to the output color and density by a time-conditioned latent vector φt\bm{\varphi}_{t}, which we refer to as T-NeRF:

Note that since T-NeRF implicitly handles deformation, it is difficult to compute correspondences over time. NSFF augments T-NeRF’s implicit function FtF_{t} to output an explicit scene flow field Wt→t+δW_{t\rightarrow t+\delta} between adjacent time steps tt and t+δt+\delta,

This explicit flow field is used to regularize motion and, as shown below, can be chained to compute long-range point correspondences across views and times.

D.2 Correspondence readout

Our goal is to find view-to-view correspondences such that given a set of key-points on a source image at time t1t_{1}, we can find their correspondence on a target image at time t2t_{2}.

For clarity, we start with assuming a known 3D view-to-view warp Wt1→t2W_{t_{1}\rightarrow t_{2}}, outlined in the last sub-section. The 2D correspondence ut2\mathbf{u}_{t_{2}} given ut1\mathbf{u}_{t_{1}} can be obtained by three steps, which we describe as “warp-integrate-project”. In the “warp” step, given the pixel location ut1\mathbf{u}_{t_{1}} and camera πt1\pi_{t_{1}}, we sample points on the ray passing from the camera center through the pixel πt1−1(ut1)\pi_{t_{1}}^{-1}(\mathbf{u}_{t_{1}}). Then, we warp the sampled points toward their 3D correspondences in the target frame using the known 3D warp Wt1→t2W_{t_{1}\rightarrow t_{2}}. In the “integrate” step, we compute the expected 3D location for the source samples weighted by the probability mass wt1w_{t_{1}} by volume rendering, as per NeRF . We can use densities from either source or target frame, a choice that we find insensitive in practice. In our formulation, we use the densities from the source frame. Finally, in the “project” step, we project the expected 3D location to the target frame through the target camera πt2\pi_{t_{2}}. The “warp-integrate-project” process can be written as

Note that there are also other alternatives such as “warp-project-integrate” where integration happens after projecting warped points to 2D. We find in practice that these different approaches make little difference to the final results when the surface is dense such that wtw_{t} is concentrated near one point (almost one-hot) for each ray.

We solve for the forward map given the inverse map through optimization:

We use the Broyden solver for root-finding, as per SNARF , and initialize xc\mathbf{x}_{c} with xt\mathbf{x}_{t}.

NSFF [4].

We can compose Wt1→t2W_{t_{1}\rightarrow t_{2}} by chaining the scene flow predictions through time. Concretely we have

Appendix E Summary of the capture setup and data processing for our iPhone dataset

Our capture setup has 77 multi-camera captures (MV) and 77 single camera captues (SV). We evaluate novel-view synthesis on the multi-camera captures and correspondence on all captures.

For multi-camera captures, we employ three cameras: one hand-held camera to capture monocular video for training and two stationary mounted cameras for validation. The two validation cameras face inward from two distinct viewpoints with large baseline. This wide-baseline setup enables us to better evaluate the shape modeling quality for novel-view synthesis. We use the “Record3D” app on iPhone to record both RGB and depth information at each time step. Note that we only collect depth information for training views given that we will only use the depths for supervision. We discuss the preprocessing procedure for the training video sequence in the “Single-camera captures” paragraph below.

To synchronize multiple cameras, we leverage the “audio-based multi-camera synchronization” functionality in Adobe Premiere Pro, as per , which achieves millisecond-level accuracy. In Figure 13, we show visualizations of our multi-camera captures after time synchronization. To ensure that our input sequence covers most of the scene regions in evaluation, we intentionally move the training camera in front of each test camera at certain frames. When we do so, that particular test frame is excluded due to severe occlusion (shown as “Excluded” in the figure). For the Wheel sequence (last row), we only employ the right camera due to the limited physical space to set up the multi-camera rig in that scene.

After time synchronization, we calibrate the multi-camera system. The Record3D app provides camera parameters and poses at each time step, but the poses only relate to each other within each capture sequence. In fact, each camera pose is recorded as relative pose to the first frame in each sequence, with the first pose being identity. We therefore need to solve the relative SE(3)SE(3) transforms between the first frame in each test sequence with respect to the first training frame. This problem can be formulated as a Perspective-n-Point (PnP) problem where, given a set of 3D points and their corresponding 2D pixels in two sequences, we aim to solve the camera pose. In practice, given a training RGBD frame and a testing RGB frame, we compute a set of 2D correspondences by SIFT feature matching and obtain their 3D point positions (in the training sequence’s world space) by back-projecting the 2D keypoints with the training frame depth map. This process is repeated for all time steps. We exploit our problem structure by constraining the camera poses within each test sequence to be the same, i.e., static camera. We use the RANSAC PnP solver in OpenCV .

Single-camera captures.

We treat the single-camera captures as the training sequence in our multi-camera capture setup. In effect, the single-camera capture setup will not have validation data for novel-view synthesis evaluation. We preprocess the depth data for the training sequence by applying a Sobel filter to filter out inaccurate depth values around object edges. In Figure 14, we visualize our depth data before and after filtering. We find that NeRF is particularly sensitive to depth noise and this filtering step is necessary. Finally, we manually annotate keypoints for correspondence evaluation. For sequences of humans and quadrupeds (dogs or cats), we annotate keypoints based on the skeleton defined in the COCO challenge and StanfordExtra . For sequences that focus on more general objects (e.g., our Block and Teddy sequences), we manually identify and annotate 5 to 15 trackable keypoints across frames. We visualize keypoint annotations (with skeleton if available) for both our proposed iPhone dataset and the Nerfies-HyperNeRF dataset in Figure 15.

Note that both Nerfies and HyperNeRF use background regularization which requires a point cloud of the background static scene. We first extract the object mask over time by MTTR, an off-the-shelf video segmentation network , which takes a text prompt of the foreground object as input. Since our foreground objects are quite diverse (e.g., backpack and block), the segmentation results are usually noisy. Thus we apply TSDF Fusion to the background point clouds over the whole sequence to get a completed background point cloud. We find that this point cloud can be noisy when segmentation fails, and that it is necessary to manually filter the background point cloud to make sure that it does not include any foreground regions. We consider this manual process a weakness of the previous background regularization .

Appendix F Summary of the implementation details and remaining differences

To ensure a fair comparison, we align numerous training details between the models that we investigate in this paper: T-NeRF, NSFF , Nerfies and HyperNeRF . Code and checkpoints are available on our project page.

To start with, we align the total number of rays seen during training. We add support of ray undistortion in the third-party implementation of NSFF to make sure that the training rays are the same across codebases. All models are trained with view-dependency modeling turned on. We did not find appearance encoding helpful in terms of quantitative results. This might due to the lighting difference between training and validation captures – a common issue in evaluation discussed in mip-NeRF 360 .

Due to no publicly available code to train NSFF on the Nerfies-HyperNeRF dataset. We adapt and extend the third-party implementation of NSFF (which we find to perform better than the official repo ). We confirm the finding from HyperNeRF that the default hyper-parameters in the NSFF paper are not suitable for long video sequences, and use their hyper-parameters instead. In Table 6, we check on one sequence that our modified re-implementation of NSFF can reproduce the numbers from the ones we obtain by running the released code. On 11 NVIDIA RTX A4000 or NVIDIA A100 GPU, it takes roughly 7272 to train a NSFF. With better implementation, we hypothesize that the training process can be largely accelerated.

While we try to ensure the fairness in our comparison, there are still four main remaining differences, namely: (1) static scene stablization, (2) sampling and rendering, (3) NeRF coordinates, and (4) flow supervision. First, Nerfies and HyperNeRF use additional background points from SfM system as supervision to stabilize the static region of the scene, which we find sensitive to foreground segmentation errors as mentioned in Section E. On the other hand, NSFF stabilizes the static region by composing the samples from a time-invariant static NeRF and a time-varying dynamic NeRF. Second, Nerfies and HyperNeRF sample S=128S=128 points during the coarse stage, and another 2S2S points during the fine stage, evaluating 3S=3843S=384 points in total. NSFF, on the other hand, only samples SS points for dynamic NeRF and another SS points for static NeRF, without coarse-to-fine sampling, evaluating 2S=2562S=256 points in total. Third, Nerfies and HyperNeRF sample points in world space, while NSFF samples in normalized device coordinates (NDC), which can cause issues when applying to non-forward-facing scenes like the ones we use in this paper. Finally, NSFF uses additional optical flow supervision, while Nerfies and HyperNeRF do not. In fact, we consider the fact that NSFF can leverage correspondence supervision as a merit in the sense that it is non-trivial to apply optical flow supervision to Nerfies and HyperNeRF since their warp representation is not fully invertible.

Appendix G Additional results on the impact of effective multi-view

In Figure 16, we provide more qualitative comparisons between models that are trained with and without camera teleportation on the Nerfies-HyperNeRF dataset.

Appendix H Additional results on per-sequence quantitative performance breakdown

We document the per-sequence quantitative performances of different models on both the Nerfies-HyperNeRF dataset (under non-teleporting setting) in Table 7 and the proposed iPhone dataset in Table 8.

Appendix I Additional results on novel-view synthesis

We provide additional novel-view synthesis qualitative results under the non-teleporting setting. In Figure 17, we show qualitative results on the Nerfies-HyperNeRF dataset. In Figure 18, we show qualitative results on the multi-camera captures from the proposed iPhone dataset. All models except NSFF are trained with all the additional regularizations that we find helpful through ablation, denoted with “++” to distinguish with the original models. In Figure 19, we show qualitative results on the single-camera captures from the proposed iPhone dataset. We render novel views using the camera pose from the first captured frame. Finally, in Figure 20, we show the rendering results with and without co-visibility mask applied.

Appendix J Additional results on inferred correspondence

In Table 9, we provide additional quantitative results of the inferred correspondence on the single-camera captures from the proposed iPhone dataset. In Figure 21 and 22, we provide additional qualitative results of the inferred correspondence on both the Nerfies-iPhone dataset and the proposed iPhone dataset. Note that all models are trained with additional regularizations on the proposed iPhone dataset except NSFF.