Generative Image Dynamics
Zhengqi Li, Richard Tucker, Noah Snavely, Aleksander Holynski
Introduction
The natural world is always in motion, with even seemingly static scenes containing subtle oscillations as a result of factors such as wind, water currents, respiration, or other natural rhythms. Motion is one of the most salient visual signals, and humans are particularly sensitive to it: captured imagery without motion (or even with slightly unrealistic motion) can often seem uncanny or unreal.
While it is easy for humans to interpret or imagine motion in scenes, training a model to learn realistic scene motion is far from trivial. The motion we observe in the world is the result of a scene’s underlying physical dynamics, i.e., forces applied to objects that respond according to their unique physical properties — their mass, elasticity, etc. These properties and forces are hard to measure and capture at scale, but fortunately, in many cases measuring them is unnecessary: the necessary signals for producing plausible motion can often be extracted from observed 2D motion . While real-world observed motion is multi-modal and grounded in complex physical effects, it is nevertheless often predictable: candles will flicker in certain ways, trees will sway, and their leaves will rustle. This predictability is ingrained in our human perception of real scenes: by viewing a still image, we can imagine plausible motions that might have been ongoing as the picture was captured — or, since there might have been many possible such motions, a distribution of natural motions conditioned on that image. Given the facility with which humans are able to imagine these possible motions, a natural research problem is to model this same distribution computationally.
Recent advances in generative models, in particular conditional diffusion models , have enabled us to model rich distributions, including distributions of real images conditioned on text . This capability has enabled a number of previously impossible applications, such as text-conditioned generation of diverse and realistic image content. Following the success of these image models, recent work has extended these models to other domains, such as videos and 3D geometry .
In this paper, we explore modeling a generative prior for image-space scene motion, i.e., the motion of all pixels in a single image. This model is trained on motion trajectories automatically extracted from a large collection of real video sequences. In particular, from each training video we compute motion in the form of a spectral volume , a frequency-domain representation of dense, long-range pixel trajectories. This motion representation is well-suited to scenes that exhibit oscillatory dynamics such as trees and flowers moving in the wind. We find that this representation is also highly efficient and effective as an output of a diffusion model for modeling scene motions. We train a generative model that, conditioned on a single image, can sample spectral volumes from its learned distribution. A predicted spectral volume can then be directly transformed into a motion texture—a set of per-pixel, long-range pixel motion trajectories—that can be used to animate the image. Further, the spectral volume can be interpreted as an image-space modal basis that can be used to simulate interactive dynamics, using the modal analysis technique of Davis et al. . In this paper, we refer to the underlying frequency-space representation as either a spectral volume or an image-space modal basis, depending on whether it is used to encode a specific motion texture, or used to simulate dynamics.
We predict spectral volumes from input images using a diffusion model that generates coefficients one frequency at a time, but coordinates these predictions across frequency bands through a shared attention module. The predicted motions can be used to synthesize future frames (via an image-based rendering model)—turning still images into realistic animations, as illustrated in Fig. 1.
Compared with priors over raw RGB pixels, priors over motion capture more fundamental, lower-dimensional underlying structure that efficiently explains long-range variations in pixel values. Hence, generating intermediate motion leads to more coherent long-term generation and more fine-grained control over animations when compared with methods that perform image animation via synthesis of raw video frames. We demonstrate the use of our trained model in several downstream applications, such as creating seamless looping videos, editing the generated motions, and enabling interactive dynamic images via image-space modal bases, i.e., simulating the response of object dynamics to user-applied forces .
Related Work
Recent advances in generative models have enabled photorealistic synthesis of images conditioned on text prompts . These generative text-to-image models can be augmented to synthesize video sequences by extending the generated image tensors along a time dimension . While these methods are effective at producing plausible video sequences that capture the spatiotemporal statistics of real footage, the resulting videos often suffer from artifacts such as incoherent motion, unrealistic temporal variation in textures, and violations of physical constraints like preservation of mass.
Animating images.
Instead of generating videos entirely from text, other techniques take as input a still picture and animate it. Many recent deep learning methods adopt a 3D-Unet architecture to produce video volumes directly from an input image . Because these models are effectively the same video generation models (but conditioned on image information instead of text), they exhibit similar artifacts to those mentioned above. One way to overcome these limitations is to not directly generate the video content itself, but instead animate an input source image through explicit or implicit image-based rendering, i.e., moving the image content around according to motion derived from external sources such as a driving video , motion or 3D geometry priors , or user annotations . Animating images according to motion fields yields greater temporal coherence and realism, but these prior methods require additional guidance signals or user input, or otherwise utilize limited motion representations (e.g., optical flow fields, as opposed to full-video dense motion trajectories).
Motion models and motion priors.
In computer graphics, natural, oscillatory 3D motion (e.g., water rippling or trees waving in the wind) has long been modeled with noise that is shaped in the Fourier domain and then converted via an inverse Fourier transform to time-domain motion fields . Some of these methods rely on a modal analysis of the underlying dynamics of the systems being simulated . These spectral techniques were adapted to animate plants, water, and clouds from single 2D pictures by Chuang et al. with additional user annotations. Our work is particularly inspired by that of Davis , who showed how to connect modal analysis of a scene with the motions observed in a video of that scene, and how to use this analysis to simulate interactive dynamics from a video. We adopt the frequency-space motion representation of the spectral volume from Davis et al., extract this representation from a large set of training videos, and show that this representation is suitable for predicting motion from single images with diffusion models.
Other methods have also used various motion representations in prediction tasks — where an image or video is used to inform a deterministic future motion estimate , or a more rich distribution of possible motions (which can be modeled explicitly or by predicting the pixel values that would be induced by some implicit motion estimate) . However, many of these methods predict an optical flow motion estimate (i.e., the instantaneous motion of each pixel), not full per-pixel motion trajectories. In addition, much of this prior work is focused on tasks like activity recognition, not on synthesis tasks. More recent work has demonstrated the advantages of modeling and predicting motion using generative models in a number of closed-domain settings such as humans and animals .
Videos as textures.
Certain moving scenes can be thought of as a kind of texture—termed dynamic textures by Doretto et al. —that model videos as space-time samples of a stochastic process. Dynamic textures can represent smooth, natural motions such as waves, flames, or moving trees, and have been widely used for video classification, segmentation or encoding . A related kind of texture, called a video texture, represents a moving scene as a set of input video frames along with transition probabilities between any pair of frames . A large body of work exists for estimating and producing dynamic or video textures through analysis of scene motion and pixel statistics, with the aim of generating seamlessly looping or infinitely varying output videos . In contrast to much of this previous work, our method learns priors in advance that can then be applied to single images.
Overview
Given a single picture , our goal is to generate a video of length featuring oscillation dynamics such as those of trees, flowers, or candle flames moving in the breeze. Our system consists of two modules, a motion prediction module and an image-based rendering module. Our pipeline begins by using a latent diffusion model (LDM) to predict a spectral volume for the input image . The predicted spectral volume is then transformed to a sequence of motion displacement fields (a motion texture) using an inverse discrete Fourier transform. This motion determines the position of each input pixel at each future time step.
Given a predicted motion texture, our rendering module animates the input RGB image using an image-based rendering technique that splats encoded features from the input image and decodes these splatted features into an output frame with an image synthesis network (Sec. 5). We explore applications of this method, including producing seamless looping animations and simulating interactive dynamics, in Sec. 6.
Predicting motion
Formally, a motion texture is a sequence of time-varying 2D displacement maps , where the 2D displacement vector at each pixel coordinate from input image defines the position of that pixel at a future time . To generate a future frame at time , one can splat pixels from using the corresponding displacement map , resulting in a forward-warped image :
If our goal is to produce a video via a motion texture, then one choice would be to predict a time-domain motion texture directly from an input image. However, the size of the motion texture would need to scale with the length of the video: generating output frames implies predicting displacement fields. To avoid predicting such a large output representation for long output videos, many prior animation methods either generate video frames autoregressively , or predict each future output frame independently via an extra time embedding . However, neither strategy ensures long-term temporal consistency of generated video frames.
Fortunately, many natural motions, can be described as a superposition of a small number of harmonic oscillators represented with different frequencies, amplitude and phases . Because the underlying motions are quasi-periodic, it is natural to model them in the frequency domain, from which it is convenient to generate a video of arbitrary length.
Hence, we adopt an efficient frequency space representation of motion in a video from Davis et al. called a spectral volume, visualized in Figure 1. A spectral volume is the temporal Fourier transform of pixel trajectories extracted from a video, organized into images called modal images. Davis et al. further shows that, under certain assumptions, the spectral volume, evaluated at certain frequencies, forms an image-space modal basis that is a projection of the vibration modes of the underlying scene (or, more generally, captures spatial correlations in motion) . We use the term spectral volume to refer to a frequency-space encoding of a specific motion texture (with high frequencies removed). Later, we also refer to a spectral volume as an “image-space modal basis” when it is used for simulation.
Given this motion representation, we formulate the motion prediction problem as a multi-modal image-to-image translation task: from an input image to an output spectral motion volume. We adopt latent diffusion models (LDMs) to generate spectral volumes comprised of a -channel 2D motion spectrum map, where is the number of frequencies modeled, and where at each frequency we need four scalars to represent the complex Fourier coefficients for the and dimensions. Note that the motion trajectory of a pixel at future time steps and its representation as a spectral volume are related by the Fast Fourier transform (FFT):
How should we select the output frequencies? Prior work in real-time animation has observed that most natural oscillation motions are composed primarily of low-frequency components . To validate this observation, we computed the average power spectrum of the motion extracted from 1,000 randomly sampled 5-second real video clips. As shown in the left plot of Fig. 2, the power spectrum of the motion decreases exponentially with increasing frequency. This suggests that most natural oscillation motions can indeed be well represented by low-frequency terms. In practice, we found that the first Fourier coefficients are sufficient to realistically reproduce the original natural motion in a range of real videos and scenes.
2 Predicting motion with a diffusion model
We choose a latent diffusion model (LDM) as the backbone for our motion prediction module, as LDMs are more computationally efficient than pixel-space diffusion models, while preserving generation quality. A standard LDM consists of two main modules: (1) a Variational Autoencoder (VAE) that compresses the input image to a latent space through an encoder , then reconstructs the input from the latent features via a decoder , and (2) a U-Net based diffusion model that learns to iteratively denoise latent features starting from Gaussian random noise. Our training applies this not to input images but to motion spectra from real video sequences, which are encoded and then diffused for steps with a pre-defined variance schedule to produce noisy latents . The 2D U-Nets are trained to denoise the noisy latents by iteratively estimating the noise used to update the latent feature at each step . The training loss for the LDM is written as
where is the embedding of any conditional signal, such as text, semantic labels, or, in our case, the first frame of the training video sequence, . The clean latent features are then passed through the decoder to recover the spectral volume.
One issue we observed is that motion textures have particular distribution characteristics across frequencies. As visualized in the left plot of Fig. 2, the amplitude of our motion textures spans a range of 0 to 100 and decays approximately exponentially with increasing frequency. As diffusion models require that output values lie between 0 and 1 for stable training and denoising, we must normalize the coefficients of extracted from real videos before using them for training. If we scale the magnitudes of coefficients to based on image width and height as in prior work , almost all the coefficients at higher frequencies will end up close to zero, as shown in Fig. 2 (right-hand side). Models trained on such data can produce inaccurate motions, since during inference, even small prediction errors can lead to large relative errors after denormalization when the magnitude of the normalized coefficients are very close to zero.
To address this issue, we employ a simple but effective frequency adaptive normalization technique. In particular, we first independently normalize Fourier coefficients at each frequency based on statistics computed from the training set. Namely, at each individual frequency , we compute the percentile of the Fourier coefficient magnitudes over all input samples and use that value as a per-frequency scaling factor . Furthermore, we apply a power transformation to each scaled Fourier coefficient to pull it away from extremely small or large values. In practice, we found that a square root transform performs better than other transformations, such as log or reciprocal. In summary, the final coefficient values of spectral volume at frequency (used for training our LDM) are computed as
As shown on the right plot of Fig. 2, after applying frequency adaptive normalization the spectral volume coefficients no longer concentrate in a range of extremely small values.
Frequency-coordinated denoising.
The straightforward way to to predict a spectral volume with frequency bands is to output a tensor of channels from a standard diffusion U-Net. However, as in prior work , we observe that training a model to produce such a large number of channels tends to produce over-smoothed and inaccurate output. An alternative would be to independently predict a motion spectrum map at each individual frequency by injecting an extra frequency embedding to the LDM, but this results in uncorrelated predictions in the frequency domain, leading to unrealistic motion.
Therefore, we propose a frequency-coordinated denoising strategy as shown in Fig. 3. In particular, given an input image , we first train an LDM to predict a spectral volume texture map with four channels to represent each individual frequency , where we inject extra frequency embedding along with time-step embedding to the LDM network. We then freeze the parameters of this LDM model and introduce attention layers and interleave them with 2D spatial layers of across frequency bands. Specifically, for a batch size of input images, the 2D spatial layers of treat the corresponding noisy latent features of channel size as independent samples with shape . The cross-attention layer then interprets these as consecutive features spanning the frequency axis, and we reshape the latent features from previous 2D spatial layers to before feeding them to the attention layers. In other words, the frequency attention layers are used to coordinate the pre-trained motion latent features across all frequency channels in order to produce coherent spectral volumes. In our experiments, we observed that the average VAE reconstruction error improves from to when we switch from a standard 2D U-Net to a frequency-coordinated denoising module, suggesting an improved upper bound on LDM prediction accuracy; in our ablation study in Sec. 7.6, we also demonstrate that this design choice improves video generation quality compared with simpler configurations mentioned above.
Image-based rendering
We now describe how we take a spectral volume predicted for a given input image and render a future frame at time . We first derive motion trajectory fields in the time domain using the inverse temporal FFT applied at each pixel . The motion trajectory fields determine the position of every input pixel at every future time step. To produce a future frame , we adopt a deep image-based rendering technique and perform splatting with the predicted motion field to forward warp the encoded , as shown in Fig. 4. Since forward warping can lead to holes, and multiple source pixels can map to the same output 2D location, we adopt the feature pyramid softmax splatting strategy proposed in prior work on frame interpolation .
Specifically, we encode through a feature extractor network to produce a multi-scale feature map . For each individual feature map at scale , we resize and scale the predicted 2D motion field according to the resolution of . As in Davis et al. , we use predicted flow magnitude, as a proxy for depth, to determine the contributing weight of each source pixel mapped to its destination location. In particular, we compute a per-pixel weight, as the average magnitude of the predicted motion trajectory fields. In other words, we assume large motions correspond to moving foreground objects, and small or zero motions correspond to background objects. We use motion-derived weights instead of learnable ones because we observe that in the single-view case, learnable weights are not effective for addressing disocclusion ambiguities, as shown in the second column of Fig. 5.
With the motion field and weights , we apply softmax splatting to warp feature map at each scale to produce a warped feature , where is the softmax splatting operation. The warped features are then injected into intermediate blocks of an image synthesis decoder network to produce a final rendered image .
We jointly train the feature extractor and synthesis networks with start and target frames randomly sampled from real videos, using the estimated flow field from to to warp encoded features from , and supervising predictions against with a VGG perceptual loss . As shown in Fig. 5, compared to direct average splatting and a baseline deep warping method , our motion-aware feature splatting produces a frame without holes or artifacts around disocclusions.
Applications
We demonstrate applications that add dynamics to single still images using our proposed motion representations and animation pipeline.
Our system enables the animation of a single still picture by first predicting a spectral volume from the input image and generating an animation by applying our image-based rendering module to the motion displacement fields derived from the spectral volume. Since we explicitly model scene motion, this allows us to produce slow-motion videos by linear interpolating the motion displacement fields and to magnify (or minify) animated motions by adjusting the amplitude of predicted spectral volume coefficients.
2 Seamless looping
It is sometimes useful to generate videos with motion that loops seamlessly, meaning that there is no discontinuity in appearance or motion between the start and end of the video. Unfortunately, it is hard to find a large collection of seamlessly looping videos for training diffusion models. Instead, we devise a method to use our motion diffusion model, trained on regular non-looping video clips, to produce seamless looping video. Inspired by recent work on guidance for image editing , our method is a motion self-guidance technique that guides the motion denoising sampling processing using explicit looping constraints. In particular, at each iterative denoising step during the inference stage, we incorporate an additional motion guidance signal alongside standard classifier-free guidance , where we enforce each pixel’s position and velocity at the start and end frames to be as similar as possible:
where is the predicted 2D motion displacement field at time and denoising step . is the classifier-free guidance weight, and is the motion self-guidance weight.
3 Interactive dynamics from a single image
As shown in Davis et al. , the image-space motion spectrum from an observed video of an oscillating object, under certain assumptions, is proportional to the projections of vibration mode shapes of that object, and thus a spectral volume can be interpreted as an image-space modal basis. The modal shapes capture underlying oscillation dynamics of the object at different frequencies, and hence can be used to simulate the object’s response to a user-defined force such as poking or pulling. Therefore, we adopt the modal analysis technique from prior work , which assumes that the motion of an object can be explained by the superposition of a set of harmonic oscillators. This allows us to write the image-space 2D motion displacement field for the object’s physical response as a weighted sum of Fourier spectrum coefficients modulated by the state of complex modal coordinates at each simulated time step :
We simulate the state of the modal coordinates via a forward Euler method applied to the equations of motion for a decoupled mass-spring-damper system (in modal space) . We refer readers to the original work for a full derivation. Note that our method produces an interactive scene from a single picture, whereas these prior methods required a video as input.
Experiments
We use an LDM as the backbone for predicting spectral volumes, for which we use a variational auto-encoder (VAE) with a continuous latent space of dimension . We train the VAE with an reconstruction loss, a multi-scale gradient consistency loss , and a KL-divergence loss with respective weights of . We train the same 2D U-Net used in the original LDM work to perform iterative denosing with a simple MSE loss , and adopt the attention layers from for frequency-coordinated denoising. For quantitative evaluation, we train the VAE and LDM on images of size , which takes around 6 days to converge using 16 Nvidia A100 GPUs. For our main quantitative and qualitative results, we run the motion diffusion model with DDIM for 250 steps. We also show generated videos of up to a resolution of , created by fine-tuning our models on corresponding higher resolution data.
We adopt ResNet-34 as feature extractor in our IBR module. Our image synthesis network is based on a co-modulation StyleGAN architecture, which is a prior conditional image generation and inpainting model . Our rendering module runs in real-time at 25FPS on a Nvidia V100 GPU during inference. Additionally, we adopt the universal guidance to produce seamless looping videos, where we set weights , and 500 DDIM steps with 2 self-recurrence iterations.
2 Data and baselines
Since our focus is on natural scenes exhibiting oscillatory motion such as trees, flowers, and candles moving in the wind, we collect and process a set of 3,015 videos of such phenomena from online sources as well as from our own captures, where we withhold 10% of the videos for testing and use the remainder for training. To generate ground truth spectral volumes for training our motion prediction module, we found the choice of optical flow method to be crucial. In particular, we observed that deep-learning based flow estimators tend to produce over-smoothed flow fields. Instead, we apply a coarse-to-fine image pyramid-based optical flow algorithm between selected starting image and every future frame within a video sequence. We treat every 10th frame from each training video as a starting image and generate corresponding ground truth spectral volumes using the following 149 frames. We filter out samples with incorrect motion estimates or significant camera motion by removing examples with an average flow motion magnitude 8 pixels, or where all pixels have an average motion magnitude larger than one pixel. In total, our data consists of more than 150K samples of image-motion pairs.
3 Metrics
We compare our approach to several recent single-image animation and video prediction methods. Both Endo et al. and DMVFN predict instantaneous 2D motion fields and future frames in an auto-regressive manner. We also compare with Holynsky et al. which animate a picture through predicted Eulerian Motion Fields. Other recent work such as Stochastic Image-to-Video (Stochastic-I2V) , TATS , and MCVD adopt either VAEs, temporal transformers, or diffusion models to directly predict video frames. LFDM predicts flow fields in latent space with a diffusion model, then uses those flow fields to warp the encoded input image, generating future frames via a decoder. For methods that predict videos of short length, we apply them autoregressively to generate longer videos by taking the last output frame and using it as the input to another round of generation until the video reaches a length of 150 frames. We train all the above methods on our data using their respective open-source implementationsWe use open-source reimplementation from Fan et al. for the method of Holynsky et al. ..
We evaluate the quality of the videos generated by our approach and by prior baselines in two main ways. First, we evaluate the quality of individual synthesized frames using metrics designed for image synthesis tasks. We adopt the Fréchet Inception Distance (FID) and Kernel Inception Distance (KID) to measure the average distance between the distribution of generated frames and the distribution of ground truth frames.
Second, to evaluate the quality and temporal coherence of synthesized videos, we adopt the Fréchet Video Distance with window size 16 (FVD) and 32 (), based on an I3D model trained on the Human Kinetics datasets . To more faithfully reflect synthesis quality for the natural oscillation motions we seek to generate, we also adopt the Dynamic Texture Frechet Video Distance proposed by Dorkenwald et al. , which measures the distance from videos of window size 16 (DTFVD) and size 32 (), using a I3D model trained on the Dynamic Textures Database , a dataset consisting primarily of natural motion textures.
We further use a sliding window FID of a window size of 30 frames, and a sliding window DTFVD with window size 16 frames, as proposed by , to measure how generated video quality degrades over time.
For all the methods, we evaluate each error metric on videos generated without performing temporal interpolation, at resolution.
4 Quantitative results
Table 1 shows quantitative comparisons between our approach and baselines on our test set of unseen video clips. Our approach significantly outperforms prior single-image animation baselines in terms of both image and video synthesis quality. Specifically, our much lower FVD and DT-FVD distances suggest that the videos generated by our approach are more realistic and more temporally coherent. Further, Fig. 6 shows the sliding window FID and sliding window DT-FVD distances of generated videos from different methods. Thanks to the global spectral volume representation, videos generated by our approach are more temporally consistent and do not suffer from drift or degradation over time.
5 Qualitative results
We visualize qualitative comparisons between videos generated by our approach and by baselines in two ways. First, we show spatio-temporal - slices of the generated videos, a standard way of visualizing small or subtle motions in a video . As shown in Fig. 7, our generated video dynamics more strongly resemble the motion patterns observed in the corresponding real reference videos (second column), compared to other methods. Baselines such as Stochastic I2V and MCVD fail to model both appearance and motion realistically over time. Endo et al. produces video frames with fewer artifacts but exhibits over-smooth or non-oscillation motions.
We also qualitatively compare the quality of individual generated frames and motions across different methods by visualizing the predicted image and its corresponding motion displacement field at time . Fig. 8 shows that the frames generated by our approach exhibit fewer artifacts and distortions compared to other methods, and our corresponding 2D motion fields most resemble the reference displacement fields estimated from the corresponding real videos. In contrast, the background content generated by other methods tend to drift, as shown in the flow visualizations in the even-numbered rows. Moreover, the video frames generated by other methods exhibit significant color distortion or ghosting artifacts, suggesting that the baselines are less stable when generating videos with long time duration.
6 Ablation study
We conduct an ablation study to validate the major design choices in our motion prediction and rendering modules, comparing our full configuration with different variants. Specifically, we evaluate results using different numbers of frequency bands , , , and . We observe that increasing the number of frequency bands improves video prediction quality, but the improvement is marginal when using more than 16 frequencies. Next, we remove adaptive frequency normalization from the ground truth spectral volumes, and instead just scale them based on input image width and height (Scale w/ resolution). Additionally, we remove the frequency coordinated-denoising module (Independent pred.), or replace it with a simpler module where a tensor volume of channel spectral volumes are predicted jointly via a standard 2D U-net diffusion model (Volume pred.). Finally, we compare results where we render video frames using average splatting (Average splat), or use a baseline rendering method that applies softmax splatting over single-scale features subject to learnable weights used in Holynski et al. (Baseline splat). From Table 2, we observe that all simpler or alternative configurations lead to worse performance compared with our full model.
Discussion and conclusion
Since our approach only predicts spectral volumes at lower frequencies, it might fail to model general non-oscillating motions or high-frequency vibrations such as those of musical instruments. Furthermore, the quality of our generated videos relies on the quality of the motion trajectories estimated from the real video sequences. Thus, we observed that animation quality can degrade if the motion in the training videos consists of large displacements. Moreover, since our approach is based on image-based rendering from input pixels, the animation quality can also degrade if the generated videos require the creation of large amounts of content unseen in the input frame.
Conclusion.
We present a new approach for modeling natural oscillation dynamics from a single still picture. Our image-space motion prior is represented with spectral volumes , a frequency representation of per-pixel motion trajectories, which we find to be highly suitable for prediction with diffusion models, and which we learn from collections of real world videos. The spectral volumes are predicted using our frequency-coordinated latent diffusion model and are used to animate future video frames using a neural image-based rendering module. We show that our approach produces photo-realistic animations from a single picture and significantly outperforms prior baseline methods, and that it can enable other downstream applications such as creating interactive animations.
Acknowledgements.
We thank Abe Davis, Rick Szeliski, Andrew Liu, Boyang Deng, Qianqian Wang, Xuan Luo, and Lucy Chai for fruitful discussions and helpful comments.