TRIP: Temporal Residual Learning with Image Noise Prior for Image-to-Video Diffusion Models
Zhongwei Zhang, Fuchen Long, Yingwei Pan, Zhaofan Qiu, Ting Yao, Yang Cao, Tao Mei
Introduction
In recent years, deep generative models have demonstrated impressive capabilities to create high-quality visual content. In between, Generative Adversarial Networks (GANs) and diffusion models have brought forward milestone improvement for a series of generative tasks in computer vision field, e.g., text-to-image generation , image editing , and text-to-video generation . In this work, we are interested in image-to-video generation (I2V), i.e., animating a static image by endowing it with motion dynamics, which has great potential real-world applications (e.g., entertainment and creative media). Conventional I2V techniques commonly target for generating video from a natural scene image with random or very coarse motion dynamics (such as clouds moving or fluid flowing). In contrast, our focus is a more challenging text-driven I2V scenario: given the input static image and text prompt, the animated frames should not only faithfully align with the given first frame, but also be temporally coherent and semantically matched with text prompt.
Several recent pioneering practices for I2V is to remould the typical latent diffusion model of text-to-video generation by directly leveraging the given static image as additional condition during diffusion process. Figure 1 (a) illustrates such conventional noise prediction strategy in latent diffusion model. Concretely, by taking the given static image as the first frame, 2D Variational Auto-encoder (VAE) is first employed to encode the first frame as image latent code. This encoded image latent code is further concatenated with noised video latent code (i.e., a sequence of Gaussian noise) to predict the backward diffusion noise of each subsequent frame via a learnable 3D-UNet. Nevertheless, such independent noise prediction strategy of each frame leaves the inherent relation between the given image and each subsequent frame under-exploited, and is lacking in efficacy of modeling temporal coherence among adjacent frames. As observed in I2V results of Figure 1 (a), both the foreground content (a glass) and the background content (e.g., a table) in the synthesized second frame are completely different from the ones in the given first frame. To alleviate this issue, our work shapes a new way to formulate the noise prediction in I2V diffusion model as temporal residual learning on the basis of amplified guidance of given image (i.e., image noise prior). As illustrated in Figure 1 (b), our unique design is to integrate the typical noise prediction with an additional shortcut path that directly estimates the backward diffusion noise of each frame by solely referring image noise prior. Note that this image noise prior is calculated based on both input image and noised video latent codes, which explicitly models the inherent correlation between the given first frame and each subsequent frame. Meanwhile, such residual-like scheme re-shapes the typical noise prediction via 3D-UNet as residual noise prediction to ease temporal modeling among adjacent frames, thereby leading to more temporally coherent I2V results.
By materializing the idea of executing image-conditioned noise prediction in a residual manner, we present a novel diffusion model, namely Temporal Residual learning with Image noise Prior (TRIP), to facilitate I2V. Specifically, given the input noised video latent code and the corresponding static image latent code, TRIP performs residual-like noise prediction along two pathways. One is the shortcut path that first achieves image noise prior based on static image and noised video latent codes through one-step backward diffusion process, which is directly regarded as the reference noise of each frame. The other is the residual path which concatenates static image and noised video latent codes along temporal dimension and feeds them into 3D-UNet to learn the residual noise of each frame. Eventually, a Transformer-based temporal noise fusion module is leveraged to dynamically fuse the reference and residual noises of each frame, yielding high-fidelity video that faithfully aligns with given image.
Related Work
Text-to-Video Diffusion Models. The great success of diffusion models to generate fancy images based on text prompts has been witnessed in recent years. The foundation of the basic generative diffusion models further encourages the development of customized image synthesis, such as image editing and personalized image generation . Inspired by the impressive results of diffusion models on image generation, a series of text-to-video (T2V) diffusion models start to emerge. VDM is one of the earlier works that remoulds text-to-image (T2I) 2D-UNet architecture with temporal attention for text-to-video synthesis. Later in and , the temporal modules (e.g., temporal convolution and self-attention) in diffusion models are solely optimized to emphasize motion learning. To further alleviate the burden on temporal modeling, Ge et al. present a mixed video noise model under the assumption that each frame should correspond to a shared common noise and its own independent noise. Similar recipe is also adopted by VideoFusion which estimates the basic noise and residual noise via two image diffusion models for video denoising. Other works go one step further by applying diffusion models for customized video synthesis, such as video editing and personalized video generation . In this work, we choose the basic T2V diffusion model with 3D-UNet as our backbone for image-to-video generation task.
Image-to-Video Generation. Image-to-Video (I2V) generation is one kind of conditional video synthesis, where the core condition is a given static image. According to the availability of motion cues, I2V approaches can be grouped into two categories: stochastic generation (solely using input image as condition) and conditional generation (using image and other types of conditions) . Early attempts of I2V generation mostly belong to the stochastic models which usually focus on the short motion of fluid elements (e.g., water flow) or human poses for landscape or human video generation. To further enhance the controllability of motion modeling, the conditional models exploit more types of conditions like text prompts in the procedure of video generation. Recent advances start to explore diffusion models for conditional I2V task. One pioneering work (VideoComposer ) flexibly composes a video with additional conditions of texts, optical flow or depth mapping sequence for I2V generation. Nevertheless, VideoComposer independently predicts the noise of each frame and leaves the inherent relation between input image and other frames under-exploited, thereby resulting in temporal inconsistency among frames.
Deep Residual Learning. The effectiveness of learning residual components with additional shortcut connections in deep neural networks has been verified in various vision tasks. Besides typical convolution-based structure, recent advances also integrate the shortcut connection into the emerging transformer-based structure . Intuitively, learning residual part with reference to the optimal identity mapping will ease the network optimization. In this work, we capitalize on such principle and formulate noise prediction in I2V diffusion as temporal residual learning on the basis of the image noise prior derived from static image to enhance temporal coherence.
In summary, our work designs a novel image-to-video diffusion paradigm with residual-like noise prediction. The proposed TRIP contributes by studying not only how to excavate the prior knowledge of given image as noise prior to guide temporal modeling, and how to strengthen temporal consistency conditioned on such image noise prior.
Our Approach
In this section, we present our Temporal Residual learning with Image noise Prior (TRIP) for I2V generation. Figure 2 illustrates an overview of our TRIP. Given a video clip at training, TRIP first extracts the per-frame latent code via a pre-trained VAE and groups them as a video latent code. The image latent code of the first frame (i.e., the given static image) is regarded as additional condition, which is further concatenated with the noised video latent code along temporal dimension as the input of the 3D-UNet. Next, TRIP executes a residual-like noise prediction along two pathways (i.e., shortcut path and residual path), aiming to trigger inter-frame relational reasoning. In shortcut path, the image noise prior of each frame is calculated based on the correlation between the first frame and subsequent frames through one-step backward diffusion process. In residual path, the residual noise is estimated by employing 3D-UNet over the concatenated noised video latent code plus text prompt. Finally, a transformer-based Temporal Noise Fusion (TNF) module dynamically merges the image noise prior and the residual noise as the predicted noise for video generation.
T2V diffusion models are usually constructed by remoulding T2I diffusion models . A 3D-UNet with text prompt condition is commonly learnt to denoise a randomly sampled noise sequence in Gaussian distribution for video synthesis. Specifically, given an input video clip with frames, a pre-trained 2D VAE encoder first extracts latent code of each frame to constitute a latent code sequence . Then, this latent sequence is grouped via concatenation along temporal dimension as a video latent code . Next, the Gaussian noise is gradually added to through forward diffusion procedure (length: , noise scheduler: ). The noised latent code at each time step is thus formulated as:
where and is the adding noise vector sampled from standard normal distribution . 3D-UNet with parameter aims to estimate the noise based on the multiple inputs (i.e., noised video latent code , time step , and text prompt feature extracted by CLIP ). The Mean Square Error (MSE) loss is leveraged as the final objective :
In the inference stage, the video latent code is iteratively denoised by estimating noise in each time step via 3D-UNet, and the following 2D VAE decoder reconstructs each frame based on the video latent code .
2 Image Noise Prior
Different from typical T2V diffusion model, I2V generation model goes one step further by additionally emphasizing the faithful alignment between the given first frame and the subsequent frames. Most existing I2V techniques control the video synthesis by concatenating the image latent code of first frame with the noised video latent code along channel/temporal dimension. In this way, the visual information of first frame would be propagated across all frames via the temporal modules (e.g., temporal convolution and self-attention). Nevertheless, these solutions only rely on the temporal modules to calibrate video synthesis and leave the inherent relation between given image and each subsequent frame under-exploited, which is lacking in efficacy of temporal coherence modeling. To alleviate this issue, we formulate the typical noise prediction in I2V generation as temporal residual learning with reference to an image noise prior, targeting for amplifying the alignment between synthesized frames and first static image.
Formally, given the image latent code of first frame, we concatenate it with the noised video latent code sequence ( denotes time steps in forward diffusion) along temporal dimension, yielding the conditional input for 3D-UNet. Meanwhile, we attempt to excavate the image noise prior to reflect the correlation between the first frame latent code and -th noised frame latent code . According to Eq. (1), we can reconstruct the -th frame latent code in one-step backward diffusion process as follows:
where denotes Gaussian noise which is added to the -th frame latent code. Considering the basic assumption in I2V that all frames in a short video clip are inherently correlated to the first frame, the -th frame latent code can be naturally formulated as the combination of the first frame latent code and a residual item as:
Next, we construct a variable by adding a scale ratio to the residual item as follows:
Thus Eq. (5) can be re-written as follows:
After integrating Eq. (3) and Eq. (6) into Eq. (4), the first frame latent code is denoted as:
where is defined as the image noise prior which can be interpreted as the relationship between the first frame latent code and -th noised frame latent code . The image noise prior is thus measured as:
According to Eq. (8), the latent code of first frame can be directly reconstructed through the one-step backward diffusion by using the image noise prior and the noised latent code of -th frame. In this way, such image noise prior acts as a reference noise for Gaussian noise which is added in -th frame latent code, when is close to (i.e., -th frame is similar to the given first frame).
3 Temporal Residual Learning
Note that the computation of Eq. (9) is operated through the Temporal Noise Fusion (TNF) module that estimates the backward diffusion noise of each frame. Moreover, since the temporal correlation between -th frame and the first frame will decrease when increasing frame index with enlarged timespan, we shape the trade-off parameter as a linear decay parameter with respect to frame index.
4 Temporal Noise Fusion Module
Our residual-like dual-path noise prediction explores the image noise prior as a reference to facilitate temporal modeling among adjacent frames. However, the handcrafted design of TNF module with simple linear fusion (i.e., Eq. (9)) requires a careful tuning of hyper-parameter , leading to a sub-optimal solution. Instead, we devise a new Transformer-based temporal noise fusion module to dynamically fuse reference and residual noises, aiming to further exploit the relations in between and boost noise fusion.
where denotes the operation of TNF module. Compared to the simple linear fusion strategy for noise estimation, our Transformer-based TNF module sidesteps the hyper-parameter tuning and provides an elegant alternative to dynamically merge the reference and residual noises of each frame to yield high-fidelity videos.
Experiments
Datasets. We empirically verify the merit of our TRIP model on WebVid-10M , DTDB and MSR-VTT datasets. Our TRIP is trained over the training set of WebVid-10M, and the validation sets of all three datasets are used for evaluation. The WebVid-10M dataset consists of about video-caption pairs with a total of video hours. videos are used for validation and the remaining videos are utilized for training. We further sample videos from the validation set of WebVid-10M for evaluation. DTDB contains more than dynamic texture videos of natural scenes. We follow the standard protocols in to evaluate models on a video subset derived from 4 categories of scenes (i.e., fires, clouds, vegetation, and waterfall). The texts of each category are used as the text prompt since there is no text description. For the MSR-VTT dataset, there are clips in training set and clips in validation set. Each video clip is annotated with text descriptions. We adopt the standard setting in to pre-process the video clips in validation set for evaluation.
Implementation Details. We implement our TRIP on PyTorch platform with Diffusers library. The 3D-UNet derived from Stable-Diffusion v2.1 is exploited as the video backbone. Each training sample is a 16-frames clip and the sampling rate is fps. We fix the resolution of each frame as 256256, which is centrally cropped from the resized video. The noise scheduler is set as linear scheduler ( and ). We set the number of time steps as in model training. The sampling strategy for video generation is DDIM with steps. The 3D-UNet and TNF module are trained with AdamW optimizer (learning rate: for 3D-UNet and for TNF module). All experiments are conducted on 8 NVIDIA A800 GPUs.
Evaluation Metrics. Since there is no standard evaluation metric for I2V task, we choose the metrics of frame consistency (F-Consistency) which is commonly adopted in video editing and Frechet Video Distance (FVD) widely used in T2V generation for the evaluation on WebVid-10M. Note that the vision model used to measure F-Consistency is the CLIP ViT-L/14 . Specifically, we report the frame consistency of the first 4 frames (F-Consistency4) and all the 16 frames (F-Consistencyall). For the evaluation on DTDB and MSR-VTT, we follow and report the frame-wise Frechet Inception Distance (FID) and FVD performances.
Human Evaluation. The user study is also designed to assess the temporal coherence, motion fidelity, and visual quality of the generated videos. In particular, we randomly sample testing videos from the validation set in WebVid-10M for evaluation. Through the Amazon MTurk platform, we invite evaluators, and each evaluator is asked to choose the better one from two synthetic videos generated by two different methods given the same inputs.
2 Comparisons with State-of-the-Art Methods
We compare our TRIP with several state-of-the-art I2V generation methods, e.g., VideoComposer and T2V-Zero , on WebVid-10M. For the evaluation on DTDB, we set AL and cINN as baselines. Text-to-Video generation advances, e.g., CogVideo and Make-A-Video are included for comparison on MSR-VTT.
Evaluation on WebVid-10M. Table 1 summarizes performance comparison on WebVid-10M. Overall, our TRIP consistently leads to performance boosts against existing diffusion-based baselines across different metrics. Specifically, TRIP achieves the F-Consistency4 of , which outperforms the best competitor T2V-Zero by . The highest frame consistency among the first four frames of our TRIP generally validates the best temporal coherence within the beginning of the whole video. Such results basically demonstrate the advantage of exploring image noise prior as the guidance to strengthen the faithful visual alignment between given image and subsequent frames. To further support our claim, Figure 4 visualizes the frame consistency of each run which is measured among different numbers of frames. The temporal consistencies of TRIP across various frames constantly surpass other baselines. It is worthy to note that TRIP even shows large superiority when only calculating the consistency between the first two frames. This observation again verifies the effectiveness of our residual-like noise prediction paradigm that nicely preserves the alignment with given image. Considering the metric of FVD, TRIP also achieves the best performance, which indicates that the holistic motion dynamics learnt by our TRIP are well-aligned with the ground-truth video data distribution in WebVid-10M.
Figure 5 showcases four I2V generation results across three methods (VideoComposer, T2V-Zero, and our TRIP) based on the same given first frame and text prompt. Generally, compared to the two baselines, our TRIP synthesizes videos with higher quality by nicely preserving the faithful alignment with the given image & text prompt and meanwhile reflecting temporal coherence. For example, T2V-Zero produces frames that are inconsistent with the second prompt “the ship passes close to the camera” and VideoComposer generates temporally inconsistent frames for the last prompt “the portrait of the girl in glasses”. In contrast, the synthesized videos of our TRIP are more temporally coherent and better aligned with inputs. This again confirms the merit of temporal residual learning with amplified guidance of image noise prior for I2V generation.
Evaluation on DTDB. Next, we evaluate the generalization ability of our TRIP on DTDB dataset in a zero-shot manner. Table 2 lists the performances of averaged FID and FVD over four scene-related categories. In general, our TRIP outperforms two strong stochastic models (AL and cINN) over both two metrics. Note that cINN also exploits the residual information for I2V generation, but it combines image and residual features as clip feature by learning a bijective mapping. Our TRIP is fundamentally different in that we integrate residual-like noise estimation into the video denoising procedure to ease temporal modeling in the diffusion model. Figure 6 shows the I2V generation results of four different scenes produced by our TRIP. As shown in this figure, TRIP manages to be nicely generalized to scene-related I2V generation task and synthesizes realistic videos conditioned on the given first frame plus class label.
Evaluation on MSR-VTT. We further compare our TRIP with recent video generation advances on MSR-VTT. Note that here we also evaluate TRIP in a zero-shot manner. Table 3 details the frame-wise FID and FVD of different runs. Similar to the observation on DTDB, our TRIP again surpasses all baselines over both metrics. The results basically validate the strong generalization ability of our proposal to formulate the open-domain motion dynamics.
3 Human Evaluation
In addition to the evaluation over automatic metrics, we also perform human evaluation to investigate user preference with regard to three perspectives (i.e., temporal coherence, motion fidelity, and visual quality) across different I2V approaches. Table 4 shows the comparisons of user preference on the generated videos over WebVid-10M dataset. Overall, our TRIP is clearly the winner in terms of all three criteria compared to both baselines (VideoComposer and T2V-Zero). This demonstrates the powerful temporal modeling of motion dynamics through our residual-like noise prediction with image noise prior.
4 Ablation Study on TRIP
In this section, we perform ablation study to delve into the design of TRIP for I2V generation. Here all experiments are conducted on WebVid-10M for performance comparison.
First Frame Condition. We first investigate different condition strategies to exploit the given first frame as additional condition for I2V generation. Table 5 summarizes the performance comparisons among different variants of our TRIP. TRIPC follows the common recipe to concatenate the clean image latent code of first frame with the noise video latent code along channel dimension. TRIPTE is implemented by concatenating image latent code at the end of the noise video latent code along temporal dimension, while our proposal (TRIP) temporally concatenates image latent code at the beginning of video latent code. In particular, TRIPTE exhibits better performances than TRIPC, which indicates that temporal concatenation is a practical choice for augmenting additional condition in I2V generation. TRIP further outperforms TRIPTE across both F-Consistency and FVD metrics. Compared to concatenating image latent code at the end of video, augmenting video latent code with input first frame latent code at the start position basically matches with the normal temporal direction, thereby facilitating temporal modeling.
Temporal Residual Learning. Next, we study how each design of our TRIP influences the performances of both temporal consistency and visual quality for I2V generation. Here we include two additional variants of TRIP: (1) TRIP- removes the shortcut path of temporal residual learning and only capitalizes on 3D-UNet for noise prediction; (2) TRIPW is a degraded variant of temporal residual learning with image noise prior by simply fusing reference and residual noise via linear fusion strategy. As shown in Table 6, TRIPW exhibits better performances than TRIP- in terms of F-consistency. The results highlight the advantage of leveraging image noise prior as the reference to amplify the alignment between the first frame and subsequent frames. By further upgrading temporal residual learning (TRIPW) with Transformer-based temporal noise fusion, our TRIP achieves the best performances across all metrics. Figure 8 further showcases the results of different variants, which clearly confirm the effectiveness of our designs.
5 Application: Customized Image Animation
In this section, we provide two interesting applications of customized image animation by directly extending TRIP in a tuning-free manner. The first application is text-to-video pipeline that first leverages text-to-image diffusion models (e.g., Stable-Diffusion XL ) to synthesize fancy images and then employs our TRIP to animate them as videos. Figure 7 shows six text-to-video generation examples by using Stable-Diffusion XL and TRIP. The results generally demonstrate the generalization ability of our TRIP when taking text-to-image results as additional reference for text-to-video generation. To further improve controllability of text-to-video generation in real-world scenarios, we can additionally utilize image editing models (e.g., InstructPix2Pix and ControlNet ) to modify the visual content of input images and then feed them into TRIP for video generation. As shown in Figure 9, our TRIP generates vivid videos based on the edited images with promising temporal coherence and visual quality.
Conclusions
This paper explores inherent inter-frame correlation in diffusion model for image-to-video generation. Particularly, we study the problem from a novel viewpoint of leveraging temporal residual learning with reference to the image noise prior of given image for enabling coherent temporal modeling. To materialize our idea, we have devised TRIP, which executes image-conditioned noise prediction in a residual fashion. One shortcut path is novelly designed to learn image noise prior based on given image and noised video latent code, leading to reference noise for enhancing alignment between synthesized frames and given image. Meanwhile, another residual path is exploited to estimate the residual noise of each frame through eased temporal modeling. Experiments conducted on three datasets validate the superiority of our proposal over state-of-the-art approaches.