COVE: Unleashing the Diffusion Feature Correspondence for Consistent Video Editing
Jiangshan Wang, Yue Ma, Jiayi Guo, Yicheng Xiao, Gao Huang, Xiu Li
Introduction
Diffusion models have shown exceptional performance in image generation , thereby inspiring their application in the field of image editing . These approaches typically leverage a pre-trained Text-to-Image (T2I) stable diffusion model , using DDIM inversion to transform source images into noise, which is then progressively denoised under the guidance of a prompt to generate the edited image.
Despite satisfactory performance in image editing, achieving high-quality video editing remains challenging. Specifically, unlike the well-established open-source T2I stable diffusion models , comparable T2V diffusion models are not as mature due to the difficulty of modeling complicated temporal motions, and training a T2V model from scratch demands substantial computational resources . Consequently, there is a growing focus on adapting the pre-trained T2I diffusion for video editing . In this case, maintaining temporal consistency in edited videos is one of the biggest challenges, which requires the generated frames to be stylistically coherent and exhibit smooth temporal transitions, rather than appearing as a series of independent images. Numerous methods have been working on this topic while still facing various limitations, such as the inability to ensure fine-grained temporal consistency (leading to flickering or blurring in generated videos), requiring additional components or needing extra training or optimization , etc.
In this work, our goal is to achieve highly consistent video editing by leveraging the intra-frame correspondence relationship among tokens, which is intuitively closely related to the temporal consistency of videos: If corresponding tokens across frames exhibit high similarity, the resulting video will thus demonstrate high temporal consistency. Taking a video of a man as an example, if the token representing his nose has high similarity across frames, his nose will be unlikely to deform or flicker throughout the video. However, how to obtain accurate correspondence information among tokens is still largely under-explored in existing works, although the intrinsic characteristic of the video editing task (i.e., the source video and edited video are expected to share similar motion and semantic layout) determines that it naturally exists in the source video. Some previous methods leverage a pre-trained optical-flow model to obtain the flowing trajectory of each token across frames, which can be seen as a kind of coarse correspondence information. Despite the self-attention among tokens in the same trajectory can enhance the temporal consistency of the edited video, it still encounters two primary limitations: Firstly, these methods heavily rely on a highly accurate pre-trained optical-flow model to obtain the correspondence relationship of tokens, which is not available in many scenarios . Secondly, supposing we have access to an extremely accurate optical-flow model, it is still only able to obtain the coarse one-to-one correspondence among tokens in different frames (Figure 2a), which would lead to the loss of information because one token is highly likely to correspond to multiple tokens in other frames in most cases (Figure 2b).
Addressing these problems, we notice that the inherent diffusion features naturally contain precise correspondence information. For instance, it is easy to find the corresponding points between two images by extracting their diffusion features and calculating the cosine similarity between tokens . However, until now none of the existing works have successfully utilized this characteristic in more complicated and challenging tasks such as video editing. In this paper, we propose COVE, which is the first work unleashing the potential of inherent diffusion feature correspondence to significantly enhance the quality and temporal consistency in video editing. Given a source video, we first extract the diffusion feature of each frame. Then for each token in the diffusion feature, we obtain its corresponding tokens in other frames based on their similarity. Within this process, we propose a sliding-window-based approach to ensure computational efficiency. In our sliding-window-based method, for each token, it is only required to calculate the similarity between it and the tokens in the next frame located within a small window, identifying the tokens with the top () highest similarity. After the correspondence calculation process, for each token, the coordinates of its corresponding tokens in each other frame can be obtained. During the inversion and denoising process, we sample the tokens in noisy latents based on the obtained coordinates. To reduce the redundancy and accelerate the editing process, token merging is applied in the temporal dimension, which is followed by self-attention. Our method can be seamlessly integrated into the off-the-shelf T2I diffusion model without extra training or optimization. Extensive experiments demonstrate that COVE significantly improves both the quality and the temporal consistency of generated videos, outperforming a wide range of existing methods and achieving state-of-the-art results.
Related Works
Diffusion Models have recently showcased impressive results in image generation, which generates the image through gradual denoising from the standard Gaussian noise. A large number of efforts on diffusion models has enabled it to be applied to numerous scenarios . With the aid of large-scale pretraining , text-to-image diffusion models exhibit remarkable progress in generating diverse and high-quality images . ControlNet enables users to provide structure or layout information for precise generation. Naturally, diffusion models have found application in video synthesis, often by integrating temporal layers into image-based DMs . Despite successes in unconditional video generation , text-to-video diffusion models lag behind their image counterparts.
2 Text-to-Video Editing.
There are increasing works adopting the pre-trained text-to-image diffusion model to the video editing task , where keeping the temporal consistency in the generated video is the most challenging. Recently, a large number of works focusing on zero-shot video editing has been proposed. FateZero proposes to use attention blending to achieve high-quality edited videos while struggling to edit long videos. TokenFlow reduces the effects of flickering through the linear combinations between diffusion features, while the smoothing strategy can cause blurring in the generated video. RAVE proposes the randomized noise shuffling method, suffering the problem of fine details flickering. There are also a large number of methods that enhance the temporal consistency with the aid of pre-trained optical-flow models . Although the effectiveness of them, all of them severely rely on a pre-trained optical-flow model. Recent works illustrate that the diffusion feature contains rich correspondence information. Although VideoSwap adopts this characteristic by tracking the key points across frames, it still needs users to provide the key points as the extra addition manually.
Method
In this section, we will introduce COVE in detail, which can be seamlessly integrated into the pre-trained T2I diffusion model for high-quality and consistent video editing without the need for training or optimization (Figure 3). Specifically, given a source video, we first extract the diffusion feature of each frame using the pre-trained T2I diffusion model. Then, we calculate the one-to-many correspondence of each token across frames based on cosine similarity (Figure 3a). To reduce resource consumption during correspondence calculation, we further introduce an efficient sliding-window-based strategy (Figure 4). During each timestep of inversion and denoising in video editing, the tokens in the noisy latent are sampled based on the correspondence and then merged. Through the self-attention among merged tokens (Figure 3b), the quality and temporal consistency of edited videos are significantly enhanced.
Diffusion Models. DDPM is the latent generative model trained to reconstruct a fixed forward Markov chain . Given the data distribution , the Markov transition is defined as a Gaussian distribution with a variance schedule .
where indicates the textual prompt. and are predicted by the denoising model . Since the diffusion and denoising process in the pixel space is computationally extensive, latent diffusion is proposed to address this issue by performing these processes in the latent space of a VAE .
DDIM Inversion. DDIM can convert random noise to a deterministic during sampling . The inversion process in deterministic DDIM can be formulated as follows:
where denotes . The inversion process of DDIM is utilized to transform the input into , facilitating subsequent tasks such as reconstruction and editing.
2 Correspondence Acquisition
As discussed in Section 1, intra-frame correspondence is crucial for the quality and temporal consistency of edited videos while remaining largely under-explored in existing works. In this section, we introduce our method for obtaining correspondence relationships among tokens across frames.
One-to-many Correspondence Calculation. For each token within the diffusion feature , its corresponding tokens in other frames are identified based on the cosine similarity. Without loss of generality, we could consider a specific token in the th frame with the coordinate . Unlike previous methods where only one corresponding token of can be identified in each frame (Figure 2a), our method can obtain the one-to-many correspondences simply by selecting tokens with the top highest similarity in each frame. We record their coordinates, which are used for sampling the tokens for self-attention in the subsequent inversion and denoising process. To implement this process, the most straightforward method is through a direct matrix multiplication of the normalized diffusion feature .
The similarity between and all tokens in the feature is given by . The coordinates of the corresponding tokens in the th frame () are then obtained by selecting the tokens with the top similarities in the th frame.
Here the top--argmax() denotes the operation to find coordinates of the top biggest values in a matrix, where . represents the coordinates of the token in th frame which has highest similarity with . A similar process can be conducted for each token of , thereby obtaining their correspondences among frames.
Sliding-window Strategy. Although the one-to-many correspondence among tokens can be effectively obtained through the above process, it requires excessive computational resources because is always a huge number, especially in long videos. As a result, the computational complexity of this process is extremely high, which can be represented as . At the same time, multiplication between these two huge matrices consumes a substantial amount of GPU memory in practice. These limitations severely limit its applicability in many real-world scenarios, such as on mobile devices.
To address the above problem, we further propose the sliding-window-based strategy as an alternative, which not only effectively obtains the one-to-many correspondences but also significantly reduces the computational overhead (Figure 4). Firstly, for the token , it is only necessary to calculate its similarity with the tokens in the next frame instead of in all frames, i.e.,
For tokens in th frame, instead of considering , we identify the tokens in th frame which have the top largest similarity with the token through the . Similarly, we can obtain the corresponding token in other future or previous frames.
Through the above process, the overall complexity is reduced to . Furthermore, it is noteworthy that frames in a video exhibit temporal continuity, implying that the spatial positions of corresponding tokens are unlikely to change significantly between consecutive frames. Consequently, for the token , it is enough to only calculate the similarity within a small window of length in the adjacent frame, where is much smaller than and ,
3 Correspondence-guided Video Editing.
In this section, we explain how to apply the correspondence information to the video editing process (Figure 3c). In the inversion and denoising process of video editing, we sample the corresponding tokens from the noisy latent for each token based on the coordinates obtained in Figure 4. For the token , the set of corresponding tokens in other frames at a timestep is:
We merge these tokens following , which can accelerate the editing process and reduce GPU memory usage without compromising the quality of editing results:
Then, the self-attention is conducted on the merged tokens,
where is the scale factor. The above process of correspondence-guided attention is illustrated in Figure 3b. Following the previous methods , we also retain the spatial-temporal attention in the U-Net. In spatial-temporal attention, considering a query token, all tokens in the video serve as keys and values, regardless of their relevance to the query. This correspondence-agnostic self-attention is not enough to maintain temporal consistency, introducing irrelevant information into each token, and thus causing serious flickering effects . Our correspondence-guided attention can significantly alleviate the problems of spatial-temporal attention, increasing the similarity of corresponding tokens and thus enhancing the temporal consistency of the edited video.
Experiment
In the experiment, we adopt Stable Diffusion (SD) 2.1 from the official Huggingface repository for COVE, employing 100 steps of DDIM inversion and 50 steps of denoising. To extract the diffusion feature, the noise of the specific timestep is added to each frame of the source video following . The feature is then extracted from the intermediate layer of the 2D Unet decoder during a single step of denoising. The window size is set to 9 for correspondence calculation, and is set to 3 for correspondence-guided attention. The merge ratio for token merging is 50%. For both qualitative and quantitative evaluation, we select 23 videos from social media platforms such as TikTok and other publicly available sources . Among these 23 videos, 3 videos have a length of 10 frames, 15 videos have a length of 20 frames, and 5 videos have a length of 32 frames. The experiments are conducted on a single RTX 3090 GPU for our method unless otherwise specified. We compare COVE with 5 baseline methods: FateZero , TokenFlow , FLATTEN , FRESCO and RAVE . For all of these baseline methods, we follow the default settings from their official Github repositories. The more detailed experimental settings of our method are provided in Appendix A.
2 Qualitative Results
We evaluate COVE on various videos under different types of prompts including both global and local editing (Figure 5). Global editing mainly involves background editing and style transferring. For background editing, COVE can modify the background while keeping the subject of the video unchanged (e.g. Third row, first column. “a car driving in milky way”). For style transfer, COVE can effectively modify the global style of the source video according to the prompt (e.g. Third row, second column. “Van Gogh style”). Our prompts for local editing include changing the subject of the video to another one (e.g. Third row, third column. “A cute raccoon”) and making local edits to the subject (e.g. fifth row, third column. “A sorrow woman”). For all of these editing tasks, COVE demonstrates outstanding performance, generating frames with high visual quality while successfully preserving temporal consistency. We also compare COVE with a wide range of state-of-the-art video editing methods (Figure 6). The experimental results illustrate that COVE effectively edits the video with high quality, significantly outperforming the previous methods.
3 Quantitative Results
For quantitative comparison, we follow the metrics proposed in VBench , including Subject Consistency, Motion Smoothness, Aesthetic Quality, and Imaging Quality. Among them, Subject Consistency assesses whether the subject (e.g., a person) remains consistent throughout the whole video by calculating the similarity of DINO feature across frames. Motion Smoothness utilizes the motion priors of the video frame interpolation model to evaluate the smoothness of the motion in the generated video. Aesthetic Quality uses the LAION aesthetic predictor to assess the artistic and beauty value perceived by humans on each frame. Imaging Quality evaluates the degree of distortion in the generated frames (e.g., blurring, flickering) through the MUSIQ image quality predictor. Each video undergoes editing with 3 global prompts (such as style transferring, background editing, etc.) and 2 local prompts (such as editing the appearance of the subject in the video), generating a total of 115 text-video pairs. For each metric, we report the average score of these 115 videos. We further conducted a user study with 45 participants following . Participants are required to choose the most preferable results among these methods. The result is shown in Table 1. Among various methods, COVE achieves outstanding performance in both qualitative metrics and user studies, further demonstrating its superiority.
4 Ablation Study
We conduct an ablation study to illustrate the effectiveness of the Correspondence-guided attention and the number of tokens selected in each frame (i.e., the value of ). The experimental results (Table 2 and Figure 7) illustrate that without correspondence-guided attention, the edited video exhibits obvious temporal inconsistency and flickering effects (which is marked in yellow and orange boxes in Figure 7), thus severely impairing the visual quality. As increases from 1 to 3, the generated video contains more fine-grained details, exhibiting better visual quality. However, further increasing to 5 does not significantly improve the video quality. We also illustrate the effectiveness of temporal dimensional token merging. By merging the tokens with high correspondence across frames, the editing process becomes more efficient (Table 3) while there is no significant decrease in the quality of the edited video (Figure 8). The ablation of the sliding-window size is shown in Appendix B. If the window size is too small, the actual corresponding token may not be included within the window, resulting in suboptimal correspondence and poor editing results. On the other hand, a too-large window size is not necessary for identifying the corresponding tokens, which would lead to high computational complexity and excessive memory usage. The experiment results illustrate that is suitable to strike a balance. Additionally, we also visualize the correspondence obtained by COVE, which is shown in Appendix C.
Conclusion
In this paper, we propose COVE, which is the first to explore how to employ inherent diffusion feature correspondence in video editing to enhance editing quality and temporal consistency. Through the proposed efficient sliding-window-based strategy, the one-to-many correspondence relationship among tokens across frames is obtained. During the inversion and denoising process, self-attention is performed within the corresponding tokens to enhance temporal consistency. Additionally, we also apply token merging in the temporal dimension to improve the efficiency of the editing process. Both quantitative and qualitative experimental results demonstrate the effectiveness of our method, which outperforms a wide range of previous methods, achieving state-of-the-art editing quality.
Limitaions. The limitation of our method is discussed in Appendix E.
References
Appendix
Appendix A Detailed Experimental Settings
In the experiment, the size of all source videos is . We adopt Stable Diffusion (SD) 2.1 from the official Huggingface repository for our method. To extract the diffusion feature, following , the noise of the timestep is added to each frame of the source video. The noisy frames of video are fed into the U-net, the feature is extracted from the intermediate layer of the 2D Unet decoder. The height and weight of the diffusion feature is 64. Following previous works, at the first 40 timesteps, the diffusion features are saved during DDIM inversion and are further injected during denoising. For Spatial-temporal attention, we use the xFormers to reduce memory consumption, while it is not used in correspondence-guided attention.
Appendix B Ablation Study on the Window Size
To illustrate the influence of window size , we conduct the experiment on a video with 20 frames on a single A100 GPU. During the correspondence calculation process, we calculate the theoretical computational complexity, which is the total number of multiplications and additions required. We also record actual GPU memory consumed under different window sizes, the result is shown in Table 4. With our sliding window strategy, the computational complexity and the GPU memory in the correspondence calculation process are significantly reduced. The visualization result is shown in Fig. 9. If the window size is too small, the motion in the video cannot be tracked, causing unsatisfying results. We choose for the experiments in other sections, which can achieve a balance between the memory consumed and the quality of the edited video.
Appendix C Visualisation of the Correspondence
We visualize the correspondence calculated by our sliding-window-based method to illustrate its effectiveness (Fig. 10). To be specific, we calculate the correspondence based on the diffusion feature, which is extracted at the final layer of the U-net decoder. We select the token representing the left eye of the wolf in the 7th frame (marked in red) and visualize its corresponding tokens in other frames (marked in yellow). The result illustrates that our method can effectively identify the corresponding tokens.
Appendix D Broader Impacts
Our work enables high-quality video editing, which is in high demand across various social media platforms, especially short video websites like TikTok. Using our method, people can easily create high-quality and creative videos, significantly reducing production costs. However, there is a potential for misuse, such as replacing the characters in videos with celebrities, which may infringe upon the celebrities’ image rights. Therefore, it is also necessary to improve relevant laws and regulations to ensure the legal use of our method.
Appendix E Limitations
Despite achieving outstanding results, our methods still encounter several limitations. First, although the correspondence calculation process is efficient through the proposed sliding window strategy, the implementation of correspondence-guided attention is still not efficient enough, leading to the extra usage of GPU memory and time (Table 3). This problem is expected to be alleviated largely through the use of xFormers. We will work on it in the future.
Second, further exploration is required to optimize the application of the obtained correspondence information. In this study, we utilize the correspondence information to sample tokens during the inversion and denoising processes and do the self-attention. However, we believe that there may be more effective alternatives to self-attention that could further unleash the potential of the correspondence information.
Appendix F More Qualitative Results
We provide more qualitative results of our method to illustrate its effectiveness, which is shown in Fig. 12 and Fig. 11.