FRESCO: Spatial-Temporal Correspondence for Zero-Shot Video Translation

Shuai Yang, Yifan Zhou, Ziwei Liu, Chen Change Loy

Introduction

In today’s digital age, short videos have emerged as a dominant form of entertainment. The editing and artistic rendering of these videos hold considerable practical importance. Recent advancements in diffusion models have revolutionized image editing by enabling users to manipulate images conveniently through natural language prompts. Despite these strides in the image domain, video manipulation continues to pose unique challenges, especially in ensuring natural motion with temporal consistency.

Temporal-coherent motions can be learned by training video models on extensive video datasets or finetuning refactored image models on a single video , which is however neither cost-effective nor convenient for ordinary users. Alternatively, zero-shot methods offer an efficient avenue for video manipulation by altering the inference process of image models with extra temporal consistency constraints. Besides efficiency, zero-shot methods possess the advantages of high compatibility with various assistive techniques designed for image models, e.g., ControlNet and LoRA , enabling more flexible manipulation.

Existing zero-shot methods predominantly concentrate on refining attention mechanisms. These techniques often substitute self-attentions with cross-frame attentions , aggregating features across multiple frames. However, this approach ensures only a coarse-level global style consistency. To achieve more refined temporal consistency, approaches like Rerender-A-Video and FLATTEN assume that the generated video maintains the same inter-frame correspondence as the original. They incorporate the optical flow from the original video to guide the feature fusion process. While this strategy shows promise, three issues remain unresolved. 1) Inconsistency. Changes in optical flow during manipulation may result in inconsistent guidance, leading to issues such as parts of the foreground appearing in stationary background areas without proper foreground movement (Figs. 2(a)(f)). 2) Undercoverage. In areas where occlusion or rapid motion hinders accurate optical flow estimation, the resulting constraints are insufficient, leading to distortions as illustrated in Figs. 2(c)-(e). 3) Inaccuracy. The sequential frame-by-frame generation is restricted to local optimization, leading to the accumulation of errors over time (missing fingers in Fig. 2(b) due to no reference fingers in previous frames).

To address the above critical issues, we present FRamE Spatial-temporal COrrespondence (FreSCo). While previous methods primarily focus on constraining inter-frame temporal correspondence, we believe that preserving intra-frame spatial correspondence is equally crucial. Our approach ensures that semantically similar content is manipulated cohesively, maintaining its similarity post-translation. This strategy effectively addresses the first two challenges: it prevents the foreground from being erroneously translated into the background, and it enhances the consistency of the optical flow. For regions where optical flow is not available, the spatial correspondence within the original frame can serve as a regulatory mechanism, as illustrated in Fig. 2.

In our approach, FreSCo is introduced to two levels: attention and feature. At the attention level, we introduce FreSCo-guided attention. It builds upon the optical flow guidance from and enriches the attention mechanism by integrating the self-similarity of the input frame. It allows for the effective use of both inter-frame and intra-frame cues from the input video, strategically directing the focus to valid features in a more constrained manner. At the feature level, we present FreSCo-aware feature optimization. This goes beyond merely influencing feature attention; it involves an explicit update of the semantically meaningful features in the U-Net decoder layers. This is achieved through gradient descent to align closely with the high spatial-temporal consistency of the input video. The synergy of these two enhancements leads to a notable uplift in performance, as depicted in Fig. 1. To overcome the final challenge, we employ a multi-frame processing strategy. Frames within a batch are processed collectively, allowing them to guide each other, while anchor frames are shared across batches to ensure inter-batch consistency. For long video translation, we use a heuristic approach for keyframe selection and employ interpolation for non-keyframe frames. Our main contributions are:

A novel zero-shot diffusion framework guided by frame spatial-temporal correspondence for coherent and flexible video translation.

Combine FreSCo-guided feature attention and optimization as a robust intra-and inter-frame constraint with better consistency and coverage than optical flow alone.

Long video translation by jointly processing batched frames with inter-batch consistency.

Related Work

Image diffusion models. Recent years have witnessed the explosive growth of image diffusion models for text-guided image generation and editing. Diffusion models synthesize images through an iterative denoising process . DALLE-2 leverages CLIP to align text and images for text-to-image generation. Imagen cascades diffusion models for high-resolution generation, where class-free guidance is used to improve text conditioning. Stable Diffusion builds upon latent diffusion model to denoise at a compact latent space to further reduce complexity.

Text-to-image models have spawned a series of image manipulation models . Prompt2Prompt introduces cross-attention control to keep image layout. To edit real images, DDIM inversion and Null-Text Inversion are proposed to embed real images into the noisy latent feature for editing with attention control .

Besides text conditioning, various flexible conditions are introduced. SDEdit introduces image guidance during generation. Object appearances and styles can be customized by finetuning text embeddings , model weights or encoders . ControlNet introduces a control path to provide structure or layout information for fine-grained generation. Our zero-shot framework does not alter the pre-trained model and, thus is compatible with these conditions for flexible control and customization as shown in Fig. 1.

Zero-shot text-guided video editing. While large video diffusion models trained or fine-tuned on videos have been studied , this paper focuses on lightweight and highly compatible zero-shot methods. Zero-shot methods can be divided into inversion-based and inversion-free methods.

Inversion-based methods apply DDIM inversion to the video and record the attention features for attention control during editing. FateZero detects and preserves the unedited region and uses cross-frame attention to enforce global appearance coherence. To explicitly leverage inter-frame correspondence, Pix2Video and TokenFlow match or blend features from the previous edited frames. FLATTEN introduces optical flows to the attention mechanism for fine-grained temporal consistency.

Inversion-free methods mainly use ControlNet for translation. Text2Video-Zero simulates motions by moving noises. ControlVideo extends ControlNet to videos with cross-frame attention and inter-frame smoothing. VideoControlNet and Rerender-A-Video warps and fuses the previous edited frames with optical flow to improve temporal consistency. Compared to inversion-based methods, inversion-free methods allow for more flexible conditioning and higher compatibility with the customized models, enabling users to conveniently control the output appearance. However, without the guidance of DDIM inversion features, the inversion-free framework is prone to flickering. Our framework is also inversion-free, but further incorporates intra-frame correspondence, greatly improving temporal consistency while maintaining high controllability.

Methodology

We follow the inversion-free image translation pipeline of Stable Diffusion based on SDEdit and ControlNet , and adapt it to video translation. An input frame II is first mapped to a latent feature x0=E(I)x_{0}=\mathcal{E}(I) with an Encoder E\mathcal{E}. Then, SDEdit applies DDPM forward process to add Gaussian noise to x0x_{0}

where αˉt\bar{\alpha}_{t} is a pre-defined hyperparamter at the DDPM step tt. Then, in the DDPM backward process , the Stable Diffusion U-Net ϵθ\epsilon_{\theta} predicts the noise of the latent feature to iteratively translate xT′=xTx^{\prime}_{T}=x_{T} to x0′x^{\prime}_{0} guided by prompt cc:

where αt\alpha_{t} and βt=1−αt\beta_{t}=1-\alpha_{t} are pre-defined hyperparamters, ztz_{t} is a randomly sampled standard Guassian noise, and x^0′\hat{x}^{\prime}_{0} is the predicted x0′x^{\prime}_{0} at the denoising step tt,

and ϵθ(xt,t′,c,e)\epsilon_{\theta}(x_{t},t^{\prime},c,e) is the predicted noise of xt′x^{\prime}_{t} based on the step tt, the text prompt cc and the ControlNet condition ee. The ee can be edges, poses or depth maps extracted from II to provide extra structure or layout information. Finally, the translated frame I′=D(x0′)I^{\prime}=\mathcal{D}(x^{\prime}_{0}) is obtained with a Decoder D\mathcal{D}. SDEdit allows users to adjust the transformation degree by setting different initial noise level with TT, i.e., large TT for greater appearance variation between I′I^{\prime} and II. For simplicity, we will omit the denoising step tt in the following.

2 Overall Framework

The proposed zero-shot video translation pipeline is illustrated in Fig. 3. Given a set of video frames I={Ii}i=1N\mathbf{I}=\{I_{i}\}^{N}_{i=1}, we follow Sec. 3.1 to perform DDPM forward and backward processes to obtain its transformed I′={Ii′}i=1N\mathbf{I}^{\prime}=\{I^{\prime}_{i}\}^{N}_{i=1}. Our adaptation focuses on incorporating the spatial and temporal correspondences of I\mathbf{I} into the U-Net. More specifically, we define temporal and spatial correspondences of I\mathbf{I} as:

Temporal correspondence. This inter-frame correspondence is measured by optical flows between adjacent frames, a pivotal element in keeping temporal consistency. Denoting the optical flow and occlusion mask from IiI_{i} to IjI_{j} as wijw^{j}_{i} and MijM^{j}_{i} respectively, our objective is to ensure that Ii′I^{\prime}_{i} and Ii+1′I^{\prime}_{i+1} share wii+1w^{i+1}_{i} in non-occluded regions.

Spatial correspondence. This intra-frame correspondence is gauged by self-similarity among pixels within a single frame. The aim is for Ii′I^{\prime}_{i} to share self-similarity as IiI_{i}, i.e., semantically similar content is transformed into a similar appearance, and vice versa. This preservation of semantics and spatial layout implicitly contributes to improving temporal consistency during translation.

Our adaptation focuses on the input feature and the attention module of the decoder layer within the U-Net, since decoder layers are less noisy than encoder layers, and are more semantically meaningful than the xtx_{t} latent space:

Feature adaptation. We propose a novel FreSCo-aware feature optimization approach as illustrated in Fig. 3. We design a spatial consistency loss Lspat\mathcal{L}_{spat} and a temporal consistency loss Ltemp\mathcal{L}_{temp} to directly optimize the decoder-layer features f={fi}i=1N\mathbf{f}=\{f_{i}\}_{i=1}^{N} to strengthen their temporal and spatial coherence with the input frames.

Attention adaptation. We replace self-attentions with FreSCo-guided attentions, comprising three components, as shown in Fig. 3. Spatial-guided attention first aggregates features based on the self-similarity of the input frame. Then, cross-frame attention is used to aggregate features across all frames. Finally, temporal-guided attention aggregates features along the same optical flow to further reinforce temporal consistency.

The proposed feature adaptation directly optimizes the feature towards high spatial and temporal coherence with I\mathbf{I}. Meanwhile, our attention adaptation indirectly improves coherence by imposing soft constraints on how and where to attend to valid features. We find that combining these two forms of adaptation achieves the best performance.

3 FreSCo-Aware Feature Optimization

The input feature f={fi}i=1N\mathbf{f}=\{f_{i}\}_{i=1}^{N} of each decoder layer of U-Net is updated by gradient descent through optimizing

The updated f^\hat{\mathbf{f}} replaces f\mathbf{f} for subsequent processing.

For the temporal consistency loss Ltemp\mathcal{L}_{temp}, we would like the feature values of the corresponding positions between every two adjacent frames to be consistent,

4 FreSCo-Guided Attention

A FreSCo-guided attention layer contains three consecutive modules: spatial-guided attention, efficient cross-frame attention and temporal-guided attention, as shown in Fig. 3.

Spatial-guided attention. In contrast to self-attention, patches in spatial-guided attention aggregate each other based on the similarity of patches before translation rather than their own similarity. Specifically, consistent with calculating Lspat\mathcal{L}_{spat} in Sec. 3.3, we perform a single-step DDPM forward and backward process over IiI_{i}, and extract its self-attention query vector QirQ^{r}_{i} and key vector KirK^{r}_{i}. Then, spatial-guided attention aggregate QiQ_{i} with

where λs\lambda_{s} is a scale factor and dd is the query vector dimension. As shown in Fig. 4, the foreground patch will mainly aggregate features in the C-shaped foreground region, and attend less to the background region. As a result, Q′Q^{\prime} has better spatial consistency with the input frame than QQ.

Efficient cross-frame attention. We replace self-attention with cross-frame attention to regularize the global style consistency. Rather than using the first frame or the previous frame as reference (V1, Fig. 4), which cannot handle the newly emerged objects (e.g., fingers in Fig. 2(b)), or using all available frames as reference (V2, Fig. 4), which is computationally inefficient, we aim to consider all frames simultaneously but with as little redundancy as possible. Thus, we propose efficient cross-frame attentions: Except for the first frame, we only reference to the areas of each frame that were not seen in its previous frame (i.e., the occlusion region). Thus, we can construct a cross-frame index pup_{u} of all patches within the above region. Keys and values of these patches can be sampled as K[pu]K[p_{u}], V[pu]V[p_{u}]. Then, cross-frame attention is applied

Temporal-guided attention. Inspired by FLATTEN , we use flow-based attention to regularize fine-level cross-frame consistency. We trace the same patches in different frames as in Fig. 4. For each optical flow, we build a cross-frame index pfp_{f} of all patches on this flow. In FLATTEN, each patch can only attend to patches in other frames, which is unstable when a flow contains few patches. Different from it, the temporal-guided attention has no such limit,

where λt\lambda_{t} is a scale factor. And HH is the final output of our FreSCo-guided attention layer.

5 Long Video Translation

The number of frames NN that can be processed at one time is limited by GPU memory. For long video translation, we follow Rerender-A-Video to perform zero-shot video translation on keyframes only and use Ebsynth to interpolate non-keyframes based on translated keyframes.

Keyframe selection. Rerender-A-Video uniformly samples keyframes, which is suboptimal. We propose a heuristic keyframe selection algorithm as summized in Algorithm 1. We relax the fixed sampling step to an interval [smin,smax][s_{\text{min}},s_{\text{max}}], and densely sample keyframes when motions are large (measured by L2L_{2} distance between frames).

Keyframe translation. With over NN keyframes, we split them into several NN-frame batches. Each batch includes the first and last frames in the previous batch to impose inter-batch consistency, i.e., keyframe indexes of the kk-th batch are {1,(k−1)(N−2)+2,(k−1)(N−2)+3,...,k(N−2)+2}\{1,(k-1)(N-2)+2,(k-1)(N-2)+3,...,k(N-2)+2\}. Besides, throughout the whole denoising steps, we record the latent features xt′x^{\prime}_{t} (Eq. (2)) of the first and last frames of each batch, and use them to replace the corresponding latent features in the next batch.

Experiments

Implementation details. The experiment is conducted on one NVIDIA Tesla V100 GPU. By default, we set batch size N∈N\in based on the input video resolution, the loss weight λspat=50\lambda_{\text{spat}}=50, the scale factors λs=λt=5\lambda_{s}=\lambda_{t}=5. For feature optimization, we update f\mathbf{f} for K=20K=20 iterations with Adam optimizer and learning rate of 0.40.4. We find optimization mostly converges when K=20K=20 and larger KK does not bring obvious gains. GMFlow is used to estimate optical flows and occlusion masks. Background smoothing is applied to improve temporal consistency in the background region.

We compare with three recent inversion-free zero-shot methods: Text2Video-Zero , ControlVideo , Rerender-A-Video . To ensure a fair comparison, all methods employ identical settings of ControlNet, SDEdit, and LoRA. As shown in Fig. 5, all methods successfully translate videos according to the provided text prompts. However, the inversion-free methods, relying on ControlNet conditions, may experience a decline in video editing quality if the conditions are of low quality, due to issues like defocus or motion blur. For instance, ControlVideo fails to generate a plausible appearance of the dog and the boxer. Text2Video-Zero and Rerender-A-Video struggle to maintain the cat’s pose and the structure of the boxer’s gloves. In contrast, our method can generate consistent videos based on the proposed robust FreSCo guidance.

For quantitative evaluation, adhering to standard practices , we employ the evaluation metrics of Fram-Acc (CLIP-based frame-wise editing accuracy), Tmp-Con (CLIP-based cosine similarity between consecutive frames) and Pixel-MSE (averaged mean-squared pixel error between aligned consecutive frames). We further report Spat-Con (LspatL_{spat} on VGG features) for spatial coherency. The results averaged across 23 videos are reported in Table 1. Notably, our method attains the best editing accuracy and temporal consistency. We further conduct a user study with 57 participants. Participants are tasked with selecting the most preferable results among the four methods. Table 1 presents the average preference rates across the 11 test videos, revealing that our method emerges as the most favored choice.

2 Ablation Study

To validate the contributions of different modules to the overall performance, we systematically deactivate specific modules in our framework. Figure 6 illustrates the effect of incorporating spatial and temporal correspondences. The baseline method solely uses cross-frame attention for temporal consistency. By introducing the temporal-related adaptation, we observe improvements in consistency, such as the alignment of textures and the stabilization of the sun’s position across two frames. Meanwhile, the spatial-related adaptation aids in preserving the pose during translation.

In Fig. 7, we study the effect of attention adaptation and feature adaption. Clearly, each enhancement individually improves temporal consistency to a certain extent, but neither achieves perfection. Only the combination of the two completely eliminates the inconsistency observed in hair strands, which is quantitatively verified by the Pixel-MSE scores of 0.037, 0.021, 0.018, 0.015 for Fig. 7(b)-(e), respectively. Regarding attention adaptation, we further delve into temporal-guided attention and spatial-guided attention. The strength of the constraints they impose is determined by λt\lambda_{t} and λs\lambda_{s}, respectively. As shown in Figs. 8-9, an increase in λt\lambda_{t} effectively enhances consistency between two transformed frames in the background region, while an increase in λs\lambda_{s} boosts pose consistency between the transformed cat and the original cat. Beyond spatial-guided attention, our spatial consistency loss also plays an important role, as validated in Fig. 10. In this example, rapid motion and blur make optical flow hard to predict, leading to a large occlusion region. Spatial correspondence guidance is particularly crucial to constrain the rendering in this region. Clearly, each adaptation makes a distinct contribution, such as eliminating the unwanted ski pole and inconsistent snow textures. Combining the two yields the most coherent results, as quantitatively verified by the Pixel-MSE scores of 0.031, 0.028, 0.025, 0.024 for Fig. 10(b)-(e), respectively.

Table 2 provides a quantitative evaluation of the impact of each module. In alignment with the visual results, it is evident that each module contributes to the overall enhancement of temporal consistency. Notably, the combination of all adaptations yields the best performance.

Figure 11 ablates the proposed efficient cross-frame attention. As with Rerender-A-Video in Fig. 2(b), sequential frame-by-frame translation is vulnerable to new appearing objects. Our cross-frame attention allows attention to all unique objects within the batched frames, which is not only efficient but also more robust, as demonstrated in Fig. 12.

FreSCo uses diffusion features before the attention layers for optimization. Since U-Net is trained to predict noise, features after attention layers (near output layer) are noisy, leading to failure optimization (Fig. 13(b)). Meanwhile, the four-channel x^0′\hat{x}^{\prime}_{0} (Eq. (3)) is highly compact, which is not suitable for warping or interpolation. Optimizing x^0′\hat{x}^{\prime}_{0} results in severe blurs and over-saturation artifacts (Fig. 13(c)).

3 More Results

Long video translation. Figure 1 presents examples of long video translation. A 16-second video comprising 400400 frames are processed, where 3232 frames are selected as keyframes for diffusion-based translation and the remaining 368368 non-keyframes are interpolated. Thank to our FreSCo guidance to generate coherent keyframes, the non-keyframes exhibit coherent interpolation as in Fig. 14.

Video colorization. Our method can be applied to video colorization. As shown in Fig. 15, by combining the L channel from the input and the AB channel from the translated video, we can colorize the input without altering its content.

4 Limitation and Future Work

In terms of limitations, first, Rerender-A-Video directly aligns frames at the pixel level, which outperforms our method given high-quality optical flow. We would like to explore an adaptive combination of these two methods in the future to harness the advantages of each. Second, by enforcing spatial correspondence consistency with the input video, our method does not support large shape deformations and significant appearance changes. Large deformation makes it challenging to use the optical flow of the original video as a reliable prior for natural motion. This limitation is inherent in zero-shot models. A potential future direction is to incorporate learned motion priors .

Conclusion

This paper presents a zero-shot framework to adapt image diffusion models for video translation. We demonstrate the vital role of preserving intra-frame spatial correspondence, in conjunction with inter-frame temporal correspondence, which is less explored in prior zero-shot methods. Our comprehensive experiments validate the effectiveness of our method in translating high-quality and coherent videos. The proposed FreSCo constraint exhibits high compatibility with existing image diffusion techniques, suggesting its potential application in other text-guided video editing tasks, such as video super-resolution and colorization.

Acknowledgments. This study is supported under the RIE2020 Industry Alignment Fund Industry Collaboration Projects (IAF-ICP) Funding Initiative, as well as cash and in-kind contribution from the industry partner(s). This study is also supported by NTU NAP and MOE AcRF Tier 2 (T2EP20221-0012).

References