MagicStick: Controllable Video Editing via Control Handle Transformations

Yue Ma, Xiaodong Cun, Sen Liang, Jinbo Xing, Yingqing He, Chenyang Qi, Siran Chen, Qifeng Chen

Introduction

Due to the remarkable progress of text-to-image (T2I) generation , video editing has recently achieved significant progress leveraging the generation prior of text-to-image models . Previous works have studied various editing effects, such as visual style transfer or modification of the character and object to a similar one . However, many straightforward edits in the video remain out of reach currently. For example, can we enlarge or resize the specific objects in the video? What if we replace the localization of the specific objects? Or even, can we edit the human motion in the video?

These editings seem possible for images but it is hard for video. First, each frame in the original video is different so the editing needs to be applied to all the frames individually. Second, the temporal consistency will deteriorate if we apply frame-wise editing to the video directly. Finally, taking resizing as an example, if we directly segment the object and then edit, the generated video still relies on the abilities of additional inpainting and harmonization . To handle these problems, our key insight is inspired by ControlNet , where the structure can be generated following additional control. Editing the control signal is a relatively easy and clear way other than directly editing appearance. Thus, we first segment the specific object and then perform the editing on the one specific frame of the video, where other frames can utilize the same transformations by simple propagation.

To this end, we propose MagicStick, a universal video editing framework for geometry editing rather than appearance. To keep appearance, we first train a controllable video generation network following based on the pre-trained Stable Diffusion and ControlNet to get the specific appearance of the given video with the control signal and text prompt. After that, we involve the transformed conditional in both video inversion and generation. Finally, we propose a novel attention remixing block by remixing the attention in the inversion and generation with the guidance of fine-tuned ControlNet signals for editing. Thanks to our unified framework, we can successfully make changes to the specific video in terms of shape, size, and localization. To the best of our knowledge, we are the first to demonstrate the ability of such video editing via the pure text-to-image model without any 3D information. Besides the mentioned attributes, we can also alter the motion in the input human video, which is also absent in previous T2I-based video editing approaches. To justify the effectiveness of our novel editing framework, we conduct extensive experiments on various videos, demonstrating the superiority of our approach.

Our contributions are summarized as follows:

We demonstrate the ability of a series of important video editing aspects (including shape, size, location, and human motion) for the first time by introducing a new unified controllable video editing framework using pretrained T2I models.

We propose an effective attention remix module utilizing the attention from control signals to faithfully retain the edit-unrelated information from the source video.

The experiments show the advantages of the proposed methods over the previous similar topics, e.g., Shape-aware Video Editing and handcrafted motion controls .

Related Work

Video Editing. Editing natural videos is a vital task that has drawn the attention of current researchers in the field of computer vision. Before the advent of diffusion models, many GAN-based approaches have achieved good performance. The emergency of diffusion models delivers higher quality and more diverse editing results. Text2live and StableVideo present layer-atlas-based methods and edit the video on a flattened texture map. FateZero and Video-p2p , guided by original and target text prompts, perform semantic editing by blending cross-attention activations. There are also some approaches edit the appearance of the generated images by powerful yet private video diffusion models. However, Most of these methods mainly focus on the texture of content rather than the shape editing, which is a more challenging topic. They show obvious artifacts even with the optimization of generative priors. In contrast, our framework can achieve the editing of complex properties including shape, size, and location, while maintaining both appearance and temporal consistency.

Image and Video Generation. Text-to-image generation is a popular topic with extensive research in recent years. Many approaches have developed based on transformer architectures to achieve textual control for generated content. Currently, diffusion-based models have gained remarkable attention due to the impressive generation performance. DALLE-2 enhances text-image alignments by leveraging the representation space of CLIP and Imagen employs cascaded diffusion models to achieve high-definition image generation. Self-guidance conducts controllable image generation using attention and activations However, it operates attention during generation and struggles to maintain the consistency with input. Moreover, its focus is on image generation and cannot be applied to video editing directly.

To address a similar problem in video generation, various works try to extend LDM to video domain. The approaches such as MagicVideo and GEN1 adopt a distinct strategy by initializing the model from text-to-image . These methods then generate continuous content by incorporating additional temporal layers. Tune-A-Video proposes a method that specializes in one-shot text-to-video generation tasks. The model has the ability to generate video with similar motion to the source video. However, how to edit real-world video content using this model is still unclear. Inspired by the image controllable generation methods and Tune-A-Video, our method can achieve video property editing in real-world videos by leveraging the pretrained text-to-image model.

Image Editing has a long history in computer vision. Early works mainly leverage GAN or VAE to generate results in some specific domain. Thanks to the development of diffusion model , many current works adopt the pre-trained diffusion model for editing. SDEdit adds noise and corruption to the image for generating the content. DiffEdit and Blended Diffusion use the additional mask to blend the inversion noises or representations during the image generation process. Prompt-to-Prompt achieve semantic editing by mixing activations from original and target text prompts. Similar work has also been proposed by Plug-and-Play , Pix2pix-Zero and Masactrl . There are also some methods achieving better editing performance by finetuning on a single image. However, applying these image editing approaches to video frames directly will lead to serious issues such as flickering and inconsistency among frames.

Method

We aim to edit the property changes (e.g., shape, size, location, motion) in a video through transformations on one specific control signal as shown in Fig. 2. Below, we first give the basic knowledge of the latent diffusion model and inversion in Sec. 3.1. Then, we introduce our video customization method in Sec. 3.2 to keep the appearance. Finally, we present the details of control handle transformation (Sec. 3.3) and the editing details in inference (Sec. 3.4).

Derived from diffusion Models, Latent Diffusion Models reformulate the diffusion and denoising procedures within a latent space. First, an encoder E\mathcal{E} compresses a pixel space image xx to a low-resolution latent z=E(x)z=\mathcal{E}(x) , which can be reconstructed from latent feature to image D(z)≈x\mathcal{D}(z)\approx x by decoder D\mathcal{D}. Second, a U-Net εθ\varepsilon_{\theta} with self-attention and cross-attention is optimized to eliminate the artificial noise using the objective during training:

where pp is the embedding of the text prompt and ztz_{t} is a noisy sample of z0z_{0} at timestep tt. After training, we can generate the image from a random noise ε\varepsilon and a text embedding pp to recover the latent zz and then decoded by D\mathcal{D}.

DDIM Inversion. During inference, we can use deterministic DDIM sampling to transform a random noise zTz_{T} to a clean latent z0z_{0} across a sequence of timesteps from TT to 11:

where αt\alpha_{t} is the parameter for noise scheduling . Thus, DDIM Inversion is proposed to inverse the above progress from a clean latent space z0z_{0} to a noised latent space zTz_{T} by adding noising:

In this context, the z0z_{0} can be reconstructed by inverted latent z^T\hat{z}_{T} using DDIM and classifier-free guidance.

2 Controllable Video Customization

Since the generation process of the generative model is too stochastic to keep the appearance, for our tasks, we first tune the network to satisfy our requirements, which is similar to model customization in text-to-image generation. Differently, we also involve the pretrained ContorlNet to add additional correspondence between the condition and output, with additional designs of temporal attention layer , trainable LoRA and token embedding for better video customization finetuning.

In detail, as shown in Fig. 2, the low-rank matrices are injected into pre-trained linear layers within the cross-attention modules of the denoising UNet. LoRA employs a low-rank learnable factorization technique to update the attention weight matrix WqW_{q}, WkW_{k}, WvW_{v}:

where i=q,k,vi=q,k,v denotes the different part in cross-attention. Wi0∈Rd×kW_{i_{0}}\in R^{d\times k} represents the original weights of the pre-trained T2I model , B∈Rd×rB\in R^{d\times r} and A∈Rr×kA\in R^{r\times k} represent the low-rank factors, where rr is much smaller than original dimensions d and kk. Remarkably, this operation does not hurt the ability of the pre-trained T2I model in concept generation and composition. In addition, we leverage inflated ControlNet inspired by Tune-A-Video to extract temporal information among the condition sequences. The original condition of the object is encoded by structure-guided module, where we convert the self-attention to spatial-temporal self-attention.

3 Control Handle Transformation

Since naive object editing in a video may deteriorate the surrounding background, we opt to edit the first condition to ensure a more consistent editing performance. Then they are incorporated into UNet after being encoded by ControlNet-based structure guided module. As shown in Fig. 2, our control handle editing involves three steps.

1) Extract. We first segment the interested objects using Segment-and-Track-Anything . Then the annotator (e.g., pose detector, hed-detector, depth estimator) is utilized to extract conditions in the clean segmented object sequences.

2) Transformation. After getting the intermediate representations of the specific frame, the user can perform single or multiple geometric editing operations on objects in the first condition frame, such as resizing and moving.

3) Propagation. We compute the transformation parameters in the first edited frame and then apply these transformations across all frames, ensuring consistent guidance throughout the video. This propagation can also be adjusted frame-by-frame by the users.

4 Controllable Video Editing

With guidance from the user and the appearance customization, we can finally edit the video in our framework. In detail, we conduct the video inversion to obtain the intermediate features and then inject saved features during the denoising stage as shown in Fig. 2. Different from previous unconditional video inversion and editing pipeline , our method guides both processes with the control signals and then performs editing via the proposed attention remix module.

Most current works perform the inversion and editing pipeline to achieve appearance editing. For our task, we aim to achieve more complex property editing by incorporating a control handle during the inversion and generation. To be more specific, the same edited conditions in two stages are employed to play distinct roles. During inversion, we inject the representation of the edited condition into the UNet as well as modify the self-attention and cross-attention from the source video. As shown in Fig. 3, we observe that both the original and target areas are activated by edited words after injection. During generation, the edited condition is reintroduced into the network to serve as guidance for regenerating specific highlighted areas. In both stages, the injection of the edited condition is indispensable to accomplish our editing tasks through their mutual coordination.

Attention Remix Module.

With only the guidance of structure can not perform the editing well since the original objects will also influence the results. To achieve our ultimate goal, we propose the Attention ReMix module to modify the attention during the generation process. As shown in Fig. 3, we store the intermediate self-attention maps {stsrc}t=1T\{s_{t}^{\text{src}}\}_{t=1}^{T} and cross-attention maps {ctsrc}t=1T\{c_{t}^{\text{src}}\}_{t=1}^{T} at each timestep tt and the last noising latent maps zTz_{T} as:

where Inv donates for the DDIM inversion pipeline and Cedit{\mathcal{C}}^{edit} represents edited condition. As shown in Fig. 3, during the denoising stage, the activation areas of cross-attention by edited words provide significant assistance for mask generation. The binary mask MtM_{t} is obtained from thresholding the cross-attention map {ctsrc}t=1T\{c_{t}^{\text{src}}\}_{t=1}^{T}. Then we blend the {stsrc}t=1T\{s_{t}^{\text{src}}\}_{t=1}^{T} with MtM_{t} and use the structure-guided module to generate the target object. Formally, this process is implemented as:

where τ\tau stands for the threshold, GetMask(⋅)GetMask\left(\cdot\right) and G(⋅)G\left(\cdot\right) represent the operation to get mask and structure guidance.

Experiments

We implement our method based on the pretrained model of Stable Diffusion and ControlNet at 100 iterations. We sample 8 uniform frames at the resolution of 512×512512\times 512 from the input video and finetune the models for about 5 minutes on one NVIDIA RTX 3090Ti GPU. The learning rate is 3×10−53\times 10^{-5}. During editing, the attentions are fused in the DDIM step at t∈[0.5×T,T]t\in[0.5\times T,T] with total timestep T=50T=50. We use DDIM sampler with classifier-free guidance in our experiments. The source prompt of the video is generated via the image caption model and then replaced or modified some words manually.

Object size editing. Using the pretrained text-to-image diffusion model , our method supports object size editing through manual modification of the specific sketch in the first frame. As shown in the 1st and 2nd rows of Fig. 4, our method achieves consistent editing of the foreground “swan” by scaling down its sketch, while preserving the original background content.

Object position editing. One of the applications of our method is to edit the position of objects by adjusting the object’s position in the first frame of the control signal. This task is challenging because the model needs to generate the object at the target area while inpainting the original area simultaneously. In the 3rd and 4th rows of Fig. 4, thanks to the proposed Attention ReMix module, our approach enables moving the “parrot” from the right of the branch to left by shifting the position of the parrot sketch in 1st condition frame. The background and the object successfully remain consistent in different video frames.

Human motion editing. Our approach is also capable of editing human motion just replacing the source skeleton condition to the target sequences extracted from certain videos. For instance, we can modify man’s motion from “jumping” to “raising hand” in the 5th and 6th rows in Fig. 4. The result shows that we can generate new content using target pose sequences while maintaining the human appearance in the source video.

Object appearance editing. Thanks to the broad knowledge of per-trained T2I models , we modify the object appearance through the editing of text prompts which is similar to current mainstream video editing methods. Differently, we can further fulfill the property and appearance editing simultaneously. The 7th and 8th rows in Fig. 4 showcase this powerful capability. We replace “bear” with “lion” simply by modifying the corresponding words and enlarging its sketch.

2 Comparisons

Qualitative results. Our method enables multiple possibilities for editing. Here we give two types since they can be directly compared with other methods. The first is Shape-aware editing. As shown in Fig 5, the most naive method is to segment the objects and direct inpainting and paste. However, we find this naive method struggles with the video harmonization and temporal artifacts caused by inpainting and segmentation as shown in Fig. 5. Another relevant method is . However, its appearance is still limited by the optimization, causing blur and unnatural results. We also compare the method by naively using Tune-A-Video with T2I-Adapter . It accomplishes local editing after overfitting but still exhibits temporal inconsistency. Differently, the proposed methods manage to resize the “swan” and edit its appearance to “duck” simultaneously, while maintaining the temporal coherence. On the other hand, we compare our method on conditional video generation from handcrafted motion conditions and a single appearance. As shown in Fig. 5, VideoComposer struggles to generate temporally consistent results. In contrast, our method showcases a natural appearance and better temporal consistency.

Quantitative results. We also conduct the quantitative results of the experiments to show the advantage of the proposed framework with the following metrics: 1) Tem-Con: Following the previous methods , we perform the quantitative evaluation using the pretrained CLIP model. Specifically, we evaluate the temporal consistency in frames by calculating the cosine similarity between all pairs of consecutive frames. 2) Fram-Acc: We execute the appearance editing and report the matrix about frame-wise editing accuracy, which is the percentage of frames. In detail, the edited frame will exhibit a higher CLIP similarity to the target prompt than the source prompt. 3) Four user studies metrics: Following FateZero , we assess our approach using four user studies metrics (“Edit”, “Image”, “Temp” and “ID”). They are editing quality, overall frame-wise image fidelity, temporal consistency of the video, and object appearance consistency, respectively. For a fair comparison, we ask 20 subjects to rank different methods. Each study displays four videos in random order and requests the evaluators to identify the one with superior quality. From Tab. 1, the proposed method achieves the best temporal consistency against baselines and shows a comparable object appearance consistency as framewise inpainting and pasting. As for the user studies, the average ranking of our method earns user preferences the best in four aspects.

3 Ablation Studies

In the right column of Fig. 6, we present the cases without LoRA in cross-attention or token embedding tunning. The visualized result in 3rd column illustrates the significance of cross-attention LoRA. The preservation of object appearance is unattainable without it. We also observe a decline in performance when token embedding is not tuned, further indicating that this operation is crucial for preserving content. In contrast, our framework ensures the consistency of texture and appearance by finetuning both components.

Attention Remix Module

is studied in Fig. 7, where we remove the structure-guided module to ablate its role during inversion and generation. The 3rd column shows that removing the structure-guided module in inversion led to the failure of moving editing. Since the lack of target area mask (left bottom of 3rd column), the self-attention in this region entirely derives from that of the background. The 4th column demonstrates the absence of guidance during generation. It also fails to achieve the task even with the target area mask (left bottom of 4th column). We analyze that there is no guidance for the target area during the generation (left top of 4th column). Finally, we also visualize the result without attention remix module both two stages in the 5th column. The framework severely degrades to be a self-attention reconstruction. We observe that it fails to accomplish the moving task and maintain background consistency (red rectangles in the 5th column). In contrast, when we equip the guidance both in inversion and generation, the “cup” can be shifted successfully, which further emphasizes the significance of our module.

Temporal modules.

We ablate the temporal modules in our framework, including spatial-temporal self-attention and the temporal layer. In the 3rd column of Fig. 8, we notice that removing the spatial-temporal self-attention results in temporal artifacts among frames, emphasizing its importance for temporal consistency. The 4th column illustrates the impact when the temporal layer is removed, showing its effects on the temporal stability of the generated video.

Quantitative Ablation.

We employ a similar setting with the comparison with baseline in Sec. 4.2. In Tab. 2, we note that the ID performance experiences a significant decline without token embedding tunning. When removing spatial self-attention, the temporal coherence of the generated video is significantly disrupted, further underscoring the crucial role of modules.

Conclusion

In this paper, we propose a new controllable video editing method MagicStick that performs temporal consistent video property editing such as shape, size, location, motion, or all that can also be edited in videos. To the best of our knowledge, our method is the first to demonstrate the capability of video editing using a trained text-to-image model. To achieve this, we make an attempt to study and utilize the transformations on one control signal (e.g., edge maps of objects or human pose) using customized ControlNet. A new Attention ReMix module is further proposed for more complex video property editing. Our framework leverages pre-trained image diffusion models for video editing, which we believe will contribute to a lot of new video applications.

While our method achieves impressive results, it still has some limitations. During geometry editing, it struggles to modify the motion of the object to follow new trajectories significantly different from those in the source video. For instance, making a swan take flight or a person perform a full Thomas spin is quite challenging for our framework. We will test our method on the more powerful pretrained video diffusion model for better editing abilities.

Acknowledgments.

We thank Jiaxi Feng, Yabo Zhang for their helpful comments. This project was supported by the National Key R&D Program of China under grant number 2022ZD0161501.

References