MagicEdit: High-Fidelity and Temporally Coherent Video Editing

Jun Hao Liew, Hanshu Yan, Jianfeng Zhang, Zhongcong Xu, Jiashi Feng

Introduction

Video editing plays an ubiquitous role in creating fascinating visual effects for films, short videos, etc. However, professional editing is not only complex and time-consuming, but also challenging for novice users. As a result, there is an increasing demand for easy-to-use and performant video editing tools. Recently, we have witnessed a rapid development of video editing algorithms thanks to the introduction of powerful text-conditioned diffusion models trained on large-scale datasets (e.g., DALL-E 2 , Imagen , Stable Diffusion ). In general, there are two ways to extend image diffusion models for the video editing tasks: per-frame methods and per-clip methods.

Per-frame methods treat a video clip as a sequence of frames and run image editing on each frame independently. These methods often require some ad-hoc tricks to reduce temporal inconsistency, such as cross-frame attention , flow warping , latents matching/ fusion etc. However, these strategies can only maintain high-level styles and coarse shapes, and are less effective in preserving fine-grained details and texture across frames. In addition, these methods often struggle when there exists large motion.

Per-clip methods, on the other hand, treat a video clip as a 3D spatio-temporal volume and directly edit the entire video. These methods typically inflate the image diffusion model into a video model by adding temporal layers. Among these, a popular line of research is to either fine-tune the pre-trained model on the input video to generate videos with similar motion, or utilize Null-text Inversion for video inversion. However, since fine-tuning or optimization is needed for every input video, these methods suffer from low efficiency. Gen-1 , on the other hand, incorporates temporal-aware structures and learns motion priors from large-scale video datasets, demonstrating remarkable video editing performance. In general, compared to per-frame methods, temporal inconsistency is typically less of an issue for per-clip methods due to the explicit modeling of motion signal. Nevertheless, since these methods update the whole networks, the domain knowledge of the original text-to-image model is inevitably hurt, resulting in degradation of per-frame quality.

In this report, we discover a surprisingly simple yet effective recipe for text-guided video editing, i.e., by explicitly disentangling the learning of content, structure and motion during training, we can easily achieve high-fidelity, temporally consistent video-to-video translation. With this, we present MagicEdit, which supports a variety of downstream editing tasks, including video stylization, local editing and video-MagicMix and video outpainting.

MagicEdit

Given a video sequence of dimension F×H×W×3F\times H\times W\times 3, where FF is the number of frames, H,WH,W are height and width, respectively, and a prompt description c{\bm{c}} (e.g., “a pretty girl, pink dress” in Fig. 1), our goal is to edit the content of the video while preserving its structure. Specifically, we solve this task by learning a generative model p(x∣c,s)p({\bm{x}}|{\bm{c}},{\bm{s}}) of videos x=[x1,⋯ ,xF]{\bm{x}}=[{\bm{x}}_{1},\cdots,{\bm{x}}_{F}], conditioned on text prompt c{\bm{c}} and structure representation s=[s1,⋯ ,sF]{\bm{s}}=[{\bm{s}}_{1},\cdots,{\bm{s}}_{F}]. We mathematically formulate this as:

where L\mathcal{L} refers to the noise estimation loss and Θ={θc\mathbf{\Theta}=\{\theta_{\rm c}, θs\theta_{\rm s}, θm}\theta_{\rm m}\}. In specific, θc\theta_{\rm c} represents the UNet parameters of the text-to-image generation model; θs\theta_{\rm s} refers to the parameters of the structure conditioning module; and θm\theta_{\rm m} denotes the parameters of temporal/motion layers. We explicitly disentangle the modeling of content, structure and motion via stage-wise training as follows:

Stage I: Text-to-Image generation. In the first stage, we train a base text-to-image (T2I) diffusion model to encourage each generated frame xi{\bm{x}}_{i} to adhere to the given text prompt c{\bm{c}}.

In this work, we choose stable-diffusion-v1-5 as our base T2I generation model.

Stage II-A: Structure-conditioned generation. An important property of video editing is to ensure that each video frame follows the structure or trajectory of the source video (e.g., depth/ pose/ shape, etc.). For example, given a video of a girl swinging her arms (Fig. 1), each edited frame xi{\bm{x}}_{i} should follow the corresponding pose si{\bm{s}}_{i} in the given video.

To achieve this, we train a structure-conditioned module parameterized by θs\theta_{\rm s} while freezing the pre-trained UNet parameters θˉc\bar{\theta}_{\rm c}. We follow the approach of ControlNet for per-frame structure-preserving generation.

Stage II-B: Temporally consistent video generation. Lastly, to ensure the video frames remain temporally coherent, we train a motion module to enforce cross-frame consistency. For simplicity, we choose the vanilla temporal transformers from as the design of our motion module. Once again, we train the motion module, parameterized by θm\theta_{\rm m} while freezing the learned UNet parameters θˉc\bar{\theta}_{\rm c}.

Inference. During inference, we simply combine the three individually trained modules. It is worth noting that, since the base T2I weights remain frozen when training the structure-conditioned and motion module, during inference, we can simply swap the base T2I weights (stable-diffusion-v1-5) with different personalized Stable Diffusion models from CivitAI https://civitai.com/ for different styles and better appearance (Fig. 2 bottom right).

Implementation details. Vertical and horizontal videos are resized such that its shorter size is 320. Square videos are resized to 512×\times512. For each video clip, we sample 16 frames with fixed interval. We use 25 step DDIM sampler and set the classifier-free guidance scale to 7.5. Following , we employ a linear beta schedule.

Discussion. We argue that the most important key to high-fidelity and temporally coherent video editing lies in the explicit disentanglement of the three modules during training. More specifically, the text-to-image diffusion UNet should remain frozen when training the structure-following module and motion module. This is because video training data is of lower quality as compared to image data and often consists of motion blur. In other words, jointly modeling all the three components, as done in most existing works, would lead to degraded per-frame quality. To counter this, one needs to collect large-scale high quality video data, which is prohibitively expensive.

Note that, while none of these components/ modules are new, the main contribution of this work is to showcase that, explicitly disentangling the three sources of signal is the key towards high-quality temporally smooth video editing. We hope that this finding could shed lights on future video generation and editing research.

Applications

Next, we discuss the possible applications of MagicEdit, including stylization, local editing, video-MagicMix and video outpainting.

Video stylization. Video stylization enables one to (1) transform the source video into a new video with a style-of-interest (e.g., realistic, cartoon), or (2) creating a new scene with different subject (e.g., dog →\rightarrow cat) and different background (e.g., living room →\rightarrow beach). Given a source video, we first extract its structure representation (e.g., we extract disparity maps with MiDaS or human pose with OpenPose ). Next, we swap base T2I weights with different personalized models from CivitAI (e.g., RealisticVision https://civitai.com/models/4201, majicMix Realistic https://civitai.com/models/43331, Disney Pixar Cartoon Type A https://civitai.com/models/65203) for different styles. Following , for each personalized model, we follow the prompts format provided at the model homepage. We show some examples in Fig. 3.

Local editing. There are cases when a user only wants to make local modification to the video while leaving other regions untouched (e.g., make the young lady wear glasses as shown in Fig. 1). To achieve this, following SDEdit , we first invert the source video via DDIM inversion with a source prompt csrc{\bm{c}}_{\rm src} describing the original video content. Then, we run the denoising process as usual but with the target prompt c{\bm{c}}. Some examples can be found in Fig. 4

Video-MagicMix. Liew et al. previously demonstrated that two different concepts can be mixed to construct a new concept (e.g., “rabbit” + “tiger” →\rightarrow a rabbit-alike tiger). Similarly, we show that MagicMix can be applied to video domain to create a moving rabbit-alike tiger (Fig. 5).

Video Outpainting. We found that MagicEdit can also be applied for video outpainting task without any re-training. Given an input video of spatial size H×WH\times W, let the size of the outpainted video be h×wh\times w. We first invert the source video via DDIM inversion, obtaining a sequence of latents of size H/8×W/8H/8\times W/8. Then, we randomly sample FF Gaussian noises of size ×h/8×w/8\times h/8\times w/8 and run denoising. At each denoising step, we replace the known regions with the inverted latents above to ensure the known areas remain unchanged. To ensure smooth transition across image borders, we do not replace the known latents for the last few steps. Some examples of video outpainting are shown in Fig. 6.

In Fig. 7, we can also see that the model can handle various ratios, including horizontal, vertical, and even large outpainting ratio (e.g., bottom + 100%). More interestingly, as shown in Fig. 8, our model is also capable to generate different contents by giving different text prompts (e.g., short or long pants), allowing the users to outpaint a video flexibly.

Conclusion

In this technical report, we present MagicEdit, a surprisingly simple recipe for effective training of a video editing tool. Our findings show that, high-fidelity and temporally coherent video-to-video translation can be obtained by explicitly disentangling the learning of content, structure and motion signals during training. As a result, MagicEdit supports a wide variety of downstream editing applications, including video stylization, local editing, video-MagicMix and video outpainting.

References