LivePhoto: Real Image Animation with Text-guided Motion Control

Xi Chen, Zhiheng Liu, Mengting Chen, Yutong Feng, Yu Liu, Yujun Shen, Hengshuang Zhao

Introduction

Image and video content synthesis has become a burgeoning topic with significant attention and broad real-world applications. Fueled by the diffusion model and extensive training data, image generation has witnessed notable advancements through powerful text-to-image models and controllable downstream applications . In the realm of video generation, a more complex task requiring spatial and temporal modeling, text-to-video has steadily improved . Various works also explore enhancing controllability with sequential inputs like optical flows, motion vectors, depth maps, etc.

This work explores utilizing a real image as the initial frame to guide the “content” and employ the text to control the “motion” of the video. This topic holds promising potential for a wide range of applications, including meme generation, production advertisement, film making, etc. Previous image-to-video methods mainly focus on specific subjects like humans or could only animate synthetic images. GEN-2 and Pikalabs animate real images with an optional text input, however, an overlooked issue is that the text could only enhance the content but usually fails to control the motions.

Facing this challenge, we propose LivePhoto, an image animation framework that truly listens to the text instructions. We first establish a powerful image-to-video baseline. The initial step is to equip a text-to-image model (i,e., Stable Diffusion) with the ability to refer to a real image. Specifically, we concatenate the image latent with input noise to provide pixel-level guidance. In addition, a content encoder is employed to extract image patch tokens, which are injected via cross-attention to guide the global identity. During inference, a noise inversion of the reference image is introduced to offer content priors. Afterward, following the contemporary methods , we freeze stable diffusion models and insert trainable motion layers to model the inter-frame temporal relations.

Although the text branch is maintained in this strong image-to-video baseline, the model seldom listens to the text instructions. The generated videos usually remain nearly static, or sometimes exhibit overly intense movements, deviating from the text. We identify two key issues for the problem: firstly, the text is not sufficient to describe the desired motion. Phrases like “shaking the head” or “camera zooms in” lack important information like moving speed or action magnitude. Thus, a starting frame and a text may correspond to diverse motions with varying intensities. This ambiguity leads to difficulties in linking text and motion. Facing this challenge, we parameterize the motion intensity using a single coefficient, offering a supplementary condition. This approach eases the optimization and allows users to adjust motion intensity during inference conveniently. Another issue arises from the fact that the text contains both content and motion descriptions. The content descriptions translated by stable diffusion may not perfectly align with the reference image, while the image is prioritized for content control. Consequently, when the content descriptions are learned to be suppressed to mitigate conflicts, motion descriptions are simultaneously under-weighted. To address this concern, we propose text re-weighting, which learns to accentuate the motion descriptions, enabling the text to work compatibly with the image for better motion control.

As shown in Fig. 1, equipped with motion intensity guidance and text re-weighting, LivePhoto demonstrates impressive abilities for text-guided motion control. LivePhoto is able to deal with real images from versatile domains and subjects, and adequately decodes the motion descriptions like actions and camera movements. Besides, it shows fantastic capacities of conjuring new contents from thin air, like “pouring water into a glass” or simulating “lightning and thunder”. In addition, with motion intensity guidance, LivePhoto supports users to customize the motion with the desired intensity.

Related Work

Image animation. To realize content controllable video synthesis, image animation takes a reference image as content guidance. Most of the previous works depend on another video as a source of motion, transferring the motion to the image with the same subject. Other works focus on specific categories like fluide or nature objects . Make-it-Move uses text control but it only manipulates simple geometries like cones and cubes. Recently, human pose transfer methods convert the human images to videos with extra controls like dense poses, depth maps, etc. VideoComposer could take image and text as controls, however, the text shows limited controllability for the motion and it usually requires more controls like sketches and motion vectors. In general, existing work either requires more controls than text or focuses on a specific subject. In this work, we explore constructing a generalizable framework for universal domains and use the most flexible control (text) to customize the generated video.

Text-to-video generation. Assisted by the diffusion model , the field of text-to-video has progressed rapidly. Early attempts train the entire parameters, making the task resource-intensive. Recently, researchers have turned to leveraging the frozen weights of pre-trained text-to-image models tapping into robust priors. Tune-A-Video inflates the text-to-video model and tuning attention modules to construct an inter-frame relationship with a one-shot setting. Align-Your-Lantens inserts newly designed temporal layers into frozen text-to-image models to make video generation. AnimateDiff proposes to freeze the stable diffusion blocks and add learnable motion modules, enabling the model to incorporate with subject-specific LoRAs to make customized generation. A common issue is that the text could only control the spatial content of the video but exert limited effect for controlling the motions.

Method

We first give a brief introduction to the preliminary knowledge for diffusion-based image generation in Sec. 3.1. Following that, our comprehensive pipeline is outlined in Sec. 3.2. Afterward, Sec. 3.3 delves into image content guidance to make the model refer to the image. In Sec. 3.4 and Sec. 3.5, we elaborate on the novel designs of motion intensity guidance and text re-weighting to better align the text conditions with the video motion.

Text-to-image with diffusion models. Diffusion models show promising abilities for both image and video generation. In this work, we opt for the widely used Stable Diffusion as the base model, which adapts the denoising procedure in the latent space with lower computations. It initially employs VQ-VAE as the latent encoder to transform an image x0\mathbf{x}_{0} into the latent space: z0=E(x0)\mathbf{z}_{0}=\mathcal{E}(\mathbf{x}_{0}). During training, Stable Diffusion transforms the latent into Gaussian noise as follows:

where the noise ϵ∼U()\mathbf{\epsilon}\sim\mathcal{U}(), and αtˉ\bar{\alpha_{t}} is a cumulative products of the noise coefficient αt\alpha_{t} at each step. Afterward, it learns to predict the added noise as:

tt is the diffusion timestep, c\mathbf{c} is the condition of text prompts. During inference, Stable Diffusion is able to recover an image from Gaussian noise step by step by predicting the noise added for each step. The denoising results are fed into a latent decoder to recover the colored images from latent representations as x^0=D(z^0)\mathbf{\hat{x}}_{0}=\mathcal{D}(\mathbf{\hat{z}}_{0}).

2 Overall Pipeline

The framework of LivePhoto is demonstrated in Fig. 2. The model takes a reference image, a text, and the motion intensity as input to synthesize the desired video. When the ground truth video is provided during training, the reference image is picked from the first frame, and the motion intensity is estimated from the video. During inference, users could customize the motion intensity or directly use the default level. LivePhoto utilizes a 4-channel tensor of zB×F×C×H×W\mathbf{z}^{B\times F\times C\times H\times W} to represent the noise latent of the video, where the dimensions mean batch, frame, channel, height, and width, respectively. The reference latent is extracted by VAE encoder to provide local content guidance. Meanwhile, the motion intensity is transformed to a 1-channel intensity embedding. We concatenate the noise latent, the reference latent, the intensity embedding, and a frame embedding to form a 10-channel tensor for the input of UNet. At the same time, we use a content encoder to extract the visual tokens of the reference image and inject them via cross-attention. A text re-weighting module is added after the text encoder , which learns to assign different weights to each part of the text to accentuate the motion descriptions of the text. Following modern text-to-video models . We freeze the stable diffusion blocks and add learnable motion modules at each stage to capture the inter-frame relationships.

3 Image Content Guidance

The most essential step is enabling LivePhoto to keep the identity of the reference image. Thus, we collect local guidance by concatenating the reference latent at the input. Moreover, we employ a content encoder to extract image tokens for global guidance. Additionally, we introduce the image inversion in the initial noise to offer content priors.

Reference latent. We extract the reference latent and incorporate it at the UNet input to provide pixel-level guidance. Simultaneously, a frame embedding is introduced to impart temporal awareness to each frame. Thus, the first frame could totally trust the reference latent. Subsequent frames make degenerative references and exhibit distinct behavior. The frame embedding is represented as a 1-channel map, with values linearly interpolated from zero (first frame) to one (last frame).

Content encoder. The reference latent effectively guides the initial frames due to their higher pixel similarities. However, as content evolves in subsequent frames, understanding the image and providing high-level guidance becomes crucial. Drawing inspiration from , we employ a frozen DINOv2 to extract patch tokens from the reference image. We add a learnable linear layer after DINOv2 to project these tokens, which are then injected into the UNet through newly added cross-attention layers.

Prior inversion. Previous methods prove that using an inverted noise of the reference image, rather than a pure Gaussian noise, could effectively provide appearance priors. During inference, we add the inversion of the reference latent r0\mathbf{r}_{0} to the noise latent zTn\mathbf{z}_{T}^{n} of frame nn at the initial denoising step (T), following Eq. 3.

where αn\alpha^{n} is a descending coefficient from the first frame to the last frame. We set αn\alpha^{n} as a linear interpolation from 0.033 to 0.016 by default.

4 Motion Intensity Estimation

It is challenging to align the motion coherently with the text. We analyze the core issue is that the text lacks descriptions for the motion speed and magnitude. Thus, the same text leads to various motion intensities, creating ambiguity in the optimization process. To address this, we leverage the motion intensity as an additional condition. We parameterize the motion intensity using a single coefficient. Thus, the users could adjust the intensity conveniently by sliding a bar or directly using the default value.

In our pursuit of parameterizing motion intensity, we experimented with various methods, such as calculating optical flow magnitude, computing mean square error between adjacent frames, and leveraging CLIP/DINO similarity between frames. Ultimately, we found that Structural Similarity (SSIM) produces results the most aligned with human perceptions. Concretely, given a training video clip Xn\mathbf{X}^{n} with n frames, we determine its motion intensity I\mathbf{I} by computing the average value for the SSIM between each adjacent frame as in Eq. 4 and Eq. 5:

The structure similarity considers the luminance (ll), contrast (cc), and structure (ss) differences between two images. By default, α\alpha, β\beta, and γ\gamma are set as 1.

We compute the motion intensity on the training data to determine the overall distribution and categorize the values into 10 levels. We create a 1-channel map filled with the level numbers and concatenate it with the input of UNet. During inference, users can utilize level 5 as the default intensity or adjust it between levels 1 to 10. Throughout this paper, unless specified, we use level 5 as the default.

5 Text Re-weighting

Another challenge in instructing video motions arises from the fact that the text prompt encompasses both “content descriptions” and “motion descriptions”. The “content descriptions”, translated by the frozen Stable Diffusion, often fail to perfectly align with the reference images. When we expect the text prompts to guide the motion, the “content descriptions” are inherently accentuated simultaneously. However, as the reference image provides superior content guidance, the effect of the whole text would be suppressed when content conflicts appear.

To accentuate the part related to the “motion descriptions”, we explore manipulating the CLIP text embeddings. Recognizing that directly tuning the text encoder on limited samples might impact generalization, we assign different weights for each embedding without disrupting the CLIP feature space. Concretely, we add three trainable transformer layers and a linear projection layer after the CLIP text embeddings. Afterward, the predicted weights are normed from 0 to 1 with a sigmoid function. These weights are then multiplied with the corresponding text embeddings, thereby providing guidance that focuses on directing the motions. The comprehensive structure of the text re-weighting module and actual examples are depicted in Fig. 3. The numerical results prove that the module successfully learns to emphasize the “motion descriptions”. This allows signals from images and texts to integrate more effectively, resulting in stronger text-to-motion control.

Experiments

Training configurations. We implement LivePhoto based on the frozen Stable Diffusion v1.5 . The structure of our Motion Module aligns with AnimateDiff . Our model is trained on the WebVID dataset employing 8 A100 GPUs. We sample training videos with 16 frames, perform center-cropping, and resize each frame to 256×256256\times 256 pixels. For classifier-free guidance, we utilize a 0.5 probability of dropping the text prompt during training. We only use a simple MSE loss to train the model.

Evaluation protocols. We conduct user studies to compare our approach with previous methods and analyze our newly designed modules. To validate the generalization ability, we gather images from various domains encompassing real images and cartoons including humans, animals, still objects, natural sceneries, etc. For quantitative assessment, we utilize the validation set of WebVID . The first frame and prompt are used as controls to generate videos. We measure the average CLIP similarity and DINO similarity between adjacent frames to evaluate the frame consistency following previous works .

2 Ablation Studies

In this section, we thoroughly analyze each of our proposed modules to substantiate their effectiveness. We first analyze how to add content guidance with the reference image, which is an essential part of our framework. Following that, we delve into the specifics of our newly introduced motion intensity guidance and text re-weighting.

Image content guidance. As introduced in Sec. 3.2, we concatenate the reference latent with the input as the pixel-wise guidance and use a content encoder to provide the holistic identity information. Besides, the prior inversion further assists the generation of details. In Fig. 4, we illustrate the step-by-step integration of these elements. In row 1, the reference latent could only keep the identity for the starting frames as the contents are similar to the reference image. After adding the content encoder in row 2, the identity for the subsequent frames could be better preserved but the generation quality for the details is not satisfactory. With the inclusion of prior inversion, the overall quality sees further improvement. The quantitative results in Tab. 1 consistently confirm the effectiveness of each module. These three strategies serve as the core of our strong baseline for real image animation.

Motion intensity guidance. As introduced in Sec. 3.4, we parameterize the motion intensity as a coefficient, and use it to indicate the motion speed and ranges. We carry out ablation studies in Fig. 5. The absence of motion intensity guidance often leads to static or erratic video outputs, as depicted in the first row. However, with the introduction of intensity guidance, the subsequent rows display varying motion levels, allowing for the production of high-quality videos with different motion ranges. Notably, lower levels like level 2 generate almost static videos, while higher levels like 10 occasionally produce overly vigorous motions. Users could directly use the default value (level 5) or tailor the intensity according to specific preferences.

Text re-weighting. In Fig. 6, we demonstrate the efficacy of text re-weighting. In the given examples, the content description “baby dinosaur” would conflict with the reference image. In the first three rows, without the assistance of re-weighting, the frozen Stabel Diffusion tends to synthesize the content through its understanding of the text. Thus, the produced video tends to ignore the text and follow the reference image as in row 1. In other cases, it has risks of becoming a “baby” (row 2) or a “dinosaur” (row 3). As visualized in the bottom of Fig. 6, text re-weighting elevates emphasis on motion descriptions like “waving its hand”. This approach enables our model to faithfully follow text-based instructions for motion details while upholding image-consistent content with the reference image.

The quantitative results are listed in Tab. 2. The motion intensity guidance and text re-weighting both contribute to the frame consistency.

3 Comparisons with Existing Alternatives

We compare LivePhoto with other works that support image animation with text control. VideoComposer is a strong compositional generator covering various conditions including image and text. GEN-2 and Pikalabs are famous products that support image and text input. I2VGEN-XL , AnimateDiff-I2V , Talesofai are open-source projects claiming similar abilities.

Qualitative analysis. In Fig. 7, we compare LivePhoto with VideoComposer , Pikalabs , and GEN-2 with representative examples. The selected examples cover animals, humans, cartoons, and natural scenarios. To reduce the randomness, we ran each method 8 times to select the best result for more fair comparisons. VideoComposer demonstrates proficiency in creating videos with significant motion. However, as not specifically designed for photo animation, the identity-keeping ability is not satisfactory. The identities of the reference images are lost, especially for less commonly seen subjects. Additionally, it shows a lack of adherence to the provided text instructions. Pikalabs and GEN-2 produce high-quality videos. However, as a trade-off, the generated videos own limited motion ranges. Although they support text as supplementary, the text descriptions seldom work. The motions are generally estimated from the content of the reference image.

In contrast, LivePhoto adeptly preserves the identity of the reference image and generates consistent motions with the text instructions. It performs admirably across various domains, encompassing animals, humans, cartoon characters, and natural sceneries. It not only animates specific actions (examples 1-4) but also conjures new effects from thin air (examples 5-6).

We also compare LivePhoto with open-sourced project in Fig. 8. I2VGEN-XL does not set the reference image as the first frame but generates videos with similar semantics. AnimateDiff-I2V and Talsofai are extensions of AnimateDiff . However, the former produces quasi-static videos. The latter fails to keep the image identity unless using SD-generated images with the same prompt and corresponding LoRA .

User studies. Metrics like DINO/CLIP scores have limitations in thoroughly evaluating the model, thus, we carry out user studies. We ask the annotators to rate the generated videos from 4 perspectives: Image consistency evaluates the identity-keeping ability of the reference image. Text consistency measures whether the motion follows the text descriptions. Content quality considers the general quality of videos like the smoothness, the resolution, etc. Motion quality assesses the reasonableness of generated motion, encompassing aspects such as speed and deformation.

We construct a benchmark with five tracks: humans, animals, cartoon characters, still objects, and natural sceneries. We collect 10 reference images per track and manually write 2 prompts per image. Considering the variations that commonly exist in video generation, each method is required to predict 8 results. Thus, we get 800 samples for each method. We first ask 4 annotators to pick the best ones out of 8 predictions according to the aforementioned four perspectives. Then, we ask 10 annotators to further rate the filtered samples. As the projects demonstrates evidently inferior results, we only compare LivePhoto with VideoComposer , GEN-2 , and Pikalabs .

Results in Tab. 3 demonstrate that GEN-2 and Pikalabs own slightly better image consistency because their generated video seldom moves. LivePhoto shows significantly better text consistency and motion quality compared with other works. We admit that GEN-2 and Pikalabs own superior smoothness and resolution. We infer that they might collect much better training data and leverage super-resolution networks as post-processing. However, as an academic method, LivePhoto shows distinguishing advantages over mature products in certain aspects. We have reasons to believe its potential for future applications.

Limitations

LivePhoto is implemented on SD-1.5 with 256×256256\times 256 output considering the training cost. We believe that with higher resolution and stronger models like SD-XL , the overall performance could be further improved significantly.

Conclusion

We introduce LivePhoto, a novel framework for photo animation with text control. We propose a strong baseline that gathers the image content guidance from the given image and utilizes motion intensity as a supplementary to better capture the desired motions. Besides, we propose text re-weighting to accentuate the motion descriptions. The whole pipeline illustrates impressive performance for generalized domains and instructions.

References