SparseCtrl: Adding Sparse Controls to Text-to-Video Diffusion Models
Yuwei Guo, Ceyuan Yang, Anyi Rao, Maneesh Agrawala, Dahua Lin, Bo Dai
Introduction
With the advance of text-to-image (T2I) generation and large-scale text-video paired datasets , there has been a surge of progress in the field of text-to-video (T2V) generative models . These developments enable users to generate compelling videos through textual descriptions of the desired content. Nonetheless, textual prompts, being inherently abstract expressions, struggle to accurately define complex structural attributes such as spatial layouts, poses, and shapes. This lack of precise control impedes its practical application in more demanding and professional contexts, such as anime creation and filmmaking. Consequently, users often find themselves engrossed in numerous rounds of random trial-and-error to achieve their desired outputs. This process can be time-consuming, especially since there is no straightforward method to guide the synthetic results toward the expected direction during the iterative trying process.
To unlock the potential of T2V generation, efforts have been made to incorporate more precise control through structural information. For instance, Gen-1 pioneers using monocular depth maps as structural guidance. VideoComposer and DragNUWA investigate the domain of compositional video generation, employing diverse modalities such as depth, sketch, and initial image as control signals. Furthermore, prior studies utilize the image ControlNet to introduce various controlling modalities to video generation. By harnessing additional structural sequences, these approaches provide enhanced control capabilities. However, for precise output control, existing works necessitate temporally dense structural map sequences, which means that users need to furnish condition maps for each frame in the generated video, thereby increasing the practical costs. Additionally, most approaches towards controllable T2V typically redesign the model architecture to accommodate the extra condition input, which demands costly model retraining. Such practice is inefficient when a well-trained T2V model is already available or when there is a requirement to incorporate a new control modality into a pre-trained generator.
In this paper, we introduce SparseCtrl, an efficient approach that targets controlling text-to-video generation via temporally sparse condition maps with an add-on encoder. More specifically, in order to control the synthesis, we apply the philosophy of ControlNet , which implements an auxiliary encoder while preserving the integrity of the original generator. This design allows us to incorporate additional conditions by merely training the encoder network on top of the pre-trained T2V model, thereby eliminating the need for comprehensive model retraining. Additionally, this design facilitates control over not only the original T2V but also the derived personalized models when combined with the plug-and-play motion module of AnimateDiff . To achieve this, we design a condition encoder equipped with temporal-aware layers that propagate the sparse condition signals from conditioned keyframes to unconditioned frames. Significantly, we find that purging the noised sample input in the vanilla ControlNet further prevents potential quality degradation in our scenario. Moreover, we apply widely used masking strategies during training to accommodate varying degrees of sparsity and tackle a broad range of application scenarios.
We evaluate SparseCtrl by training three encoders on sketches, depth, and RGB images. Experimental results show that users can manipulate the structure of the synthetic videos by providing just one or a few input condition maps. Comprehensive ablation studies are performed to investigate the contribution of each component. We additionally show that by integrating with plug-and-play video generation backbone such as AniamteDiff , our method exhibits compatibility and excellent visual quality with various personalized text-to-image models. Leveraging this sparse control approach, SparseCtrl enables a broad range of applications. For instance, the sketch encoder empowers users to transform hand-drawn storyboards into dynamic videos; The depth encoder provides the ability to render videos by supplying a minimum number of depth maps; Furthermore, the RGB image encoder unifies multiple tasks, including image animation, keyframe interpolation, video prediction, etc. We anticipate that this work will contribute towards bridging the gap between text-to-video research and real-world content creation processes.
Related Works
Text-to-video diffusion models. The field of text-to-video (T2V) generation has witnessed significant progression recently, driven by advancements in diffusion models and large-scale text-video paired datasets . Initial attempts in this area focus on training a T2V model from scratch. For example, Video Diffusion Model expands the standard image architecture to accommodate video data and trains on both image and video together. Imagen Video employs a cascading structure for high-resolution T2V generation, while Make-A-Video uses a text-image prior model to reduce reliance on text-video paired data. Others turn to build T2V models upon powerful text-to-image (T2I) models such as Stable Diffusion , by incorporating additional layers to model cross-frame motion and consistency . Among these, MagicVideo utilizes a causal design and executes training in a compressed latent space to mitigate computational demands. Align-Your-Latents efficiently turns T2I into video generators by aligning independently sampled noise maps. AnimateDiff utilizes a pluggable motion module to enable high-quality animation creation on personalized image backbones . Other contributions include noise prior modeling , training on high-quality datasets , and latent-pixel hybrid space denoising , all leading to remarkable pixel quality. However, current text-conditioned video generation techniques lack fine-grained controllability over synthetic results. In response to this challenge, our work aims to enhance the control of T2V models via an add-on encoder.
Controllable text-to-video generation. Given that a text prompt can often result in ambiguous guidance to the video motion, content, and spatial structure, such controllabilities become crucial factors in T2V generation. For high-level video motion control, several studies propose learning LoRA layers for specific motion patterns , while others employ extracted trajectories , motion vectors , or pose sequence . To manage specific synthetic keyframes for animation or interpolation, recent explorations include encoding the image separately to the generator , concatenating with the noise input , or utilizing multi-level feature injection . For fine-grained spatial structure control, some low-level representations are introduced. Gen-1 is the first to use monocular depth sequences as structural guidance. VideoComposer encodes sketch and depth sequences via a shared encoder, facilitating flexible combinations at inference. Additionally, some approaches utilize readily available image controlling models for controllable video generation . Though these methods achieve fine-grained controllability, they necessitate providing conditions for every synthetic frame, which incurs prohibitive costs in practical applications. In this study, we aim to control video generation through temporally sparse conditions by inputting only a few condition maps, thus making T2V more practical in a broader range of scenarios.
Add-on network for additional control. Training foundational T2I/T2V generative models is computationally demanding. Therefore, a preferred approach to incorporate extra control into these models is to train an additional condition encoder while maintaining the integrity of the original backbone . ControlNet pioneered the potential of training plug-and-play condition encoders for pre-trained T2I models. It involves creating a trainable duplicate of the pre-trained layers that accommodates the condition input. The encoder output is then reintegrated into the T2I model through zero-initialized layers. Similarly, T2I-Adapter utilizes a lightweight structure to infuse control. IP-Adapter , integrates the style condition by transposing the reference image into supplementary embeddings, which are subsequently concatenated with the text embeddings. Our approach aligns with the principles of these works and aims to achieve sparse control through an auxiliary encoder module.
SparseCtrl
To enhance the controllability of a pre-trained text-to-video (T2V) model with temporally sparse signals, we introduce add-on sparse encoders to control the video generation process, leaving the original T2V generator untouched. This section is thus organized as follows: Sec. 3.1 presents the background of T2V diffusion models; Sec. 3.2 discuss the design of our sparse condition encoder, followed by the supported modalities and applications in Sec. 3.3.
Leveraging powerful text-to-image generators. Text-to-image (T2I) generation has been dramatically advanced by powerful image generators like Stable Diffusion . A practical path for T2V tasks is to leverage such powerful T2I priors. Recent T2V model typically extend a pre-trained T2I generator for videos by incorporating temporal layers between the 2D image layers, as illustrated in the lower part of Fig. 2 (a). This arrangement enables cross-frame information exchange, thereby effectively modeling the cross-frame motion and temporal consistency.
Training objectives. The training objectives of T2V models are generally aligned with their image counterparts. Specifically, the model tries to predict the noise scale added to the clean RGB video (or latent features) with frames, encouraged by an MSE loss:
where is the embeddings of the text description, is the sampled Gaussian noise in the same shape of , and are terms that control the added noise strength, is a uniformly sampled diffusion step, is the number of total steps. In the following context, we also adopt this objective for training.
2 Sparse Condition Encoder
To enable efficient sparse control, we introduce an add-on encoder capable of accepting sparse condition maps as inputs, which we call sparse condition encoders. In the T2I domain, ControlNet successfully adds structure control to the pre-trained image generator by partially replicating a copy of the pre-train model and its input, then adding the conditions and reintegrating the output back to the original model through zero-initialized layers, as shown in the left of Fig. 2 (b). Inspired by its success, we start with a similar design to enable sparse control in the T2V setting.
Limited controllability of frame-wise encoder. We start with a straightforward solution: training a ControlNet-like encoder to incorporate sparse condition signals. To this end, we build a frame-wise encoder akin to ControlNet, replicate it across the temporal dimension, and add the conditions to the desired keyframes through this auxiliary structure. For frames that are not directly conditioned, we input a zero image to the encoder and indicate the unconditioned state through an additional mask channel. However, experimental results in Sec. 4.4.1 show that such frame-wise conditions sometimes fail to maintain temporal consistency when used with sparse input conditions, e.g., in the image animation scenario where only the first frame is conditioned. In such cases, only keyframes react to the condition, leading to abrupt content changes between the conditioned and unconditioned frames.
Condition propagation across frames. Considering the sparsity and temporal relationship of given inputs, we hypothesize that the above problem arises because the T2V backbone has difficulty inferring the intermediate condition states for the unconditioned frames. To solve this, we propose to add temporal layers (e.g., temporal attention with position encoding) to the sparse condition encoders that allow the conditional signal to propagate from frame to frame. Intuitively, although not identical, different frames within a video clip share similarities in both appearance and structure. The temporal layers can thus propagate such implicit information from the conditioned keyframes to the unconditioned frames, thereby enhancing consistency. Our experiments confirm that this design significantly improves the robustness and consistency of the generated results.
Quality degradation caused by manually noised latents. Although the sparse condition encoder with temporal layers could tackle the sparsity of inputs, it sometimes leads to visual quality degradation of the generated videos, as shown in Sec. 4.4.1. When examining the design of the vanilla ControlNet, we find that simply applying the ControlNet in our scenario is unsuitable due to the copying of noised sample inputs. Concretely, as illustrated in Fig. 2 (b), the original ControlNet copies not only the UNet encoder but also the noised sample input . Namely, the input for the ControlNet encoder is the sum between the condition (after zero-initialized layers) and the noised sample. This design stabilizes the training and accelerates the model convergence in its original scenario. However, in terms of the unconditioned frames in our setting, the informative input of the sparse encoder becomes only the noised sample. This might encourage the sparse encoder to overlook the condition maps and rely on the noised sample during training, which contradicts our goal of controllability enhancement. Accordingly, as shown in Fig. 2 (b), our proposed sparse encoder eliminates the noised sample input and only accepts the condition maps after concatenation. This straightforward yet effective method eliminates the observed quality degradation in our experiments.
Unifying sparsity via masking. In practice, to unify different sparsity with a single model, we use zero images as the input placeholder for unconditioned frames and concatenate a binary mask sequence to the input conditions, which is a common practice in video reconstruction and prediction . As shown in Fig. 2 (a), we concatenate a mask channel-wise in addition to the condition signals at each frame to form the input of the sparse encoder. Setting indicates the current frame is unconditioned and vice versa. In this way, different sparse input cases can be represented with a unified input format.
3 Multiple Modalities and Applications
In this paper, we implement SparseCtrl with three modalities: sketches, depth maps, and RGB images. Notably, our method is potentially compatible with other modalities, such as skeleton and edge map, which we leave for future developments.
Sketch-to-video generation. Sketches can serve as an efficient guiding tool for T2V due to their ease of creation by non-professional users. With SparseCtrl, users can supply any number of sketches to shape the video content. For instance, a single sketch can establish the overall layout of the video, while sketches of the first, last, and selected intermediate frames can define coarse motion, making the method highly beneficial for storyboarding.
Depth guided generation. Integrating depth conditions with the pre-trained T2V enables depth-guided generation. Consequently, users can render a video by directly exporting sparse depth maps from engines or 3D representations or conduct video translation using depth as an intermediate representation.
Image animation and transition; video prediction and interpolation. Within the context of RGB video, numerous tasks can be unified into a single problem of video generation with RGB image conditions. In this scheme, image animation corresponds to video generation conditioned on the first frame; Transition is conditioned by the first and last frames; Video prediction is conditioned on a small number of beginning frames; Interpolation is conditioned on uniformly sparsed keyframes.
Experiments
In this section, we evaluate SparseCtrl under various settings. Sec. 4.1 present the detailed implementations. Sec. 4.2 showcases the results and applications given one or few conditions. Sec. 4.3 suggests that SparseCtrl could achieve comparable performances on chosen popular tasks with baseline methods, e.g., sparse depth-to-video generation and image animation. Sec. 4.4 present comprehensive ablation studies and evaluate SparseCtrl’s response to textual prompts and unrelated conditions.
Text-to-video generator. We implement SparseCtrl upon AnimateDiff , which can serve as a general T2V generator when integrated with its pretraining image backbone, Stable Diffusion V1.5 , or function as a personalized generator when combined with personalized image backbones such as RealisticVision and ToonYou . We test with both settings and showcase the results.
Training. The training objective of SparseCtrl aligns with Eq. 1. The only difference is the integration of the proposed sparse condition encoder into the pre-trained text-to-video (T2V) backbone. To help the condition encoder learn robust controllability, we adopted a simple strategy to mask out conditions during training. In each iteration, we first randomly sample a number between and to determine how many frames will receive the condition. Subsequently, we draw indices without repeating from and keep the conditions for the corresponding frames. We train SparseCtrl on WebVid-10M and extract the corresponding conditions on the fly. More details can be found in the supplementary material.
2 Main Results
We showcase the qualitative results and applications of SparseCtrl with three modalities in Fig. 1, 3, and the supplementary material, covering original and personalized T2V settings. As shown in the figure, with SparseCtrl, the synthetic videos closely adhere to control signals and maintain an excellent temporal consistency, being robust to different numbers of conditioning frames.
Remarkably, by drawing a single sketch, we can trigger the capability of the pre-trained T2V model to generate rare semantic compositions, such as a panda standing on a surfboard shown in the first row of Fig. 3. In contrast, the pre-trained T2V model struggles to generate such complex samples using textual descriptions alone. This suggests that the full potential of the T2V, pre-trained on large-scale datasets, may not be fully unlocked with only textual guidance. Additionally, we show that with well-learned real-world motion knowledge, the pre-trained T2V is capable of inferring the intermediate states with as few as two conditions, as illustrated in 3/5-th rows in Fig. 3. This indicates that temporally dense control might not be necessary.
3 Comparisons on Popular Tasks
Since it is challenging to compare SparseCtrl against prior efforts on all applications that we could enable, we choose two popular tasks for evaluation: sparse depth-to-video generation and image animation. For the first task, dense depth condition mode of VideoComposer (VC) and Text2Video-Zero (Zero) serve as the baseline. We also implement a baseline by combining AnimateDiff (AD) with ControlNet via applying frame-wise control signals to the conditioned keyframes. For image animation, we compare SparseCtrl against two open-sourced image animation baselines: DynamiCrafter (DC) and VideoComposer’s initial frame mode.
Providing a dense depth sequence for video generation helps specify structural information to some extent. We thus evaluate our method on this task with much more challenging yet practical settings: only a few depths are given for the synthesis. The controlling fidelity under different sparsity of input is measured for the quantitative comparison. Specifically, we first select 20 videos from the validation set of WebVid-10M that are not seen during training. Thereafter, we estimate the corresponding depth sequences with the off-the-shelf MiDaS model, evenly mask out some of them with a ratio , and use the remaining depth maps as conditions to generate videos. We then estimate the depth maps from the conditioned keyframes in generated videos and, following the metrics in previous work, we perform scale shift realignment and compute the mean absolute error (MAE) against the depth maps extracted from the original videos. On the other hand, to prevent the model from learning a shot cut by solely controlling the keyframes and ignoring temporal consistency, we also report cross-frame CLIP similarity following previous works .
The quantitative results are shown in Tab. 1. To stay close to the original implementation, we only report results of for VideoComposer and Text2Video-Zero, where the controls for every frame are provided. As shown in the table, as the control sparsity, i.e., the masking rate , increases, our method maintains a comparable error rate with dense control baselines. In contrast, the error of AnimateDiff with frame-wise ControlNet increases, indicating that this baseline method tends to ignore the condition signals when the control becomes sparser.
3.2 Image animation
By providing the RGB image as the first frame condition, SparseCtrl can handle the task of image animation. To validate our method’s effectiveness, we further compare it with two baselines in this domain. We collect eight in-the-wild images and animate them using the three methods to generate 24 samples in total. Similar in Sec. 4.3.1, our metrics lie in two aspects: the first frame fidelity to the input image measured by LPIPS , and temporal consistency measured by CLIP similarity. Additionally, we invited 20 users to rank the results individually in terms of the fidelity to the given image and the overall quality preference. We obtained 160 ranking results for each aspect. We use average human ranking (AHR) as a preference metric and report the results in Tab. 2. The result shows that our method can achieve comparable performance with specifically designed animation pipelines while being favored in terms of fidelity to the first frame.
4 Ablative Study
We ablate on the sparse encoder architecture to verify our choice. Specifically, we experiment with four designs: (1) frame-wise condition encoder, where we repeat the 2D ControlNet along the temporal axis and encode the control signals to the keyframes, as depicted in Sec. 3.2; (2) condition encoder with propagation layers, where we add temporal layers upon (1) to propagate conditions across frames, as discussed in Sec. 3.2; (3) our full model, where we further eliminate the noised sample input to the condition encoder in (2). To better compare the effectiveness of these three choices, we consider the most challenging case, i.e., the RGB image conditions, because compared to other abstract modalities, here the synthetic results need to faithfully demonstrate the fine-grained details of the condition signal and propagate it to other unconditioned frames to ensure temporal consistency. With AnimateDiff , we additionally show the result on personalized image backbone, which further assists us in distinguishing the merits and shortcomings of different choices.
In Fig. 4, we show the qualitative image animation results. According to the figure, with all three variations, the first frame in the generated videos is fidelity to the input image control. The frame-wise encoder, under the personalized generation setting, fails to propagate the control to the unconditioned frames (1st row, right), leading to temporal inconsistency where the details of the character (e.g., hair, and clothes color) change over time. Upon the pre-trained T2V, the encoder with propagation layers, as stated in Sec. 3.2, suffers quality degradation (2nd row, left), and we hypothesize that this is because the noised sample input to the encoder provides misleading information for the condition tasks. Finally, with propagation layers and eliminating the noised sample input, our full model works well under the two settings (3rd row), maintaining both fidelity to condition and temporal consistency.
4.2 Unrelated Conditions
Besides the common usages, we experiment with an extreme case where the input conditions are unrelated or contradicted. Regarding this, we input two unrelated images to the RGB image encoder and require the model to interpolate between them, as shown in the first-row in Fig. 5. Surprisingly, the sparse encoder can still help generate smooth transitions between the input images, which further verifies the robustness of the SparseCtrl and shows potential in visual effects synthesis.
4.3 Response to Textual Prompt
Another interesting question is, with the additional information provided by the sparse condition encoder, to what extent does the final generated outcome respond to the input text description? To answer this, we experiment with different textual prompts with the same input and demonstrate the results in Fig. 5. In the image animation setting, we compare the prompt that faithfully describes the image content (2nd row) and the prompt that describes a slightly different content (3rd row). The results show that the input text prompts do influence the outcome by leading the contents towards the corresponding directions.
In the sketch-to-video setting, we construct three types of prompts: (1) insufficient prompt with no useful information (4th row), e.g., “an excellent video, best quality, masterpieces”; (2) incomplete prompt that partially describes the desired content (5th row), e.g., “sea, sunlight, …”, ignoring the central object “sailboat”; (3) completed prompt that describes every content (6th row). As shown in Fig. 5, with the sketch condition, the content can be properly generated only when the prompt is completed, showing that the text input still plays a significant role when the provided condition is highly abstract and insufficient to infer the content.
Discussion and Conclusion
We present SparseCtrl, a unified approach of adding temporally sparse controls to pre-trained text-to-video generators via an add-on encoder network. It can accommodate various modalities, including depth, sketches, and RGB images, greatly enhancing practical control for video generation. This flexibility proves invaluable in diverse applications like sketch-to-video, image animation, keyframe interpolation, etc. Extensive experiments have validated method’s effectiveness and generalizability across original and personalized text-to-video generators, making it a promising tool for real-world usage.
Limitations. Though with SparseCtrl, the visual quality, semantic composition ability, and domain of the generated results are limited by the pre-trained T2V backbone and the training data. In experiments, we find that the failure cases mostly come from out-of-domain input, such as anime image animation, since such data is scarce in the T2V and sparse encoder’s pre-training dataset WebVid-10M , whose contents are mainly real-world videos. Possible solutions for enhancing the generalizability could be improving the training dataset’s domain diversity and utilizing some domain-specific backbone, such as integrating SparseCtrl with AnimateDiff .
Acknowledgement. The project is supported by the Shanghai Artificial Intelligence Laboratory (P23KN00601, P23KS00020, 2022ZD0160201), CUHK Interdisciplinary AI Research Institute, and the Centre for Perceptual and Interactive Intelligence (CPIl) Ltd under the Innovation and Technology Commission (ITC)’s InnoHK.