DreamVideo: Composing Your Dream Videos with Customized Subject and Motion

Yujie Wei, Shiwei Zhang, Zhiwu Qing, Hangjie Yuan, Zhiheng Liu, Yu Liu, Yingya Zhang, Jingren Zhou, Hongming Shan

Introduction

The remarkable advances in diffusion models have empowered designers to generate photorealistic images and videos based on textual prompts, paving the way for customized content generation . While customized image generation has witnessed impressive progress , the exploration of customized video generation remains relatively limited. The main reason is that videos have diverse spatial content and intricate temporal dynamics simultaneously, presenting a highly challenging task to concurrently customize these two key factors.

Current existing methods have effectively propelled progress in this field, but they are still limited to optimizing a single aspect of videos, namely spatial subject or temporal motion. For example, Dreamix and Tune-A-Video optimize the spatial parameters and spatial-temporal attention to inject a subject identity and a target motion, respectively. However, focusing only on one aspect (i.e., subject or motion) may reduce the model’s generalization on the other aspect. On the other hand, AnimateDiff trains temporal modules appended to the personalized text-to-image models for image animation. It tends to pursue generalized video generation but suffers from a lack of motion diversity, such as focusing more on camera movements, making it unable to meet the requirements of customized video generation tasks well. Therefore, we believe that effectively modeling both spatial subject and temporal motion is necessary to enhance video customization.

The above observations drive us to propose the DreamVideo, which can synthesize videos featuring the user-specified subject endowed with the desired motion from a few images and videos respectively, as shown in Fig. 1. DreamVideo decouples video customization into subject learning and motion learning, which can reduce model optimization complexity and increase customization flexibility. In subject learning, we initially optimize a textual identity using Textual Inversion to represent the coarse concept, and then train a carefully designed identity adapter with the frozen textual identity to capture fine appearance details from the provided static images. In motion learning, we design a motion adapter and train it on the given videos to capture the inherent motion pattern. To avoid the shortcut of learning appearance features at this stage, we incorporate the image feature into the motion adapter to enable it to concentrate exclusively on motion learning. Benefiting from these two-stage learning, DreamVideo can flexibly compose customized videos with any subject and any motion once the two lightweight adapters have been trained.

To validate DreamVideo, we collect 20 customized subjects and 30 motion patterns as a substantial experimental set. The extensive experimental results unequivocally showcase its remarkable customization capabilities surpassing the state-of-the-art methods.

We propose DreamVideo, a novel approach for customized video generation with any subject and motion. To the best of our knowledge, this work makes the first attempt to customize both subject identity and motion.

We propose to decouple the learning of subjects and motions by the devised identity and motion adapters, which can greatly improve the flexibility of customization.

We conduct extensive qualitative and quantitative experiments, demonstrating the superiority of DreamVideo over the existing state-of-the-art methods.

Related Work

Text-to-video generation. Text-to-video generation aims to generate realistic videos based on text prompts and has received growing attention . Early works are mainly based on Generative Adversarial Networks (GANs) or autoregressive transformers . Recently, to generate high-quality and diverse videos, many works apply the diffusion model to video generation . Make-A-Video leverages the prior of the image diffusion model to generate videos without paired text-video data. Video Diffusion Models and ImagenVideo model the video distribution in pixel space by jointly training from image and video data. To reduce the huge computational cost, VLDM and MagicVideo apply the diffusion process in the latent space, following the paradigm of LDMs . Towards controllable video generation, ModelScopeT2V and VideoComposer incorporate spatiotemporal blocks with various conditions and show remarkable generation capabilities for high-fidelity videos. These powerful video generation models pave the way for customized video generation.

Customized generation. Compared with general generation tasks, customized generation may better accommodate user preferences. Most current works focus on subject customization with a few images . Textual Inversion represents a user-provided subject through a learnable text embedding without model fine-tuning. DreamBooth tries to bind a rare word with a subject by fully fine-tuning an image diffusion model. Moreover, some works study the more challenging multi-subject customization task . Despite the significant progress in customized image generation, customized video generation is still under exploration. Dreamix attempts subject-driven video generation by following the paradigm of DreamBooth. However, fine-tuning the video diffusion model tends to overfitting and generate videos with small or missing motions. A concurrent work aims to customize the motion from training videos. Nevertheless, it fails to customize the subject, which may be limiting in practical applications. In contrast, this work proposes DreamVideo to effectively generate customized videos with both specific subject and motion.

Parameter-efficient fine-tuning. Drawing inspiration from the success of parameter-efficient fine-tuning (PEFT) in NLP and vision tasks , some works adopt PEFT for video generation and editing tasks due to its efficiency . In this work, we explore the potential of lightweight adapters, revealing their superior suitability for customized video generation.

Methodology

In this section, we first introduce the preliminaries of Video Diffusion Models. We then present DreamVideo to showcase how it can compose videos with the customized subject and motion. Finally, we analyze the efficient parameters for subject and motion learning while describing training and inference processes for our DreamVideo.

Video diffusion models (VDMs) are designed for video generation tasks by extending the image diffusion models to adapt to the video data. VDMs learn a video data distribution by the gradual denoising of a variable sampled from a Gaussian distribution. This process simulates the reverse process of a fixed-length Markov Chain. Specifically, the diffusion model ϵθ\epsilon_{\theta} aims to predict the added noise ϵ\epsilon at each timestep tt based on text condition cc, where t∈U(0,1)t\in\mathcal{U}(0,1). The training objective can be simplified as a reconstruction loss:

2 DreamVideo

Given a few images of one subject and multiple videos (or a single video) of one motion pattern, our goal is to generate customized videos featuring both the specific subject and motion. To this end, we propose DreamVideo, which decouples the challenging customized video generation task into subject learning and motion learning via two devised adapters, as illustrated in Fig. 2. Users can simply combine these two adapters to generate desired videos.

Subject learning. To accurately preserve subject identity and mitigate overfitting, we introduce a two-step training strategy inspired by for subject learning with 3∼\sim5 images, as illustrated in the upper left portion of Fig. 2.

The first step is to learn a textual identity using Textual Inversion . We freeze the video diffusion model and only optimize the text embedding of pseudo-word “S∗S^{*}” using Eq. (1). The textual identity represents the coarse concept and serves as a good initialization.

Motion learning. Another important property of customized video generation is to make the learned subject move according to the desired motion pattern from existing videos. To efficiently model a motion, we devise a motion adapter with a structure similar to the identity adapter, as depicted in Fig. 3(b). Our motion adapter can be customized using a motion pattern derived from a class of videos (e.g., videos representing various dog motions), multiple videos exhibiting the same motion, or even a single video.

where h^t′\hat{h}_{t}^{\prime} is the output of motion adapter. At inference time, we randomly take a training image provided by the user as the appearance condition input to the motion adapter.

3 Model Analysis, Training and Inference

Where to put these two adapters. We address this question by analyzing the change of all parameters within the fine-tuned model to determine the appropriate position of the adapters. These parameters are divided into four categories: (1) cross-attention (only exists in spatial parameters), (2) self-attention, (3) feed-forward, and (4) other remaining parameters. Following , we use Δl=∥θl′−θl∥2/∥θl∥2\Delta_{l}=\left\|\theta_{l}^{\prime}-\theta_{l}\right\|_{2}/\left\|\theta_{l}\right\|_{2} to calculate the weight change rate of each layer, where θl′\theta_{l}^{\prime} and θl\theta_{l} are the updated and pre-trained model parameters of layer ll. Specifically, to compute Δ\Delta of spatial parameters, we only fine-tune the spatial parameters of the UNet while freezing temporal parameters, for which the Δ\Delta of temporal parameters is computed in a similar way.

We observe that the conclusions are different for spatial and temporal parameters. Fig. 4(a) shows the mean Δ\Delta of spatial parameters for the four categories when fine-tuning the model on “Chow Chow” images (dog in Fig. 1). The result suggests that the cross-attention layers play a crucial role in learning appearance compared to other parameters. However, when learning motion dynamics in the “bear walking” video (see Fig. 7), all parameters contribute close to importance, as shown in Fig. 4(b). Remarkably, our findings remain consistent across various images and videos. This phenomenon reveals the divergence of efficient parameters for learning subjects and motions. Therefore, we insert the identity adapter to cross-attention layers while employing the motion adapter to all layers in temporal transformer.

Decoupled training strategy. Customizing the subject and motion simultaneously on images and videos requires training a separate model for each combination, which is time-consuming and impractical for applications. Instead, we tend to decouple the training of subject and motion by optimizing the identity and motion adapters independently according to Eq. (1) with the frozen pre-trained model.

Inference. During inference, we combine the two customized adapters and randomly select an image provided during training as the appearance guidance to generate customized videos. We find that choosing different images has a marginal impact on generated results. Besides combinations, users can also customize the subject or motion individually using only the identity adapter or motion adapter.

Experiment

Datasets. For subject customization, we select subjects from image customization papers for a total of 20 customized subjects, including 9 pets and 11 objects. For motion customization, we collect a dataset of 30 motion patterns from the Internet, the UCF101 dataset , the UCF Sports Action dataset , and the DAVIS dataset . We also provide 42 text prompts used for extensive experimental validation, where the prompts are designed to generate new motions of subjects, new contexts of subjects and motions, and etc.

Implementation details. For subject learning, we take ∼\sim3000 iterations for optimizing the textual identity following with learning rate 1.0×10−41.0\times 10^{-4}, and ∼\sim800 iterations for learning identity adapter with learning rate 1.0×10−51.0\times 10^{-5}. For motion learning, we train motion adapter for ∼\sim1000 iterations with learning rate 1.0×10−51.0\times 10^{-5}. During inference, we employ DDIM with 50-step sampling and classifier-free guidance to generate 32-frame videos with 8 fps. Additional details of our method and baselines are reported in Appendix A.

Baselines. Since there is no existing work for customizing both subjects and motions, we consider comparing our method with three categories of combination methods: AnimateDiff , ModelScopeT2V , and LoRA fine-tuning . AnimateDiff trains a motion module appended to a pre-trained image diffusion model from Dreambooth . However, we find that training from scratch leads to unstable results. For a fair comparison, we further fine-tune the pre-trained weights of the motion module provided by AnimateDiff and carefully adjust the hyperparameters. For ModelScopeT2V and LoRA fine-tuning, we train spatial and temporal parameters/LoRAs of the pre-trained video diffusion model for subject and motion respectively, and then merge them during inference. In addition, we also evaluate our generation quality for customizing subjects and motions independently. We evaluate our method against Textual Inversion and Dreamix for subject customization while comparing with Tune-A-Video and ModelScopeT2V for motion customization.

Evaluation metrics. We evaluate our approach with the following four metrics, three for subject customization and one for video generation. (1) CLIP-T calculates the average cosine similarity between CLIP image embeddings of all generated frames and their text embedding. (2) CLIP-I measures the visual similarity between generated and target subjects. We compute the average cosine similarity between the CLIP image embeddings of all generated frames and the target images. (3) DINO-I , another metric for measuring the visual similarity using ViTS/16 DINO . Compared to CLIP, the self-supervised training model encourages distinguishing features of individual subjects. (4) Temporal Consistency , we compute CLIP image embeddings on all generated frames and report the average cosine similarity between all pairs of consecutive frames.

2 Results

In this section, we showcase results for both joint customization as well as individual customization of subjects and motions, further demonstrating the flexibility and effectiveness of our method.

Arbitrary combinations of subjects and motions. We compare our DreamVideo with several baselines to evaluate the customization performance, as depicted in Fig. 5. We observe that AnimateDiff preserves the subject appearances but fails to model the motion patterns accurately, resulting in generated videos lacking motion diversity. Furthermore, ModelScopeT2V and LoRA suffer from fusion conflicts during combination, where either subject identities are corrupted or motions are damaged. In contrast, our DreamVideo achieves effective and harmonious combinations that the generated videos can retain subject identities and motions under various contexts; see Appendix B.1 for more qualitative results about combinations of subjects and motions.

Tab. 1 shows quantitative comparison results of all methods. DreamVideo outperforms other methods across CLIP-T, CLIP-I, and DINO-I, which is consistent with the visual results. Although AnimateDiff achieves highest Temporal Consistency, it tends to generate videos with small motions. In addition, our method remains comparable to Dreamix in Temporal Consistency but requires fewer parameters.

Subject customization. To verify the individual subject customization capability of our DreamVideo, we conduct qualitative comparisons with Textual Inversion and Dreamix , as shown in Fig. 6. For a fair comparison, we employ the same baseline model, ModelScopeT2V, for all compared methods. We observe that Textual Inversion makes it hard to reconstruct the accurate subject appearances. While Dreamix captures the appearance details of subjects, the motions of generated videos are relatively small due to overfitting. Moreover, certain target objects in the text prompts, such as “pizza” in Fig. 6, are not generated by Dreamix. In contrast, our DreamVideo effectively mitigates overfitting and generates videos that conform to text descriptions while preserving precise subject appearances.

The quantitative comparison for subject customization is shown in Tab. 2. Regarding the CLIP-I and Temporal Consistency, our method demonstrates a comparable performance to Dreamix while surpassing Textual Inversion. Remarkably, our DreamVideo outperforms alternative methods in CLIP-T and DINO-I with relatively few parameters. These results demonstrate that our method can efficiently model the subjects with various contexts. Comparison with Custom Diffusion and more qualitative results are reported in Appendix B.2.

Motion customization. Besides subject customization, we also evaluate the motion customization ability of our DreamVideo by comparing it with several competitors, as shown in Fig. 7. For a fair comparison, we only fine-tune the temporal parameters of ModelScopeT2V to learn a motion. The results show that ModelScopeT2V inevitably fuses the appearance information of training videos, while Tune-A-Video suffers from discontinuity between video frames. In contrast, our method can capture the desired motion patterns while ignoring the appearance information of the training videos, generating temporally consistent and diverse videos; see Appendix B.3 for more qualitative results about motion customization.

As shown in Tab. 3, our DreamVideo achieves the highest CLIP-T and Temporal Consistency compared to baselines, verifying the superiority of our method.

User study. To further evaluate our approach, we conduct user studies for subject customization, motion customization, and their combinations respectively. For combinations of specific subjects and motions, we ask 5 annotators to rate 50 groups of videos consisting of 5 motion patterns and 10 subjects. For each group, we provide 3∼\sim5 subject images and 1∼\sim3 motion videos; and compare our DreamVideo with three methods by generating videos with 6 text prompts. We evaluate all methods with a majority vote from four aspects: Text Alignment, Subject Fidelity, Motion Fidelity, and Temporal Consistency. Text Alignment evaluates whether the generated video conforms to the text description. Subject Fidelity and Motion Fidelity measure whether the generated subject or motion is close to the reference images or videos. Temporal Consistency measures the consistency between video frames. As shown in Tab. 4, our approach is most preferred by users regarding the above four aspects. More details and user studies of subject customization as well as motion customization can be found in the Appendix B.4.

3 Ablation Studies

We conduct an ablation study on the effects of each component in the following. More ablation studies on the effects of parameter numbers and different adapters are reported in Appendix C.

Effects of each component. As shown in Fig. 8, we can observe that without learning the textual identity, the generated subject may lose some appearance details. When only learning subject identity without our devised motion adapter, the generated video fails to exhibit the desired motion pattern due to limitations in the inherent capabilities of the pre-trained model. In addition, without proposed appearance guidance, the subject identity and background in the generated video may be slightly corrupted due to the coupling of spatial and temporal information. These results demonstrate each component makes contributions to the final performance. More qualitative results can be found in Appendix C.1.

The quantitative results in Tab. 5 show that all metrics decrease slightly without textual identity or appearance guidance, illustrating their effectiveness. Furthermore, we observe that only customizing subjects leads to the improvement of CLIP-I and DINO-I, while adding the motion adapter can increase CLIP-T and Temporal Consistency. This suggests that the motion adapter helps to generate temporal coherent videos that conform to text descriptions.

Conclusion

In this paper, we present DreamVideo, a novel approach for customized video generation with any subject and motion. DreamVideo decouples video customization into subject learning and motion learning to enhance customization flexibility. We combine textual inversion and identity adapter tuning to model a subject and train a motion adapter with appearance guidance to learn a motion. With our collected dataset that contains 20 subjects and 30 motion patterns, we conduct extensive qualitative and quantitative experiments, demonstrating the efficiency and flexibility of our method in both joint customization and individual customization of subjects and motions.

Limitations. Although our method can efficiently combine a single subject and a single motion, it fails to generate customized videos that contain multiple subjects with multiple motions. One possible solution is to design a fusion module to integrate multiple subjects and motions, or to implement a general customized video model. We provide more analysis and discussion in Appendix D.

References

Appendix

Appendix A Experimental Details

In this section, we supplement the experimental details of each baseline method and our method. To improve the quality and remove the watermarks of generated videos, we further fine-tune ModelScopeT2V for 30k iterations on a randomly selected subset from our internal data, which contains about 30,000 text-video pairs. For a fair comparison, we use the fine-tuned ModelScopeT2V model as the base video diffusion model for all methods except for AnimateDiff and Tune-A-Video , both of which use the image diffusion model (Stable Diffusion ) in their official papers. Here, we use Stable Diffusion v1-5https://huggingface.co/runwayml/stable-diffusion-v1-5 as their base image diffusion model. During training, unless otherwise specified, we default to using AdamW optimizer with the default betas set to 0.9 and 0.999. The epsilon is set to the default 1.0×10−81.0\times 10^{-8}, and the weight decay is set to 0. During inference, we use 50 steps of DDIM sampler and classifier-free guidance with a scale of 9.0 for all baselines. We generate 32-frame videos with 256 ×\times 256 spatial resolution and 8 fps. All experiments are conducted using one NVIDIA A100 GPU. In the following, we introduce the implementation details of baselines from subject customization, motion customization, and arbitrary combinations of subjects and motions (referred to as video customization).

For all methods, we set batch size as 4 to learn a subject.

DreamVideo (ours). In subject learning, we take ∼\sim3000 iterations for optimizing the textual identity following with learning rate 1.0×10−41.0\times 10^{-4}, and ∼\sim800 iterations for learning identity adapter with learning rate 1.0×10−51.0\times 10^{-5}. We set the hidden dimension of the identity adapter to be half the input dimension. Our method takes ∼\sim12 minutes to train the identity adapter on one A100 GPU.

Textual Inversion . According to their official codehttps://github.com/rinongal/textual_inversion, we reproduce Textual Inversion to the video diffusion model. We optimize the text embedding of pseudo-word “S∗S^{*}” with prompt “a S∗S^{*}” for 3000 iterations, and set the learning rate to 1.0×10−41.0\times 10^{-4}. We also initialize the learnable token with the corresponding class token. These settings are the same as the first step of our subject-learning strategy.

Dreamix . Since Dreamix is not open source, we reproduce its method based on the codehttps://modelscope.cn/models/damo/text-to-video-synthesis of ModelScopeT2V. According to the descriptions in the official paper, we only train the spatial parameters of the UNet while freezing the temporal parameters. Moreover, we refer to the third-party implementationhttps://github.com/XavierXiao/Dreambooth-Stable-Diffusion of DreamBooth to bind a unique identifier with the specific subject. The text prompt used for target images is “a [V] [category]”, where we initialize [V] with “sks”, and [category] is a coarse class descriptor of the subject. The learning rate is set to 1.0×10−51.0\times 10^{-5}, and the training iterations are 100 ∼\sim 200.

Custom Diffusion . We refer to the official codehttps://github.com/adobe-research/custom-diffusion of Custom Diffusion and reproduce it on the video diffusion model. We train Custom Diffusion with the learning rate of 4.0×10−54.0\times 10^{-5} and 250 iterations, as suggested in their paper. We also detach the start token embedding ahead of the class word with the text prompt “a S∗S^{*} [category]”. We simultaneously optimize the parameters of the key as well as value matrices in cross-attention layers and text embedding of S∗S^{*}. We initialize the token S∗S^{*} with the token-id 42170 according to the paper.

A.2 Motion Customization

To model a motion, we set batch size to 2 for training from multiple videos while batch size to 1 for training from a single video.

DreamVideo (ours). In motion learning, we train motion adapter for ∼\sim1000 iterations with learning rate 1.0×10−51.0\times 10^{-5}. Similar to the identity adapter, the hidden dimension of the motion adapter is set to be half the input dimension. On one A100 GPU, our method takes ∼\sim15 and ∼\sim30 minutes to learn a motion pattern from a single video and multiple videos, respectively.

ModelScopeT2V . We only fine-tune the temporal parameters of the UNet while freezing the spatial parameters. We set the learning rate to 1.0×10−51.0\times 10^{-5}, and also train 1000 iterations to learn a motion.

Tune-A-Video . We use the official implementationhttps://github.com/showlab/Tune-A-Video of Tune-A-Video for experiments. The learning rate is 3.0×10−53.0\times 10^{-5}, and training iterations are 500. Here, we adapt Tune-A-Video to train on both multiple videos and a single video.

A.3 Video Customization

DreamVideo (ours). We combine the trained identity adapter and motion adapter for video customization during inference. No additional training is required. We also randomly select an image provided during training as the appearance guidance. We find that choosing different images has a marginal impact on generated videos.

AnimateDiff . We use the official implementationhttps://github.com/guoyww/AnimateDiff of AnimateDiff for experiments. AnimateDiff trains the motion module from scratch, but we find that this training strategy may cause the generated videos to be unstable and temporal inconsistent. For a fair comparison, we further fine-tune the pre-trained weights of the motion module provided by AnimateDiff and carefully adjust the hyperparameters. The learning rate is set to 1.0×10−51.0\times 10^{-5}, and training iterations are 50. For the personalized image diffusion model, we use the third-party implementation4 code to train a DreamBooth model. During inference, we combine the DreamBooth model and motion module to generate videos.

ModelScopeT2V . We train spatial/temporal parameters of the UNet while freezing other parameters to learn a subject/motion. Settings of training subject and motion are the same as Dreamix in Sec. A.1 and ModelScopeT2V in Sec. A.2, respectively. During inference, we combine spatial and temporal parameters into a UNet to generate videos.

LoRA . In addition to fully fine-tuning, we also attempt the combinations of LoRAs. According to the conclusions in Sec. 3.3 of our main paper and the method of Custom Diffusion , we only add LoRA to the key and value matrices in cross-attention layers to learn a subject. For motion learning, we add LoRA to the key and value matrices in all attention layers. The LoRA rank is set to 32. Other settings are consistent with our DreamVideo. During inference, we merge spatial and temporal LoRAs into corresponding layers.

Appendix B More Results

In this section, we conduct further experiments and showcase more results to illustrate the superiority of our DreamVideo.

We provide more results compared with the baselines, as shown in Fig. A1. The videos generated by AnimateDiff suffer from little motion, while other methods still struggle with the fusion conflict problem of subject identity and motion. In contrast, our method can generate videos that preserve both subject identity and motion pattern.

B.2 Subject Customization

In addition to the baselines in the main paper, we also compare our DreamVideo with another state-of-the-art method, Custom Diffusion . Both the qualitative comparison in Fig. A2 and the quantitative comparison in Tab. A1 illustrate that our method outperforms Custom Diffusion and can generate videos that accurately retain subject identity and conform to diverse contextual descriptions with fewer parameters.

As shown in Fig. A3, we provide the customization results for more subjects, further demonstrating the favorable generalization of our method.

B.3 Motion Customization

To further evaluate the motion customization capabilities of our method, we show more qualitative comparison results with baselines on multiple training videos and a single training video, as shown in Fig. A4. Our method exhibits superior performance than baselines and ignores the appearance information from training videos when modeling motion patterns.

We showcase more results of motion customization in Fig. A5, providing further evidence of the robustness of our method.

B.4 User Study

For subject customization, we generate 120 videos from 15 subjects, where each subject includes 8 text prompts. We present three sets of questions to participants with 3∼\sim5 reference images of each subject to evaluate Text Alignment, Subject Fidelity, and Temporal Consistency. Given the generated videos of two anonymous methods, we ask each participant the following questions: (1) Text Alignment: “Which video better matches the text description?”; (2) Subject Fidelity: “Which video’s subject is more similar to the target subject?”; (3) Temporal Consistency: “Which video is smoother and has less flicker?”. For motion customization, we generate 120 videos from 20 motion patterns with 6 text prompts. We evaluate each pair of compared methods through Text Alignment, Motion Fidelity, and Temporal Consistency. The questions of Text Alignment and Temporal Consistency are similar to those in subject customization above, and the question of Motion Fidelity is like: “Which video’s motion is more similar to the motion of target videos?” The human evaluation results are shown in Tab. A2 and Tab. A3. Our DreamVideo consistently outperforms other methods on all metrics.

Appendix C More Ablation Studies

We provide more qualitative results in Fig. A6 to further verify the effects of each component in our method. The conclusions are consistent with the descriptions in the main paper. Remarkably, we observe that without appearance guidance, the generated videos may learn some noise, artifacts, background, and other subject-unrelated information from training videos.

C.2 Effects of Parameters in Adapter and LoRA

To measure the impact of the number of parameters on performance, we reduce the hidden dimension of the adapter to make it have a comparable number of parameters as LoRA. For a fair comparison, we set the hidden dimension of the adapter to 32 without using textual identity and appearance guidance. We adopt the DreamBooth paradigm for subject learning, which is the same as LoRA. Other settings are the same as our DreamVideo.

As shown in Fig. A7, we observe that LoRA fails to generate videos that preserve both subject identity and motion. The reason may be that LoRA modifies the original parameters of the model during inference, causing conflicts and sacrificing performance when merging spatial and temporal LoRAs. In contrast, the adapter can alleviate fusion conflicts and achieve a more harmonious combination.

The quantitative comparison results in Tab. A4 also illustrate the superiority of the adapter compared to LoRA in video customization tasks.

C.3 Effects of Different Adapters

To evaluate which adapter is more suitable for customization tasks, we design 4 combinations of adapters and parameter layers for motion customization, as shown in Tab. A5. We consider the serial adapter and parallel adapter along with self-attention layers and feed-forward layers. The results demonstrate that using parallel adapters on all layers achieves the best performance. Therefore, we uniformly employ parallel adapters in our approach.

Appendix D Social Impact and Discussions

Social impact. While training large-scale video diffusion models is extremely expensive and unaffordable for most individuals, video customization by fine-tuning only a few images or videos provides users with the possibility to use video diffusion models flexibly. Our approach allows users to generate customized videos by arbitrarily combining subject and motion while also supporting individual subject customization or motion customization, all with a small computational cost. However, our method still suffers from the risks that many generative models face, such as fake data generation. Reliable video forgery detection techniques may be a solution to these problems.

Discussions. We provide some failure examples in Fig. A8. For subject customization, our approach is limited by the inherent capabilities of the base model. For example, in Fig. A8(a), the basic model fails to generate a video like “a wolf riding a bicycle”, causing our method to inherit this limitation. The possible reason is that the correlation between “wolf” and “bicycle” in the training set during pre-training is too weak. For motion customization, especially fine single video motion, our method may only learn the similar (rough) motion pattern and fails to achieve frame-by-frame correspondence, as shown in Fig. A8(b). Some video editing methods may be able to provide some solutions . For video customization, some difficult combinations that contain multiple objects, such as “cat” and “horse”, still remain challenges. As shown in Fig. A8(c), our approach confuses “cat” and “horse” so that both exhibit “cat” characteristics. This phenomenon also exists in multi-subject image customization . One possible solution is to further decouple the attention map of each subject.