Identity-Preserving Text-to-Video Generation by Frequency Decomposition
Shenghai Yuan, Jinfa Huang, Xianyi He, Yunyuan Ge, Yujun Shi, Liuhan Chen, Jiebo Luo, Li Yuan
Introduction
Large-scale pre-trained video diffusion models have facilitated a variety of downstream applications , particularly in identity-preserving text-to-video (IPT2V) . However, existing methods face significant challenges, particularly the high overhead associated with the need for case-by-case finetuning, which diminishes their applicability. Within the open-source community, only the ID-Animator can implement tuning-free IPT2V, but it can only generate videos similar to talking head and has poor id preservation.
Additionally, the above efforts are predominantly based on U-Net and cannot be adapted to the emerging DiT-based video model . This challenge may stem from the inherent limitations of DiT compared to U-Net, including greater difficulty in training convergence and weakness in perceiving facial details. From some prior findings in frequency analysis of vision/diffusion transformers , we can know that the reason is: Finding 1: Shallow (e.g., low-level, low-frequency) features are essential for pixel-level prediction tasks in diffusion models, as they ease model training. U-Net facilitates model convergence by aggregating shallow features to the decoder via long skip connections, a mechanism that DiT does not incorporate; Finding 2: Transformers have limited perception of high-frequency information, which is important for preserving facial features. The encoder-decoder architecture of U-Net naturally possesses multi-scale features (e.g., richness in high-frequency), while DiT lacks a comparable structure. To develop a DiT-based control model, it is necessary to address these issues first.
For ID-preserving video generation, the challenges stem from the requirement for each frame to incorporate both high-frequency (e.g., age- and make-up-independent identity markers) and low-frequency information (e.g., facial shape) derived from the reference image, which can just be used to make up for the DiT defects mentioned above. Therefore, we propose ConsisID, to keep the identity consistency in video generation by frequency decomposition, based on the previously Findings of DiT in frequency analysis. Thanks to the large-scale pre-trained DiT, we can use its powerful capabilities to achieve tuning-free effects. ConsisID decouples identity features into high- and low-frequency signals, which are injected into specific locations within the DiT, facilitating efficient IPT2V generation. Specifically, in line with Finding 1, we first convert the reference image and the facial key points to the low-frequency signal, then concatenate them with input noise latent to ease the training. Following Finding 2, we utilize a dual-tower feature extractor to capture high-frequency facial information, which is integrated with vision tokens within the transformer block, thereby enhancing the DiT’s high-frequency perception capabilities. Finally, to transform the pre-trained model into an IPT2V model and improve its generalization, we further introduce a hierarchical training strategy.
Our contributions can be summarized as follows:
We introduce ConsisID, a tuning-free identity-preserving DiT-based IPT2V model, which preserves the identity of the main subject of the video using control signals from frequency decomposition.
We propose a hierarchical training strategy, including coarse-to-fine training, dynamic mask loss, and dynamic cross-face loss, which work together to facilitate training and enhance generalization effectively.
Extensive experiments demonstrate our ConsisID can generate high-quality, editable, consistent identity-preserving videos, benefiting from our frequency-aware identity-preserving T2V DiT-based control scheme.
Related Work
Tuning-based Identity-preserving T2V Models. Diffusion models are widely recognized for their strong generative capabilities , significantly advancing the development of identity-preserving generative models . Initially, the researchers used tuning-based methods to generate content that matched the input ID. This process requires finetuning pretrained model for each new person during inference. For example, DreamBooth introduced a novel loss function to fine-tune the entire network, embedding identity information while preserving the original generative capabilities. LoRA , similar to DreamBooth , requires training only a small subset of network parameters. In contrast, Textual Inversion freezes the pretrained network and embeds identity information into a trainable word embedding. Subsequent tuning-based methods, including both image and video models based on U-Net or DiT architectures , generally follow three main approaches. While these models demonstrate substantial effectiveness, the requirement to fine-tune for each new identity restricts their practical applicability.
Tuning-free Identity-preserving T2V Models. To address the issue of high resource consumption, several tuning-free diffusion models have recently emerged in the field of image generation . These models do not require finetuning parameters for newly introduced IDs during inference. For instance, IP-Adapter utilizes the CLIP features of the identity image through cross-attention to guide the pretrained model in generating identity-preserving images. InstantID extends this approach by replacing CLIP features with Arcface features and integrating a pose network to adjust facial proportions. Unlike these initial methods, which introduce control signals via visual tokens, PhotoMaker and Imagine Yourself leverage text tokens. Specifically, PhotoMaker concatenates identity features obtained from the CLIP encoder to the text embedding, while Imagine Yourself uses element-wise addition for feature fusion. In the domain of video generation, only MovieGen and ID-Animator currently support ID-preserving text-to-video (IPT2V) generation. MovieGen is closed-source, whereas ID-Animator is open-source but uses a methodology similar to image models, leading to lower-quality identity preservation in the generated videos. We select the emerging DiT architecture and optimize it for IPT2V, drawing on conclusions from prior frequency analyses . This enables high-quality, editable, and consistent ID-preserving video generation.
Methodology
Diffusion Model. Text-to-video generation models usually utilize the diffusion paradigm, which gradually transforms noise into a video . Originally, denoising was conducted directly within the pixel space ; however, due to significant computational overheads, recent methods predominantly employ latent space . The optimization process is defined:
where is text condition, is sampled from a standard normal distribution (e.g., ), and is the text encoder. By replacing with , the latent diffusion is derived, which is used by ConsisID.
Diffusion Transformer. The DiT-based video generation model shows significant potential in simulating the physical world . Despite being a novel architecture, research on controllable generation has been limited, and current methods largely resemble U-Net based approaches . However, no study has yet examined why this approach works with DiT. Drawing from prior analyses of Diffusion and Transformer from a frequency domain perspective , we conclude that: (1) Low-level (e.g., shallow-layer) features are essential for pixel-level prediction tasks in diffusion models, which helps facilitate model training; (2) Transformers have limited perception for high-frequency information, which is important for controllable generation. Based on these, we decouple ID features into high- and low-frequency parts and inject them into specific locations, achieving effective identity-preserving text-to-video generation.
2 ConsisID: Keep Your Identity Consistent
The overview is illustrated in Figure 2. Given a reference image, the global facial extractor and local facial extractor inject both high- and low-frequency facial information into model, which then generates identity-preserving videos with the assistance of the consistency training strategy.
In light of Finding 1, enhancing low-level (e.g., shallow, low-frequency) features accelerates model convergence. To easily adapt a pre-trained model for the IPT2V task, the most direct approach is concatenating the reference face with the noise input latent . However, the reference face contains both high-frequency details (e.g., eye and lip textures) and low-frequency information (e.g., facial proportions and contours). From Finding 2, prematurely injecting high-frequency information into the Transformer is inefficient and may hinder the model’s processing of low-frequency information, as the Transformer focuses primarily on low-frequency features. In addition, feeding the reference face directly into the model could introduce irrelevant noise such as lighting and shadows. To mitigate this, we extract facial key points, convert them to an RGB image, and then concatenate it with the reference image, as shown in Figure 2. This strategy focuses the model’s attention on the low-frequency signals in the face, while minimizing the impact of extraneous features. We found that when this component is discarded, the model has a gradient explosion. The objective function is changed to:
where is the global facial extractor, represents the reference image.
2.2 High-frequency View: Local Facial Extractor
In light of Finding 2, we recognize that Transformers have limited sensitivity to high-frequency information. It can be concluded that relying solely on global facial features is insufficient for IPT2V generation, as global facial features primarily consist of low-frequency information and lack the intrinsic features necessary for editing. This task requires not only maintaining identity consistency, but also incorporating editing capabilities, such as generating videos of faces with the same identity but varying age and makeup. Achieving this requires the extraction of facial features that are unaffected by non-ID attributes (e.g., expression, posture, and shape), since age and makeup do not alter a person’s core identity. We define these features as intrinsic identity features (e.g., high-frequency).
Previous research use local features from the CLIP image encoder as intrinsic features to improve editing capabilities. However, since CLIP is not specifically trained on face datasets, the extracted features contain harmful non-face information . Therefore, we choose to use a face recognition backbone to extract intrinsic identity features. Instead of using the output of the backbone as the intrinsic identity feature, we use the penultimate layer, which retains more spatial information related to identity. However, these features still lack sufficient semantic information , which is crucial for personalized video generation.
To address these issues, we first use a facial recognition backbone to extract features that are strong in the intrinsic identity representation, and a CLIP image encoder to capture features that are strong in semantics. We then use the Q-former to fuse these two features, producing intrinsic identity features enriched with high-frequency semantic information. To reduce the impact of irrelevant features from CLIP, dropout is applied before entering into Q-Former. Additionally, we concatenate the shallow, multi-scale features from the facial recognition backbone, after interpolation, with the CLIP features. This method ensures that the model effectively captures essential intrinsic identity features while filtering out external noise unrelated to identity. After extracting the intrinsic id features, we apply cross-attention to interact with the visual tokens produced by each attention block of the pre-trained model, effectively enhancing the high-frequency information in DiT:
where represents the layer number of the attention block, , , and , where is the visual token, represents the intrinsic identity features, and , , and are trainable parameters. The objective function is changed to:
where is the local facial extractor.
2.3 Consistency Training Strategy
During training, we randomly select a frame from the training frames and apply the Crop & Align to extract the facial region as reference images, which is subsequently used as an identity-control signal, alongside the text as control.
Coarse-to-Fine Training. Compared to Identity-preserving image generation, video generation requires maintaining consistency in both spatial and temporal dimensions, ensuring that high and low-frequency facial information matches the reference image. To mitigate the complexity of training, we propose a hierarchical strategy where the model learns information globally before refining it locally. In the coarse-grained phase (e.g., corresponding Finding 1), we employ the global facial extractor, enabling the model to prioritize low-frequency features, such as facial contours and proportions, thereby ensuring rapid acquisition of identity information from the reference image and consistency across the video sequence. In the fine-grained phase (e.g. corresponding to Finding 2), the local facial extractor shifts the model’s focus to high-frequency details, such as the texture details of eyes and lips (e.g., intrinsic identification), improving the fidelity of facial expressions and the overall similarity of the generated face.
Dynamic Mask Loss. The objective of our task is to ensure that the identity of the person in the generated video remains consistent with the input reference image. However, Equation 1 considers the entire scene, encompassing both high- and low-frequency identity information as well as redundant background content, which introduces noise that interferes with model training. To address this, we propose to focus the model’s attention on face regions. Specifically, we first extract the facial mask from the video, apply trilinear interpolation to map it to the latent space, and finally use this mask to constrain the computation of :
where represents a mask with the same shape as . However, if Equation 5 is used as the supervisory signal for all training data, the model may fail to generate a natural background during inference. To mitigate this issue, we apply Equation 5 with a probability of , resulting in:
Dynamic Cross-face Loss. After training with Equation 6, we observed that the model struggled to generate satisfactory results for persons not present in the data domain during inference. This issue arises because the model, trained exclusively on faces from the training frames, tends to overfit by adopting a "copy-paste" shortcut—essentially replicating the reference image without alteration. To improve the model’s generalization capability, we introduce slight Gaussian noise to the reference images and use cross-face (e.g., reference images are sourced from video frames outside the training frames) as inputs with probability :
where is the reference image extracted from the training frames, and is extracted from outside the training frames.
Experiments
Implementation details. ConsisID selects DiT-based generation architectures CogVideoX-5B as our baseline for validation. We use an in-house human-centric dataset for training, which differs from previous datasets that focus only on the face. In the training phase, we set the resolution to 480720 and extracted 49 consecutive frames at a stride of 3 from each video as training data. We set the batch size to 80, the learning rate to , and the total number of training steps to 1.8k. The classify free guidance random null text ratio is set to , with AdamW serving as an optimizer and cosine_with_restarts as a learning rate scheduler. The training strategy is the same as Section 3.2.3. We set and in the dynamic cross-face loss () and dynamic mask loss () to , respectively. In the inference phase, we employ DPM with a sampling step of , and a text-guidance ratio of .
Benchmark. Since there is an absence of an evaluation dataset, we select 30 persons who were not included in the training data and sourced five high-quality images for each ID from the internet. We then design 90 distinct prompts, encompassing a variety of expressions, actions, and backgrounds for evaluation. Building on previous works , we evaluate four dimensions: (1). Identity Preservation: We use FaceSim-Arc and introduce FaceSim-Cur, which assesses identity preservation by measuring feature differences between face regions in the generated videos and those in real face images within the ArcFace and CurricularFace feature spaces. (2) Visual Quality: We utilize FID by calculating feature differences in the face regions between the generated frames and real face images within the InceptionV3 feature space. (3) Text Relevance: We utilize CLIPScore to measure the similarity between the generated videos and the input prompts. (4). Motion Amplitude: Due to the lack of reliable metrics, we evaluate through the user study.
2 Qualitative Analysis
In this section, we compare our method, ConsisID, with ID-Animator (e.g., the only available open-source model) for tuning-free IPT2V tasks. We randomly select images and text prompts of four individuals for qualitative analysis, all of which are absent from the training data. As shown in Figure 5, ID-Animator cannot generate human body parts beyond the face and is unable to generate complex actions or backgrounds in response to text prompts (e.g., action, attribute, background), which significantly limits its practical application. In addition, the preservation of the identity is inadequate; for example, in case 1, the reference image appears to be processed with skin smoothing. In case 2, wrinkles have been introduced which detract from the aesthetic quality. In cases 3 and 4, the face is distorted due to the lack of low frequency information, which compromises identity consistency. In contrast, the proposed ConsisID consistently produces high-quality, realistic videos that accurately match the reference identity and adhere to prompt.
3 Quantitative Analysis
We present a comprehensive quantitative evaluation of different methods, with results displayed in Table 1. Consistent with Figure 5, our method outperforms state-of-the-art methods across five metrics. For identity preservation, ConsisID achieves a higher score by designing appropriate identity signals for DiT from a frequency perspective. By contrast, ID-Animator is not optimized for IPT2V and only partially retains facial features, resulting in lower FaceSim-Arc and FaceSim-Cur scores. For Text Relevance, ConsisID not only controls expressions via prompts but also adjusts actions and backgrounds, achieving higher CLIPScore . Regarding visual quality, the FID is presented solely as a reference due to its limited alignment with human perception. Please refer to Figure 5 and 4 for qualitative analysis of the visual quality.
4 User Study
Building on previous work, we conduct a human evaluation using a binary voting strategy, with each questionnaire containing only 80 questions. Participants are required to view 40 video clips, a setup designed to improve both engagement and questionnaire validity. For the IPT2V task, each question requires participants to separately judge which option performs better in terms of Identity Preservation, Visual Quality, Text Alignment, and Motion Amplitude. This composition ensures the accuracy of the human evaluation. Owing to the extensive participant base required for this evaluation, we successfully gathered 103 valid questionnaires. The results, depicted in Figure 4, demonstrate a significant superiority of our method over ID-Animator , verifying the effectiveness of the designed DiT for IPT2V generation.
5 Effect of the Identity Signal Injection in DiT
To assess the effectiveness of Finding 1 and Finding 2, we perform ablation experiments on different methods of injecting control signals into DiT. Specifically, these experiments involved (a) injecting only low-frequency face information with key points into the noise latent, (b) injecting only high-frequency face signals within the attention block, (c) combining (a) and (b), (d) based on (c), but the low-frequency face information does not contain key points, and (e - f) based on (c), but the high-frequency signal is injected at the output or input of the attention block. (g) injecting only high-frequency face signals before the attention block. The results are shown in Figure 7 and Table 3. For Finding 1, we observe that only injecting high-frequency signals (a) greatly increases the training difficulty, causing the model to fail to converge due to the lack of low-frequency signal injection. In addition, the inclusion of facial key points (d) allows a greater focus on low-frequency information, thereby facilitating training and improving model performance. For Finding 2, when only low-frequency signals are injected (b), the model lacks high-frequency information. This reliance on low-frequency signals causes the generated face in the video to copy the reference image, making it difficult to control facial expressions, movements, and other features through prompts. Furthermore, injecting identity signals into the attention block input (f - g) disrupts the intended frequency domain distribution of DiT, resulting in a gradient explosion. Embedding control signals in the attention block (c) is preferable to embedding them in the output (e) because attention block processes predominantly low-frequency information. By embedding high-frequency information internally, the attention block is guided to highlight intrinsic facial features, whereas injecting it into the output merely concatenates features without directing focus, reducing DiT’s modeling capacity. Moreover, we apply a Fourier transform to the generated videos (only the face region) to visually compare the influence of different components to extract facial information. As shown in Figure 3, the Fourier spectrum and the log amplitude of the Fourier transform reveal that injecting high or low-frequency signals can indeed enhance the corresponding frequency information of the generated face. Moreover, the low-frequency signal can be further enhanced by matching with the face key points, and injecting the high-frequency signal into the attention block has the highest feature utilization rate. Our method (c) shows strongest high and low frequency, further validating the efficiency benefit from Findings 1 and 2. To reduce overhead, for each identity, we only select 2 reference images for the evaluation.
6 Ablation on the Consistency Training Strategy
To reduce overhead, for each identity, we only select 2 reference images for the following experiments. To demonstrate the benefits of the proposed consistency training strategy, we perform ablation experiments on coarse-to-fine training (CFT), dynamic mask loss (DML), and dynamic cross-face loss (DCL), with the results presented in Figure 6 and Table 2. When CFT is removed, GFE and LFE exhibit competing behaviors, complicating the model’s ability to prioritize high and low-frequency information accurately, leading to convergence at suboptimal points. Removing DML required the model to simultaneously focus on both foreground and background elements, with background noise negatively affecting training and reducing facial consistency. Similarly, the exclusion of DCL impaired the generalization capability, reducing fidelity for faces, not in the training set and reducing its effectiveness in generating identity-preserving videos as intended.
7 Ablation on the Number of Inversion Steps
To assess the impact of varying the number of inversion steps on model performance, we conduct an ablation study within the inference phase of ConsisID. Given constraints on computing resources, 60 prompts are randomly selected from the evaluation dataset. Each prompt is paired with a unique reference image, leading to the generation of 60 videos for each setting. Using a fixed random seed, we vary the inversion step parameter across values of 25, 50, 75, 100, 125, 150, 175, and 200. The results are illustrated in Figure 8 and Table 4. Although theoretical expectations suggest that increasing the number of inversion steps would continuously enhance the generation quality, our findings indicate a non-linear relationship where quality peaks at and subsequently declines. Specifically, at , the model produces incomplete garlands; at , it fails to generate upper body clothing; beyond , it loses critical low-frequency facial information, resulting in distorted facial features; and beyond , the visual clarity progressively deteriorates. We infer that the initial stages of denoising process are dominated by low-frequency information, such as generating the outline of a face, while the later stages focus on high-frequency details, such as intrinsic facial features. is just the optimal setting to balance these two stages.
Conclusion
In this paper, we present ConsisID, a unified framework for keeping faces consistent in video generation by frequency decomposition. It can seamlessly integrate into existing DiT-based text-to-video models, for generating high-quality, editable, consistent identity-preserving videos. Extensive experiments show that ConsisID outperforms the current state-of-the-art identity-preserving T2V models. It reveals that our frequency-aware heuristic DiT-based control scheme is an optimal solution for IPT2V generation.
Limitations and Future Work. Existing metrics do not accurately measure the capabilities of different ID preservation models. Although ConsisID can generate realistic and natural videos following a text prompt, metrics such as CLIPScore and FID show little difference from previous methods. A viable direction is to find a metric that is more in line with human perception.