BlockDance: Reuse Structurally Similar Spatio-Temporal Features to Accelerate Diffusion Transformers

Hui Zhang, Tingwei Gao, Jie Shao, Zuxuan Wu

Introduction

Diffusion models have been recognized as a pivotal advancement for both image and video generation tasks due to their impressive capabilities. Recently, there has been a growing interest in shifting the architecture of diffusion models from U-Net to transformers ,i.e. DiT. This refined architecture empowers these models not just to generate visually convincing and artistically compelling images and videos, but also to better adhere to scaling laws.

Despite their remarkable performance, DiTs are limited in real-time scenarios due to slow inference speeds caused by their iterative denoising process. Existing acceleration approaches mainly focus on two paradigms: \@slowromancapi@) reducing the number of sampling steps through novel scheduler designs or step distillation ; \@slowromancapii@) minimizing computational overhead per step through the employment of model pruning , model distillation , or the mitigation of redundant calculations . This paper aims to accelerate DiTs by mitigating redundant computation, as this paradigm can be plug-and-play into various models and tasks. Although feature redundancy is widely recognized in visual tasks , and recent works have identified its presence in the denoising process of diffusion models , the issue of feature redundancy within DiT models—and the potential strategies to mitigate this redundant computation —remains obscured from view.

To this end, we revisit the inter-feature distances between the blocks of DiTs at adjacent time steps in Figure 2 (a) and propose BlockDance, a training-free acceleration approach by caching and reusing highly similar features to reduce redundant computation. Previous feature reuse methods lack tailor-made reuse strategies for features at different scales with varying levels of similarity. Consequently, the reused set often includes low-similarity features, leading to structural distortions in the image and misalignment with the prompt. In contrast, BlockDance enhances the reuse strategy and focuses on the most similar features, i.e. Structurally Similar Spatio-Temporal (STSS) features. To be specific, during the denoising process, structural content is typically generated in the initial steps when noise levels are high, whereas texture and detail content are frequently generated in subsequent steps characterized by lower noise levels . Thus, we hypothesize that once the structure is stabilized, the structural features will undergo minimal changes. To validate this hypothesis, we decouple the features of DiTs at different scales, as illustrated in Figure 2 (b). The observation reveals that the shallow and middle blocks, which concentrate on coarse-grained structural content, exhibit minimal variation across adjacent steps. In contrast, the deep blocks, which prioritize fine-grained textures and complex patterns, demonstrate more noticeable variations. Thus, we argue that allocating computational resources to regenerate these structural features yields marginal benefits while incurring significant costs. To address this issue, we propose a strategy of caching and reusing highly similar structural features subsequent to the stabilization of the structure to accelerate DiTs while maximizing consistency with the generated results of the original model.

Considering the diverse nature of generated content and their varying distributions of redundant features, we introduce BlockDance-Ada, a lightweight decision-making network tailored for BlockDance. In simpler content with a limited number of objects, we observe a higher presence of redundant features. Therefore, frequent feature reuse in such scenarios is adequate to achieve satisfactory results while offering increased acceleration benefits. Conversely, in the context of intricate compositions characterized by numerous objects and complex interrelations, there are fewer high-similarity features available for reuse. Learning this adaptive strategy is a non-trivial task, as it involves non-differentiable decision-making processes. Thus, BlockDance-Ada is built upon a reinforcement learning framework. BlockDance-Ada utilizes policy gradient methods to drive a strategy for caching and reusing features based on the prompt and intermediate latent, maximizing a carefully designed reward function that encourages minimizing computation while maintaining quality. Accordingly, BlockDance-Ada is capable of adaptively allocating resources.

BlockDance has been validated across diverse datasets, including ImageNet, COCO, and MSR-VTT. It has been tested in tasks such as class-conditioned generation, text-to-image, and text-to-video using models such as DiT-XL/2, Pixart-α\alpha, and Open-Sora. Experimental results reveal that our method can achieve a 25%-50% acceleration in inference speed with training-free while maintaining comparable generated quality. Furthermore, the proposed BlockDance-Ada generates higher-quality content at the same acceleration ratio.

Related Work

Diffusion models have emerged as key players in the field of generation due to their impressive capabilities. Previously, U-Net-based diffusion models have demonstrated remarkable performance across various applications, including image generation and video generation . Recently, some research has transitioned to transformer-based architectures, i.e. Diffusion Transformers (DiTs). This framework excels in generating visually convincing and artistically compelling content, better adheres to scaling laws, and shows promise in efficiently integrating and generating multi-modality content. However, DiT models are still hindered by the inherent iterative nature of the diffusion process, limiting their real-time applications.

Efforts have been made to accelerate the inference process of diffusion models, which can be summarized into two paradigms: reducing the number of sampling steps and reducing the computation per step. The first paradigm often involves designing faster samplers or step distillation . The second paradigm focuses on model-level distillation , pruning , or reducing redundant computation . Several studies have unearthed the existence of redundant features in U-Net-based diffusion models, but their coarse-grained feature reuse strategies include those low-similarity features, leading to structural distortions and text-image misalignment. In contrast, we investigate the feature redundancy in DiTs and propose reusing structurally similar spatio-temporal features to achieve acceleration while maintaining high consistency with the base model’s results.

Efforts have been dedicated to fine-tuning diffusion models using reinforcement learning to align their outputs with human preferences or meticulously crafted reward functions. Typically, these models aim to improve the prompt alignment and visual aesthetics of the generated content. This paper explores learning instance-specific acceleration strategies through reinforcement learning.

Methodology

Diffusion models gradually add noise to the data and then learn to reverse this process to generate the desired noise-free data from noise. In this paper, we focus on the formulation introduced by that performs noise addition and denoising in latent space. In the forward process, the posterior probability of the noisy latent zt\mathbf{z}_{t} at time step tt has a closed form:

where αˉt=∏i=0tαi=∏i=0t(1−βi)\bar{\alpha}_{t}=\prod_{i=0}^{t}\alpha_{i}=\prod_{i=0}^{t}(1-\beta_{i}) and βi∈(0,1)\beta_{i}\in(0,1) represents the noise variance schedule. The inference process, i.e. the reverse process of generating data from noise, is a crucial part of the diffusion model framework. Once the diffusion model ϵθ(zt,t)\epsilon_{\theta}(\mathbf{z}_{t},t) is trained, during the reverse process, traditional sampler DDPM denoise zT∼N(0,I)\mathbf{z}_{T}\sim\mathcal{N}(\mathbf{0},\mathbf{I}) step by step for a total of TT steps. One can also use a faster sampler like DDIM to speed up the sampling process via the following process:

In the denoising process, the model generates rough structures of the image in the early stages and gradually refines it by adding textures and detailed information in later stages.

The number of denoising steps is related to the number of network inferences in the DiT architecture, which typically features multiple blocks stacked together. Each block sequentially computes its output based on the input from the previous block. The shallow blocks, closer to the input, are inclined to capture the global structures and rough outlines of the data. In contrast, the deep blocks, closer to the output, gradually refine specific details to achieve outputs that are both realistic and visually appealing .

2 Feature Similarity and Redundancy in DiTs

The inference speed of DiTs is limited by its inherently iterative nature of inference. This paper aims to reduce redundant computation to accelerate DiTs.

Upon revisiting the denoising process in various DiT models, including DiT-XL/2 , PixArt-α\alpha , and Open-Sora , two key findings emerged: \@slowromancapi@) There are significant feature similarities between consecutive steps, indicating redundant computation in the denoising process, as illustrated in Figure 2 (a); \@slowromancapii@) This high similarity is primarily manifested in the shallow and middle blocks (between 0 and 20 blocks) of the transformer, while deeper blocks (between 21 and 27 blocks) exhibit more variations, as depicted in Figure 2 (b). We attribute this phenomenon to the fact that structural content is generally produced in the initial steps, while textures and details are generated in the later steps.

To confirm this, we visualize the block features of PixArt-α\alpha using Principal Components Analysis (PCA), as shown in Figure 3. At the initial stages of denoising, the network primarily focuses on generating structural content, such as human poses and other basic forms. As the denoising process progresses, the shallow and middle blocks of the network still concentrate on generating low-frequency structural content, while the deeper blocks shift their focus towards generating more complex high-frequency texture information, such as clouds and crowds within depth of field. Consequently, after the structure is established, the feature maps highlighted by blue boxes in Figure 3 exhibit high consistency across adjacent steps. We define this computation as redundant computation, which relates to the low-level structures that the shallow and middle blocks of the transformer focus on. Based on these observations, we argue that allocating substantial computational resources to regenerate these similar features yields marginal benefits but leads to higher computational costs. Thus, our goal is to design a strategy that leverages these highly similar features to reduce redundant computation and accelerate the denoising process.

3 Training-free acceleration approach

We introduce BlockDance, a straightforward yet effective method to accelerate DiTs by leveraging feature similarities between steps in the denoising process. By strategically caching the highly similar structural features and reusing them in subsequent steps, we reduce redundant computation.

Specifically, we design the denoising steps into two types: cache step and reuse step, as illustrated in Figure 4. During consecutive time steps, a cache step first conducts a standard network forward based on zt+1\mathbf{z}_{t+1}, outputs zt\mathbf{z}_{t}, and saves the features FtiF_{t}^{i} of the ii-th block. For the following time step—the reuse step—we do not perform full network forward computation; instead, we carry out partial inference. More specifically, we reuse the cached features FtiF_{t}^{i} from the cache step as the input for the (i+1)(i+1)-th block in the reuse step. Therefore, the computation of the first ii blocks in the reuse step can be saved due to the sequential inference characteristic of the transformer blocks, and only the blocks deeper than ii require recalculation.

To this end, it is crucial to determine the optimal block index and the stage of the denoising process where reuse should be concentrated. Based on the insights from Figure 2 and Figure 3, we set the index as 20 and focus the reuse on the latter 60% of the denoising process, after the structure has stabilized. These settings enable the decoupling of feature reuse and specifically reuse the structurally similar spatio-temporal features. Thus, we set the first 40% of denoising steps as cache steps and evenly divide the remaining 60% of denoising steps into several groups, each comprising NN steps. The first step of each group is designated as a cache step, while the subsequent N−1N-1 steps are reuse steps to accelerate inference. With the arrival of a new group, a new cache step updates the cached features, which are then utilized for the reuse steps within that group. This process is repeated until the denoising process concludes. A larger NN represents a higher reuse frequency. We term this cache and reuse strategy as BlockDance-NN, which operates in a training-free paradigm and can effectively accelerate multiple types of DiTs while maintaining the quality of the generated content.

4 Instance-specific acceleration approach

However, the generated content exhibits varying distributions of feature similarity, as shown in Figure 6. We visualize the cosine similarity matrix of features from block index i≤20i\leq 20 at each denoising step compared to other steps. We find that the distribution of similar features is related to the structural complexity of the generated content. In Figure 6, as structural complexity increases from left to right, the number of similar features suitable for reuse decreases. To improve the BlockDance strategy, we introduce BlockDance-Ada, a lightweight framework that learns instance-specific caching and reuse strategies.

Here, each entry in m\mathbf{m} is normalized to be in the range , indicating the likelihood of performing a cache step. We define a reuse policy πf(u∣zρ,c)\pi^{f}(\mathbf{u}\mid\mathbf{z}_{\rho},\mathbf{c}) with an (s−ρs-\rho)-dimensional Bernoulli distribution:

where u∈{0,1}(s−ρ)\mathbf{u}\in\{0,1\}^{(s-\rho)} are actions based on m\mathbf{m}, and ut=1\mathbf{u}_{t}=1 indicates the tt-th step is a cache step, and zero entries in u\mathbf{u} are reuse steps. During training, u\mathbf{u} is produced by sampling from the corresponding policy, and a greedy approach is used at test time. With this approach, DiTs generate the latent z0\mathbf{z}_{0} based on the reuse policy, following the decoder DD decodes the latent into a pixel-level image x\mathbf{x}.

Based on this, we design a reward function to incentivize fdf_{d} to maximize computational savings while maintaining quality. The reward function consists of two main components: an image quality reward and a computation reward, balancing generation quality and inference speed. For the image quality reward Q(u)\mathcal{Q}(\mathbf{u}), we use the quality reward model fqf_{q} to score the generated images based on visual aesthetics and prompt adherence, i.e. Q(u)=fq(x)\mathcal{Q}(\mathbf{u})=f_{q}(\mathbf{x}). The computation reward C(u)\mathcal{C}(\mathbf{u}) is defined as the normalized number of reuse steps, given by the formula:

We use samples in mini-batches to compute the expected gradient and approximate Eqn. 6 to:

where B is the number of samples in the mini-batch. The gradient is then propagated back to train the fdf_{d} with Adam optimizer. Following this training process, the decision network perceives instance-specific cache and reuse strategies, thereby achieving efficient dynamic inference.

Experiments

We conduct evaluations on class-conditional image generation, text-to-image generation, and text-to-video generation. For class-conditional image generation, we use DiT-XL/2 to generate 50 512 ×\times 512 images per class on the ImageNet dataset via DDIM sampler , with a guidance scale of 4.0. For text-to-image generation, we used PixArt-α\alpha to generate 1024 ×\times 1024 images on the 25K validation set of COCO2017 via DPMSolver sampler , with a guidance scale of 4.5. For text-to-video generation, we used Open-Sora 1.0 to generate 16-frame videos at 512 ×\times 512 resolution on the 2990 test set of MSR-VTT via DDIM sampler, with a 7.0 guidance scale. We follow previous works to evaluate these tasks and report IQS score and Pickscore for text-to-image generation. We measure inference speed on A100 GPU by generation latency, i.e. time per image/video).

1.2 Implementation Details

For BlockDance, the cache and reuse steps in PixArt-α\alpha are primarily between 40% and 95% of the denoising process, while in DiT-XL/2 and Open-Sora, they are mainly between 25% and 95% of the denoising process. The sizes of the cached features for generating each content with these three models are 18MB, 4.5MB, and 72MB, respectively. The default block index ii is set to 20. For BlockDance-Ada, we design the decision network as a lightweight architecture consisting of three transformer blocks and a multi-layer perceptron. The parameters of the decision network amount to 0.08B. We set ρ\rho to 40% of the total number of denoising steps. The parameter λ\lambda in the reward function is set to 2. For PixArt-α\alpha, we train the step selection network on 10,000 subset from its training dataset for 100 epochs with a batch size of 16. We use Adam with a learning rate of 10−510^{-5}.

2 Main Results

The results on the 25k COCO2017 validation set, as shown in Table 1, demonstrate the efficacy of BlockDance. We extend ToMe and DeepCache to PixArt-α\alpha as baselines. For ToMe, we reduce the computational cost by removing 25% of the tokens through the merge operation. For DeepCache, we reuse features at intervals of 2 throughout the denoising process, specifically reusing the outputs from the first 14 blocks out of the 28 blocks in PixArt-α\alpha. With N=2N=2, BlockDance accelerates PixArt-α\alpha by 25.4% with no significant degradation in image quality, both in terms of visual aesthetics and prompt following. Different speed-quality trade-offs can be modulated by NN.

Compared to ToMe, BlockDance consistently outperforms ToMe by a clear margin regardless of the reuse frequency NN. This can be attributed to DiTs featuring a more attention-intensive architecture than the U-Net-based one, thus the continuous use of token merging in DiTs exacerbates quality degradation. Compared to DeepCache, BlockDance achieves better performance across all image metrics at comparable speeds by focusing on high-similarity features. We specifically reduce redundant structural computation in the later stages of denoising, avoiding dissimilar feature reuse and minimizing image quality loss. However, DeepCache reuses features throughout the entire denoising process and does not specifically aim at highly similar features for reuse. This leads to the inclusion of dissimilar features in the reused set, resulting in structural distortions and a decline in prompt alignment. Compared to TGATE , which accelerates by reducing the redundancy in cross-attention calculations. Blockdance supports DiTs that do not incorporate cross-attention, such as SD3 and Flux . Besides, experimental results show that with the same acceleration benefit, BlockDance outperforms TGATE across various metrics. Compared to PixArt-LCM obtained through consistency distillation training, BlockDance, although requiring more inference time, achieves higher generation quality across multiple metrics without additional training. It is worth noting that BlockDance’s generated images exhibit higher consistency with the base model, as evidenced by significantly better SSIM performance compared to baselines, thanks to our targeted reuse strategy.

The results of the 50k ImageNet images are shown in Table 2. We extend ToMe and DeepCache to DiT/XL-2 as the baselines. At N=2N=2, BlockDance not only accelerates DiT/XL-2 by 37.4% while maintaining image quality but also outperforms DeepCache while preserving higher consistency with the base model’s generated images. By increasing, BlockDance’s acceleration ratio can reach up to 57.5%, and the image quality consistently outperforms ToMe.

BlockDance is also effective for accelerating video generation tasks. The accelerated results on MSR-VTT are shown in Table 3. At N=2N=2, BlockDance speeds up Open-Sora by 34.8% while maintaining video quality, both in terms of visual quality and temporal consistency. Increasing NN achieves various quality-speed trade-offs. In contrast, DeepCache suffers from significant quality degradation, as evidenced by the deterioration in metrics such as FVD. This is attributed to DeepCache reusing low-similarity features, such as structural information at the early stages of denoising.

We further qualitatively analyze our approach as shown in Figure 7. ToMe merges adjacent similar tokens to save self-attention computation, but this method is not very friendly for transformer-intensive architectures, resulting in low-quality images with “blocky artifacts”. While DeepCache and TGATE achieve approximately 27% acceleration, they may cause significant structural differences from the original images and present artifacts and semantic loss in some complex cases. PixArt-LCM accelerates PixArt-α\alpha through additional consistency distillation training, yielding significant acceleration but with noticeable declines in visual aesthetics and prompt following. In contrast, BlockDance achieves a 25.4% acceleration without extra training costs, maintaining high consistency with the original image in structure and detail.

2.2 Evaluation on BlockDance-Ada

Table 4 provides a detailed breakdown of the performance of BlockDance-Ada. By dynamically allocating computational resources for each sample based on instance-specific strategies, BlockDance-Ada effectively reduces redundant computation, achieving acceleration close to that of BlockDance (N=3N=3) while delivering superior image quality. Compared to BlockDance (N=2N=2), BlockDance-Ada offers greater acceleration benefits with similar quality.

3 Discussion.

As shown in Figure 8, we illustrate how the generated images evolve as NN increases. With the reduction in generation time, the main subject of the image remains consistent, but the fidelity of the details gradually decreases, aligning with the insights presented in Table 1. Different values of NN offer flexible choices for various speed-quality trade-offs.

As shown in Figure 9, we investigate the impact of applying BlockDance in different stages of the denoising process: the initial stage (0%-40%) and the later stage (40%-95%). BlockDance primarily reuses structural features; therefore, applying it at the initial stage that focuses on the structure may result in structural changes or artifacts (highlighted by the red box in Figure 9 (a)), as the structure has not yet stabilized. Conversely, in the later stage, where structural information has stabilized and the focus shifts towards texture details, reusing structural features accelerates inference with minimal quality loss, as shown in Figure 9 (b).

We investigate the impact of reusing only the shallow and middle blocks versus reusing deeper blocks as well in the transformer, as shown in Figure 10. Due to the low similarity of features in the deeper blocks, reusing them results in the loss of computation related to details, leading to degradation in texture details, as highlighted by the red boxes in Figure 10 (b). Conversely, reusing the higher similarity shallow and middle blocks, which focus on structural information, results in minimal quality degradation.

Conclusion

In this paper, we introduced BlockDance, a novel training-free acceleration approach by caching and reusing structure-level features after the structure has stabilized, i.e., structurally similar spatio-temporal features, BlockDance significantly accelerates DiTs with minimal quality loss and maintains high consistency with the base model. Additionally, we introduced BlockDance-Ada, which enhances BlockDance by dynamically allocating computational resources according to instance-specific reuse policies.

Although BlockDance accelerates various DiT models in a plug-and-play manner, it shows limited benefits in scenarios with very few denoising steps (e.g. 1 to 4 steps), due to reduced similarity between adjacent steps.

This work was supported by National Science and Technology Major Project (No. 2021ZD0112805) and ByteDance (No.CT20230914000147).

References

1 More justification for the design in BlockDance

We supplement the (step, block, L2 distance) 3D surface plot, as shown in (a) of Figure 11. BlockDance focuses on reusing high-similarity spatio-temporal features within the blue box. In contrast, other methods reuse features indiscriminately across all spatio-temporal zones framed by the purple box, which leads to quality degradation caused by reusing low-similarity features. To further demonstrate that the features within the blue box are primarily structurally similar, we directly calculate x0t=VAEdecoder(zt−1−αˉt⋅ϵθ(zt,t)αˉt)x_{0}^{t}={VAE}_{decoder}(\frac{z_{t}-\sqrt{1-\bar{\alpha}_{t}}\cdot\epsilon_{\theta}(z_{t},t)}{\sqrt{\bar{\alpha}_{t}}}) based on the noise predicted at each step and then compute the SSIM score of the images x0tx_{0}^{t} at adjacent time steps to measure structural similarity. As shown in (b) of Figure 11, compared to other methods that reuse features indiscriminately, we reuse the computations focusing on structural features (i.e. shallow and middle blocks in DiT) after the structure stabilizes, thereby maintaining high consistency with the original content.

2 Additional Experiments

The acceleration paradigm we proposed is complementary to other acceleration techniques and can be used on top of them for further enhancement. Here, we validate the performance of BlockDance across different sampling steps for each model. As demonstrated in Tables 5, 6, and 7, BlockDance effectively accelerates the process across various step counts while maintaining the quality of the generated content.

To validate the effectiveness of our proposed paradigm across different DiT architecture variants, we apply BlockDance to MMDiT-based DiT models , such as Stable Diffusion 3 . The results are conducted on the 25k COCO2017 validation set, as shown in Table 8. The experimental results indicate that with N=2N=2, BlockDance accelerates SD3 by 25.3% while maintaining comparable image quality, both in terms of visual aesthetics and prompt following. Different speed-quality trade-offs can be modulated by N.

To comprehensively verify the method we proposed, we present additional qualitative results for each DiT model, as indicated in Figures 12, 13, and 14. Our method maintains high-quality content with a high degree of consistency with the content generated by the original models, while achieving significant acceleration.