Matten: Video Generation with Mamba-Attention
Yu Gao, Jiancheng Huang, Xiaopeng Sun, Zequn Jie, Yujie Zhong, Lin Ma
Introduction
Recent advancements in diffusion models have demonstrated impressive capabilities in video generation . It has been observed that breakthroughs in architectural design are crucial for the efficient application of these models . Contemporary studies largely concentrate on CNN-based U-Net architectures and Transformer-based frameworks , both of which employ attention mechanisms to process spatio-temporal dynamics in video content. Spatial attention, which involves computing self-attention among image tokens within a single frame, is extensively utilized in both U-Net-based and Transformer-based video generation diffusion models as shown in Fig. 1 (a). Prevailing techniques typically apply local attention within the temporal layers as illustrated in Fig. 1 (b), where attention calculations are confined to identical positions across different frames. This approach fails to address the critical aspect of capturing interrelations across varying spatial positions in successive frames. A more effective method for temporal-spatial analysis would involve mapping interactions across disparate spatial and temporal locations, as depicted in Fig. 1 (c). Nonetheless, this global-attention method is computationally intensive due to the quadratic complexity involved in computing attention, thus requiring substantial computational resources.
There has been a rise in fascination with state space models (SSMs) across a variety of fields, largely due to their ability to deal with long sequences of data . In the field of Natural Language Processing (NLP), innovations such as the Mamba model have significantly improved both the efficiency of data inference processes and the overall performance of models by introducing dynamic parameters into the SSM structure and by building algorithms tailored for better hardware compatibility. The utility of the Mamba framework has been successfully extended beyond its initial applications, demonstrating its effectiveness in areas such as vision and multimodal applications . Given the complexity of processing video data, we propose to use the Mamba architecture to explore spatio-temporal interactions in video content, as shown in Fig. 1 (d). However, unlike the self-attention layer, it’s important to note that Mamba scans, which do not inherently compute dependencies between tokens, struggle to effectively detect localised data patterns, a limitation pointed out by .
Regarding the advantages of Mamba and Attention, we introduce a latent diffusion model for video generation with a Mamba-Attention architecture, namely Matten. Specifically, we investigated the impact of various combinations of Mamba and Attention mechanisms on video generation. Our findings demonstrate that the most effective approach is to utilize the Mamba module to capture global temporal relationships (Fig. 1 (d)) while employing the Attention module for capturing spatial and local temporal relationships (Fig. 1 (a) and Fig. 1 (b)).
We conducted experimental evaluations to examine the performance and effects of Matten in both unconditional and conditional video generation tasks. Across all test benchmarks, Matten consistently exhibits the comparable FVD score and efficiency with SOTAs. Furthermore, our results indicate that Matten is scalable, evidenced by the direct positive relationship between the model’s complexity and the quality of generated samples.
In summary, our contributions are as follows:
We propose Matten, a novel video latent diffusion model integrated with the mamba block and attention operations, which enables efficient and superior video generation.
We design four model variants to explore the optimal combination of Mamba and attention in video generation. Based on these variants, we find that the most favorable approach is adopting attention mechanisms to capture local spatio-temporal details and utilizing the Mamba module to capture global information.
Comprehensive evaluations show that our Matten achieves comparable performance to other models with lower computational and parameter requirements and exhibits strong scalability.
Related Work
The task of video generation primarily focuses on producing realistic video clips characterized by high-quality visuals and fluid movements. Previous video generation work can be grouped into 3 types . Initially, a number of researchers focused on adapting powerful GAN-based image generation techniques for video creation . Nonetheless, GAN-based methods may lead to problems such as mode collapse, reducing diversity and realism.
In addition, certain models suggest the learning of data distributions via autoregressive models . These methods typically yield high-quality videos and demonstrate more reliable convergence, but they are hindered by their substantial computational demands. Finally, the latest strides in video generation are centered on the development of systems that utilize diffusion models , which have shown considerable promise. These methods primarily use CNN-based U-Net or Transformer as the model architecture. Distinct from these works, our method concentrates on investigating the underexplored area of the combination of mamba and attention within video diffusion.
2 Mamba
Mamba, a new State-Space Model, has recently gained prominence in deep learning for its universal approximation capabilities and efficient modeling of long sequences, with applications in diverse fields such as medical imaging, image restoration, graphs, NLP, and image generation . Drawing from control systems and leveraging HiPPO initialization , these models, like LSSL , address long-range dependencies but are limited by computational demands. To overcome this, S4 and other structured state-space models introduce various configurations and mechanisms that have been integrated into larger representation models for tasks in language and speech. Mamba, and its iterations like VisionMamba , S4ND , and Mamba-ND , exhibit a range of computational strategies, from bidirectional SSMs to local convolution and multi-dimensionality considerations. For 3D imaging, T-Mamba tackles the challenges in orthodontic diagnosis with Mmaba to handle long-range dependencies. For video understanding, VideoMamba and Video Mamba Suite adapt Mamba to the video domain and address the challenges of local redundancy and global dependencies prevalent in video data. In the domain of diffusion applications using mamba, Zigzag Mamba advances the scalability and efficiency of generating visual content. It tackles the crucial problem of spatial continuity with an innovative scanning approach, incorporates text-conditioning features, and shows enhanced performance across high-resolution image and video datasets. closely relates to our work, employing the mamba block in the temporal layer of video diffusion. Diverging from previous research focused mainly on local temporal modeling, our method, Matten, is uniquely designed to encompass global temporal dimensions.
Methodology
Our discussion starts with a brief overview of the latent space diffusion model and state space model in Sec. 3.1. This is followed by an in-depth description of the Matten model variants in Sec. 3.2. We then explore conditional ways related to timestep or class in Sec. 3.3. Lastly, a theoretical analysis comparing Mamba with Attention mechanisms is presented in Sec. 3.4.
Latent Space Diffusion Models. . For an input data sample , Latent Diffusion Models (LDMs) initially utilize the pre-trained VAE or VQ-VAE encoder to transform the data sample into a latent representation . This transformation is followed by a learning phase where the data distribution is modeled through diffusion and denoising steps.
During the diffusion phase, noise is incrementally added to the latent encoding, producing a series of increasingly perturbed latent states , where the intensity of additive noise is denoted by the timesteps . A specialized model such as U-Net is utilized as the noise estimate network to estimate the noise perturbations affecting the latent representation during the denoising phase, aiming to minimize the latent diffusion objective.
Furthermore, the diffusion models are enhanced with a learned reverse process covariance , optimized using as outlined by .
In our research, is designed using a Mamba-based framework. Both and are employed to refine the model’s effectiveness and efficiency.
State Space Backbone. State space models (SSMs) have been rigorously validated both theoretically and through empirical evidence to adeptly manage long-range dependencies, demonstrating linear scaling with the length of data sequences. Conventionally, a linear state space model is represented as the following type:
The process of discretization, essential for applying state space models as detailed in Eq. 2 to real-world deep learning tasks, converts continuous system parameters like and into their discrete equivalents and . This critical step typically utilizes the zero-order hold (ZOH) method, a technique well-established in academic research for its efficacy. The ZOH method uses the timescale parameter to bridge the gap between continuous and discrete parameters, thereby facilitating the application of theoretical models within computational settings.
With these discretized parameters, the model outlined in Eq. 2 is then adapted to a discrete framework using a timestep :
This approach allows for the seamless integration of state space models into digital platforms. The traditional Mamba block, initially crafted for 1D sequence processing as shown in Fig. 2, is not ideally suited for visual tasks that demand spatial cognizance. To address this limitation, Vision Mamba has developed a bidirectional Mamba block specifically tailored for vision-related applications. This innovative block is engineered to handle flattened visual sequences by employing both forward and backward SSMs concurrently, significantly improving its ability to process with spatial awareness.
Mamba employs a work-efficient parallel scan that effectively reduces the sequential dependencies typically associated with recurrent computations. This optimization, coupled with the strategic utilization of GPU operations, eliminates the necessity to explicitly manage the expanded state matrix. In our study, we explore the integration of the Mamba architecture within a video generation framework, leveraging its efficiency and scalability.
2 The Model Variants of Matten
Adopting a strategy similar to Latte, we assign , , and to structure the data effectively. Furthermore, a spatio-temporal positional embedding, denoted as , is incorporated into the token sequence . The input for the Matten model thus becomes , facilitating complex model interactions. As illustrated in Fig. 3, we introduce four distinct variants of the Matten model to enhance its versatility and effectiveness in video processing.
Global-Sequence Mamba Block with Spatial-Temporal Attention Interleaved. Although Mamba demonstrates efficient performance in long-distance modeling, its advantages in shorter sequences modeling are not as pronounced , compared to the attention operation in Transformer. Consequently, we have developed a hybrid block that leverages the strengths of both the attention mechanism and Mamba as illustrated in Fig. 3 (c), which integrates Mamba and Attention computations for both short and long-range modeling. Each block is composed of Spatial Attention computation, Temporal Attention computation, and a Global-Sequence Mamba scan in series. This design enables our model to effectively capture both the global and local information present in the latent space of videos.
Global-Sequence Mamba Block with Temporal Attention Interleaved.
The scanning in the Global-Sequence Mamba block is continuous in the spatial domain but discontinuous in the temporal domain . Thus, this variant has removed the Spatial Attention component, while retaining the Temporal Attention block. Consequently, by concentrating on a Spatial-First scan augmented with Temporal Attention shown in Fig. 3 (d), we strive to enhance our model’s efficiency and precision in processing the dynamic facets of video data, thereby assuring robust performance in a diverse range of video processing tasks.
3 Conditional Way of Timestep or Class
Drawing from the frameworks presented by Latte and DiS, we perform experiments on two distinct methodologies for embedding timestep or class information into our model. The first method, inspired by DiS, involves treating as tokens, a strategy we designate as conditional tokens. The second method adopts a technique akin to adaptive normalization (AdaN) , specifically tailored for integration within the Mamba block. This involves using MLP layer to compute parameters and from , formulating the operation , where denotes the feature maps in the Mamba block. Further, this adaptive normalization is implemented prior to residual connections of the Mamba block, implementing by the transformation , with representing the Bidirectional-Mamba scans within the block. We refer to this advanced technique as Mamba adaptive normalization (M-AdaN), which seamlessly incorporates class or timestep information to enhance model responsiveness and contextual relevance.
4 Analysis of Mamba and Attention
In summary, the hyperparameters of our proposed block encompass hidden size , expanded state dimension , and SSM dimension . All the settings of Matten are detailed in Table 2, covering different numbers of parameters and computation cost to thoroughly evaluate scalability performance. Specifically, the Gflop metric is analyzed during the generation of 16256256 unconditional videos, employing a patch size of . Consistent with , we standardize the SSM dimension across all models at 16.
involves the calculation with , , and , while denotes the calculation with . It demonstrates that self-attention’s computational demand scales quadratically with the sequence length , whereas SSM operations scale linearly. Notably, with typically fixed at 16, this linear scalability renders the Mamba architecture particularly apt for handling extensive sequences typical in scenarios like global relationship modeling in video data. When comparing the terms and , it is clear that the Mamba block is more computationally efficient than self-attention, particularly when the sequence length significantly exceeds . For shorter sequences that focus on spatial and localized temporal relationships, the attention mechanism offers a more computationally efficient alternative when the computational overhead is manageable, as corroborated by empirical results.
Experiments
This part first describes the experimental settings, including details about the datasets we used, evaluation metrics, compared methods, configurations of the Matten model, and specific implementation aspects. Following this, ablation studies are conducted to identify optimal practices and assess the impact of model size. The section concludes with a comparative analysis of our results on 4 common datasets against advanced video generation methods.
Datasets Overview. We engage in extensive experiments across four renowned and common datasets: FaceForensics , SkyTimelapse , UCF101 , and Taichi-HD . Following protocols established in Latte, we utilize predefined training and testing divisions. From these datasets, we extract video clips consisting of 16 frames, applying a sampling interval of 3, and resize each frame to a uniform resolution of 256x256 for our experiments.
Evaluation Metrics. For robust quantitative analysis, we adopt the Fréchet Video Distance (FVD) , recognized for its correlation with human perceptual evaluation. In compliance with the methodologies of StyleGAN-V, we determine FVD scores by examining 2,048 video clips, each containing 16 frames.
Baseline Comparisons. Our study includes comparisons with advanced methods to assess the performance of our approach quantitatively, including MoCoGAN , VideoGPT , MoCoGAN-HD , DIGAN , StyleGAN-V , PVDM , MoStGAN-V , LVDM , and Latte . Unless explicitly stated otherwise, all presented values are obtained from the latest relevant studies: Latte, StyleGAN-V, PVDM, or the original paper.
Matten Model Configurations. Our Matten model is structured using a series of Mamba blocks, with each block having a hidden dimension of . Inspired by the Vision Transformer (ViT) approach, we delineate four distinct configurations varying in parameter count, detailed in Table 3.
Implementation Specifics. All ablation experiments adopt the AdamW optimizer, set at a fixed learning rate of . The sole augmentation technique applied is horizontal flipping. Consistent with prevailing strategies in generative modeling , we employ the exponential moving average (EMA) of the model weights with a decay rate of 0.99 at the first 50k steps and the other 100k steps during the training process. The results reported are derived directly using the EMA-enhanced models. Additionally, the architecture benefits from the integration of a pre-trained variational autoencoder, sourced from Stable Diffusion v1-4.
2 Ablation study
In this part, we detail our experimental investigations using the SkyTimelapse dataset to assess the impact of various design modifications, model variations, and model sizes on performance, as previously introduced in Secs. 3.3 and 3.2.
Timestep-Class Information Injection Illustrated in Fig. 8b, the M-AdaN approach markedly outperforms conditional tokens. We surmise this difference stems from the method of integration of timestep or class information. Conditional tokens are introduced directly into the model’s input, potentially creating a spatial disconnect within the Mamba scans. In contrast, M-AdaN embeds both timestep and class data more cohesively, ensuring uniform dissemination across all video tokens, and enhancing the overall synchronization within the model.
Exploring Model Variants Our analysis of Matten’s model variants, as detailed in Sec. 3.2, aims to maintain consistency in parameter counts to ensure equitable comparisons. Each variant is developed from the ground up. As depicted in Fig. 8a, Variant 3 demonstrates superior performance with increasing iterations, indicating its robustness. Conversely, Variants 1 and 2, which focus primarily on local or global information, respectively, lag in performance, underscoring the necessity for a balanced approach in model design.
Assessment of Model Size We experiment with four distinct sizes of the Matten model—XL, L, B, and S as listed in Tab. 3 on the SkyTimelapse dataset. The progression of their Fréchet Video Distances (FVDs) with training iterations is captured in Fig. 9. There is a clear trend showing that larger models tend to deliver improved performance, echoing findings from other studies in image and video generation , which highlight the benefits of scaling up model dimensions.
3 Comparison Experiment
According to the findings from the ablation studies presented in Sec. 4.2, we have pinpointed the settings about how to design our Matten, notably highlighting the efficacy of model variant 3 equipped with M-AdaN. Leveraging these established best practices, we proceed to conduct comparisons against contemporary state-of-the-art techniques.
Qualitative Assessment of Results Figures 4 through 7 display the outcomes of video synthesis using various methods across datasets such as UCF101, Taichi-HD, FaceForensics, and SkyTimelapse. Across these different contexts, our method consistently delivers realistic video generations at a high resolution of 256x256 pixels. Notable achievements include accurately capturing facial motions and effectively handling dynamic movements of athletes. Our model particularly excels in generating high-quality videos on the UCF101 dataset, an area where many other models frequently falter. This capability underscores our method’s robustness in tackling complex video synthesis challenges.
Quantitative results. Tab. 1 presents the quantitative results of each comparative method. Overall, our method surpasses prior works and matches the performance of methods with image-pretrained weights, demonstrating our method’s superiority in video generation. Furthermore, our model attains roughly a 25% reduction in flops compared to Latte, the latest Transformer-based model. Given the abundance of released pre-trained U-Net-based (Stable Diffusion, SDXL) or Transformer-based (DiT, PixArt) image generation models, these U-Net-based or Transformer-based video generation models can leverage these pre-trained models for training. However, there are no released, pre-trained Mamba-based image generation models yet, so our model has to be trained from scratch. We believe that once Mamba-based image generation models become available, they will be of great help in training our Matten.
Conclusion
This paper proposes a simple diffusion method for video generation, Matten, with the Mamba-Attention structure as the backbone for generating videos. To explore the quality of Mamba for generating videos, we explore different configurations of the model, including four model variants, time step and category information injection, and model size. Extensive experiments demonstrate that Matten excels in four standard video generation benchmarks and displays impressive scalability.