M-VAR: Decoupled Scale-wise Autoregressive Modeling for High-Quality Image Generation
Sucheng Ren, Yaodong Yu, Nataniel Ruiz, Feng Wang, Alan Yuille, Cihang Xie
Introduction
Autoregressive models (Radford et al., 2018; Brown et al., 2020) have been instrumental in advancing the field of natural language processing (NLP). By modeling the probability distribution of a token given the preceding ones, these models can generate coherent and contextually relevant text. Prominent examples like GPT-3 (Brown et al., 2020) and its successors (OpenAI, 2022; 2023) have demonstrated remarkable capabilities in language understanding and generation, setting new benchmarks across various NLP applications.
Building upon the success in NLP, the autoregressive modeling paradigm (Yu et al., 2022; Sun et al., 2024; Van den Oord et al., 2016) has also been extended to computer vision for image generation tasks, aiming to generate high-fidelity images by predicting visual content in a sequential manner. Recently, VAR (Tian et al., 2024) has further enhanced this image autoregressive pipeline by structurally reformulating the learning target into a coarse-to-fine “next-scale prediction”, which innately introduces strong semantics to interconnecting tokens along scales. As demonstrated in the VAR paper, this pipeline exhibits much stronger scalability and can achieve competitive, sometimes even superior, performance compared to advanced diffusion models.
This paper aims to further optimize VAR’s computation structure. Our key insight lies in decoupling VAR’s cross-scale autoregressive modeling into two distinct parts: intra-scale modeling and inter-scale modeling. Specifically, intra-scale modeling involves bidirectionally modeling multiple tokens within each scale, capturing intricate spatial dependencies and preserving the 2D structure of images. In contrast, inter-scale modeling focuses on unidirectional causality between scales by sequentially progressing from coarse to fine resolutions — each finer scale is generated conditioned on all preceding coarser scales, ensuring that global structures guide the refinement of local details. Notably, the sequence length involved in inter-scale modeling is much longer than that of intra-scale modeling, resulting in significantly higher computational costs. But meanwhile, our analysis of attention scores for both intra-scale and inter-scale interactions (as discussed in Sec. 3.2) suggests a contrasting reality: intra-scale interactions dominate the model’s attention distribution, while inter-scale interactions contribute significantly less.
Motivated by the observations above, we propose to develop a more customized computation configuration for VAR. For the intra-scale component, given the much shorter sequence lengths within each scale and its significant contribution to the model’s attention distribution, we retain the bidirectional attention mechanism to fully capture comprehensive spatial dependencies. This ensures that local spatial relationships and fine-grained details are effectively modeled at a reasonable computational overhead. Conversely, for the inter-scale component, which involves much longer sequences but demands relatively less comprehensiveness in modeling global relationships, we adopt Mamba (Gu & Dao, 2023; Dao & Gu, 2024), a linear-complexity mechanism, to handle such inter-scale dependencies efficiently.
By segregating these two modeling modules and applying appropriate mechanisms to each, our approach significantly reduces computational complexity while preserving the model’s ability to maintain 2D spatial coherence and unidirectional coarse-to-fine consistency, making it well-suited for high-quality image generation. As shown in Figure 2, our proposed framework, which we term M-VAR, outperforms existing models in both image quality and inference speed. For instance, comparing to the largest VAR model, which has 2B parameters and attains an FID score of 1.97, our 1.5B parameter M-VAR model achieves an FID score of 1.93 with much fewer parameters and 1.2 faster inference speed. M-VAR also scales well — our largest model, M-VAR-d32, achieves an impressive FID score of 1.78 on ImageNet at resolution, outperforming the prior best autoregressive models LlamaGen by 0.4 and VAR by 0.19, respectively, and well-known diffusion models LDM by 1.82 and DiT by 0.49, respectively.
Related Work
Visual generation can generally be split into three categories: 1) Diffusion models (Dhariwal & Nichol, 2021; Rombach et al., 2022) which treat visual generation as the reverse process of the diffusion process; 2) Mask prediction models (Chang et al., 2022) which follow BERT-style (Devlin, 2018) language model to generate images by predicting mask tokens; and 3) Autoregressive models which generate images by predicting the next pixel/token/scale in a sequence. We focus on the last one in this paper.
The pioneering method that brings autoregressive modeling into visual generation is PixelCNN (Van den Oord et al., 2016), which models images by predicting the discrete probability distributions of raw pixel values, effectively capturing all dependencies within an image. Building on this foundation, VQGAN (Esser et al., 2021) advances the field by applying autoregressive learning within the latent space of VQVAE (Razavi et al., 2019), simplifying the data representation for more efficient modeling. The RQ Transformer (Lee et al., 2022) introduces a novel technique using a fixed-size codebook to approximate an image’s feature map with stacked discrete codes, forecasting the next quantized feature vectors by predicting subsequent code stacks. Parti (Yu et al., 2022) takes a different route by framing image generation as a sequence-to-sequence modeling task akin to machine translation, using sequences of image tokens as targets instead of text tokens, and thus capitalizing on the significant advancements made in large language models through data and model scaling. LlamaGen (Sun et al., 2024) further extends this concept by applying the traditional ”next-token prediction” paradigm of large language models to visual generation, demonstrating that standard autoregressive models like Llama can achieve state-of-the-art image generation performance when appropriately scaled, even without considering specific inductive biases for visual signals. Lastly, departing from the conventional raster-scan “next-token prediction” method, the recent work VAR (Tian et al., 2024) offers a new perspective by developing a coarse-to-fine “next-scale prediction” strategy for visual autoregressive modeling. Our work is a followup of VAR, aiming to making the whole framework more efficient.
2 Mamba
State-space models (SSMs) (Gu et al., 2021a; b) have recently emerged as a compelling alternative to Transformers (Vaswani, 2017), which employ hidden states to capture long-range dependencies efficiently. The latest advancement in this domain is Mamba (Gu & Dao, 2023; Dao & Gu, 2024), a sophisticated SSM that introduces data-dependent layers with expanded hidden states. Empirically, Mamba is able to construct a versatile language model backbone that not only rivals Transformers across various scales but also maintains linear scalability with respect to sequence length.
Building on Mamba’s success in NLP, its application has been quickly extended to computer vision tasks. Vision Mamba (Vim) (Zhu et al., 2024) utilizes pure Mamba layers within Vim blocks, leveraging both forward and backward scans to model bidirectional representations. This approach effectively addresses the direction-sensitive limitations of the original Mamba model. Additionally, ARM (Ren et al., 2024) pioneers the integration of autoregressive pretraining with Mamba in the vision domain.
For image generation, Diffusion Mamba (DiM) (Fei et al., 2024) combines the efficiency of the Mamba sequence model with diffusion processes to achieve high-resolution image synthesis. Specifically, DiM employs multi-directional scans, introduces learnable padding tokens, and enhances local features to adeptly manage two-dimensional signal processing. AiM (Li et al., 2024) further advances this by replacing Transformers with Mamba for autoregressive image generation, following methodologies similar to LlamaGen (Sun et al., 2024). However, these existing methods typically apply Mamba to sequences with the length merely up to 256. Our proposed M-VAR model further unleash Mamba’s true capability in capturing super-long-sequence dependencies, empirically up to 2,240 visual tokens. This increase in sequence length can further highlight the benefits brought by Mamba’s efficiency in modeling long sequences within vision applications.
Method
Given a set of corpus , autoregressive modeling predicts next words based on all preceding words:
Autoregressive modeling minimize the negative log-likelihood of each word given all preceding words from to :
This strategy leads to the success of a large language model Touvron et al. (2023); OpenAI (2022; 2023)
Token-wise autoregressive modeling in computer vision.
From language to image, to apply autoregressive pertaining, image tokenization via vector-quantization transfers 2D images into 2D tokens and then flatten these tokens into 1D token sequences :
However, such the flatten operation breaks the inherent 2D structure of an image. To address this issue, VAR (Tian et al., 2024) proposes to perform scale-wise autoregressive modeling to keep the 2D structure, as detailed next.
Scale-wise autoregressive modeling.
Instead of tokenizing image into a sequence of tokens, VAR tokenizes the image into multi-scale token maps , where is the token map with the resolution of downsampling from . Compared to which only contains one token and breaks the 2D structure, contains tokens and is able to maintain the 2D structure. Then, the corresponding autoregressive model is reformed to:
Note that the sequence of multiple scales is much longer than each scale (). VAR by default utilizes attention and Transformer to instantiate this modeling — i.e., to generate the scale, VAR needs to attend all preceding scales, ranging from the first one to the one, and predict all the tokens in parallel inside the scale.
2 Decouple Scale-wise Autoregressive Modeling
We note the attention in VAR can actually be separated into two parts: 1) the first part bidirectionally attends the intra-scale information, where the sequence length is much shorter; 2) the second part unidirectionally attends from the coarse scale to the fine scale, where the sequence length is much longer. We show the statistics of these two parts of attention modeling, including the attention score and the associated computation cost, in Table 1.
Firstly, we note that intra-scale attention scores account for 79.6% of the total attention scores in 256256 image generation and 77.1% in 512512 image generation. This dominance of intra-scale attention suggests that capturing fine-grained details within the same scale is crucial for high-quality image synthesis. More interestingly, we observe that the associated computation cost presents a contrasting scenario — despite intra-scale attention contributing the most to the attention scores, it only consumes 23.9% of the computation cost for 256256 images and 30.3% for 512512 images; In stark contrast, inter-scale attention, which accounts for a much smaller portion of the attention scores (20.4% for 256256 and 22.9% for 512512 images), is responsible for the majority of the computation cost, i.e., 76.1% and 69.7% respectively. The disparity between the attention scores and computation cost highlights an inefficiency in the current attention configuration in VAR.
Based on this observation, we propose a novel approach to optimize the computational efficiency of the scale-wise autoregressive image generation models. Specifically, we retain the standard attention mechanisms for intra-scale interactions, but, importantly, employ Mamba for inter-scale interactions. The main motivation of this change is that Mamba is designed to handle long-range interactions efficiently that scales linearly with the sequence length, as opposed to the quadratic scaling of traditional attention mechanisms. This property makes Mamba particularly suitable for modeling inter-scale relationships, where the computational cost is otherwise prohibitive. We name this new method M-VAR.
Formally, as illustrated in Figure 3, M-VAR models inputs in the following two steps. First, given an image with multiple scales , we apply the standard attention mechanism independently to each scale for capturing the fine-grained details and local dependencies:
Here, represents the standard attention, which produces the intra-scale representation, and is the condition token. All attention share the same parameters but process each scale independently. This design choice ensures consistency across scales and reduces the overall model complexity. For efficient implementation, we adopt FlashAttention (Dao et al., 2022; Dao, 2024) to perform the intra-scale attention in parallel.
After obtaining the intra-scale representations , modeling the relation between different scales becomes crucial for ensuring global coherence and coarse-to-fine consistency in the generated images. However, as previously discussed, traditional attention mechanisms are computationally expensive for inter-scale interactions due to their quadratic complexity, and we adopt the Mamba model with linear complexity, as
By concatenating from all scales into a single sequence, Mamba efficiently processes the combined representations, capturing the essential inter-scale interactions without incurring heavy computational burden.
Experiment
Following the same setup in Tian et al. (2024), we first train M-VAR on ImageNet (Deng et al., 2009) for 256256 conditional generation. We design multiple model variants with depths of 12, 16, 20, 24, and 32 layers.
We compare our M-VAR with previous state-of-the-art generative adversarial nets (GANs), diffusion models, autoregressive models, mask prediction models. As shown in Table 2, Our M-VAR offers a balanced synergy of high image quality and computational efficiency. For example, M-VAR outperforms GANs in terms of image fidelity and diversity while maintaining comparable inference speeds. Compared to diffusion models, M-VAR delivers comparable or even superior image quality with significantly reduced inference time. Against token-wise autoregressive and mask prediction models, our M-VAR achieves better performance metrics with fewer steps and faster inference times.
Compared with the most related VAR, our M-VAR demonstrates significant advancements in both performance and efficiency. Across various depths, M-VAR consistently achieves lower Fréchet Inception Distance (FID) scores and higher Inception Scores (IS), indicating superior image quality and diversity. Specifically, M-VAR-d24 attains an FID of 1.93 and an IS of 320.7 with 1.5 billion parameters, which already surpasses the largest VAR model, VAR-d30, with 25% fewer parameters and 14% faster inference speed. M-VAR also shows compelling scalability — our largest model, M-VAR-d32, achieves state-of-the-art performance with an FID of 1.78 and an IS of 331.2, utilizing 3 billion parameters. These results highlight the effectiveness of our approach in integrating intra-scale attention with Mamba for inter-scale modeling, leading to superior image generation quality and computational efficiency compared to existing models. In additional to these quantitative numbers, we also show more qualitative results in Figure 4.
To further enhance the generation results, as shown in Table 3, we compare our M-VAR-d32 with other state-of-the-art methods using rejection sampling Grover et al. (2018); Azadi et al. (2018) on class-conditional ImageNet 256256. Our M-VAR-d32 achieves an FID of 1.63 and an IS of 361.5, outperforming all compared models. Specifically, it surpasses the prior art VAR-d30 by 0.1 in FID and by 11.3 in IS. Additionally, M-VAR-d32 demonstrates significant improvements over ViTVQ, RQTransformer, and VQGAN by an FID of 1.41, 2.17, 3.57 respectively. These results highlight the effectiveness of our approach in achieving superior image generation quality under rejection sampling, affirming the advancements of M-VAR in image generation.
2 ImageNet 512×512 conditional generation
We hereby train M-VAR on ImageNet (Deng et al., 2009) for 512512 conditional generation. As reported in Table 4, our M-VAR-d24 exhibits competitive performance in class-conditional ImageNet 512512 generation when compared to VAR. Specifically, M-VAR-d24 achieves an FID of 2.65 and an IS of 305.1, closely matching the performance of VAR-d36. More importantly, M-VAR-d24 accomplishes this with half the inference time of VAR-d36, highlighting the efficiency gains from our decoupled intra-scale and inter-scale modeling strategy. Compared to other advanced generative models, such as BigGAN, DiT-XL/2, MaskGIT, and VQGAN, M-VAR-d24 consistently outperforms them in both FID and IS metrics while maintaining a comparable or lower inference cost. We also show more qualitative results in Figure 4, which supports that M-VAR consistently produces images with fine details, enhanced texture fidelity, and improved structural coherence.
3 Ablation Study
We reduce M-VAR’s parameters by adjusting its width or depth, aiming for a fairer comparison while assessing performance advantage and computational efficiency. As reported in Table 5, we present three variants of M-VAR alongside the baseline VAR under similar parameter constraints. Firstly, M-VAR-W reduces the width of the model from 1024 to 768 while keeping the depth constant at 16 layers. This reduction leads to a decrease in the total number of parameters to 260 million, which is lower than VAR’s 310 million parameters. Remarkably, even with fewer parameters, M-VAR-W achieves a better FID score of 3.20 compared to VAR’s 3.55; Additionally, the training cost is reduced to 0.9 times that of VAR, showcasing enhanced efficiency. Similarly, M-VAR-D maintains the original width of 1024 but reduces the depth from 16 to 12 layers. M-VAR-D attains an FID score of 3.19, outperforming VAR while also reducing the training/inference cost to 0.8/0.7 times that of the VAR. These results corroborate that our proposed M-VAR models can achieve superior image generation quality compared to the baseline VAR, even when operating under similar or reduced parameter budgets.
From VAR to MAR.
We gradually replaced the global attention layers in VAR with our proposed intra-scale attention and inter-scale Mamba modules to evaluate their impact on image generation quality. As shown in Figure 5, by progressively increasing the number of layers replaced, from 0 in the original VAR to all 16 layers in our M-VAR, we can observe a consistent improvement in FID scores from 3.55 to 3.07. This phenomenon suggests that decoupling the modeling of intra-scale and inter-scale dependencies positively impacts image synthesis quality — by effectively capturing local spatial details within each scale and efficiently modeling hierarchical relationships between scales, our approach successfully leads to more coherent and detailed image generation.
Effectiveness and efficiency of Attention and Mamba.
Table 6 illustrates the impact of different attention mechanisms on image generation quality, as measured by FID. The baseline VAR employs global attention, capturing both intra-scale and inter-scale dependencies simultaneously, and achieves an FID of 3.55. Then, when exclusively using intra-scale attention (i.e., without inter-scale modeling, Method 1), the FID significantly deteriorates to 7.17, indicating that inter-scale dependencies are crucial for high-quality image generation. Method 2, which additionally introduces token-wise relationship modeling along all scales, improves the FID to 4.12, yet still falls short of the baseline VAR performance. Our proposed M-VAR model combines intra-scale attention with Mamba for efficient inter-scale modeling — by decoupling the two types of dependencies and applying Mamba’s linear-complexity approach for inter-scale interactions, M-VAR achieves the best FID of 3.07. This demonstrates that effectively capturing intra-scale dependencies with attention and efficiently modeling inter-scale relationships with Mamba leads to superior image quality.
Conclusion
This paper develops a novel scale-wise autoregressive image generation method, which decouples intra-scale and inter-scale modeling to enhance both efficiency and performance. Our key idea is that, for inter-scale modeling, we replace the standard attention with Mamba for global-sequence modeling. This strategic separation allows our model to maintain spatial coherence and hierarchical consistency while significantly reducing computational complexity. Our experiments demonstrate that this decoupled framework outperforms existing autoregressive models and diffusion models, achieving superior image quality with fewer parameters and faster inference speeds.