FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching

Sucheng Ren, Qihang Yu, Ju He, Xiaohui Shen, Alan Yuille, Liang-Chieh Chen

Introduction

Autoregressive (AR) models have significantly advanced natural language processing (NLP) by modeling the probability distribution of each token given its preceding tokens, allowing for coherent and contextually relevant text generation. Prominent models like GPT-3 and its successors have demonstrated remarkable language understanding and generation capabilities, setting new standards across diverse NLP applications.

Building on the success of autoregressive modeling in NLP, this paradigm has been adapted to computer vision, particularly for generating high-fidelity images through sequential content prediction . In these approaches, images are discretely tokenized, with tokens flattened into 1D sequences, enabling autoregressive models to generate images token-by-token. This approach leverages the sequence modeling strengths of AR architectures to capture intricate visual details. However, directly applying 1D token-wise autoregressive methods to images presents notable challenges. Images are inherently two-dimensional (2D), with spatial dependencies across height and width. Flattening image tokens disrupts this 2D structure, potentially compromising spatial coherence and causing artifacts in generated images. Recently, VAR addresses these issues by introducing scale-wise autoregressive modeling, which progressively generates images from coarse to fine scales, preserving the spatial hierarchies and dependencies essential for visual coherence. This scale-wise approach allows autoregressive models to retain the 2D structure during image generation, capturing the layered complexity of visual content more naturally.

Despite its effectiveness, VAR faces two significant limitations: (1) a complex and rigid scale design, and (2) a dependency between the generator and a tokenizer that shares this intricate scale structure. Specifically, VAR employs a non-uniform scale sequence, {1,2,3,4,5,6,8,10,13,16}\{1,2,3,4,5,6,8,10,13,16\}, where the coarsest scale tokenizes a 256×256256\times 256 image into a single 1×11\times 1 token and the finest scale into 16×1616\times 16 tokens. This intricate sequence constrains both the tokenizer and generator to operate exclusively at these predefined scales, limiting adaptability to other resolutions or granularities. Consequently, the model struggles to represent or generate features that fall outside this fixed scale sequence. Additionally, the tight coupling between VAR’s generator and tokenizer restricts flexibility in independently updating the tokenizer, as both components must adhere to the same scale structure.

To address these limitations, we introduce FlowAR, a flexible and generalized approach to scale-wise autoregressive modeling for image generation, enhanced with flow matching . Unlike VAR , which relies on a complex multi-scale VQGAN discrete tokenizer , we utilize any off-the-shelf VAE continuous tokenizer with a simplified scale design, where each subsequent scale is simply double the previous one (e.g., {1,2,4,8,16}\{1,2,4,8,16\}), and coarse scale tokens are obtained by directly downsampling the finest scale tokens (i.e., the largest resolution token map). This streamlined design eliminates the need for a specially designed tokenizer and decouples the tokenizer from the generator, allowing greater flexibility to update the tokenizer with any modern VAE .

To further enhance image quality, we incorporate the flow matching model to learn the probability distribution at each scale. Specifically, given the class token and tokens from previous scales, we use an autoregressive Transformer to generate continuous semantics that condition the flow matching model, progressively transforming noise into the target latent representation for the current scale. Conditioning is achieved through the proposed spatially adaptive layer normalization (Spatial-adaLN), which adaptively adjusts layer normalization on a position-by-position basis, capturing fine-grained details and improving the model’s ability to generate high-fidelity images. This process is repeated across scales, capturing the hierarchical dependencies inherent in natural images. The final image is then produced by de-tokenizing the predicted latent representation at the finest scale.

The seamless integration of the scale-wise autoregressive Transformer and scale-wise flow matching model enables FlowAR to capture both the sequential and probabilistic aspects of images with multi-scale information, resulting in improved image synthesis performance. We demonstrate FlowAR’s effectiveness on the challenging ImageNet-256 benchmark , where it achieves state-of-the-art results.

Related Work

Autoregressive Models. Autoregressive modeling began in natural language processing, where language Transformers are pretrained to predict the next word in a sequence. This concept was first introduced to computer vision by PixelCNN , which utilized a CNN-based model to predict raw pixel probabilities. With the advent of Transformers, iGPT extended this approach by modeling raw pixels using Transformer architectures. VQGAN further advanced the field by applying autoregressive learning within the latent space of VQ-VAE , thereby simplifying data representation for more efficient modeling . Taking a different direction, Parti framed image generation as a sequence-to-sequence task akin to machine translation, using sequences of image tokens as targets instead of text tokens, and leveraging significant advancements in large language models through data and model scaling. LlamaGen expanded on this by applying the traditional “next token prediction” paradigm of large language models to visual generation, demonstrating that standard autoregressive models like Llama can achieve state-of-the-art image generation performance when appropriately scaled, even without specific inductive biases for visual signals. The work most similar to ours is VAR , which transitioned from token-wise to scale-wise autoregressive modeling by developing a coarse-to-fine next scale prediction. However, VAR faces significant challenges due to its complex scale designs and deep dependency on a scale residual discrete tokenizer. In contrast, our proposed FlowAR employs a simple scale design and maintains compatibility with any VAE tokenizer .

Flow- and Diffusion-based Models. Diffusion models have surpassed earlier image generation methods like GANs by utilizing multistep diffusion and denoising processes. Latent Diffusion Models (LDMs) advance this approach by transitioning diffusion from pixel space to latent space, enhancing efficiency and scalability. Building on this foundation, DiT and U-ViT replace the traditional convolution-based U-Net with Transformer architectures within the latent space, further improving performance. Flow matching redefines the forward process as direct paths between the data distribution and a standard normal distribution, offering a more straightforward transition from noise to target data compared to conventional diffusion methods. SiT leverages the backbone of DiT and employs flow matching to more directly connect these distributions. Scaling this concept, SD3 introduces a novel Transformer-based architecture trained with flow matching for text-to-image generation. MAR presents a diffusion-based strategy to model per-token probability distributions in a continuous space, enabling autoregressive models without relying on discrete tokenizers and utilizing a specialized diffusion loss function instead of the traditional categorical cross-entropy loss. In contrast to MAR, our proposed FlowAR employs a scale-wise flow matching model to capture per-scale probabilities, utilizing coarse-to-fine scale-wise conditioning derived from a scale-wise autoregressive model.

Discussion. Our FlowAR provides a more flexible framework for next scale prediction, enhanced with flow matching. In Figure 2, we compare our FlowAR model with VAR , highlighting key differences in both the tokenizer and generator components. For the tokenizer, VAR relies on a multi-scale VQGAN that is tightly integrated with its generator and trained on a complex set of scales ({1, 2, 3, 4, 5, 6, 8, 10, 13, 16}). In contrast, FlowAR can use any off-the-shelf VAE as its tokenizer, offering greater flexibility by constructing coarse scale token maps through direct downsampling of the finest scale token map. For the generator, VAR is constrained by a complex and rigid scale structure, whereas FlowAR benefits from a simpler and more general scale design, allowing the integration of the modern flow matching model .

Method

In this section, we begin with an overview of autoregressive modeling in Sec. 3.1, followed by a detailed introduction of the proposed method in Sec. 3.2.

Autoregressive Modeling in NLP. Consider a corpus represented as a sequence of words U={w1,w2,…,wn}\mathcal{U}=\{w_{1},w_{2},\ldots,w_{n}\}. In NLP, autoregressive models predict each word based on all preceding words in the sequence:

where the autoregressive model is parameterized by Θ\Theta. The objective function to minimize is the negative log-likelihood over the entire corpus:

where <k{<k} denotes all positions preceding kk. This approach serves as the foundation for successful large-scale language models, as demonstrated in .

However, flattening the token grid disrupts the intrinsic two-dimensional spatial structure of the image. To preserve this spatial information, VAR introduces a scale-wise autoregressive modeling approach, as described below.

Scale-wise Autoregressive Modeling for Image Generation. Rather than flattening images into token sequences, VAR decomposes the image into a series of token maps across multiple scales, S={s1,s2,…,sn}S=\{s_{1},s_{2},\ldots,s_{n}\}. Each token map sks_{k} has dimensions hk×wkh_{k}\times w_{k} and is obtained by a specially designed multi-scale VQGAN with residual structure . In contrast to single flattened tokens tkt_{k} that lose spatial context, each sks_{k} maintains the two-dimensional structure with hk×wkh_{k}\times w_{k} tokens. The autoregressive loss function is then reformulated to predict each scale based on all preceding scales:

In this framework, generating the kk-th scale in VAR requires attending to all previous scales s<ks_{<k} (i.e., s1s_{1} to sk−1s_{k-1}) and simultaneously predicting all hk×wkh_{k}\times w_{k} tokens in sks_{k} via categorical distributions. The chosen scales, S={1,2,3,4,5,6,8,10,13,16}S=\{1,2,3,4,5,6,8,10,13,16\}, introduce significant complexity to the scale design and constrain the model’s generalization capabilities. This is due to the tight coupling between the generator and tokenizer with the scale design, reducing flexibility in updating the tokenizer or supporting alternative scale configurations. Furthermore, VAR’s discrete tokenizer relies on a complex multi-scale residual structure, complicating the training process, with essential training code and details remaining publicly unavailable at the time of our submission.

2 Proposed Method: FlowAR

Overview. To address the issues outlined above, we introduce FlowAR seen in Figure 3, a more general scale-wise autoregressive modeling, enhanced with flow matching . Our method incorporates two primary improvements over existing next scale prediction : (1) replacing the multi-scale VQGAN discrete tokenizer with any off-the-shelf VAE continuous tokenizer , and (2) modeling the per-scale prediction (i.e., predicting all hk×wkh_{k}\times w_{k} tokens in kk-th scale sks_{k}) using flow matching to learn the probability distribution. The first change enables the flexibility to leverage any existing VAE tokenizer, benefiting from recent advances in VAE technology without being constrained by a complex scale sequence design. The second improvement enhances generation quality by utilizing the modern flow matching algorithm.

where Down(F,r)\text{Down}(F,r) denotes downsampling the latent FF by a factor of rr, and no downsampling is applied when r=1r=1.

With this simplified scale sequence and the flexibility to use any off-the-shelf VAE tokenizer, we introduce our improvement for next scale prediction. Unlike VAR , which models the categorical distribution of each scale using an autoregressive Transformer, we utilize the Transformer to generate conditioning information for each scale, while a scale-wise flow matching model captures the scale’s probability distribution based on this information. Below, we outline how the autoregressive Transformer generates conditioning information for each scale, followed by details of the scale-wise flow matching module.

Generating Conditioning Information via Scale-wise Autoregressive Transformer. To produce the conditioning information for each subsequent scale, we utilize conditions obtained from all previous scales:

where T(⋅)T(\cdot) represents the autoregressive Transformer model, CC is the class condition, and Up(s,r)\text{Up}(s,r) denotes the upsampling of latent ss by a factor of rr. For the initial scale (i=1i=1), only the class condition CC is used as input. We set r=2r=2, following our simple scale design, where each scale is double the size of the previous one. We refer to the resulting output s^i\hat{s}^{i} as the semantics for the ii-th scale, which is then used to condition a flow matching module to learn the per-scale probability distribution.

Scale-wise Flow Matching Model Conditioned by Autoregressive Transformer Output. Flow matching generates samples from the target data distribution by gradually transforming a source noisy distribution, such as a Gaussian. For each ii-th scale, FlowAR extends flow matching to generate the scale latent sis^{i}, conditioned on the autoregressive Transformer’s output s^i\hat{s}^{i}. Specifically, during training, given the scale latent sis^{i} from the target data distribution, we sample a time step t∈t\in and a source noise sample F0iF^{i}_{0} from the source noisy distribution, typically setting F0i∼N(0,1)F^{i}_{0}\sim\mathcal{N}(0,1) to match the shape of conditioned latent s^i\hat{s}^{i}, analogous to the “noise” in diffusion models . We then construct the interpolated input FtiF^{i}_{t} as:

The model is trained to predict the velocity VtiV^{i}_{t} using FtiF^{i}_{t}:

where VtiV^{i}_{t} indicates the direction to move from FtiF^{i}_{t} toward sis^{i}, guiding the transformation from the source to the target distribution at each scale. Unlike prior approaches that condition velocity prediction on class or textual information, we condition on the scale-wise semantics s^i\hat{s}^{i} from the autoregressive Transformer’s output. Notably, in prior methods , the conditions and image latents often have different lengths, whereas FlowAR shares the same length (i.e., sis^{i} and s^i\hat{s}^{i} have the same shape, ∀i=1,⋯ ,n\forall i=1,\cdots,n). The training objective for scale-wise flow matching is:

Scale-wise Injection of Semantics via Spatial-adaLN. A key design choice is determining how best to inject the semantic information s^i\hat{s}^{i}, generated by the autoregressive Transformer, into the flow matching module. A straightforward approach would be to concatenate the semantics s^i\hat{s}^{i} with the flattened input FtiF_{t}^{i}, similar to in-context conditioning where the class condition is concatenated with the input sequence. However, this approach has two main drawbacks: (1) it increases the sequence length input to the flow matching model, raising computational costs, and (2) it provides indirect semantic injection, potentially weakening the effectiveness of semantic guidance. To address these issues, we propose using spatially adaptive layernorm for position-by-position semantic injection, resulting in the proposed Spatial-adaLN. Specifically, given the semantics s^i\hat{s}^{i} from the scale-wise autoregressive Transformer and the intermediate feature Fti′F^{i^{\prime}}_{t} in the flow matching model, we inject the semantics to the scale γ\gamma, shift β\beta, and gate α\alpha parameters of the adaptive normalization, following the standard adaptive normalization procedure :

Inference Pipeline. At the beginning of inference, the autoregressive Transformer generates the initial semantics s^1\hat{s}^{1} using only the class condition CC. This semantics s^1\hat{s}^{1} conditions the flow matching module, which gradually transforms a noise sample into the target distribution for s1s^{1}. The resulting token map is upsampled by a factor of 2, combined with the class condition, and fed back into the autoregressive Transformer to generate the semantics s^2\hat{s}^{2}, which conditions the flow matching module for the next scale. This process is iterated nn scales until the final token map sns^{n} is estimated and subsequently decoded by the VAE decoder to produce the generated image. Notably, we use the KV cache in the autoregressive Transformer to efficiently generate each semantics s^i\hat{s}^{i}.

Experimental Results

In this section, we present our main results on the challenging ImageNet-256 generation benchmark (Sec. 4.1), followed by ablation studies (Sec. 4.2).

Following the settings in VAR , we train FlowAR on ImageNet-256 for class-conditional image generation. We evaluate the model using Fréchet Inception Distance (FID) and inception score (IS) , Precision (Pre.) and Recall (rec.) as metrics.

Quantitative Results. As shown in Table 1, when compared to previous generative adversarial models , autoregressive models , diffusion-based methods , and flow matching methods , FlowAR achieves significant performance gains. Specifically, our best model variant, FlowAR-H, attains an FID of 1.65, outperforming StyleGAN (2.30), LlamaGen-3B (2.18), DiT (2.27), and SiT (2.06).

Compared to the closely related VAR , FlowAR provides superior image quality at similar model scales. For example, FlowAR-L, with 589M parameters, achieves an FID of 1.90—surpassing both VAR-d20 (FID 2.95) of comparable size and even largest VAR-d30 (FID 1.97), which has 2B parameters. Furthermore, our largest model, FlowAR-H (1.9B parameters, FID 1.65), sets a new state-of-the-art benchmark for scale-wise autoregressive image generation.

We visualize samples generated by FlowAR using different tokenizers in Figure 4, showing that FlowAR is capable of producing high-quality images with impressive visual fidelity and is compatible with various off-the-shelf VAEs. More samples are provided in the appendix.

2 Ablation Studies

Tokenizer Compatibility. VAR relies on a complex multi-scale residual tokenizer that compresses images into discrete tokens at different scales, with the scale structure of the tokenizer directly mirroring VAR’s architectural scales. This tight coupling between the tokenizer and VAR limits the framework’s flexibility and adaptability. In contrast, our proposed FlowAR is compatible with a wide range of variational autoencoders (VAEs), enhancing versatility and ease of integration. As shown in Table 2, FlowAR achieves superior performance across various VAE architectures, including DC-AE (FID of 4.22), SD-VAE (FID of 3.94), and MAR-VAE (FID of 3.61), compared to VAR’s multi-scale residual discrete tokenizer, which yields an FID of 5.81. These results underscore FlowAR ’s effectiveness and adaptability, highlighting its advantage over VAR’s more rigid tokenizer dependency.

Construction of Scale Sequence. Instead of using a multi-scale VQGAN with residual connections, as in VAR , we propose a simpler approach by directly downsampling the latent representations extracted by any off-the-shelf continuous VAE tokenizer. An alternative design choice would be to downsample the image before feeding it into a VAE. We explore this design choice in Table 3, where downsampling the latents (FID 3.61) significantly outperforms downsampling the image (FID 12.19).

Scale Configurations. To demonstrate FlowAR’s flexibility with respect to scale design, we perform an ablation study by progressively reducing the number of scales used in the model. Table 4 presents the results of this study. VAR relies on a complex scale configuration with ten scales ({1, 2, 3, 4, 5, 6, 8, 10, 13, 16}) to achieve its reported performance. Simplifying VAR’s scale configuration to {1, 2, 4, 8, 16} leads to training failure, indicating a strong dependency between its tokenizer and generator. In contrast, FlowAR demonstrates strong robustness to scale reduction. With the simplified sequence {1, 2, 4, 8, 16}, FlowAR achieves an FID of 3.61, outperforming VAR even with its full scale sequence. Reducing the scales further to {1, 4, 8, 16}, FlowAR still maintains competitive performance with an FID of 4.88. Even with just three scales ({1, 4, 16}), FlowAR achieves an FID of 6.10, comparable to VAR’s FID of 5.81.

Scale-wise Flow Matching Model. The flow matching model is used to learn the per-scale probability distribution, predicting all hk×wkh_{k}\times w_{k} tokens in the kk-th scale sks_{k}. We consider two design alternatives. First, per-scale prediction could be replaced with per-token prediction using Multi-Layer Perceptrons (MLPs) , which, however, lacks the ability to capture interactions between tokens. Second, we could substitute the flow matching approach with a diffusion framework . These design choices are explored in Table 5, where per-scale prediction consistently outperforms per-token prediction, regardless of whether flow matching or diffusion is used. Additionally, flow matching provides marginal improvements over the diffusion framework. Our final model configuration employs per-scale prediction with the flow matching framework.

Injection of Semantics. There are several methods to inject the semantics s^i\hat{s}^{i}, generated by the autoregressive Transformer, into the flow matching module. We summarize these methods in Table 6: (1) ‘addition’: Element-wise addition of the semantics and the flattened input. (2) ‘cross attention’: Using cross-attention where the flattened input serves as the query and the semantics act as the key and value. (3) ‘sequence concatenation’: Concatenating the semantics with the flattened input along the sequence dimension. (4) ‘channel concatenation’: Concatenating the semantics with the flattened input along the channel dimension. (5) ‘adaLN’: Adaptive LayerNorm conditioned on the spatially averaged semantics. (6) ‘Spatial-adaLN’: The proposed spatial adaptive LayerNorm, injecting semantics position-by-position.

As shown in Table 6, the choice of semantic injection method significantly impacts performance. The proposed ‘Spatial-adaLN’ achieves the best results, with an FID of 3.61 and an Inception Score (IS) of 234.1, outperforming all other methods. These results indicate that methods preserving spatial structures and offering position-wise semantic guidance yield superior image generation quality. The exceptional performance of Spatial-adaLN can be attributed to its ability to inject semantics directly into the normalization layers in a spatially adaptive manner, effectively capturing fine-grained details and enhancing the model’s capacity to generate high-fidelity images.

Conclusion

In this work, we presented FlowAR, a flexible and generalized approach to scale-wise autoregressive modeling for image generation, enhanced with flow matching for improved quality. By adopting a streamlined scale design and compatible with any VAE tokenizer, FlowAR addresses limitations of prior models, offering greater adaptability and superior image quality. With spatially adaptive layer normalization, it effectively captures fine-grained details, achieving state-of-the-art results on ImageNet-256 generation benchmark. We hope FlowAR will inspire more future research in autoregressive image modeling.

References

Appendix

The appendix includes the following additional information:

Sec. A lists the hyper-parameters of FlowAR.

Sec. B provides the architectural details of FlowAR model variants.

Sec. C provides more visualization results.

A Hyper-parameters

We list the hyper-parameters of our FlowAR in Table 7.

B Model Variants

In Table 8, we provide four kinds of different configurations of FlowAR for a fair comparison under similar parameters with VAR . The proposed FlowAR contains two main modules: Autoregressive Model and Flow Matching Model, both build on top of Transformer architectures .

C More Visualization Results

Additional visualization results generated by FlowAR-H are provided from Figure 5 to Figure 12.