SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL
Junke Wang, Zhi Tian, Xun Wang, Xinyu Zhang, Weilin Huang, Zuxuan Wu, Yu-Gang Jiang
Introduction
Recent years have seen the rapid advancements of deep generative models , which offer innovative approaches for the creation of visual content. Among these models, diffusion models and autoregressive models have emerged as the leading paradigms.
Diffusion models synthesize visual outputs by iteratively refining random noise through a learned denoising process. This manner has been proven highly effective in producing high-fidelity images and videos , thus gaining remarkable popularity in the field of multimodal generation.
Another line of work is autoregressive (AR) models , which formulate visual generation as a sequential process, i.e., each pixel or token is generated based on the preceding ones. This autoregressive process naturally excels in precise and coherent prediction, making it particularly effective for tasks that demand fine-grained control. Moreover, AR visual generation models are naturally compatible with diverse modalities (e.g., language and audio), which facilitates native multimodal understanding and generation .
Despite these strengths, autoregressive visual generation still underperforms diffusion models currently. Several hypotheses have been proposed to explain this gap. One posits that discrete visual tokenizer imposes a fundamental limit on the quality of autoregressive generation, while another points to the considerably longer length of visual sequences compared to text, and long-range dependency modeling leads to significantly more challenges. To mitigate these issues, various variants have been proposed, such as MAR and VAR . Although these approaches have achieved promising results on academic benchmarks , they compromise the intrinsic pattern of “next-token prediction” in language models, resulting in a trade-off between performance and simplicity.
In this work, we aim to push the frontier of autoregressive visual generation by preserving the simplicity of the autoregressive framework while carefully optimizing its training paradigm through pretraining, supervised fine-tuning (SFT) and GRPO -based reinforcement learning (RL). With this, we show that a vanilla AR model with only 0.5B parameters is capable of generating 10241024 images with superior aesthetics, and achieving competitive results on existing text-to-image benchmarks, e.g., 0.59 on GenEval. When scaled with more computes (i.e., parameters and tokens), our model consistently demonstrates improved generation quality with higher fidelity and more coherent structures. These results highlight the potential of autoregressive visual generation to compete with diffusion models.
In addition, we also investigate the inference acceleration of autoregressive visual generation models using vLLM and speculative sampling . When deployed with vLLM, our model can generate a 10241024 image in approximately 14 seconds, showcasing its potential for real-world applications.
Related Work
Similar to language models , autoregressive visual generation methods formulate image and video generation as a next-token prediction process. These approaches rely on a tokenizer to tokenize visual inputs into discrete tokens, and train an autoregressive transformer to model the sequential dependencies among these tokens using causal attention. Representative work, including DALL-E , Parti , and LlamaGen , exhibit strong instruction-following and high-fidelity image generation capabilities.
2 Unified Models for Multimodal Understanding and Generation
Recent advancements in multimodal learning have led to the development of unified models that seamlessly integrate vision and language for both understanding and generation . Chameleon and Emu3 adopt token-based architecture by handling diverse modalities with a unified autoregressive transformer. Transfusion and Show-o , on the other hand, integrate next-token prediction for text and diffusion processes for images within a single model, enabling effective handling of both discrete and continuous data. These unified approaches demonstrate the potential of leveraging large-scale transformer architectures to enhance the synergy between vision and language, setting the stage for more versatile and generalizable multimodal AI systems.
Method
Our goal is to advance autoregressive visual generation by preserving the simplicity of the AR framework while optimizing its training pipeline and inference efficiency. To this end, we propose SimpleAR, which consists of a pretrained visual tokenizer that discretizes images into compact visual tokens, a text tokenizer, and a decoder-only transformer that autoregressively models the joint distribution of text and image tokens. Note that, unlike diffusion models or previous AR models that require an additional text encoder , our model integrates text encoding and visual generation within a unified transformer architecture, gaining remarkable advantages in both efficiency and multimodal coherence.
After this, we concatenate text and image tokens along the sequence dimension, and input them to a transformer decoder to model their dependency.
2 Three-stage Training
To improve the generation quality and training efficiency, we employ a three-stage learning paradigm: large-scale pretraining on diverse visual datasets to capture generalizable patterns, supervised fine-tuning (SFT) on high-quality data to enhance fidelity and instruction-following, and reinforcement learning (RL) to further refine multimodal alignment and alleviate exposure bias .
Pretraining and SFT: during pretraining and supervised-finetuning (SFT), our model is both trained with language modeling loss:
where the prediction of is conditioned on both text tokens and its preceeding visual tokens .
RL with GRPO: recently, reinforcement learning has gained increasing attention in the post-training of large language models (LLMs) to improve reasoning capability and diffusion models for better alignment with human preferences. Among these approaches, Group Relative Policy Optimization (GRPO) has emerged as a particularly promising technique by offering better training efficiency and stability.
Using the checkpoint after SFT, GRPO initializes a trainable policy model and a frozen reference model . For a given text prompt , it samples a group of outputs from the policy model , and maximizes the following object to optimize :
where and denote clipping hyper-parameter and coefficient controlling the Kullback–Leibler (KL) penalty, and is the advantage computed in one group. We adopt CLIP as the reward and find it surprisingly useful during experiments.
3 Inference
Since token prediction in autoregressive (AR) models must be performed sequentially, inference latency can be a significant bottleneck. However, various optimizations developed in the LLM community for inference accelerate, such as KV cache and paged attention , offer promising solutions to this. In this work, we explore the application of these techniques to accelerate AR-based visual generation, which will be introduced below beriefly.
KV Cache stores previously computed key-value embeddings from the attention layers and, reuse them across autoregressive decoding steps to reduce redundant computation. It is widely used in LLM inference and could decrease the complexity from to .
vLLM Serving leverages optimized memory management and efficient attention mechanisms, such as paged attention, to enable high-throughput and low-latency inference for autoregressive models on modern hardware.
Speculative Jacobi Decoding accelerates autoregressive generation by first sampling multiple candidate token sequences from a draft model and then efficiently verifying them using the target model. We use the pretrained AR model to serve as both draft model and target model for acceleration.
Experiments
Experimental Setup: we adopt the same architecture as Qwen for our transformer model, and intialize it with LLM weights. While for the visual tokenizer, we use Cosmos-Tokenizer , whose codebook size is 64k and downsample ratio is 16.
As mentioned in Sec. 3.2, the training consists of three stages: 1) 512 resolution pretraining, 2) 1024 resolution SFT, and 3) 1024 resolution RL. During pretraining / SFT, the learning rate is set to 1e-4 / 2e-5, and batch size is 256 in total. The RL training is conducted using trl framework, the learning rate is 1e-5, and batch size is 28. We do not use warm up and learning rate decay in all stages. AdamW is employed for optimization. All experiments are conducted on 32 NVIDIA A100 GPUs.
Training data: the pretraining data involves CC3M , CC12M , OpenImages , SAM1B , and Megalith-huggingface (around 43M in total). For SFT, we use JourneyDB , Synthetic-dataset-1M , and 10M internal data. We adopt a simple data filtering strategy for supervised fine-tuning (SFT) data by removing all images whose short edge is smaller than 1024 pixels. We recaption all the images using Qwen2-VL , and randomly choose from long and short prompts during training.
2 Comparing with State-of-the-arts
We present a comprehensive comparison between SimpleAR and existing state-of-the-art visual generation models in Table 1. Remarkably, with only 0.5B parameters, our model achieves superior performance on established text-to-image benchmarks, i.e., 0.59 overall score on GenEval and 79.66 on DPG-Bench , outperforming all comparable-scale methods (those with fewer than 1B parameters) by significant margins. This includes both diffusion-based approaches (e.g., SDv2.1 ) and autoregressive alternatives (e.g., LlamaGen ). Notably, diffusion models also require auxiliary text encoders (e.g., Flan-T5-XL with 3B parameters) that effectively double their parameter footprint, but SimpleAR process both modalities using a unified transformer, enjoying more efficient parameter utilization and native support for conditional generation.
Moreover, scaling SimpleAR to 1.5B yields consistent improvements across benchmarks (+0.04 on GenEval, +1.85 on DPG-Bench), demonstrating predictable scaling behavior analogous to large language models. Although our model still falls behind Infinity on GenEval, we believe this gap is primarily due to the disparity in the volume of training data and can be narrowed with data scaling.
3 Ablation Studies
In this section, we conduct ablation studies to explore the training and inference of autoregressive visual geneation models, including model initialization, position encodings, etc.
Effects of LLM initialization: as shown in Table 3, whether or not using LLM intialization does not have remarkable effect on the DPG-Bench performance. We believe this is due to the text prompts in existing benchmarks is simple and lack sufficient linguistic complexity or reasoning requirements. The results may also indicate that training on text-to-image generation data will lead to drastic forgetting of initial language capabilities. Therefore, incorporating text data during the training phase is essential for maintaining the text understanding and generation capabilities of LLMs.
Effects of Position Encodings: it is also interesting to comapre 1D and 2D rotary position encoding for image generation using autoregressive framework. We compare the pretraining results on DPG-Bench, and the results in Table 3 demonstrate that replacing 1D positional encoding in LLM with 2D will not improve visual generation significantly. However, we believe 2D positional encoding is quite necessary for dynamic resolution generation and video generation.
Effects of RL: for GRPO-based reinforcement learning, we compare two different reward modules: CLIP-ViT-H-14 and HPSv2 . Notably, HPSv2 is a fine-tuned version of CLIP-ViT-H-14, trained on a specially constructed human preference dataset. The GenEval performance is quantitatively compared in Table 4, where we observe that using both reward modules results in performance improvements. Specifically, the CLIP reward achieves a greater performance gain, with a +0.6 increase on GenEval for the 0.5B model. The qualitative results, as shown in Figure 4, further demonstrate that CLIP could enhance the text rendering capability, along with improved perception of quantifiers and spatial descriptions.
We also plot the reward value and GenEval performance during training in Figure 3- 3. As can be seen from the left, the reward values exhibit a gradual increase during training. Interestingly, the GenEval performance demonstrates a positive correlation with the reward progression, suggesting that a simple reward function, i.e., CLIP-ViT-H-14, is effective in providing consistent feedback that aligns well with the desired task performance.
Inference Speedup: the sequential prediction nature of autoregressive (AR) models could result in high inference latency, making it challenging to deploy them in real-time applications. However, recent advancements in LLM optimization techniques, such as KV cache, paged attention, and speculative decoding, offer opportunities for accelerating inference in AR visual generation models. Therefore, we conduct experiments to apply these approaches to speed up the inference of SimpleAR. The throughput is calculated on an Nvidia A100 node, with CFG enabled.
Table 6 demonstrates that using KV cache can effectivelly save 34% inference time, while serving with vLLM can lead to more sigficant inference acceleration, reducing the time to generate a 10241024 image to 13.55 sceconds.
We also try speculative jacobi decoding (SJD), which speculatively decoding multiple tokens in parallel at inference time to reduce the autoregressive generation steps. The results are shown in Table 6, we can see that SJD can lead to around 2 reduction in steps and slightly better performance on DPG. We also compare the sliding-window design proposed by . Although SJD does not practically reduce the testing latency of the autoregressive (AR) model (unable to use KV cache and need to forward the entire sequence each time), it still presents many possibilities for optimizing the AR inference process.
4 Visualizations and Failure Cases
We visualize image generation results of SimpleAR in Figure 4 and 5. It can be observed that our model could not only generate high-fidelity, aesthetically pleasing images but also demonstrate strong instruction-following capabilities.
Several failure cases are also shown in Figure 6. The limited data scale and parameter size constrain SimpleAR to generate complex poses, objects, and text. Additionally, our model may synthesize content that does not adhere to physical laws.
Conclusion and Future Work
This work presented SimpleAR, a vanilla autoregressive framework for visual generation that discards complex architecture modifications. We focus on the optimization of two fundamental components: 1) Training pipeline, through large-scale pretraining, high-quality supervised finetuning, and GRPO training, SimpleAR achieves competitive performance on existing text-to-image generation benchmarks with only 0.5B parameters. 2) Inference efficiency, we explore various inference acceleration techniques, and show SimpleAR could generate a 10241024 image in around 14 seconds when served with vLLM. We hope this work can inspire further exploration into autoregressive visual generation and firmly believe that it is a promising alternative to diffusion models.
Despite the superior results achieved, we admit that there still exist many limitations that are worth deeper improvements or exploring:
Stronger visual tokenizers: the reconstruction performance of Cosmos-Tokenizer is limited, especially in capturing fine-grained visual details, e.g., faces and texts. This leaves room for better visual generation results with improved tokenization methods.
Text-to-video generation: compared to text-to-image generation, text-to-video generation presents significantly more challenges since the model has to generate coherent outputs that are contextually and temporally consistent.
Native multimodal understanding and generation: recently, the native multimodal understanding and generation capabilities of GPT-4o have captured widespread attention. It offers a glimpse into the future of models that can seamlessly integrate vision, text, and even audio without relying on modality-specific encodings. Moving SimpleAR forward, building truly native large multimodal models that can perform end-to-end reasoning across images, text, and other modalities is a promising and crucial research direction.