Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation

Sucheng Ren, Qihang Yu, Ju He, Xiaohui Shen, Alan Yuille, Liang-Chieh Chen

Introduction

Autoregressive (AR) models have driven major advances in natural language processing (NLP) through next-token prediction, where each token is generated from its preceding tokens. This framework enables coherent, context-aware text generation, with landmark models like GPT-3 and its successors setting new benchmarks across diverse NLP applications.

Building on the successes of AR modeling in NLP, researchers have extended this framework to computer vision, particularly for high-fidelity image generation . In these approaches, image patches are discretized into tokens and reshaped into 1D sequences, allowing AR models to predict each token sequentially. However, unlike language, where tokens correspond to semantically meaningful units such as words, vision lacks a universally agreed-upon token definition. This naturally raises the question: How can “next-token prediction” be generalized to “next-X prediction,” and what constitutes the most suitable X for image generation?

Additionally, beyond token design, traditional AR models rely on teacher forcing during training, where ground truth tokens are provided at each step instead of the model’s own predictions. While this stabilizes training, it introduces exposure bias , since the model is never exposed to potential errors. Consequently, during inference, without ground truth guidance, errors accumulate over time, leading to cascading errors and context drift as the model conditions solely on its past predictions.

To address these challenges, we propose xAR, a general next-X prediction framework that reformulates discrete token classification (conditioned on all preceding discrete ground truth tokens) into a continuous entity regression problem conditioned on all previous noisy entities. The regression process is guided by flow-matching at each AR step. As illustrated in Fig. 1, within this framework, X serves as a flexible representation that can correspond to an individual patch token, a cell (a group of surrounding tokens), a subsample (a non-local grouping), a scale (coarse-to-fine resolution), or even an entire image.

Unlike teacher forcing , which always provides ground truth inputs, xAR deliberately exposes the model to noisy contexts during training, allowing it to learn from imperfect, corrupted, or partially inaccurate conditions. We refer to this approach as Noisy Context Learning (NCL), a reformulation that reduces reliance on ground truth inputs, improving robustness and mitigating exposure bias by enabling the model to generalize better during inference.

We demonstrate the effectiveness of xAR on the challenging ImageNet generation benchmark . Through systematic experimentation with different X configurations, we find that next-cell prediction—where neighboring tokens are grouped into moderately sized cells (e.g., 8×\times8 tokens)—yields the best performance by capturing richer spatial-semantic relationships. Leveraging both next-cell prediction and Noisy Context Learning, our base model xAR-B (172M) outperforms the large DiT-XL and SiT-XL (675M) while achieving 20×\times faster inference. Additionally, our largest model, xAR-H (1.1B), sets a new state-of-the-art with an FID of 1.24 and runs 2.2×\times faster than the previous best-performing model on ImageNet-256 , without relying on vision foundation models (e.g., DINOv2 ) or extra guidance interval sampling .

Related Work

AR Modeling in NLP. Autoregressive language models have driven significant progress toward general-purpose AI. Their core principle is simple yet powerful: predicting the next token based on preceding context. This approach has demonstrated impressive scalability, guided by scaling laws, and adaptability, enabling zero-shot generalization. These strengths have extended AR modeling beyond traditional language tasks, influencing a wide range of modalities.

AR Modeling in Vision. Inspired by the success of AR modeling in NLP, researchers have explored its application in vision . A pioneering effort in this direction was PixelCNN , which factorized the joint pixel distribution into a product of conditionals, enabling the model to learn complex image distributions. This idea was further refined in PixelRNN , which incorporated recurrent layers to capture richer context in both horizontal and vertical directions. iGPT extended this pixel-level approach by leveraging Transformers for next-pixel prediction. Beyond next-pixel modeling, AR methods have shifted toward more abstract token representations. VQ-VAE introduced discrete latent codes that could be modeled autoregressively, offering a compressed yet expressive representation of images. Later models like Parti and LlamaGen combined these learned tokens with Transformer-based architectures to generate high-fidelity images while maintaining scalable training. Recently, MAR introduced a diffusion-based approach to model per-token probability distributions in a continuous space, replacing categorical cross-entropy with a diffusion loss. VAR extended next-token prediction to a coarse-to-fine scale prediction paradigm, progressively refining image details. Our work unifies these approaches under a general next-X prediction framework, where X can flexibly represent tokens, scales, or our newly introduced cells, providing a more flexible and generalizable formulation for autoregressive visual modeling.

Diffusion and Flow Matching. Beyond autoregressive modeling, diffusion and flow matching have surpassed Generative Adversarial Networks (GANs) by employing multi-step denoising. Latent Diffusion Models (LDMs) improve speed and scalability by operating in a compressed latent space instead of raw pixels. Building on this, DiT and U-ViT replace the traditional convolution-based U-Net with Transformers in latent space, further enhancing performance. Simple Diffusion introduces a streamlined approach for scaling pixel-space diffusion models to high-resolution outputs, while DiMR progressively refines features across multiple scales, improving detail from low to high resolution. In parallel, flow matching reformulates the generative process by directly mapping data distributions to a standard normal distribution, simplifying the transition from noise to structured data. SiT builds on this by integrating flow matching into DiT’s Transformer backbone for more efficient distribution alignment. Extending this approach, SD3 introduces a Transformer-based architecture that leverages flow matching for text-to-image generation. REPA refines denoising by aligning noisy intermediate states with clean image embeddings extracted from pretrained visual encoders .

Method

In this section, we first provide an overview of autoregressive modeling with the next-token prediction paradigm in Sec. 3.1, followed by our proposed xAR framework with next-X prediction and Noisy Context Learning in Sec. 3.2.

Autoregressive modeling with next-token prediction is a fundamental approach in language modeling where the joint probability of a token sequence is factorized into a product of conditional probabilities. Formally, given a sequence x={x1,x2,…,xN}\boldsymbol{x}=\{x_{1},x_{2},\dots,x_{N}\}, the model estimates

In practice, an autoregressive language model predicts the next token xnx_{n} through token classification, conditioned on all preceding tokens {x1,x2,…,xn−1}\{x_{1},x_{2},\dots,x_{n-1}\}. This process proceeds sequentially from left to right (i.e., n={1,…,N}n=\{1,\dots,N\}) until the full sequence is generated. For visual generation, a VQ tokenizer discretizes an image into a sequence of tokens. An autoregressive visual generation model then follows the next-token prediction paradigm, sequentially predicting tokens through classification conditioned on previously generated tokens. However, directly applying the next-token prediction paradigm to visual generation introduces several challenges:

Information Density. In NLP, each token (e.g., a word) carries rich semantic meaning. In contrast, visual tokens typically represent small image patches, which may not be as semantically meaningful in isolation. A single patch can contain fragments of different objects or textures, making it difficult for the model to infer meaningful relationships between consecutive patches. Additionally, the quantization process in VQ-VAE can discard fine details, leading to lower-quality reconstructions. As a result, even if the model predicts the next token correctly, the generated image may still appear blurry or lack detail.

Accumulated Errors. Teacher forcing , a common training strategy, feeds the model ground truth tokens to stabilize learning. However, this reliance on perfect context causes exposure bias —the model never learns to recover from its potential mistakes. During inference, when it must condition on its own predictions, small errors can accumulate over time, leading to compounding artifacts and degraded output quality.

To address these challenges, we extend next-token prediction to next-X prediction, transitioning from traditional AR to xAR. This is accomplished by introducing a more expressive prediction entity X and training the model with noisy entities for improved robustness.

2 The Proposed xAR

We introduce xAR, which consists of two key components: next-X prediction (Sec. 3.2.1) and Noisy Context Learning (Sec. 3.2.2). We first detail each component, then describe the inference strategy (Sec. 3.2.3), followed by a discussion on how xAR enhances visual generation (Sec. 3.2.4).

Given an image, we use an off-the-shelf VAE (instead of VQ-VAE to avoid quantization loss) to convert it into a continuous latent I∈RHf×Wf×CI\in\mathcal{R}^{\frac{H}{f}\times\frac{W}{f}\times C}, where HH and WW denote image height and width, ff is the downsampling rate (we use f=16f=16 ), and CC represents the number of channels. We then construct a sequence of prediction entities X={X1,X2,…,XN}\boldsymbol{X}=\{X_{1},X_{2},\dots,X_{N}\} based on II. Each XiX_{i} is a flexible entity that can represent an individual token (an image patch), a cell (a group of surrounding tokens), a subsample (a non-local grouping), a scale (coarse-to-fine resolution), or even an entire image. We outline common choices for X below and refer readers to Fig. 1 for visualization and Algorithm 1 for a PyTorch pseudo-code implementation.

Individual Patch Token (Fig. 1 (a)). When XiX_{i} corresponds to a single image patch, xAR reduces to standard AR modeling, where each token is predicted sequentially.

Cell (Fig. 1 (b)). The image is divided into an m×mm\times m grid, where each cell has k×kk\times k spatially adjacent tokensWe also experimented with rectangular cells (e.g., cells with shape k/2×2kk/2\times 2k or 2k×k/22k\times k/2), but observed no significant difference compared to squared cells. Thus, we adopt the simpler squared cell design..

Subsample (Fig. 1 (c)). Entities are created by spatially and uniformly subsampling the image grid .

Entire Image (Fig. 1 (d)). As an extreme case, all tokens are grouped into a single entity, i.e., X=X1=IX=X_{1}=I, transforming xAR into a flow matching method .

Default Choice of X. Extensive ablation studies in Sec. 4.2 show that cell (with a size of 8×\times8 tokens) achieves the best performance among all X designs. Therefore, unless specified otherwise, xAR adopts 8×\times8 cells as the default X.

2.2 Noisy Context Learning

xAR transitions the paradigm from “discrete token classification” (conditioned on all preceding ground truth tokens) to “continuous entity regression” (conditioned on all previous noisy entities). Specifically, unlike traditional AR modeling, which directly classifies XnX_{n} based on all preceding ground truth entities {X1,…,Xn−1}\{X_{1},\dots,X_{n-1}\}, xAR predicts XnX_{n} by minimizing a regression loss derived from flow matching , conditioned on all previous noisy entities.

During training, we randomly sample nn noise time steps {t1,…,tn}⊂\{t_{1},\dots,t_{n}\}\subset, and draw nn noise samples {ϵ1,…,ϵn}\{\epsilon_{1},\dots,\epsilon_{n}\} from the source Gaussian noise distribution. Specifically, at the nn-th AR step, the noise samples are drawn as ϵn∼N(0,I)\epsilon_{n}\sim\mathcal{N}(0,I), where ϵn\epsilon_{n} and XnX_{n} share the same shape . We construct the interpolated input FntnF_{n}^{t_{n}} as:

where VntnV_{n}^{t_{n}} represents the directional flow from FntnF_{n}^{t_{n}} toward XnX_{n}, guiding the transformation from the source to the target distribution.

The model is trained to predict the velocity VntnV_{n}^{t_{n}} using all preceding and current noisy entities {F1t1,…,Fntn}\{F_{1}^{t_{1}},\dots,F_{n}^{t_{n}}\}:

We refer to this scheme as Noisy Context Learning (NCL), where the model is trained by conditioning on all previous noisy entities rather than perfect ground truth inputs. This effectively reduces reliance on clean training signals, improving robustness and mitigating exposure bias . Fig. 3 (Training) provides an illustration of NCL. Notably, when sampling the time steps {t1,…,tn}⊂\{t_{1},\dots,t_{n}\}\subset, no constraints are imposed (e.g., we do not enforce t1>t2t_{1}>t_{2}), allowing the model to experience varying degrees of noise in preceding entities, strengthening its adaptability during inference.

2.3 Inference Scheme

xAR performs autoregressive prediction at the level of entity X. Since “cell” is the default choice for X, we use it as a concrete example. As illustrated in Fig. 3 (Inference), xAR begins by predicting an initial cell X^1\hat{X}_{1} from a Gaussian noise sample ϵ1∼N(0,I)\epsilon_{1}\sim\mathcal{N}(0,I) (where ϵ1\epsilon_{1} has the same shape as X^1\hat{X}_{1}) via flow matching . Conditioned on the clean estimate X^1\hat{X}_{1}, xAR generates the next cell X^2\hat{X}_{2} from another Gaussian noise sample ϵ2\epsilon_{2}. This process continues autoregressively, where at the ii-th AR step, the model predicts the next cell X^i\hat{X}_{i} based on all previously generated clean cells {X^1,⋯ ,X^i−1}\{\hat{X}_{1},\cdots,\hat{X}_{i-1}\} and the newly drawn Gaussian noise sample ϵi\epsilon_{i}. This iterative approach progressively refines the image, ensuring structured and context-aware generation at the cell level.

2.4 Discussion

As discussed in Sec. 3.1, traditional AR modeling for visual generation faces two key challenges: information density and accumulated errors. The proposed xAR is designed to address these limitations.

Semantic-Rich Prediction Entity. A cell (i.e., a k×kk\times k grouping of spatially contiguous tokens) aggregates neighboring tokens, effectively capturing both local structures (e.g., edges, textures) and regional contexts (e.g., small objects or parts of larger objects). This leads to richer semantic representations compared to single-token predictions. By modeling relationships within the cell, the model learns to generate coherent local and regional features, shifting from isolated token-level predictions to holistic patterns. Additionally, predicting a cell rather than an individual token allows the model to reason at a higher abstraction level, akin to how NLP models predict words instead of characters. The larger receptive field per prediction step contributes more semantic information, bridging the gap between low-level visual patches and high-level semantics.

Robustness to Previous Prediction Errors. The Noisy Context Learning (NCL) strategy trains the model on noisy entities instead of perfect ground truth inputs, reducing over-reliance on pristine contexts. This alignment between training and inference distributions enhances the model’s ability to handle errors in self-generated predictions. By conditioning on imperfect contexts, xAR learns to tolerate minor inaccuracies, preventing small errors from compounding into cascading errors. Additionally, exposure to noisy inputs encourages smoother representation learning, leading to more stable and consistent generations.

Experimental Results

In this section, we present the main results in Sec. 4.1, followed by ablation studies on key design choices in Sec. 4.2.

We conduct experiments on ImageNet at 256×\times256 and 512×\times512 resolutions. Following prior works , we evaluate model performance using FID , Inception Score (IS) , Precision, and Recall. xAR is trained with the same hyper-parameters as (e.g., 800 training epochs), with model sizes ranging from 172M to 1.1B parameters. See Appendix Sec. A for hyper-parameter details.

ImageNet-256. In Tab. 1, we compare xAR with previous state-of-the-art generative models. Out best variant, xAR-H, achieves a new state-of-the-art-performance of 1.24 FID, outperforming the GAN-based StyleGAN-XL by 1.06 FID, masked-prediction-based MaskBit by 0.28 FID, AR-based RAR by 0.24 FID, VAR by 0.73 FID, MAR by 0.31 FID, and flow-matching-based REPA by 0.18 FID. Notably, xAR does not rely on vision foundation models or guidance interval sampling , both of which were used in REPA , the previous best-performing model. Additionally, our lightweight xAR-B (172M), surpasses DiT-XL (675M) by 0.55 FID while achieving an inference speed of 9.8 images per second—20×\times faster than DiT-XL (0.5 images per second). Detailed speed comparison can be found in Appendix B.

ImageNet-512. In Tab. 2, we report the performance of xAR on ImageNet-512. Similarly, xAR-L sets a new state-of-the-art FID of 1.70, outperforming the diffusion based DiT-XL/2 and DiMR-XL/3R by a large margin of 1.34 and 1.19 FID, respectively. Additionally, xAR-L also surpasses the previous best autoregressive model VAR-d36 and flow-matching-based REPA by 0.93 and 0.38 FID, respectively.

Qualitative Results. Fig. 4 presents samples generated by xAR (trained on ImageNet) at 512×\times512 and 256×\times256 resolutions. These results highlight xAR’s ability to produce high-fidelity images with exceptional visual quality.

2 Ablation Studies

In this section, we conduct ablation studies using xAR-B, trained for 400 epochs to efficiently iterate on model design.

Prediction Entity X. The proposed xAR extends next-token prediction to next-X prediction. In Tab. 3, we evaluate different designs for the prediction entity X, including an individual patch token, a cell (a group of surrounding tokens), a subsample (a non-local grouping), a scale (coarse-to-fine resolution), and an entire image.

Among these variants, cell-based xAR achieves the best performance, with an FID of 2.48, outperforming the token-based xAR by 1.03 FID and surpassing the second best design (scale-based xAR) by 0.42 FID. Furthermore, even when using standard prediction entities such as tokens, subsamples, images, or scales, xAR consistently outperforms existing methods while requiring significantly fewer parameters. These results highlight the efficiency and effectiveness of xAR across diverse prediction entities.

Cell Size. A prediction entity cell is formed by grouping spatially adjacent k×kk\times k tokens, where a larger cell size incorporates more tokens and thus captures a broader context within a single prediction step. For a 256×256256\times 256 input image, the encoded continuous latent representation has a spatial resolution of 16×1616\times 16. Given this, the image can be partitioned into an m×mm\times m grid, where each cell consists of k×kk\times k neighboring tokens. As shown in Tab. 4, we evaluate different cell sizes with k∈{1,2,4,8,16}k\in\{1,2,4,8,16\}, where k=1k=1 represents a single token and k=16k=16 corresponds to the entire image as a single entity. We observe that performance improves as kk increases, peaking at an FID of 2.48 when using cell size 8×88\times 8 (i.e., k=8k=8). Beyond this, performance declines, reaching an FID of 3.13 when the entire image is treated as a single entity. These results suggest that using cells rather than the entire image as the prediction unit allows the model to condition on previously generated context, improving confidence in predictions while maintaining both rich semantics and local details.

Noisy Context Learning. During training, xAR employs Noisy Context Learning (NCL), predicting XnX_{n} by conditioning on all previous noisy entities, unlike Teacher Forcing. The noise intensity of previous entities is contorlled by noise time steps {t1,…,tn−1}⊂\{t_{1},\dots,t_{n-1}\}\subset, where t=0t=0 corresponds to pure Gaussian noise. We analyze the impact of NCL in Tab. 5. When conditioning on all clean entities (i.e., the “clean” variant, where ti=0,∀i<nt_{i}=0,\forall i<n), which is equivalent to vanilla AR (i.e., Teacher Forcing), the suboptimal performance is obtained. We also evaluate two constrained noise schedules: the “increasing noise” variant, where noise time steps increase over AR steps (t1<t2<⋯<tn−1t_{1}<t_{2}<\cdots<t_{n-1}), and the “ decreasing noise” variant, where noise time steps decrease (t1>t2>⋯>tn−1t_{1}>t_{2}>\cdots>t_{n-1}). While both settings improve over the “clean” variant, they remain inferior to our final “random noise” setting, where no constraints are imposed on noise time steps, leading to the best performance.

Conclusion

In this work, we introduced xAR, a general next-X prediction framework for autoregressive visual generation. Unlike traditional next-token prediction, xAR reformulates discrete token classification as continuous entity regression, enabling more flexible and semantically meaningful prediction units. Through systematic exploration, we found that next-cell prediction provides the best balance between local structure and global coherence. To mitigate exposure bias, we proposed Noisy Context Learning (NCL), which trains the model on noisy entities instead of pristine ground truth inputs, improving robustness and reducing cascading errors. As a result, xAR achieves state-of-the-art performance on ImageNet-256 and ImageNet-512.

References

Appendix

The supplementary material includes the following additional information:

Sec. A details the hyper-parameters used for xAR.

Sec. B provides a comprehensive speed comparison.

Sec. C discusses the limitations and future directions.

Sec. D presents visualization samples generated by xAR.

A Hyper-parameters for xAR

We list the detailed training and inference hyper-parameters in Tab. 6.

B Speed Comparison.

We compare xAR with diffusion-, flow matching-, and autoregressive-based models in Tab. 7. Our most lightweight variant, xAR-B (172M), outperforms DiT-XL (diffusion-based), SiT-XL (flow matching-based), and MAR (autoregressive-based), while achieving a 20×\times speedup (9.8 vs. 0.5 images/sec). Additionally, xAR-L surpasses the recent state-of-the-art model REPA, running 5.3×\times faster (3.2 vs. 0.6 images/sec). Finally, our largest model, xAR-H, achieves 1.24 FID on ImageNet-256, setting a new state-of-the-art, while still running 2.2×\times faster than REPA.

C Discussion and Limitations

Our empirical evaluations indicate that a square 8×\times8 cell configuration achieves the best performance, with no noticeable difference when using rectangular cells (e.g., k/2×2kk/2\times 2k or 2k×k/22k\times k/2), which introduce additional complexity without clear benefits. Given that different regions in an image contain varying levels of semantic information (e.g., dense object areas vs. uniform sky regions), future research could explore whether dynamically shaped prediction entities provide additional benefits. However, in this work, we adopt a simple yet effective square cell design, demonstrating state-of-the-art results on the challenging ImageNet generation benchmark.

D Visualization of Generated Samples

Additional visualization results generated by xAR-H are provided from Fig. 5 to Fig. 13.