WorldVLA: Towards Autoregressive Action World Model

Jun Cen, Chaohui Yu, Hangjie Yuan, Yuming Jiang, Siteng Huang, Jiayan Guo, Xin Li, Yibing Song, Hao Luo, Fan Wang, Deli Zhao, Hao Chen

Introduction

The development of Vision-Language-Action (VLA) models has emerged as a significant focus within robotics action model research (Brohan et al., 2023; Kim et al., 2024; Black et al., 2024). These models are constructed by augmenting large-scale pre-trained Multimodal Large Language Models (MLLMs) (Liu et al., 2023b; Li et al., 2024; Zhang et al., 2025; Bai et al., 2025) with with either an action head or additional action expert module to generate actions. MLLMs contribute robust capabilities in perception and decision making, enabling VLA models to exhibit enhanced generalization across a wide range of robotic tasks (Black et al., 2024; Intelligence et al., 2025). Nevertheless, a notable limitation persists: these models often lack a comprehensive understanding of actions, as actions are treated solely as outputs but not being integrated as inputs for deeper analysis. In contrast, world models demonstrate the ability to predict future visual states based on current observations and actions, thereby achieving a dual understanding of both visual information and behavioral dynamics (Ha and Schmidhuber, 2018; Agarwal et al., 2025; Wu et al., 2025). Despite this advantage, world models are constrained by their inability to directly generate action outputs, resulting in a functional gap that limits their application in scenarios requiring explicit action planning.

To address the constraints inherent in both Vision-Language-Action (VLA) models and world models, we introduce WorldVLA, an autoregressive action world model for unified action and image understanding and generation. As depicted in Fig. 1, WorldVLA employs three separate tokenizers to encode images, text, and actions. The tokens from different modalities are set to share the same vocabulary so that understanding and generation across these modalities can be unified within a single LLM architecture. The world model component captures the underlying physical dynamics of the environment by generating visual representations based on input actions. This process of action interpretation and environmental physics learning is essential for enabling effective decision making within the action model. Concurrently, the action model embedded within WorldVLA refines the understanding of visual data, thereby improving the precision of image generation performed by the world model. This bidirectional enhancement creates a more robust and comprehensive model capable of both understanding and generating actions and images.

Action chunking and parallel decoding have been demonstrated to significantly influence the performance of action models (Kim et al., 2025). However, we find that generating multiple actions in sequence leads to performance drop in autoregressive models. The primary reason for this is that pretrained multimodal language models have predominantly been exposed to images and text rather than actions, resulting in limited action generalization capabilities. In autoregressive models where subsequent actions are conditioned on preceding ones, error propagation becomes a critical issue, as the earlier incorrect predictions influence subsequent actions over time. To alleviate this issue, we propose an action attention masking strategy that selectively masks prior actions during the generation of current actions. This approach effectively mitigates error accumulation and yields substantial improvements in the task of action chunk generation.

The experiments on LIBERO benchmark show that our WorldVLA outperforms the action model with the same backbone by 4% grasping success rate. Further, compared to vanilla world model, our WorldVLA shows superior video generation capability and reduces Fréchet Video Distance (FVD) on LIBERO dataset by 10%. These results underscore the mutual benefits derived from integrating world and action models, highlighting the advantages of a unified framework for image and action comprehension and generation. In the context of action chunk generation, the grasping success rate decreseas by 10% to 50% when employing a conventional autoregressive approach. However, the implementation of our attention masking strategy significantly mitigates this decrease, yielding a 4% to 23% improvement in grasping success rate.

In summary, our contributions are as follows:

We propose WorldVLA, an autoregressive action world model that unifies action and image understanding and generation.

We introduce an action attention masking strategy for the action chunk generation task in autoregressive models, addressing the challenge of action error accumulation when generating multiple actions in sequence.

Our experiments demonstrate that WorldVLA outperforms the standalone action and world models, highlighting the mutual enhancement between the world model and action model. Additionally, the action attention masking strategy solves the performance degradation when generating action chunks and significantly improves grasping performance.

Related Works

Our proposed WorldVLA is related to the action model, video prediction model, and world model. The difference between them are summaried in Table LABEL:tab:diff_models.

Behavior cloning (Pomerleau, 1988) is a classic imitation learning approach for robot manipulation, which learns a policy by mimicking expert observation-action pairs. Conventional architectures typically combine a vision backbone, such as ResNet (He et al., 2016) or Vision Transformer (Dosovitskiy et al., 2020), with an action head. The action head may consist of multilayer perceptrons (MLPs) (Rumelhart et al., 1986), query-based transformer decoders (Zhao et al., 2023), or diffusion-based policy heads (Chi et al., 2023). Recently, Vision-Language-Action (VLA) models have been proposed, utilizing pre-trained multi-modality large language models (MLLM) as the backbone (Brohan et al., 2022, 2023; Li et al., 2023; Huang et al., 2023; Belkhale and Sadigh, 2024; Wen et al., 2025; Zhen et al., 2024). These frameworks are equiped with either discrete action decoders (Kim et al., 2024; Pertsch et al., 2025) or continuous diffusion policy heads (Black et al., 2024; Wen et al., 2024) to predict actions. The internet-scale prior knowledge in MLLM enables effective generalization to unseen scenarios tasks for VLA models. Our proposed WorldVLA advances this paradigm by jointly generating actions and predicting future video frames, providing a comprehensive solution for understanding and generation.

Video Generation.

Video generation plays a dual role in robotics. On one hand, some policy models generate the future video first and then generate the corresponding actions based on the generated video (Du et al., 2023; Ajay et al., 2023; Bu et al., 2024). Large-scale video data could be used for pre-training the future video generation part, as seen in approaches (Wu et al., 2023; Cheang et al., 2024). Here, video generation serves as a mechanism for visual imagination and planning, providing valuable insights that improve downstream policy generation (Cen et al., 2024). On the other hand, video generation models can act as world models, simulating diverse future scenarios (Ha and Schmidhuber, 2018). Such world models are widely utilized to generate varied training data (Agarwal et al., 2025), support model-based reinforcement learning algorithms (Wu et al., 2025), and aid in selecting the most suitable policies from a pool of generated options (Li et al., 2025; Bar et al., 2024). In this work, we show that our WorldVLA enables precise control over video generation through action inputs, while also demonstrating that video generation significantly enhances the quality of action generation.

Unified Understanding and Generation Model.

Most multi-modality large language models (MLLMs) are designed to perform visual understanding tasks, where the model generates textual responses based on combined image and language inputs (Liu et al., 2023b; Li et al., 2024; Zhang et al., 2025; Bai et al., 2025). Recently, there has been a growing interest in unifying visual understanding and visual generation within a single framework (Team, 2024; Zhou et al., 2024). One line of work tokenizes images into discrete tokens akin to text, enabling large language models (LLMs) to both interpret and generate visual content seamlessly (Team, 2024; Wang et al., 2024). Another approach integrates diffusion processes into LLMs for image generation while relying on additional visual encoders, such as CLIP (Radford et al., 2021; Zhai et al., 2023), for image understanding (Chen et al., 2025; Tong et al., 2024). In the robotics domain, the Unified Video Action Model (Li et al., 2025) proposes a unified architecture that generates images and actions through distinct diffusion heads. In contrast, our WorldVLA explores an alternative direction by employing a discrete autoregressive architecture to build a unified model capable of handling both perception and action generation.

Methods

In this work, we address the challenge of learning a unified model capable of simultaneously performing action prediction and world state forecasting. Specifically, we define two primary components: an action model (or policy model) πθ\pi_{\theta} and a world model fϕf_{\phi}. The action model πθ\pi_{\theta} is responsible for generating an action ata_{t} conditioned on a history of image observations {ot−h,ot−h+1,…,ot}\{o_{t-h},o_{t-h+1},\dots,o_{t}\} and a language instruction ll, which can be formally expressed as:

Meanwhile, the world model fϕf_{\phi} predicts the next frame oto_{t} based on the historical sequence of observations {ot−h,ot−h+1,…,ot−1}\{o_{t-h},o_{t-h+1},\dots,o_{t-1}\} and the corresponding sequence of actions {at−h,at−h+1,…,at−1}\{a_{t-h},a_{t-h+1},\dots,a_{t-1}\}. This relationship is formulated as:

Our objective is to develop an integrated action-world model MψM_{\psi} that unifies these two functionalities. The model MψM_{\psi} should be capable of both predicting actions as a policy model and forecasting future states as a world model. Formally, the unified model MψM_{\psi} is defined as:

where MψpolicyM_{\psi}^{\text{policy}} represents the action generation component and MψworldM_{\psi}^{\text{world}} denotes the world state prediction component. By learning such a unified model, we aim to achieve a compact and efficient framework that leverages shared representations for both decision-making and environment modeling.

2 Architecture

The overall architecuture of autoregressive action world model is shown in Fig. 2. We initilize the model from Chameleon (Team, 2024) since it is a unified model for image understanding and generation. Three tokenizers are involved, including an image tokenizer, a text tokenizer, and an action tokenizer. The image tokenizer is a VQ-GAN model (Esser et al., 2021) with additional perceptual losses to specific image regions, e.g., faces and salient objects (Gafni et al., 2022). The compression ratio of the image tokenizer is 16 and the codebook size is 8192. The image tokenizer generates 256 tokens for 256×256256\times 256 images and 1024 tokens for 512×512512\times 512 images. The action tokenizer discretizes each dimension of continuous robot actions into one of 256 bins, with bin widths determined by the range of the training data (Kim et al., 2024; Brohan et al., 2023). The actions are represented as 7 tokens, including 3 relative positions, 3 relative angles and 1 absolute gripper states. The text tokenizer is a trained BPE tokenizer (Sennrich et al., 2015) with a vocabulary size of 65,536, which includes 8192 image tokens and 256 action tokens. All texts, actions, and images are discretized into tokens and are trained under the autoregressive manner.

3 Training Strategy

We mix the action model data and world model data to train our WorldVLA. There are three primary reasons for incorporating world model data to enhance action generation. First, the world model acquires an understanding of the environmental physics by learning to predict future observations based on the current state and applied actions. This learned representation of environmental physics is helpful for manipulation tasks. Second, the world model enables the system to simulate and evaluate potential outcomes of candidate actions, thereby facilitating the avoidance of actions that may lead to unfavorable states. Third, the world model requires a precise interpretation of the action inputs, which in turn supports the action model in producing more effective and contextually appropriate actions. On the other hand, action model enhances visual understanding and in turn supports the visual generation capability of the world model.

Action model is to generate the action given the text instruction and image observations. The text inputs are "What action should the robot take to + task instruction + ?". The overall token sequence is:

[BOS]{text}[BOI]{image}…{image}[EOI]⏟×M\underbrace{\texttt{[BOI]\{image\}\ldots\{image\}[EOI]}}_{\times M}[EOS][BOA]{action}…{action}[EOA]⏟×K[EOS]⏞Laction\overbrace{\underbrace{\texttt{[BOA]\{action\}\ldots\{action\}[EOA]}}_{\times K}\texttt{[EOS]}}^{\mathcal{L}_{action}},

where {text}, {image}, and {action} refer to the discret text, image, and action tokens. [BOS], [EOS], [BOI], [EOI], [BOA], [EOA] refer to the beginning of sentence, end of sentence, beginning of image, end of image tokens, beginning of action, and end of action tokens. The input contains MM images and the output contains KK actions. We only calculate the loss of action tokens Laction\mathcal{L}_{action}.

World Model Data.

World model is to generate the next image frame given the current image observation and action. It does not need the task instruction since the action itself could totally determine the next state. The text inputs are "Generate the next frame based on the current image and the action.". The overall token sequence is:

[BOS]{text} [BOI]{image}[EOI][BOA]{action}[EOA][EOS][BOI]{image}[EOI][EOS]⏞Lworld⏟×N\underbrace{\texttt{[BOI]\{image\}[EOI][BOA]\{action\}[EOA][EOS]}\overbrace{\texttt{[BOI]\{image\}[EOI][EOS]}}^{\mathcal{L}_{world}}}_{\times N} .

The next frame prediction conditioned on the action repeats NN times, and we only calculate the loss of generated image tokens Lworld\mathcal{L}_{world}.

Attention Mask.

The standard attention mechanism in autoregressive models typically employs a causal attention mask, which restricts the current token’s access to information exclusively from preceding tokens, excluding any subsequent ones, as illustrated in Fig. 3 (a). Nevertheless, this conventional configuration proves inadequate for generating action chunks, i.e., multiple consecutive actions. While the foundational MLLM demonstrates robust generalization capabilities across image and text domains due to the large-scale pretraining on diverse datasets, its capacity to generalize effectively in the action domain is comparatively limited. Consequently, errors originating from earlier actions propagate to subsequent actions under the default attention mask, resulting in performance degradation. To address this limitation, we introduce an alternative attention mask tailored for action generation, depicted in Fig. 3 (b). This modified mask ensures that current actions rely solely on textual and visual inputs, while prohibiting access to prior actions. Such a design enables the autoregressive framework to generate multiple actions in parallel, aligning with methodologies presented in (Kim et al., 2025; Black et al., 2024). The world model part adheres to the conventional causal attention mask, as shown in Fig. 3 (c).

Training Objective.

We mix the action model data and world model data so that the autoregressive action world model could behave as both action model and world model. The loss function is:

where Laction\mathcal{L}_{action} and Lworld\mathcal{L}_{world} refer to the cross-entropy loss of the action model data and world model data. Since the image tokens (256 tokens for 256×256256\times 256 images and 1024 tokens for 512×512512\times 512 images) are much more than the action tokens (7 tokens), we use α\alpha to balance the loss contribution.

Experiments

We use LIBERO benchmark (Liu et al., 2023a) in our experiments. LIBERO benchmark contains LIBERO-Spatial, LIBERO-Object, LIBERO-Goal, LIBERO-Long and LIBERO-90. LIBERO-Spatial focuses on spatial relationships by requiring the robot to place a bowl based on its location. LIBERO-Object emphasizes object recognition by having the robot pick and place unique objects. LIBERO-Goal tests procedural learning through varying task goals with fixed objects. LIBERO-Long includes 10 long-horizon tasks. LIBERO-90 provides 90 short-horizon tasks for pretraining.

Datasets.

We first filter out the unsuccessful recorded trajectories and no-operation actions like OpenVLA (Kim et al., 2024). Considering the world model evaluation needs ground truth-paired video and action data, we split the 90% of the trajectories as the training set and the 10% remaining trajectories as the validation set. The training set is used for model training by default, with the exception of Table 2, where all available data are utilized during training to ensure a fair comparison.

Baselines.

There are two kinds of action models including the continous action model and discrete action model. Continous action model generates multiple actions in parellel and uses l1 regression loss for training. Diffusion-based action model like Diffusion Policy (Chi et al., 2023), Octo (Team et al., 2024), DiT Policy (Hou et al., 2024), and UVA (Li et al., 2025) use diffusion process to generate the actions. Seer (Tian et al., 2024) and OpenVLA-OFT (Kim et al., 2025) use an action head to directly output multiple actions in one time. Discrete action models like OpenVLA (Kim et al., 2024) considers the action as tokens just like texts, and the actions are generated in an autoregressive manner. Discrete models inherently exhibit inferior performance, as the tokenization process of actions may lead to information loss.

Training Setting.

The action model utilizes a default input image count of M=2M=2. The action chunk size is set to K=10K=10 for the LIBERO Long task and K=5K=5 for the remaining three LIBERO tasks under default configuration. To minimize computational expenditure, the world model operates with a single round N=1N=1. The parameter α\alpha is fixed at 0.04 in the experimental setup.

Metrics.

For action model evaluation, each task is evaluated for 50 rollouts under different initial states and we record the success rates (SR). For world model evaluation, we use the validation set and record the FVD, PSNR, SSIM, and LPIPS values.

2 Evaluation Results and Discussion

Table 2 indicates that the proposed WorldVLA model exhibits superior performance compared to the discrete OpenVLA model, even in the absence of pretraining. This outcome shows the effectiveness of the WorldVLA’s design. Furthermore, a positive correlation is observed between image resolution and model performance. Specifically, the 512 ∗* 512 pixel resolution yielded enhanced results compared to the 256 ∗* 256 pixel resolution. This phenomenon is primarily attributable to the pretraining regimen of the Chameleon backbone (Team, 2024), whose image tokenization module and the large language model components are inherently optimized at 512 ∗* 512 resolution. Additionally, higher resolution naturally provides a greater level of detailed visual information, which is particularly crucial for robotic grasping tasks since it demands high operational precision.

World Model Helps Action Model.

Quantitative results in Table LABEL:tab:ab_action, including row 2 vs. row 1, or row 5 vs. row 4, demonstrate that the integration of a world model significantly enhances the performance of the action model. The world model’s fundamental function involves predicting the subsequent state of the environment, conditioned on the current state and a given action. This generative process inherently promotes the acquisition of an understanding of the system’s underlying physical dynamics, which is a critical prerequisite for successful execution in dexterous manipulation tasks such as grasping. Furthermore, the world model endows the system with the capacity for prospective simulation, enabling it to anticipate the consequences of potential actions. This predictive foresight facilitates more informed decision-making, thereby optimizing action selection to maximize the probability of task success. Fig. 4 shows that the action model directly move to the destination without successfully grasping the cheese or bottle. In contrast, our action world model repeatedly attempts to grasp the objects until successful manipulation is achieved before proceeding to the target location.

Action Model Helps World Model.

Table LABEL:tab:ab_world demonstrates that the action world model outperforms the pure world model in terms of generation quality, particularly when producing longer video sequences. The action model derives actions based on the input images. On one hand, this contributes to more accurate visual interpretation; on the other, the process of generating actions enhances the understanding of the underlying behavioral patterns. Both aspects support the overall performance of the world model, which relies on robust comprehension of both visual and action-related information to predict future states effectively. As illustrated in Fig. 5, the pure world model fails in several scenarios: it is unable to open the drawer (a), causes the bowl to disappear after moving the disk (b), and fails to lift the bowl onto the stove (c). In contrast, the action world model produces coherent and physically plausible subsequent states in these cases.

Action Chunking Generation with Proposed Attention Mask.

Simultaneous generation of multiple actions is essential for achieving effective and efficient grasping. However, we observe that a naive autoregressive approach—where actions are generated sequentially—can degrade model performance, as evidenced by the results in row 3 of Table LABEL:tab:ab_action and Fig. 6. The grasping success rate gradullay decreases with longer action chunks. This degradation arises because later actions become overly dependent on preceding ones since they share the same space, rather than being grounded in visual input which is a distinct modality. The generalization of the action is not that strong as this modality was not involved during pretraining the MLLM. Consequently, errors tend to accumulate as the sequence of generated actions increases. The proposed attention masking mechanism ensures that each action is generated independently and solely determined by the visual input, thereby mitigating the issue of error propagation within the action sequence. As illustrated in Fig. 6, the model incorporating the proposed attention mask demonstrates superior performance compared to the naive attention mask, particularly under conditions of longer chunk lengths. This highlights the efficacy of the introduced masking approach. If the length of the action chunk is excessively prolonged, the robot’s ability to timely adapt its policy becomes constrained, leading to a decline in overall performance, as demonstrated in Fig. 6.

World Model versus Video Prediction Model.

Video prediction model is to generate the next frames based on the current frame and the task instruction. Video prediction has been used for pretraining the action model in prior research, such as GR-1 (Wu et al., 2023) and GR-2 (Cheang et al., 2024). Both video prediction model and world model belong to visual generation models, so we conduct a comparison to assess which framework provides greater utility for the action model. The text inputs of video prediction model are "Generate the future image based on the task and current image. + task instruction". The overall token sequence is:

[BOS]{text}[BOI]{image}…{image}[EOI]⏟×M\underbrace{\texttt{[BOI]\{image\}\ldots\{image\}[EOI]}}_{\times M}[EOS][BOI]{image}[EOI]⏟×K[EOS]⏞Lvideo\overbrace{\underbrace{\texttt{[BOI]\{image\}[EOI]}}_{\times K}\texttt{[EOS]}}^{\mathcal{L}_{video}}.

The difference between video prediction model and world model is that the world model is conditioned on the action while video prediction model is not. As illustrated in Fig. 7, the integration of a world model enhances the performance of the action model across all evaluated tasks. The video prediction model, however, demonstrates beneficial effects for two tasks while negatively impacting performance on one task. This discrepancy may arise from the inherent ambiguity in video prediction when action inputs are absent, as the subsequent frame cannot be uniquely determined from the initial frame alone. Consequently, multiple plausible future frames or ground truth sequences may correspond to a single starting frame, potentially introducing noise or inconsistency during training. Furthermore, the incorporation of a world model necessitates an understanding of actions which could contribute to more effective action generation.

Historical Image Input.

Unified models for understanding and generation, such as Chameleon (Team, 2024), employ the discrete image tokenizer VQGAN (Esser et al., 2021) for image interpretation. However, their capacity for semantic comprehension is comparatively limited when contrasted with vision-based perceptual models like CLIP (Radford et al., 2021). As demonstrated in Table LABEL:tab:his_input, the use of a single-frame input results in suboptimal performance. To enhance the model’s access to visual context, we incorporate multiple historical image frames, which leads to a progressive improvement in performance. Furthermore, the results indicate that the performance is saturated with two frames when generating action chunks. Consequently, we adopt a two-frame input configuration as the default in our experiments, optimizing the trade-off between task success rate and computational efficiency.

Pretrain Action Model using World Model.

Our WorldVLA framework integrates both action model data and world model data during training. We further investigate the possibility of utilizing the world model as a source of pretraining weights for the action model. This form of pretraining necessitates that the model develop an understanding of visual inputs, actions, and the underlying physical dynamics governing state transitions. As presented in Table LABEL:tab:pretrain, employing the world model for pretraining leads to notable improvements in grasping performance. These findings highlight the potential of leveraging world model pretraining in robotic applications, particularly in enhancing task-specific performance through prior exposure to general world knowledge.

Conclusion and Future Work

This study introduces WorldVLA, a novel autoregressive framework that unifies action and visual understanding with generation capabilities. We demonstrate that the integration of world modeling and action modeling within this architecture can lead to mutual enhancement in performance. An attention mask mechanism has been proposed to enable autoregressive generation of action sequences. Scaling of both data and model size emerges as a promising avenue for further development of the WorldVLA framework. Additionally, the current image tokenizer, which relies on discrete representations, exhibits limitations in perceptual expressiveness; hence, the design of a unified tokenizer capable of both understanding and generating high-quality visual content is an important direction for improvement. The incorporation of an auxiliary action head presents another potential strategy to enhance grasping performance. We anticipate that this work will contribute to and inspire future research in robotics, particularly in the domains of world modeling and unified models for action and image understanding and generation.

References