Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
Moo Jin Kim, Chelsea Finn, Percy Liang
I Introduction
Recent vision-language-action models (VLAs)—robot policies built by fine-tuning pretrained vision-language models on large-scale robot datasets for low-level robotic control—have demonstrated strong task performance, semantic generalization, and language following abilities across diverse robots and tasks . Despite their strengths, fine-tuning is crucial for satisfactory deployment of VLAs on novel robots and tasks, yet it is unclear what the most effective approach for adaptation is given the large design space. A robotics practitioner who wishes to fine-tune a VLA to a new robot setup and task may default to using the same training recipe used for pretraining (or a parameter-efficient variant), but it is not obvious whether this would yield the best policy, and there is limited empirical analysis of alternative fine-tuning approaches in the literature.
Prior work has begun exploring VLA adaptation strategies, with Kim et al. 2024 proposing parameter-efficient fine-tuning via LoRA. However, their autoregressive action generation remains too slow (3-5 Hz) for high-frequency control (25-50+ Hz), and both LoRA and full fine-tuning of autoregressive VLAs often yield unsatisfactory performance in bimanual manipulation tasks . While recent approaches improve efficiency through better action tokenization schemes , achieving 2 to 13 speedups, significant latency between action chunks (e.g., 750 ms for the recent FAST approach ) still limits real-time deployment on high-frequency bimanual robots. Exploring alternative VLA adaptation approaches that achieve both satisfactory speed and quality remains an underexplored area of research.
In this work, we study key design decisions for adapting VLAs to novel robots and tasks using OpenVLA, a representative autoregressive VLA, as our base model. We examine three key design choices: action decoding scheme (autoregressive vs. parallel generation), action representation (discrete vs. continuous), and learning objective (next-token prediction vs. L1 regression vs. diffusion). Our study reveals several key insights that build on each other: (1) parallel decoding with action chunking not only boosts inference efficiency but also improves success rates on downstream tasks while enabling greater flexibility in the model’s input-output specifications; (2) continuous action representations further improve model quality compared to discrete representations; and (3) fine-tuning the VLA with an L1 regression objective yields comparable performance to diffusion-based fine-tuning while offering faster training convergence and inference speed.
Building on these insights, we introduce OpenVLA-OFT: an instantiation of an Optimized Fine-Tuning (OFT) recipe that integrates parallel decoding and action chunking, continuous action representations, and an L1 regression objective to enhance inference efficiency, task performance, and model input-output flexibility while maintaining algorithmic simplicity. We conduct experiments on both the standardized LIBERO simulation benchmark and dexterous tasks on a real bimanual ALOHA robot. In LIBERO, OpenVLA-OFT establishes a new state of the art by achieving 97.1% average success rate across four task suites, outperforming both fine-tuned OpenVLA policies (76.5%) and policies (94.2%) while achieving a 26 speedup in action generation with 8-step action chunks. For real-world ALOHA tasks , we augment our recipe with FiLM for enhanced language grounding, denoting the augmented recipe as OFT+. OpenVLA-OFT+ successfully executes dexterous bimanual tasks like folding clothes and manipulating target food items based on the user’s prompt (see Figure 1), outperforming both fine-tuned VLAs ( and RDT-1B) and prominent imitation learning policies trained from scratch (Diffusion Policy and ACT) by up to 15% (absolute) in average success rate. With 25-timestep action chunks, OpenVLA-OFT+ achieves 43 faster throughput than base OpenVLA, demonstrating that our new fine-tuning recipe enables real-time robot control with strong task performance and language following ability.
II Related Work
Prior works have leveraged language and vision foundation models to enhance robotic capabilities, using them as pretrained visual representations that accelerate robotic policy learning , for object localization in robotics tasks , and for high-level planning and reasoning . More recently, researchers have explored fine-tuning vision-language models (VLMs) to directly predict low-level robotic control actions, producing “vision-language-action” models (VLAs) , which have demonstrated effective language following and generalization to out-of-distribution test conditions and unseen semantic concepts. These works focus primarily on model development, while we focus on developing a recipe for fine-tuning such models, justifying individual design decisions with insights that we gain from our empirical analysis.
Despite the importance of fine-tuning for real-world VLA deployment, empirical analysis of effective fine-tuning recipes remains limited. While Kim et al. 2024 study various parameter update strategies and from their findings show that LoRA fine-tuning enables effective adaptation to single-arm robots operating at low control frequencies ( Hz), their analysis does not extend to bimanual robots with high control frequencies (25-50+ Hz), a more complex control scenario. We address this gap by exploring VLA adaptation design decisions for fast inference and reliable task execution on a real-world bimanual manipulator with a 25 Hz controller.
Recent works by Belkhale and Sadigh 2024 and Pertsch et al. 2025 improve VLA efficiency through new action tokenization schemes, using vector quantization or discrete cosine transform-based compression to represent action chunks (sequences of actions) with fewer tokens than simple per-dimension binning (as used in RT-2 and OpenVLA ). While these approaches achieve 2 to 13 speedups for autoregressive VLAs, we explore design decisions beyond autoregressive modeling, which remains inherently limited by iterative generation. Our parallel decoding approach, when paired with action chunking, achieves significantly greater speedups: 26 to 43 throughput with much lower latency (0.07 ms for single-arm tasks with one input image and 0.321 ms for bimanual tasks with three input images).
Another line of research demonstrates effective VLA fine-tuning for high-frequency, bimanual manipulation using generative approaches like diffusion or flow matching. While these diffusion-based VLAs achieve higher action throughput than autoregressive VLAs by generating multi-timestep action chunks simultaneously, they introduce computational trade-offs through slower training and multiple denoising or integration steps at inference time. Furthermore, these diffusion VLAs vary considerably in architecture, learning algorithm, vision-language fusion approach, and input-output specifications—and which design elements most significantly impact performance remains unclear. Through controlled experiments, we show that policies fine-tuned with a simpler L1 regression objective can match more complex approaches in task performance while achieving significantly greater inference efficiency.
Lastly, unlike prior work that relies on a separate faster low-level control policy in addition to a slower VLA , in this work we fine-tune the base VLA end-to-end with techniques that enable high efficiency without needing a separate controller. Further, unlike other works that collect online interaction data to continually update specialist policies via reinforcement learning , here we focus on the simpler imitation learning paradigm in which we develop high-performing policies just by training once offline on a fixed dataset of expert task demonstrations.
III Preliminaries
Original OpenVLA formulation. We use OpenVLA as our representative base VLA, a 7B-parameter manipulation policy created by fine-tuning the Prismatic VLM on 1M episodes from the Open X-Embodiment dataset . See Appendix A-A for architecture details. OpenVLA’s original training formulation uses autoregressive prediction of 7 discrete robot action tokens per timestep: 3 for position control, 3 for orientation control, and 1 for gripper control. It employs next-token prediction with cross-entropy loss as its learning objective, similar to language models. We explore alternative formulations including parallel decoding, continuous action representations, and learning objectives like L1 regression and diffusion modeling in the next few sections.
Action chunking. Prior works have shown that action chunking—i.e., predicting and executing a sequence of future actions without intermediate replanning—improves policy success rates across many manipulation tasks . However, OpenVLA’s autoregressive generation scheme makes action chunking impractical, as generating even a single-timestep action takes 0.33 seconds on an NVIDIA A100 GPU. For a chunk size of timesteps and action dimensionality , OpenVLA requires sequential decoder forward passes versus just passes without chunking. This -fold increase in latency makes action chunking impractical for high-frequency robots under the original formulation. In the next section, we present a parallel generation scheme that enables efficient action chunking.
IV Studying Key VLA Fine-Tuning Design Decisions
In this section, we first outline key design decisions for adapting VLAs to novel robot setups and tasks and provide details on their implementation.
Existing approaches that fine-tune VLAs using the base model’s autoregressive training recipe face two key limitations: slow inference speed (3-5 Hz) unsuitable for high-frequency control, and unreliable task execution on bimanual manipulators .
To address these challenges, we investigate three key design components for VLA fine-tuning:
Action generation strategy (Figure 2, left): We compare autoregressive generation, which requires sequential token-by-token processing, with parallel decoding, which generates all actions simultaneously and enables efficient action chunking.
Action representation (Figure 2, right): We examine discrete actions (256-bin discretization of normalized actions) processed through softmax-based token prediction, versus continuous actions directly generated by an MLP action head. For discrete actions, the final hidden states of the language model decoder are linearly projected into logits, which are processed by a softmax operation to form the probability distribution over action tokens. For continuous actions, the final hidden states are instead mapped directly to normalized continuous actions by a separate action head MLP.
Learning objective (Figure 2, right): We compare policies fine-tuned with next-token prediction for discrete actions, L1 regression for continuous actions , and conditional denoising diffusion for continous actions (similar to Chi et al. 2023).
We conduct our study using OpenVLA as the base model, adapting it via LoRA fine-tuning due to our relatively small training datasets (500 demonstrations versus 1M demonstrations for pretraining).
IV-B Implementing Alternative Design Components
The base OpenVLA model originally employs autoregressive generation of discrete action tokens optimized via next-token prediction. We implement alternative design decisions for fine-tuning while keeping the original pretraining unchanged. We describe the key implementation aspects below, with further details explained in Appendix A-B.
Parallel decoding and action chunking. Unlike autoregressive generation which requires sequential token prediction, parallel decoding enables the model to map input embeddings to the predicted output sequence in a single forward pass. We modify the model to receive empty action embeddings as input and replace the causal attention mask with bidirectional attention, allowing the decoder to predict all actions simultaneously. This reduces action generation from sequential passes to a single pass, where is the action dimensionality.
Parallel decoding naturally extends to action chunking: to predict actions for multiple future timesteps, we simply insert additional empty action embeddings in the decoder’s inputs, which are then mapped to a chunk of future actions. For chunk size , the model predicts actions in one forward pass, increasing throughput -fold with minimal latency impact. While parallel decoding may theoretically be less expressive than autoregressive approaches, our experiments show no performance degradation across diverse tasks.
Continuous action representations. OpenVLA originally uses discrete action tokens where each action dimension is normalized to and uniformly discretized into 256 bins. While this approach is convenient since it requires no architectural modifications to the underlying VLM, the discretization process can sacrifice fine-grained action details. We study continuous action representations with two learning objectives drawn from prominent imitation learning approaches:
First, similar to Zhao et al. 2023, we implement L1 regression by replacing the decoder’s output embedding layer with an MLP action head that directly maps final decoder layer hidden states to continuous action values. The model is trained to minimize the mean L1 difference between predicted and ground-truth actions, maintaining the efficiency benefits of parallel decoding while potentially improving action precision.
Second, inspired by Chi et al. 2023, we implement conditional denoising diffusion modeling where the model learns to predict noise added to action samples during forward diffusion. During inference, the policy gradually denoises noisy action samples via reverse diffusion to produce real actions. While this approach offers potentially more expressive action modeling, it requires multiple forward passes during inference (50 diffusion steps in our implementation), impacting deployment latency even with parallel decoding.
Additional model inputs and outputs. While the original OpenVLA processes a single camera view, some robot setups include multiple viewpoints and additional robot state information. We implement a flexible input processing pipeline: For camera images, we use OpenVLA’s dual vision encoder to extract 256 patch embeddings per view, which are projected into the language embedding space with a shared projector network. For low-dimensional robot state inputs (e.g., joint angles and gripper state), we employ a separate projection network to map these into the same embedding space as one additional input embedding.
All input embeddings—visual features, robot state, and language tokens—are concatenated along the sequence dimension before being passed to the decoder. This unified latent representation enables the model to attend to all available information when generating actions. Combined with parallel decoding and action chunking, this architecture can efficiently process rich multimodal inputs while generating multiple timesteps of actions, as illustrated in Figure 1.
IV-C Augmenting OpenVLA-OFT with FiLM for Enhanced Language Grounding.
Challenges with language following. When deploying on the ALOHA robot setup with multiple viewpoints including from wrist-mounted cameras, we observe that policies can struggle with language following due to spurious correlations in visual inputs. During training, policies may learn to latch onto such spurious correlations when predicting actions, rather than properly attending to the language instructions, resulting in poor adherence to the user’s commands at test time. Furthermore, language inputs may only be critical at specific moments in a task—for example, after grasping the spoon and deciding which ingredient to scoop in the “scoop X into bowl” task discussed in Section VI. Therefore, without special techniques, training the model to appropriately focus on language inputs can be particularly challenging.
FiLM. To enhance language following, we employ feature-wise linear modulation (FiLM) , which infuses language embeddings into the visual representations so that the model pays more attention to the language inputs. We compute the average of the language embeddings from the task description and project it to obtain scaling and shifting vectors and . These vectors modulate the visual features through an affine transformation:
A crucial implementation detail is the choice of what represents a “feature” for modulation in vision transformers. While one might naturally consider treating individual patch embeddings as features to be modulated, we find that this approach results in poor language following. Instead, drawing from how FiLM operates in convolutional networks, where modulation applies spatially-agnostically by scaling and shifting entire feature maps, we apply each element of and to the corresponding hidden unit across all visual patch embeddings so that and influence all patch embeddings. Concretely, this makes and -dimensional vectors, where is the number of hidden dimensions (i.e., the number of elements in each of the patch embeddings in the vision transformer’s latent representations).
We apply FiLM after the self-attention layer and before the feedforward layer in each vision transformer block, with separate projectors for each block (see Figure 8). Additional implementation details are provided in Appendix A-C. We only use FiLM for the ALOHA experiments discussed in Section VI, where multiple camera viewpoints lead to a larger presence of spurious correlations in visual inputs.
V Experiments: Evaluating VLA Fine-Tuning Design Decisions
In this section, we evaluate the effects of our proposed VLA adaptation design decisions through controlled experiments aimed at answering three key questions:
How does each design decision affect the fine-tuned policy’s success rate on downstream tasks?
How does each design decision affect model inference efficiency (action generation throughput and latency)?
How do the alternative fine-tuning formulations affect flexibility in model input-output specifications?
We evaluate on the LIBERO simulation benchmark , which features a Franka Emika Panda arm in simulation with demonstrations containing camera images, robot state, task annotations, and delta end-effector pose actions. We use four task suites—LIBERO-Spatial, LIBERO-Object, LIBERO-Goal, and LIBERO-Long—each providing 500 expert demonstrations across 10 tasks to assess policy generalization to different spatial layouts, objects, goals, and long-horizon tasks.
Following Kim et al. 2024, we filter unsuccessful demonstrations and fine-tune OpenVLA via LoRA on each task suite independently. We train for 50-150K gradient steps for non-diffusion methods and 100-250K steps for diffusion methods (which converge slower), using a batch size of 64-128 across 8 A100/H100 GPUs. We test checkpoints every 50K steps and report the best performance for each run. Unless specified otherwise, policies receive one third-person image and language instruction as input. For methods using action chunking, we set chunk size to to match the Diffusion Policy baseline , and execute full chunks before replanning, which we find improves both speed and performance. See Appendix A-D for hyperparameter details.
Our primary baseline in this study is the base OpenVLA model fine-tuned using the original fine-tuning recipe. However, for broader comparison, we also include LIBERO results from prior state-of-the-art imitation learning methods, such as Diffusion Policy , Octo , DiT Policy , Seer , MDT , and . Note that Seer uses additional LIBERO-90 pretraining data.
V-B LIBERO Task Performance Comparisons
For satisfactory deployment, robot policies must demonstrate reliable task execution. We first assess how different VLA fine-tuning design decisions affect success rates on the LIBERO benchmark.
Our efficiency analysis (which we discuss later) reveals that parallel decoding (PD) and action chunking (AC) together are necessary for high-frequency control (25-50+ Hz), especially for bimanual robots with double the amount of action dimensions. We therefore evaluate OpenVLA policies with both techniques used jointly, comparing variants using discrete actions, continuous actions with L1 regression, and continuous actions with diffusion.
Results in Table I show that parallel decoding and action chunking not only increase throughput but also improve performance significantly, raising average success rates by 14% (absolute) over autoregressive OpenVLA policies. This improvement is particularly pronounced in LIBERO-Long, suggesting that action chunking helps capture temporal dependencies and reduce compounding errors , which ultimately leads to smoother and more reliable task execution. In addition, we find that using continuous action variants further improves success rates by 5% (absolute) over the discrete action variant, likely due to higher precision in the action predictions. L1 regression and diffusion variants achieve comparable performance, indicating that the high-capacity OpenVLA model can effectively model the multi-task action distribution even with simple L1 regression.
V-C LIBERO Inference Efficiency Comparisons
Efficient inference is crucial for deploying VLAs on high-frequency control robots. We now evaluate how parallel decoding (PD), action chunking (AC), and continuous action representations affect model inference speed. We measure average latency (time to generate one robot action or action chunk) and throughput (total actions generated per second) by querying each model variant 100 times on an NVIDIA A100 GPU. Each query processes a 224 x 224 px image and a sample LIBERO language instruction (“pick up the alphabet soup and place it in the basket”).
Results in Table II show that parallel decoding reduces latency and increases throughput by 4 by replacing 7 sequential forward passes through the decoder portion of the policy with a single pass. Adding action chunking () increases latency by 17% due to longer attention sequences in the decoder, but when combined with parallel decoding, it dramatically improves throughput, achieving a 26 speedup over baseline OpenVLA. The continuous actions variant with L1 regression shows negligible difference in efficiency since the additional MLP action head adds minimal computational cost compared to the base model. The primary diffusion variant requires 50 denoising steps (since this is the number of steps specified at train time) and thus suffers from high latency. However, it still achieves the same effective throughput as baseline OpenVLA due to parallel decoding and action chunking. This means that despite longer pauses between action chunks, the 50-step diffusion variant still completes robot episodes at the same speed as the original autoregressive variant. We also test the OpenVLA diffusion variant using less than 50 denoising steps at inference time (enabled by the DDIM sampler ), matching state-of-the-art diffusion and flow matching methods that use 5-10 steps at inference time . These configurations result in much greater inference efficiency, as shown in Table II. However, the success rates decrease with fewer denoising steps due to reduced model quality.
V-D Model Input-Output Flexibility
As explained in Section IV-B and validated by our efficiency evaluations in the prior section, parallel decoding enables OpenVLA to generate action chunks with minimal latency increase, thereby enhancing flexibility in model outputs. The significant speedup realized by parallel decoding and action chunking creates headroom for processing additional model inputs as well. We demonstrate this by fine-tuning OpenVLA with additional inputs such as robot proprioceptive state and a robot wrist-mounted camera image, which doubles the number of visual patch embeddings being passed into the language model decoder, from 256 to 512. Despite this substantial increase in input sequence length, the fine-tuned OpenVLA policy maintains high throughput (71.4 Hz) and low latency (0.112 sec), as shown in Table II.
Evaluating these policies with additional inputs on the LIBERO benchmark reveals further improvements in average success rate across all task suites (Table I). Notably, our enhanced fine-tuned OpenVLA policies outperform even the best fine-tuned policies —which benefit from a base model with larger-scale pretraining and a more sophisticated learning objective (flow matching )—as well as Multimodal Diffusion Transformer (MDT) and Seer policies. Even with a simpler base model pretrained on less data than more recent VLAs, we find that our alternative VLA adaptation design decisions empower fine-tuned OpenVLA policies to establish a new state of the art on the LIBERO benchmark.
V-E Optimized Fine-Tuning Recipe
Based on the demonstrated improvements in task performance, inference efficiency, and model input-output flexibility, we propose an Optimized Fine-Tuning (OFT) recipe for VLA adaptation that combines three key components:
These design choices work together to produce strong policies that can be deployed at high frequencies while maintaining algorithmic simplicity. We denote policies fine-tuned from the OpenVLA base model using our OFT recipe as OpenVLA-OFT. In Section VI, we evaluate OpenVLA-OFT’s capabilities on dexterous, bimanual manipulation tasks in the real world using a high-frequency control robot.
V-F Additional Experiments
Given that the alternative fine-tuning formulation, along with additional model inputs and outputs, induces a distribution shift between the base VLA’s pretraining and fine-tuning, one reasonable question is whether the base VLA’s pretrained representations are helpful and have any influence on the results we have reported. We conduct an ablation study in Appendix A-G2 to address this question, ablating the OpenVLA pretraining phase and directly fine-tuning the underlying pretrained VLM with the OFT recipe. As shown in Table XV, the base OpenVLA pretrained representations are indeed still beneficial for robotic policy learning, as removing them leads to a 5.2% drop in average success rate (absolute) in our LIBERO evaluation suite.
(See Appendix A-G for other additional experiments, including scaling OpenVLA-OFT up to larger datasets—in both the LIBERO simulation setting as well as a real-world single-arm robot manipulation setting with over 50K demonstrations from the BridgeData V2 dataset .)
VI Experiments: Adapting OpenVLA to a Real-World ALOHA Robot
While our experimental results in the prior section demonstrate OpenVLA-OFT’s effectiveness in simulation, successful deployment in the real world, on robot platforms that differ substantially from those seen during pretraining, is crucial for showing broad applicability. We thus assess the efficacy of our optimized fine-tuning recipe on the ALOHA robot setup , a real bimanual manipulation platform operating at a high control frequency. We evaluate on novel dexterous manipulation tasks that have never been encountered before during OpenVLA’s pretraining (which only involves single-arm robot data).
Prior works have shown that vanilla LoRA fine-tuning with autoregressive VLAs is impractical for such tasks, as its throughput (3-5 Hz for single-arm robots and even lower for bimanual tasks) falls well below the 25-50 Hz required for real-time deployment. We therefore exclude this baseline from our experiments and compare more effective methods that we discuss shortly.
In this section, we use an augmented version of our VLA fine-tuning recipe (OFT+) that additionally includes feature-wise linear modulation (FiLM) for enhanced language grounding, as described in Section IV-C. We denote the OpenVLA policy instantiated through this augmented fine-tuning recipe as OpenVLA-OFT+.
The ALOHA platform comprises two ViperX 300 S arms, three camera viewpoints (one top-down, two wrist-mounted), and robot state inputs (14-dimensional joint angles). It operates at 25 Hz (reduced from the original 50 Hz to enable faster training while still maintaining smooth robotic control), with actions representing target absolute joint angles. This setup differs significantly from OpenVLA’s pretraining, which includes single-arm robot data only, a single camera viewpoint from a third-person camera, no robot state inputs, low-frequency control (3-10 Hz), and relative end-effector pose actions. The distribution shift poses a challenge to the adaptation of this model.
We design four representative tasks testing deformable object manipulation, long-horizon skills, tool usage, and language-driven control:
“fold shorts”: Fold white shorts on a table with two consecutive bimanual folds. Training: 20 demonstrations. Evaluation: 10 trials.
“fold shirt”: Fold white T-shirt through multiple synchronized bimanual folds, testing contact-rich, long-horizon manipulation. Training: 30 demonstrations. Evaluation: 10 trials.
“scoop X into bowl”: Move bowl to center of table with left arm, scoop specified ingredient (“raisins,” “almonds and green M&Ms,” or “pretzels”) with right arm using metal spoon. Training: 45 demonstrations (15 per ingredient). Evaluation: 12 trials (4 per ingredient).
“put X into pot”: Open pot with left arm, place specified item (“green pepper,” “red pepper,” or “yellow corn”) with right arm, close pot. Training: 300 demonstrations (100 per object). Evaluation: 24 trials (12 in-distribution, 12 out-of-distribution).
We fine-tune OpenVLA using OFT+ on each task independently for 50-150K gradient steps (total batch size 32 with 8 A100/H100-80GB GPUs) with action chunk size . At inference time, we execute the full action chunk before requerying the model for the next chunk.
VI-B Methods in Comparison
The ALOHA tasks present a significant adaptation challenge for OpenVLA as the base model, given the substantial differences from its pretraining platforms in terms of control frequency, action space, and input modalities. For this reason, we compare OpenVLA-OFT+ against more recent VLAs—RDT-1B and —that were pretrained on bimanual manipulation data and might reasonably be expected to perform better on these downstream tasks. We evaluate these models after fine-tuning them using their authors’ recommended recipes, and these methods serve as important points of comparison. Additionally, to provide comparisons with computationally efficient alternatives, we evaluate two popular imitation learning baselines: ACT and Diffusion Policy , trained from scratch on each task.
To enable language following in these baseline methods, we use language-conditioned implementations. For ACT, we modify EfficientNet-B0 to process CLIP language embeddings via FiLM . We use this FiLM-EfficientNet implementation only for language-dependent tasks (“scoop X into bowl” and “put X into pot”). For clothes folding tasks, we use the original ResNet-18 backbone as in . For Diffusion Policy, we use the DROID dataset implementation that conditions action denoising on DistilBERT language embeddings, modified to support bimanual control and multiple image inputs.
VI-C ALOHA Task Performance Results
We evaluate all methods—ACT, Diffusion Policy, RDT-1B, , and OpenVLA-OFT+—on our four ALOHA tasks. To provide fine-grained assessment, we use a predetermined rubric that assigns scores for partial task completion (see Appendix A-F for details). Figure 4 shows aggregate performance scores, while Figure 5 specifically tracks language following ability for the language-dependent tasks.
Performance of non-VLA baselines. The baseline methods trained from scratch show varying levels of success. ACT, while able to complete basic tasks, produces less precise actions and achieves the lowest overall performance. Diffusion Policy demonstrates stronger capabilities, matching or exceeding RDT-1B’s reliability on the clothes folding and scooping tasks. However, it struggles with the “put X into pot” task which has a larger training dataset, suggesting limited scalability compared to VLA-based approaches.
Performance of fine-tuned VLAs. Fine-tuned VLA policies generally outperform the from-scratch baselines in both task execution and language following, consistent with prior findings . Among VLAs, we observe distinct characteristics: RDT-1B achieves good language following through its “Alternating Condition Injection” scheme , but shows a limitation in handling closed-loop feedback. As visualized in Figure 6, it often fails to correct mistakes in the “scoop X into bowl” task—for instance, continuing to pour ingredients into an imaginary bowl after missing the actual bowl, suggesting over-reliance on proprioceptive state over visual feedback. On the other hand, demonstrates more robust execution with smoother motions and better reactivity to feedback, often successfully recovering from initial failures (as shown in Figure 6). While its language following slightly trails RDT-1B’s, achieves better overall task completion, making it the strongest baseline. Finally, OpenVLA-OFT+ achieves the highest performance across both task execution and language following (see Figure 7 for examples of successful task rollouts). This is particularly noteworthy given that the base OpenVLA model was pretrained only on single-arm data, while RDT-1B and were pretrained on substantial bimanual datasets (6K episodes and 8K hours of bimanual data, respectively). This suggests that the fine-tuning technique can be more crucial than pretraining data coverage for downstream performance.
Ablation study of FiLM. We evaluate the importance of FiLM in our OpenVLA-OFT+ approach by ablating it and assessing the policies’ language following ability on the last two tasks, which require good language grounding for successful execution. As shown in Figure 5, language following drops to 33% in both tasks—equal to randomly choosing the correct instruction. This demonstrates that FiLM is essential for preventing the model from overfitting to spurious visual features and ensuring proper attention to language inputs.
Please see the project website for ALOHA robot rollout videos and an in-depth qualitative analysis of all methods: https://openvla-oft.github.io
VI-D ALOHA Inference Efficiency Comparisons
We evaluate inference efficiency by measuring action generation throughput and latency across 100 queries for each method. We report the results in Table III. The original OpenVLA formulation, even with just the additional wrist camera inputs, shows poor efficiency with 1.8 Hz throughput and 0.543 sec latency. In contrast, OpenVLA-OFT+ achieves 77.9 Hz throughput, though its latency is higher compared to the policies in the prior LIBERO experiments since it must process two additional input images.
Other methods demonstrate higher throughput than OpenVLA-OFT+ due to their smaller architectures: ACT (84M parameters), Diffusion Policy (157M), RDT-1B (1.2B), and (3.3B)—while OpenVLA has 7.5B parameters. ACT achieves the highest speed by combining L1 regression-based single-pass action generation (like OpenVLA-OFT+) with its compact architecture. Also, despite its larger size, outperforms both RDT-1B and Diffusion Policy in speed thanks to its optimized JAX implementation (all other methods are implemented in PyTorch).
Notably, OpenVLA-OFT+’s throughput (77.9 Hz) approaches RDT-1B’s (84.1 Hz) despite being 7 larger, as it generates actions in a single forward pass rather than requiring multiple denoising steps as in RDT-1B.
VII Discussion
Our study on VLA fine-tuning design decisions reveals how different components impact inference efficiency, task performance, model input-output flexibility, and language following ability. These insights lead to our Optimized Fine-Tuning (OFT) recipe, which enables effective VLA adaptation to novel robots and tasks through parallel decoding, action chunking, continuous actions, L1 regression, and (optionally) FiLM language conditioning. The success of OFT is particularly noteworthy with OpenVLA: despite having no exposure to bimanual robots or multi-view image inputs during pretraining, OpenVLA fine-tuned with OFT can adapt to such configurations and match or even outperform more recent diffusion-based VLAs ( and RDT-1B) which have encountered bimanual manipulators and multiple input images during pretraining. This demonstrates that a well-designed fine-tuning recipe can have a significant impact on final performance, and existing VLAs can be successfully adapted to new robotic systems without extensive retraining from scratch. Moreover, our results show that a simple L1 regression-based approach with a high-capacity model such as OpenVLA is quite effective for adapting to novel robots and tasks. This approach offers practical advantages over diffusion-based methods: the simpler algorithm leads to faster training convergence and inference speed while maintaining strong performance, making it particularly suitable for real-world robotics applications.
VIII Limitations
While our Optimized Fine-Tuning (OFT) recipe shows promise for adapting VLAs to novel robots and tasks, several important questions remain.
Handling multimodal demonstrations. Our experiments use focused demonstration datasets with a consistent strategy per task. While L1 regression may help smoothen out noise in training demonstrations by encouraging the policy to learn the median mode in demonstrated actions, it may struggle to accurately model truly multimodal action distributions where multiple valid actions exist for the same input, which may not be ideal in cases where the ability to generate alternative action sequences would be beneficial for task completion. Conversely, diffusion-based approaches may better capture such multimodality but risk overfitting to suboptimal modes in training data (see our website for discussions and video illustrations of these nuances). Understanding OFT’s effectiveness with multimodal demonstrations remains an important direction for future work.
Pretraining versus fine-tuning. Our study focuses specifically on fine-tuning VLAs for downstream tasks. Whether OFT’s benefits extend effectively to pretraining, or whether more expressive algorithms like diffusion are necessary for large-scale training, requires further investigation.
Inconsistent language grounding. Our ALOHA experiments reveal that OpenVLA without FiLM exhibits poor language grounding, despite showing no such issues in LIBERO simulation benchmark experiments. The source of this discrepancy—whether from the lack of bimanual data in pretraining or other factors—remains unclear and warrants further study.
Acknowledgments
We thank the Stanford Center for Research on Foundation Models (CRFM) and Stanford Institute for Human-Centered AI (HAI) for providing computational resources that supported this research. Toyota Research Institute (TRI) provided funds to assist the authors with their research, but this article solely reflects the opinions and conclusions of its authors and not TRI or any other Toyota entity. This work was also in part supported by the Robotics and AI Institute and ONR grant N00014-22-1-2621. We thank Physical Intelligence for providing beta access to the model used in our ALOHA robot evaluations. We also thank Dan Fu, Karl Pertsch, and Alec Lessing for insightful discussions which contributed to the development of this work. Lastly, we are grateful to Moritz Reuss for providing additional experimental results for MDT and engaging in helpful discussions related to this work.
References
Appendix A Appendix
Base OpenVLA Architecture. OpenVLA combines a fused vision backbone (with both SigLIP and DINOv2 vision transformers), a Llama-2 7B language model , and a 3-layer MLP projector with GELU activation for projecting visual features into the language embedding space.
The original model processes a single third-person image and a language instruction (e.g., “put eggplant into pot”). The fused vision encoder extracts 256 patch embeddings from each vision transformer, concatenates them along the hidden dimension, and projects them into the language embedding space. These projected features are concatenated with language embeddings along the sequence dimension before being processed by the Llama-2 decoder to output a 7-dimensional robot action representing delta end-effector pose, represented by a string of discrete action tokens.
OpenVLA-OFT architecture modifications. OpenVLA-OFT introduces six key changes:
processes multiple input images (e.g., third-person image plus wrist camera images) through the shared SigLIP-DINOv2 backbone
projects robot proprioceptive state to language embedding space via a 2-layer MLP with GELU activation
replaces causal attention with bidirectional attention for parallel decoding
substitutes the language model decoder output layer with a 4-layer MLP (ReLU activation) for generation of continuous actions (instead of discrete actions)
outputs chunks of K actions instead of single-timestep actions
(for OpenVLA-OFT+) adds FiLM modules that use the average task language embedding to modulate visual features in both SigLIP and DINOv2 vision transformers (see Appendix A-C for details)
The complete OpenVLA-OFT+ architecture is illustrated in Figure 1.
A-B Implementation Details
In the original OpenVLA autoregressive training scheme, the model receives ground-truth action tokens shifted right by one position as input (a setup known as teacher forcing). A causal attention mask ensures the model only attends to current and previous tokens. At test time, each predicted token is fed back as input for the next prediction.
For parallel decoding, we replace this input with empty action embeddings that differ only in their positional encoding values (similar to ). We also use a bidirectional attention mask (instead of causal), enabling the model to leverage all intermediate features non-causally when predicting each element in the action chunk.
A-B2 Continuous Action Representations
For discrete actions, increasing the number of bins used for discretization improves precision but reduces the frequency of individual tokens in the training data, potentially hurting generalization. On the other hand, with a continuous action representation, the VLA can directly model the action distribution without lossy discretization.
Our continuous representation implementations use the following specifications:
L1 regression: The MLP action head consists of 4 layers with ReLU activation, mapping final Llama-2 decoder layer hidden states directly to continuous actions.
4-layer noise predictor with same MLP architecture as the L1 regression head
A-B3 Input Processing Details
Passing each input image through the OpenVLA fused vision encoder produces 256 patch embeddings, which are projected to the langauge model embedding space via a 3-layer MLP with GELU activation . Low-dimensional robot states are also projected to the language embedding space through a 2-layer MLP with GELU activation.
A-C Feature-wise Linear Modulation (FiLM) Implementation Details
FiLM schematic. Section IV-C describes how we implement feature-wise linear modulation (FiLM) for OpenVLA. Figure 8 illustrates our implementation. OpenVLA has a fused vision encoder with both SigLIP and DINOv2 vision transformers, and we apply FiLM to both transformers.
Design considerations. In our implementation, following Perez et al. 2018, we multiply by instead of since and are near zero at initialization. This helps preserve the visual encoder’s original activations at the start of fine-tuning, minimizing perturbation in the pretrained representation.
Implementation specifics. The functions and that project language embeddings to obtain and are implemented as simple affine transformations. Separate projectors are learned for each transformer block to allow for block-specific modulation patterns. This design enables the model to learn different modulation patterns at different levels of visual feature processing.
One might initially consider modulating each patch embedding independently, as opposed to each hidden dimension of each embedding as discussed in Section IV-C. However, our spatially-agnostic modulation approach more closely mirrors FiLM’s operation in convolutional networks, where modulation applies globally across spatial dimensions since entire feature maps are scaled and shifted by individual elements of and . This design choice better maintains the benefits of FiLM and improves the policy’s language grounding substantially. We find that an alternative formulation that modulates each patch embedding independently leads to weaker language grounding.
A-D OpenVLA-OFT Hyperparameters and Training Details
OpenVLA-OFT training details for LIBERO. Hyperparameters for OpenVLA-OFT fine-tuning on LIBERO are listed in Table IV. We train until the mean L1 loss between predicted and ground-truth normalized actions (scaled between ) falls below 0.01. For faster convergence, we decay the learning rate from 5e-4 to 5e-5 after 100K gradient steps. We evaluate checkpoints every 50K steps, with the 150K checkpoint achieving best performance in all task suites except for LIBERO-Goal. Note that we do not use FiLM for LIBERO experiments since the fine-tuned policies without it already demonstrate good language grounding.
OpenVLA-OFT+ training details for ALOHA. Hyperparameters for OpenVLA-OFT+ training on ALOHA tasks (with FiLM in the augmented OFT+ recipe) are shown in Table V. We maintain the same convergence criterion as in the LIBERO experiments (training until mean normalized L1 loss falls below 0.01) and similar learning rate decay strategy (again 10 reduction, but after 50K gradient steps instead of 100K, since the ALOHA datasets are smaller). For the two clothes folding tasks which do not require language grounding, we still include FiLM to verify that these additional parameters do not impair task execution.
A-E Baseline Methods Hyperparameters and Training Details
ACT training details for ALOHA. Table VI lists hyperparameters for ACT trained from scratch on each task. For non-language-dependent tasks (“fold shorts”, “fold shirt”), we use the default ResNet-18 backbone, which does not include language conditioning. For language-dependent tasks (“scoop X into bowl”, “put X into pot”), we implement EfficientNet-B0 with FiLM , similar to . While the authors of ACT recommend training for at least 5K epochs, we extend training to 10K-70K epochs per task to improve performance.
Diffusion Policy training details for ALOHA. For Diffusion Policy training, we use the DROID implementation , which conditions action predictions on DistilBERT language embeddings of the task description. We list hyperparameters in Table VII.
RDT-1B training details for ALOHA. Hyperparameters for RDT-1B fine-tuning are shown in Table VIII. The authors of RDT-1B recommend training for 150K gradient steps, but we observe that training converges in significantly fewer steps since our fine-tuning datasets are much smaller than the RDT-1B fine-tuning dataset. Therefore, we observe that it is unnecessary to train for such a large amount of time. In fact, on “scoop X into bowl”, the earlier 18K step checkpoint (73.3% success) outperforms the later 40K step checkpoint (70.0%) (we report the former in our ALOHA experiments).
training details for ALOHA. Table IX lists hyperparameters for fine-tuning. We use full fine-tuning (the default option in their codebase) and train until convergence.
A-F ALOHA Evaluation Details
Below are detailed specifications for each task in our ALOHA experiments:
Task: Bimanual folding of white shorts with two synchronized folds
Dataset: 20 demonstrations (19 training, 1 validation)
Episode length: 1000 timesteps (40 seconds)
Task: Long-horizon T-shirt folding with multiple synchronized bimanual folds
Dataset: 30 demonstrations (29 training, 1 validation)
Episode length: 1250 timesteps (50 seconds)
Task: Move bowl to center, scoop specified ingredient (raisins, almonds and green M&Ms, or pretzels) into bowl
Dataset: 45 demonstrations (15 per target; 42 training, 3 validation)
Episode length: 900 timesteps (36 seconds)
Task: Open pot, place specified item (green pepper, red pepper, or yellow corn) into pot, close pot
Dataset: 300 demonstrations (100 per target; 285 training, 15 validation) This relatively large number of demonstrations for the “put X into pot” task is not necessary for satisfactory performance. It simply reflects an earlier investigative phase of this work during which we encountered difficulties with language grounding in learned policies and initially hypothesized that increasing demonstration quantity might improve language following capabilities. However, merely expanding the training set proved insufficient for achieving satisfactory language grounding, and we discovered that additional techniques were needed for more reliable language grounding. We still decided to fine-tune on the full dataset with 300 demonstrations nonetheless.
Initial variation: 45 cm horizontal, 20 cm vertical for food items; fixed pot pose
Episode length: 400 timesteps (16 seconds)
Evaluation: 24 trials (12 in-distribution evaluations, 12 out-of-distribution evaluations)
Initial states: See Figures 12 (in-distribution) and 13 (out-of-distribution)
A-F2 ALOHA Task Scoring Rubric
The scoring rubrics and detailed results for the four ALOHA tasks are shown in Tables X, XI, XII, and XIII.
A-G Additional Experiments
In Section V and Table I, we report results with OpenVLA-OFT policies trained on each task suite independently. In this section, we assess whether our method scales to larger fine-tuning datasets by training one OpenVLA-OFT policy on all four task suites combined. As shown in Table XIV, this new policy achieves comparable average task performance as the task suite-specific policies—confirming that our method scales to larger fine-tuning datasets.
A-G2 Ablating FiLM in LIBERO
The FiLM ablation study in Section VI suggests that FiLM is crucial for enabling strong language following in real-world ALOHA robot tasks. In this section, we assess whether FiLM is similarly important in the LIBERO simulation tasks. We train a single OpenVLA-OFT+ policy (with FiLM) on all LIBERO task suites combined (similar to the previous section) and report task performance results in Table XIV, comparing performance against the OpenVLA-OFT policy trained without FiLM. In LIBERO, FiLM leads to slightly higher average success rate, though the difference here is minor compared to the findings in the real-world ALOHA experiments, where language following is more challenging due to the reasons discussed in Section IV-C.
A-G3 Ablating the OpenVLA Pretrained Representation
We evaluate the performance of OpenVLA-OFT policies produced by fine-tuning the underlying Prismatic VLM directly on the LIBERO downstream datasets without OpenVLA’s Open X-Embodiment robot pretraining. This ablation study investigates whether OpenVLA’s robot-pretrained representation remains valuable when subjected to a substantially different fine-tuning approach such as OFT. The results in Table XV demonstrate that the variant without the pretrained OpenVLA representation consistently underperforms compared to the full OpenVLA-OFT model, confirming the benefits of using the pretrained representation for downstream policy learning.
A-G4 Scaling Up OpenVLA-OFT to a Larger Real-World Dataset (BridgeData V2)
In Appendix A-G1, we observe that a single OpenVLA-OFT policy can effectively fit all four LIBERO task suite datasets combined, confirming that the proposed method scales to larger fine-tuning datasets. In this section, we scale up the fine-tuning data further and assess whether OpenVLA-OFT can also fit a real-world robotic manipulation dataset that is significantly larger and more diverse than both the LIBERO datasets and the ALOHA robot datasets discussed in Section VI. Specifically, we train OpenVLA-OFT (without FiLM) on the BridgeData V2 dataset , which contains 50,365 real WidowX robot demonstrations (25 more than the four LIBERO task suites combined). Although the base OpenVLA model was already pretrained on Bridge data, we note that the model must still be fine-tuned when using the new OFT recipe since the architecture, learning algorithm, and action representation and decoding scheme have been modified significantly from the pretraining setup. In fact, the initial action regression L1 loss in the beginning of OpenVLA-OFT training on Bridge is large (roughly 0.5 when actions are normalized to the scale ).
We evaluate OpenVLA-OFT on a subset of BridgeData V2 WidowX robot tasks from the evaluation suite used in the original OpenVLA work . This representative subset covers the four types of generalization (visual, motion, physical, semantic) as well as language grounding tasks. We compare the task performance of OpenVLA-OFT with that of the public OpenVLA checkpoint, scoring both methods using the same criteria used in the OpenVLA work. As shown in Table XVI, OpenVLA-OFT surpasses OpenVLA on average across these tasks (69.2% versus 65.8%, respectively)—confirming scalability of the proposed OFT approach to much larger and more diverse real robot datasets. Further, we observe that FiLM is not necessary for satisfactory language following in Bridge tasks, likely because the base model has shown effective language following in Bridge tasks.