AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning

Bo Jiang, Shaoyu Chen, Qian Zhang, Wenyu Liu, Xinggang Wang

Introduction

Autonomous driving has witnessed rapid advances in recent years, with end-to-end autonomous driving emerging as one of the most representative models . They take sensor data as input and leverage learnable neural networks to plan the vehicle’s future trajectory. Benefiting from large-scale driving demonstrations, end-to-end models continuously improving their planning capabilities by expanding training data and increasing model parameters.

However, due to their black-box nature and lack of common sense, end-to-end models still face significant challenges when handling complex and long-tail driving scenarios. For instance, consider a situation where the vehicle ahead is carrying traffic cones while driving. An end-to-end model may fail to comprehend the relationship between the leading vehicle and the traffic cones, mistakenly assuming that the road ahead is under construction and thus impassable, leading to an incorrect decision to brake. Therefore, relying solely on end-to-end models to achieve high-level autonomous driving remains challenging.

With the success of GPT , large language models (LLMs) show remarkable comprehension and reasoning abilities . Furthermore, their capabilities have evolved from unimodal text understanding to multimodal vision-language processing. . The commonsense and reasoning abilities of VLMs hold great potential to mitigate the limitations of end-to-end models.

Recently, OpenAI o1 , which incorporates reasoning techniques, achieves performance comparable to or even surpassing that of human experts in fields such as programming. Additionally, DeepSeek R1 , which leverages reinforcement learning, not only demonstrates “emergent abilities”and achieves top-tier performance but also requires significantly lower training costs compared to other models. These advances underscore the immense potential of reasoning techniques and RL in the development of large models.

Existing research on applying VLMs to autonomous driving can be broadly categorized into two directions. The first focuses on leveraging VLMs for the understanding of driving scenes . The second explores the use of VLMs for planning, where some studies treat VLMs as end-to-end systems that process driving images and other inputs to directly predict trajectories . However, unlike end-to-end models which are specifically designed for trajectory planning, VLMs operate in a language space and are not inherently suited for precise numerical predictions . Consequently, directly employing VLMs for trajectory planning may result in suboptimal performance and even pose safety risks.

Some studies leverage VLMs for high-level planning by formulating the ego vehicle’s future actions in natural language, such as “slow down and turn right” . Although this approach circumvents the aforementioned drawbacks, existing works still lack further exploration of training methodologies. Most of them primarily rely on SFT, overlooking the impact of different training strategies on planning performance and the associated training costs.

In this paper, we explore the following question: How can RL and reasoning — which achieves remarkable success in general large models — be applied to autonomous driving, particularly in planning, to enhance the performance of VLMs in autonomous driving while reducing training costs?

Through preliminary experiments, we find that directly applying existing RL and reasoning techniques to planning results in suboptimal performance. We attribute this to three main factors. First, the reward design in RL for general tasks is not well-suited for planning. For example, in visual object counting, the reward can be simply determined based on whether the model predicts the correct answer. However, in autonomous driving, while high-level planning can be formulated as a multi-class classification problem, the varying significance of different driving behaviors makes it inappropriate to assign equal weights to all actions.

Second, unlike mathematical or counting, the solution of planning are usually not unique. For instance, on an open, straight road, one may choose to maintain a constant speed or accelerate, both of which are valid decisions. Therefore, rigidly assessing whether the model’s planning output exactly matches the ground truth in the training data may not be the optimal approach.

Finally, while domains such as mathematics have abundant reasoning data, including textbooks and solution manuals that can be easily utilized, autonomous driving lacks readily available datasets that capture the reasoning process. Collecting such data is highly costly and requires extensive manual annotation. As a result, directly applying existing reasoning techniques to planning remains challenging.

To address the aforementioned challenges, this paper introduces AlphaDrive, a VLM-based reinforcement learning and reasoning framework specifically designed for autonomous driving planning. In particular, AlphaDrive employs a RL strategy based on Group Relative Policy Optimization (GRPO) . Compared to Proximal Policy Optimization (PPO) and Direct Preference Optimization (DPO) , GRPO exhibits better training stability and performance. Furthermore, the group relative optimization strategy in GRPO is well-suited for planning, as planning often involves multiple valid solutions, making relative optimization across multiple solutions a natural fit. Our experiments show that AlphaDrive exhibits some emergent multimodal planning capabilities, which we think can be attributed to the use of GRPO.

AlphaDrive introduces four GRPO rewards tailored for planning. The first is the planning accuracy reward, which evaluates the consistency between the model’s planning actions and the ground truth actions. The second is the action-weighted reward, which assigns different weights to various actions based on their importance to safety. For instance, actions such as braking and steering are critical for safety, so weighting them accordingly helps the model achieve better performance in planning key actions. The third is the planning diversity reward, which encourages the model to generate multiple diverse solutions. This prevents mode collapse and enhances overall planning performance. The last one is the planning format reward, where we define a specific output format and encourage the model to follow it. This ensures more structured outputs and contributes to more stable training.

In addition to RL, we propose a planning reasoning technique. Our approach employs a two-stage training strategy based on knowledge distillation, integrating SFT and RL. In the first stage, we leverage a large model, such as GPT-4o, to generate a small yet high-quality dataset containing planning reasoning processes derived from real driving actions. This dataset is then used to fine-tune our model via SFT, effectively distilling knowledge from the large model. In the second stage, we further refine the model using RL. Introducing the SFT stage as a warm-up step effectively mitigates hallucinations and instability commonly observed in the early stages of reinforcement learning, while also enhancing planning performance.

Our contributions are summarized as follows:

We propose AlphaDrive, a VLM tailored for high-level planning in autonomous driving. To the best of our knowledge, AlphaDrive is the first to integrate GRPO-based RL with planning reasoning to autonomous driving, significantly boosting both performance and training efficiency.

AlphaDrive introduces four GRPO rewards for planning: planning accuracy reward, action-weighted reward, planning diversity reward, and planning format reward. These optimized rewards make GRPO more suitable for autonomous driving.

We propose a two-stage reasoning training strategy based on knowledge distillation, integrating SFT and RL. Our approach achieves better planning performance compared to training with RL alone or without reasoning.

Experiments on a large-scale driving dataset validate the superiority of AlphaDrive. Compared to the SFT-trained model, AlphaDrive significantly improves the planning accuracy by 25.52% and, with only 20% of the training data, outperforms the SFT-trained model by 35.31%. We are also excited to discover that, following RL training, AlphaDrive exhibits some emergent multimodal planning capabilities, which is promising for improving driving safety and efficiency.

Related Work

Vision Language Models. Since the release of GPT , the capabilities of large models have gradually expanded from single modality to multi-modalities. Large vision language models now demonstrate superior abilities in visual understanding and reasoning. Early works attempt to integrate visual models with large language models (LLMs), Flamingo uses a visual encoder to process visual signals and adds attention layers in the LLM decoder to interact with the visual features. BLIP introduces the Q-Former architecture and cross-modal contrastive learning tasks to bridge the vision encoder with LLMs. LLaVA propose using vanilla MLP as the connector between the visual encoder and LLMs, which achieves impressive visual understanding capabilities with relatively limited data. The QwenVL series continuously improve the visual module, offering better support for high-resolution and dynamic resolution images, while also demonstrate excellent performance in multilingual tasks and spatial perception.

Reinforcement Learning and Reasoning. Autoregressive learning is currently the mainstream pre-training strategy for LLMs. Besides, RL and reasoning techniques further enhance the capabilities of large models . For instance, GPT employs RL with Human Feedback (RLHF) , which incorporates human feedback into the training process. By integrating human intentions and behavioral preferences, RLHF enables LLMs to generate outputs that align more closely with human habits and preferences. Direct Preference Optimization (DPO) enhances the model’s performance by directly optimizing preference feedback. Building on this, Group Relative Policy Optimization (GRPO) introduces a strategy of group relative optimization, which considers the relative superiority or inferiority between multiple output groups, further improving the stability and effectiveness of the training process.

The recent DeepSeek R1 experiences an “Aha Moment”during training based on GRPO, where, without any explicit guidance, the model autonomously allocates more thinking to the problem and re-evaluates its initial approach. This highlights the potential of RL in enabling large models to evolve from mere imitation to emergent intelligence. In our experients, we are also excited to discover that, after GRPO-based RL training, AlphaDrive demonstrates some emergent multimodal planning capabilities, enabling it to generate multiple reasonable driving plans. We believe it has great potential to improve driving safety and efficiency.

In terms of reasoning, Chain-of-thought has demonstrated great performance in solving complex problems by breaking them down and reasoning step by step. OpenAI o1 , which is based on Chain-of-thought, introduces inference-time scaling. By increasing the computational cost during inference and combining search strategies such as Monte Carlo Tree Search (MCTS) and Beam Search , significant improvements have been achieved in areas such as science and programming that require complex reasoning. This also shows that, beyond scaling model parameters and training data, scaling the inference-time computation is also a promising direction for exploration.

Autonomous Driving Planning. Planning is the ultimate task of autonomous driving. The earliest planning algorithms are rule-based , which have significant limitations in terms of generalizability and efficiency. Recently, end-to-end models has gained popularity, where a unified neural network is used to directly output planning trajectories or control signals from sensor data. By leveraging large-scale driving demonstrations, end-to-end models are trained in a data-driven manner, achieving impressive planning performance. However, since end-to-end models are black-box models that lack common-sense and reasoning capabilities, they still struggle to address the long-tailed problems in autonomous driving.

VLMs and Autonomous Driving. The common-sense and reasoning abilities of large models can effectively compensate for the limitations of end-to-end models in autonomous driving. In the field of robotics, Vision-Language-Action (VLA) models have made significant progress in understanding language instructions and executing complex actions. A common approach is to use VLMs as the planning module to generate planning instructions, which are then translated into control signals through an action model. There have also been some works based on large models in the field of autonomous driving. DriveGPT4 utilizes a VLM that takes front-view videos as input, and the model directly predicts control signals. ELM leverages large-scale, cross-domain video training for VLMs, showing that using data from various domains can effectively enhance the performance of VLMs in driving-related tasks. OmniDrive proposes the use of sparse 3D tokens to represent driving scenes, which are then input into VLMs for scene understanding and planning.

In addition to the above works that directly apply VLMs to driving, DriveVLM combines VLMs with end-to-end models for the first time, where VLMs predict low-frequency trajectories and an end-to-end model generates high-frequency trajectories. Senna proposes a framework where VLMs handle high-level planning, while end-to-end models are responsible for low-level trajectory prediction. Additionally, several datasets and benchmarks have been proposed , which promote the application of VLMs in autonomous driving. However, most of the current works on VLMs in the field of autonomous driving involves directly using pre-trained models and then utilizing SFT on driving data, which lacks in-depth exploration on training strategies specifically designed for planning. Further effort is needed to adapt the impressive RL and reasoning techniques from general tasks to autonomous driving.

AlphaDrive

AlphaDrive is a VLM designed for autonomous driving planning. Unlike previous approaches that rely solely on SFT, we explore the incorporation of RL and reasoning techniques to better align with the unique characteristics of driving planning: (1) the varying importance of different driving behaviors; (2) the existence of multiple feasible solutions; and (3) the scarcity of readily available reasoning data for planning decisions.

We propose four GRPO-based RL rewards tailored for planning, along with a two-stage planning-reasoning training strategy that integrates SFT with RL. Our experiments demonstrate that, compared to using SFT alone or training without reasoning, AlphaDrive achieves significant improvements in both planning performance and training efficiency. In the following sections, we will detail the design of each component.

2 Planning-oriented Reinforcement Learning

Current commonly used RL algorithms include PPO , DPO , and GRPO . Given a query qq, GRPO samples a group of outputs {o1,o2,⋯ ,oG}\{o_{1},o_{2},\cdots,o_{G}\} from the old policy πθold\pi_{\theta_{old}} and optimizes the new policy πθ\pi_{\theta} by maximizing:

where wi=πθ(oi∣q)πθold(oi∣q)w_{i}=\frac{\pi_{\theta}(o_{i}|q)}{\pi_{\theta_{old}}(o_{i}|q)}, ϵ\epsilon and β\beta are hyper-parameters, and the advantage AiA_{i} is computed using the normalized reward within the group.

We ultimately choose GRPO as the RL algorithm for AlphaDrive for two key reasons: (1) DeepSeek R1 has demonstrated the effectiveness of GRPO in general domains. Compared to other algorithms, GRPO provides higher training stability and efficiency; (2) Moreover, the group relative optimization strategy introduced by GRPO is particularly well-suited for planning, as planning often involves multiple valid solutions, making relative optimization across multiple solutions is a natural fit. Experimental results further confirm that models trained with GRPO exhibit strong planning capabilities.

2.2 Planning Reward Modeling

Planning Accuracy Reward. In fields such as mathematics or programming, the reward in GRPO can be intuitively determined based on whether the final answer is correct. However, planning is more complex, as it involves both lateral (direction) and longitudinal (speed) components. Furthermore, the set of possible actions is constrained. As a result, we use the F1-Score to evaluate the accuracy of both lateral and longitudinal decisions separately, and assign rewards accordingly.

Initially, we evaluate accuracy by checking whether the model’s prediction exactly matches the ground truth. However, due to imperfect format in the model’s early training phase, such as discrepancies in case sensitivity or the presence of extraneous outputs, this approach results in poor stability during the early stages of training. We then attempt to extract all the words from the prediction and check whether the ground truth is included among the words. This introduces a new issue where the model sometimes learns shortcut solutions, such as outputting all possible actions, which causes mode collapse. Ultimately, we adopt the F1-score for evaluation, as it not only prevents the model from learning shortcut solutions (where outputting all decisions could result in high recall but low accuracy) but also improves the stability during the early training phase.

Action-Weighted Reward. As mentioned above, the importance of different behaviors in planning varies. For instance, decelerating and stopping are more critical for safety than maintaining speed. Therefore, we assign different importance weights to various actions, incorporating them as weighted components in the final reward.

Planning Diversity Reward. Since planning is inherently multimodal, during GRPO-based RL training, the model generates multiple solutions for group relative optimization. In the later stages of training, we observe that the model’s outputs tend to converge to the same solution. Our goal is to encourage the model to generate a variety of feasible solutions, rather than merely aligning with the ground truth actions in the training data. To achieve this, we propose the Planning Diversity Reward. When the model’s outputs differ, we assign a higher reward; otherwise, we reduce the reward.

Planning Format Reward. The last reward is used to regularize the output, making it easier to extract both the reasoning process and the final answer. This approach is inspired by R1. The reasoning process is encapsulated within the tags, while the planning result is enclosed within the tags. If the final output does not conform to this format, the format reward will be set to 0.

The Planning Accuracy Reward, the Action-Weighted Reward, and the Planning Diversity Reward are multiplied to compute the Planning Quality Reward. We calculate the Planning Quality Reward separately for speed planning and direction planning. Finally, the Planning Quality Reward and the Planning Format Reward are used to calculate the GRPO loss and update the model parameters. For details about Planning Reward Modeling, please refer to Alg. 1.

3 Reasoning: Distillation from Large Models

Unlike fields such as mathematics or science, which have abundant high-quality reasoning data available for training, the planning process in autonomous driving is difficult to record, and the cost of manual annotation is high. As a result, there is currently no large-scale, readily available planning reasoning dataset. We initially attempt to incorporate reasoning steps directly into the RL training process, but the final results are suboptimal, mainly due to the following shortcomings: (1) insufficient perception of key elements, such as traffic lights; (2) disorganized reasoning process with weak causal relationships; (3) reasoning outputs that are overly lengthy and ineffective.

Therefore, we adopt a more capable cloud-based large model, such as GPT-4o, to generate high-quality planning reasoning data from a small set of driving clips. Specifically, we provide the model with prompts that include the real driving actions in a given scenario, along with the vehicle’s current state and navigation information, prompting the model to generate a concise decision-making process. We find that the quality of the generated reasoning process is pretty good. After conducting a manual quality check and filtering out samples with obvious errors, we obtain a batch of high-quality planning reasoning data. Subsequently, our model can improve its planning reasoning ability through knowledge distillation based on this data.

4 Training: SFT Warm-Up, RL Exploration

RL relies on sparse reward signals, whereas SFT is based on dense supervision, making it more suitable for knowledge distillation. Additionally, we find that relying solely on RL can lead to instability in the early stages of training. Therefore, we use a small amount of data for a warm-up phase based on SFT, followed by RL training with the full dataset. We discover that this approach improves stability in the early stages of training and enhances the model’s planning reasoning performance, ultimately leading to better overall planning capabilities.

Experiments

Dataset. We adopt MetaAD, a large-scale real-world driving dataset, as our training and evaluation benchmark. This dataset consists of a total of 120k driving clips, each lasting three seconds. MetaAD is a high-quality dataset specifically designed for planning, supporting multi-sensor data and perception annotations. Furthermore, it maintains a well-balanced distribution across various driving environments and planning actions. The dataset is divided into 110k clips for training and 10k clips for validation. As for reasoning, we sample 30k data from the training dataset to generate the planning reasoning process. All reported results are obtained by training on the training set and evaluating on the validation set.

Training. We use Qwen2VL-2B as the base model. Qwen2VL is currently one of the best-performing open-source models, and it offers a smaller 2B version that better meets the latency requirements for autonomous driving. Additionally, Qwen2VL provides better support for RL. The model’s inputs include a front-view image and a planning prompt, which contains the vehicle’s current speed and navigation information. The navigation data, consistent with real-world driving, is obtained from sparse navigation points via AMap (similar to Google Maps) and is converted into text form for inclusion in the prompt, such as “Go straight for 100m, then turn right”. Training is conducted using 16 NVIDIA A800 GPUs.

Evaluation. The evaluation metrics consist of two aspects. First, the accuracy of meta-action planning is measured by calculating the F1-Score for all categories of lateral and longitudinal meta-actions, followed by the overall planning accuracy. Additionally, for planning reasoning, we compute the similarity between the generated planning reasoning process and the annotated reasoning process in the dataset using BLEU-4 , CIDEr , and METEOR scores.

2 Main Results

Tab. 1 presents the performance of AlphaDrive in high-level planning. The first four rows show the results obtained by directly evaluating the corresponding pretrained models. It can be observed that, while these models demonstrate stronger general capabilities, their performance in planning is suboptimal, highlighting the need for further training with driving data. The subsequent five rows display the results of models fine-tuned on the MetaAD dataset. As shown, AlphaDrive significantly outperforms the other models. Compared to Qwen2VL-7B, the second-best performing model after AlphaDrive, the planning accuracy significantly improves by 25.5%. There is a noticeable enhancement in key decisions such as steering and acceleration/deceleration. Additionally, the quality of planning reasoning is the best among all models, demonstrating the effectiveness of our proposed two-stage RL training and reasoning strategies.

3 Ablation Study

Planning Rewards. In Tab. 2, we validate the effectiveness of the four proposed GRPO planning rewards. The Base Accuracy reward directly determines the reward based on whether the response exactly matches the ground truth, a common approach in general domains. As shown, the model using the Base Accuracy reward lags significantly behind across all metrics (ID 1). The combination with the Planning Format Reward yields a slight improvement. (ID 2). A significant improvement is seen with the adoption of our proposed Planning Accuracy Reward (ID 3). Further enhancement in acceleration/deceleration decisions is achieved by incorporating the Action-Weighted Reward (ID 4). Finally, by combining the Planning Diversity Reward, the best planning performance is achieved (ID 5-6).

Reasoning Training Strategies. The ablation study of the reasoning training strategies is shown in Tab. 3. As observed, introducing planning reasoning under different training strategies effectively enhances model performance. Notably, the improvement is especially significant for complex actions such as acceleration and deceleration, demonstrating that reasoning can greatly enhance decision-making in complex scenarios. Furthermore, the model trained exclusively with RL performs worse in reasoning compared to the model trained with SFT. We attribute this to the limited parameter size of smaller models, which results in insufficient perception and reasoning capabilities. Therefore, incorporating SFT as a warm-up phase and using knowledge distillation to learn the reasoning process from a larger model can effectively address this issue. By combining SFT and RL, the model achieves the best planning reasoning capabilities.

Amount of Training Data. Tab. 4 shows the impact of training data size on different training strategies. As observed, when the training data size decreases, SFT is more affected. With only 20k training samples, the model trained with RL reaches a planning accuracy of 46.08%, which is significantly higher than that of the SFT-trained model. When using nearly half of the data, with 50k samples, AlphaDrive already achieves a planning accuracy of 70.83%, demonstrating the efficiency of our training strategy.

4 Emergence of Multimodal Planning Capability

Fig. 3 illustrates the multimodal planning capability of AlphaDrive after RL training. In complex scenarios, it can effectively generate multiple feasible solutions, whereas the SFT-trained model can only produce a single planning decision. AlphaDrive can be integrated with a downstream action model to dynamically select the optimal solution from multiple options.

Conclusions and Limitations

In this work, we propose AlphaDrive, a VLM for high-level planning in autonomous driving. Compared to previous models that solely employed the SFT, we explore the integration of advanced RL and reasoning in planning. Specifically, AlphaDrive introduces a planning-oriented RL strategy based on GRPO and further designs a two-stage planning reasoning training paradigm. To the best of our knowledge, AlphaDrive is the first to introduce the RL and reasoning to autonomous driving planning, significantly boosting both performance and training efficiency.

Currently, due to a lack of rich data annotation, AlphaDrive is still unable to output more complex driving behaviors such as lane changes or nudges. Additionally, the current planning reasoning data come from pseudo-labels generated by large models based on ground-truth driving actions, which still suffer from inaccurate perception and a failure to capture key factors. Therefore, further systematic validation is required to improve data quality and verify the performance upper bound of AlphaDrive.

Acknowledgments

We sincerely thank Hao Gao, Tianheng Cheng, Bencheng Liao, Haoyi Jiang, and Dongli Hu for their valuable feedback on the draft.

References