RAD: Training an End-to-End Driving Policy via Large-Scale 3DGS-based Reinforcement Learning

Hao Gao, Shaoyu Chen, Bo Jiang, Bencheng Liao, Yiang Shi, Xiaoyang Guo, Yuechuan Pu, Haoran Yin, Xiangyu Li, Xinbang Zhang, Ying Zhang, Wenyu Liu, Qian Zhang, Xinggang Wang

Introduction

End-to-end autonomous driving (AD) is currently a trending topic in both academia and industry. It replaces a modularized pipeline with a holistic one by directly mapping sensory inputs to driving actions, offering advantages in system simplicity and generalization ability. Most existing end-to-end AD algorithms follow the Imitation Learning (IL) paradigm, which trains a neural network to mimic human driving behavior. However, despite their simplicity, IL-based methods face significant challenges in real-world deployment.

One key issue is causal confusion. IL trains a network to replicate human driving policies by learning from demonstrations. However, this paradigm primarily captures correlations rather than causal relationships between observations (states) and actions. As a result, IL-trained policies may struggle to identify the true causal factors behind planning decisions, leading to shortcut learning , e.g., merely extrapolating future trajectories from historical ones . Furthermore, since IL training data predominantly consists of common driving behaviors and does not adequately cover long-tailed distributions, IL-trained policies tend to converge to trivial solutions, lacking sufficient sensitivity to safety-critical events such as collisions.

Another major challenge is the gap between open-loop training and closed-loop deployment. IL policies are trained in an open-loop manner using well-distributed driving demonstrations. However, real-world driving is a closed-loop process where minor trajectory errors at each step accumulate over time, leading to compounding errors and out-of-distribution scenarios. IL-trained policies often struggle in these unseen situations, raising concerns about their robustness.

A straightforward solution to these problems is to perform closed-loop Reinforcement Learning (RL) training, which requires a driving environment that can interact with the AD policy. However, using real-world driving environments for closed-loop training poses prohibitive safety risks and operational costs. Simulated driving environments with sensor data simulation capabilities (which are required for end-to-end AD) are typically built on game engines but fail to provide realistic sensor simulation results.

In this work, we establish a 3DGS-based closed-loop RL training paradigm. Leveraging 3DGS techniques, we construct a photorealistic digital replica of the real world, where the AD policy can extensively explore the state space and learn to handle out-of-distribution situations through large-scale trial and error. To ensure effective responses to safety-critical events and a better understanding of real-world causations, we design specialized safety-related rewards. However, RL training presents several critical challenges, which this paper addresses.

One significant challenge is the Human Alignment Problem. The exploration process in RL can lead to policies that deviate from human-like behavior, disrupting the smoothness of the action sequence. To address this, we incorporate imitation learning as a regularization term during RL training, helping to maintain similarity to human driving behavior. As illustrated in Fig. 1, RL and IL work together to optimize the AD policy: RL enhances IL by addressing causation and the open-loop gap, while IL improves RL by ensuring better human alignment.

Another major challenge is the Sparse Reward Problem. RL often suffers from sparse rewards and slow convergence. To alleviate this issue, we introduce dense auxiliary objectives related to collisions and deviations, which help constrain the full action distribution. Additionally, we streamline and decouple the action space to reduce the exploration cost associated with RL.

To validate the effectiveness of our approach, we construct a closed-loop evaluation benchmark comprising diverse, unseen 3DGS environments. Our method, RAD, outperforms IL-based approaches across most closed-loop metrics, notably achieving a collision rate that is 3×3\times lower.

The contributions of this work are summarized as follows:

We propose the first 3DGS-based RL framework for training end-to-end AD policy. The reward, action space, optimization objective, and interaction mechanism are specially designed to enhance training efficiency and effectiveness.

We combine RL and IL to synergistically optimize the AD policy. RL complements IL by modeling the causations and narrowing the open-loop gap, while IL complements RL in terms of human alignment.

We validate the effectiveness of RAD on a closed-loop evaluation benchmark consisting of diverse, unseen 3DGS environments. RAD achieves stronger performance in closed-loop evaluation, particularly a 3×3\times lower collision rate, compared to IL-based methods.

Related Work

Dynamic Scene Reconstruction. Implicit neural representations have dominated novel view synthesis and dynamic scene reconstruction, with methods like UniSim , MARS , and NeuRAD leveraging neural scene graphs to enable structured scene decomposition. However, these approaches rely on implicit representations, leading to slow rendering speeds that limit their practicality in real-time applications. In contrast, 3D Gaussian Splatting (3DGS) has emerged as an efficient alternative, offering significantly faster rendering while maintaining high visual fidelity. Recent works have explored its potential for dynamic scene reconstruction, particularly in autonomous driving scenarios. StreetGaussians , DrivingGaussians , and HUGSIM have demonstrated the effectiveness of Gaussian-based representations in modeling urban environments. These methods achieve superior rendering performance while maintaining controllability by explicitly decomposing scenes into structured components. However, these works primarily leverage 3DGS for closed-loop evaluation. In this work, we incorporate 3DGS into an RL training framework.

End-to-End Autonomous Driving. Learning-based planning has shown great potential recently due to its data-driven nature and impressive performance with increasing amounts of data. UniAD demonstrates the potential of end-to-end autonomous driving by integrating multiple perception tasks to enhance planning performance. VAD further explores the use of compact vectorized scene representations to improve efficiency. A series of works also adopt the single-trajectory planning paradigm and further enhance planning performance. VADv2 shifts the paradigm towards multi-mode planning by modeling the probabilistic distribution of a planning vocabulary. Hydra-MDP improves the scoring mechanism of VADv2 by introducing additional supervision from a rule-based scorer. SparseDrive explores an alternative BEV-free solution. DiffusionDrive proposes a truncated diffusion policy that denoises an anchored Gaussian distribution to a multi-mode driving action distribution. Most of the end-to-end methods follow the data-driven IL-training paradigm. In this work, we present an RL-training paradigm based on 3DGS.

Reinforcement Learning. Reinforcement Learning is a promising technique that has not been fully explored. AlphaGo and AlphaGo Zero have demonstrated the power of Reinforcement Learning in the game of Go. Recently, OpenAI O1 and Deepseek-R1 have leveraged Reinforcement Learning to develop reasoning abilities. Several studies have also applied Reinforcement Learning in autonomous driving . However, these studies are based on non-photorealistic simulators (such as CARLA ) or do not involve end-to-end driving algorithms, as they require perfect perception results as input. To the best of our knowledge, RAD is the first work to train an end-to-end AD agent using Reinforcement Learning in a photorealistic 3DGS environment.

RAD

The overall framework of RAD is depicted in Fig. 2. RAD takes multi-view image sequences as input, transforms the sensor data into scene token embeddings, outputs the probabilistic distribution of actions, and samples an action to control the vehicle.

BEV Encoder. We first employ a BEV encoder to transform multi-view image features from the perspective view to the Bird’s Eye View (BEV), obtaining a feature map in the BEV space. This feature map is then used to learn instance-level map features and agent features.

Map Head. Then we utilize a group of map tokens to learn the vectorized map elements of the driving scene from the BEV feature map, including lane centerlines, lane dividers, road boundaries, arrows, traffic signals, etc.

Agent Head. Besides, a group of agent tokens is adopted to predict the motion information of other traffic participants, including location, orientation, size, speed, and multi-mode future trajectories.

Image Encoder. Apart from the above instance-level map and agent tokens, we also use an individual image encoder to transform the original images into image tokens. These image tokens provide dense and rich scene information for planning, complementary to the instance-level tokens.

Action Space. To accelerate the convergence of RL training, we design a decoupled discrete action representation. We divide the action into two independent components: lateral action and longitudinal action. The action space is constructed over a short 0.50.5-second time horizon, during which the vehicle’s motion is approximated by assuming constant linear and angular velocities. Under this assumption, the lateral action axa^{x} and longitudinal action aya^{y} can be directly computed based on the current linear and angular velocities. By combining decoupling with a limited temporal scope and simplified motion model, our approach effectively reduces the dimensionality of the action space, accelerating training convergence.

Planning Head. We use EsceneE_{\text{scene}} to denote the scene representation, which consists of map tokens, agent tokens, and image tokens. We initialize a planning embedding denoted as EplanE_{\text{plan}}. A cascaded Transformer decoder ϕ\phi takes the planning embedding EplanE_{\text{plan}} as the query and the scene representation EsceneE_{\text{scene}} as both key and value.

The output of the decoder ϕ\phi is then combined with navigation information EnaviE_{\text{navi}} and ego state EstateE_{\text{state}} to output the probabilistic distributions of the lateral action axa^{x} and the longitudinal action aya^{y}:

where EplanE_{\text{plan}}, EnaviE_{\text{navi}}, EstateE_{\text{state}}, and the output of MLP are all of the same dimension (1×D1\times D).

The planning head also outputs the value functions Vx(s)V_{x}(s) and Vy(s)V_{y}(s), which estimate the expected cumulative rewards for the lateral and longitudinal actions, respectively:

The value functions are used in RL training (Sec. 3.5).

2 Training Paradigm

We adopt a three-stage training paradigm: perception pre-training, planning pre-training, and reinforced post-training, as shown in Fig. 2.

Perception Pre-Training. Information in the image is sparse and low-level. In the first stage, the map head and the agent head explicitly output map elements and agent motion information, which are supervised with ground-truth labels. Consequently, map tokens and agent tokens implicitly encode the corresponding high-level information. In this stage, we only update the parameters of the BEV encoder, the map head, and the agent head.

Planning Pre-Training. In the second stage, to prevent the unstable cold start of RL training, IL is first performed to initialize the probabilistic distribution of actions based on large-scale real-world driving demonstrations from expert drivers. In this stage, we only update the parameters of the image encoder and the planning head, while the parameters of the BEV encoder, map head, and agent head are frozen. The optimization objectives of perception tasks and planning tasks may conflict with each other. However, with the training stage and parameters decoupled, such conflicts are mostly avoided.

Reinforced Post-Training. In the reinforced post-training, RL and IL synergistically fine-tune the distribution. RL aims to guide the policy to be sensitive to critical risky events and adaptive to out-of-distribution situations. IL serves as the regularization term to keep the policy’s behavior similar to that of humans.

We select a large amount of risky dense-traffic clips from collected driving demonstrations. For each clip, we train an independent 3DGS model that reconstructs the clip and serves as a digital driving environment. As shown in Fig. 3, we set NN parallel workers. Each worker randomly samples a 3DGS environment and begins rollout, i.e., the AD policy controls the ego vehicle to move and iteratively interacts with the 3DGS environment. After the rollout process of this 3DGS environment ends, the generated rollout data (st,at,rt+1,st+1,...)(s_{t},a_{t},r_{t+1},s_{t+1},...) are recorded in a rollout buffer, and the worker will sample a new 3DGS environment for another round of rollout.

As for policy optimization, we iteratively perform RL-training steps and IL-training steps. For RL-training steps, we sample data from the rollout buffer and follow the Proximal Policy Optimization (PPO) framework to update the AD policy. For IL-training steps, we use real-world driving demonstrations to update the policy. After a fixed number of training steps, the updated AD policy is sent to every worker to replace the old one, to avoid a distribution shift between data collection and optimization. We only update the parameters of the image encoder and the planning head. The parameters of the BEV encoder, the map head, and the agent head are frozen. The detailed RL design is presented below.

3 Interaction Mechanism between AD Policy and 3DGS Environment

In the 3DGS environment, the ego vehicle acts according to the AD policy. Other traffic participants act according to real-world data in a log-replay manner. A simplified kinematic bicycle model is employed to iteratively update the ego vehicle’s pose at every Δt\Delta t seconds as follows:

where xtwx_{t}^{w} and ytwy_{t}^{w} denote the position of the ego vehicle relative to the world coordinate; ψtw\psi_{t}^{w} is the heading angle that defines the vehicle’s orientation with respect to the world xx-coordinate; vtv_{t} is the linear velocity of the ego vehicle; δt\delta_{t} is the steering angle of the front wheels; and LL is the wheelbase, i.e., the distance between the front and rear axles.

During the rollout process, the AD policy outputs actions (atx,aty)(a_{t}^{x},a_{t}^{y}) for a 0.50.5-second time horizon at time step tt. We derive the linear velocity vtv_{t} and steering angle δt\delta_{t} based on (atx,aty)(a_{t}^{x},a_{t}^{y}). Based on the kinematic model in Eq. 3, the pose of the ego vehicle in the world coordinate system is updated from pt=(xtw,ytw,ψtw){p}_{t}=(x_{t}^{w},y_{t}^{w},\psi_{t}^{w}) to pt+1=(xt+1w,yt+1w,ψt+1w){p}_{t+1}=(x_{t+1}^{w},y_{t+1}^{w},\psi_{t+1}^{w}).

Based on the updated pt+1{p}_{t+1}, the 3DGS environment computes the new ego vehicle’s state st+1s_{t+1}. The updated pose pt+1{p}_{t+1} and state st+1s_{t+1} serve as the input for the next iteration of the inference process.

The 3DGS environment also generates rewards R\mathcal{R} (Sec. 3.4) according to multi-source information (including trajectories of other agents, map information, the expert trajectory of the ego vehicle, and the parameters of Gaussians), which are used to optimize the AD policy (Sec. 3.5).

4 Reward Modeling

The reward is the source of the training signal, which determines the optimization direction of RL. The reward function is designed to guide the ego vehicle’s behavior by penalizing unsafe actions and encouraging alignment with the expert trajectory. It is composed of four reward components: (1) collision with dynamic obstacles, (2) collision with static obstacles, (3) positional deviation from the expert trajectory, and (4) heading deviation from the expert trajectory:

As illustrated in Fig. 4, these reward components are triggered under specific conditions. In the 3DGS environment, dynamic collision is detected if the ego vehicle’s bounding box overlaps with the annotated bounding boxes of dynamic obstacles, triggering a negative reward rdcr_{\text{dc}}. Similarly, static collision is identified when the ego vehicle’s bounding box overlaps with the Gaussians of static obstacles, resulting in a negative reward rscr_{\text{sc}}. Positional deviation is measured as the Euclidean distance between the ego vehicle’s current position and the closest point on the expert trajectory. A deviation beyond a predefined threshold dmaxd_{\text{max}} incurs a negative reward rpdr_{\text{pd}}. Heading deviation is calculated as the angular difference between the ego vehicle’s current heading angle ψt\psi_{t} and the expert trajectory’s matched heading angle ψexpert\psi_{\text{expert}}. A deviation beyond a threshold ψmax\psi_{\text{max}} results in a negative reward rhdr_{\text{hd}}.

Any of these events, including dynamic collision, static collision, excessive positional deviation, or excessive heading deviation, triggers immediate episode termination. Because after such events occur, the 3DGS environment typically generates noisy sensor data, which is detrimental to RL training.

5 Policy Optimization

In the closed-loop environment, the error in each single step accumulates over time. The aforementioned rewards are not only caused by the current action but also by the actions of the preceding steps. The rewards are propagated forward with Generalized Advantage Estimation (GAE) to optimize the action distribution of the preceding steps.

Specifically, for each time step tt, we store the current state sts_{t}, action ata_{t}, reward rtr_{t}, and the estimate of the value V(st)V(s_{t}). Based on the decoupled action space, and considering that different rewards have different correlations to lateral and longitudinal actions, the reward rtr_{t} is divided into lateral reward rtxr_{t}^{x} and longitudinal reward rtyr_{t}^{y}:

Similarly, the value function V(st)V(s_{t}) is decoupled into two components: Vx(st)V_{x}(s_{t}) for the lateral dimension and Vy(st)V_{y}(s_{t}) for the longitudinal dimension. These value functions estimate the expected cumulative rewards for the lateral and longitudinal actions, respectively. The advantage estimates A^tx\hat{A}_{t}^{x} and A^ty\hat{A}_{t}^{y} are then computed as follows:

where δtx\delta_{t}^{x} and δty\delta_{t}^{y} are the temporal difference errors for the lateral and longitudinal dimensions, γ\gamma is the discount factor, and λ\lambda is the GAE parameter that controls the trade-off between bias and variance.

To further clarify the relationship between the advantage estimates and the reward components, we decompose A^tx\hat{A}_{t}^{x} and A^ty\hat{A}_{t}^{y} based on the reward decomposition in Eq. 5 and the advantage estimation in Eq. 6. Specifically, we derive the following decomposition:

where A^tsc\hat{A}_{t}^{\text{sc}} is the advantage estimate for avoiding static collisions, A^tpd\hat{A}_{t}^{\text{pd}} is the advantage estimate for minimizing positional deviations, A^thd\hat{A}_{t}^{\text{hd}} is the advantage estimate for minimizing heading deviations, and A^tdc\hat{A}_{t}^{\text{dc}} is the advantage estimate for avoiding dynamic collisions.

These advantage estimates are used to guide the update of the AD policy πθ\pi_{\theta}, following the PPO framework . By leveraging the decomposed advantage estimates A^tx\hat{A}_{t}^{x} and A^ty\hat{A}_{t}^{y}, we can independently optimize the lateral and longitudinal dimensions of the policy. This is achieved by defining separate objective functions LxCLIP(θ)\mathcal{L}_{x}^{\text{CLIP}}(\theta) and LyCLIP(θ)\mathcal{L}_{y}^{\text{CLIP}}(\theta) for each dimension, as follows:

where ρtx=πθ(atx∣st)πθold(atx∣st)\rho_{t}^{x}=\frac{\pi_{\theta}(a_{t}^{x}\mid s_{t})}{\pi_{\theta_{\text{old}}}(a_{t}^{x}\mid s_{t})} is the importance sampling ratio for the lateral dimension, ρty=πθ(aty∣st)πθold(aty∣st)\rho_{t}^{y}=\frac{\pi_{\theta}(a_{t}^{y}\mid s_{t})}{\pi_{\theta_{\text{old}}}(a_{t}^{y}\mid s_{t})} is the importance sampling ratio for the longitudinal dimension, ϵx\epsilon_{x} and ϵy\epsilon_{y} are small constants that control the clipping range for the lateral and longitudinal dimensions, ensuring stable policy updates.

The clipped objective function LPPO(θ)\mathcal{L}^{\text{PPO}}(\theta) prevents excessively large updates to the policy parameters θ\theta, thereby maintaining training stability.

6 Auxiliary Objective

RL usually faces the challenge of sparse rewards, which makes the convergence process unstable and slow. To speed up convergence, we introduce auxiliary objectives that provide dense guidance to the entire action distribution.

The auxiliary objectives are designed to penalize undesirable behaviors by incorporating specific reward sources, including dynamic collisions, static collisions, positional deviations, and heading deviations. These objectives are computed based on the actions atx,olda_{t}^{x,\text{old}} and aty,olda_{t}^{y,\text{old}} selected by the old AD policy πθold\pi_{\theta_{\text{old}}} at time step tt. To facilitate the evaluation of these actions, we separate the probability distribution of the action into four parts:

Here, Δπydec\Delta\pi_{y}^{\text{dec}} represents the total probability of deceleration actions, Δπyacc\Delta\pi_{y}^{\text{acc}} represents the total probability of acceleration actions, Δπxleft\Delta\pi_{x}^{\text{left}} represents the total probability of leftward steering actions, and Δπxright\Delta\pi_{x}^{\text{right}} represents the total probability of rightward steering actions.

Dynamic Collision Auxiliary Objective. The dynamic collision auxiliary objective adjusts the longitudinal control action atya_{t}^{y} based on the location of potential collisions relative to the ego vehicle. If a collision is detected ahead, the policy prioritizes deceleration actions (aty<aty,olda_{t}^{y}<a_{t}^{y,\text{old}}); if a collision is detected behind, it encourages acceleration actions (aty>aty,olda_{t}^{y}>a_{t}^{y,\text{old}}). To formalize this behavior, we define a directional factor fdcf_{\text{dc}}:

The auxiliary objective for dynamic collision avoidance is defined as:

where A^tdc\hat{A}_{t}^{\text{dc}} is the advantage estimate for dynamic collision avoidance.

Static Collision Auxiliary Objective. The static collision auxiliary objective adjusts the steering control action atxa_{t}^{x} based on the proximity to static obstacles. If the static obstacle is detected on the left side, the policy promotes rightward steering actions (atx>atx,olda_{t}^{x}>a_{t}^{x,\text{old}}); if the static obstacle is detected on the right side, it promotes leftward steering actions (atx<atx,olda_{t}^{x}<a_{t}^{x,\text{old}}). To formalize this behavior, we define a directional factor fscf_{\text{sc}}:

The auxiliary objective for static collision avoidance is defined as:

where A^tsc\hat{A}_{t}^{\text{sc}} is the advantage estimate for static collision avoidance.

Positional Deviation Auxiliary Objective. The positional deviation auxiliary objective adjusts the steering control action atxa_{t}^{x} based on the ego vehicle’s lateral deviation from the expert trajectory. If the ego vehicle deviates leftward, the policy promotes rightward corrections (atx>atx,olda_{t}^{x}>a_{t}^{x,\text{old}}); if it deviates rightward, it promotes leftward corrections (atx<atx,olda_{t}^{x}<a_{t}^{x,\text{old}}). We formalize this with a directional factor fpdf_{\text{pd}}:

The auxiliary objective for positional deviation correction is:

where A^tpd\hat{A}_{t}^{\text{pd}} estimates the advantage of trajectory alignment.

Heading Deviation Auxiliary Objective. The heading deviation auxiliary objective adjusts the steering control action atxa_{t}^{x} based on the angular difference between the ego vehicle’s current heading and the expert’s reference heading. If the ego vehicle deviates counterclockwise, the policy promotes clockwise corrections (atx>atx,olda_{t}^{x}>a_{t}^{x,\text{old}}); if it deviates clockwise, it promotes counterclockwise corrections (atx<atx,olda_{t}^{x}<a_{t}^{x,\text{old}}). To formalize this behavior, we define a directional factor fhdf_{\text{hd}}:

The auxiliary objective for heading deviation correction is then defined as:

where A^thd\hat{A}_{t}^{\text{hd}} is the advantage estimate for heading alignment.

Overall Auxiliary Objectives. The overall auxiliary objectives are a weighted sum of the individual objectives:

where λ1\lambda_{1}, λ2\lambda_{2}, λ3\lambda_{3}, and λ4\lambda_{4} are weighting coefficients that balance the contributions of each auxiliary objective.

Optimization Objective. The final optimization objective combines the clipped PPO objective with the auxiliary objective:

Experiments

Dataset and Benchmark. We collect 2000h2000h of expert human driving demonstrations in the real physical world. We get ground-truths of maps and agents in these driving demonstrations through a low-cost automated annotation pipeline. We use the map and agent labels as supervision for the first-stage perception pre-training. For the second-stage planning pre-training, we use the odometry information of the ego vehicle as supervision. For the third-stage reinforced post-training, we select 4305 critical dense-traffic clips of high collision risks from collected driving demonstrations and reconstruct these clips into 3DGS environments. Of these, 3968 3DGS environments are used for RL training, and the other 337 3DGS environments are used as closed-loop evaluation benchmarks.

Metric. We evaluate the performance of the AD policy using nine key metrics. Dynamic Collision Ratio (DCR) and Static Collision Ratio (SCR) quantify the frequency of collisions with dynamic and static obstacles, respectively, with their sum represented as the Collision Ratio (CR). Positional Deviation Ratio (PDR) measures the ego vehicle’s deviation from the expert trajectory with respect to position, while Heading Deviation Ratio (HDR) evaluates the ego vehicle’s consistency to the expert trajectory with respect to the forward direction. The overall deviation is quantified by the Deviation Ratio (DR), defined as the sum of PDR and HDR. Average Deviation Distance (ADD) quantifies the mean closest distance between the ego vehicle and the expert trajectory before any collisions or deviations occur. Additionally, Longitudinal Jerk (Long. Jerk) and Lateral Jerk (Lat. Jerk) assess driving smoothness by measuring acceleration changes in the longitudinal and lateral directions. CR, DCR, and SCR mainly reflect the policy’s safety, and ADD reflects the trajectory consistency between the AD policy and human drivers.

2 Ablation Study

To evaluate the impact of different design choices in RAD, we conduct three ablation studies. These studies examine the balance between RL and IL, the role of different reward sources, and the effect of auxiliary objectives.

RL-IL Ratio Analysis. We first analyze the effect of different RL-to-IL step mixing ratios (Tab. 1). A pure IL policy (0:1) results in the highest CR (0.229) but the lowest ADD (0.238), indicating strong trajectory consistency but poor safety. In contrast, a pure RL policy (1:0) significantly reduces CR (0.143) but increases ADD (0.345), suggesting improved safety at the cost of trajectory deviation. The best balance is achieved at a 4:1 ratio, which results in the lowest CR (0.089) while maintaining a relatively low ADD (0.257). Further increasing RL dominance (e.g., 8:1) leads to a deteriorated ADD (0.323) and higher jerk, implying reduced trajectory smoothness.

Reward Source Analysis. We analyze the influence of different reward components (Tab. 2). Policies trained with only partial reward terms (e.g., ID 1, 2, 3, 4, 5) exhibit higher collision rates (CR) compared to the full reward setup (ID 6), which achieves the lowest CR (0.089) while maintaining a stable ADD (0.257). This demonstrates that a well-balanced reward function, incorporating all reward terms, effectively enhances both safety and trajectory consistency. Among the partial reward configurations, ID 2, which omits the dynamic collision reward term, exhibits the highest CR (0.238), indicating that the absence of this term significantly impairs the model’s ability to avoid dynamic obstacles, resulting in a higher collision rate.

Auxiliary Objective Analysis. Finally, we examine the impact of auxiliary objectives (Tab. 3). Compared to the full auxiliary setup (ID 8), omitting any auxiliary objective increases CR, with a significant rise observed when all auxiliary objectives are removed. This highlights their collective role in enhancing safety. Notably, ID 1, which retains all auxiliary objectives but excludes the PPO objective, results in a CR of 0.187. This value is higher than that of ID 8, indicating that while auxiliary objectives help reduce collisions, they are most effective when combined with the PPO objective.

Our ablation studies highlight the importance of combining RL and IL, using a comprehensive reward function, and implementing structured auxiliary objectives. The optimal RL-IL ratio (4:1) and the full reward and auxiliary setups consistently yield the lowest CR while maintaining stable ADD, ensuring both safety and trajectory consistency.

3 Comparisons with Existing Methods

As presented in Tab. 4, we compare RAD with other end-to-end autonomous driving methods in the proposed 3DGS-based closed-loop evaluation. For fair comparisons, all the methods are trained with the same amount of human driving demonstrations. The 3DGS environments for the RL training in RAD are also based on these data. RAD achieves better performance compared to IL-based methods in most metrics. Especially in terms of CR, RAD achieves 3×3\times lower collision rate, demonstrating that RL helps the AD policy learn general collision avoidance ability.

4 Qualitative Comparisons

We provide qualitative comparisons between the IL-only AD policy (without reinforced post-training) and RAD, as shown in Fig. 5. The IL-only method struggles in dynamic environments, frequently failing to avoid collisions with moving obstacles or manage complex traffic situations. In contrast, RAD consistently performs well, effectively avoiding dynamic obstacles and handling challenging tasks. These results highlight the benefits of closed-loop training in the hybrid method, which enables better handling of dynamic environments. Additional visualizations are included in the Appendix (Fig. A1).

Limitation and Conclusion

In this work, we propose the first 3DGS-based RL framework for training end-to-end AD policy. We combine RL and IL, with RL complementing IL to model the causations and narrow the open-loop gap, and IL complementing RL in terms of human alignment. This work also has some limitations. Currently the used 3DGS environments are running in a non-reactive manner, i.e., other traffic participants do not react according to ego vehicle’s behavior but act with log replay. And the effect of 3DGS still has room for improvement, especially for rendering non-rigid pedestrians, unobserved views, and low-light scenarios. Future works will focus on solving these problems and scaling up RL to the next level.

Acknowledgement

We would like to acknowledge Qingjie Wang, Yongjun Yu, Zehua Li, Peng Wang, Nuoya Zhou, Songlin Yang, Ruiqi Wang, Tianheng Cheng, Changze Li, Zhe Chen, and Tong Qin for discussion and assistance.

References

Appendix A Appendix

Here, we provide a more comprehensive explanation of the action space design. To ensure stable control and efficient learning, we define the action space over a short time horizon of 0.5 seconds. The ego vehicle’s movement is modeled using discrete displacements in both the lateral and longitudinal directions.

The lateral displacement, denoted as axa^{x}, represents the vehicle’s movement in the lateral direction over the 0.50.5-second horizon. We discretize this dimension into NxN_{x} options, symmetrically distributed around zero to allow leftward and rightward movements, with an additional option to maintain the current trajectory. The set of possible lateral displacements is:

In our implementation, we use Nx=61N_{x}=61, with dminx=−0.75d^{x}_{\text{min}}=-0.75 m, dmaxx=0.75d^{x}_{\text{max}}=0.75 m, and intermediate values sampled uniformly.

Longitudinal Displacement.

The longitudinal displacement, denoted as aya^{y}, represents the vehicle’s movement in the forward direction over the 0.50.5-second horizon. Similar to the lateral component, we discretize this dimension into NyN_{y} options, covering a range of forward displacements, including an option to maintain the current position:

In our setup, we use Ny=61N_{y}=61, with dmaxy=15d^{y}_{\text{max}}=15m, and intermediate values sampled uniformly.

A.2 Implementation Details

In this section, we summarize the training settings, configurations, and hyperparameters used in our approach.

Planning Pre-Training. The action space is discretized using predefined anchors A={(aix,ajy)}i=1,j=1Nx,Ny\mathcal{A}=\{(a^{x}_{i},a^{y}_{j})\}_{i=1,j=1}^{N_{x},N_{y}}. Each anchor corresponds to a specific steering-speed combination within the 0.5-second planning horizon. Given the ground truth vehicle position at t=0.5 st=0.5\ \text{s} denoted as pgt=(pgtx,pgty)p_{\text{gt}}=(p_{\text{gt}}^{x},p_{\text{gt}}^{y}), we implement normalized nearest-neighbor matching over predefined anchor positions:

Based on the matched anchor indices (i^,j^)(\hat{i},\hat{j}), we formulate the imitation learning objective as a dual focal loss :

where Lfocal\mathcal{L}_{\text{focal}} is focal loss for discrete action classification, and π(ax∣s)\pi(a^{x}\mid s) and π(ay∣s)\pi(a^{y}\mid s) are predicted action distributions from Eq. 1.

Training configurations. We provide detailed hyperparameters for the two main stages, Planning Pre-Training and Reinforced Post-Training, in Tab. A1 and Tab. A2, respectively.

A.3 Metric Details

We evaluate the performance of the autonomous driving policy using eight key metrics.

Dynamic Collision Ratio (DCR). DCR quantifies the frequency of collisions with dynamic obstacles. It is defined as:

where NdcN_{dc} is the number of clips in which collisions with dynamic obstacles occur, and NtotalN_{total} is the total number of clips.

Static Collision Ratio (SCR). SCR measures the frequency of collisions with static obstacles and is defined as:

where NscN_{sc} is the number of clips with static obstacle collisions.

Collision Ratio (CR). CR represents the total collision frequency, given by:

Positional Deviation Ratio (PDR). PDR evaluates the ego vehicle’s adherence to the expert trajectory in terms of position. It is defined as:

where NpdN_{pd} is the number of clips in which the positional deviation exceeds a predefined threshold.

Heading Deviation Ratio (HDR). HDR assesses orientation accuracy by computing the proportion of clips where heading deviations surpass a predefined threshold:

where NhdN_{hd} is the number of clips where the heading deviation exceeds the threshold.

Deviation Ratio (DR). captures the overall deviation from the expert trajectory, given by:

Average Deviation Distance (ADD). ADD quantifies the mean closest distance between the ego vehicle and the expert trajectory during time steps when no collisions or deviations occur. It is defined as:

where TsafeT_{safe} represents the total number of time steps in which the ego vehicle operates without collisions or deviations, and dmin(t)d_{min}(t) denotes the minimum distance between the ego vehicle and the expert trajectory at time step tt.

Finally, Longitudinal Jerk (Long. Jerk) and Lateral Jerk (Lat. Jerk) quantify the smoothness of vehicle motion by measuring acceleration changes. Longitudinal jerk is defined as:

where vlongv_{long} represents the longitudinal velocity. Similarly, lateral jerk is defined as:

where vlatv_{lat} is the lateral velocity. These metrics collectively capture abrupt changes in acceleration and steering, providing a comprehensive measure of passenger comfort and driving stability.

A.4 More Qualitative Results

This section presents additional qualitative comparisons across various driving scenarios, including detours, crawling in dense traffic, traffic congestion, and U-turn maneuvers. The results highlight the effectiveness of our approach in generating smoother trajectories, enhancing collision avoidance, and improving adaptability in complex environments.