Hydra-NeXt: Robust Closed-Loop Driving with Open-Loop Training
Zhenxin Li, Shihao Wang, Shiyi Lan, Zhiding Yu, Zuxuan Wu, Jose M. Alvarez
Introduction
End-to-end autonomous driving (E2E AD) has emerged as a trending alternative to traditional perception-planning pipelines. Current research in E2E AD is divided into open-loop and closed-loop environment-oriented approaches.
An open-loop environment refers to an autonomous system’s performance being assessed without environmental feedback. The main advantage of this approach is its simplicity and efficiency, as it avoids the complexity of real-time feedback, making it cost-effective and suitable for large-scale training and testing. Benchmarks like NAVSIM provide controlled environments where agents follow predefined trajectories, allowing for precise evaluation of metrics like collision rates, progress, and traffic-rule following. However, the lack of reactive behaviors and dynamic interactions with other agents or changing conditions makes it difficult to deploy models trained in open-loop environments in real-world driving.
A closed-loop environment in autonomous driving involves real-time feedback between the agent’s actions and the environment, typically via a simulator . The primary advantage of closed-loop methods is their ability to simulate dynamic, interactive environments, allowing for more realistic testing of driving policies. However, they face significant challenges: These systems often rely on synthetic simulator data, leading to a domain gap between training and real-world deployment. For instance, the previous state-of-the-art E2E AD method depends heavily on reinforcement learning (RL) teacher feature embeddings, unavailable in real-world data. In addition, the driving behavior learned may deviate from human patterns , and evaluation metrics focus mainly on collision rates, overlooking other important factors such as smoothness and driving comfort .
The gap between open-loop and closed-loop environment has led to different approaches in each domain, and this separation continues to grow. Fig. 1 shows that closed-loop methods learn to control through reinforcement learning in simulations , while open-loop methods focus on imitating expert trajectories . The gap between the two domains stem from their training methods and output representations, which can be potentially addressed through two approaches. We could either adapt models trained in an open-loop environment to a closed-loop environment or bring closed-loop models to open environments. We argue that the first option, From Open-loop to Closed-loop, is more promising due to the availability of massive real-world data. In contrast, the second option, From Closed-loop to Open-loop, faces sim-to-real domain gaps , which pose the biggest challenge to deploy a model in a real car. Hence, we focus on addressing the limitations of open-loop models when driving in closed-loop environments, which involve handling reactive agents and adhering to kinematic constraints.
As the state-of-the-art open-loop E2E method on the NAVSIM benchmark , Hydra-MDP shows promising results for non-reactive agents, so the challenge is to extend this ability to reactive agents. Non-reactive agents maintain their behavior even if the ego vehicle deviates from the ground truth trajectory. In contrast, reactive agents adjust in response to the ego’s predicted trajectory, potentially causing the E2E AD model’s collision predictions to fail. The general issue caused by open-loop training is that most open-loop metrics focus on waypoints rather than control signals. While waypoints are useful for general navigation, they do not account for immediate changes in the behavior of other agents, leading to delays in reaction time when unexpected events occur. In contrast, control signals directly influence the vehicle’s actions, such as braking or steering, allowing for faster adjustments in response to sudden changes. Although existing literature also predicts control with waypoints directly, it performs poorly in interactive situations , which may stem from a lack of multi-modal decisions and insufficient exploration of their nuanced relationships. Further, the predicted waypoints may not satisfy kinematic constraints in the closed-loop environment, which can cause compounding errors and result in unsafe behaviors in dynamic scenarios.
Therefore, we introduce Hydra-NeXt , featuring a multi-branch planning framework. Hydra-NeXt consists of Multi-head Motion Decoders for trajectory and control prediction, and a Trajectory Refinement network for kinematics-based proposal selection. In addition to the trajectory decoder to handle general-case planning , the control decoder enhances the capability for short-term actions. Additionally, the Trajectory Refinement module applies kinematic constraints to improve predictions from both decoders, effectively combining and refining the planning decisions.
Our contributions can be summarized as follow:
We propose Hydra-NeXt , which is the first framework to unify control, trajectory prediction, and trajectory refinement. Hydra-NeXt demonstrates robust closed-loop driving performance with only open-loop learning.
We benchmark Hydra-NeXt against other E2E AD solutions on the Bench2Drive dataset based on the simulator CARLA v2 . Hydra-NeXt outperforms previous state-of-the-art methods by a significant margin (+22.9 Driving Score, +17.5 Success Rate) by the CARLA v2 evaluation protocol without relying on external experts for data collection. Hydra-NeXt also generalizes well to real-world data, achieving a new state-of-the-art on the open-loop planning benchmark NAVSIM .
Related Work
To explore end-to-end planning with real-world vehicle data , researchers have integrated modularized neural networks into end-to-end fully-differentiable stacks . This approach quickly gained significant interest in the AV community as it shows a promising direction to scalable autonomy. However, recent works discover heavy biases of current E2E planning datasets , the imitation learning paradigm , and evaluation protocols , rendering open-loop E2E methods as unreliable. To address these issues, a more advanced benchmark NAVSIM is proposed to filter homogeneous planning data and benchmark with a series of rule-based metrics for collision avoidance, map adherence, and more. The state-of-the-art open-loop planner Hydra-MDP , which learns these rule-based metrics via knowledge distillation, achieves greater reliability than imitation-based E2E planners. Building on Hydra-MDP, we develop a model suited for dynamic scenarios, transitioning from open-loop trajectory planning to a closed-loop framework using control prediction and kinematics-based trajectory refinement.
2 Closed-loop Driving Benchmarks and Policies
Compared with open-loop E2E autonomous driving, closed-loop driving has a longer history. Benchmarks such as nuPlan , CARLA v1 , CARLA v2 , and Bench2Drive focus on different aspects of driving policies, from trajectory planning to E2E driving . Recently, CARLA v2 poses significant challenges as it requires the handling of interactive and dynamic scenarios. Bench2Drive optimizes the variances produced by the CARLA v2 evaluation protocol and collects data from an RL-based expert on diverse, shorter routes, making it suitable to assess policies in varied conditions.
On these benchmarks, expert-level policies include rule-based path planners and RL-based policies . These policies take privileged inputs such as ground-truth perception and map data, while E2E AD methods only use sensor observations. These E2E policies, as well as previous open-loop E2E methods, fall behind expert-level policies when it comes to dynamic and interactive scenarios . Moreover, Bench2Drive shows that these methods rely on expert feature embeddings for better performance. In our paper, we aim to close the gap between the two worlds: open-loop planning and closed-loop driving by enhancing a top open-loop planner with control and kinematic components.
3 Diffusion-based Driving Policies
Diffusion policies have been used widely for the generation of robot behaviors. They prove effective in capturing multi-modal action distributions and generating smooth trajectories. In autonomous driving, recent works leverage diffusion models to predict trajectory waypoints given privileged input of surrounding traffic scenes . Our approach differs by employing the diffusion policy to generate multiple smooth and high-frequency control sequences, which serve as additional proposals for trajectory refinement.
Trajectory Decoder: Hydra-MDP
Before introducing Hydra-NeXt, we start from Hydra-MDP , a multi-modal planner with a perception network and a trajectory decoder, which learns from both human demonstrations and rule-based open-loop metrics.
Perception Network: Given front-view and back-view images, we use an image backbone to extract multi-view features. These features are flattened into a sequence of environment tokens to represent the surroundings. To focus on motion planning, we apply a minimal design in the perception network without auxiliary perception tasks , although they can facilitate planning.
Trajectory Decoder: The trajectory decoder generates a trajectory for routing based on environment tokens and a discrete trajectory vocabulary . Following Hydra-MDP , we create a trajectory vocabulary with 4096 discrete trajectories, embed them into latent queries , and attend to environment tokens in a transformer decoder . With only imitation learning, this approach achieves 49.0 DS as shown in Fig. 2, outperforming the previous state-of-the-art DriveAdapter by 6.1 DS. Further, we implement an open-loop metric system on Bench2Drive using the following metrics:
Collision: The collision metric checks if the ego vehicle intersects with other agents in the Bird’s-Eye View (BEV) space . We compute collisions at a higher frequency (i.e. 10Hz) by interpolating trajectory waypoints following NAVSIM .
Soft Lane Keeping: To ensure lane adherence while allowing lane changing, we design a soft lane keeping metric that bounds the ego’s driving area following . We penalize trajectories with excessive angular differences between trajectory segments and lane segments.
Ego Progress: The Ego Progress metric is included to discourage passive driving behaviors and promote driving progress. Similar to , we project the trajectory waypoints onto the lane segments that the expert travels through, and normalize the ego distance by the expert’s.
These metrics are used as rule-based teachers in Hydra-MDP. Fig. 2 shows the performance enhancement brought by rule-based heads, which appear to be limited compared with observations on NAVSIM . This is likely caused by randomly disappearing agents in Bench2Drive https://github.com/Thinklab-SJTU/Bench2Drive/issues/31, which confuse the rule-based heads. e.g. A trajectory that extends far into the future may be labeled as safe if its future waypoints do not intersect with a disappeared agent, even though the agent is currently visible and close to the ego vehicle. In the next section, we propose Hydra-Next, an extended version of Hydra-MDP addressing the main limitations in closed-loop driving.
Hydra-NeXt
In this section, we elaborate on Hydra-NeXt, an E2E framework with robust closed-loop abilities. Apart from the Trajectory Decoder , Hydra-NeXt has two other policies: Control Decoder , and Trajectory Refinement .
Our goal is to enable quick responses to reactive agents through direct control output, instead of relying on trajectory waypoints for rapid decisions. This is because waypoints are typically transformed into control signals through specific controllers such as a PID Controller or Model Predictive Control , which can lead to errors during path following. Further, TCP points out that ensembling the trajectory and the control signal can boost driving performance, but it only considers a single modality for both the trajectory and the control prediction branches. This limitation makes it challenging to address uncertainties in environments and makes it susceptible to interpolate between different action modes .
Based on these findings, we add a second classification-based control decoder to generate control signals for timesteps. In the CALRA simulator , each control signal is a tuple of used to direct the vehicle. Meanwhile, we incorporate the idea of discretization into the control decoder to handle uncertainty, following the practice in the RL-based driving policy Think2Drive . The use of learning-based control prediction sidesteps the need to convert trajectories into control signals through traditional controllers, enabling faster responses to reactive agents in the closed-loop environment.
In particular, we first randomly initialize control queries , which correspond to the current and future steps for which we want to predict control signals. Similarly, attends to in a transformer decoder. After this, is processed by three separate MLP layers for each control signal (i.e. ). Since the expert data collected by Think2Drive only consists of discrete control signals, the control decoder can be easily trained using cross-entropy loss functions without further discretizing the control signal ground truths. We select the control tuples with the highest likelihood at each step as the final output of . The design choices of the control decoder are discussed in Sec. 5.4, such as architectures and loss functions.
2 Trajectory Refinement πdp\pi_{dp}
Another important aspect of closed-loop driving is to ensure the smoothness of the driving process. Benchmarks like nuPlan and NAVSIM post-process trajectories using an LQR tracker and a kinematic bicycle model to adhere to kinematic constraints. Nevertheless, this idea has been neglected by existing open-loop E2E AD methods, which mainly focuses on cloning trajectory waypoints without considering kinematic constraints in closed-loop environments. This oversight becomes particularly critical when the vehicle experiences control loss in closed-loop environments or encounters large compounding errors when quickly recovering to predicted waypoints is difficult.
With these smooth control sequence proposals, we devise a process named Nearest Neighbor Matching for choosing the best proposal in a kinematically feasible way. Algorithm 1 depicts the nearest neighbor matching process for proposal selection in detail. Specifically, we first use a kinematic bicycle model to transform control sequences from and into trajectory waypoints. After the transformation, we select two nearest control candidates to match current candidates based on L2 distances. The new control candidates can be viewed as a kinematically feasible version of the predictions given by the Multi-head Motion Decoder. Finally, we ensemble the candidates into the final control signal by averaging the throttle and steer values, which is a simplified ensembling method from TCP . The brake is set to 1 if the condition holds; otherwise, it is set to 0. Note that the brake values produced by the candidates are binary and is a predefined brake threshold.
Experiments
We use the E2E driving benchmark Bench2Drive for training and evaluating Hydra-NeXt. The training data in this paper uses the official training dataset of Bench2Drive, which includes 2 million frames encompassing 44 interactive scenarios. These data are collected by Think2Drive , an RL-based expert model on CARLA v2 . For evaluation, the E2E AD model is deployed in the CARLA simulator to perform closed-loop driving on 220 short routes designed by Bench2Drive. These short routes assess the AD model’s abilities on different scenarios, while ensuring low variance in the final score.
Bench2Drive includes several metrics, including Driving Score (DS), Success Rate (SR), Efficiency, Comfort, and Multi-ability Results. The calculation of DS can be based on two evaluation protocols: CARLA v2 and Bench2Drive. The former accumulates the minimum speed infractions into the DS, whereas the Bench2Drive protocol separates this into the metric Efficiency. Meanwhile, the latter protocol relaxes the time constraint of the closed-loop evaluation, which reduces the difficulty for the model to complete the routes. Without further notations, we default to the CARLA v2 evaluation protocol in ablation studies to reflect the comprehensive performance of the method using the DS metric.
Furthermore, we evaluate Hydra-NeXt on the real-world NAVSIM Benchmark , which evaluates a 4-second trajectory using open-loop metrics: No at-fault Collisions (NC), Drivable Area Compliance (DAC), Time-to-collision (TTC), Comfort (C), and Ego Progress (EP). The PDM score (PDMS) is an aggregate of these sub-metrics.
2 Implementation Details
The implementation of Hydra-NeXt is largely consistent with open-loop E2E baselines on Bench2Drive. First, we train Hydra-NeXt on the training data for 20 epochs with a total batch size of 256, using 8 NVIDIA V100 GPUs. The AdamW optimizer is used with the Cosine Annealing Scheduler at a learning rate of and a weight decay of 0.01. During training, data augmentations such as random cropping and photometric distortion are applied to the input images, which are first resized to a resolution of .
Hydra-NeXt employs an ImageNet-pretrained ResNet-50 as the image backbone to extract front-view and back-view image features. Following VAD , the trajectory decoder predicts a 3-second trajectory at 2Hz. The frequencies of and are set to 2Hz and 10Hz by default, which will be further analyzed in Sec. 5.4. The ego status feature includes the current longitudinal velocity and a one-hot navigation command. In Trajectory Refinement, the brake threshold is set to half of the candidates set size 2 for moderate behaviors, while the kinematic bicycle model is based on the implementation of PDM-Lite and the denoising timestep of is set to 100 . For NAVSIM, we extend Hydra-MDP by incorporating acceleration and steering rate prediction as a substitute for control signals. Details can be found in the appendix.
3 Main Results
As shown in Tab. 1 and Tab. 2, Hydra-NeXt surpasses all E2E methods on Bench2Drive in key metrics such as the DS and SR distinctively. Specifically, under the CARLA v2 protocol, Hydra-NeXt outperforms the previous state-of-the-art, DriveAdapter , by 22.98 DS and 17.49% SR. Under the Bench2Drive protocol, it achieves improvements of 9.64 DS and 16.92 SR, despite DriveAdapter utilizing expert features from Think2Drive for training. Additionally, Hydra-NeXt consistently demonstrates higher efficiency than all baselines, especially those using expert features. Nevertheless, Hydra-NeXt falls behind TCP in terms of Comfort, which is likely due to frequent braking during interactive situations.
For the Multi-ability Results shown in Tab. 3, Hydra-NeXt shows an improvement of 11.14% over DriveAdapter in the average performance. It is worth noting that Hydra-NeXt exhibits superior performance in interactive scenarios such as merging (+11.18%), overtaking (+38.06%), and emergency braking (+12.91%), which indicates its proficiency in handling reactive agents. However, Hydra-NeXt falls behind DriveAdapter in scenarios where the vehicle should adhere to traffic signs. This phenomenon may be caused by causal confusion when encountering mixed expert behaviors before traffic signs, and is worth future investigations. Notably, all methods achieve less satisfactory results in yielding to specialized vehicles (e.g. ambulances), which may be caused by the scarcity of such training data.
Moreover, Tab. 4 shows that Hydra-NeXt achieves a new state-of-the-art on the NAVSIM Benchmark. Though NAVSIM only evaluates the predicted trajectory from , the extension of learning targets in and can benefit the trajectory prediction, leading to an improvement of 2.1 PDMS over Hydra-MDP .
4 Ablation Study
Ablation on Different Components and their Designs. Fig. 2 shows the detailed ablations of each component in Hydra-NeXt. Operating on perspective image tokens, Hydra-MDP achieves a 52.8 DS, which already surpasses DriveAdapter by around 10 DS. Nevertheless, the trajectory prediction remains insufficient when facing more complex agent interactions, which promotes the utilization of a control prediction network for rapid reactions. The output of is ensembled with the trajectory in a similar fashion to Sec. 4.2, where the brake threshold is correspondingly set to 1 given two proposals. This gives us a large increase of 12.7% in the SR, indicating fewer collisions with reactive agents. Furthermore, using multi-step control predictions into future frames as in provides auxiliary supervision (+1.8 DS). Applying a binary focal loss for brake prediction balances the training (+1.8% SR) and simply post-processing the final control output, such as slowing down during turns, leads to further enhancements (+2.5% SR). Finally, we build the Trajectory Refinement network on top of the previous network. Involving the Diffusion Policy and applies a straightforward ensembling leads to an increase of around 3 DS, while applying the Nearest Neighbor Matching (Algorithm 1) to follow kinematic constraints boosts the final performance to 65.9 DS and 48.2% SR.
Design of Trajectory Refinement. To illustrate the necessity of using the Diffusion Policy as the proposal generator, we first replace the Diffusion Policy with a discrete control decoder with a similar architecture to , which can also generate proposals based on top- confidence scores. This leads to a performance degradation of 5.78 DS as shown in Tab. 5, , which potentially stems from the poor ability of to generate long and smooth control sequences. Further, we investigate the ability of in modeling high-frequency control sequences. Tab. 6 indicates a performance degradation when using a low-frequency . On the contrary, a low-frequency is better than a high-frequency one. This echos the previous finding and proves the effectiveness of employing as the generator. Finally, Tab. 7 examines the impact of varying the number of proposals. Increasing the proposal number from 5 to 20 leads to saturating performance on DS, while SR mildly fluctuates with a larger . The fluctuation may result from over-confident matching between and the other decoders, while we can potentially benefit from small differences between their predictions for exploration.
5 Runtime Efficiency and Visualization
Runtime Efficiency. Tab. 8 compares the runtime efficiency of Hydra-NeXt with E2E baselines. Since performs planning based on perspective image tokens without constructing an explicit BEV feature , it achieves higher efficiency than the previous UniAD. The control decoder is also lightweight. Nevertheless, involving the Diffusion Policy for refining trajectories results in an efficiency degradation, which is an inherent limitation of the iterative denoising process and deserves further optimization.
Visualization. Fig. 4 depicts the process of Trajectory Refinement in a highway merging scenario and an unprotected left turn. Multiple control sequences are proposed by and transformed into trajectories. The best-matching proposal can smooth the trajectory produced by with a large curvature, ensuring kinematic feasibility.
Conclusion
In this paper, we first analyzed the development of current end-to-end autonomous driving systems, focusing on the discrepancy between open-loop and closed-loop environments and the challenges of bringing an open-loop trained model to closed-loop driving. These challenges include handling reactive agents and adhering to kinematic constraints. To address these issues, we propose Hydra-NeXt, a unified approach for trajectory and control prediction. By integrating Hydra-MDP with the Multi-head Motion Decoder and the Trajectory Refinement network, Hydra-NeXt obtains stronger closed-loop driving performance. Specifically, Hydra-NeXt demonstrates superior performance of 65.89 Driving Score (DS) and 48.20% Success Rate (SR) on the closed-loop driving benchmark Bench2Drive and outperforms previous state-of-the-art methods by a significant margin. Hydra-NeXt also achieves a new state-of-the-art on the real-world E2E planning benchmark NAVSIM.
References
Details on Training Hydra-NeXt
In this section, we provide the details on training Hydra-NeXt such as loss functions and the formulation of the diffusion policy . The three policies used in Hydra-NeXt (, , and ) are all trained end-to-end simultaneously on the Bench2Drive dataset , which is collected by the RL-based expert Think2Drive . External data from different experts such as PDMLite are not used.
The loss of is consistent with Hydra-MDP , which consists of two terms: an imitation loss and a knowledge distillation loss . We denote trajectory anchors used in as . The imitation loss is calculated as a cross entropy based on the expert trajectory and the predicted imitation scores for each trajectory anchor :
measures the similarity between the expert trajectory and each trajectory anchor . The knowledge distillation loss aims to learn open-loop objectives via binary cross-entropy between predicted metric scores and the ground-truth metric scores :
correspond to the Collision, Soft Lane Keeping, and Ego Progress metrics defined in Sec. 3. The ground-truth metric scores are binary values, derived from the privileged information and each trajectory anchor. The overall loss of is formulated as:
2 πctrl\pi_{ctrl}: Control Decoder
The loss of computes a classification-based loss for each control signal (i.e. brake, throttle, and steer) across timesteps. For simplicity, a control signal tuple is abbreviated as , while the expert demonstration is denoted as . Specifically, these loss functions are formulated as:
where stands for the focal loss and is the cross entropy loss. We empirically find that applying a focal loss to throttle and steer predictions leads to a performance degradation, possibly due to their more balanced distribution in the dataset.
3 πdp\pi_{dp}: Diffusion Policy
is based on the standard Diffusion Policy trained on continuous data. Given expert control signals across frames, is trained to predict the noise added to the expert control signals, where is the denoising iteration. MSE loss is applied for noise prediction:
Finally, the overall loss of Hydra-NeXt becomes
Performance of Individual Policies
We conduct experiments on how each individual policy performs on the Bench2Drive Benchmark. As shown in Tab. 10, using alone leads to serious performance degradation compared with the full version of Hydra-NeXt (-16.56 DS and -29.33 SR). Moreover, achieves better results when its control candidate follows the predicted trajectory from rather than being randomly selected (+5.36 DS and +4.28 SR), highlighting the importance of trajectory guidance. Finally, the full version of Hydra-NeXt achieves substantial improvements (+13.09 DS and +17.47 SR) compared with the baseline Hydra-MDP ().
Implementation on NAVSIM
Our implementation of Hydra-NeXt on NAVSIM uses a different perception network Transfuser following Hydra-MDP . Transfuser features two backbones for camera and lidar feature extraction, a BEV segmentation head, a 3D object detection head, and transformer layers for multi-modal feature interaction. This setting helps to make a fair comparison to baselines such as Hydra-MDP and DiffusionDrive . For and , we incorporate the same transformer architectures used on Bench2Drive for acceleration and steering rate predictions since NAVSIM does not utilize control signals like CARLA and only evaluates the trajectory. Therefore, these auxiliary predictions only act as extra learning targets. As a result, Hydra-NeXt surpasses the state-of-the-art planner DiffusionDrive by 0.5 PDMS (see Tab. 9) when adopting the grid search trick among different metric scores in .
Visualization Results
Fig. 5 shows more visualization results in interactive scenarios (Merging, Overtaking, and Give Way). The Diffusion Policy can capture multiple planning modes such as following other agents or overtaking them (see the second row and the fourth row of the figure).
Limitations.
Although Hydra-NeXt shows outstanding closed-loop driving performance compared with E2E methods, it still falls behind RL-based experts using privileged input. The runtime efficiency of the Diffusion Policy also deserves optimization. We expect these to be addressed in future research.