VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning

Bo Jiang, Shaoyu Chen, Hao Gao, Bencheng Liao, Qian Zhang, Wenyu Liu, Xinggang Wang

Introduction

End-to-end autonomous driving is an important and popular field recently. Mass of human driving demonstrations are easily available. It seems promising to learn a human-like driving policy from large-scale demonstrations.

However, the uncertainty and non-deterministic nature of planning make it challenging to extract the driving knowledge from driving demonstrations. To demonstrate such uncertainty, two scenarios are presented in Fig. 1. 1) Following another vehicle. The human driver has diverse reasonable driving maneuvers, keeping following or changing lanes to overtake. 2) Interaction with the coming vehicle. The human driver has two possible driving maneuvers, yield or overtake. From the perspective of statistics, the action (including the timing and speed) is highly stochastic, affected by many latent factors that can not be modeled.

Existing learning-based planning methods follow a deterministic paradigm to directly regress the action. The regression target a^\hat{a} is the future trajectory in and control signal (acceleration and steering) in . Such a paradigm assumes there exists a deterministic relation between environment and action, which is not the case. The variance of human driving behavior causes the ambiguity of the regression target. Especially when the feasible solution space is non-convex (see Fig. 1), the deterministic modeling cannot cope with non-convex cases and may output an in-between action, causing safety problems. Besides, such deterministic regression-based planner tends to output the dominant trajectory, which appears the most in the training data (like stop or go straight), and results in undesirable planning performance.

In this work, we propose probabilistic planning to cope with the uncertainty of planning. As far as we know, VADv2 is the first work to use probabilistic modeling to fit the continuous planning action space, which is different from previous practices that use deterministic modeling for planning. We model the planning policy as an environment-conditioned non-stationary stochastic process, formulated as p(a∣o)p(a|o), where oo is the historical and current observations of the driving environment, and aa is a candidate planning action. Compared with deterministic modeling, probabilistic modeling can effectively capture the uncertainty in planning and achieve more accurate and safe planning performance.

The planning action space is a high-dimensional continuous spatiotemporal space. We resort to a probabilistic field function to model the mapping from the action space to the probabilistic distribution. Since directly fitting the continuous planning action space is not feasible, we discretize the planning action space to a large planning vocabulary and use mass driving demonstrations to learn the probability distribution of planning actions based on the planning vocabulary. For discretization, we collect all the trajectories in driving demonstrations and adopt the furthest trajectory sampling to select NN representative trajectories which serve as the planning vocabulary.

Probabilistic planning has two other advantages. First, probabilistic planning models the correlation between each action and environment. Unlike deterministic modeling which only provides sparse supervision for the target planning action, probabilistic planning can provide supervision not only for the positive sample but also for all candidates in the planning vocabulary, which brings richer supervision information. Besides, probabilistic planning is flexible in the inference stage. It outputs multi-mode planning results and is easy to combine with rule-based and optimization-based planning methods. And we can flexibly add other candidate planning actions to the planning vocabulary and evaluate them because we model the distribution over the whole action space.

Based on the probabilistic planning, we present VADv2, an end-to-end driving model, which takes surround-view image sequence as input in a streaming manner, transforms sensor data into token embeddings, outputs the probabilistic distribution of action, and samples one action to control the vehicle. Only with camera sensors, VADv2 achieves state-of-the-art closed-loop performance on the CARLA Town05 benchmark, significantly outperforming all existing methods. Abundant closed-loop demos are presented at https://hgao-cv.github.io/VADv2. VADv2 runs stably in a fully end-to-end manner, even without the rule-based wrapper.

Our contributions are summarized as follows:

We propose probabilistic planning to cope with the uncertainty of planning. We design a probabilistic field to map from the action space to the probabilistic distribution, and learn the distribution of action from large-scale driving demonstrations.

Based on the probabilistic planning, we present VADv2, an end-to-end driving model, which transforms sensor data into environmental token embeddings, outputs the probabilistic distribution of action, and samples one action to control the vehicle.

In CARLA simulator, VADv2 achieves state-of-the-art closed-loop performance on Town05 benchmark. Closed-loop demos show it runs stably in an end-to-end manner.

Related Work

Perception. Perception is the first step in achieving autonomous driving, and a unified representation of driving scenes is beneficial for easy integration into downstream tasks. Bird’s Eye View (BEV) representation has become a common strategy in recent years, enabling effective scene feature encoding and multimodal data fusion. LSS is a pioneering work that achieves the perspective view to BEV transformation by explicitly predicting depth for image pixels. BEVFormer , on the other hand, avoids explicit depth prediction by designing spatial and temporal attention mechanisms, and achieves impressive detection performance. Subsequent works continuously improve performance in downstream tasks by optimizing temporal modeling and BEV transformation strategies. In terms of vectorized mapping, HDMapNet converts lane segmentation into vector maps through post-processing. VectorMapNet predicts vector map elements in an autoregressive manner. MapTR introduces permutation equivalence and hierarchical matching strategies, significantly improving mapping performance. LaneGAP introduces path-wise modeling for lane graphs.

Motion Prediction. Motion prediction aims to forecast the future trajectories of other traffic participants in driving scenes, assisting the ego vehicle in making informed planning decisions. Traditional motion prediction task utilizes input such as historical trajectories and high-definition maps to predict future trajectories. However, recent developments in end-to-end motion prediction methods perform perception and motion prediction jointly. In terms of scene representation, some works adopt rasterized image representations and employ CNN networks for prediction . Other approaches utilize vectorized representations and employ Graph Neural Networks or Transformer models for feature extraction and motion prediction. Some works see future motion as dense occupancy and flow instead of agent-level future waypoints. Some motion prediction methods adopt Gaussian Mixture Model (GMM) to regress multi-mode trajectories. It can be applied in planning to model uncertainty. But the number of modes is limited.

Planning. Learning-based planning has shown great potential recently due to its data-driven nature and impressive performance with increasing amounts of data. Early attempts use a completely black-box spirit, where sensor data is directly used to predict control signals. However, this strategy lacks interpretability and is difficult to optimize. In addition, there are numerous studies combining reinforcement learning and planning . By autonomously exploring driving behavior in closed-loop simulation environments, these approaches achieve or even surpass human-level driving performance. However, bridging the gap between simulation and reality, as well as addressing safety concerns, poses challenges in applying reinforcement learning strategies to real driving scenarios. Imitation learning is another research direction, where models learn expert driving behavior to achieve good planning performance and develop a driving style close to that of humans. In recent years, end-to-end autonomous driving has emerged, integrating perception, motion prediction, and planning into a single model, resulting in a fully data-driven approach that demonstrates promising performance. UniAD cleverly integrates multiple perception and prediction tasks to enhance planning performance. VAD explores the potential of vectorized scene representation for planning and getting rid of dense maps.

Large Language Model in Autonomous Driving. The interpretability and logical reasoning abilities demonstrated by large language models (LLMs) can greatly assist in the field of autonomous driving. Recent research has explored the combination of LLMs and autonomous driving. One line of work utilizes LLMs for driving scene understanding and evaluation through question-answering (QA) tasks. Another approach goes a step further by incorporating planning on top of LLM-based scene understanding. For instance, DriveGPT4 takes inputs such as historical video and text (including questions and additional information like historical control signals). After encoding, these inputs are fed into an LLM, which predicts answers to the questions and control signals. LanguageMPC , on the other hand, takes in historical ground truth perception results and HD maps in the form of language descriptions. It then utilizes a Chain of Thought analysis approach to understand the scene, and the LLM finally predicts planning actions from a predefined set as output. Each action corresponds to a specific control signal for execution. VADv2 draws inspiration from GPT to cope with the uncertainty problem. Uncertainty also exists in language modeling. Given a specific context, the next word is non-deterministic and probabilistic. LLM learns the context-conditioned probabilistic distribution of the next word from a large-scale corpus, and samples one word from the distribution. Inspired by LLM, VADv2 models the planning policy as an environment-conditioned nonstationary stochastic process. VADv2 discretizes the action space to generate a planning vocabulary, approximates the probabilistic distribution based on large-scale driving demonstrations, and samples one action from the distribution at each time step to control the vehicle.

Method

The overall framework of VADv2 is depicted in Fig. 2. VADv2 takes multi-view image sequences as input in a streaming manner, transforms sensor data into environmental token embeddings, outputs the probabilistic distribution of action, and samples one action to control the vehicle. Large-scale driving demonstrations and scene constraints are used to supervise the predicted distribution.

Information in the image is sparse and low-level. We use an encoder to transform the sensor data into instance-level token embeddings EenvE_{\rm env}, to explicitly extract high-level information. EenvE_{\rm env} includes four kinds of token: map token, agent token, traffic element token, and image token. VADv2 utilizes a group of map tokens to predict the vectorized representation of the map (including lane centerline, lane divider, road boundary, and pedestrian crossing). Besides, VADv2 uses a group of agent tokens to predict other traffic participants’ motion information (including location, orientation, size, speed, and multi-mode future trajectories). Traffic elements also play a vital role in planning. VADv2 transforms sensor data into traffic element tokens to predict the states of traffic elements. In CARLA, we consider two types of traffic signals: traffic light signals and stop signs. Map tokens, agent tokens, and traffic element tokens are supervised with corresponding supervision signals to make sure they explicitly encode corresponding high-level information. We also take image tokens as scene representations for planning, which contain rich information and are complementary to the instance-level tokens above. Besides, navigation information and ego state are also encoded into embeddings {Enavi,Estate}\{E_{\rm navi},E_{\rm state}\} with an MLP.

2 Probabilistic Planning

We propose probabilistic planning to cope with the uncertainty of planning. We model the planning policy as an environment-conditioned nonstationary stochastic process, formulated as p(a∣o)p(a|o). We approximate the planning action space as a probabilistic distribution based on large-scale driving demonstrations, and sample one action from the distribution at each time step to control the vehicle.

3 Training

We train VADv2 with three kinds of supervision, distribution loss, conflict loss, and scene token loss,

Distribution Loss. We learn the probabilistic distribution from large-scale driving demonstrations. KL divergence is used to minimize the difference between the predicted distribution and the distribution of the data.

In the training phase, the ground truth trajectory is added to the planning vocabulary as the positive sample. Other trajectories are regarded as negative samples. We assign different loss weights to negative trajectories. Trajectories close to the ground truth trajectory are less penalized.

Conflict Loss. We use the driving scene constraints to help the model learn important prior knowledge about driving and further regularize the predicted distribution. Specifically, if one action in the planning vocabulary conflicts with other agents’ future motion or road boundary, the action is regarded as a negative sample, and we impose a significant loss weight to reduce the probability of this action.

Scene Token Loss. Map tokens, agent tokens, and traffic element tokens are supervised with corresponding supervision signals to make sure they explicitly encode corresponding high-level information.

The loss of map tokens is the same with MapTRv2 . l1\textit{l}_{1} loss is adopted to calculate the regression loss between the predicted map points and the ground truth map points. Focal loss is used as the map classification loss.

The loss of agent tokens is composed of the detection loss and the motion prediction loss, which is the same with VAD . l1\textit{l}_{1} loss is used as the regression loss to predict agent attributes (location, orientation, size, etc.), and focal loss to predict agent classes. For each agent who has matched with a ground truth agent, we predict KK future trajectories and use the trajectory that has the minimum final displacement error (minFDE) as a representative prediction. Then we calculate l1\textit{l}_{1} loss between this representative trajectory and the ground truth trajectory as the motion regression loss. Besides, focal loss is adopted as the multi-modal motion classification loss.

Traffic element tokens consist of two parts: the traffic light token and the stop sign token. On one hand, we send the traffic light token to an MLP to predict the state of the traffic light (yellow, red, and green) and whether the traffic light affects the ego vehicle. On the other hand, the stop sign token is also sent to an MLP to predict the overlap between the stop sign area and the ego vehicle. Focal loss is used to supervise these predictions.

4 Inference

In closed-loop inference, it’s flexible to get the driving policy πmodel\pi_{model} from the distribution. Intuitively, we sample the action with the highest probability at each time step, and use the PID controller to convert the selected trajectory to control signals (steer, throttle, and brake).

In real-world applications, there are more robust strategies to make full use of the probabilistic distribution. A good practice is, sampling top-K actions as proposals, and adopting a rule-based wrapper for filtering proposals and an optimization-based post-solver for refinement. Besides, the probability of the action reflects how confident the end-to-end model is, and can be regarded as the judgment condition to switch between conventional PnC and learning-based PnC.

Experiments

The widely used CARLA simulator is adopted to evaluate the performance of VADv2. Following common practice, we use Town05 Long and Town05 Short benchmarks for closed-loop evaluation. Specifically, each benchmark contains several pre-defined driving routes. Town05 Long consists of 10 routes, each route is about 1km in length. Town05 Short consists of 32 routes, each route is 70m in length. Town05 Long validates the comprehensive capabilities of the model, while Town05 Short focuses on evaluating the model’s performance in specific scenarios, such as lane changing before intersections.

We use the official autonomous agent of CARLA to collect training data by randomly generating driving routes in Town03, Town04, Town06, Town07, and Town10. The data is sampled at a frequency of 2Hz, and we collect approximately 3 million frames for training. For each frame, we save 6-camera surround-view images, traffic signals, information about other traffic participants, and the state information of the ego vehicle. Additionally, we obtain the vectorized maps for training the online mapping module by preprocessing the OpenStreetMap format maps provided by CARLA. It is important to note that the map information was only provided as ground truth during training, and VADv2 does not utilize any high-definition map in closed-loop evaluation.

2 Metrics

For closed-loop evaluation, we use the official metrics of CARLA. Route Completion indicates the percentage of the route distance completed by an agent. Infraction Score indicates the degree of infractions happening along the route. Typical infractions include running red lights, collisions with pedestrians, etc.. Each type of infraction has a corresponding penalty coefficient, with more infractions happening, Infraction Score becomes lower. Driving Score serves as the product between the Route Completion and the Infraction Score, which is the main metric for evaluation. In benchmark evaluation, most works adopt a rule-based wrapper to reduce the infraction. For fair comparisons with other methods, we follow the common practice of adopting a rule-based wrapper over the learning-based policy.

For open-loop evaluation, L2 distance and collision rate are adopted to show which degree the learned policy drives similar to the expert demonstrations. In ablation experiments, we adopt open-loop metrics for evaluation, considering open-loop metrics are fast to calculate and more stable. We use the official autonomous agent of CARLA to generate the validation set on the Town05 Long benchmark for open-loop evaluation, and the results are averaged over all validation samples.

3 Comparisons with State-of-the-Art Methods

On the Town05 Long benchmark, VADv2 achieved a Drive Score of 85.1, a Route Completion of 98.4, and an Infraction Score of 0.87, as shown in Tab. 1. Compared to the previous state-of-the-art method , VADv2 achieves a higher Route Completion while significantly improving Drive Score by 9.0. It is worth noting that VADv2 only utilizes cameras as perception input, whereas utilizes both cameras and LiDAR. Furthermore, compared to the previous best method which only relies on cameras, VADv2 demonstrates even greater advantages, with a remarkable increase in Drive Score of up to 16.8.

We present the results for all publicly available works on the Town05 Short benchmark in Tab. 2. Compared to the Town05 Long benchmark, the Town05 Short benchmark focuses more on evaluating the ability of models to perform specific driving behaviors, such as lane changing in congested traffic flow and lane changing before intersections. In comparison to the previous result , VADv2 significantly improves Drive Score and Route Completion by 25.3 and 5.7 respectively, demonstrating the comprehensive driving ability of VADv2 in complex driving scenarios.

4 Ablation Study

Tab. 3 shows the ablation experiments of the key modules in VADv2. The model performs poorly in terms of planning accuracy without the supervision of expert driving behavior provided by the Distribution Loss (ID 1). The Conflict Loss provides critical prior information about driving, so without the Conflict Loss (ID 2), the model’s planning accuracy is also affected. Scene tokens encode important scene elements into high-dimensional features, and the planning tokens interact with the scene tokens to learn both dynamic and static information about the driving scene. When any type of scene token is missing, the model’s planning performance will be affected (ID 3-ID 6). The best planning performance is achieved when the model incorporates all of the aforementioned designs (ID 7).

5 Visualization

Fig. 3 presents some qualitative results of VADv2. The first image showcases multi-modal planning trajectories predicted by VADv2 at different driving speeds. The second image showcases VADv2’s predictions of both forward creeping and multi-modal left-turn trajectories in a lane-changing scenario. The third image depicts a right lane-changing scenario at an intersection, where VADv2 predicts multiple trajectories for both going straight and changing lanes to the right. The final image demonstrates a lane-changing scenario where there is a vehicle in the target lane, and VADv2 predicts multiple reasonable lane-changing trajectories.

Conclusion

In this work, we present VADv2, an end-to-end driving model based on probabilistic planning. In the CARLA simulator, VADv2 runs stably and achieves state-of-the-art closed-loop performance. The feasibility of this probabilistic paradigm is primarily validated. However, its effectiveness in more complicated real-world scenarios remains unexplored, which is the future work.

References