Enhancing End-to-End Autonomous Driving with Latent World Model
Yingyan Li, Lue Fan, Jiawei He, Yuqi Wang, Yuntao Chen, Zhaoxiang Zhang, Tieniu Tan
Introduction
End-to-end autonomous driving is increasingly recognized for its potential advantages over traditional methods. The traditional planners cannot access the original sensor data. This leads to information loss and error accumulation . In contrast, end-to-end planners process sensor data to directly output planning decisions, which is shown as a promising area for further exploration.
Most end-to-end autonomous driving methods , though operating in an end-to-end fashion, leverage a variety of auxiliary tasks such as detection, tracking, and map segmentation. These auxiliary tasks help the model learn better scene representations. However, they require a large amount of manual annotations, which is quite expensive and limits the data scalability. In contrast, a few end-to-end methods do not adopt perception tasks and only learn from recorded driving videos and trajectories. These approaches can leverage a large amount of available data, making it a promising direction. However, using only limited guidance from trajectories makes it difficult for the network to learn effective scene representations and achieve optimal driving performance.
To address this issue, we enhance end-to-end driving through self-supervised learning, as illustrated in Fig. 1. Traditional self-supervised methods in imaging typically concentrate on static, single-frame images. However, autonomous driving involves a dynamic series of inputs, making it essential to use temporal data effectively. A key skill in driving is predicting future conditions based on the current surroundings. Inspired by this, we propose a self-supervised task aimed at forecasting latent features. Specifically, a latent world model is developed to forecast future states based on the current states and ego actions, where the states are represented as the latent scene features within the network. During training, we extract the latent feature of the future frame to supervise the predicted latent feature from the latent world model. As a result, we jointly optimize the latent feature learning and trajectory prediction of the current frame.
Moreover, we establish a simple yet strong planner to extract view-wise latent features and serve as the testbed of the proposed latent world model. Unlike previous methods, this planner does not incorporate ad-hoc modules and perception-related branches, making it more straightforward to understand the inner workings of the latent world model. Given this planner and the latent world model, we have side products. Since the latent world model is capable of predicting future view latent features, we can skip the feature extraction process of some views in the future frame and use the predicted futures of these views as a substitution. By skipping the feature extraction for certain views, we enhance the efficiency of the entire pipeline. To determine which views should be substituted, we propose a view selection strategy. Combined with view latent substitution, this strategy significantly speeds up the whole pipeline with minimal performance loss.
In summary, our main contributions are as follows:
We propose a LAtent World model for self-supervised learning that enhances the training of end-to-end autonomous driving framework.
Based on the latent world model, we further propose a view selection strategy, which greatly accelerates the pipeline while incurring minimal performance loss.
Our framework LAW achieves state-of-the-art results on both open-loop and closed-loop benchmarks without manual annotations.
Related Works
We divide end-to-end autonomous driving methods into two categories, explicit methods and implicit methods, depending on whether performing traditional perception tasks.
Explicit end-to-end methods perform multiple perception tasks simultaneously, such as detection , tracking , map segmentation and occupancy prediction . As a pioneering work, P3 employs a differentiable semantic occupancy representation as a cost factor in the motion planning process. Following this, ST-P3 introduces a spatial-temporal feature learning approach to generate more representative features for perception, prediction, and planning tasks concurrently. Then, many works focus on performing detection and BEV map segmentation tasks based on the BEV feature map. As a representative, UniAD integrates multiple modules, including tracking and motion prediction, to support goal-driven planning. VAD explores vectorized scene representation for planning purposes.
Implicit end-to-end methods present a promising direction as they avoid utilizing a large number of perception annotations. Early implicit end-to-end methods primarily relied on reinforcement learning. For instance, MaRLn designed a reinforcement learning algorithm based on implicit affordances, while LBC trained a reinforcement learning expert using privileged (ground-truth perception) information. Using trajectory data generated by the reinforcement learning expert, TCP combined a trajectory waypoint branch with a direct control branch, achieving good performance. However, implicit end-to-end methods often suffer from inadequate scene representation capabilities. Our work aims to address this issue through latent prediction.
2 World Model in Autonomous Driving
Existing world models in autonomous driving can be categorized into two types: image-based world models and occupancy-based world models. Image-based world models aim to enrich the autonomous driving dataset through generative approaches. GAIA-1 is a generative world model that utilizes video, text, and action inputs to create realistic driving scenarios. MILE produces urban driving videos by leveraging 3D geometry as an inductive bias. Drive-WM utilizes a diffusion model to predict future images and plans based on these predicted images. Copilot4D tokenizes sensor observations with VQVAE and then predicts the future via discrete diffusion. Another category involves occupancy-based world models . OccWorld and DriveWorld use the world model to predict the occupancy, which requires occupancy annotations. On the contrary, our proposed latent world model requires no manual annotations.
Preliminary
End-to-End Autonomous Driving In the task of end-to-end autonomous driving, the objective is to estimate the future trajectory of the ego vehicle in the form of waypoints. Formally, let be the set of surrounding multi-view images captured at time step . We expect the model to predict a sequence of waypoints , where each waypoint represents the predicted BEV position of the ego vehicle at time step . represents the number of future positions of the ego vehicle that the model aims to predict.
World Model In autonomous driving tasks, a world model aims to predict future states based on the current state and actions. To be specific, let represent the features extracted from the current frame at time step , denote the sequence of planned waypoints by the planner, the world model predicts features of the future frame using and .
Method
The overall methodology is divided into three parts. First, we develop a strong and general end-to-end planner in Sec. 4.1 to extract latentThe terms ”latent” and ”latent feature” convey the same meaning.. Next, based on the end-to-end planner, we introduce a world model to predict latent in Sec. 4.2. Finally, the predicted latents can substitute for some unimportant latents so we propose a view selection approach in Sec. 4.3.
To extract effective latent feature, we introduce a general and strong end-to-end planner. Initially, -view images are processed through an image backbone to extract their respective feature representations. Following PETR , we generate 3D position embeddings for these image features. These position embeddings are integrated with the image features to uniquely identify each view. The enriched image features are denoted as .
Then, we employ a view attention mechanism to compress into observed view latent . Here, we use the term "observed" to distinguish this view latent from others discussed later in the paper. To be specific, for views, there are learnable view queries . Each view query undergoes a cross-attention with its corresponding image feature , resulting in observed view latent , where
serves as the key and value of the cross attention. Next, we perform the temporal aggregation to the observed view latent. The observed view latent is enhanced by a historical view latent , which is generated from the previous frame (the details will be discussed in Sec. 4.2). In this way, we have
where is named enhanced view latent. Given , we develop a waypoint decoder to decode waypoints. This module uses waypoint queries to extract relevant information from . To be specific, we initialize waypoint queries, , where each query is a learnable embedding. These waypoint queries interact with through a cross-attention mechanism. The updated waypoint queries are then passed through an MLP head to output the waypoints , which is formulated as:
During training, we use the L1 loss to measure the discrepancy between the predicted waypoints and the ground truth waypoints as:
The proposed end-to-end planner extracts the latent simply and effectively, which serves as a good testbed of the latent world model.
2 World Model for Latent Prediction
where denotes the concatenating operation. The overall action-based view latent is . Subsequently, given , we obtain the predicted view latent of the frame by the latent world model:
The network architecture of the latent world model is a transformer decoder, which consists of two blocks. Each block contains a self-attention and FFN module. The self-attention is performed in the view dimension. During training, we use the end-to-end planner to extract the observed view latent of frame . serves as the supervision of using an L2 loss function:
where and .
Besides, given , we encode the temporal information into the history view latent . is used to enhance the observed view latent through Eq. (2). To be specific, we conduct self-attention on in the view dimension, obtaining
and serve distinct functions. aims to encode temporal information as a residual, whereas is designed to predict the view latent of the future frame. In addition, serves as a good substitute for the observed view latent of future frames, which inspires us to propose the concept of view selection with latent substitution.
3 View Selection via Latent Substitution
We propose a view selection approach thanks to the effective view latent predicted by the world model. Taking multi-view videos as input, this approach dynamically selects some informative views to extract features. The other views are not processed and their corresponding view latents are substituted by the predicted view latent from the world model. As shown in Fig. 3, this section consists of three components. First, given several potential view selection strategies, the Selection Reward Prediction component predicts the rewards of these strategies and chooses the strategy with the highest reward. Then, the Planner with Selected Views predicts the trajectory given the selected views. During training, we propose a Selection Reward Labeling module, which assigns a reward label to each selection strategy.
Selection Reward Prediction As illustrated in Fig. 3(a) and (b), we introduce a reward prediction module designed to estimate the reward associated with each selection strategy. The reward quantitatively reflects the effectiveness of the planning outcomes obtained using each strategy. In detail, we define selection queries. These selection queries correspond to potential selection strategies. Each selection query is a learnable embedding. Each strategy selects specific views for processing while discarding the rest. Then, we update the selection queries by performing cross-attention between the queries and the view latent predicted by the world model. The updated selection queries are fed into an MLP head to predict the rewards. Given these rewards, we choose the strategy with the highest predicted reward. The strategy is chosen at the frame and the views selected by this strategy serve as the input of the Planner with Selected Views at frame .
Planner with Selected Views This planner takes selected views as input to produce waypoints as Fig. 3(c) shows. It shares the same weights as the planner in Sec. 4.1. To be specific, Let represent the total number of views. Define as the set of indices for views that are selected. is denoted as the complement of . Given and , the combined view latents are formulated as:
where denotes the concatenation operation over the latents in the view dimension. Then, the combined view latents are passed through the same pipeline in Sec. 4.1 to predict trajectories. Based on the Planner with Selected Views, we propose a selection reward labeling module to label the selection strategies.
Selection Reward Labeling We introduce the reward labeling approach as Fig 3(d) shows. Specifically, for the -th strategy, the corresponding selected views are fed into the Planner with Selected Views to predict waypoints . The reward label of -th strategy is defined as the L2 distance between the predicted waypoints and the ground truth waypoints as:
The larger is, the closer are to the ground truth waypoints . During training, we use the L1 loss to learn the rewards, formulated as
In summary, the total loss of our framework is:
where is an optional loss depending on whether using the view selection approach. The weight of each loss is discussed in the implementation details.
Experiments
Open-loop Benchmark The open-loop benchmark uses recorded video streams of expert drivers along with the corresponding trajectories of the ego vehicle. We conduct our experiments on the nuScenes dataset , which comprises 1,000 driving scenes. In line with previous works , we employ Displacement Error (DE) and Collision Rate (CR) to comprehensively evaluate the planning performance. The Displacement Error measures the L2 distance between the predicted trajectory and the GT trajectory. The Collision Rate quantifies the rate of collisions that occur with other objects when following the predicted trajectory.
Closed-loop Benchmark Closed-loop evaluation is essential to autonomous driving as it constantly updates the sensor inputs based on the driving actions. The training dataset is collected from the CARLA simulator (version 0.9.10.1) using the teacher model Roach following , resulting in 189K frames. We use the widely-used Town05 Long benchmark to assess the closed-loop driving performance. We use the official metrics: Route Completion (RC) represents the percentage of the route completed by the autonomous agent. Infraction Score (IS) quantifies the number of infractions as well as violations of traffic rules. A higher Infraction Score indicates better adherence to safe driving practices. Driving Score (DS) is the primary metric used to evaluate overall performance. It is calculated as the product of Route Completion and Infraction Score.
Implementation Details The default configuration of LAW does not include view selection unless specified. For the open-loop benchmark, we use Swin-Transformer-Tiny (Swin-T) as the backbone. The input image is resized to 800 320. We employ a Cosine Annealing learning rate schedule with an initial learning rate of 5e-5. AdamW optimizer is utilized with a weight decay of 0.01 and we train the model with batch size 8 for 12 epochs on 8 RTX 3090 GPUs. The weight of the waypoint loss and latent prediction loss are set to 1.0. As for the planner with selected views, we finetune it with the reward loss based on the LAW. We set the initial learning rate to 5e-6 and train for an additional 6 epochs. The weight of the reward loss is set to 1.0. For the closed-loop benchmark. And we use ResNet-34 as the backbone following for a fair comparison. We use the TCP head following . The size of the input image is 900 256. The optimizer is Adam. The learning rate is set to 1e-4 and the weight decay is 1e-7. We train the model with batch size 128 for 60 epochs. The learning rate is reduced by a factor of 2 after 30 epochs.
2 Comparison with State-of-the-art Methods
For the open-loop benchmark, we compare LAW with several state-of-the-art methods, including UniAD , VAD on the nuScenes dataset. The results are summarized in Table 1. LAW outperforms UniAD and VAD in terms of the average L2 displacement error over 1s, 2s, and 3s prediction horizons. Moreover, our method achieves remarkable real-time performance with a latency of 30.9 ms, highlighting the efficiency of our approach. For the closed-loop benchmark, as shown in Table 2, our proposed method outperforms all existing methods. Notably, our approach surpasses previous leading methods such as ThinkTwice and DriveAdapter , which incorporate extensive supervision from depth estimation, semantic segmentation, and map segmentation.
3 Ablation Study
In this ablation study, we investigate the effectiveness of the latent prediction. The results are presented in Table 3. Initially, we only use the view latent as input of the world model, which means omitting the predicted trajectory component.
As shown in the third row of Table 3, this approach results in a slight performance improvement compared with the model without the latent prediction. When we include the predicted trajectory as part of the input (fourth row of Table 3), performance is significantly enhanced. It shows that an accurate prediction of future latents requires the incorporation of driving actions, highlighting the rationality of using the latent world model. Additionally, we provide an ablation study on the latent prediction in the closed-loop setting, as depicted in Table 4. Notably, we observed substantial improvements in the Infraction Score. This indicates that the capability to predict future scenarios effectively aids in mitigating potential collisions.
Network Architecture of Latent World Model To validate the impact of the network architecture of the latent world model, we conduct experiments as shown in Table 5. Firstly, it is evident that a single-layer neural network, represented as Linear Projection, is not adequate for fulfilling the functions of the world model, resulting in significantly degraded performance. The two-layer MLP shows considerable improvement in performance. However, it lacks the capability to facilitate interactions among latents from different views. Therefore, we use the transformer decoder as our default network architecture, which achieves the best results among the tested architectures. This suggests that for any particular view, incorporating information from multiple adjacent views can enhance the prediction of its future latent.
The Time Horizon of Latent World Model In this experiment, the world model predicts latent features at three distinct future time horizons: 0.5 seconds, 1.5 seconds, and 3.0 seconds. This corresponds to the first, third, and sixth future frames from the current frame, given that keyframes occur every 0.5 seconds in the nuScenes dataset. The results, displayed in Table 6, show that the model achieves the best performance at the 1.5-second horizon. The reasons are as follows. The 0.5-second interval typically presents scenes with minimal changes, providing insufficient dynamic content to improve feature learning. In contrast, the 3.0-second interval increases the complexity of the prediction task, which hinders better feature learning. This conclusion aligns with observations from MAE , where both excessively low and high mask ratios negatively impact the ability of the network.
View Selection To ablate the effectiveness of our view selection approach, we conduct the experiments shown in Table 7. We train our model with the view selection module and then test it with several strategies: 1) front view and a random view, 2) front view and a view selected by our view selection module, 3) front view and a view selected based on the rewards label in Sec. 4.3. This reward label is generated with the help of GT trajectory and this experiment serves as the upper bound. 4) all six views. The reason for fixing the front view will be discussed in the appendix A.1. The results demonstrate that the selection made by our view selection module significantly outperforms random selections and closely approaches the upper bound set by the GT.
Conclusion and Limitation
In conclusion, this paper introduces a novel self-supervised approach using the latent world model. This approach enhances the learning of scene representations in end-to-end autonomous driving systems without costly annotations. Although our method has demonstrated promising outcomes on current benchmarks, it is constrained by the limited volume of data utilized. In future work, we aim to enhance the scalability of our approach by applying it to larger and more diverse datasets. Leveraging large-scale data, we intend to employ the latent world model for pretraining.
References
Appendix A Appendix
In our view selection approach, we always use the front view and dynamically choose one additional view from the other five views. The reason for doing this is driven by the following experiments.
Fixing Front View Helps This section presents the experimental justification for always choosing the front camera view as one of the input views. Specifically, given the trained end-to-end planner without reward loss, we conduct two view-selection strategies on it. The first strategy always uses the front camera and selects one random view from the remaining five cameras. The second strategy randomly chooses two cameras. The results are shown in Table 8. This experiment shows that the strategy of fixing the front view and randomly selecting an additional view outperforms the random selection of two views. This superiority can be attributed to the fact that the front view can provide more crucial information for the planning task, especially in forward-driving scenarios.
Reducing the number of fixed cameras can also alleviate the computational burden. Therefore, it is natural to question how the view selection approach compares to a configuration utilizing a reduced number of fixed cameras. To investigate this, we carry out experiments as illustrated in Fig. 4 (a). The specific settings of the experiments are as follows. We set up four groups of experiments with the number of views used being 1, 2, 4, and 6. When using only one view, we always use the front camera. For two views, the fixed-view model is trained and tested using the front and back cameras, while the dynamic-view model (i.e., our method) fixes the front camera and then selects the most informative view from the remaining five cameras. For four views, the fixed-view model uses the front, front-left, front-right, and back cameras, while the dynamic-view model fixes the front camera and chooses three additional cameras from the remaining five. Finally, for six views, we use all available cameras. The results demonstrate that our dynamic selection method consistently outperforms fixed-view settings with the same number of views. This indicates that our method is capable of selecting the informative views from multiple options.
Latency Breakdown Our Planner with Selected Views attains an impressive speed of 32.4 FPS. Here we present the detailed latency of each module. We test it on the NVIDIA Geforce RTX 3090 GPU with batch size 1. The code is based on the mmdetection3d . The specific latency associated with each module in our model is detailed in Fig. 4 (b). As illustrated in the figure, the backbone constitutes the majority of the model’s latency. Reducing the number of input views leads to a linear decrease in the cost of the backbone. Since we only have two views passing through the backbone, our view-selection method substantially boosts the efficiency.
As depicted in Fig. 5, we present a visualization analysis of four typical cases. From these visualizations, we derive two key insights: 1) Our method has a preference for views with visually salient objects. As demonstrated in cases (a), (b), and (c), the model tends to select views that feature groups of people, vehicles with the potential to cut in, or vehicles at risk of a rear-end collision. This preference arises because such objects have a significant impact on driving behavior. Through reward learning, our model can identify the view that is most important at a certain moment, by the clues offered by these salient objects. 2) Our model learns prior knowledge of specific scenarios. Human drivers have prior knowledge about the world, such as the expectation of pedestrians suddenly appearing at crosswalks, which necessitates slower driving. Our view latent reconstruction module can also learn similar priors. For instance, in case (d), the view selection model, aided by the reconstruction module, reasonably focuses on the crosswalk.