NEAT: Neural Attention Fields for End-to-End Autonomous Driving
Kashyap Chitta, Aditya Prakash, Andreas Geiger
Introduction
Navigating large dynamic scenes for autonomous driving requires a meaningful representation of both the spatial and temporal aspects of the scene. Imitation Learning (IL) by behavior cloning has emerged as a promising approach for this task . Given a dataset of expert trajectories, a behavior cloning agent is trained through supervised learning, where the goal is to predict the actions of the expert given some sensory input regarding the scene . To account for the complex spatial and temporal scene structure encountered in autonomous driving, the training objectives used in IL-based driving agents have evolved by incorporating auxiliary tasks. Pioneering methods, such as CILRS , use a simple self-supervised auxiliary training objective of predicting the ego-vehicle velocity. Since then, more complex training signals aiming to reconstruct the scene have become common, e.g. image auto-encoding , 2D semantic segmentation , Bird’s Eye View (BEV) semantic segmentation , 2D semantic prediction , and BEV semantic prediction . Performing an auxiliary task such as BEV semantic prediction, which requires the model to output the BEV semantic segmentation of the scene at both the observed and future time-steps, incorporates spatiotemporal structure into the intermediate representations learned by the agent. This has been shown to lead to more interpretable and robust models . However, so far, this has only been possible with expensive LiDAR and HD map-based network inputs which can be easily projected into the BEV coordinate frame.
The key challenge impeding BEV semantic prediction from camera inputs is one of association: given a BEV spatiotemporal query location in the scene (e.g. 2 meters in front of the vehicle, 5 meters to the right, and 2 seconds into the future), it is difficult to identify which image pixels to associate to this location, as this requires reasoning about 3D geometry, scene motion, ego-motion, and intention, as well as interactions between scene elements. In this paper, we propose NEural ATtention fields (NEAT), a flexible and efficient feature representation designed to address this challenge. Inspired by implicit shape representations , NEAT represents large dynamic scenes with a fixed memory footprint using a multi-layer perceptron (MLP) query function. The core idea is to learn a function from any query location to an attention map for features obtained by encoding the input images. NEAT compresses the high-dimensional image features into a compact low-dimensional representation relevant to the query location , and provides interpretable attention maps as part of this process, without attention supervision . As shown in Fig. 1, the output of this learned MLP can be used for dense prediction in space and time. Our end-to-end approach predicts waypoint offsets to solve the main trajectory planning task (described in detail in Section 3), and uses BEV semantic prediction as an auxiliary task.
Using NEAT intermediate representations, we train several autonomous driving models for the CARLA driving simulator . We consider a more challenging evaluation setting than existing work based on the new CARLA Leaderboard with CARLA version 0.9.10, involving the presence of multiple evaluation towns, new environmental conditions, and challenging pre-crash traffic scenarios. We outperform several strong baselines and match the privileged expert’s performance on our internal evaluation routes. On the secret routes of the CARLA Leaderboard, NEAT obtains competitive driving scores while incurring significantly fewer infractions than existing methods.
Contributions: (1) We propose an architecture combining our novel NEAT feature representation with an implicit decoder for joint trajectory planning and BEV semantic prediction in autonomous vehicles. (2) We design a challenging new evaluation setting in CARLA consisting of 6 towns and 42 environmental conditions and conduct a detailed empirical analysis to demonstrate the driving performance of NEAT. (3) We visualize attention maps and semantic scene interpolations from our interpretable model, yielding insights into the learned driving behavior. Our code is available at https://github.com/autonomousvision/neat.
Related Work
Implicit Scene Representations: The geometric deep learning community has pioneered the idea of using neural implicit representations of scene geometry. These methods represent surfaces as the boundary of a neural classifier or zero-level set of a signed distance field regression function . They have been applied for representing object texture , dynamics and lighting properties . Recently, there has been progress in applying these representations to compose objects from primitives , and to represent larger scenes, both static and dynamic . These methods obtain high-resolution scene representations while remaining compact, due to the constant memory footprint of the neural function approximator. While NEAT is motivated by the same property, we use the compactness of neural approximators to learn better intermediate features for the downstream driving task.
End-to-End Autonomous Driving: Learning-based autonomous driving is an active research area . IL for driving has advanced significantly and is currently employed in several state-of-the-art approaches, some of which predict waypoints , whereas others directly predict vehicular control . While other learning-based driving methods such as affordances and Reinforcement Learning could also benefit from a NEAT-based encoder, in this work, we apply NEAT to improve IL-based autonomous driving.
BEV Semantics for Driving: A top-down view of a street scene is powerful for learning the driving task since it contains information regarding the 3D scene layout, objects do not occlude each other, and it represents an orthographic projection of the physical 3D space which is better correlated with vehicle kinematics than the projective 2D image domain. LBC exploits this representation in a teacher-student approach. A teacher that learns to drive given BEV semantic inputs is used to supervise a student aiming to perform the same task from images only. By doing so, LBC achieves state-of-the-art performance on the previous CARLA version 0.9.6, showcasing the benefits of the BEV representation. NEAT differs from LBC by directly learning in BEV space, unlike the LBC student model which learns a classical image-to-trajectory mapping.
Other works deal with BEV scenes, e.g., obtaining BEV projections or BEV semantic predictions from images, but do not use these predictions for driving. More recently, LSS and OGMs demonstrated joint BEV semantic reconstruction and driving from camera inputs. Both methods involve explicit projection based on camera intrinsics, unlike our learned attention-based feature association. They only predict semantics for static scenes, while our model includes a time component, performing prediction up to a fixed horizon. Moreover, unlike us, they only evaluate using offline metrics which are known to not necessarily correlate well with actual downstream driving performance . Another related work is P3 which jointly performs BEV semantic prediction and driving. In comparison to P3 which uses expensive LiDAR and HD map inputs, we focus on image modalities.
Method
A common approach to learning the driving task from expert demonstrations is end-to-end trajectory planning, which uses waypoints as outputs. A waypoint is defined as the position of the vehicle in the expert demonstration at time-step , in a BEV projection of the vehicle’s local coordinate system. The coordinate axes are fixed such that the vehicle is located at at the current time-step , and the front of the vehicle is aligned along the positive y-axis. Waypoints from a sequence of future time-steps form a trajectory that can be used to control the vehicle, where is a fixed prediction horizon.
As our agent drives through the scene, we collect sensor data into a fixed-length buffer of time-steps, where each comes from one of sensors. The final frame in the buffer is always the current time-step (). In practice, the sensors are RGB cameras, the standard input modality in existing work on CARLA . By default, we use cameras, one oriented forward and the others 60 degrees to the left and right. After cropping these camera images to remove radial distortion, these images together provide a full 180∘ view of the scene in front of the vehicle. While NEAT can be applied with different buffer sizes, we focus in our experiments on the setting where the input is a single frame (), as several studies indicate that using historical observations can be detrimental to the driving task .
In addition to waypoints, we use BEV semantic prediction as an auxiliary task to improve driving performance. Unlike waypoints which are small in number (e.g. ) and can be predicted discretely, BEV semantic prediction is a dense prediction task, aiming to predict semantic labels at any spatiotemporal query location bounded to some spatial range and the time interval . Predicting both observed () and future () semantics provides a holistic understanding of the scene dynamics. Dynamics prediction from a single input frame is possible since the orientation and position of vehicles encodes information regarding their motion .
The coordinate system used for BEV semantic prediction is the same as the one used for waypoints. Thus, if we frame waypoint prediction as a dense prediction task, it can be solved simultaneously with BEV semantic prediction using the proposed NEAT as a shared representation. Therefore, we propose a dense offset prediction task to locate waypoints as visualized in Fig. 2 using a standard optical flow color wheel . The goal is to learn the field of 2-dimensional offset vectors from query locations to the waypoint (e.g. when and ). In certain situations, future waypoints along different trajectories are plausible (e.g. taking a left or right turn at an intersection), thus it is important to adapt based on the driver intention. We do this by using provided target locations as inputs. Target locations are GPS coordinates provided by a navigational system along the route to be followed. They are transformed to the same coordinate system as the waypoints before being used as inputs. These target locations are sparse and can be hundreds of meters apart. In Fig. 2, the target location to the right of the intersection helps the model decide to turn right rather than proceeding straight. We choose target locations as the method for specifying driver intention as they are the default intention signal in the CARLA simulator since version 0.9.9. In summary, the goal of dense offset prediction is to output for any 5-dimensional query point .
As illustrated in Fig. 3, our architecture consists of three neural networks that are jointly trained for the BEV semantic prediction and dense offset prediction tasks: an encoder , neural attention field , and decoder . In the following, we go over each of the three components in detail.
Encoder: Our encoder takes as inputs the sensor data buffer and a scalar , which is the vehicle velocity at the current time-step . Formally, it is denoted as
Neural Attention Field: While the transformer aggregates features globally, it is not informed by the query and target location. Therefore, we introduce NEAT (Fig. 1), which identifies the patch features from the encoder relevant for making predictions regarding any query point in the scene . It introduces a bottleneck in the network and improves interpretability (Fig. 6). Its operation can be formally described as
The feature is used as the input of along with at the next attention iteration, implementing a recurrent attention loop (see Fig. 3). Note that the dimensionality of is significantly smaller than that of the transformer output , as aggregates information (via Eq. (3)) across sensors , time-steps and patches . For the initial iteration, is set to the mean of (equivalent to assuming a uniform initial attention). We implement as a fully-connected MLP with 5 ResNet blocks of 128 hidden units each, conditioned on using conditional batch normalization (details in supplementary). We share the weights of across all iterations which works well in practice.
Decoder: The final network in our model is the decoder:
2 Training
Sampling: An important consideration is the choice of query samples during training, and how to acquire ground truth labels for these points. Among the 5 dimensions of , and are fixed for any , but , , and can all be varied to access different positions in the scene. Note that in the CARLA simulator, the ground truth waypoint is only available at discrete time-steps, and the ground truth semantic class only at discrete locations. However, this is not an issue for NEAT as we can supervise our model using arbitrarily sparse observations in the space-time volume. We consider semantic classes by default: none, road, obstacle (person or vehicle), red light, and green light. The location and state of the traffic light affecting the ego-vehicle are provided by CARLA. We use this to set the semantic label for points within a fixed radius of the traffic light pole to the red light or green light class, similar to . In our work, we focus on the simulated setting where this information is readily available, though BEV semantic annotations of objects (obstacles, red lights, and green lights) can also be obtained for real scenes using projection. The only remaining labels required by NEAT (for the road class) can be obtained by fitting the ground plane to LiDAR sweeps in a real dataset or more directly from localized HD maps.
We acquire these BEV semantic annotations from CARLA up to a fixed prediction horizon after the current time-step and register them to the coordinate frame of the ego-vehicle at . is a hyper-parameter that can be used to modulate the difficulty of the prediction task. From the aligned BEV semantic images, we only consider points approximately in the field-of-view of our camera sensors. We use a range of 50 meters in front of the vehicle and 25 meters to either side (detailed sensor configurations are provided in the supplementary).
Since semantic class labels are typically heavily imbalanced, simply using all observations for training (or sampling a random subset) would lead to several semantic classes being under-represented in the training distribution. We use a class-balancing sampling heuristic during training to counter this. To sample points for semantic classes, we first group all the points from all the time-steps in the sequence into bins based on their semantic label. We then attempt to randomly draw points from each bin, starting with the class having the least number of available points. If we are unable to sample enough points from any class, this difference is instead sampled from the next bin, always prioritizing classes with fewer total points.
We obtain the offset ground truth for each of these sampled points by collecting the ground truth waypoints for the time-steps around each frame. The offset label for each of the sampled points is calculated as its difference from the ground truth waypoint at the corresponding time-step . Being a regression task, we find that offset prediction does not benefit as much from a specialized sampling strategy. Therefore, we use the same points for supervising the offsets even though they are sampled based on semantic class imbalance, improving training efficiency.
Loss: For each of the sampled points, the decoder makes predictions and at each of the attention iterations. The encoder, NEAT and decoder are trained jointly with a loss function applied to these predictions:
where is the distance between the true offset and predicted offset , is the cross-entropy between the true semantic class and predicted class , is a weight between the semantic and offset terms, and is used to down-weight predictions made at earlier iterations (). These intermediate losses improve performance, as we show in our experiments.
3 Controller
To drive the vehicle at test time, we generate a red light indicator and waypoints from our trained model; and convert them into steer, throttle, and brake values. For the red light indicator, we uniformly sample a sparse grid of points in at the current time-step , in the area 50 meters to the front and 25 meters to the right side of the vehicle. We append the target location to these grid samples to obtain 5-dimensional queries that can be used as NEAT inputs. From the semantic prediction obtained for these points at the final attention iteration , we set the red light indicator as 0 if none of the points belongs to the red light class, and 1 otherwise. In our ablation study, we find this indicator to be important for performance.
To generate waypoints, we sample a uniformly spaced grid of points in a square region of side meters centered at the ego-vehicle at each of the future time-steps . Note that predicting waypoints with a single query point () is possible, but we use a grid for robustness. After encoding the sensor data and performing attention iterations, we obtain for each of the query points at each of the future time-steps. We offset the location coordinates of each query point towards the waypoint by adding , effectively obtaining the waypoint prediction for that sample, i.e. . After this offset operation, we average all waypoint predictions at each future time instant, yielding the final waypoint predictions . To obtain the throttle and brake values, we compute the vectors between waypoints of consecutive time-steps and input the magnitude of these vectors to a longitudinal PID controller along with the red light indicator. The relative orientation of the waypoints is input to a lateral PID controller for turns. Please refer to the supplementary material for further details on both controllers.
Experiments
In this section, we describe our experimental setting, showcase the driving performance of NEAT in comparison to several baselines, present an ablation study to highlight the importance of different components of our architecture, and show the interpretability of our approach through visualizations obtained from our trained model.
Task: We consider the task of navigation along pre-defined routes in CARLA version 0.9.10 . A route is defined by a sequence of sparse GPS locations (target locations). The agent needs to complete the route while coping with background dynamic agents (pedestrians, cyclists, vehicles) and following traffic rules. We tackle a new challenge in CARLA 0.9.10: each of our routes may contain several pre-defined dangerous scenarios (e.g. unprotected turns, other vehicles running red lights, pedestrians emerging from occluded regions to cross the road).
Routes: For training data generation, we store data using an expert agent along routes from the 8 publicly available towns in CARLA, randomly spawning scenarios at several locations along each route. We evaluate NEAT on the official CARLA Leaderboard , which consists of 100 secret routes with unknown environmental conditions. We additionally conduct an internal evaluation consisting of 42 routes from 6 different CARLA towns (Town01-Town06). Each route has a unique environmental condition combining one of 7 weather conditions (Clear, Cloudy, Wet, MidRain, WetCloudy, HardRain, SoftRain) with one of 6 daylight conditions (Night, Twilight, Dawn, Morning, Noon, Sunset). Additional details regarding our training and evaluation routes are provided in the supplementary. Note that in this new evaluation setting, the multi-lane road layouts, distant traffic lights, high density of background agents, diverse daylight conditions, and new metrics which strongly penalize infractions (described below) make navigation more challenging, leading to reduced scores compared to previous CARLA benchmarks .
Metrics: We report the official metrics of the CARLA Leaderboard, Route Completion (RC), Infraction Score (IS)The Leaderboard refers to this as infraction penalty. We use the terminology ‘score’ since it is a multiplier for which higher values are better. and Driving Score (DS). For a given route, RC is the percentage of the route distance completed by the agent before it deviates from the route or gets blocked. IS is a cumulative multiplicative penalty for every collision, lane infraction, red light violation, and stop sign violation. Please refer to the supplementary material for additional details regarding the penalty applied for each kind of infraction. Finally, DS is computed as the RC weighted by the IS for that route. After calculating all metrics per route, we report the mean performance over all 42 routes. We perform our internal evaluation three times for each model and report the mean and standard deviation for all metrics.
Baselines: We compare our approach against several recent methods. CILRS learns to directly predict vehicle controls (as opposed to waypoints) from visual features while being conditioned on a discrete navigational command (follow lane, change lane left/right, turn left/right). It is a widely used baseline for the old CARLA version 0.8.4, which we adapted to the latest CARLA version. LBC is a knowledge distillation approach where a teacher model with access to ground truth BEV semantic segmentation maps is first trained using expert supervision to predict waypoints, followed by an image-based student model which is trained using supervision from the teacher. It is the state-of-the-art approach on CARLA version 0.9.6. We train LBC on our dataset using the latest codebase provided by the authors for CARLA version 0.9.10. AIM is an improved version of CILRS, where a GRU decoder regresses waypoints. To assess the effects of different forms of auxiliary supervision, we create 3 multi-task variants of AIM (AIM-MT). Each variant adds a different auxiliary task during training: (1) 2D semantic segmentation using a deconvolutional decoder, (2) BEV semantic segmentation using a spatial broadcast decoder , and (3) both 2D depth estimation and 2D semantic segmentation as in . We also replace the CILRS backbone of Visual Abstractions with AIM, to obtain AIM-VA. This approach generates 2D segmentation maps from its inputs which are then fed into the AIM model for driving. Finally, we report results for the privileged Expert used for generating our training data.
Implementation: By default, NEAT’s transformer uses layers with 4 parallel attention heads. Unless otherwise specified, we use , , , , , , , , and . We use a weight of on the loss, set for the intermediate iterations (), and set . For a fair comparison, we choose the best performing encoders for each model among ResNet-18, ResNet-34, and ResNet-50 (NEAT uses ResNet-34). Moreover, we chose the best out of two different camera configurations ( and ) for each model, using a late fusion strategy for combining sensors in the baselines when we set . Additional details are provided in the supplementary.
Our results are presented in Table 1. Table 1(a) focuses on our internal evaluation routes, and Table 1(b) on our submissions to the CARLA Leaderboard. Note that we could not submit all the baselines from Table 1(a) or obtain statistics for multiple evaluations of each model on the Leaderboard due to the limited monthly evaluation budget (200 hours).
Importance of Conditioning: We observe that in both evaluation settings, CILRS and LBC perform poorly. However, a major improvement can be obtained with a different conditioning signal. CILRS uses discrete navigational commands for conditioning, and LBC uses target locations represented in image space. By using target locations in BEV space and predicting waypoints, AIM and NEAT can more easily adapt their predictions to a change in driver intention, thereby achieving better scores. We show this adaptation of predictions for NEAT in Fig. 4, by predicting semantics and waypoint offsets for different target locations and time steps . The waypoint offset formulation of NEAT introduces a bias that leads to smooth trajectories between consecutive waypoints (red lines in ) towards the provided target location in blue.
AIM-MT and Expert: We observe that AIM-MT is a strong baseline that becomes progressively better with denser forms of auxiliary supervision. The final variant which incorporates dense supervision of both 2D depth and 2D semantics achieves similar performance to NEAT on our 42 internal evaluation routes but does not generalize as well to the unseen routes of the Leaderboard (Table 1(b)). Interestingly, in some cases, AIM-MT and NEAT match or even exceed the performance of the privileged expert in Table 1(a). Though our expert is an improved version of the one used in , it still incurs some infractions due to its reliance on relatively simple heuristics and driving rules.
Leaderboard Results: While NEAT is not the best performing method in terms of DS, it has the safest driving behavior among the top three methods on the Leaderboard, as evidenced by its higher IS. WOR is concurrent work that supervises the driving task with a Q function obtained using dynamic programming, and MaRLn is an extension of the Reinforcement Learning (RL) method presented in . WOR and MaRLn require 1M and 20M training frames respectively. In comparison, our training dataset only has 130k frames, and can potentially be improved through orthogonal techniques such as DAgger .
2 Ablation Study
In Fig. 5, we compare multiple variants of NEAT, varying the following parameters: training seed, semantic class count (), attention iterations (), prediction horizon (), input sensor count (), transformer layers (), and loss weights (). While a detailed analysis regarding each factor is provided in the supplementary, we focus here on four variants in particular: Firstly, we observe that different random training seeds of NEAT achieve similar performance, which is a desirable property not seen in all end-to-end driving models . Second, as observed by , 2D semantic models (such as AIM-VA and AIM-MT) rely heavily on lane marking annotations for strong performance. We observe that these are not needed by NEAT for which the default configuration with 5 classes outperforms the variant that includes lane markings with 6 classes. Third, in the shorter horizon variant () with only 2 predicted waypoints, we observe that the output waypoints do not deviate sharply enough from the vertical axis for the PID controller to perform certain maneuvers. It is also likely that the additional supervision provided by having a horizon of in our default configuration has a positive effect on performance. Fourth, the gain of the default NEAT model compared to its version without the semantic loss () is 30%, showing the benefit of performing BEV semantic prediction and trajectory planning jointly.
Runtime: To analyze the runtime overhead of NEAT’s offset prediction task, we now create a hybrid version of AIM and NEAT. This model directly regresses waypoints from NEAT’s encoder features using a GRU decoder (like AIM) instead of predicting offsets. We still use a semantic decoder at train time supervised with only the cross-entropy term of Eq. (5). At test time, the average runtime per frame of the hybrid model (with the semantics head discarded) is 15.92 ms on a 3080 GPU. In comparison, the default NEAT model takes 30.37 ms, i.e., both approaches are real-time even with un-optimized code. Without the compute-intensive red light indicator, NEAT’s runtime is only 18.60 ms. Importantly, NEAT (DS = 65.10) significantly outperforms the AIM-NEAT hybrid model (DS = 33.63). This shows that NEAT’s attention maps and location-specific features lead to improved waypoint predictions.
3 Visualizations
Our supplementary videohttps://www.youtube.com/watch?v=gtO-ghjKkRs contains qualitative examples of NEAT’s driving capabilities. For the first route in the video, we visualize attention maps for different locations on the route in Fig. 6. For each frame in the video, we randomly sample BEV locations and pass them through the trained NEAT model until one of the locations corresponds to the class obstacle, red light, or green light. Four such frames are shown in Fig. 6. We observe a common trend in the attention maps: NEAT focuses on the image corresponding to the object of interest, albeit sometimes at a slightly different location in the image. This can be attributed to the fact that NEAT’s attention maps are over learned image features that capture information aggregated over larger receptive fields. To quantitatively evaluate this property, we extract the image patch which NEAT assigns the highest attention weight for one random location in each scene of our validation set and analyze its ground truth 2D semantic segmentation labels. The semantic class predicted by NEAT for is present in the 2D patch in 79.67% of the scenes.
Conclusion
In this work, we take a step towards interpretable, high-performance, end-to-end autonomous driving with our novel NEAT feature representation. Our approach tackles the challenging problem of joint BEV semantic prediction and vehicle trajectory planning from camera inputs and drives with the highest safety among state-of-the-art methods on the CARLA simulator. NEAT is generic and flexible in terms of both input modalities and output task/supervision and we plan to combine it with orthogonal ideas (e.g., DAgger, RL) in the future.
Acknowledgements: This work is supported by the BMBF through the Tübingen AI Center (FKZ: 01IS18039B) and the BMWi in the project KI Delta Learning (project number: 19A19013O). We thank the International Max Planck Research School for Intelligent Systems (IMPRS-IS) for supporting Kashyap Chitta. The authors also thank Micha Schilling for his help in re-implementing AIM-VA.