End-to-end Contextual Perception and Prediction with Interaction Transformer

Lingyun Luke Li, Bin Yang, Ming Liang, Wenyuan Zeng, Mengye Ren, Sean Segal, Raquel Urtasun

I Introduction

Self-driving vehicles (SDVs) have the potential to change the way we live by providing safe, reliable and cost-effective transportation for everyone everywhere. In order to plan a safe maneuver, SDVs have to not only understand the past and the present but also predict the potential future for at least the duration of the motion planning horizon.

Traditional software engineering stacks for self-driving are composed of a set of processes which are executed serially. Objects are first perceived in 3D via object detection to explain the current frame. These objects are then associated with detections from previous frames, and the trajectories are refined typically using a probabilistic filter. The future positions and velocities of these objects are estimated by unrolling the state for the next few seconds, where a dynamical model (e.g., the bicycle model) is typically employed to produce physically realistic trajectories. An alternative approach is to associate an actor with a lane to form a goal (e.g., going straight, turning) and then predict its future trajectory towards that goal. These models are, however, simplistic and have very little information as input. As a consequence, their estimates are often not very accurate.

More sophisticated approaches based on deep learning representations have been developed to increase the accuracy of the predictions. These approaches assume that the vehicle is localized and rasterize in bird’s-eye view (BEV) both a high-definition map (HD map) as well as the past and current detections. A convolutional neural network then learns to produce future trajectories from these rasters. However, all these approaches are slow as computation is not shared between the detection and motion forecasting networks. Instead, the forecasting network takes the raster images as input and performs many layers of computation. As a consequence real-time inference is very hard under this setting. Moreover, as the prediction network does not have access to raw data, errors in the detection phase are very hard if not impossible to recover from due to information loss.

In contrast, FAF developed a single network to jointly perform detection and motion forecasting. This results in better trajectories as the prediction module has access to the raw data. Furthermore, faster inference is achieved as the computation is shared between tasks. IntentNet proposed to output additionally a probability distribution over driving intentions (e.g., stopping, parked, lane change). took this a step further and jointly perform detection, forecasting and motion planning. This resulted in the first end-to-end trainable motion planner that also produces interpretable intermediate representations. While effective, all the aforementioned approaches ignore the statistical dependencies between actors, and instead predict each trajectory (or its probability) independently given the features.

Various approaches have been proposed to model interactions between traffic actors and forecast motion given object detections. Ma et al. formulated the interaction between pedestrians under a game-theoretic framework with fictitious play . introduced a courtesy term to an inverse reinforcement learning (IRL) objective, which better explains real-world human behavior. Other approaches model the interaction implicitly by message passing between actors . However, these methods use past ground truth trajectories of all actors as input. Self-driving cars do not have access to ground truth object trajectories and thus need to handle the uncertainty and noise of perception systems. Therefore, we tackle the end-to-end setting, i.e., from raw sensor data to detection and motion forecasting.

Towards this goal, we design a multi-sensor detector followed by an efficient interaction module that is inspired by the Transformer . We adapt the original architecture, which was developed for sequence modeling, to our problem domain by incorporating a novel pairwise attention mechanism with spatial reasoning via relative positions of the actors. We also observe that the evolution of a trajectory largely depends on the behaviors of other actors in the recent past. To model this fact we propose a recurrent structure with per-time-step refinement to capture temporal dependencies between different prediction time steps. Importantly, our architecture is very scalable since most of the computation scales linearly with the number of traffic actors. As a consequence, our end-to-end model runs at 10 FPS on a 1080Ti GPU. We demonstrate the effectiveness of our approach on the challenging ATG4D and nuScenes datasets. Our experiments show that we outperform the state-of-the-art by a large margin in detection, trajectory accuracy and collision rate.

II Related Work

In this section we review previous research on 3D object detection, trajectory prediction and interaction modeling.

Camera-based approaches take either monocular or stereo images as input. However, their performance is limited due to the difficulty of inferring depth. Recent 3D detectors rely on LiDAR. Different data representations have been developed: 3D voxels , bird’s-eye view (BEV) , range view , and point-wise operators . Approaches that exploit multiple sensors achieve superior performance compared to single-sensor ones. Besides sensors, HD maps provide useful priors .

II-2 Trajectory Prediction

Dynamic models have been used to estimate the future states by propagating the current state over time . However, these approaches are usually too simple to handle complexities in longer horizon prediction. Recent data-driven approaches have shown promising results in generating realistic future trajectories. feeds the past vehicle locations in BEV into an LSTM network. generates BEV raster images that encode surrounding vehicle states and HD maps, and then feeds them into a CNN. These approaches are however slow as the computation is not shared with the perception module.

II-3 Joint Perception and Prediction

End-to-end models, on the contrary, are trained from raw sensor data and the perception and forecasting modules share features. Existing models take multi-sweep LiDAR data and HD maps as input, and solve the two tasks using a single network. In addition, IntentNet further reasons about semantic intentions. NeuralMP further estimates jointly motion planning. However, these methods do not explicitly take into account the interaction between actors, which plays an important role in real-world traffic.

II-4 Interaction Modeling

Various frameworks have been proposed to model multi-actor interactions, e.g., game theory , Graph Neural Networks . An alternative approach is feature aggregation between multiple actors. encodes the pedestrians in a 2D map and uses a CNN to capture the local interactions. propose a fully-connected layer (SocialPool) to aggregate information of neighboring pedestrians, while further modifies the layer to be convolutional (ConvSocialPool). propose explicit attention modules to identify important physical constraints and social neighbors. In contrast, we propose to capture both temporal and spatial dependencies between actors by incorporating a Transformer-like module into a recurrent neural network. Towards this goal, we propose several important architecture modifications of the Transformer to handle our problem domain.

III Interaction Transformer

In this paper, we propose an end-to-end model to solve the tasks of detection and motion forecasting from raw sensor data. Importantly, our model reasons about the interactions between the actors in the scene. We refer the reader to Fig. 1 for the overall model architecture. Our model has two main components: 1) a detection network that fuses information from multi-sensor data (i.e., LiDAR, images, HD maps) and outputs a set of bounding boxes, and 2) a recurrent interactive prediction module that processes interactions between actors and predicts their future trajectories. In the subsequent sections, we first explain our model’s input and output parameterizations, then describe each component of the model in detail, including our learning procedure.

Given a LiDAR sweep, we voxelize the point cloud into a 3D occupancy grid in BEV with a fixed resolution, centered at the ego vehicle. We define each LiDAR point feature to be 11, and compute the voxel feature by summing all nearby point features weighted by their relative positions to the voxel center. This helps preserve fine-grained information. To capture past motion for the task of motion forecasting, we aggregate multiple past LiDAR sweeps (registered to the current frame by ego-motion compensation) by concatenating their voxel representations along the ZZ axis. Following , we exploit both geometric and semantic map priors for better reasoning. In particular, we subtract the ground height from the ZZ value of each LiDAR point before voxelization. In this way, we can remove the variation caused by the ground slope. We also extract semantic priors in the form of road and lane masks. Each of them is a one-channel raster image in BEV depicting the drivable surface and all the lanes respectively (see Fig. 1 left). We augment the LiDAR voxel representation with semantic map priors by simple concatenation.

III-A2 Output Representation

We define the output space in BEV as it enables efficient feature sharing between perception and forecasting. We parameterize the detection as a set of oriented bounding boxes. A detection box is represented with (x,y,w,l,θ)(x,y,w,l,\theta), where (x,y)(x,y) is the box center, (w,l)(w,l) is the box size, and θi\theta_{i} is its orientation. Note that the missing ZZ dimension can be recovered from the ground prior in the HD map. Additionally, we represent the trajectory as a sequence of boxes at future time steps, denoted as {(xi(t),yi(t),θi(t))}\{(x_{i}^{(t)},y_{i}^{(t)},\theta_{i}^{(t)})\}, with t=1,...,Tt=1,...,T. We assume that the objects are rigid and thus their sizes are kept the same across all time steps. Below we use the terms of trajectory prediction and motion forecasting interchangeably.

III-B Object Detection from Multi-Sensors

The sensor-fusion backbone network follows the two-stream architecture of . The BEV stream is a customized 2D CNN that extracts features in BEV space from joint LiDAR and map representation. Inception-like blocks are stacked sequentially with residual connections to extract multi-scale feature maps. The image stream is a ResNet-18 pre-trained on ImageNet , and receives RGB camera images as input. We aggregate multi-scale image feature maps from each ResNet-18 residual block with a feature pyramid network , and fuse the aggregated feature map with the BEV stream via a continuous fusion layer , providing dense fusion. Specifically, image features are first back-projected to BEV space according to the existing LiDAR observation. At BEV locations with no LiDAR points, the image features are interpolated from nearby occupied locations via an MLP. The image and BEV features are then fused by element-wise addition in BEV space. The output feature map from the BEV stream is used to provide multi-sensor features for the detection and prediction modules.

III-B2 Dense BEV Object Detection

As vehicles are relatively similar and do not overlap in BEV space, following RetinaNet , we formulate object detection as dense prediction without any object anchors. We apply several 1×11\times 1 convolutions on top of the BEV feature map, which outputs an 8-channel vector per voxel, representing a confidence score ss and a bounding box parameterized as (dx,dy,w,l,sin⁡2θ,cos⁡2θ,clsθ)(dx,dy,w,l,\sin 2\theta,\cos 2\theta,cls_{\theta}), where (dx,dy)(dx,dy) are the relative position offsets from the voxel center to the box center, (w,l)(w,l) are the box size, and (sin⁡2θ,cos⁡2θ,clsθ)(\sin 2\theta,\cos 2\theta,cls_{\theta}) are used to decode the orientation. Following , we regress to (sin⁡2θ,cos⁡2θ)(\sin 2\theta,\cos 2\theta) to get the orientation without distinguishing flips. To effectively model interaction we need to understand which direction each vehicle is facing. We use an additional classification step clsθcls_{\theta} to classify whether the orientation is in (14π\frac{1}{4}\pi, 34π\frac{3}{4}\pi] . This classification target is better than (0, π\pi] since orientations of many vehicles are similar or opposite to the orientation of the ego vehicle, which are at the boundaries of (0, π\pi]. We finally apply oriented non-maximum suppression (NMS) to remove the duplicates and keep all remaining boxes whose score is above a threshold.

III-C Recurrent Interactive Motion Forecasting

We based our design on two observations. First, the behaviors of actors heavily depend on each other. For example, drivers control the vehicle speed to keep a safe distance to the vehicle ahead. At intersections, drivers typically wait for those that have the right of way. Second, the output at each time step depends on the outputs at previous time steps. We thus propose a recurrent interactive motion forecasting model that 1) jointly reasons about all actors modeling their interactions and 2) iteratively infers the trajectory to capture its sequential nature. Our interaction module is inspired by the Transformer , an architecture developed for machine translation. We adapt it to our motion forecasting task. Below we first review the Transformer module, then describe the differences with our novel Interaction Transformer module, and finally explain the recurrent inference process.

Sequence to sequence models with an encoder-decoder architecture have been predominant in natural language processing. In this context, the Transformer was proposed as an attention mechanism that can be used to draw global dependencies between inputs and outputs, especially for long sequences. The Transformer projects each feature to a query and a key-value pair, which are all vectors. For each query, it computes a set of attentional weights using a compatibility function between the query and the set of keys. The output feature is then the sum of values weighted by the attention, plus some nonlinear transformations.

where a softmax function is used to add a sum-to-one normalization to the attentional weights of a query (each row of QKTQK^{T}). The scaling factor 1dk\frac{1}{\sqrt{d_{k}}} is used to prevent the dot product from being numerically too large. Finally, the Transformer uses a set of non-linear transformations with shortcut connections to get the output features:

where MLP is a Multi-Layer Perception applied to each row of AA, and ResBlock is a residual block also applied on each row. The output FoutF^{out} has the same shape as FinF^{in}.

III-C2 Our Interaction Transformer

The Transformer was designed to process an input sequence. In our task, the input is a set of actors and their representations. We represent the state of each actor with features extracted from the BEV feature map as well as the actor’s spatial information, which contains the center location, size, and orientation of the actor. In our task, the queries Q represent the actors we want to forecast their motion of, and keys K and values V represent neighbor contextual information from other actors. Fig. 2 shows a diagram of our Interaction Transformer.

xjix_{j}^{i}, yjiy_{j}^{i}, and θji\theta_{j}^{i} are the locations and orientation of actor jj after it has been transformed into actor ii’s local coordinate system, sgnsgn is the sign function, and (wj,lj)(w_{j},l_{j}) encode the actor’s size. A two-layer MLP transforms the 7-channel input to the 16-channel embedding.

where Concat denotes concatenation along the second dimension, and an MLP is applied on each row vector. We then compute the attention as:

Note that we change the softmax function in Eq. (2) to a sigmoid function. For each prediction target, the softmax function forces the attention scores of all other actors to sum to 1, which is not a reasonable assumption in our context, as this sum is expected to be low when the target actor is not interacting with others, and high when the target actor has strong interactions. Finally, we compute the output features following Eq. (3).

III-C3 Recurrent Temporal Prediction

To capture the sequential nature of the trajectory outputs, we use a recurrent model to predict the motion in an autoregressive fashion (see Fig. 1 right), where the number of recurrent steps equals the number of predicted time steps. At time step tt, the relative location embedding R(t)R^{(t)} is computed from the output waypoints at t−1t-1, and the input actor feature is denoted as Fin(t)F^{in(t)}. We use detection bounding boxes to compute R(0)R^{(0)}, and set Fin(0)F^{in(0)} to be the bi-linearly interpolated output BEV features extracted at the detection box centers. The Interaction Transformer takes Fin(t)F^{in(t)} and R(t)R^{(t)} as input and outputs Fout(t)F^{out(t)}, which is then fed into a two-layer MLP to compute the next set of output waypoints. We then use Fout(t)F^{out(t)} as Fin(t+1)F^{in(t+1)}.

III-C4 Per-Time-Step refinement

For each time step, we use an extra refinement step to obtain more accurate waypoints. The Interaction Transformer first takes Fint−1F_{\text{in}}^{t-1} and relative location embedding Rt−1R^{t-1} to compute Foutt−1F_{\text{out}}^{t-1} and predict a waypoint proposal PproposaltP_{\text{proposal}}^{t}. The updated feature FintF_{\text{in}}^{t} (same as Foutt−1F_{\text{out}}^{t-1}) and RtR^{t} (computed from PproposaltP_{\text{proposal}}^{t}) are then fed into the Interactive Transformer again to obtain the refined waypoints PrefinetP_{\text{refine}}^{t} as our final motion forecasting output. Compared to the proposal step, the refinement step encodes more up-to-date spatial relations between actors. For computational efficiency, PrefineP_{\text{refine}} and PproposalP_{\text{proposal}} are computed in parallel by a single forward pass of the Interaction Transformer. i.e., we obtain both PrefinetP_{\text{refine}}^{t} and Pproposalt+1P_{\text{proposal}}^{t+1} from FouttF_{\text{out}}^{t}, by two different MLPs. Both PrefineP_{\text{refine}} and PproposalP_{\text{proposal}} are parameterized as (dx,dy,sin⁡2θ,cos⁡2θ,clsθ)(dx,dy,\sin 2\theta,\cos 2\theta,cls_{\theta}), where (dxdx, dydy) are relative to the detection box center, and (sin⁡2θ,cos⁡2θ,clsθ)(\sin 2\theta,\cos 2\theta,cls_{\theta}) are defined in the same way as the detection output parameterizations.

III-D Learning

Our model is fully differentiable and can be trained end-to-end by minimizing the weighted sum of detection and prediction losses

For object detection, we use the distance between BEV voxels and their closest ground-truth box center to determine positive and negative samples. Samples whose distance is smaller than a threshold are considered as positive, and negative otherwise. As a large proportion of the samples are negative in dense object detection, we adopt online hard negative mining, where we only keep the most difficult negative samples (with the largest loss) and ignore the easy negative ones. Classification loss is computed over both positive and negative samples while regression loss is computed over positive samples only.

We perform online associations between detection results and ground-truth labels to compute the motion forecast loss. For each detection, we assign the ground-truth box with the maximum (oriented) IoU. If a ground truth box is assigned to multiple detections, only the detection with maximum IoU is kept while other detections are ignored. Regression on future motion is then averaged over those detections with associated ground-truth. Compared to regression on all ground truth actors, our motion forecast loss will not be dominated by actors which are not detected.

IV Experimental Evaluation

We evaluate our approach on two large-scale real-world driving datasets: ATG4D and nuScenes. ATG4D was collected by driving in multiple North American cities with a multi-sensor kit mounted on top of a fleet of vehicles. In total it contains 5,500 video snippets of 25 seconds each with a frame sampling rate of 10 Hz. We use 5,000 snippets for training and 500 snippets for evaluation. Although scenario diversity has already been considered when creating the ATG4D dataset, interactions between traffic actors still happen relatively rarely. To address this problem, we also sample an additional evaluation subset ATG4D-interact from ATG4D-eval which contains interacting actors. Specifically, we focus on blocking interactions by defining actors that: 1) are not parked 2) have a velocity smaller than 0.2 m/s; 3) and have another vehicle nearby (<<5m) in front and in the same lane. We search in the whole dataset for keyframes when there is a vehicle entering or leaving the blocking status. We extend each keyframe temporally by 5 seconds from the past or into the future depending on whether the actor is entering or leaving the blocking status. Using these heuristics, we collect around 800 5-second snippets that are considered “interactive”. nuScenes contains 850 scenes of 20 seconds of driving data, with a frame sampling rate of 20 Hz. Positions of all actors are labeled at 2Hz. We follow the official splits.

IV-2 Evaluation Metrics

IV-3 Implementation Details

We use a voxel resolution of 0.1560.156m for the input BEV representation. We aggregate the past 5 LiDAR sweeps for ATG4D and 10 sweeps for nuScenes to encode a 0.5-second history. We train our model end-to-end using the Adam optimizer . We ignore object labels that are not observable by the LiDAR (i.e., points inside the box). We apply NMS with 0.05 IoU threshold to the detection results to avoid object collision at t=0t=0. We also remove detections with <0.1<0.1 confidence score. The prediction horizon of our model is 3 seconds, with a time-step interval of 0.5 second (therefore T=6T=6). For ATG4D, we train and evaluate the model in the front region of the ego-car within a 100m range. For nuScenes, we follow the official evaluation range, which is 50m around the ego-car. We do not use images or maps in nuScenes because not every frame has well-aligned images and some map data has large alignment error with LiDAR. Due to the limited number of labels in nuScenes, we also conduct augmentation during training. Since labels are only available at 2Hz, we use linear interpolation to estimate actors’ bounding boxes for frames without labels. We also conduct spatial augmentation by randomly scaling (uniformly sampling from [0.95, 1.05]), rotating the ground plane between [−π6,π6]\left[-\frac{\pi}{6},\frac{\pi}{6}\right], and translating in the range [1.0, 1.0, 0.2] m for x, y, and z axes respectively.

IV-4 Comparison with State-of-the-Art

Table I and II show the comparison with the state-of-the-art on object detection, trajectory prediction and interaction modeling. Since none of the compared methods uses any form of collision loss, for fair comparison, we show results of our method both with and without collision loss. We first compare with previous methods that solve the same end-to-end joint perception and prediction task as ours, i.e., FAF and NeuralMP , on the ATG4D dataset. As shown in Table I, our model outperforms both approaches by a large margin in all metrics. Specifically, compared to the previous best joint detection and prediction model NeuralMP, we achieve 4.8%/6.0% AP gain at 0.5/0.7 IoU, and 17.9%/19.3% relative reduction in ADE/FDE at 90% recall. In terms of the interaction specific metric TCR, we show that by introducing the Interaction Transformer, we achieve 80.8% relative reduction at 90% recall. Adding collision loss further reduces TCR by 56.3% on a relative basis. We then compare to SpAGNN , which is the state-of-the-art end-to-end motion forecasting model on nuScenes. SpAGNN takes LiDAR and map as input, and models interaction with graph neural networks, while our model only uses LiDAR as input for nuScenes. For fair comparisons, in Table II, we follow the same metrics as SpAGNN, and demonstrate 34.1%/22.3%/74.3% relative gains on L2@1s, L2@3s, and TCR with IoU 0.1, even without using collision loss.

We also compare with other state-of-the-art interaction modeling approaches, SocialPooling and its convolutional variant on both ATG4D and nuScenes, by replacing our Interaction Transformer with these modules. Note that we exploit our powerful multi-sensor fusion backbone on all these baselines. As shown in Table I and II, on ATG4D, even without collision loss, our method is 3.8%/3.7%/59.4% better compared to ConvSocialPool in ADE/FDE/TCR at 90% detection recall. On nuScenes, at 80% detection recall, our method is 4.2%/5.0%/62.7% better on L2@1s, L2@3s and TCR with IoU 0.1. For detection, in addition to the IoU-AP metrics, we also evaluate with the official nuScenes metrics. Our detection model achieves an mAP of 81.4% on car detection, which is on par with the state-of-the-art.

IV-5 Ablation Studies

To further analyze the contribution of each module proposed in the paper, we conduct several ablation studies on both the ATG4D-interact testset and the full nuScenes dataset. We first provide an ablation study on the high-level model architecture. Towards this goal, our baseline model has the same single-stage structure as previous end-to-end models , with the only difference being the multi-sensor backbone network. We then add our two-stage architecture, the Interaction Transformer, recurrent prediction, and additional refinement on top of the baseline model sequentially. As shown in Table III, using a two-stage architecture with more specialized feature representation and loss computation for each task brings over 1% gain in detection AP. Adding the Interaction Transformer improves the detection performance slightly, but significantly improves all prediction metrics on both datasets. The recurrent architecture further pushes the prediction performance; this gain is largely from updated relative location embedding in each prediction time step. Finally, we add the additional refinement step to recover our full model (without collision loss). This allows the Interaction Transformer to be aware of the spatial relations at the current prediction time step. Compared to the baseline with the same backbone network, our model achieves 6.8%/7.6%/83.1% relative reduction on ADE/FDE/TCR on the ATG4D-interact testset. On nuScenes, we also achieved 4.9%/7.0%/54.6% relative gain on ADE/FDE/TCR.

Next, we analyze the collision loss. The 2nd and 3rd rows of Table III show results of using collision loss without the Interaction Transformer. Collision loss reduces TCR slightly, but at the same time also harms ADE and FDE. In particular, ADE and FDE on ATG4D-interact are 2.0% and 1.3% worse compared to the baseline without collision loss. On the other hand, as shown in the last two rows of Table III, adding collision loss to our recurrent Interaction Transformer achieves much better results. TCR is further reduced by 55.7% and 86.6% on the two datasets, while ADE and FDE are almost unaffected. This indicates that collision loss works better if the network understands spacial relations between actors, as in our Interaction Transformer.

Our last ablation focuses on the design of the Interaction Transformer. For a fair comparison, collision loss is not used here because it harms the baseline model. Our baseline model follows the original Transformer’s architecture, which uses a softmax function to normalize the pairwise attention over all other actors. We first add a relative location feature branch to the attention computation, and then improve it by reasoning spatially in each actor’s own coordinate system at each prediction time step, instead of the global coordinate system. Finally, we replace the softmax function from the original Transformer with the sigmoid function to remove the constraint that attention scores on all actors must sum to 1. As shown in Table IV, each modification provides significant gains on all prediction metrics on both datasets.

IV-6 Qualitative Results

Fig. 3 shows qualitative results of five driving scenarios, comparing the proposed model and the one-stage baseline without the interaction module (baseline in Table III). For fair comparison, neither method uses collision loss. We visualize both detection bounding boxes and estimated future trajectories, and highlight the actors undergoing interactions in red. With our approach, the target vehicle in the first four scenarios successfully senses the existence of the front vehicle and adjusts its future motion accordingly to avoid collisions. In contrast, for the baseline, the future prediction of the target vehicle is similar to a simple extrapolation of its past states. As a result, this leads to a collision with the front vehicle. In the last scenario, we show that the target vehicle reacts to the parked vehicle and makes a sharper turn to avoid collision.

V Conclusion

We have proposed a novel approach to joint detection and future motion forecasting that takes into account the interactions among the actors. We validate our approach on the challenging ATG4D and nuScene datasets and show very significant improvements over the state-of-the-art. In the future, we plan to investigate interactions involving other types of actors such as bicyclists and pedestrians, as well as an extension of our approach to jointly perform motion planning in a similar spirit to NeuralMP .

References