Learning to Communicate and Correct Pose Errors

Nicholas Vadivelu, Mengye Ren, James Tu, Jingkang Wang, Raquel Urtasun

Introduction

Despite the powerful capabilities of deep neural networks in fitting raw, high dimensional data, they are limited by the computational power and sensory input available to a single agent. Thus, combining the sensory information and computational power of multiple agents to cooperatively accomplish a goal can greatly amplify the effectiveness of these systems . For example, V2VNet has recently shown that by allowing multiple self-driving vehicles (SDVs) to communicate through a set of learned spatially-aware feature maps, we can obtain significant gains in detecting obstacles that would have otherwise been occluded or far away from a single-agent perspective.

The success of V2VNet depends on the precise localization of each participating vehicle, which is used to warp the feature maps so they can be spatially aligned. Localization noise, however, is common in the real world. While V2VNet exhibits some implicit tolerance, the performance degrades below single-agent performance under realistic amounts of noise. Due to the safety critical nature of self-driving, it is paramount to study the robustness against pose noise in a vehicle-to-vehicle communication system and to design models that can explicitly reason under such noise.

In this paper, we propose end-to-end learnable neural reasoning layers that learn to communicate, to estimate pose errors, and finally, to reach a consensus about those errors. First, the pose regression module predicts the relative pose noise between a pair of vehicles. Second, to ensure globally consistent poses, we propose a consistency module based on a Markov random field with Bayesian reweighting. Lastly, in the communicated messages aggregation step, we propose using predicted attention weights to attenuate the outlier messages among vehicles.

Our evaluation under the same setting as the original V2VNet shows that our model can maintain the same level of performance under strong translation and heading localization noise, while V2VNet eventually suffers from such input noise, even if the network is trained with data augmentation. Our framework also outperforms other competitive pose synchronization methods.

Related Work

In this section, we describe the literature in the area of collaborative self-driving. We also give an overview on related problem formulations such as transformation synchronization, visual odometry, and point-cloud registration.

Existing literature studies how to leverage multiple self-driving vehicles (SDVs) to perform vehicle-to-vehicle communication (V2V) to enhance perception, prediction, and motion planning. The benefits of multiple agents can be exploited by aggregating raw sensor data , communicating intermediate feature maps , or combining the outputs of multiple vehicles . show limited robustness to localization error, with no explicit steps to address it. We follow the setting of V2VNet by communicating intermediate feature maps since it achieves better performance and more efficient communication.

Transformation synchronization is the process of extracting absolute poses given relative poses. Methods include spectral solutions , semidefinite relaxations , probabilistic approaches , sparse matrix decomposition , and/or learned approaches . While these methods could be used to refine our pairwise estimates, they are only shown to be robust when there are many more views to synchronize than in our setting (e.g., 30 views per scene in vs. up to 7 in our setting). Hence, they are susceptible to outliers which strongly influence the final synchronized poses. Our approach can certainly be used in standard transformation synchronization problems, but more importantly, we propose an end-to-end system for robust multi-agent perception and motion forecasting.

Visual odometry is the process of determining the pose of an agent given images from the agent’s view. In our setting, when correcting pose error, we extract the relative poses given pairs of views. Yousif et al. provide a survey on several visual odometry methods, including feature-based and stereo-based . More recently, approaches based on RCNN learn this task end-to-end. These approaches are optimized for images, LiDAR, or other raw sensory inputs, whereas in our setting, we aim to align intermediate feature maps.

Point cloud registration is the task of finding a (typically rigid) transformation to align two point clouds. propose robust methods for registration, while propose deep learning based approaches. provides a full review of traditional point cloud registration methods. These methods are not suited in our setting due to the high communication overhead required to transmit LiDAR point clouds to neighboring self-driving vehicles.

Outside of self-driving, there is broad literature on multi-agent deep learning systems. communicate actions and state to other agents, while use a controller network for communication. Our setting is more similar to the former, where each vehicle communicates an intermediate representation of its view to nearby vehicles. uses a learned graph neural network for communication and cooperative routing. However, many of these methods are typically studied in toy settings, whereas we evaluate our model on a realistic self-driving dataset.

Learning to Communicate and Correct Pose Errors

Pose noise has been shown to severely detriment existing collaborative multi-agent self-driving systems. In this section, we describe our novel approach to correct pose errors in such settings. In the following, we first review V2VNet , the collaborative self-driving framework that we base our models on. We then propose a pose error correction network composed of i) a pose regression module to predict pairwise relative poses, ii) a consistency module to reach global consensus, and iii) an attention aggregation module to filter out outlier messages. These modules are learned end-to-end jointly to improve object detection and motion forecasting.

Our pose correction approach is based on V2VNet , a state-of-the-art collaborative multi-vehicle self-driving network which has been shown to provide significant improvements in both object detection and motion forecasting over single vehicle systems. We call the combined detection and forecasting task perception and prediction (PnP). We first review the background of V2VNet—an overview diagram is illustrated in Figure 1.

Given multiple LiDAR sweeps, V2VNet voxelizes the point cloud into 15 cm3 voxels, and concatenates them along the height dimension to form a birds-eye view input representation. It then processes this representation using a 2D CNN, denoted FF, to produce a spatial feature map of shape c×l×wc\times l\times w (channels, length, width). To facilitate cooperation, each self-driving vehicle (SDV) compresses and broadcasts these spatial feature maps to nearby SDVs. We thus call these spatial feature maps messages and denote the message from vehicle i as mi\mathbf{m}_{i}.

Vehicle ii collects all incoming messages and aggregates them via a graph neural network (GNN) GG . The set of vehicles which communicate with vehicle i is denoted adj(i)adj(i). When vehicle ii receives message mj\mathbf{m}_{j} from vehicle j∈adj(i)j\in adj(i), it warps mj\mathbf{m}_{j} from the perspective of vehicle jj to its own. Vehicle i uses its own pose ξi\boldsymbol{\xi}_{i} and the other vehicle’s pose ξj\boldsymbol{\xi}_{j} to compute the relative pose ξji\boldsymbol{\xi}_{ji}. The message from vehicle j (mj\mathbf{m}_{j}) is transformed via ξji\boldsymbol{\xi}_{ji} to produce the warped message mji\mathbf{m}_{ji}, which is aligned to the perspective of vehicle i. Let the aggregated message for agent ii be hi:=G({mji}j∈adj(i))\mathbf{h}_{i}:=G\left(\{\mathbf{m}_{ji}\}_{j\in adj(i)}\right) We refer to for details on the aggregation algorithm.

Finally, vehicle ii uses a CNN HH to process aggregated messages to predict the final outputs which consist of object detections represented with their 3D position, width, height, and orientation, as well as prediction outputs representing the locations of objects at future time steps.

2 Robust V2V communication against pose noise

V2VNet has been shown to be vulnerable to pose noise because misaligned incoming messages will result in unusable features for the network. Under realistic noise, V2VNet’s performance can be worse than single vehicle PnP. In this section we introduce details of our approach to improve robustness against pose noise. An illustration is shown in Figure 2.

We now refine the relative pose estimates from the regression module by finding a set of globally consistent absolute poses among all our SDVs. By allowing the SDVs to reach a global consensus about eachothers absolute pose, we can further mitigate pose error.

To perform inference on our MRF, we would like to estimate the values of our absolute poses ξi\boldsymbol{\xi}_{i}, the scale parameters Σi\Sigma_{i}, and the weights wjiw_{ji} that maximize the product of all our pairwise potentials. We achieve this via Iterated Conditional Modes , described in Algorithm 1. The maximization step on Line 4 happens simultaneously for all nodes via weighted expectation-maximization (EM) for the tt distribution . We provide the EM algorithm in the Supplementary Material. The maximization step on Line 5 can be computed using the following closed form :

We then use these estimated poses to update the relative transformations needed to warp the messages.

The aggregated message is then used by the network HH to predict bounding boxes for object detection and waypoints at future timesteps for motion forecasting.

3 Learning

This function produces smooth labels to temper the attention module’s predictions so the attention weights are not just 0 or 1. We define the loss for our joint training task as follows:

where LCE\mathcal{L}_{CE} is binary cross entropy loss. This additional supervision was paramount to training the attention mechanism—training with LPnPL_{PnP} alone produced a significantly less effective model.

Then, we freeze V2VNet and the attention and train only the regression module using Lc\mathcal{L}_{c}. In this stage, all the SDVs get noise from the strong noise distribution Ds\mathcal{D}_{s}. We train this network using a loss which is a sum of losses over each coordinate:

Finally, we fine-tune the entire network end-to-end with the combined loss: L=Lc+Ltask\mathcal{L}=\mathcal{L}_{c}+\mathcal{L}_{task}, which is possible because our MRF inference algorithm is differentiable via backpropogation.

Experiments

We evaluate our method on detection, prediction, and pose correction in various noise settings, including noise not seen during training. The specific architectures and hyperparameters are provided in the supplementary material.

Our model is trained on the V2V-Sim dataset , which is generated from a high-fidelity LiDAR simulator . The simulator uses real-world snippets to first reconstruct 3D scenes with static and dynamic objects, then simulates LiDAR point clouds from the viewpoints of multiple self-driving vehicles. Each scene contains up to 7 SDVs. There are 46,796/4,404 frames for the train/test split, where each frame contains 5 LiDAR sweeps. We refer readers to for more details.

Throughout training and evaluation, the noise is sampled and applied independently to the pose of each SDV. This can be applied as a post-processing step on the data, or can be simulated directly within LiDARSim . During training, the positional noise is drawn from a Gaussian with μ=0\mu=0, σ=0.4\sigma=0.4 for Ds\mathcal{D}_{s} and σ=0.01\sigma=0.01 for Dw\mathcal{D}_{w}; the rotational noise is drawn from a von Mises distribution with μ=0\mu=0, σ=4∘\sigma=4^{\circ} for Ds\mathcal{D}_{s} and σ=0.1∘\sigma=0.1^{\circ} for Dw\mathcal{D}_{w}. During evaluation, the parameters of these distributions are varied as described for each experiment. We show experiments on both noise similar or greater than the noise levels seen during training. Self-driving cars utilize geometric registration algorithms that localize the vehicle online with respect to a 3D HD map. These methods are very precise, with 99% of the errors being much smaller than 0.2m, which informed the evaluation ranges chosen.

We compare our method to a competitive transformation synchronization method Learn2Sync , which considers the pairs of depth maps to iteratively reweight pairwise registrations when finding globally consistent poses. To process pairs of messages instead of depth maps, a larger version of the Learn2Sync architecture is used (see supplementary material). During evaluation, Learn2Sync is used in place of our consistency module. Our pretrained pose correction module produces the initial pairwise registrations for Learn2Sync.

For a simple baseline to our method, V2VNet is trained with noisy poses as a form of input data augmentation, which asks the network to implicitly handle pose noise instead of explicitly correcting the noise. We refer to this network as Data Aug.

2 Experimental results

As shown in Figure 3, V2VNet is quite vulnerable to pose noise, especially heading noise. When trained with data augmentation, the model becomes significantly more robust, however, this is at the cost of worse performance in less noisy conditions. The original model trusts incoming messages too much, whereas the data augmented model trusts them too little and discards too much information. Using the correction provides significant benefits: there is little drop in performance when faced with the noise seen in the training set (0.4 m, 4.0∘4.0^{\circ} std.). The model generalizes well to noise stronger than seen in the training set. Our consistency method shows considerable improvement over Learn2Sync, which is expected in this case as synchronization algorithms are commonly designed and evaluated on far larger graphs. Having so few transformations to synchronize renders these methods vulnerable to outliers.

Table 1 shows that our consistency module further enhances the correction performance. Note that while the RMSE decreases significantly with other methods, the MAE only decreases marginally. This implies that, while the outliers are corrected, the average correction is not improved significantly. Also, this means outliers “poison” the good predictions, resulting in relative pose estimates that are mediocre. Improving the average case is more important than dealing with outliers as our model with attention can ignore outliers and focus on well-aligned messages.

Table 2 shows that all the components provide significant benefits. Interestingly, using the attention module provides improvement over V2VNet even when no noise is present.

There will always be a domain gap between the noise seen during training and the noise an agent may experience in the real world. In our setting, the pose regression is trained on unbiased Gaussian noise, however, in the real world, a vehicle may experience systematic, biased error. Figure 4 evaluates the generalization ability of our method on noise that is biased and stronger than what the model may face in reality. The performance of the model stays well above single vehicle PnP. Furthermore, outliers become more prevalent in this setting, which affects the performance of consistency methods not designed to deal with outliers in small graphs.

Strong performance independent of the number of nearby SDVs is important for safe operation of an SDV. Figure 5 shows that V2VNet’s performance drops as soon as we introduce another SDV due to the pose noise affecting messages, even after Data Augmentation. This is not the case with our correction: increasing the number of SDVs improves performance, almost matching the original model evaluated with no noise. The consistency also maintains reliable performance even with few nearby SDVs.

Conclusion

Collaborative self-driving cars will bring the safety of self-driving to the next level. In this paper, we propose a collaborative self-driving framework that is made robust to pose errors in vehicle-to-vehicle communication. Unlike traditional pose synchronization methods, our model is end-to-end learned to improve detection and motion forecasting. We demonstrate the effectiveness of our method under various levels of pose noise using V2V simulation. In the future, we can extend our work to exploit the temporal consistency of the pose error in incoming messages to improve performance and efficiently reuse computation. We also aim to expand our neural reasoning framework to correct more general types of communication noises to make collaborative self-driving more robust.

We would like to thank Andrei Ba^\hat{\text{a}}rsan and Pranav Subramani for insightful discussions. We would also like to thank all the reviewers for their helpful comments.

References

Appendix A EM for weighted t𝑡t-distribution

Recall in Algorithm 1 on line 4 from the main manuscript we maximize the following quantity for each ii:

This is equivalent to finding the weighted maximum likelihood estimate (MLE) of ξi,Σi\boldsymbol{\xi}_{i},\Sigma_{i} given observations {ξ^ji∘ξj}j∈adj(i)∪{ξ^ij−1∘ξj}j∈adj(i)\{\widehat{\boldsymbol{\xi}}_{ji}\circ\boldsymbol{\xi}_{j}\}_{j\in adj(i)}\cup\{\widehat{\boldsymbol{\xi}}_{ij}^{-1}\circ\boldsymbol{\xi}_{j}\}_{j\in adj(i)}. Recall that ξi,Σi\boldsymbol{\xi}_{i},\Sigma_{i} are the location and scale of the tt distribution with ν\nu degrees of freedom. We modify the EM algorithm given in to compute the weighted MLE.

The student tt distribution can be defined as follows:

where 1 is the mean of the Gamma, 2/ν2/\nu is the shape parameter k, and N\mathcal{N} denotes the multivariate normal distribution. We provide the full expressions for the tt and Gamma distributions in section E. For the expectation step, we compute the expectation of our latent parameter ηji\eta_{ji}. For the maximization step, we compute ξi,Σi\boldsymbol{\xi}_{i},\Sigma_{i} given ηji\eta_{ji}. We use δji\boldsymbol{\delta}_{ji} to denote the difference between observation jiji and the current estimate of ξi\boldsymbol{\xi}_{i} for convenience. The full algorithm is described in Algorithm 2.

When there are only two vehicles communicating, we a simple average instead of EM to estimate ξi\boldsymbol{\xi}_{i}. Notice on line 19 we do not use the weights wjiw_{ji}, as the small size of our graph often leads to a singular Σi\Sigma_{i} when using these weights. 15 iterations is sufficient for convergence and 2 degrees of freedom worked well.

Appendix B Additional Experiments

We analyze the effects of positional and heading noise seperately in Figure 6. Heading noise is far more detrimental than positional noise, as objects far from the vehicle can be displaced significantly even with slight heading error.

Appendix C Qualitative Examples

Figure 7 shows PnP outputs from five scenes in the validation set when the agents are subject to pose noise. As shown, the misaligned messages causes many detections to be innacurate, particularly detections farther away from the ego vehicle. We also see that forecasting predictions are skewed without the correction module.

Appendix D Implementation Details

In this section, we provide the implementation details for the training procedure and architectures used.

V2VNet and the attention network are trained using the Adam optimizer with a one-cycle learning rate for 6 epochs starting from the pre-trained LiDAR backbone with a peak one-cycle learning rate of 0.0004. Then, V2VNet and the attention network are frozen and only the regression module is trained for 12 epochs with a peak one-cycle learning rate of 0.002. For the loss, we use λpos=2/3\lambda_{pos}=2/3 and λrot=1/3\lambda_{rot}=1/3. Finally, the entire network is fine tuned with the combined loss L\mathcal{L} for 3 epochs with a peak learning rate of 0.0001. For the consistency module, using a tt-distribution with 2 degrees of freedom, k=120k=120 for the prior worked well, 15 iterations of EM for the tt-distribution, 15 steps of ICM, and 10 reweighting steps worked well. The attention module is trained with γ=0.9\gamma=0.9, p=0.5p=0.5, λPnP=0.9\lambda_{PnP}=0.9, and λattn=0.1\lambda_{attn}=0.1 without significant tuning. We make slight modifications to V2VNet detailed in the supplementary materials. These modifications resulted in virtually no change in PnP performance.

D.2 Changes to V2VNet

Due to GPU memory limitations, we use a slightly altered V2VNet with near identical performance to the architecture from . V2VNet originally performed 3 rounds of message passing between vehicles per inference; we reduce this to 2. Our correction system only operates during the first round of propagation. The second round uses the corrected localization and attention weights from the first round. When receiving messages, V2VNet uses a convolutional neural network to process each incoming message before aggregating and passing to the ConvGRU in the GNN. We remove this processing step and aggregate the messages directly before passing them to the ConvGRU. Finally, V2VNet uniformly samples between 1 and 7 SDVs per training example. We sample exactly 4 SDVs per training example when training V2VNet and the attention, for more consistent GPU memory utilization. We sample up to 7 SDVs per scene when training only the regression module (as some training examples have fewer than 7 vehicles).

D.3 Architecture for our Method

The dimensions of a message are (c,l,w)=(80,128,320)(c,l,w)=(80,128,320). Therefore, the dimensions of the input to the regression and attention modules are (160,128,320)(160,128,320). Architectures are described in terms of PyTorch modules. All convolutional layers have a padding and stride of (1,1)(1,1) unless otherwise specified. We annotate each layer with the output activation shape.

We describe our attention architecture below.

The use of AdaptiveMaxPool2d is important: it allows our computed attention weights to be invariant to the amount of spatial overlap between two messages.

We describe the architecture of our regression module below.

D.4 Architecture for Learn2Sync

We train Learn2Sync for 10 epochs using the Adam optimizer and a one-cycle learning with a maximum learning rate of 0.01. We searched for the optimal learning rate from the set {0.1,0.01,0.001,0.0001}\{0.1,0.01,0.001,0.0001\}. Learn2Sync originally used a modified AlexNet architecture . We simply increased the size as detailed below. The rest of the hyperparameters were kept from .

Appendix E Distributions