Latency-Aware Collaborative Perception
Zixing Lei, Shunli Ren, Yue Hu, Wenjun Zhang, Siheng Chen
Introduction
Collaborative perception considers a multi-agent system to perceive a scene, where multiple agents collaborate through a communication network . With the observation from multiple agents, collaborative perception can fundamentally overcome the physical limits of single-agent perception, such as over-the-horizon and occlusion. Such collaborative perception models can be widely applied to practical applications, such as autonomous driving and robotics mapping. Previous collaborative perception methods have achieved remarkable success in multiple perception tasks, including 2D/3D object detection, and semantic segmentation . focus on semantic segmentation for drones and discuss the 3D object detection based on the vehicle-to-vehicle-communication-aided autonomous driving. Considering the trade-off between communication bandwidth and perception performance, previous works achieve the collaboration in the intermediate-feature space and leverage attentive mechanisms to fuse the collaboration features.
However, none of these previous collaborative perception methods consider a realistic communication setting where latency is inevitable. As stated in , in a real-time LTE-V2X communication system, the latency time is up to an average of (498 communication periods) + 131.30 ms. Besides, the varying latency times of various communication channels would cause severe time asynchronous issues. Experimentally, latency issue severely damages the collaborative perception system, resulting in even worse performance than single-agent perception. From Fig. 1, we see that: i) the detected vehicles in the purple box in (a) with collaboration is missed in the (b); ii) the correctly detected vehicles in the blue box in (c) are incorrect in (b). The reason is that the received collaborative data with latency represents the situation 1s ago, it misleads the detector to output boxes with significant deviation. This motivates us to consider a collaborative perception system robust to the inevitable communication latency.
To tackle the latency issue, from a machine learning perspective, we propose the first latency-aware collaborative perception system, which actively adapts asynchronous perceptual features from multiple agents to the same time stamp, promoting the robustness and effectiveness of collaboration. As shown in Fig.2, our latency-aware collaborative perception system follows the intermediate collaboration framework and consists of five components: i) encoding module, extracting perceptual features from the raw data; ii) communication module, transmitting the perceptual features across agents under varying communication latency; iii) latency compensation module, synchronizing multiple agents’ features to the same time stamp; iv) fusion module, aggregating all the synchronized features and producing the fusion feature; v) decoding module, adopting the fusion feature to get the final perception output. The main advantage of our latency-aware collaborative perception system is that it’s able to synchronize the collaboration features before aggregation, mitigating the effect caused by latency instead of directly aggregate the received asynchronous features.
The key component of the proposed system is the latency compensation module, aiming to achieve feature-level synchronization. To realize this, we propose a novel SyncNet, which leverages historical collaboration information to simultaneously estimate the current feature and the corresponding collaboration attention, both of which are unknown due to latency. As the attention weight between two agents during collaboration, the proposed collaboration attention has the same spatial resolution with the features and indicates the informative level of each spatial region in the features. It thus provides informative hints for the collaboration partner about how to exploit the collaboration features. Intuitively, the feature and the corresponding collaboration attention are coupling together. Based on this design rationale, the proposed SyncNet leverages a feature-attention symbiotic estimation, which simultaneously infers the collaboration features and the corresponding collaboration attention unknown due to latency, mutually enhancing each other and avoiding the cascading error;
Compared with common time-series prediction methods, the proposed SyncNet has two main differences: i) feature-level estimation, instead of output-level prediction; ii) estimation of coupling features and the associated collaboration attention, instead of predicting a single output.
We extensively evaluate the novel latency-aware collaborative perception system with SyncNet on V2X-Sim dataset on collaborative 3D object detection for autonomous driving. The results verify the robustness of our system and show substantial improvements over state-of-the-art approaches. With SyncNet, our latency-aware collaborative perception system significantly and consistently outperforms single-agent perception under varying communication latency.
To summarize, our contributions are as follows:
We formulate the communication latency challenge in collaborative perception for the first time and propose a novel latency-aware collaborative perception system, which promotes robust multi-agents perception by mitigating the effect of inevitable communication latency.
We propose a novel latency compensation module, termed SyncNet, to achieve feature-level synchronization. It achieves symbiotic estimation of two types of critical collaboration information, including intermediate features and collaboration attention, mutually enhancing each other.
We conduct comprehensive experiments to show that our proposed SyncNet achieves huge performance improvement in latency scenarios compared with the previous method and keeps collaborative perception being superior to single-agent perception under severe latency.
Related Work
V2V communication has two major protocols: IEEE 802.11p protocol and cellular network standards . In IEEE 802.11p protocol, there is a Wireless Access in Vehicular Environment mode to allow users to skip a Basic Service Set, which reduces the overhead in connection setup . In the cellular network, the Long Term Evolution(LTE) standard has derived LTE-V2X . Though the achieved progress in the V2V network, the communication latency issues are still far from perfect and extremely risky for collaborative perception, the latency time is up to an average of (498 communication periods) + 131.30 ms . Instead of avoiding latency from the communication perspective, we aim to mitigate the effect caused by inevitable communication latency from a machine learning perspective, leading to a novel latency-aware collaborative perception system.
2 Collaborative perception
Collaborative perception enables agents to share perceived information through the communication network, fundamentally upgrades perception capabilities over single-agent perception. uses a handshake mechanism to determine which two agents should communicate; introduces a multi-round message passing graph neural networks; proposes a graph-based collaborative perception system with knowledge distillation to balance the communication cost and perception performance. Most previous works focus on the collaboration strategy learning under ideal scenarios. Recently, more realistic scenarios are considered. exploits a pose error regression module to correct errors in the received noisy posture. However, none of the previous works consider the realistic imperfect communication in the collaboration system. To fill this gap, we address the unavoidable communication latency issue, which is extremely risky to the collaboration system, and build a latency-aware collaborative perception system to mitigate the effect caused by latency.
3 Time-series prediction
Time-series prediction targets to predict the future signal according to the historical data. proposes a conv-LSTM architecture in precipitation now-casting. Video prediction, a universal and representative time-series type, has been actively studied . By leveraging prediction techniques, our work recovers the missing information due to latency from historical collaboration information. However, unlike standard prediction, our goal is to maximize the final perception performance, instead of precisely estimating the current state.
Methodology
To tackle the latency issue, we propose a latency-aware collaborative perception system in Section 3.1. As the key of the entire system, the latency compensation module is realized by the proposed SyncNet; see Section 3.2. Finally, Section 3.3 introduces the loss function for training supervision.
Collaborative perception enables multiple agents to perceive a scene together by sharing the perceived data through a communication network. Since communication latency is inevitable in a realistic communication system, here we focus on a latency-aware collaborative perception system; that is, given a non-ideal communication channel with uncontrollable latency, we aim to optimize the perception ability of each agent by mitigating the effect of latency.
where is the estimated feature of the th agent at time stamp after synchronization, is the estimated collaboration attention between the th agent and the th agent at time stamp , is the estimated feature of the th agent at time stamp after aggregating estimated collaboration information, is the neighbors of the th agent and is a hyper parameter.
Step (1a) considers perceptual feature extraction from observation data, where is the encoding network. In Step (1b), we receive perceptual features from other agents with varying latency times. To compensate for latency, Step (1c) estimates the feature and collaborative attention at time stamp by leveraging historical features from the same agent and the real-time feature perceived by ego agent , where denotes the estimation network. Here we assume that each agent can store frames of historical features in memory. Step (1d) fuses all the estimated collaboration information. Finally, Step (1e) outputs the final perceptual output, where is the decoder network. To correspond to Figure 2, Steps (1a) and (1b) contribute to the encoding module; Steps (1c) contributes to the latency compensation module; Step (1d) contributes to the latency fusion module; and Step (1e) contributes to the decoding module.
The proposed latency-aware system has four advantages: i) we explicitly include the communication latency into the design of a collaborative perception system, which has never been done in previous works; see (1b) (1c); ii) we mitigate the effect of latency by estimating missing information from historical collaboration information; see (1c). Instead of synchronizing the perceptual output, we consider feature-level synchronization, because it allows an end-to-end learning framework with more learning flexibility; iii) In (1c), we estimate the coupling collaboration feature and attention simultaneously. If we only estimate the features, we would need to calculate the collaboration attention based on the estimated features. This would amplify the estimation error, causing cascading failures; and iv) we adopt the attention-based estimation, which leverages the collaboration attention in (1c) to promote more precise estimation on more perceptual-sensitive area; see (1d).
2 SyncNet: latency compensation module
Since the latency compensation module is the key of the latency-aware collaborative perception system, we specifically design the estimation networks in (1c), and propose the novel SyncNet. Its functionality is to leverage historical information to achieve latency compensation. SyncNet includes two parts: feature-attention symbiotic estimation, which adopts a dual-branch pyramid LSTM to estimate the real-time features and collaboration attention simultaneously, and time modulation, which uses the latency time to adaptively adjust the final estimation of the collaboration features.
Feature-attention symbiotic estimation. Feature-attention symbiotic estimation (FASE) simultaneously estimates the feature and its corresponding collaboration attention by leveraging a novel dual-branch architecture, including feature estimation branch and attention estimation branch. Both branches of the dual-LSTM network share the same input including real-time features perceived by the ego-agent and frames of the historical features perceived by its collaborator . Each branch is implemented by a pyramid LSTM, which models the series of historical collaborative information and estimates the current state. Pyramid LSTM is specifically designed to capture spatially correlated collaboration features. As shown in Fig. 4, when the vehicle group in the red box relatively moves to the right compared with the central vehicle, the same movements occur in similar areas on the feature. The fact shows that the spatial information is significant for our estimation task. We modify the matrix multiplication in LSTM to a multi-scale convolution architecture; see details in Fig. 5(a). The main differences between the proposed pyramid LSTM and the ordinary ones are that LSTM does not specifically consider extracting spatial features; conv-LSTM extracts spatial features at a single scale; while the proposed pyramid LSTM is designed to capture local-to-global features at multiple scales.
The feature estimation branch aims to obtain the most informative features for collaboration at current time. To achieve this, the feature estimation branch should be attention-aware. And the attention estimation branch aims to find the most informative areas for collaboration at current time, besides, it has to suppress the areas with large estimation errors. To achieve this, the attention estimation branch should be feature-aware. To allow the estimation of feature and the corresponding attention be aware of each other, we recurrently leverage both the estimated feature map and collaboration attention from the previous time stamp to be the input of the following time stamp for either branch.
The entire process is shown in Algorithm 1, where is the latency time, be the historical frames, be the current time, and are the collaboration attention and feature from th agent to th agent at time stamp , respectively, and are the estimation of collaboration attention and feature at time stamp , respectively, is the input of pyramid LSTM at time stamp , , , and are the the hidden states and cell states of the pyramid LSTM in each branch, respectively.
The proposed feature-attention symbiotic estimation has three characteristics: i) the dual-branch structure simultaneously infers the collaboration features and the corresponding collaboration attention, keeping independent and eliminating the cascaded failure; ii) the estimation networks take the collaboration attention as input so that to focus on more informative areas, promoting more effective estimation; iii) the learnable attention estimation network obtains the information of the feature and gets supervision from attention and fusion feature under ideal collaboration. During the end-to-end optimization, it can not only imitate the weight distribution calculated without latency, but also actively learn to reduce the attention of the spatial position with large noise in features.
Time modulation. Although FASE achieves the basic functionality of , we find, when latency is low, the performance degradation caused by latency is relatively minor than the estimation noise led by FASE. To handle this, we propose time modulation, it attentively fuses the raw (working well at low latency) and estimated (working well at high latency) features conditioned on the latency time, generating more comprehensive and reliable estimation.
3 Loss function
Let be the ground truth of final perception output of the th agent at time stamp , be the ground truth feature of the th agent at time stamp after aggregating real-time collaboration information, be the ground truth feature map of the th agent at time stamp , and be the ground truth collaboration attention from the th agent to the th agent at time stamp . We consider minimizing the following objective to optimize the overall latency-aware collaborative perception system:
Experiments
We validate our SyncNet on LIDAR-based 3D object detection task with a multi-agent dataset, V2X-Sim. V2X-Sim is built with the co-simulation of SUMO and CARLA. V2X-Sim includes 80 scenes in the training set and 11 scenes in the test set. Each sample contains 2.67 agents on average and includes 3D point clouds input and 3D bounding box annotations. The 3D point clouds are generated by a LIDAR with 32 channels and 70m max range, 20Hz rotation frequency, and 5Hz recorded frequency. To simulate collaborative perception under a latency scenario, we load data in asynchronous time stamps, and the latency time is randomly generated from an exponential distribution.
2 Implementation details
Experimental setting. We crop the point clouds which locate in the region of defined in the ego-vehicle Cartesian coordinate system. We set the size of each voxel as . After crop and voxelization, we get a Bird’s-Eyes view map with dimension . The encoded features to be transmitted has a dimension of . The latency between two agents can be a fixed or random number generated by exponential distribution rounded to an integer. We train our model using NVIDIA RTX 3090 GPU with Pytorch. The evaluation metric is the Average Precision(AP) metric at Intersection-over-Union threshold of 0.5 and 0.7.
Baselines. Our proposed latency-aware collaborative perception system adopts one of the state-of-the-art collaborative perception frameworks, DiscoNet , and leverages the proposed SyncNet as the latency compensation module to handle various latency settings. To validate our latency-aware collaborative perception system, DiscoNet+SyncNet, we compare with three baselines: i) single-agent perception system, No Collaboration; ii) latency-unaware collaborative perception, DiscoNet ; iii) naive latency-aware late-fusion-based collaborative perception by using Kalman filter, Late collaboration+Kalman Filter. Note that SyncNet can also work as a plugin latency compensation module for other intermediate collaborative perception methods, such as V2VNet . SyncNet is equivalent to feature-attention symbiotic estimation (FASE) + time modulation(TM). Corresponding to FASE with the dual-branch structure, a simplified variation is Vanilla Estimation(VE), which adopts a single-branch LSTM to estimate collaborative features only. In ablation study, we will compare the performances of DiscoNet, DiscoNet+FASE, DiscoNet+VE and DiscoNet+SyncNet.
Training strategy. We use a curriculum learning strategy in the training stage. Curriculum learning starts with easy samples and then gradually increases the difficulty. To handle the flexible latency time, we train the model under various latency settings. However, the training loss sharply increasing with the latency time causes an unstable and vulnerable training process. To tackle this issue, we employ the curriculum learning technology and gradually increase the latency time by 1 every 10 epochs until 10. Afterward, we randomly sample the latency time with an exponential distribution averaging 5 to further upgrade the model to accommodate the flexible communication latency.
3 Quantitative evaluation
Fig. 6 compares the detection performances among our latency-aware collaborative perception system, no collaboration, DiscoNet without latency compensation, and late collaboration with a Kalman filter as a function of latency time. We see that: i) DiscoNet is vulnerable to latency, whose performance is even lower than No Collaboration at high latency; ii) our DiscoNet+SyncNet is robust to latency and outperforms No Collaboration even in a terrible communication condition with a communication latency of up to 10 frames; iii) our DiscoNet+SyncNet consistently outperforms DiscoNet under varying communication latency, and improves performance in AP@0.5/0.7 by up to 15.6%/12.6%.
Fig. 7 shows the performances of other frameworks, including V2VNet and a transformer-based fusion module, with and without SyncNet. The transformer-based fusion module deploys a multi-head attention architecture to fuse the collaborative features at each spatial position. The SyncNet module improves the performance up to 11.8%/8.7% in AP@0.5, respectively. It shows that various collaborative perception models are vulnerable to latency and the proposed compensation module consistently and significantly benefits those frameworks.
4 Ablation study
We first study the effect of historical frames in Fig. 8. We see that, significantly outperforms , but only brings marginal benefit. The default choice in the paper is , achieving a reasonable balance between computation efficiency and performance. We further validate the effectiveness of the two major components of our proposed latency compensation module (SyncNet): FASE and TM. Vanilla Estimation(VE) adopts a single-branch structure only estimating collaborative features. Fig. 9 compares DiscoNet, DiscoNet + FASE, DiscoNet + VE and DiscoNet + SyncNet as a function of latency time. We see that: i) comparing green line with blue line, our latency-aware collaborative perception system only needs a vanilla LSTM compensation module to achieve significant performance improvement in latency scenarios. ii) comparing red line with blue line, FASE architecture can improve performance in AP@0.7 metric; iii) comparing red line with yellow line, TM can improve performance when latency is low. Table 1 further discusses the effectiveness of the compensation model, multi-scale convolution and time modulation module at low() and high latency(). We see that: i) D surpasses A, E surpasses B, F surpasses C, reflecting FASE is consistently effective in AP@0.7 metric; ii) C surpasses B, F surpasses E, reflecting TM is consistently effective when latency is high.
5 Qualitative evaluation
Fig.10 shows the detection results of DiscoNet without latency, DiscoNet with latency, DiscoNet+VE and DiscoNet + SyncNet. Comparing (a) with (b), we see that the correctly detected vehicles in the purple box in (a) are missed or incorrectly detected in (b) due to the latency. (c) shows that the vanilla estimation (without FASE) partially compensates latency error in the blue box but fails to achieve accurate estimation in the orange box, while our SyncNet could precisely recover the true position of both vehicles, shown in purple box of (d). Plot (d) shows that SyncNet achieves the best compensation and precisely recovers the true position of vehicles.
Fig. 11 shows the attention weight of the collaboration feature from the neighbor agent in the example shown in the first row of Fig. 10. We can see that: (b), (c) both have a similar large weights in the red box, which introduce noise into the collaboration, and (d) has a small weight like (a), here to capture the truly informative area and avoid the cascading errors caused by the inaccurate feature estimation because the attention estimation branch in SyncNet under the supervision of the ground truth of collaboration attention. These qualitative results suggests the effectiveness of SyncNet.
Conclusions
We introduce latency-aware collaborative perception and propose a novel latency compensation module, SyncNet, for time-domain synchronization, which fits in existing intermediate collaboration methods. SyncNet adopts a novel symbiotic estimation architecture, which jointly estimates intermediate features and attention weights, as well as the time modulation, which significantly improves the overall performance at low-latency range. Comprehensive quantitative and qualitative experiments show that the proposed SyncNet can improve the perception performance in the communication latency scenario and effectively address the latency issue in the collaborative perception.
Acknowledgenments
This research is partially supported by the National Key R&D Program of China under Grant 2021ZD0112801, National Natural Science Foundation of China under Grant 62171276, the Science and Technology Commission of Shanghai Municipal under Grant 21511100900 and CALT Grant 2021-01.