Fooling LiDAR Perception via Adversarial Trajectory Perturbation

Yiming Li, Congcong Wen, Felix Juefei-Xu, Chen Feng

Introduction

Autonomous driving systems are generally equipped with all kinds of sensors to perceive the complex environment . Among the sensors, LiDAR has played a crucial role due to its plentiful geometric information sampled by incessantly spinning a set of laser emitters and receivers. However, LiDAR scans are easily distorted by vehicle’s motion, \ie, the points in a full sweep are sampled at different timestamps when vehicle is at different locations and orientations, as shown in Fig. 1. Imagine that a self-driving car is on a highway at a speed of 30m/s, its LiDAR with a 20Hz scanning frequency would move 1.5 meters during a full sweep, severely distorting the captured point cloud.

Such distortions are typically compensated by querying the vehicle/LiDAR’s pose at any time from a continuous vehicle ego-motion tracking module that fuses pose estimation from Global Navigation Satellite System (GNSS, \eg, GPS, GLONASS, and BeiDou), Inertial Navigation System (INS), and SLAM-based localization using LiDAR or cameras. Well-known LiDAR datasets like KITTI and nuScenes have already corrected such motion distortions prior to the dataset release. Researchers then made impressive progresses by processing those distortion-free point clouds using deep neural networks (DNNs) for many tasks, \eg, 3D object detection, semantic/instance segmentation, motion prediction, multiple object tracking.

However, using DNNs on LiDAR point clouds creates a potentially dangerous and less cognizant vulnerability in self-driving systems. First, the above perception tasks are functions of LiDAR point clouds implemented via DNNs. Second, the motion compensation makes LiDAR point clouds a function of the vehicle trajectories. This functional composition leads to a simple but surprising fact that those perception tasks are now also functions of the trajectories. Thus, such a connection exposes the well-known adversarial robustness issue of DNNs to hackers who could now fool a self-driving car’s safety-critical LiDAR-depending perception modules by calculatedly spoofing the area’s wireless GNSS signals, which is still a serious and unresolved security problem demonstrated on many practical systems . Luckily, given that the aforementioned non-GNSS pose sources such as INS and visual SLAM are fused together for vehicle ego-motion estimation, large variations (meter-level) of GNSS trajectory spoofing could be detected and filtered, ensuring a safe localization and mapping for self-driving cars. However, what if a spoofed trajectory only has dozens of centimeters offset? Would current point cloud DNNs be robust enough under such small variations?

In this paper, we initiate the first work to reveal and investigate such a vulnerability. Different from existing works directly attacking point cloud coordinates with 3D point perturbations or adversarial object generation , we propose to fool LiDAR-based perception modules by attacking the vehicle trajectory, which could be detrimental and easily deployable in the physical world. Our investigation includes how to obtain LiDAR sweeps with simulated motion distortion from real-world datasets, convert LiDAR point clouds as a differentiable function of the vehicle trajectories, and eventually calculate the adversarial trajectory perturbation and make them less perceptible. Our principal contributions are as follows:

We propose an effective approach for simulating motion distortion using a sequence of real-world LiDAR sweeps from existing dataset.

We propose a novel view of LiDAR point clouds as a differentiable function of the vehicle trajectories, based on the real-world motion compensation process.

We propose to Fool LiDAR perception with Adversarial Trajectory (FLAT), which has better feasibility and transferability (code will be released).

We conduct extensive experiments on 3D object detection as a downstream task example, and show that the advanced detectors can be effectively blinded.

Related Work

GNSS/INS and LiDAR Motion Compensation. Motion distortion is also known as motion blur or rolling shutter effect of LiDAR on ego-motion vehicles . To compensate for such distortions, GNSS/INS is often used to provide the pose of the LiDAR at any moment when a point is scanned. This opens backdoors for a self-driving system. First, space weather is a substantial error source for GNSS and also significantly affects systems such as differential GPS. The influence of ionosphere disturbances on GPS kinematic precise point positioning (PPP) can be larger than 2 to 10 meters at different latitudes , while solar radio burst could cause GPS positioning errors as large as 300 meters vertically and 50 meters horizontally . Besides those naturally occurring events, malicious attacks such as GPS jamming or spoofing could also be used to arbitrarily modify the GPS trajectory . Moreover, when GNSS and INS are fused, such attacks can affect not only positions but also the rotational component (gyroscope bias compensation) of a vehicle’s trajectory . Of course, nowadays for motion compensation, LiDAR poses are usually fused between GNSS/INS and LiDAR-/camera-based localization, so large pose variations from extreme space weather or “urban canyon” effect could be filtered. But as long as GNSS is a part of the equation, the backdoors could remain open, especially when the spoofed trajectories only have small variations from the ground truth, as we show later.

Image-based Adversarial Attack. Despite the great success achieved by deep learning in both academia and industry, researchers have found that deep networks are susceptible to carefully-designed adversarial perturbation which is hard to distinguish. Since such vulnerability is firstly pointed out in image classification task , broad attentions have been paid to adversarial robustness in various downstream tasks, \eg, semantic segmentation , object detection , visual tracking , \etc. Adversarial attack is divided into white box and black box based on whether the model parameters are known. Besides, attacks can be categorized as targeted and untargeted according to whether the adversary has a particular goal. Meanwhile, many attempts have been made in defense mechanisms, such as adversarial training , certified defense , adversarial example detector , and ensemble diversity . Extensive studies in image-based adversarial attack and defense have largely promoted the development of trustworthy machine learning in 2D computer vision and inspired similar investigations in 3D vision.

Point Cloud Attack and Defense. Recently, researchers have explored the vulnerability of DNNs taking point clouds of 3D objects as input. For the object-level point cloud attack, Xiang et al. proposed point perturbation as well as cluster generation to attack the widely-used PointNet . Besides, critical points removal , adversarial deformation , geometric-level attack are proposed for fooling the point cloud-based deep model. However, none of them directly target LiDAR point clouds of self-driving scenes which have domain gaps than object-level point clouds. Moreover, to implement the above point cloud attacks towards a self-driving car requires tampering with its software for altering point cloud coordinates. Differently, our paper reveals a simple yet dangerous possibility of spoofing the trajectory to attack deep modules through the LiDAR motion correction process yet without any need of software hacking. As for the works on scene-level point cloud attack, Tu et al. proposed to generate 3D adversarial shapes placed on the rooftop of a target vehicle, making the target invisible for the detectors. Some works created spoofing obstacles/faked points in front of the car to influence the vehicle’s decision, but they can only modify the limited area of the scene. Differently, our method does not need to physically alter any shapes in the scene, and can easily scale up the attack to the whole scene. In a word, existing works directly manipulate on the point coordinates either physically or virtually, while our attack is realized by spoofing the vehicle trajectories.

Affected Downstream Tasks. Theoretically, every downstream task that require LiDAR point cloud as input could be affected by the attacks proposed in this paper. This could include geometric vision tasks such as registration, pose estimation, and mapping, as well as pattern recognition tasks such as 3D object detection , semantic segmentation , motion prediction , and multiple object tracking . While the first group of tasks could be less severely affected by GNSS spoofing via data fusion as mentioned above, the second group has higher vulnerability, because small but calculated perturbations in the point coordinates could affect deep networks as demonstrated in the above related works. In this paper, without loss of generality, we choose to focus on the 3D object detection task to illustrate the severity of this backdoor, because miss detection of safety-critical objects surrounding a self-driving car could be a matter of life or death.

Motion Distortion in LiDAR

LiDAR measurements are obtained along with the rotation of its beams, so the measurements in a full sweep are captured at different timestamps, introducing motion distortion which jeopardizes the vehicle perception. Autonomous systems generally utilize LiDAR’s location and orientation obtained from the localization system to correct distortion . Most LiDAR-based datasets have finished synchronization before release. Hence, the performance of current 3D perception algorithms in the distorted point cloud remains unexplored. We briefly introduce the nomenclature in this work before detailed illustrations.

World Frame. We use a coordinate frame WW fixed in the world with the orthonormal basis {xW,yW,zW}\{\mathbf{x}_{W},\mathbf{y}_{W},\mathbf{z}_{W}\} to describe the global displacement of the self-driving car.

Object Frame. The car can be associated with a right-handed, orthonormal coordinate frame which can describe its rigid body motion. Such a frame attached to the car is called object frame . In the following, we use frame to denote the object frame of the car at different timestamps.

Sweep&Packet. LiDAR points in a complete 360∘ is called a sweep, and the points are emitted as a stream of packets, each covering a sector of the 360∘ coverage .

where NN is the total interpolation steps. For the orientation, we implement spherical linear interpolation (slerp) . Using qA=[qwA,qxA,qyA,qzA]⊤\mathbf{q}_{A}=\left[q^{A}_{w},q^{A}_{x},q^{A}_{y},q^{A}_{z}\right]^{\top} and qB=[qwB,qxB,qyB,qzB]⊤\mathbf{q}_{B}=\left[q^{B}_{w},q^{B}_{x},q^{B}_{y},q^{B}_{z}\right]^{\top} to represent quaternion at frame AA and BB, then the quaternion at the nn-th interpolated frame is:

where θ=cos⁡−1(qA⋅qB)\theta=\cos^{-1}\left(\mathbf{q}_{A}\cdot\mathbf{q}_{B}\right) is the rotation angle between AA and BB. Convert the quaternion q(n)=[qwn,qxn,qyn,qzn]⊤\mathbf{q}(n)=\left[q^{n}_{w},q^{n}_{x},q^{n}_{y},q^{n}_{z}\right]^{\top} to the rotation matrix R(n)∈SO(3)\mathbf{R}(n)\in SO(3), then

the homogeneous transformation matrix from the world frame WW to the nn-th interpolated frame denoted as TnW∈SE(3)\mathbf{T}^{W}_{n}\in SE(3) is:

2 Motion Distortion Simulation

where [⋅;⋅;...;⋅][\cdot;\cdot;...;\cdot] is the concatenation operation along the row. It is noted that m=∑n=0N−1mnm=\sum_{n=0}^{N-1}m_{n}.

3 Motion Compensation with Ego-Pose

where TnA{\mathbf{T}^{A}_{n}} is the transformation from frame AA to frame nn, which is the inverse matrix of TAn{\mathbf{T}^{n}_{A}} in Eq. (5).

Adversarial Trajectory Perturbation

The studies of object-level point cloud generally treat point cloud as a set of 3D points sampled from mesh models instantaneously . In autonomous driving, however, the 3D points are captured in a dynamic setting through raycasting, thus LiDAR not only records points’ (x,y,z)(x,y,z) coordinates, but also the timestamps at which the points are captured. To this end, we propose a novel representation of point cloud as a function of vehicle trajectory through Eq. (7) which can be written as a general form:

Physical Feasibility. Since the motion compensation is naturally occurring in self-driving, it is physically-realizable and straightforward to attack the trajectory, \eg, by wireless GNSS spoofing. In contrast, coordinate attack requires software hacking which is infeasible in practice.

Better Transferability. Our learned trajectory perturbation with the same size (N×4×4N\times 4\times 4) can be easily transferred to different sweeps. Yet, coordinate attacks cannot be transferred across sweeps because different sweeps could have different numbers of points, and the perturbation in the point space could have different dimensions.

Novel Parameterization. We attack the 6-DoF pose of each packet, but coordinate attack modifies a single point’s xyz position without orientations. Hence, our method has better performance due to new attacked parameters.

2 Objective Function

Since the point cloud is represented as a differentiable function of the trajectory, the gradient can be back-propagated to the trajectory smoothly for adversarial learning. The adversarial objective is to minimize the negative loss function of the deep model with parameter θ\theta:

Polynomial Trajectory Perturbation. To make the attack highly imperceptible, the polynomial trajectory perturbation is proposed: we define the perturbation in translations as a third-order polynomial of time as follows:

where n=[1,n,n2,n3]⊤\mathbf{n}=[1,n,n^{2},n^{3}]^{\top} and β=[βx,βy,βz]\bm{\beta}=[\bm{\beta_{x}},\bm{\beta_{y}},\bm{\beta_{z}}] is the polynomial coefficients. Now δ\bm{\delta} is a differentiable function of β\bm{\beta}, and the gradient will be back-propagated to β\bm{\beta}, so the adversarial coefficients are calculated by:

hence, we only need to manipulate several key points to bend a polynomial-parameterized trajectory which can be easily achieved in reality, achieving a real-time attack.

3 Attack with Regularization

where λt\lambda_{\mathbf{t}} and λR\lambda_{\mathbf{R}} are used to balance the influence of translations and rotations, pp denotes the norm (p=2p=2 in this work). With trajectory smoothness regularization, the objective of adversarial attacks is as follows:

where λs\lambda_{s} is for adjusting the smoothness degree. By optimizing over Eq. (15), we aim to find an imperceptible adversarial perturbation δ\bm{\delta} with desirable smoothness.

where D(δ)\mathcal{D}(\bm{\delta}) can be either DL(δ)\mathcal{D}_{L}(\bm{\delta}) or DC(δ)\mathcal{D}_{C}(\bm{\delta}). λd\lambda_{d} is to control the degree of distortion. By optimizing over Eq. (16), we try to search for a powerful adversarial perturbation δ\bm{\delta} leading to subtle distortion in the point cloud.

Experiments

In this work, we select widely-studied LiDAR-based 3D detection, which aims to estimate 3D bounding boxes of the objects in point cloud, as a downstream task example to verify our attack pipeline. Currently, LiDAR-based detection has two mainstreams: 1) point-based method directly consuming raw point cloud data, 2) voxel-based method which requires non-differentiable voxelization in the preprocessing stage. For the white box attack, we use point-based PointRCNN . For the black box transferability test, we adopt voxel-based PointPillar++ .

PointRCNN. Our white box model, PointRCNN, uses PointNet++ as its backbone and includes two stages: stage-1 for proposal generation based on each foreground point, and stage-2 for proposal refinement in the canonical coordinate. Since PointRCNN uses raw point cloud as the input, the gradient can smoothly reach the point cloud, then arrive at vehicle trajectory. In this work, we individually attack the classification as well as regression branches in stage-1 and stage-2, with four attack targets in total.

PointPillar++. PointPillar proposes a fast point cloud encoder using pseudo-image representation. It divides point cloud into bins and uses PointNet to extract the feature for each pillar. Due to the non-differentiable preprocesssing stage, the gradient cannot reach the point cloud. Peiyun et al. proposed to augment PointPillar with the visibility map, achieving better precision. In this work, we use PointPillar++ to denote PointPillar with visibility map in . We use perturbation learned from the white box PointRCNN to attack black box PointPillar++, in order to examine the transferability of our attack pipeline.

2 Dataset and Evaluation Metrics

Dataset. nuScenes is a large-scale multimodel autonomous driving dataset captured by a real SDV with a full 360∘360^{\circ} field of view in various challenging urban driving scenarios. Including 1000 scenes collected in Boston and Singapore in different weather, nuScenes has much more annotations (7 times) and images (100 times) than the pioneer KITTI . Besides, nuScenes provides a temporal sequence of samples in each scene, facilitating linear pose interpolation for motion distortion simulation, yet the 3D detection dataset in KITTI solely offers independent frames without temporal connection. Considering the above factors, nuScenes is employed in this work. We use PointRCNN model released in , and the open-source PointPillar++ model in . Both models are trained on nuScenes training set. We report the results of white box on 1,000 samples from the validation set, and the results of black box on the whole validation set.

Metrics. For the white box PointRCNN, we report the 3D bounding boxes average precision (AP) with the IoU thresholds at 0.7 on car category. Following , we evaluate the detector in three scenarios (Easy, Moderate, and Hard) based on the difficulty level of the surrounding cars. In addition, the performance within three depth ranges, \ie, 0∼300\sim 30, 30∼5030\sim 50, and 50∼7050\sim 70 meters, are also assessed. For the black box PointPillar++, we follow the original paper to employ official 3D detection evaluation protocol in nuScenes , \ie, average mAP over ten categories at four distance thresholds. For evaluating attack quality, we utilize the performance drop after attack.

3 Experimental Setup

Implementation details. nuScenes samples keyframes at 2Hz from the original 20Hz data, so we assume that a sweep consumes 0.5 secondIn fact, the LiDAR rotation period is 0.05 second, however, we only have access to annotations on keyframes at 2Hz, so we have to assume the capture frequency of LiDAR is also 2Hz. and implement linear pose interpolation between two adjacent keyframes. We set total interpolation step NN as 100. For PGD, we restrict the perturbation to 10cm in translation, and to 0.01 in rotation. The step size for each attack iteration is 0.1/0.01 in translation/rotation, and the number of iterations is 20. More experimental settings like the step size and the number of iterations are reported in the supplementary.

Baselines. To demonstrate the superiority of our attack pipeline, we employ two baseline methods for comparison.

Random Attack. We add Gaussian noise with a 10cm standard deviation to the original point cloud, Gaussian noise with a 10cm standard deviation to the translations and Gaussian noise with 0.01 standard deviation to rotations.

Coordinate Attack. Coordinate attack is directly manipulating the point set to deceive the detector. We attack classification of stage-2 and restrict the change in point coordinate to 10cm. The binary search step number is 10 and the number of iteration for each binary search step is 100.

Attack Settings. The polynomial trajectory attack is temporally-smooth and the others are temporally-discrete:

Attack Translation Only. We merely modify the translation vector. The perturbation is a 100×3100\times 3 matrix.

Polynomial Perturbation. We add a polynomial perturbation into the vehicle trajectory (translation part).

Attack Rotation Only. We only interfere with the rotation matrix. The perturbation is a 100×3×3100\times 3\times 3 tensor.

Attack Full Trajectory. We tamper with the transformation matrix. The perturbation is a 100×4×4100\times 4\times 4 tensor.

4 White Box Attack

To explore the vulnerability of different stages and branches in PointRCNN, we separately attack its four modules, \ie, classification/regression branch of stage-1/2. Qualitative results are displayed in Fig. 3 and more examples can be found in the supplementary.

Attack Translation Only. The quantitative results are exhibited in Table 1. Simply adding random noise with a small standard deviation can largely decrease the performance, \eg, AP in the easy case is reduced by 30.4430.44 (around 64%64\%). This phenomenon should stand as a warning for the autonomous driving community. Also, attacking classification branch of stage-2 is the most effective way to fool the detector: AP is respectively lowered by 35.7235.72 (75.3%75.3\%), 15.6915.69 (72.8%72.8\%) and 14.8514.85 (71.0%71.0\%) in the scenarios of easy, moderate and hard. This is because the stage-2 outputs the final predictions and is extremely safety-relevant. Moreover, we can find that attacking the classification branch is more effective than attacking regression, which is reasonable because classifying the objects takes priority over estimating their sizes in the detection task. Compared to the random attack, our method is more detrimental thanks to our making use of the vulnerability pointed out by the gradient. In easy, moderate, and hard scenarios, adversarial perturbation generated by attacking stage-2’s classification can yield additional drops of 5.285.28, 2.712.71, and 2.842.84 in AP, compared to the random perturbation. Besides, our attack is better than the coordinate attack due to the parameterization, validating the superiority of the trajectory attack.

Polynomial Trajectory Perturbation. When attacking polynomial coefficients instead of the individual trajectory point, the performance is still on par with the discrete setting, as shown in Table 1, yet the attack is highly imperceptible especially in the point cloud space as shown in Fig. 4.

Attack Rotation Only. Coordinate attack is manipulating each point’s xyz position. In contrast, our method treat each LiDAR packet as a rigid body, therefore we can attack the rotation. From Table 1 we can find that attacking classification in stage-2 is still the most powerful way to fool PointRCNN. Besides, attacking rotation achieves more performance drop compared to attacking translation, \eg, in easy, moderate, and hard situations, AP is respectively decreased by 45.0945.09 (95.0%95.0\%), 20.7620.76 (96.3%96.3\%) and 20.3020.30 (97.1%97.1\%) in comparison with original PointRCNN. Moreover, fooling regression in stage-1 achieves the second best attacking quality, proving the effectiveness of attacking the fundamental proposal generation. Besides, attacking regression of stage-2 has the worst attacking quality, proving that attacking the refinement of box size has no significant effect. Compared to the random attack, attacking stage-2’s classification has realized additional AP drop (9.959.95, 4.074.07, 4.524.52), validating the merits of adversarial learning.

Attack Full Trajectory. As shown in Table 1, fooling the full trajectory has achieved the best attacking quality, \eg, while attacking classification of stage-2, AP can be decreased to nearly zero. Compared to original PointRCNN, AP is respectively lowered by 47.2547.25 (99.6%99.6\%) in easy scenarios, indicating that the detector is completely blinded.

5 Black Box Attack

We choose voxel-based PointPillar++ to test the transferability across different input representation. The quantitative results on ten categories are shown in Table 2: the performance is still largely dropped by our attack, \eg, mAP on ten categories can be reduced by 20.620.6 (58.2%58.2\%), and the AP on car is decreased by 35.035.0 (43.8%43.8\%). Meanwhile, our method has demonstrated satisfactory transferability across categories. With only adversarial learning against the car detector, the detection of other categories are also deceived. For small object like pedestrian/motorcycle, AP can be dropped by 48.348.3 (72.2%72.2\%)/18.018.0 (97.3%97.3\%), which has demonstrated the superior attacking quality of our method against small object detector. This superiority is mainly because the small object with less points is more susceptible to the perturbation compared to the large object like bus (dropped by 32.132.1, 59.3%59.3\%) or truck (dropped by 17.517.5, 48.9%48.9\%). Several qualitative examples are displayed in Fig. 4 and more examples are in the supplementary.

Conclusion

We proposed a generic and feasible DNN attack pipeline based on the trajectory against LiDAR perception. We conduct experiments on the well-studied 3D object detection task. In white box attack, even only with a 10cm perturbation in translations, the precision can be dropped by around 70%\%. While attacking the full trajectory, the precision can be decreased to nearly zero, yet the attack is less perceptible (especially the point clouds). Our attack also shows good transferability across various input representations and target categories, raising a red flag for perception systems using LiDAR and DNN jointly.

Acknowledgment. The research is supported by NSF FW-HTF program under DUE-2026479. The authors gratefully acknowledge the useful comments and suggestions from Yong Xiao, Wenxiao Wang, Chenzhuang Du, Wang Zhao, Ziyuan Huang, Hang Zhao and Siheng Chen, and also thank Yan Wang, Shaoshuai Shi and Peiyun Wu for their helpful open-source code.

References

Appendix

I Overview

In the supplementary, we first present more results of our FLAT attack against black box PointPillar++ . Then more qualitative evaluations of white box attack are shown, and the performances under different parameters (the number of attack iteration, the step size for each iteration, and the interpolation step) are discussed. Finally, more examples of polynomial trajectory perturbation (including both the point cloud after our attack and and the trajectory perturbation) are presented to validate the imperceptibility in point cloud space and the strong smoothness in trajectory.

II Black Box Attack

Additional qualitative examplesNoted that we transfer the adversarial perturbation of the white-box full trajectory attack targeting classification of stage-2 are presented in Fig. I and Fig. II. Although our FLAT attack cannot know the model parameters of PointPillar++, a subtle perturbation can make the detector lose many safety-critical objects, \eg, as for the four sweeps in Fig. I, the adversarial perturbation in the full trajectory can make the detector lose 7, 20, 8 and 6 objects respectively (from top to bottom), severely damaging the self-driving car’s perception module. Moreover, false positives are also increased by our perturbation (6 more in the third row of Fig. I), making the car mistakenly believe that there are obstacles in the free space.

II.2 Quantitative Evaluation

The nuScenes dataset employs 2D center distance (0.5, 1, 2, 4 meters) as the matching threshold when calculating the Average Precision (AP). The per category precision-recall plots of the original detector as well as four attack settings are shown from Fig. III - Fig. VII. The performance drop is the largest when the threshold is 0.5m (the highest precision standard), \eg, the AP@0.5m in car category is decreased by 23.123.1(33.1%33.1\%), 36.536.5(52.4%52.4\%), 41.541.5(59.6%59.6\%) while attacking translation, rotation and the full trajectory. As shown in Fig. VIII, the curve shifts to the lower left when increasing the attack magnitude.

III White Box Attack

Three qualitative examples are demonstrated in Fig. IX - Fig. XI. In Fig. IX, original PointRCNN can perceive three our of five cars, and attacking translation only has slightly shifted the predictions. Attacking rotation has a stronger impact and cause two objects to be undetectable, while attacking the full trajectory further creates four false positives. In Fig. X, PointRCNN can successfully detect two out of three surrounding cars, then adversarial translation perturbation make the detection of the nearest car drift. Both two detections are drifted in the scenario of adversarial rotation perturbation. Finally, when attacking both translation and rotation, five false positives have emerged, severely degrading the perception capability of the self-driving car. Similar results can be found in Fig. XI. Moreover, three qualitative examples of attacking with and without regularization are presented in Fig. XII, Fig. XIII and Fig. XIV. As the strength of regularization is enlarged, the variation in the point space is reduced, enabling less perceptible attack.

III.2 Different Parameters

The detection performances under different parameter settings are reported in Table I - Table III. In the case of attacking the translation (fooling classification in stage-2), AP of easy case in six parameter settings fluctuate between 11.7211.72 and 13.8913.89, which means that all the parameter configurations can produce a high-quality adversarial perturbation. In the situation of attacking the full trajectory (with the aim of attacking regression branch in stage-1), AP of easy case in six parameter settings fluctuate between 0.530.53 and 2.442.44, proving that all the six settings can maintain the attacking quality at a high level. To sum up, the number of iteration and the step size for each iteration has a relatively small impact on our FLAT attack. As for the interpolation step, when it is increased, there will be more sectors, \ie, more point cloud groups, and the number of trajectory points which can be attacked are also increased, easing the learning of an effective adversarial perturbation. Consequently, when the interpolation step is increased from 50 to 1000, the AP in easy scenarios can be lowered by 5.045.04 (36.7%36.7\%) while attacking the translation, by 3.703.70 (78.2%78.2\%) while attacking the rotation, and by 0.510.51 (87.9%87.9\%) while attacking the full trajectory, as shown in Table IV.

Noted that our simulation of motion distortion/correction accurately matches with the real-world scenarios, \eg, in Fig. 2 of : “the Velodyne VLP-16 software used in our electric vehicle platform produces 76 packets for each full revolution scan. Each packet covers an azimuth angle of approximately 4.74∘4.74^{\circ}.” Noted that the number of point cloud sector/packet is adjustable depending on vehicle platforms, and the attack performance keeps at a high level under different number of sector/interpolation step, as shown in Table IV.

IV Polynomial Trajectory Perturbation

Our trajectory attack is feasible, \eg, through GNSS spoofing as proven in : “Today, it is feasible to execute GNSS spoofing attacks with less than \100$ of equipment. GNSS signal generators can be programmed to transmit radio frequency signals corresponding to a static position, or simulate entire trajectories.” We have verified the effectiveness of the discrete adversarial trajectory perturbation in both white box and black box attack. To achieve a temporally-smooth attack which is less perceptible, we implement a polynomial regression before the generation of perturbation and attack the polynomial coefficients instead of the trajectory itself. In this scenario, we only need to manipulate several key points to bend a polynomial-parameterized trajectory which can be easily achieved in reality, realizing a real-time and high-quality attack. Several qualitative examples of the polynomial trajectory perturbation are shown in Fig. XV. Although the attack performance is inferior to the full trajectory attack, the detector still missed many safety-critical objects yet the perturbation in point cloud space is highly imperceptible: even for human eyes, it is quite difficult to distinguish the point cloud before and after the polynomial trajectory perturbation.

V Summary

In this supplementary, we presented more qualitative evaluations of both white box and black box attack, to validate the effectiveness of our attack pipeline. Besides, we found that the number of PGD iteration and the step size for each iteration have a relatively small impact on the attack quality, but the interpolation step, \ie, the number of LiDAR packets, can have a relatively large influence on the attack performance, because increasing trajectory points which can be deliberately modified can facilitate the adversarial learning, resulting in a stronger adversarial perturbation. Finally, more qualitative examples of the polynomial trajectory perturbation are demonstrated to validate the imperceptibility of out attack.