Beyond 3D Siamese Tracking: A Motion-Centric Paradigm for 3D Single Object Tracking in Point Clouds

Chaoda Zheng, Xu Yan, Haiming Zhang, Baoyuan Wang, Shenghui Cheng, Shuguang Cui, Zhen Li

Introduction

Single Object Tracking (SOT) is a basic computer vision problem with various applications, such as autonomous driving and surveillance system . Its goal is to keep track of a specific target across a video sequence, given only its initial state (appearance and location).

Existing LiDAR-based SOT methods all follow the Siamese paradigm, which has been widely adopted in 2D SOT since it strikes a balance between performance and speed. During the tracking, a Siamese model searches for the target in the candidate region with an appearance matching technique, which relies on the features of the target template and the search area extracted by a shared backbone (see Fig.1(a)).

Though the appearance matching for 3D SOT shows satisfactory results on KITTI dataset , we observe that KITTI has the following proprieties: i) the target’s motion between two consecutive frames is minor, which ensures no drastic appearance change; ii) there are few/no distractors in the surrounding of the target. However, the above characteristics do not hold in natural scenes. Due to self-occlusion, significant appearance changes may occur in consecutive LiDAR views when objects move fast, or the hardware only supports a low frame sampling rate. Besides, the negative samples grow significantly in dense traffic scenes. In these scenarios, it is not easy to locate a target based on its appearance alone (even for human beings).

Is the appearance matching the only solution for LiDAR SOT? Actually, Motion Matters. Since the task deals with a dynamic scene across a video sequence, the target’s movements among successive frames are critical for effective tracking. Knowing this, researchers have proposed various 2D Trackers to temporally aggregate information from previous frames . However, the motion information is rarely explicitly modeled since it is hard to be estimated under the perspective distortion. Fortunately, 3D scenes keep intact information about the object motion, which can be easily inferred from the relationships among annotated 3D bounding boxes (BBoxes)This is greatly held for rigid objects (e.g. cars), and it is approximately true for non-rigid objects (e.g. pedestrian).. Although 3D motion matters for tracking, previous approaches have greatly overlooked it. Due to the Siamese paradigm, previous methods have to transform the target template (initialized by the object point cloud in the first target 3D BBox and updated with the last prediction) from the world coordinate system to its own object coordinate system. This transformation ensures that the shared backbone extracts a canonical target feature, but it adversely breaks the motion connection between consecutive frames.

Based on the above observations, we propose to tackle 3D SOT from a different perspective instead of sticking to the Siamese paradigm. For the first time, we introduce a new motion-centric paradigm that localizes the target in sequential frames without appearance matching by explicitly modeling the target motion between successive frames (Fig. 1(b)). Following this paradigm, we design a novel two-stage tracker M2-Track (Fig. 2). During the tracking, the 1st1^{st}-stage aims at generating the target BBox by predicting the inter-frame relative target motion. Utilizing all the information from the 1st1^{st}-stage, the 2nd2^{nd}-stage refines the BBox using a denser target point cloud, which is aggregated from two partial target views using their relative motion. We evaluate our model on KITTI , NuScenes and Waymo Open Dataset (WOD) , where NuScenes and WOD cover a wide variety of real-world environments and are challenging for their dense traffics. The experiment results demonstrate that our model outperforms the existing methods by a large margin while running faster than the previous top-performer . Besides, the performance gap becomes even more significant when more distractors exist in the scenes. Furthermore, we demonstrate that our method can directly benefit from appearance matching when integrated with existing methods.

In summary, our main contributions are as follows:

A novel motion-centric paradigm for real-time LiDAR SOT, which is free of appearance matching.

A specific second-stage pipeline named M2-Track that leverages the motion-modeling and motion-assisted shape completion.

State-of-the-art online tracking performance with significant improvement on three widely adopted datasets (i.e. KITTI, NuScenes and Waymo Open Dataset).

Related Work

Single Object Tracking. A majority of approaches are built for camera systems and take 2D RGB images as input . Although achieving promising results, they face great challenges when dealing with low light conditions or textureless objects. In contrast, LiDARs are insensitive to texture and robust to light variations, making them a suitable complement to cameras. This inspires a new trend of SOT approaches which operate on 3D LiDAR point clouds. These 3D methods all inherit the Siamese paradigm based on appearance matching. As a pioneer, uses the Kalman filter to heuristically sample a bunch of target proposals, which are then compared with the target template based on their feature similarities. The proposal which has the highest similarity with the target template is selected as the tracking result. Since heuristic sampling is time-consuming and inhibits end-to-end training, propose to use a Region Proposal Network (RPN) to generate high-quality target proposals efficiently. Unlike which uses an off-the-shelf 2D RPN operating on bird’s eye view (BEV), adapts SiamRPN to 3D point clouds by integrating a point-wise correlation operator with a point-based RPN . The promising improvement brought by inspires a series of follow-up works . They focus on either improving the point-wise correlation operator by feature enhancement, or refining the point-based RPN with more sophisticated structures.

The appearance matching achieves excellent success in 2D SOT because images provide rich texture, which helps the model distinguish the target from its surrounding. However, LiDAR point clouds only contain geometric appearances that lack texture information. Besides, objects in LiDAR sweeps are usually sparse and incomplete. These bring considerable ambiguities which hinder effective appearance matching. Unlike existing 3D approaches, our work no more uses any appearance matching. Instead, we examine a new motion-centric paradigm and show its great potential for 3D SOT.

3D Multi-object Tracking / Detection. In parallel with 3D SOT, 3D multi-object tracking (MOT) focuses on tracking multiple objects simultaneously. Unlike SOT where the user can specify a target of interest, MOT relies on an independent detector to extract potential targets, which obstructs its application for unfamiliar objects (categories unknown by the detector). Current 3D MOT methods predominantly follow the “tracking-by-detection” paradigm, which first detects objects at each frame and then heuristically associates detected BBoxes based on objects’ motion or appearance . Recently, proposes to jointly perform detection and tracking by combining object detection and motion association into a unified pipeline. Our motion-centric tracker draws inspiration from the motion-based association in MOT. But unlike MOT, which applies motion estimation on detection results, our approach does not depend on any detector and can leverage the motion prediction to refine the target BBox further.

Spatial-temporal Learning on Point Clouds. Our method utilizes spatial-temporal learning to infer relative motion from multiple frames. Inspired by recent advances in natural language processing , there emerges methods that adapt LSTM , GRU , or Transformer to model point cloud videos. However, their heavy structures make them impractical to be integrated with other downstream tasks, especially for real-time applications. Another trend forms a spatial-temporal (ST) point cloud by merging multiple point clouds with a temporal channel added to each point . Treating the temporal channel as an additional feature (like RGB or reflectance), one can process such an ST point cloud using any 3D backbones without structural modifications. We adopt this strategy to process successive frames for simplicity and efficiency.

Methodology

2 Motion-centric Paradigm

Having the predicted RTM Mt−1,t\mathcal{M}_{t-1,t}, one can easily obtain the target BBox in Pt\mathcal{P}_{t} using rigid body transformation:

Following the motion-centric paradigm, we design a two-stage motion-centric tracking pipeline M2M^{2}-Track (illustrated in Fig.2). M2M^{2}-Track first coarsely localizes the target through target segmentation and motion transformation at the 1st1^{st} stage, and then refines the BBox at the 2nd2^{nd} stage using motion-assisted shape completion. More details of each module are given as follows.

Target segmentation with spatial-temporal learning

Stage I: Motion-Centric BBox prediction As shown in Fig. 3, we encode the spatial-temporal target point clouds P~t−1,t\mathcal{\widetilde{P}}_{t-1,t} into an embedding using another PointNet encoder. A multi-layer perceptron (MLP) is applied on top of the embedding to obtain the motion state of the target, which includes a 4D RTM Mt−1,t\mathcal{M}_{t-1,t} and 2D binary classification logits indicating whether the target is dynamic. To reduce accumulation errors while performing frame-by-frame tracking, we generate a refined previous target BBox B~t−1\mathcal{\widetilde{B}}_{t-1} by predicting its RTM with respect to Bt−1\mathcal{B}_{t-1} through another MLP (More details are presented in the supplementary). Finally, we can get the current target BBox Bt\mathcal{B}_{t} by applying Eqn. 2 on Mt−1,t\mathcal{M}_{t-1,t} and B~t−1\mathcal{\widetilde{B}}_{t-1} if the target is classified as dynamic. Otherwise, we simply set Bt\mathcal{B}_{t} as B~t−1\mathcal{\widetilde{B}}_{t-1}.

4 Box-aware Feature Enhancement

5 Implementation Details

Loss Functions. The loss function contains classification losses and regression losses, which is defined as L=λ1Lcls_target+λ2Lcls_motion+λ3(Lreg_motion+Lreg_refine_prev+Lreg_1st+Lreg_2nd)L=\lambda_{1}L_{\text{cls\_target}}+\lambda_{2}L_{\text{cls\_motion}}+\lambda_{3}(L_{\text{reg\_motion}}+L_{\text{reg\_refine\_prev}}+L_{\text{reg\_1st}}+L_{\text{reg\_2nd}}). Lcls_targetL_{\text{cls\_target}} and Lcls_motionL_{\text{cls\_motion}} are standard cross-entropy losses for target segmentation and motion state classification at the 1st1^{st}-stage (Points are considered as the target if they are inside the target BBoxes; A target is regarded as dynamic if its center moves more than 0.15 meter between two frames). All regression losses are defined as the Huber loss between the predicted and ground-truth RTMs (inferred from ground-truth target BBoxes), where Lreg_motionL_{\text{reg\_motion}} is for the RTM between targets in the two frames; Lreg_refine_prevL_{\text{reg\_refine\_prev}} is for the RTM between the predicted and the ground-truth BBoxes at timestamp (t−1)(t-1); Lreg_1stL_{\text{reg\_1st}} / Lreg_2ndL_{\text{reg\_2nd}} is for the RTM between the 1st1^{st} / 2nd2^{nd}-stage and ground-truth BBoxes. We empirically set λ1=λ2=0.1\lambda_{1}=\lambda_{2}=0.1 and λ3=1\lambda_{3}=1.

Input & Motion Augmentation. Since SOT only takes care of one target in a scene, we only need to consider a subregion where the target may appear. For two consecutive frames at (t−1)(t-1) and tt timestamp, we choose the subregion by enlarging the target BBox at (t−1)(t-1) timestamp by 2 meters. We then sample 1024 points from the subregion respectively at (t−1)(t-1) and tt timestamp to form Pt−1\mathcal{P}_{t-1} and Pt\mathcal{P}_{t}. To simulate testing errors during the training, we feed the model a perturbed BBox by adding a slight random shift to the ground-truth target BBox at (t−1)(t-1) timestamp. To encourage the model to learn various motions during the training, we randomly flip both targets’ points and BBoxes in their horizontal axes and rotate them around their upup-axes by Uniform[−10∘-10^{\circ},10∘10^{\circ}]. We also randomly translate the targets by offsets drawn from Uniform [-0.3, 0.3] meter.

Training & Inference. We train our models using the Adam optimizer with batch size 256 and an initial learning rate 0.001, which is decayed by 10 times every 20 epochs. The training takes ∼4\sim 4 hours to converge on a V100 GPU for the KITTI Cars. During the inference, the model tracks a target frame-by-frame in a point cloud sequence given the target BBox at the first frame.

Experiments

Datasets. We extensively evaluate our approach on three large-scale datasets: KITTI , NuScenes and Waymo Open Dataset (WOD) . We follow to adapt these datasets for 3D SOT by extracting the tracklets of annotated tracked instances from each of the scenes. KITTI contains 21 training sequences and 29 test sequences. We follow previous works to split the training set into train/val/test splits due to the inaccessibility of the test labels. NuScenes contains 1000 scenes, which are divided into 700/150/150 scenes for train/val/test. Officially, the train set is further evenly split into “train_track” and “train_detect” to remedy overfitting. Following , we train our model with “train_track” split and test it on the val set. WOD includes 1150 scenes with 798 for training, 202 for validation, and 150 for testing. We do training and testing respectively on the training and validation set. Note that NuScenes and WOD are much more challenging than KITTI due to larger data volumes and complexities. The LiDAR sequences are sampled at 10Hz for both KITTI and WOD. Though NuScenes samples at 20Hz, it only provides the annotations at 2Hz. Since only annotated keyframes are considered, such a lower frequency for keyframes introduces additional difficulties for NuScenes.

Evaluation Metrics. We evaluate the models using the One Pass Evaluation (OPE) . It defines overlap as the Intersection Over Union (IOU) between the predicted and ground-truth BBox, and defines error as the distance between two BBox centers. We report the Success and Precision of each model in the following experiments. Success is the Area Under the Curve (AUC) with the overlap threshold varying from 0 to 1. Precision is the AUC with the error threshold from 0 to 2 meters.

2 Comparison with State-of-the-arts

Results on KITTI. We compare M2M^{2}-Track with seven top-performance approaches , which have published results on KITTI. As shown in Tab. 1, our method benefits both rigid and non-rigid object tracking, outperforming current approaches under all categories except Car, where PTT and V2B surpass us by minor margins. The lack of car distractors in the scenes makes our improvement over previous appearance-matching-based methods minor for cars. But our improvement for pedestrians is significant (13.2%/13.7% in terms of success/precision) because pedestrian distractors are widespread in the scenes (see the supplementary for more details). Besides, methods using point-based RPN all perform badly on cyclists, which are relative small in size but usually move fast across time. The second row in Fig. 5 shows the case in which a cyclist moves rapidly across frames. Our method perfectly keeps track of the target while BAT almost fails. To handle such fast-moving objects, leverage BEV-based RPN to generate high-recall proposals from a larger search region. In contrast, we handle this simply by motion modeling without sophisticated architectures.

Results on NuScenes & WOD. We select three representative open-source works: SC3D , P2B and BAT as our competitors on NuScenes and WOD. The results on NuScenes except for the Pedestrian class are provided by . We use the published codes of the competitors to obtain other results absent in . SC3D is omitted for WOD comparison due to its costly training time. As shown in Tab. 2, M2M^{2}-Track exceeds all the competitors under all categories, mostly by a large margin. On such two challenging datasets with pervasive distractors and drastic appearance changes, the performance gap between previous approaches and M2M^{2}-Track becomes even larger (e.g. more than 30% precision gain on Waymo Pedestrian). Note that for large objects (i.e. Truck, Trailer, and Bus), even if the predicted centers are far from the target (reflected from lower precision), the output BBoxes of the previous model may still overlap with the ground truth (results in higher success). In contrast, the motion modeling helps to improve not only the success but also the precision by a large margin (e.g.+23.43% gain on Bus) for large objects. Visualization results are provided in Fig. 5 and the supplementary.

3 Analysis Experiments

In this section, we extensively analyze M2M^{2}-Track with a series of experiments. Firstly, we compare the behaviors of M2M^{2}-Track and previous appearance-matching-based methods in different setups. Afterward, we equip M2M^{2}-Track with the previous appearance matching approaches to show its potential. Finally, we study the effectiveness of each component in M2M^{2}-Track. All the experiments are conducted on the Car category of KITTI unless otherwise stated.

Robustness to Distractors. Though achieving promising improvement on NuScenes and WOD, M2M^{2}-Track brings little improvement on the Car of KITTI. To explain this, we look at the scenes of three datasets and find that the surroundings of most cars in KITTI are free of distractors, which are pervasive in NuScenes and WOD. Although appearance-matching-based methods are sensitive to distractors, they provide more precise results than our motion-based approach in distractor-free scenarios. But as the number of distractors increases, these methods suffer from noticeable performance degradation due to ambiguities from the distractors. To verify this hypothesis, we randomly add KK car instances to each scene of KITTI, and then re-train and evaluate different models using this synthesis dataset. As shown in Fig. 6, M2M^{2}-Track consistently outperforms the other two matching-based methods in scenes with more distractors, and the performance gap grows as KK increases. Thanks to the box-awareness, BAT can aid such ambiguities to some extent. But our performance is more stable than BAT’s when more distractors are added. Besides, the first row in Fig. 5 shows that, when the number of points decreases due to occlusion, BAT is misled by a distractor and then tracks off course, while M2M^{2}-Track keeps holding tight to the ground truth. All these observations demonstrate the robustness of our approach.

Influence of Motion Augmentation. We improve the performance of M2M^{2}-Track using the motion augmentation in training, which is not adopted in previous approaches. For a fair comparison, we re-train BAT and P2B using the same configurations in their open-source projects except additionally adding motion augmentation. Tab. 3 shows that motion augmentation instead has an adverse effect on both BAT and P2B. Our model benefits from motion augmentation since it explicitly models target motion and is robust to distractors. In contrast, motion augmentation may move a target closer to its potential distractors and thus harm those appearance-matching-based approaches.

Combine with Appearance Matching. Although our motion-centric model outperforms previous methods from various aspects, appearance-matching-based approaches still show their advantage when dealing with distractor-free scenarios. To combine the advantages of both motion-based and matching-based methods, we apply BAT/P2B as a “re-tracker” to fine-tune the results of M2M^{2}-Track. Specifically, we directly utilize BAT/P2B to search for the target in a small neighborhood of the M2M^{2}-Track’s output. Tab. 4 confirms that M2M^{2}-Track can further benefit from appearance matching, even under this naive combination. On KITTI Car, both combined models outperform the top-ranking PTT by noticeable margins. We believe that one can further boost 3D SOT by combining motion-based and matching-based paradigms with a more delicate design.

Ablations. In Tab. 5, we conduct an exhaustive ablation study on both KITTI and NuScenes to understand the components of M2M^{2}-Track. Specifically, we respectively ablate the box-aware feature enhancement, previous BBox refinement, binary motion classification and 2nd2^{nd} stage from M2M^{2}-Track. In general, the effectiveness of the components varies across the datasets, but removing any one of them causes performance degradation. The only exception is the binary motion classification used in the 1st1^{st} stage, which causes a slight drop on KITTI in terms of success. We suppose this is due to the lack of static objects for KITTI’s cars, which results in a biased classifier. Besides, Tab. 5 shows that M2M^{2}-Track keeps performing competitively even with module ablated, especially on NuScenes. This reflects that the main improvement of M2M^{2}-Track is from the motion-centric paradigm instead of the specific pipeline design.

More Discussion

Running Overheads. M2M^{2}-Track achieves exciting performance with just a simple PointNet . Compared with other hierarchical backbones (e.g. ) used in previous works, PointNet saves more computational overheads since it does not perform any sampling or grouping operations, which are not only time-consuming but also memory-intensive. Therefore, M2M^{2}-Track runs 1.67×\times faster as the previous top-performer BAT (only consider model forwarding time) but saves 31.1% memory footprint. Using a more advanced backbone (e.g. ) may further boost the performance but inevitably slows down the running speed. Since we focus on online tracking, we prefer a simpler backbone to balance performance and efficiency.

Limitations. Unlike appearance matching, our motion-centric model requires a good variety of motion in the training data to ensure its generalization on data sampled with different frequencies. For instance, our model suffers from considerable performance degradation if trained with 2Hz data but tested with 10Hz data because the motion distribution of the 2Hz and 10Hz data differs significantly. But fortunately, we can aid this using a well-design motion augmentation strategy.

Conclusions

In this work, we revisit 3D SOT in LiDAR point clouds and propose to handle it with a new motion-centric paradigm, which is proven to be an excellent complement to the matching-based Siamese paradigm. In addition to the new paradigm, we propose a specific motion-centric tracking pipeline M2M^{2}-Track, which significantly outperforms the state-of-the-arts from various aspects. Extensive analysis confirms that the motion-centric model is robust to distractors and appearance changes and can directly benefit from previous matching-based trackers. We believe that the motion-centric paradigm can serve as a primary principle to guide future architecture designs. In the future, we may try to improve M2M^{2}-Track by considering more frames and integrating it with the appearance matching under a more delicate design.

Acknowledgment

This work was supported in part by NSFC-Youth 61902335, by Key Area R&D Program of Guangdong Province with grant No.2018B030338001, by the National Key R&D Program of China with grant No.2018YFB1800800, by Shenzhen Outstanding Talents Training Fund, by Guangdong Research Project No.2017ZT07X152, by Guangdong Regional Joint Fund-Key Projects 2019B1515120039, by the NSFC 61931024&81922046, by helixon biotechnology company Fund and CCF-Tencent Open Fund.

References