3D Siamese Voxel-to-BEV Tracker for Sparse Point Clouds
Le Hui, Lingpeng Wang, Mingmei Cheng, Jin Xie, Jian Yang
Introduction
Object tracking is an essential task in computer vision and has been widely in various applications, such as autonomous vehicle, mobile robotics, and augmented reality. In the past few years, many efforts kristan2015visual ; bertinetto2016fully ; danelljan2017eco ; valmadre2017end have been made on 2D object tracking from RGB data. Recently, with the development of 3D sensor such as LiDAR and Kinect, 3D object tracking Lebeda20142DON ; Whelan2016ElasticFusionRD ; Rnz2017CofusionRS ; Kart2018HowTM ; Liu2019ContextAwareTM has attracted more attention. Lately, some pioneering works gordon2004beyond ; Feng2020ANO ; Qi2020P2BPN have focused on point cloud based 3D object tracking. However, due to the sparsity of 3D point clouds, 3D object tracking on point clouds is still a challenging task.
Few works are dedicated to 3D single object tracking (SOT) with only point clouds. As a pioneer, SC3D giancola2019leveraging is the first 3D Siamese tracker that performs matching between the template and candidate 3D target proposals generated by Kalman filtering gordon2004beyond . Furthermore, a shape completion network is used to enhance shape information of candidate proposals in sparse point clouds, thereby improving the accuracy of matching. However, SC3D cannot perform the end-to-end training, and consumes much time when matching exhaustive candidate proposals. Towards these concerns, Qi et al. Qi2020P2BPN proposed an end-to-end framework termed P2B, which first localizes potential target centers in the search area via Hough voting qi2019deep , and then aggregates vote clusters to generate target proposals. Nonetheless, when facing sparse scenes, P2B may not be able to track the object accurately, or even lose the tracked object. On the one hand, it adopts random sampling to generate initial seed points, which further exacerbates the sparsity of point clouds. On the other hand, it is difficult to generate high-quality target proposals on sparse 3D point clouds. Although SC3D has enhanced shape information of candidate proposals, the low-quality candidate proposals obtained from sparse point clouds still degrade tacking performance.
As shown in Fig. 1, we count the number of points on KITTI’s cars. It can be found that 51% of cars have less than 100 points, and only 7% of cars have more than 2500 points. When facing sparse point clouds, it is difficult to distinguish the target from the background due to the sparsity of point clouds. Therefore, how to improve the tracking performance in sparse scenes should be considered. Our intuition consists of two folds. First, enhancing shape information of target will provide discriminative information to distinguish the target from the background, especially in sparse point clouds. Second, due to the sparsity of the point cloud, it is difficult to regress the target center in 3D space. We hence consider compressing the sparse 3D space into a dense 2D space, and perform center regression in the dense 2D space to improve tracking performance.
In this paper, we propose a novel Siamese voxel-to-BEV (V2B) tracker, which aims to improve the tracking performance of 3D single object tracking, especially in sparse point clouds. We illustrate our framework in Fig. 2. We first feed the template and search area into the Siamese network to extract point features, respectively. Then, we employ the global and local template feature embedding to strengthen the correlation between the template and search area so that the potential target in the search area can be effectively localized. After that, we introduce a shape-aware feature learning module to learn the dense geometric features of the potential target, where the complete and dense point clouds of the target are generated. Thus, the geometric structures of the potential target can be captured better so that the potential target can be effectively distinguished from the background in the search area. Finally, we develop a voxel-to-BEV target localization network to localize the target in the search area. In order to avoid using the low-quality proposals on sparse point clouds for target center prediction, we directly regress the 3D center of the target with the highest response in the dense bird’s eye view (BEV) feature map, where the dense BEV feature map is generated by voxelizing the learned dense geometric features and performing max-pooling along the axis. Thus, with the constructed dense BEV feature map, for sparse point clouds, our method can more accurately localize the target center without any proposal.
In summary, we propose a novel Siamese voxel-to-BEV tracker, which can significantly improve tracking performance, especially in sparse point clouds. We develop a Siamese shape-aware feature learning network that can introduce shape information to enhance the discrimination of the potential target in the search area. We develop a voxel-to-BEV target localization network, which can accurately detect the 3D target’s center in the dense BEV space compared to sparse 3D space. Extensive results show that our method has achieved new state-of-the-art results on the KITTI dataset Geiger2012AreWR , and has a good generalization ability on the nuScenes nuscenes2019 dataset.
Related Work
2D object tracking. Numerous schemes bromley1993signature ; gordon2004beyond ; kristan2016novel ; wu2013online ; bertinetto2016staple have been presented and achieved impressive results in 2D object tracking. Early works are mainly based on correlation filtering. As a pioneer, MOSSE bolme2010visual presents stable correlation filters for visual tracking. After that, correlation-based methods use Circulant matrices henriques2012exploiting , kernelized correlation filters henriques2014high , continuous convolution filters danelljan2016beyond , factorized convolution operators danelljan2017eco to improve tracking performance. In recent years, Siamese-based methods fan2019siamese ; held2016learning have been more popular in the tracking field. In bertinetto2016fully , Bertinetto et al. proposed SiamFC, a pioneering work that combines naive feature correlation with a fully-convolutional Siamese network for object tracking. Subsequently, some improvements zhu2018distractor ; wang2019fast ; xu2020siamfc++ ; zhang2020ocean ; yan2020alpha are made to Siamese trackers, such as combining with a region proposal network Feichtenhofer2017DetectTT ; li2018high ; Yu2020DeformableSA ; Voigtlaender2020SiamRV or an anchor-free FCOS detector Choi2020VisualTB , using a deeper architecture li2019siamrpn++ or two-branch structure he2018twofold , exploiting attention wang2018learning ; Zhou2019SiamManSM or self-attention chen2021transformer , applying triplet loss dong2018triplet . However, these methods are specially designed for 2D object tracking, so they cannot be directly applied to 3D point clouds.
3D single object tracking. Early 3D single object tracking (SOT) methods focus on RGB-D information. As a pioneer, Song et al. song2013tracking first proposed a unified 100 RGB-D video dataset, which opened up a new research direction for RGB-D tracking bibi20163d ; liu2018context ; Kart2019ObjectTB . Based on RGB-D information, 3D SOT methods spinello2010layered ; luber2011people ; Rnz2017CofusionRS usually combines techniques from the 2D tracking with additional depth information. However, RGB-D tracking also relies on RGB information, and it may fail when the RGB information is degraded. Recent efforts luo2018fast ; wang2020pointtracknet begin to use LiDAR point clouds for 3D single object tracking. Among them, SC3D giancola2019leveraging is the first 3D Siamese tracker, but it is not and end-to-end framework. Following it, Re-Track Feng2020ANO is a two-stage framework that re-tracks the lost objects of the coarse stage in the fine stage. Lately, Qi et al. Qi2020P2BPN proposed P2B, which solves the problem that SC3D cannot perform end-to-end training and consumes a lot of time. P2B adopts the tracking-by-detection scheme, which uses VoteNet qi2019deep to generate proposals and selects the proposal with the highest score as the target. Based on P2B, to handle sparse and incomplete target shapes, BAT zheng2021box introduces a box-aware feature module to enhance the correlation learning between template and search area. Nonetheless, when facing very sparse scenarios, VoteNet used in P2B and BAT may be difficult to generate high-quality proposals, resulting in performance degradation.
3D multi-object tracking. Most 3D multi-object tracking (MOT) systems follow the same schemes with the 2D multi-object tracking systems, but the only difference is that 2D detection methods are replaced by 3D detection methods. Most 3D MOT methods wu20213d ; shenoi2020jrmot ; kim2021eagermot usually adopt tracking-by-detection schemes. Specifically, they first use a 3D object detector shi2019pointrcnn ; shi2020points ; shi2020pv to detect numerous objects of each frame, and then exploit the data association between detection results of two frames to match the corresponding objects. To exploit the data association, early works scheidegger2018mono use handcrafted features such as spatial distance. Instead, modern 3D trackers use motion information that can be obtained by 3D Kalman filters patil2019h3d ; chiu2020probabilistic ; weng20203d and learned deep features yin2021center ; zhang2019robust .
Deep learning on point clouds. With the introduction of PointNet qi2017pointnet , 3D deep learning on point clouds has stimulated the interest of researchers. Existing methods can be mainly divided into: point-based qi2017pointnet++ ; pointcnn2018 ; kpconv2019 ; cheng2020cascaded , volumetric-based qi2016volumetric ; liu2019point , graph-based wang2018dynamic ; landrieu2018large ; landrieu2019point ; gacnet2019 ; cheng2021sspc ; hui2021superpoint , and view-based su2015multi ; su20153d ; yang2019learning methods. However, volumetric-based and view-based methods lose fine-grained geometric information due to voxelization and projection, while graph-based methods are not suitable for sparse point clouds since few points cannot provide sufficient local geometric information for constructing a graph. Thus, existing 3D tracking networks giancola2019leveraging ; Feng2020ANO ; Qi2020P2BPN ; zheng2021box are point-based methods.
Method
Our work is specifically designed for 3D single object tracking in sparse point clouds. An overview of our framework is depicted in Fig. 2. We first present the Siamese shape-aware feature learning network to enhance the discrimination of the potential target in search area (Sec. 3.1). We then localize the target by voxel-to-BEV target localization network (Sec. 3.2).
Suppose the size of the template is , and the size of the search area is (generally, ). Before template feature embedding, we first use the Siamese network to extract point features of the template and search area, denoted by and . The Siamese network consists of template branch and detection branch. In order to reduce the inference time of the network, we just use PointNet++ qi2017pointnet++ as the backbone and share parameters. It can be replaced with a powerful network such as KPConv kpconv2019 . We then employ template feature embedding to encode the search area by learning the similarity of the global shape and local geometric structures between the template and search area. The illustration of template feature embedding is shown in the left half of Fig. 3.
Template global feature embedding. We use the multi-layer perceptron (MLP) network to adaptively learn the correlation between the template and search area. The similarity between the template and search area is formulated as:
Template local feature embedding. To characterize the local similarity between the template and search area, we first obtain the similarity map by computing the cosine distance between them. The similarity function is written as:
where indicates the similarity between points and . We then assign each point in the search area with its most similar point in the template, which is written as:
1.2 Shape-Aware Feature Learning
Due to the sparse and incomplete point clouds of the potential target in the search area, we employ shape-aware feature learning to learn dense geometric features of the target, where the dense and complete point clouds of the target can be obtained. It is expected that the learned features from the generated dense point clouds can characterize the geometric structures of the target better.
Dense ground truth processing. To obtain the dense 3D point cloud ground truth, we first crop and center points lying inside the target’s ground truth bounding box in all frames. We then concatenate all cropped and centered points to generate a dense aligned 3D point cloud, denoted by , where is the 3D position, and we fix the number of points to 2048 by randomly discarding and duplicating points.
By minimizing the CD loss, we can learn dense geometric features of the potential target in the search area by generating a dense and complete point cloud of the target. Note that shape information encoding is only performed during training and will be discarded during testing. Thus, it does not increase the inference time of object tracking in the test scheme.
Although SC3D giancola2019leveraging also uses shape completion to encode shape information of the target, the template completion model in SC3D cannot recover complex geometric structures of the potential target well due to limited templates and large variations of the potential target in the search area. However, our method is a complex point cloud generation method that learns the target completion model from the samples of search areas with the gate mechanism to enhance the feature of the potential target and suppress the background in the search area. In addition, the template completion model in SC3D only employs PointNet qi2017pointnet to extract point features of sparse point clouds, while our target completion model constructs a global-local branch to extract global shape features and local geometric features of sparse point clouds.
2 Voxel-to-BEV Target Localization Network
In order to avoid using the low-quality proposals on sparse point clouds for target center prediction, we develop a simple yet effective target center localization network without any proposal to improve the localization precision in sparse point clouds.
In order to improve the localization precision in sparse point clouds, we utilize the voxelization and max-pooling operation to convert the learned discriminative features of sparse 3D points into the dense bird’s eye view (BEV) feature map for the target localization, as shown in the right half of Fig. 2. We first convert the point features of the search area into a volumetric representation by averaging the 3D coordinates and features of the points in the same voxel bin. Then, we apply a stack of 3D convolutions on the voxelized feature map to aggregate the feature of the potential target in the search area, where the voxels lying on the target can be encoded with rich target information. However, in the sparse volume space, due to the large number of empty voxels, the differences between the responses in the voxelized feature map might not be remarkable. Thus, the highest response in the feature map is difficult to distinguish from the low responses, leading to the inaccurate regression of the 3D center of the target, including the -axis center. By performing max-pooling on the voxelized feature map along the -axis, we can obtain the dense BEV feature map, where the low responses in the voxelized feature map can be suppressed. Thus, compared to the voxelized feature map, we can more accurately localize the 2D center of the target with the highest response in the dense BEV feature map. The response of the 2D center (, max-pooling feature along the -axis) in the BEV feature map actually contains the geometric structure information of the potential target while the responses of other points in the BEV feature map do not. In addition, we apply a stack of 2D convolutions on the dense BEV feature map to aggregate the feature so that the potential target can obtain sufficient local information in the BEV feature map. Thus, with the constructed dense BEV feature map, for sparse point clouds, our method can more accurately localize the target center without any proposal.
2.2 Target Localization in BEV
Inspired by Ge2020AFDetAF , we develop a simple yet powerful network to detect the 2D center and the -axis center based on the obtained dense BEV feature map. As shown in the right half of Fig. 2, it consists of three heads: 2D-center head, offset & rotation head, and -axis head. The 2D-center head aims to localize 2D center of target on the - plane, and -axis head regresses the target center of the -axis. Since the 2D center of 2D grid is discrete, we also regress the offset between it and the continuous center. Thus, we use a offset & rotation head to regress offset plus additional rotation.
The final loss of our network is as follows: , where , , and are the hyperparameter for shape generation, 2D center and offset regression, and -axis position regression, respectively. In the experiment, we set , , and .
Experiments
Datasets. For 3D single object tracking, we use KITTI Geiger2012AreWR and nuScenes nuscenes2019 datasets for training and evaluation. Since the ground truth of the test set of KITTI dataset cannot be obtained, we follow giancola2019leveraging ; Qi2020P2BPN and use the training set to train and evaluate our method. It contains 21 video sequences and 8 types of objects. We use scenes 0-16 for training, scenes 17-18 for validation, and scenes 19-20 for testing. For nuScenes dataset, we use its validation set to evaluate the generalization ability of our method. Note that the nuScenes dataset only labels key frames, so we report the performance evaluated on the key frames.
Evaluation metrics. For 3D single object tracking, we use the Success and Precision defined in the one pass evaluation (OPE) Kristan2016ANP to evaluate the tracking performance of different methods. Success measures the IOU between the predicted and ground truth bounding boxes, while Precision measures the error AUC of the distance between the centers of two bounding boxes.
Implementation details. Following Qi2020P2BPN , we set the number of points and for the template and search area by randomly discarding and duplicating points. For the backbone network, we use a slightly modified PointNet++ qi2017pointnet++ , which consists of three set-abstraction (SA) layers (with query radius of 0.3, 0.5, and 0.7) and three feature propagation (FP) layers. For each SA layer passed, the points will be randomly downsampled by half. For the shape generation network, we generate 2048 points. The global branch is the max pooling combined with two fully connected layers, while the local branch only uses one EdgeConv layer. We use a two layer MLP network to generate 3D coordinates. For 3D center detection, the voxel size is set to 0.3 meters in volumetric space. We stack four 3D convolutions (with stride of 2, 1, 2, 1 along the -axis) and four 2D convolutions (with stride of 2, 1, 1, 2) combined with the skip connections for feature aggregation, respectively. For all experiments, we use Adam Kingma2015AdamAM optimizer with learning rate 0.001 for training, and the learning rate decays by 0.2 every 6 epochs. It takes about 20 epochs to train our model to convergence.
Training and testing. For training, we combine the points inside the first ground truth bounding box (GTBB) and the points inside the previous GTBB plus the random offset as the template of the current frame. To generate the search area, we enlarge the current GTBB by 2 meters and plus the random offset. For testing, we fuse the points inside the first GTBB and the previous result’s point cloud (if exists) as the template. Besides, we first enlarge the previous result by 2 meters in current frame, and then collect the points lying in it to generate the search area.
2 Results
Quantitative results. We compare our method with current state-of-the-art methods, including SC3D giancola2019leveraging , P2B Qi2020P2BPN , and BAT zheng2021box . The quantitative results are listed in Tab. 1. For the KITTI Geiger2012AreWR dataset, we follow giancola2019leveraging ; Qi2020P2BPN and report the performance of four categories, including car, pedestrian, van, and cyclist, and their average results. As one can see from the table, our method is significantly better than other methods on the mean results of four categories. For the car category, our method can even improve the Success from 60.5% (BAT) to 70.5% (V2B). However, for tracking-by-detection methods that rely on large amounts of training samples, it is difficult to effectively track cyclists with few training samples. Thus, P2B, BAT, and V2B are worse than SC3D in the cyclist category. However, SC3D uses exhaustive search to generate numerous candidate proposals for template matching, so it performs well with few training samples. For the nuScenes nuscenes2019 dataset, we directly apply the models, trained on the corresponding categories of the KITTI dataset, to evaluate performance on the nuScenes dataset. Specifically, the corresponding categories between KITTI and nuScenes datasets are CarCar, PedestrianPedestrian, VanTruck, and CyclistBicycle, respectively. It can be seen that our V2B can still achieve better performance on the mean results of all four categories. Due to few training samples on the van category, P2B, BAT, and V2B cannot obtain good generalization ability compared to SC3D in the truck category. The quantitative results on the nuScenes dataset further demonstrate that our V2B has a good generalization ability to adapt to different datasets.
Quantitative results on sparse scenes. To verify the effectiveness of our method for object tracking in sparse scenes, we count the performance of SC3D, P2B, BAT, and our V2B in sparse scenes of the KITTI dataset. Specifically, we filter out sparse scenes for evaluation according to the number of points lying in the target bounding boxes in the test set. Specifically, the conditions for the sparse scenes are: (car), (pedestrian), (van), and (cyclist), respectively. For the four categories, the number of selected frames are 3293 (car), 1654 (pedestrian), 734 (van), and 59 (cyclist), respectively. In Tab. 2, we report the results of Success and Precision. As one can see from the table, our V2B achieves the best performance on the mean results of all four categories. The results of the cyclist category are worse due to few training samples. Note that when switching from sparse frames (Tab. 2) to all types of frames (Tab. 1), SC3D and P2B suffer from a performance drop on the mean results of four categories. The worse tracking performance of SC3D and P2B on large amounts of sparse frames leads to the inaccurate template updates on the consecutive dense frames. Thus, SC3D and P2B cannot obtain better tracking performance on the dense frames. Although SC3D uses template shape completion, due to limited template samples and large variations of the potential target in the search area, it cannot accurately recover the complex geometric structures of the target in the sparse frames, which poses challenges on localizing the potential target with sparse points. On the contrary, our V2B employs the proposed shape-aware feature learning module to generate dense and complete point clouds of the potential target for the target shape completion, leading to more accurate localization of the target in the sparse frames. Compared with BAT, our V2B achieves the performance gain of 2% on the mean results of all four categories from sparse frames to all types of frames. Thus, the comparison results can demonstrate that our V2B can effectively improve the performance of single object tracking in sparse point clouds.
Visualization results. As shown in Fig. 4, we plot the visualization results of P2B and our V2B on the car category. Specifically, we plot a couple sparse and dense scenarios on the car category of the KITTI dataset. It can be clearly seen from the figure that compared with P2B, our V2B can track the targets more accurately in both sparse and dense scenes. Especially in sparse scenes, compared with P2B, our V2B can track the targets effectively. The visualization results can demonstrate the effectiveness of our V2B for sparse point clouds.
3 Ablation Study
Template feature embedding. We study the impact of template feature embedding on tracking performance. As shown in Tab. 3, we report the results of the car category in the KITTI dataset. It can be seen that without using the template feature embedding (dubbed “w/o template feature”), the performance will be greatly reduced from 70.5 / 81.3 to 63.9 / 73.9 by a large margin. In addition, only using the local branch or global branch cannot achieve the best performance. Since template feature embedding builds the relationship between the template and search area, it will contribute to identify the potential target from the background in the search area. Therefore, when the template feature embedding is absent, the performance will be greatly reduced, which further demonstrates the effectiveness of the proposed template feature embedding for improving tracking performance.
Shape-aware feature learning. For sparse point clouds, we further introduce shape-aware feature learning to enhance the ability to distinguish the potential target from the background in the search area. As shown in Tab. 3, we conduct experiments to demonstrate the effectiveness of the shape information. It can be seen from the table that without using shape-aware feature learning module (dubbed “w/o shape information”), the performance will reduce from 70.5 / 81.3 to 67.6 / 78.2. In addition, only using the local geometric branch or global shape branch cannot achieve the best performance. Since the shape generation network can capture 3D shape information of the object to learn the discriminative features of potential target so that it can be identified from the search area.
Voxel-to-BEV target localization. Different from SC3D giancola2019leveraging and P2B Qi2020P2BPN , our V2B adopts another route to localize potential target in object tracking. SC3D performs matching between the template and the exhaustive candidate 3D proposals to select the most similar proposal as the target. P2B and BAT use VoteNet qi2019deep to generate 3D target proposals, and select the proposal with the highest score as the target. However, when facing sparse point clouds, it is hard to generate high-quality proposals, so these methods may not be able to track the object effectively. Our V2B is an anchor-free method that does not require generating numerous 3D proposals. Therefore, our method can overcome the above concern. In order to prove this, we use VoteNet instead of voxel-to-BEV target localization to conduct experiments in the KITTI dataset. In Tab. 4, we report the results of different detection methods in the sparse scenarios. Likewise, we filter out sparse scenes in the test set for evaluation according to the number of points (refer to the setting of Tab. 2). It can be found that the results of VoteNet are lower than that of voxel-to-BEV target localization, which further demonstrates the effectiveness of our method in sparse point clouds.
Different voxel sizes. We compress the voxelized point cloud into a BEV feature map for subsequent target center detection. Since the scope of object tracking is a large area, the size of voxel will affect the size of BEV feature map, thereby affecting the tracking performance. We hence study the impact of different voxel sizes on the tracking performance. Specifically, we consider four sizes, including 0.1, 0.2, 0.3, and 0.4 meters. The Success/Precision results of the four sizes are 52.5 / 62.2 (0.1m), 68.1 / 79.4 (0.2m), 70.5 / 81.3 (0.3m), and 69.1 / 80.0 (0.4m), respectively. When the voxel size is set to 0.3 meters, we achieve the best performance. A larger voxel size will increase the sparsity and cause the loss of the target’s details. A smaller voxel size will increase the size of the BEV feature map, thereby increasing the difficulty of detecting the center.
Template generation scheme. Following giancola2019leveraging ; Qi2020P2BPN ; zheng2021box , we study the impact of different template generation schemes on tracking performance. As shown in Tab. 5, we report the results of four schemes on the car category in the KITTI dataset. It can be seen from the table that our V2B outperforms SC3D, P2B, and BAT in all schemes by a large margin. Compared with these methods, our V2B can yield stable results on four template generation schemes, which further demonstrates that our method can consistently generate accurate tracking results in all types of frames.
Conclusion
In this paper, we proposed a Siamese voxel-to-BEV (V2B) tracker for 3D single object tracking on sparse point clouds. In order to learn the dense geometric features of the potential target in the search area, we developed a Siamese shape-aware feature learning network that utilizes the target completion model to generate the dense and complete targets. In order to avoid using the low-quality proposals on sparse point clouds for target center prediction, we developed a simple yet effective voxel-to-BEV target localization network that can directly regress the center of the potential target from the dense BEV feature map without any proposal. Rich experiments on the KITTI and nuScenes datasets have demonstrated the effectiveness of our method on sparse point clouds.
Acknowledgments
This work was supported by the National Science Fund of China (Grant Nos. U1713208, 61876084).
References
Appendix A Overview
This supplementary material provides more details on network architecture, implementation, and experiments in the main paper to validate and analyze our proposed method. We will release the code after the paper is published.
In Sec. B, we provide specific network architecture and more details about the target center parameterization for the voxel-to-BEV target localization network. In Sec. C, we provide more details about template and search area generation in training and testing. In Sec. D, we show more experimental results including quantitative results, visualization, and ablation study.
Appendix B Network Architecture
In this section, we provide specific network architecture used for the voxel-to-BEV target localization network. As shown in Fig. 5, we illustrate the specific network structure. Specifically, we first use the 3D network to aggregate features in the volumetric space. Then, we present the 2D network to aggregate features in the BEV space. After that, we introduce the center network for target localization. Finally, we provide more details on the target center parameterization.
3D network. We use a stack of 3D convolutions to the volumetric space to aggregate the features so that target’s voxel can obtain rich target information. As shown on the left side of Fig. 5, we depict the specific structure of the network. Specifically, it consists of four 3D convolutions with the filter sizes 333. In order to reduce memory consumption, we set the stride of four 3D convolutions to 112, 111, 112, and 111, respectively. Note that the stride along the -axis is 2, so the feature size in the - plane will not change. Finally, the 3D network outputs a new feature map with a size of .
2D network. After projecting the voxelized point cloud into the bird’s eye view (BEV) space through the max pooling function, we obtain a new BEV feature map with a size of . We use a stack of 2D convolutions to construct a shallow encoder-decoder neural network to aggregate the features so that target’s pixel can obtain rich target information. We show the specific network structure in the middle part of Fig. 5. Specifically, we first adopt three 2D convolutions (of filter sizes 33) and one transposed convolution (of filter size 22). And the strides of four 2D convolutions are 22, 11, 11, and 22, respectively. Note that the last 2D convolution is the transposed convolution. We also use a concatenated skip connection to fuse low-level and high-level features. Finally, after using a 2D convolution, we obtain a new feature map with a size of .
Center network. Since we have known the size (length, width and height) of the object in the template, we only need to regress the target center, offset, and rotation. As shown on the right side of Fig. 5, we illustrate the specific network structure. Specifically, we use three heads to regress the 2D center, offset & rotation, and -axis location, respectively. For each head, we use two convolutions with filter sizes of 33 and 11. The output sizes are (2D-center head), (offset & rotation head) and (-axis head), respectively. Note that for the offset & rotation head, 3-dim means the 2-dim coordinate offset and 1-dim rotation.
Target center parameterization. To obtain the target center in the BEV space, we perform target center parameterization. Assuming the voxel size and the range of the search area in the - plane, we can obtain the resolution of the BEV feature map by:
where is the floor operation. Given a 3D center of the target ground truth, we can compute the 2D target center in the - plane by:
Appendix C Implementation Details
Template and search area in training. Following , we adopt the same strategy to generate the template and search area during training. For the current frame, the template is generated by fusing two frames, , the first frame and the previous frame (if exists). As shown in the left half of Fig. 7, we combine the points inside the first frame’s ground truth bounding box (GTBB) and the points inside the previous frames’ GTBB plus the random offset as the template of the current frame. During training, we sample 512 points from the template by discarding and duplicating the points. For the search area, we enlarge the current frame’s GTBB by 2 meters and plus the random offset. Likewise, we sample 1024 points from the search area by discarding and duplicating the points.
Template and search area in testing. In the right half of Fig. 7, we show the specific process of generating the template and search area. For testing, we fuse the points inside the first frame’s ground truth bounding box (GTBB) and the previous result’s point cloud (if exists) as the template of the current frame. For the search area, we first enlarge the previous predicted bounding box by 2 meters in current frame, and then collect the points lying in it to construct the search area. Note that unlike the training phase, we do not apply random offset augmentation to the previous result’s point cloud. Besides, we sample 512 points in the template and 1024 points in the search area during testing.
Template generation schemes. In the experiment, we also report the tracking performance of four different template generation schemes. They are dubbed as “The First GT”, “Previous result”, “The First GT & Previous result” and “All previous results”, respectively. “First GT” means that we only use the first frame as the template and “Previous result” means that we only use the tracked result’s point cloud in the previous frame as the template. Thus, “The First GT & Previous result” is a combination of “First GT” and ‘Previous result”. Besides, “All previous results” represents that we align all of the previous tracked results’ point clouds as a template.
Appendix D Experiments
Different factors of the shape loss. As shown in Fig. 8, we study the impact of different factors of the shape loss on the tracking performance. It can be seen that when the factor is set to , we achieve the best performance. Furthermore, according to the figure, different factors are insensitive to the performance.
Quantitative results. To better validate and analyze the proposed method, we also show the performance of different objects in different point intervals. Specifically, we divide the interval according to the number of points lying in the ground truth bounding boxes in the test set. For large-size categories such as car and van, we set four intervals, including [0, 150), [150, 1000), [1000, 2500), and [2500, +). For small-size categories such as pedestrian and cyclist, we set four intervals, including [0, 100), [100, 500), [500, 1000), and [1000, +). As shown in Tab. 6, we report the Success and Precision of SC3D , P2B , and our V2B. It can be seen that our method is superior to other methods in terms of the mean of the four categories. In addition, our method is lower than SC3D on the cyclist category. Since there are few training samples on the cyclist category, it will affect the performance of P2B and our method. However, SC3D enumerates exhaustive candidate proposals, so it can achieve higher performance.
Visualization. As shown in Fig. 9, we provide more visualization results of our method for four categories, including car, pedestrian, van, and cyclist. It can be seen that our method can accurately localize targets in both dense and sparse scenes. As shown in Fig. 10, we also compare the tracking results of SC3D , P2B , and our method.
Failure cases. As shown in Fig. 11, we provide the visualization results of the failure cases. It can be found that our method will fail in extremely sparse scenes. If the tracking fails in the previous frame, a poor-quality search area will be generated for the current frame, which will further affect the tracking results of subsequent frames.