Behind the Curtain: Learning Occluded Shapes for 3D Object Detection

Qiangeng Xu, Yiqi Zhong, Ulrich Neumann

Introduction

With high-fidelity, the point clouds acquired by LiDAR sensors significantly improved autonomous agents’s ability to understand 3D scenes. LiDAR-based models achieved state-of-the-art performance on 3D object classification (Xu et al. 2020), visual odometry (Pan et al. 2021), and 3D object detection (Shi et al. 2020). Despite being widely used in these 3D applications, LiDAR frames are technically 2.5D. After hitting the first object, a laser beam will return and leave the shapes behind the occluder missing from the point cloud.

To locate a severely occluded object (e.g., the car in Figure LABEL:fig:teaser(b)), a detector has to recognize the underlying object shapes even when most of its parts are missing. Since shape miss inevitably affects object perception, it is important to answer two questions:

What are the causes of shape miss in point clouds?

What is the impact of shape miss on 3D object detection?

To answer the first question, we study the objects in KITTI (Geiger et al. 2013) and discover three causes of shape miss.

External-occlusion. As visualized in Figure LABEL:fig:teaser(c), occluders block the laser beams from reaching the red frustums behind them. In this situation, the external-occlusion is formed, which causes the shape miss located at the red voxels.

Signal miss. As Figure LABEL:fig:teaser(c) illustrates, certain materials and reflection angles prevent laser beams from returning to the sensor after hitting some regions of the car (blue voxels). After projected to range view, the affected blue frustums in Figure LABEL:fig:teaser(c) appear as the empty pixels in Figure LABEL:fig:teaser(a).

Self-occlusion. LiDAR data is 2.5D by nature. As shown in Figure LABEL:fig:teaser(d), for a same object, its parts on the far side (the green voxels) are occluded by the parts on the near side. The shape miss resulting from self-occlusion inevitably happens to every object in LiDAR scans.

2 Impact of Shape Miss

To analyze the impact of shape miss on 3D object detection, we evaluate the car detection results of the scenarios where we recover certain types of shape miss on each object by borrowing points from similar objects (see the details of finding similar objects and filling points in Sec. 3.1).

In each scenario, after resolving certain shape miss in both the train and val split of KITTI (Geiger et al. 2013), we train and evaluate a popular detector PV-RCNN (Shi et al. 2020). The four scenarios are:

NR: Using the original data without shape miss recovery.

EO: Recovering the shape miss caused by external-occlusion (adding the red points in Figure 1(a)).

EO+SM: Recovering the shape miss caused by external-occlusion and signal miss (adding the red and blue points in Figure 1(a)).

EO+SM+SO: Recovering all the shape miss (adding the red, blue and green points in Figure 1(a)).

We report detection results on cars with three occlusion levels (level labels are provided by the dataset). As shown in Figure 1(b), without recovery (NR), it is more difficult to detect objects with higher occlusion levels. Recovering shapes miss will reduce the performance gaps between objects with different levels of occlusion . If all shape miss are resolved (EO+SM+SO), the performance gaps are eliminated and almost all objects can be effectively detected (APs >> 99%99\%).

3 The Proposed Method

The above experiment manually resolves the shape miss by filling points into the labeled bounding boxes and significantly improve the detection results. However, during test time, how do we resolve shape miss without knowing bounding box labels?

In this paper, we propose Behind the Curtain Detector (BtcDet). To the best of our knowledge, BtcDet is the first 3D object detector that targets the object shapes affected by occlusion. With the knowledge of shape priors, BtcDet estimates the occupancy of complete object shapes in the regions affected by occlusion and signal miss. After being integrated into the detection pipeline, the occupancy estimation benefits both region proposal generation and proposal refinement. Eventually, BtcDet surpasses all of the state-of-the-art methods published to date by remarkable margins.

Related Work

LiDAR-based 3D object detectors. Voxel-based methods divide point clouds by voxel grids to extract features (Zhou and Tuzel 2018). Some of them also use sparse convolution to improve model efficiency, e.g., SECOND(Yan et al. 2018). Point-based methods such as PointRCNN (Shi et al. 2019a) generate proposals directly from points. STD (Yang et al. 2019) applies sparse to dense refinement and VoteNet (Qi et al. 2019a) votes the proposal centers from point clusters. These models are supervised on the ground truth bounding boxes without explicit consideration for the object shapes.

Learning shapes for 3D object detection. Bounding box prediction requires models to understand object shapes. Some detectors learn the shape related statistics as an auxiliary task. PartA2 (Shi et al. 2020) learns object part locations. SA-SSD and AssociateDet (He et al. 2020; Du et al. 2020) use auxiliary networks to preserve structural information. Studies (Li et al. 2021; Yan et al. 2020; Najibi et al. 2020; Xu et al. 2021) such as SPG conduct point cloud completion to improve object detection. These models are shape-aware but overlook the impact of occlusion on object shapes.

Occlusion handling in computer vision. The negative impact of occlusion on various computer vision tasks, including tracking (Liu et al. 2018), image-based pedestrian detection (Zhang et al. 2018), image-based car detection (Reddy et al. 2019) and semantic part detection (Saleh et al. 2021), is acknowledged. Efforts addressing occlusion include the amodal instance segmentation (Follmann et al. 2019), the Multi-Level Coding that predicts the presence of occlusion (Qi et al. 2019b). These studies, although focus on 2D images, demonstrate the benefits of modeling occlusion to solving visual tasks. Point cloud visibility is addressed in (Hu et al. 2020) and is used in multi-frame detection and data augmentation. This method, however, does not learn and explore the visibility’s influence on object shapes. Our proposed BtcDet is the first 3D object detector that learns occluded shapes in point cloud data. We compare (Hu et al. 2020)’s approach with ours in Sec. 4.3.

Behind the Curtain Detector

Let Θ\Theta denote the parameters of a detector, {p1,p2,...,pN}\{p_{1},p_{2},...,p_{N}\} denote the LiDAR point cloud, X,D,Sob,Soc\mathcal{X},\mathcal{D},{\mathcal{S}_{ob}},{\mathcal{S}_{oc}} denote the estimated box center, the box dimension, the observed objects shapes and the occluded object shapes, respectively. Most LiDAR-based 3D object detectors (Yi et al. 2020; Chen et al. 2020; Shi and Rajkumar 2020) only supervise the bounding box prediction. These models have

while structure-aware models (Shi et al. 2020; He et al. 2020; Du et al. 2020) also supervise Sob{\mathcal{S}_{ob}}’s statistics so that

None of the above studies explicitly model the complete object shapes S=Sob∪Soc\mathcal{S}=\mathcal{S}_{ob}\cup\mathcal{S}_{oc}, while the experiments in Sec. 1.2 show the improvements if S\mathcal{S} is obtained. BtcDet estimates S\mathcal{S} by predicting the shape occupancy OS\mathcal{O_{S}} for regions of interest. After that, BtcDet conducts object detection conditioned on the estimated probability of occupancy P(OS)\mathcal{P}(\mathcal{O_{S}}). The optimization objectives can be described as follows:

Model overview. As illustrated in Figure 2, BtcDet first identifies the regions of occlusion ROC\mathcal{R_{OC}} and signal miss RSM\mathcal{R_{SM}}, and then, let a shape occupancy network Ω\Omega estimate the probability of object shape occupancy P(OS)\mathcal{P}(\mathcal{O_{S}}). The training process is described in Sec. 3.1.

Next, BtcDet extracts the point cloud 3D features by a backbone network Ψ\Psi. The features are sent to a Region Proposal Network (RPN) to generate 3D proposals. To leverage the occupancy estimation, the sparse tensor P(OS)\mathcal{P}(\mathcal{O_{S}}) is concatenated with the feature maps of Ψ\Psi. (See Sec. 3.2.)

Finally, BtcDet applies the proposal refinement. The local geometric features fgeof_{geo} are composed of P(OS)\mathcal{P}(\mathcal{O_{S}}) and the multi-scale features from Ψ\Psi. For each region proposal, we construct local grids covering the proposal box. BtcDet pools the local geometric features fgeof_{geo} onto the local grids, aggregates the grid features, and generates the final bounding box predictions. (See Sec. 3.3.)

Approximate the complete object shapes for ground truth labels. Occlusion and signal miss preclude the knowledge of the complete object shapes SS. However, we can assemble the approximated complete shapes S‾\overline{\mathcal{S}}, based on two assumptions:

Most foreground objects resemble a limited number of shape prototypes, e.g., pedestrians share a few body types.

Foreground objects, especially vehicles and cyclists, are roughly symmetric.

We use the labeled bounding boxes to query points belonging to the objects. For cars and cyclists, we mirror the object points against the middle section plane of the bounding box.

A heuristic H(A,B)\mathcal{H}(A,B) is created to evaluate if a source object BB covers most parts of a target object AA and provides points that can fill AA’s shape miss. To approximate AA’s complete shape, we select the top 3 source objects B1,B2,B3B_{1},B_{2},B_{3} with the best scores. The final approximation S‾\overline{\mathcal{S}} consists of AA’s original points and the points of B1,B2,B3B_{1},B_{2},B_{3} that fill AA’s shape miss. The target objects are the occluded object in the current training frame, while the source objects are other objects of the same class in the detection training set. Both can be extracted by the ground truth bounding boxes. Please find details of H(A,B)\mathcal{H}(A,B) in Appendix B and more visualization of assembling S‾\overline{\mathcal{S}} in Appendix G.

Identify ROC∪RSM\mathcal{R_{OC}}\cup\mathcal{R_{SM}} in the spherical coordinate system. According to our analysis in Sec. 1.1, “shape miss” only exists in the occluded regions ROC\mathcal{R_{OC}} and the regions with signal miss RSM\mathcal{R_{SM}} (see Figure LABEL:fig:teaser(c) and (d)). Therefore, we need to identify ROC∪RSM\mathcal{R_{OC}}\cup\mathcal{R_{SM}} before learning to estimate shapes.

In real-world scenarios, there exists at most one point in the tetrahedron frustum of a range image pixel. When the laser is stopped at a point, the entire frustum behind the point is occluded. We propose to voxelize the point cloud using an evenly spaced spherical grid so that the occluded regions can be accurately formed by the spherical voxels behind any LiDAR point. As shown in Figure 3(a), each point (x,y,zx,y,z) is transformed to the spherical coordinate system as (r,ϕ,θr,\phi,\theta):

ROC\mathcal{R_{OC}} includes nonempty spherical voxels and the empty voxels behind these voxels. In Figure LABEL:fig:teaser(a), the dashed lines mark the potential areas of signal miss. In range view, we can find pixels on the borders between the areas having LiDAR signals and the areas of no signal. RSM\mathcal{R_{SM}} is formed by the spherical voxels that project to these pixels.

Create training targets. In ROC∪RSM\mathcal{R_{OC}}\cup\mathcal{R_{SM}} , we predict the probability P(OS)\mathcal{P}(\mathcal{O_{S}}) for voxels if they contain points of S‾\overline{\mathcal{S}}. As illustrated in 3(b), S‾\overline{\mathcal{S}} are placed at the locations of the corresponding objects. We set OS‾\mathcal{O_{\overline{S}}} =1=1 for the spherical voxels that contain S‾\overline{\mathcal{S}}, and OS‾\mathcal{O_{\overline{S}}} =0=0 for the others. OS‾\mathcal{O_{\overline{S}}} is used as the ground truth label to approximate the occupancy OS\mathcal{O_{S}} of the complete object shape. Estimating occupancy has two advantages over generating points:

S‾\overline{\mathcal{S}} is assembled by multiple objects. The shape details approximated by the borrowed points are inaccurate and the point density of different objects is inconsistent. The occupancy OS‾\mathcal{O_{\overline{S}}} avoids these issues after rasterization.

The plausibility issue of point generation can be avoided.

Estimate the shape occupancy. In ROC∪RSM\mathcal{R_{OC}}\cup\mathcal{R_{SM}}, we encode each nonempty spherical voxel with the average properties of the points inside (x,y,z,feats), then, send them to a shape occupancy network Ω\Omega. The network consists of two down-sampling sparse-conv layers and two up-sampling inverse-convolution layers. Each layer also includes several sub-manifold sparse-convs (Graham and van der Maaten 2017) (see Appendix D). The spherical sparse 3D convolutions are similar to the ones in the Cartesian coordinate, except that the voxels are indexed along (r,ϕ,θr,\phi,\theta). The output P(OS)\mathcal{P}(\mathcal{O_{S}}) is supervised by the sigmoid cross-entropy Focal Loss (Lin et al. 2017):

Since S‾\overline{\mathcal{S}} borrows points from other objects in the shape miss regions, we assign them a weighting factor δ\delta, where δ<1\delta<1.

2 Shape Occupancy Probability Integration

Trained with the customized supervision, Ω\Omega learns the shape priors of partially observed objects and generates P(OS)\mathcal{P}(\mathcal{O_{S}}). To benefit detection, P(OS)\mathcal{P}(\mathcal{O_{S}}) is transformed from the spherical coordinate to the Cartesian coordinate and fused with Ψ\Psi, a sparse 3D convolutional network that extracts detection features in the Cartesian coordinate..

For example, a spherical voxel has a center (r,ϕ,θr,\phi,\theta) which is transformed as x=rcosθcosϕ, y=rcosθsinϕ, z=rsinθx=rcos\theta cos\phi,\ y=rcos\theta sin\phi,\ z=rsin\theta. Assume x,y,zx,y,z is inside a Cartesian voxel vi,j,kv^{i,j,k}. Since several spherical voxels can be mapped to vi,j,kv^{i,j,k}, vi,j,kv^{i,j,k} takes the max value of these voxels SV(vi,j,k)SV(v^{i,j,k}):

The occupancy probability of these Cartesian voxels forms a sparse tensor map P(OS)⊥={P(OS)v}\mathcal{P}(\mathcal{O_{S}})_{\perp}=\{\mathcal{P}(\mathcal{O_{S}})_{v}\}, which is, then, down-sampled by max-poolings into multiple scales and concatenated with Ψ\Psi’s intermediate feature maps:

where fΨiinf_{\Psi_{i}}^{in}, fΨi−1outf_{\Psi_{i-1}}^{out} and maxpool×2 i−1(⋅)maxpool^{\ i-1}_{\times 2}(\cdot) denote the input features of Ψ\Psi’s iith layer, the output features of Ψ\Psi’s i−1i-1th layer, and applying stride-22 maxpooling i−1i-1 times, respectively.

The Region Proposal Network (RPN) takes the output features of Ψ\Psi and generates 3D proposals. Each proposal includes (xp,yp,zp),(lp,wp,hp),θp,pp(x_{p},y_{p},z_{p}),(l_{p},w_{p},h_{p}),\theta_{p},p_{p}, namely, center location, proposal box size, heading and proposal confidence.

3 Occlusion-Aware Proposal Refinement

Local geometry features. BtcDet’s refinement module further exploits the benefit of the shape occupancy. To obtain accurate final bounding boxes, BtcDet needs to look at the local geometries around the proposals. Therefore, we construct a local feature map fgeof_{geo} by fusing multiple levels of Ψ\Psi’s features. In addition, we also fuse P(OS)⊥\mathcal{P}(\mathcal{O_{S}})_{\perp} into fgeof_{geo} to bring awareness to the shape miss in the local regions. P(OS)⊥\mathcal{P}(\mathcal{O_{S}})_{\perp} provides two benefits for proposal refinement:

P(OS)⊥\mathcal{P}(\mathcal{O_{S}})_{\perp} only has values in ROC∪RSM\mathcal{R_{OC}}\cup\mathcal{R_{SM}} so that it can help the box regression avoid the regions outside ROC∪RSM\mathcal{R_{OC}}\cup\mathcal{R_{SM}}, e.g., the regions with cross marks in Figure 2.

The estimated occupancy indicates the existence of unobserved object shapes, especially for empty regions with high P(OS)\mathcal{P}(\mathcal{O_{S}}) , e.g., some orange regions in Figure 2.

fgeof_{geo} is a sparse 3D tensor map with spatial resolution of 400×352×5400\times 352\times 5. The process for producing fgeof_{geo} is described in Appendix D.

RoI pooling. On each proposal, we construct local grids which have the same heading of the proposal. To expand the receptive field, we set a size factor μ\mu so that:

The grid has a dimension of 12×4×212\times 4\times 2. We pool the nearby features fgeof_{geo} onto the nearby grids through trilinear-interpolation (see Figure 2) and aggregates them by sparse 3D convolutions. After that, the refinement module predicts an IoU-related class confidence score and the residues between the 3D proposal boxes and the ground truth bounding boxes, following (Yan et al. 2018; Shi et al. 2020).

4 Total Loss

The RPN loss Lrpn\mathcal{L}_{rpn} and the proposal refinement loss Lpr\mathcal{L}_{pr} follow the most popular design among detectors (Shi et al. 2020; Yan et al. 2018). The total loss is:

More details of the losses and the network architectures can be found in Appendix C and D.

Experiments

In this section, we describe the implementation details of BtcDet and compare BtcDet with state-of-the-art detectors on two datasets: the KITTI Dataset (Geiger et al. 2013) and the Waymo Open Dataset (Sun et al. 2019). We also conduct ablation studies to demonstrate the effectiveness of the shape occupancy and the feature integration strategies. More detection results can be found in the Appendix F. The quantitative and qualitative evaluations of the occupancy estimation can be found in the Appendix E and H.

Datasets. The KITTI Dataset includes 7481 LiDAR frames for training and 7518 LiDAR frames for testing. We follow (Chen et al. 2017) to divide the training data into a train split of 3712 frames and a val split of 3769 frames. The Waymo Open Dataset (WOD) consists of 798 segments of 158361 LiDAR frames for training and 202 segments of 40077 LiDAR frames for validation. The KITTI Dataset only provides LiDAR point clouds in 3D, while the WOD also provides LiDAR range images.

Implementation and training details. BtcDet transforms the point locations (x,y,zx,y,z) to (r,ϕ,θr,\phi,\theta) for the KITTI Dataset, while directly extracting (r,ϕ,θr,\phi,\theta) from the range images for the WOD. On the KITTI Dataset, we use a spherical voxel size of (0.32m,0.52∘,0.42∘0.32m,0.52^{\circ},0.42^{\circ}) within the range [2.24m,70.72m2.24m,70.72m] for rr, [−40.69∘,40.69∘-40.69^{\circ},40.69^{\circ}] for ϕ\phi and [−16.60∘,4.00∘-16.60^{\circ},4.00^{\circ}] for θ\theta. On the WOD, we use a spherical voxel size of (0.32m,0.81∘,0.31∘0.32m,0.81^{\circ},0.31^{\circ}) within the range [2.94m,74.00m2.94m,74.00m] for rr, [−180∘,180∘-180^{\circ},180^{\circ}] for ϕ\phi and [−33.80∘,6.00∘-33.80^{\circ},6.00^{\circ}] for θ\theta. Determined by grid search, we set γ=2\gamma=2 in Eq.6, δ=0.2\delta=0.2 in Eq.7 and μ=1.05\mu=1.05 in Eq.10.

In all of our experiments, we train our models with a batch size of 8 on 4 GTX 1080 Ti GPUs. On the KITTI Dataset, we train BtcDet for 40 epochs, while on the WOD, we train BtcDet for 30 epochs. The BtcDet is end-to-end optimized by the ADAM optimizer (Kingma and Ba 2014) from scratch. We applies the widely adopted data augmentations (Shi et al. 2020; Deng et al. 2020; Lang et al. 2019; Yang et al. 2020; Ye et al. 2020), which includes flipping, scaling, rotation and the ground-truth augmentation.

We evaluate BtcDet on the KITTI val split after training it on the train split. To evaluate the model on the KITTI test set, we train BtcDet on 80%80\% of all train+val data and hold out the remaining 20% data for validation. Following the protocol in (Geiger et al. 2013), results are evaluated by the Average Precision (AP) with an IoU threshold of 0.7 for cars and 0.5 for pedestrians and cyclists.

KITTI validation set. As summarized in Table 1, we compare BtcDet with the state-of-the-art LiDAR-based 3D object detectors on cars, pedestrians and cyclists using the AP under 40 recall thresholds (R40). We reference the R40 APs of SA-SSD, PV-RCNN and Voxel R-CNN to their papers, the R40 APs of SECOND to (Pang et al. 2020) and the R40 APs of PointRCNN and PointPillars to the results of the officially released code. We also report the published 3D APs under 11 recall thresholds (R11) for the moderate car objects. On all object classes and difficulty levels, BtcDet outperforms models that only supervise bounding boxes (Eq.1) as well as structure-aware models (Eq.2). Specifically, BtcDet outperforms other models by 2.05%2.05\% 3D R11 AP on the moderate car objects, which makes it the first detector that reaches above 86%86\% on this primary metric.

KITTI test set. As shown in Table 2, we compare BtcDet with the front runners on the KITTI test leader board. Besides the official metrics, we also report the mAPs that average over the APs of easy, moderate, and hard objects. As of May. 4th, 2021, compared with all the models associated with publications, BtcDet surpasses them on car and cyclist detection by big margins. Those methods include the models that take inputs of both LiDAR and RGB images and the ones taking LiDAR input only. We also list more comparisons and the results in Appendix F.

2 Evaluation on the Waymo Open Dataset

We also compare BtcDet with other models on the Waymo Open Dataset (WOD). We report both 3D mean Average Precision (mAP) and 3D mAP weighted by Heading (mAPH) for vehicle detection. The official metrics also include separate mAPs for objects belonging to different distance ranges. Two difficulty levels are also introduced, where the LEVEL_1 mAP calculates for objects that have more than 5 points and the LEVEL_2 mAP calculates for objects that have more than 1 point.

As shown in Table 3, BtcDet outperforms these state-of-the-art detectors on all distance ranges and all difficulty levels by big margins. BtcDet outperforms other detectors on the LEVEL_1 3D mAP by 2.99%2.99\% and the LEVEL_2 3D mAP by 3.51%3.51\%. In general, BtcDet brings more improvement on the LEVEL_2 objects, since objects with fewer points usually suffer more from occlusion and signal miss. These strong results on WOD, one of the largest published LiDAR datasets, manifest BtcDet’s ability to generalize.

3 Ablation Studies

We conduct ablation studies to demonstrate the effectiveness of the shape occupancy and the feature integration strategies. All model variants are trained on the KITTI train split and evaluated on the val split.

Shape Features. As shown in Table 4, we conduct ablation studies by controlling the shape features learned by Ω\Omega and the features used in the integration. All the model variants share the same architecture and integration strategies.

Similarly to (Hu et al. 2020), BtcDet2 directly fuses the binary map of. ROC∪RSM\mathcal{R_{OC}}\cup\mathcal{R_{SM}} into the detection pipeline. Although the binary map provides the information of occlusion, the improvement is limited since that the regions with code 1 are mostly background regions and less informative.

BtcDet3 learns P(OS)\mathcal{P}(\mathcal{O_{S}})⟂ directly. The network Ω\Omega predicts probability for Cartesian voxels. One Cartesian voxel will cover multiple spherical voxels when being close to the sensor, and will cover a small portion of a spherical voxel when being located at a remote distance. Therefore, the occlusion regions are misrepresented in the Cartesian coordinate.

BtcDet4 convert the probability to hard occupancy, which cannot inform the downstream branch if a region is less likely or more likely to contain object shapes.

These experiments demonstrate the effectiveness of our choices for shape features, which help the main model improve 2.862.86 AP over the baseline BtcDet1.

Integration strategies. We conduct ablation studies by choosing different layers of Ψ\Psi to concatenate with P(OS)\mathcal{P}(\mathcal{O_{S}})⟂ and whether to use P(OS)\mathcal{P}(\mathcal{O_{S}})⟂ to form fgeof_{geo}. The former mostly affects the proposal generation, while the latter affects proposal refinement.

In Table 5, the experiment on BtcDet5 shows that we can improve the final prediction AP by 0.80.8 if we only integrate P(OS)\mathcal{P}(\mathcal{O_{S}})⟂ for proposal refinement. On the other hand, the experiment on BtcDet6 shows the integration with Ψ\Psi alone can improve the AP by 1.21.2 for proposal box and final bounding box prediction AP by 2.02.0 over the baseline.

The comparisons of BtcDet7, BtcDet8 and BtcDet (main) demonstrates integrating P(OS)\mathcal{P}(\mathcal{O_{S}})⟂ with Ψ\Psi’s first two layers is the best choice. Since P(OS)\mathcal{P}(\mathcal{O_{S}}) is a low level feature while the third layer of Ψ\Psi would contain high level features, we observe a regression when BtcDet8 also concatenates P(OS)\mathcal{P}(\mathcal{O_{S}})⟂ with Ψ\Psi’s third layer.

These experiments demonstrate both the integration with Ψ\Psi and the integration to form fgeof_{geo} can bring improvement independently. When working together, two integrations finally help BtcDet surpass all the state-of-the-art models.

Conclusion and Future Work

In this paper, we analyze the impact of shape miss on 3D object detection, which is attributed to occlusion and signal miss in point cloud data. To solve this problem, we propose Behind the Curtain Detector (BtcDet), the first 3D object detector that targets this fundamental challenge. A training method is designed to learn the underlying shape priors. BtcDet can faithfully estimate the complete object shape occupancy for regions affected by occlusion and signal miss. After the integration with the probability estimation, both the proposal generation and refinement are significantly improved. In the experiments on the KITTI Dataset and the Waymo Open Dataset, BtcDet surpasses all the published state-of-the-art methods by remarkable margins. Ablation studies further manifest the effectiveness of the shape features and the integration strategies. Although our work successfully demonstrates the benefits of learning occluded shapes, there is still room to improve the model efficiency. Designing models that expedite occlusion identification and shape learning can be a promising future direction.

Appendix A Data and Code License

The datasets we use for experiments are the KITTI Dataset (Geiger et al. 2013) and the Waymo Open Dataset (Sun et al. 2019). Both of them are well-known and licensed for academic research.

We license our code under “Apache License 2.0”. The code will be released.

Appendix B Heuristic for Source Object Selection

To approximate the complete object shapes for a target object AA, a heuristic H(A,B)\mathcal{H}(A,B) is created to evaluate if a source object BB covers most of AA and can provides points in the regions of AA’s shape miss. The lower the score, the better a object BB is for AA. The heuristic is:

where PAP_{A} and PBP_{B} are the object point sets and DAD_{A} and DBD_{B} are their bounding boxes.

The first term ∑x∈PAmin⁡y∈PB∣∣x−y∣∣\sum_{x\in P_{A}}\min_{y\in P_{B}}||x-y|| measures if AA’s points are well covered by BB’s points (half Chamfer Distance).

The second term αIoU(DA,DB)\alpha IoU(\mathcal{D}_{A},\mathcal{D}_{B}) measures the similarity of their bounding box size.

The third term \beta/\big{|}\{x:x\in Vox(P_{B}),x\notin Vox(P_{A})\}\big{|} measures the number of extra voxels that B can add to A.

Appendix C Training Target and Loss

We follow the most popular RPN design of anchor-based 3D detection models (Lang et al. 2019; Yan et al. 2018; Shi et al. 2020; Deng et al. 2020).

To generate region proposals, for each class, we first set anchor size as the size of the average 3D objects, and set anchor orientations at 0∘0^{\circ} and 90∘90^{\circ}. Then, we We adopt the box encoding for RPN, which is introduced in (Lang et al. 2019; Yan et al. 2018):

Car (KITTI) or vehicle (WOD) anchors are assigned to ground-truth objects if their IoUs are above 0.6 (fg=1f_{g}=1) or treated as in background if their IoUs are less than 0.45 (fg=0f_{g}=0). The anchors with IoUs in between are ignored in training. For pedestrians and cyclists, the foreground object matching threshold is 0.5 and the background matching threshold is 0.35.

To deal with adversarial angle problem (the orientation at or π\pi), we follow (Yan et al. 2018) and set the regression loss for orientation as:

where “p” indicates the predicted value. Since the above loss treats opposite directions indistinguishably, a direction classifier is also used and supervised by a softmax loss Ldir\mathcal{L}_{dir}.

We use Focal Loss (Lin et al. 2017) as the classification loss:

in which ppp_{p} is the predicted foreground score. The parameters of the focal loss are α=0.25\alpha=0.25 and γ=2\gamma=2. The total loss of RPN is:

where NaN_{a} is the number of sampled anchors, 1(fg≥1)\mathfrak{1}(f_{g}\geq 1) means the regression losses are only applied on the foreground anchors, Lrpnreg\mathcal{L}_{rpn}^{reg} is the SmoothL1SmoothL_{1} regression loss on the encoded x,y,z,w,l,h as described in Eq.13 and Eq.14 and Lrpndir\mathcal{L}_{rpn}^{dir} is the direction classification loss for predicting the angle bin.

C.2 Proposal Refinement

Following (Jiang et al. 2018; Li et al. 2019; Yan et al. 2018; Shi et al. 2020, 2020; Deng et al. 2020), there are two branches in the proposal refinement module, one for class confidence score and another for box regression. We follow (Jiang et al. 2018; Shi et al. 2020; Li et al. 2019; Shi et al. 2019b) and adopt the 3D IoU weighted RoI confidence training target for each RoI:

To conduct regression for bounding box refinement, We adopt the state-of-the-art residual-based box encoding functions:

where NpN_{p} is the number of sampled proposal, Lprcls\mathcal{L}_{pr}^{cls} is the binary cross entropy loss using the training targets Eq.21, 1(IoU≥0.55)\mathfrak{1}(IoU\geq 0.55) means we only apply regression loss on positive proposals, Lprθ\mathcal{L}_{pr}^{\theta} and Lprreg\mathcal{L}_{pr}^{reg} are similar to the corresponding regression losses in the RPN.

C.3 Total Loss

The total loss is the combination of Lshape\mathcal{L}_{shape} introduced in Section 3.1 of the main paper, the RPN loss Lrpn\mathcal{L}_{rpn} and the proposal refinement loss Lpr\mathcal{L}_{pr}:

where we conduct grid search and find the weighting factor of 0.3 helps BtcDet achieve the best results.

Appendix D Network Architecture

In this section, we describe the network architecture of the shape occupancy network, the detection feature backbone network, the region proposal network, and the proposal refinement network of BtcDet.

As visualized in Figure 5, we use a lightweight spherical sparse 3D convolution network with five sparse-conv layers. Two of them are down-sampling and two of them are up-sampling layers, each consists of a regular sparse-conv of stride 2 following by a sub-manifold sparse-conv (Graham and van der Maaten 2017). The dimensions of these layers’ output features are 16, 32, 64, 32, and 32, respectively.

D.2 Detection Backbone Network ΨΨ\Psi

The backbone of the detection feature extraction network follows (Shi et al. 2020; Deng et al. 2020) but has thinner network layers. The point cloud is voxelized into Cartesian voxels where the features of each occupied voxel are the mean of the points’ xyz and features (e.g., intensity). Besides, the sparse probability tensor of object occupancy in the spherical coordinate has been transformed to the Cartesian coordinate P(OS)\mathcal{P}(\mathcal{O_{S}})⟂, so that two channels from P(OS)\mathcal{P}(\mathcal{O_{S}})⟂ can be concatenated with layers of Ψ\Psi. One channel holds the occupancy probability P(OS)\mathcal{P}(\mathcal{O_{S}}) and the other holds the binary code if P(OS)\mathcal{P}(\mathcal{O_{S}}) exists in a voxel. As visualized in Figure 6, three down-sampling layers down-sample the features to 8×\times smaller, which are fed into the region proposal network. The feature maps of the second, the third, and the final layer are further integrated with P(OS)\mathcal{P}(\mathcal{O_{S}})⟂ to form a local geometric feature fgeof_{geo}, which supports the proposal refinement.

D.3 Region Proposal Network

We stack the input features to bird-eye view 2D features. Then, a thinner version of the 2D convolution networks in (Lang et al. 2019; Yan et al. 2018) propagates the features and output residues of 2 anchors per class per grid on the output feature maps. Instead of dimensions of 256 as in (Shi et al. 2020; Deng et al. 2020), the intermediate feature maps in our 2D convolution networks has feature dimension of 128.

D.4 Proposal Refinement Network

A local grid of a region proposal has a grid size of (2, 4, 12) along the locally orientated axis of Z, Y, X. As shown in Figure 7, the aggregation network of the pooled local geometric features consist of three layers with the strides of (1,1,2), (1,2,2), (2,2,3). The first two layers have zero paddings, while the last layer does not. After that, we send them to several fully connected layers.

We have 3×3×33\times 3\times 3 this kind of local grids for each proposal box. The center of a local grid (xgrid,ygrid,zgridx_{grid},y_{grid},z_{grid}) have a shift away from the proposal center (xp,yp,zpx_{p},y_{p},z_{p}) by distances (Δx∈{±λ,0}×wp,Δy={±λ,0}×lp,Δz={±λ,0}×hp\Delta_{x}\in\{\pm\lambda,0\}\times w_{p},\Delta_{y}=\{\pm\lambda,0\}\times l_{p},\Delta_{z}=\{\pm\lambda,0\}\times h_{p}), where wp,lp,hpw_{p},l_{p},h_{p} is the width, length and height of the proposal box. We find λ=0.25\lambda=0.25 achieves the best results. The proposal refinement network aggregates results from all these shifted local grids and outputs the residues regression and the class confidence score, which lead to the final bounding box predictions.

Appendix E Occupancy Estimation of Complete Object Shapes

We show the evaluation of the occupancy estimation in Table 6. The results are averaged among all voxels in the regions of ROC∪RSM\mathcal{R_{OC}}\cup\mathcal{R_{SM}}. A prediction is considered positive if P(OS)\mathcal{P}(\mathcal{O_{S}}) >> threshold. The metrics we evaluate are precision, recall, F1 score, accuracy, and object coverage. The object coverage is the percentage of all bounding boxes that contain at least one positive voxel (P(OS)\mathcal{P}(\mathcal{O_{S}}) >> threshold). We show the measures on three thresholds of 0.3,0.50.3,0.5, and 0.70.7. The accuracy results under all thresholds are very high (¿99%) since the classes are extremely imbalanced. However, no matter under which threshold, we can achieve relatively high object coverage, which means the estimation is faithful enough for RPN and other downstream networks to rely on.

Appendix F More Comparison Results on the KITTI Test Set

We show more results of comparisons between BtcDet and other state-of-the-art detectors in Table 7. Because the average point number in pedestrians is smaller than other objects, the shape estimation is sensitive to a few observed points. Therefore, if the point number distribution of pedestrians in the test split is different, our model may not be able to provide an accurate shape occupancy estimation. As a result, BtcDet’s pedestrian detection on KITTI’s test split does not perform as well as on KITTI’s val split. We consider improving the results with this situation in our future works.

Appendix G More Visualization for the Complete Object Shape Approximation

We show more results of the assembled complete object shapes of cyclists and pedestrians in this section. Figure 8 visualizes the process for cyclists which includes mirroring both source and target objects. Figure 9 and 10 visualizes the process for pedestrians which does not mirror the objects since pedestrians are less likely to be symmetric. The blue points are the points of the target object and the red points are the points of the source objects. The assembled object faithfully covers the originally partially observed parts of the target objects and provides reasonable recovery points in the shape miss regions of the target objects.

We show the qualitative results of the occupancy probability for vehicle objects on the Waymo Open Dataset (Sun et al. 2019). Figure 11 contains zoomed in views of the occupancy probability while Figure 12 contains full scene views. The higher probability one is estimated, the larger opacity we apply to the spherical voxel.

References