SE-SSD: Self-Ensembling Single-Stage Object Detector From Point Cloud
Wu Zheng, Weiliang Tang, Li Jiang, Chi-Wing Fu
Self-Ensembling Single Stage Detector
Figure LABEL:pipeline shows the framework of our self-ensembling single-stage object detector (SE-SSD), which has a teacher SSD (left) and a student SSD (right). Different from prior works on outdoor 3D object detection, we simultaneously employ and train two SSDs (of same architecture), such that the student can explore a larger data space via the augmented samples and be better optimized with the associated soft targets predicted by the teacher. To train the whole SE-SSD, we first initialize both the teacher and student with a pre-trained SSD model. Then, started from an input point cloud, our framework has two processing paths:
In the first path (blue arrows in Figure LABEL:pipeline), the teacher produces relatively precise predictions from the raw input point cloud. Then, we apply a set of global transformations on the prediction results and take them as soft targets to supervise the student SSD.
In the second path (green arrows in Figure LABEL:pipeline), we perturb the same input by the same global transformations as in the first path plus our shape-aware data augmentation (Section 1.4). Then, we feed the augmented input to the student, and train it with (i) our consistency loss (Section 1.2) to align the student predictions with the soft targets; and (ii) when we augment the input, we bring along its hard targets (Figure LABEL:pipeline (top right)) to supervise the student with our orientation-aware distance-IoU loss (Section 1.3).
In the training, we iteratively update the two SSD models: optimize the student with the above two losses and update the teacher using only the student parameters by a standard exponential moving average (EMA). Thus, the teacher can obtain distilled knowledge from student and produce soft targets to supervise student. So, we call the final trained student a self-ensembling single-stage object detector.
The model has the same structure as [zheng2020cia] to efficiently encode point clouds, but we remove the Confidence Function and DI-NMS. It has a sparse convolution network (SPConvNet), a BEV convolution network (BEVConvNet), and a multi-task head (MTHead). BEV means bird’s eye view. After point cloud voxelization, we find mean 3D coordinates and point density per voxel as the initial feature, then extract features using SPConvNet, which has four blocks ({2, 2, 3, 3} submanifold sparse convolution [graham20183d] layers) with a sparse convolution [liu2015sparse] layer at the end. Next, we concatenate the sparse 3D feature along into 2D dense ones for feature extraction with the BEVConvNet. Lastly, we use MTHead to regress bounding boxes and perform classification.
2 Consistency Loss
In 3D object detection, the patterns of point clouds in pre-defined anchors may vary significantly due to distances and different forms of object occlusion. Hence, samples of the same hard target may have very different point patterns and features. In contrast, soft targets can be more informative per training sample, and helps to reveal the difference between data samples of the same class [hinton2015distilling]. This motivates us to treat the relatively more precise teacher predictions as soft targets and employ them to jointly optimize the student with hard targets. Accordingly, we formulate a consistency loss to optimize the student network with soft targets.
We first design an effective IoU-based matching strategy before calculating the consistency loss, aiming to pair the non-axis-aligned teacher and student boxes predicted from the very sparse outdoor point clouds. To obtain high-quality soft targets from the teacher, we first filter out those predicted bounding boxes (for both teacher and student) with confidence less than threshold , which helps reduce the calculation of the consistency loss. Next, we calculate the IoU between every pair of remaining student and teacher bounding boxes, and filter out the pairs with IoUs less than threshold , thus avoiding to mislead the student with unrelated soft targets; We denote and as the initial and final number of box pairs, respectively. Thus, we keep only the highly-overlapping student-teacher pairs. Lastly, for each student box, we pair it with the teacher bounding box that has the largest IoU with it, aiming to increase the confidence of the soft targets. Compared with hard targets, the filtered soft targets are often closer to the student predictions, as they are predicted based on similar features. So, soft targets can better guide the student to fine-tune the predictions and reduce the gradient variance for better training.
Different from the IoU loss, Smooth- loss [liu2016ssd] can evenly treat all dimensions in the predictions, without biasing toward any specific one, so the features corresponding to different dimensions can also be evenly optimized. Hence, we adopt it to formulate our consistency loss for bounding boxes () to minimize the misalignment errors between each pair of teacher and student bounding boxes:
where denotes the Smooth- loss of , and and denote the sigmoid classification scores of student and teacher, respectively. Here, we adopt the sigmoid function to normalize the two predicted confidences, such that the deviation between the normalized values can be kept inside a small range. Combining Eqs \eqrefloc_cons and \eqrefcls_cons, we can obtain the overall consistency loss as
where we empirically set the same weight for both terms.
3 Orientation-Aware Distance-IoU Loss
In supervised training with hard targets, Smooth- loss [liu2016ssd] is often adopted to constrain the bounding box regression. However, due to long distances and occlusion in outdoor scenes, it is hard to acquire sufficient information from the sparse points to precisely predict all dimensions of the bounding boxes. To better exploit hard targets for regressing bounding boxes, we design the Orientation-aware Distance-IoU loss (ODIoU) to focus more attention on the alignment of box centers and orientations between the predicted and ground-truth bounding boxes; see Figure 1.
Inspired by [zheng2020distance], we impose a constraint on the distance between the 3D centers of the predicted and ground-truth bounding boxes to minimize the center misalignment. More importantly, we design a novel orientation constraint on the predicted BEV angle, aiming to further minimize the orientation difference between the predicted and ground-truth boxes. In 3D object detection, such a constraint is significant for the precise alignment between the non-axis-aligned boxes in the bird’s eye view (BEV). Also, we empirically find that this constraint is an important means to further boost the detection precision. Compared with Smooth- loss, our ODIoU loss enhances the alignment of box centers and orientations, which are easy to infer from the points distributed on the object surface, thus leading to a better performance. Overall, our ODIoU loss is formulated as
where and denote the predicted and ground-truth bounding boxes, respectively, denotes the distance between the 3D centers of the two bounding boxes (see in Figure 1), denotes the diagonal length of the minimum cuboid that encloses both bounding boxes; denotes the BEV orientation difference between and ; and is a hyper-parameter weight.
In our ODIoU loss formula, is an important term we designed specifically to encourage the predicted bounding box to rotate to the nearest direction that is parallel to the ground-truth orientation. When equals or , \ie, the orientations of the two boxes are perpendicular to each other, so the term attains its maxima. When equals , , or , the term attains its minima, which is zero. As shown in Figure 2, we can further look at the gradient of . When the training process minimizes the term, its gradient will help to bring to , , or , which is the nearest location to minimize the loss. It is because the gradient magnitude is positively correlated to the angle difference, thus promoting fast convergence and smooth fine-tuning in different training stages.
Besides, we use the Focal loss [lin2017focal] and cross-entropy loss for the bounding box classification () and direction classification (), respectively. Hence, the overall loss to train the student SSD is
where \mathcal{L}^{s}_{\text}{box} is the ODIoU loss for regressing the boxes, and the loss weights , , and are hyperparameters. Futher, our SSD can be pre-trained with the same settings as SE-SSD but without the consistency loss and teacher SSD.
4 Shape-Aware Data Augmentation
Data augmentation is important to improve a model’s generalizability. To enable the student SSD to explore a larger data space, we design a new shape-aware data augmentation scheme to boost the performance of our detector. Our insight comes from the observation that the point cloud patterns of ground-truth objects could vary significantly due to occlusions, changes in distance, and diversity of object shapes in practice. So, we design the shape-aware data augmentation scheme to mimic how point clouds are affected by these factors when augmenting the data samples.
By design, our shape-aware data augmentation scheme is a plug-and-play module. To start, for each object in a point cloud, we find its ground-truth bounding box centroid and connect the centroid with the box faces to form pyramidal volumes that divide the object points into six subsets. Observing that LiDAR points are distributed mainly on object surfaces, the division is like an object disassembly, and our augmentation scheme efficiently augments each object’s point cloud by manipulating these divided point subsets like disassembled parts.
In details, our scheme performs the following three operations with randomized probabilities , , and , respectively: (i) random dropout removes all points (blue) in a randomly-chosen pyramid (Figure 3 (top-left)), mimicking a partial object occlusion to help the network to infer a complete shape from the remained points. (ii) random swap randomly selects another input object in the current scene and swap a point subset (green) with the point subset (yellow) in the same pyramid of the other input object (Figure 3 (middle)), thus increasing the diversity of object samples by exploiting the surface similarity across objects. (iii) random sparsifying subsamples points in a randomly-chosen pyramid using farthest point sampling [qi2017pointnet++], mimicking the sparsity variation of points due to changes in distance from LiDAR camera; see the sparsified points (red) in Figure 3.
Furthermore, before the shape-aware augmentation, we perform a set of global transformations on the input point cloud, including a random translation, flipping, and scaling; see “global transformations” in Figure LABEL:pipeline.