Object DGCNN: 3D Object Detection using Dynamic Graphs

Yue Wang, Justin Solomon

Related Work

2D object detection. Object recognition research has been transitioning from models with hand-crafted components to models with limited post-processing. One-stage detectors remove the complicated region proposal networks in two-stage objectors , yielding more efficient training and testing. Anchor-free methods further simplify the one-stage pipeline by shifting from per-anchor prediction to per-pixel prediction. However, these methods still make dense predictions and rely on NMS to reduce redundancy. To alleviate this issue, DETR formulates object detection as a set-to-set prediction problem. It introduces a set-to-set loss that implicitly penalizes redundant boxes, removing the necessity of post-processing. To accelerate convergence, Deformable DETR proposes deformable self-attention and streamlines the optimization process. Our method also formulates 3D object detection as set prediction, but with a customized design for 3D.

3D object detection. VoxelNet generalizes one-stage object detection to 3D. It uses 3D dense convolutions to learn representations on voxelized point clouds, which is too inefficient to capture fine-grained features. To address that, PIXOR and PointPillars project points to a birds-eye view (BEV) and operate on 2D feature maps; PointNet aggregates features within each BEV pixel. We use a variant of PointPillars for 3D detection (§3). These methods are efficient but drop information along the vertical axis. To accompany the BEV projection, MVF adds a spherical projection. PillarOd and CenterPoint use pillar-centric object detection, making predictions per BEV pixel (pillar) rather than per anchor. These anchor-free methods simplify 3D object detection while maintaining efficiency. Beyond SSD-style one-stage models, Complex-YOLO extends YOLO to 3D for real-time perception. PointRCNN employs a two-stage architecture for high-quality detection. To improve representations of two-stage models, PVRCNN proposes a point-voxel feature set abstraction layer to leverage the flexible receptive fields of PointNet-based networks. Unlike works on point clouds, LaserNet operates on raw range scans with comparable performance. combine point clouds with camera images. Frustum-PointNet leverages 2D object detectors to form a frustum crop of points and then uses PointNet to aggregate features. describes an end-to-end learnable architecture that exploits continuous convolutions to fuse feature maps. VoteNet generalizes Hough voting for 3D object detection in point clouds. DOPS extends VoteNet and predicts 3D object shapes. In addition to visual input, shows that high-definition (HD) maps can boost performance of 3D object detectors. argues that multi-tasking can learn better representations than single-tasking. Beyond supervised learning, learns a perception model for unknown classes.

DGCNN. DGCNN pioneered learning point cloud representations via dynamic graphs. It models point clouds as connected graphs, which are dynamically built using kk-nearest neighbors in the latent space. DGCNN learns per-point features through message passing. However, it operates on point clouds for single object recognition and semantic segmentation. One of our key contributions is to generalize DGCNN to model scene-level object relations for 3D detection.

Knowledge distillation (KD). KD compresses knowledge from an ensemble of models into a single smaller model . generalizes this idea and combines it with deep learning. KD transfers knowledge from a teacher model to a student model by minimizing a loss, in which the target is the distribution of class probabilities induced by the teacher. improve knowledge distillation for classification. Beyond image classification, KD has been extended to improve object detection. leverages FitNets for object detection, addressing obstacles such as class imbalance, loss instability, and feature distribution mismatch. distills between region proposals, accelerating training with added instability. To address this issue, uses fine-grained representation imitation using object masks. uses KD to tackle a continual learning problem.

Privileged information. introduces the framework of learning with privileged information in the context of support vector machines (SVMs), wherein additional information is accessible during training but not testing. unifies KD and learning using privileged information theoretically. identifies practical applications, e.g., transferring knowledge from localized data to non-localized data, from high resolution to low resolution, from color images to edge images, and from regular images to distorted images. To mediate uncertainty and improve training efficiency, makes the variance of Dropout a function of privileged information. We extend these methods to 3D data, in which privileged information consists of dense point clouds aggregated from LiDAR sequences.

Overview

Following state-of-the-art in large-scale object detection, our pipeline learns a grid-based intermediate representation to capture local features (§3). We test two standard learning-based methods for collecting local point cloud features on a birds-eye view (BEV) grid. While in principle it might be possible to avoid grids entirely in our pipeline, this BEV representation is far more efficient and—as observed in previous work—is sufficient to find objects reliably in autonomous driving, where there is likely only one object above any given grid cell on the ground plane.

Our main architecture contribution is the Object DGCNN pipeline (§4), which transitions from this BEV grid of features to a set of object bounding boxes. Object DGCNN draws inspiration from the DGCNN architecture; its layers alternate between local feature transformations and kk-nearest neighbor aggregation to capture relationships between objects. Unlike conventional DGCNN, however, Object DGCNN incorporates features from the BEV grid in each of its layers; each layer incorporates several queries into the BEV to refine object position estimates. The output of Object DGCNN is a set of objects in the scene. We use a permutation-invariant loss (10) to measure divergence from the ground truth set of objects.

The pipeline above does not require hand-designed post-processing like NMS; our output boxes are usable directly for object detection. Beyond simplifying the object detection pipeline, this allows us to propose object detection-specific distillation procedures (§6.3) that further improve performance. These use one network to train another, e.g., to train a network operating on sparse point clouds to output features that imitate those learned by a network trained on denser, more detailed point clouds.

Local Features

As an initial step, modern 3D object detection models scatter points into either BEV pillars or 3D voxels and then use convolutional neural networks to extract features on a grid. This strategy accelerates object detection for large point clouds. We test two neural network architectures for BEV feature extraction, detailed below.

An alternative BEV embedding is SparseConv . If FV(i)F_{V}(i) returns the set of points in voxel ii, SparseConv collects point-wise features into voxel-wise features by

Object DGCNN

Desiderata. Object DGCNN uses a DGCNN-inspired architecture but incorporates grid-based BEV features, built on the philosophy that local features (§3) are reasonable to store on a dense grid, but object predictions are better modeled using sets. Hence, we require a new architecture and set-to-set loss that encourage bounding box diversity.

Object DGCNN uses LL layers that follow a series of set-based computations to produce bounding box predictions from the BEV feature maps. Each layer employs the following steps (Figure 1):

predict a set of query points and attention weights;

collect BEV features from keypoints determined by the queries; and

model object-object interactions via DGCNN.

Each layer results in a more refined set of bounding box predictions, one per query. At the end of these layers, we match the prediction set with the ground-truth set in a one-to-one fashion and evaluate a set-to-set object detection loss.

Next, we collect a BEV feature fik\boldsymbol{f}_{ik} associated to each neighbor point pik=pi+δik\boldsymbol{p}_{ik}=\boldsymbol{p}_{i}+\boldsymbol{\delta}_{ik} :

This generates scene-aware features; each object query “attends” to a certain area in the scene.

Set-to-set loss. After LL Object DGCNN layers as described above, we are left with a set of M∗M^{\ast} queries QL\mathcal{Q}_{L} used to predict our bounding boxes. For each query qLi\boldsymbol{q}_{Li}, we use a classification network to predict a categorical label ci^\hat{c_{i}} and a regression network to predict bounding box parameters bi^\hat{\boldsymbol{b}_{i}}. Our final task is to assign the predictions to the ground-truth boxes and compute a set-to-set loss.

This strategy can assign a box to multiple nearby BEV pixels. This one-to-many assignment provides dense supervision for the object detector and eases optimization. Since the training objective encourages each BEV pixel to predict the same surrounding box, however, redundant boxes are inevitable. So, NMS is usually required to remove redundant boxes at inference time.

Rather than performing dense predictions in the BEV, we make per-query predictions. Typically, M∗M^{\ast} is much larger than the number of ground-truth boxes MM. To account for this difference, we pad the set of ground-truth boxes with ∅\varnothings (no object) up to M∗M^{\ast}. Following , we use an objective built on an optimal matching between these two sets. We define the optimal bipartite matching as

Distillation

Object DGCNN enables a new set-to-set knowledge distillation (KD) pipeline. KD usually involves a teacher model T\mathcal{T} and a student model S\mathcal{S}. The common practice is to align the outputs of the student with those of the teacher using L2\mathcal{L}_{2} distance or KL-divergence. In past 3D object detection methods, since final performance heavily relies on NMS and the predictions are post-processed to be a smaller set, distilling the teacher to the student is neither efficient nor effective. Since our set-based detection model is NMS-free, we can easily distill the information between models with homogeneous detection heads (per-query object detection head in our case). First, we train a teacher T\mathcal{T} using the method above with the loss in (10). Then, we train a student S\mathcal{S} with supervision given by T\mathcal{T} and the ground-truth. The class label and box parameters predicted by the teacher for each object query are cjTc_{j}^{\mathcal{T}} and bjT\boldsymbol{b}_{j}^{\mathcal{T}}, respectively. The corresponding student outputs are cjSc_{j}^{\mathcal{S}} and bjS\boldsymbol{b}_{j}^{\mathcal{S}}. We find an optimal matching between the output set of the teacher and that of the student:

Then, the optimal matching’s KD loss is given by

Experiments

We present our experiments in four parts. We introduce the dataset, metrics, implementation, and optimization details in §6.1. Then, we demonstrate performance on the nuScenes dataset in §6.2. We present knowledge distillation results in §6.3. Finally, we provide ablation studies in §6.4.

Dataset. We experiment on the nuScenes dataset . nuScenes provides rich annotations and diverse scenes. It has 1K short sequences captured in Boston and Singapore with 700, 150, 150 sequences for training, validation, and testing, respectively. Each sequence is ∼ ⁣20\sim\!20s and contains 400 frames. This dataset provides annotation every 0.5s, leading to 28K, 6K, 6K annotated frames for training, validation, and testing. nuScenes uses 32-beam LiDAR, producing 30K points per frame. Following common practice, we use calibrated vehicle pose information to aggregate every 9 non-key frames to key frames, so each annotated frame has ∼ ⁣300\sim\!300K points. The annotations include 23 classes with a long-tail distribution, of which 10 classes are included in the benchmark.

Metrics. The major metrics are mean average precision (mAP) and the nuScenes detection score (NDS). In addition, we use a set of true positive metrics (TP metrics), which include average translation error (ATE), average scale error (ASE), average orientation error (AOE), average velocity error (AVE), and average attribute error (AAE). These metrics are computed in the physical unit.

Model architecture. Our model consists of three parts: a point-based feature extractor, a DGCNN to encode object queries and to connect the point cloud features to object queries, and a detection head to output the categorical label and bounding box parameters. We experiment with PointPillars and SparseConv as feature extractors. The three blocks of the PointPillars backbone have convolutionallayers,withdimensionsconvolutional layers, with dimensions and strides ;theinputfeaturesaredownsampledto1/2,1/4,1/8oftheoriginalfeaturemap.ForSparseConv,weusefourblocksof; the input features are downsampled to 1/2, 1/4, 1/8 of the original feature map. For SparseConv, we use four blocks of 3D sparse convolutional layers, with dimensions andstridesand strides; the input features are downsampled to 1/2, 1/4, 1/8, 1/8 of the original feature map. For SparseConv, we transform the features into BEV by collapsing the zz-axis. Both backbones use two deformable self-attention layers with dimensions totransformtheBEVfeatures.Then,weusetwoDGCNNstoencodetheobjectqueries.EachDGCNNcontainstwoEdgeConvlayerswithdimensionsto transform the BEV features. Then, we use two DGCNNs to encode the object queries. Each DGCNN contains two EdgeConv layers with dimensions, both with 16 nearest neighbors. For each object query, we predict four points in the BEV to obtain and aggregate the BEV features. The final feature for this object query is the weighted sum of features of these four BEV points. The final detection head takes the features of each object query and predicts class label and bounding box parameters w.r.t. the reference point.

Training & inference. We use AdamW to train the model. The weight decay for AdamW is 10−210^{-2}. Following a cyclic schedule , the learning rate is initially 10−410^{-4} and gradually increased to 10−310^{-3}, which is finally decreased to 10−810^{-8}. The model is initialized with a pre-trained PointPillars network on the same dataset. We train for 20 epochs on 8 RTX 3090 GPUs. During inference, we take the top 100 objects with highest classification scores as the final predictions.We do not use any post-processing such as NMS. For evaluation, we use the toolkit provided with the nuScenes dataset.

2 Object DGCNN

We compare to top-performing methods on the nuScenes dataset in Table 1. PointPillars is an anchor-based method with reasonable trade-off between performance and efficiency. FreeAnchor extends PointPillars by learning how to assign anchors to the ground-truth. RegNetX-400MF-SECFPN uses neural architecture search (NAS) to learn a flexible neural network for 3D detection; it is essentially a variant of PointPillars with an enhanced backbone network. Different from anchor-based methods, Pillar-OD makes predictions per pillar, alleviating the class imbalance issue caused by anchors. CenterPoint exploits similar detection heads, with better performance using better training scheduling and data augmentation. For these methods, we use re-implementations in MMDetection3D , which match the performances in the original papers.

We mainly compare to CenterPoint with both PointPillars and SparseConv backbones, denoted as “voxel” and “pillar” respectively. Our method outperforms other methods significantly including CenterPoint with NMS. Without NMS, the performance of CenterPoint drops considerably while our method is unaffected by NMS. This finding verifies the DGCNN implicitly models object relations and removes redundant boxes.

3 Set-to-set distillation

In this section, we present experiments involving our set-to-set distillation pipeline. We conduct three types of distillation. First, we distill a teacher model with a SparseConv backbone to a student model with a PointPillars backbone (denoted as “voxel→\rightarrowpillar”). This aligns with the common knowledge distillation setup for classification. We compare to feature-based distillation and pseudo label based methods. The objective of feature-based distillation is to align the middle-level features of the teacher model and the student model while the pseudo label based methods generate pseudo training examples with the pre-trained teacher networks. As Table 2 shows, our set-to-set distillation achieves better performance, confirming that distilling the last stage of the object detection model is more effective than distilling feature maps.

Second, we perform self-distillation (denoted as “voxel→\rightarrowvoxel” and “pillar→\rightarrowpillar”), where the teacher and the student are identical and take the same point clouds as input. As Table 3 shows, even when the teacher network and the student network have the same capacity, self-distillation still introduces a performance boost. This finding is consistent with the results in .

Finally, we try distillation with privileged information , where the teacher gets access to privileged information but the student does not. Following , the teacher takes dense point clouds, and the student takes sparse point clouds (denoted as "dense→\rightarrowsparse"). To limit computation time, we train each model over a shorter period of time. The goal is for the student model to learn the same representations as the teacher model without knowing the dense inputs. In Table 4, we compare this setup with self-distillation, where the difference is the teacher model and the student model take the same sparse point clouds in self-distillation. The student achieves better performance when the teacher takes dense point clouds. The result suggests that set-to-set knowledge distillation is an effective approach to transfer insight from privileged information.

4 Ablation

We provide ablation studies on different components of our model to verify assorted design choices. First, we study the improvements of DGCNN over its counterpart, multi-head self-attention . The multi-head self-attention has 8 heads with embedding dimension 256 and LayerNorm , following common usage. The DGCNN has two EdgeConv layers with dimensions .. The number of neighbors KK in EdgeConv is 16. In principle, DGCNN is a sparse version of multi-head self-attention; the sparse structure reduces overhead in back-propagation and leads to sharper “attention maps” as well as faster convergence.

Table 7 shows the comparisons: DGCNN consistently outperforms multi-head self-attention. This aligns with our hypothesis: objects are distributed sparsely in the scene, so dense interactions among objects are neither efficient nor effective. Furthermore, we study the effect of number of neighbors in DGCNN. When it is 1, The model reduces to an architecture without object interaction. As we increase the number, it approaches multi-head self-attention. As shown in Table 7, the sweet spot is 16, which appears to balance object interactions and sparsity.

We also investigate improvements introduced when more DGCNNs are stacked in Table 7. This result suggests it is beneficial to incorporate multiple DGCNNs to model the dynamic object relations.

Conclusion

Object DGCNN is a highly-efficient 3D object detector for point clouds. It is able to learn object interactions via dynamic graphs and is optimized through a set-to-set loss, leading to NMS-free detection. The success of Object DGCNN indicates that many post-processing operations in 3D object detection are likely unnecessary and can be replaced with suitable neural network modules. Moreover, we introduce a set-to-set knowledge distillation pipeline enabled by the Object DGCNN. This new pipeline significantly simplifies knowledge distillation for 3D object detection and may be applicable to other tasks like 3D model compression. Beyond the direct usage of our model, our experiments suggest several future directions to address current limitations. For example, our method is initialized with a pre-trained backbone network. Training the model from scratch remains elusive due to the sparse set-to-set supervision; solving this issue may yield improved generalization as in . Furthermore, studying 3D-specific feature extractors will improve the speed and generalizability of 3D object detection. Finally, the large amount of unlabeled data available at training time can serve as another type of privileged information to apply self-supervised learning to 3D domains through set-to-set distillation.

Potential impact. Our method aims to improve the object detection pipeline, which is crucial for the safety of autonomous driving systems. One potential negative impact of our work is that it still lacks theoretical guarantees, similar to many deep learning methods. Future work to improve applicability in this domain might consider challenges of explainability and transparency.

Acknowledgement

The MIT Geometric Data Processing group acknowledges the generous support of Army Research Office grants W911NF2010168 and W911NF2110293, of Air Force Office of Scientific Research award FA9550-19-1-031, of National Science Foundation grants IIS-1838071 and CHS-1955697, from the CSAIL Systems that Learn program, from the MIT–IBM Watson AI Laboratory, from the Toyota–CSAIL Joint Research Center, from a gift from Adobe Systems, from an MIT.nano Immersion Lab/NCSOFT Gaming Program seed grant, and from the Skoltech–MIT Next Generation Program.

References

Checklist

Do the main claims made in the abstract and introduction accurately reflect the paper’s contributions and scope? [Yes]

Did you describe the limitations of your work? [Yes]

Did you discuss any potential negative societal impacts of your work? [Yes]

Have you read the ethics review guidelines and ensured that your paper conforms to them? [Yes]

If you are including theoretical results…

Did you state the full set of assumptions of all theoretical results? [N/A]

Did you include complete proofs of all theoretical results? [N/A]

Did you include the code, data, and instructions needed to reproduce the main experimental results (either in the supplemental material or as a URL)? [Yes]

Did you specify all the training details (e.g., data splits, hyperparameters, how they were chosen)? [Yes]

Did you report error bars (e.g., with respect to the random seed after running experiments multiple times)? [No]

Did you include the total amount of compute and the type of resources used (e.g., type of GPUs, internal cluster, or cloud provider)? [Yes]

If you are using existing assets (e.g., code, data, models) or curating/releasing new assets…

If your work uses existing assets, did you cite the creators? [Yes]

Did you mention the license of the assets? [N/A]

Did you include any new assets either in the supplemental material or as a URL? [No]

Did you discuss whether and how consent was obtained from people whose data you’re using/curating? [N/A]

Did you discuss whether the data you are using/curating contains personally identifiable information or offensive content? [No]

If you used crowdsourcing or conducted research with human subjects…

Did you include the full text of instructions given to participants and screenshots, if applicable? [N/A]

Did you describe any potential participant risks, with links to Institutional Review Board (IRB) approvals, if applicable? [N/A]

Did you include the estimated hourly wage paid to participants and the total amount spent on participant compensation? [N/A]

Supplementary Material

In this section, we provide results when the model is distilled with 1, 6, and 20 epochs. The teacher model and the student model are both constructed with a PointPillars backbone. Table 8 shows the results. In all regimes, the models with distillation improve over their baselines, which verifies the efficacy of the set-to-set distillation.

Time complexity.

We compare the time complexity of PointPillars, Pillar-OD, CenterPoint, and our proposed model. Table 9 shows that our model is more efficient than others at inference time thanks to its NMS-free characteristic. The performance is measured on a single Nvidia RTX 3090.

Visualization.

We provide visualization of predictions by our model in Figure 3. Without NMS, our model makes a sparse set of predictions.