GrooMeD-NMS: Grouped Mathematically Differentiable NMS for Monocular 3D Object Detection
Abhinav Kumar, Garrick Brazil, Xiaoming Liu
Introduction
D object detection is one of the fundamental problems in computer vision, where the task is to infer D information of the object. Its applications include augmented reality , robotics , medical surgery , and, more recently path planning and scene understanding in autonomous driving . Most of the D object detectors are extensions of the D object detector Faster R-CNN , which relies on the end-to-end learning idea to achieve State-of-the-Art (SoTA) object detection. Some of these methods have proposed changing architectures or losses . Others have tried incorporating confidence or temporal cues .
Almost all of them output a massive number of boxes for each object and, thus, rely on post-processing with a greedy clustering algorithm called Non-Maximal Suppression (NMS) during inference to reduce the number of false positives and increase performance. However, these works have largely overlooked NMS’s inclusion in training leading to an apparent mismatch between training and inference pipelines as the losses are applied on all boxes before NMS but not on final boxes after NMS (see Fig. 1(a)).We also find that D object detection suffers a greater mismatch between classification and D localization compared to that of D localization, as discussed further in Sec. A3.2 of the supplementary and observed in . Hence, our focus is D object detection.
Earlier attempts to include NMS in the training pipeline have been made for D object detection where the improvements are less visible. Recent efforts to improve the correlation in D object detection involve calculating or predicting the scores via likelihood estimation or enforcing the correlation explicitly . Although this improves the D detection performance, improvements are limited as their training pipeline is not end to end in the absence of a differentiable NMS.
To address the mismatch between training and inference pipelines as well as the mismatch between classification and D localization, we propose including the NMS in the training pipeline, which gives a useful gradient to the network so that it figures out which boxes are the best-localized in D and, therefore, should be ranked higher (see Fig. 1(b)).
An ideal NMS for inclusion in the training pipeline should be not only differentiable but also parallelizable. Unfortunately, the inference-based classical NMS and Soft-NMS are greedy, set-based and, therefore, not parallelizable . To make the NMS parallelizable, we first formulate the classical NMS as matrix operation and then obtain a closed-form mathematical expression using elementary matrix operations such as matrix multiplication, matrix inversion, and clipping. We then replace the threshold pruning in the classical NMS with its softer version to get useful gradients. These two changes make the NMS GPU-friendly, and the gradients are backpropagated. We next group and mask the boxes in an unsupervised manner, which removes the matrix inversion and simplifies our proposed differentiable NMS expression further. We call this NMS as Grouped Mathematically Differentiable Non-Maximal Suppression (GrooMeD-NMS).
In summary, the main contributions of this work include:
This is the first work to propose and integrate a closed-form mathematically differentiable NMS for object detection, such that the network is trained end-to-end with a loss on the boxes after NMS.
We propose an unsupervised grouping and masking on the boxes to remove the matrix inversion in the closed-form NMS expression.
We achieve SoTA monocular D object detection performance on the KITTI dataset performing comparably to monocular video-based methods.
Related Work
3D Object Detection. Recent success in D object detection has inspired people to infer D information from a single D (monocular) image. However, the monocular problem is ill-posed due to the inherent scale/depth ambiguity . Hence, approaches use additional sensors such as LiDAR , stereo or radar . Although LiDAR depth estimations are accurate, LiDAR data is sparse and computationally expensive to process . Moreover, LiDARs are expensive and do not work well in severe weather .
Hence, there have been several works on monocular D object detection. Earlier approaches use hand-crafted features, while the recent ones are all based on deep learning. Some of these methods have proposed changing architectures or losses . Others have tried incorporating confidence , augmentation , depth in convolution or temporal cues . Our work proposes to incorporate NMS in the training pipeline of monocular D object detection.
Non-Maximal Suppression. NMS has been used to reduce false positives in edge detection , feature point detection , face detection , human detection as well as SoTA D and D detection . Modifications to NMS in D detection , D pedestrian detection , D salient object detection and D detection can be classified into three categories – inference NMS , optimization-based NMS and neural network based NMS .
The inference NMS changes the way the boxes are pruned in the final set of predictions. uses weighted averaging to update the -coordinate after NMS. solves quadratic unconstrained binary optimization while and use point processes and MAP based inference respectively. and formulate NMS as a structured prediction task for isolated and all object instances respectively. The neural network NMS use a multi-layer network and message-passing to approximate NMS or to predict the NMS threshold adaptively . approximates the sub-gradients of the network without modelling NMS via a transitive relationship. Our work proposes a grouped closed-form mathematical approximation of the classical NMS and does not require multiple layers or message-passing. We detail these differences in Sec. 4.2.
Background
Let denote the set of boxes or proposals from an image. Let and denote their scores (before NMS) and rescores (updated scores after NMS) respectively such that . denotes the subset of after the NMS. Let denote the matrix with denoting the D Intersection over Union (IoU) of and . The pruning function decides how to rescore a set of boxes based on IoU overlaps of its neighbors, sometimes suppressing boxes entirely. In other words, denotes the box is suppressed while denotes is kept in . The NMS threshold is the threshold for which two boxes need in order for the non-maximum to be suppressed. The temperature controls the shape of the exponential and sigmoidal pruning functions . thresholds the rescores in GrooMeD and Soft-NMS to decide if the box remains valid after NMS.
denotes the logical OR while denotes clipping of in the range $$. Formally,
2 Classical and Soft-NMS
NMS is one of the building blocks in object detection whose high-level goal is to iteratively suppress boxes which have too much IoU with a nearby high-scoring box. We first give an overview of the classical and Soft-NMS , which are greedy and used in inference. Classical NMS uses the idea that the score of a box having a high IoU overlap with any of the selected boxes should be suppressed to zero. That is, it uses a hard pruning without any temperature . Soft-NMS makes this pruning soft via temperature . Thus, classical and Soft-NMS only differ in the choice of . We reproduce them in Alg. 1 using our notations.
GrooMeD-NMS
Classical NMS (Alg. 1) uses and greedily calculates the rescore of boxes and, is thus not parallelizable or differentiable . We wish to find its smooth approximation in closed-form for including in the training pipeline.
Classical NMS uses the non-differentiable hard operation (Line of Alg. 1). We remove the by hard sorting the scores and in decreasing order (lines - of Alg. 2). We also try making the sorting soft. Note that we require the permutation of to sort . Most soft sorting methods apply the soft permutation to the same vector. Only two other methods can apply the soft permutation to another vector. Both methods use computations for soft sorting . We implement and find that is overly dependent on temperature to break out the ranks, and its gradients are too unreliable to train our model. Hence, we stick with the hard sorting of and .
1.2 NMS as a Matrix Operation
The rescoring process of the classical NMS is greedy set-based and only considers overlaps with unsuppressed boxes. We first generalize this rescoring by accounting for the effect of all (suppressed and unsuppressed) boxes as
using the relaxation of logical OR operator as . See Sec. A1 of the supplementary material for an alternate explanation of (2). The presence of on the RHS of (2) prevents suppressed boxes from influencing other boxes hugely. When outputs discretely as as in classical NMS, scores are guaranteed to be suppressed to or left unchanged thereby implying . We write the rescores in a matrix formulation as
The above two equations are written compactly as
as the solution to (5) with being the identity matrix. Intuitively, if the matrix inversion is considered division in (6) and the boxes have overlaps, the rescores are the scores divided by a number greater than one and are, therefore, lesser than scores. If the boxes do not overlap, the division is by one and rescores equal scores.
Note that the in (6) is a lower triangular matrix with ones on the principal diagonal. Hence, is always full rank and, therefore, always invertible.
1.3 Grouping
We next observe that the object detectors output multiple boxes for an object, and a good detector outputs boxes wherever it finds objects in the monocular image. Thus, we cluster the boxes in an image in an unsupervised manner based on IoU overlaps to obtain the groups . Grouping thus mimics the grouping of the classical NMS, but does not rescore the boxes. As clustering limits interactions to intra-group interactions among the boxes, we write (6) as
This results in taking smaller matrix inverses in (7) than (6).
We use a simplistic grouping algorithm, i.e., we form a group with boxes having high IoU overlap with the top-ranked box, given that we sorted the scores. As the group size is limited by , we choose a minimum of and the number of boxes in . We next delete all the boxes of this group and iterate until we run out of boxes. Also, grouping uses IoU since we can achieve meaningful clustering in D. We detail this unsupervised grouping in Alg. 3.
1.4 Masking
Classical NMS considers the IoU of the top-scored box with other boxes. This consideration is equivalent to only keeping the column of corresponding to the top box while assigning the rest of the columns to be zero. We implement this through masking of . Let denote the binary mask corresponding to group . Then, entries in the binary matrix in the column corresponding to the top-scored box are and the rest are . Hence, only one of the columns in is non-zero. Now, is a Frobenius matrix (Gaussian transformation) and we, therefore, invert this matrix by simply subtracting the second term . In other words, . Hence, we simplify (7) further to get
Thus, masking allows to bypass the computationally expensive matrix inverse operation altogether.
We call the NMS based on (8) as Grouped Mathematically Differentiable Non-Maximal Suppression or GrooMeD-NMS. We summarize the complete GrooMeD-NMS in Alg. 2 and show its block-diagram in Fig. 1(c). GrooMeD-NMS in Fig. 1(c) provides two gradients - one through and other through .
1.5 Pruning Function
As explained in Sec. 3.1, the pruning function decides whether to keep the box in the final set of predictions or not based on IoU overlaps, i.e., denotes the box is suppressed while denotes is kept in .
Classical NMS uses the threshold as the pruning function, which does not give useful gradients. Therefore, we considered three different functions for : Linear, a temperature -controlled Exponential, and Sigmoidal function.
Linear Linear pruning function is .
Sigmoidal Sigmoidal pruning function is with denoting the standard sigmoid. Sigmoidal function appears as the binary cross entropy relaxation of the subset selection problem .
We show these pruning functions in Fig. 2. The ablation studies (Sec. 5.4) show that choosing as Linear yields the simplest and the best GrooMeD-NMS.
2 Differences from Existing NMS
Although no differentiable NMS has been proposed for the monocular D object detection, we compare our GrooMeD-NMS with the NMS proposed for D object detection, D pedestrian detection, D salient object detection, and D object detection in Tab. 1. No method described in Tab. 1 has a matrix-based closed-form mathematical expression of the NMS. Classical, Soft and Distance-NMS are used at the inference time, while GrooMeD-NMS is used during both training and inference. Distance-NMS updates the -coordinate of the box after NMS as the weighted average of the -coordinates of top- boxes. QUBO-NMS , Point-NMS , and MAP-NMS are not used in end-to-end training. proposes a trainable Point-NMS. The Structured-SVM based NMS rely on structured SVM to obtain the rescores. Adaptive-NMS uses a separate neural network to predict the classical NMS threshold . The trainable neural network based NMS (NN-NMS) use a separate neural network containing multiple layers and/or message-passing to approximate the NMS and do not use the pruning function. Unlike these methods, GrooMeD-NMS uses a single layer and does not require multiple layers or message passing. Our NMS is parallel up to group (denoted by ). However, is, in general, in the NMS.
3 Target Assignment and Loss Function
Target Assignment. Our method consists of M3D-RPN and uses binning and self-balancing confidence . The boxes’ self-balancing confidence are used as scores , which pass through the GrooMeD-NMS layer to obtain the rescores . The rescores signal the network if the best box has not been selected for a particular object.
We extend the notion of the best D box to D. The best box has the highest product of IoU and gIoU with ground truth . If the product is greater than a certain threshold , it is assigned a positive label. Mathematically,
with q(b_{j},g_{l})=\text{IoU{}_{2\text{D}}}(b_{j},g_{l})~{}\left(\frac{1+\text{gIoU{}_{3\text{D}}}(b_{j},g_{l})}{2}\right). gIoU is known to provide signal even for non-intersecting boxes , where the usual IoU is always zero. Therefore, we use gIoU instead of regular IoU for figuring out the best box in D as many D boxes have a zero IoU overlap with the ground truth. For calculating gIoU, we first calculate the volume and hull volume of the D boxes. is the product of gIoU in Birds Eye View (BEV), removing the rotations and hull of the dimension. gIoU is then given by
Loss Function. Generally the number of best boxes is less than the number of ground truths in an image, as there could be some ground truth boxes for which no box is predicted. The tiny number of best boxes introduces a far-heavier skew than the foreground-background classification. Thus, we use the modified AP-Loss as our loss after NMS since AP-Loss does not suffer from class imbalance .
Vanilla AP-Loss treats boxes of all images in a mini-batch equally, and the gradients are back-propagated through all the boxes. We remove this condition and rank boxes in an image-wise manner. In other words, if the best boxes are correctly ranked in one image and are not in the second, then the gradients only affect the boxes of the second image. We call this modification of AP-Loss as Imagewise AP-Loss. In other words,
where and denote the rescores and the boxes of the image in a mini-batch respectively. This is different from previous NMS approaches , which use classification losses. Our ablation studies (Sec. 5.4) show that the Imagewise AP-Loss is better suited to be used after NMS than the classification loss.
Our overall loss function is thus given by where denotes the losses before the NMS including classification, D and D regression as well as confidence losses, and denotes the loss term after the NMS, which is the Imagewise AP-Loss with being the weight. See Sec. A2 of the supplementary material for more details of the loss function.
Experiments
Our experiments use the most widely used KITTI autonomous driving dataset . We modify the publicly-available PyTorch code of Kinematic-3D . uses DenseNet-121 trained on ImageNet as the backbone and using D-RPN settings of . As is a video-based method while GrooMeD-NMS is an image-based method, we use the best image model of henceforth called Kinematic (Image) as our baseline for a fair comparison. Kinematic (Image) is built on M3D-RPN and uses binning and self-balancing confidence.
Data Splits. There are three commonly used data splits of the KITTI dataset; we evaluate our method on all three.
Test Split: Official KITTI D benchmark consists of training and testing images .
Val 1 Split: It partitions the training images into training and validation images .
Val 2 Split: It partitions the training images into training and validation images .
Training. Training is done in two phases - warmup and full . We initialize the model with the confidence prediction branch from warmup weights and finetune using the self-balancing loss and Imagewise AP-Loss after our GrooMeD-NMS. See Sec. A3.1 of the supplementary material for more training details. We keep the weight at . Unless otherwise stated, we use as the Linear function (this does not require ) with . and are set to , and respectively.
Inference. We multiply the class and predicted confidence to get the box’s overall score in inference as in . See Sec. 5.2 for training and inference times.
Evaluation Metrics. KITTI uses AP metric to evaluate object detection following . KITTI benchmark evaluates on three object categories: Easy, Moderate and Hard. It assigns each object to a category based on its occlusion, truncation, and height in the image space. The AP performance on the Moderate category compares different models in the benchmark . We focus primarily on the Car class following .
Tab. 2 summarizes the results of D object detection and BEV evaluation on KITTI Test Split. The results in Tab. 2 show that GrooMeD-NMS outperforms the baseline M3D-RPN by a significant margin and several other SoTA methods on both the tasks. GrooMeD-NMS also outperforms augmentation based approach MoVi-3D and depth-convolution based D4LCN . Despite being an image-based method, GrooMeD-NMS performs competitively to the video-based method Kinematic (Video) , outperforming it on the most-challenging Hard set.
2 KITTI Val 1 3D Object Detection
Results. Tab. 3 summarizes the results of D object detection and BEV evaluation on KITTI Val 1 Split at two IoU thresholds of and . The results in Tab. 3 show that GrooMeD-NMS outperforms the baseline of M3D-RPN and Kinematic (Image) by a significant margin. Interestingly, GrooMeD-NMS (an image-based method) also outperforms the video-based method Kinematic (Video) on most of the metrics. Thus, GrooMeD-NMS performs best on out of the cases ( categories tasks thresholds) while second-best on all other cases. The performance is especially impressive since the biggest improvements are shown on the Moderate and Hard set, where objects are more distant and occluded.
AP at different depths and IoU thresholds. We next compare the AP performance of GrooMeD-NMS and Kinematic (Image) on linear and log scale for objects at different depths of ${}_{3\text{D}}0.3\!\relbar\joinrel\mathrel{\RHD}\!0.7{}_{3\text{D}}$ thresholds.
Comparisons with other NMS. We compare with the classical NMS, Soft-NMS and Distance-NMS in Tab. 4. More detailed results are in Tab. 8 of the supplementary material. The results show that NMS inclusion in the training pipeline benefits the performance, unlike , which suggests otherwise. Training with GrooMeD-NMS helps because the network gets an additional signal through the GrooMeD-NMS layer whenever the best-localized box corresponding to an object is not selected. Interestingly, Tab. 4 also suggests that replacing GrooMeD-NMS with the classical NMS in inference does not affect the performance.
Score-IoU Plot. We further correlate the scores with IoU after NMS of our model with two baselines - M3D-RPN and Kinematic (Image) and also the Kinematic (Video) in Fig. 4. We obtain the best correlation of exceeding the correlations of M3D-RPN, Kinematic (Image) and, also Kinematic (Video). This proves that including NMS in the training pipeline is beneficial.
Training and Inference Times. We now compare the training and inference times of including GrooMeD-NMS in the pipeline. Warmup training phase takes about hours to train on a single GB GeForce GTX Titan-X GPU. Full training phase of Kinematic (Image) and GrooMeD-NMS takes about and hours respectively. The inference time per image using classical and GrooMeD-NMS is and ms respectively. Tab. 4 suggests that changing the NMS from GrooMeD to classical during inference does not alter the performance. Then, the inference time of our method is the same as ms.
3 KITTI Val 2 3D Object Detection
Tab. 5 summarizes the results of D object detection and BEV evaluation on KITTI Val 2 Split at two IoU thresholds of and . Again, we use M3D-RPN and Kinematic (Image) as our baselines. We evaluate the released model of M3D-RPN using the KITTI metric. does not report Val 2 results, so we retrain on Val 2 using their public code. The results in Tab. 5 show that GrooMeD-NMS performs best in all cases. This is again impressive because the improvements are shown on Moderate and Hard set, consistent with Tabs. 2 and 3.
4 Ablation Studies
Tab. 6 compares the modifications of our approach on KITTI Val 1 Cars. Unless stated otherwise, we stick with the experimental settings described in Sec. 5. Using a confidence head (Conf+No NMS) proves beneficial compared to the warmup model (No Conf+No NMS), which is consistent with the observations of . Further, GrooMeD-NMS on classification scores (denoted by No Conf + NMS) is detrimental as the classification scores are not suited for localization . Training the warmup model and then finetuning also works better than training without warmup as in since the warmup phase allows GrooMeD-NMS to carry meaningful grouping of the boxes.
As described in Sec. 4.1.5, in addition to Linear, we compare two other functions for pruning function : Exponential and Sigmoidal. Both of them do not perform as well as the Linear possibly because they have vanishing gradients close to overlap of zero or one. Grouping and masking both help our model to reach a better minimum. As described in Sec. 4.3, Imagewise AP loss is better than the Vanilla AP loss since it treats boxes of two images differently. Imagewise AP also performs better than the binary cross-entropy (BCE) loss proposed in . Using the product of self-balancing confidence and classification scores instead of using them individually as the scores to the NMS in inference is better, consistent with . Class confidence performs worse since it does not have the localization information while the self-balancing confidence (Pred) gives the localization without considering whether the box belongs to foreground or background.
Conclusions
In this paper, we present and integrate GrooMeD-NMS – a novel Grouped Mathematically Differentiable NMS for monocular D object detection, such that the network is trained end-to-end with a loss on the boxes after NMS. We first formulate NMS as a matrix operation and then do unsupervised grouping and masking of the boxes to obtain a simple closed-form expression of the NMS. GrooMeD-NMS addresses the mismatch between training and inference pipelines and, therefore, forces the network to select the best D box in a differentiable manner. As a result, GrooMeD-NMS achieves state-of-the-art monocular D object detection results on the KITTI benchmark dataset. Although our implementation demonstrates monocular D object detection, GrooMeD-NMS is fairly generic for other object detection tasks. Future work includes applying this method to tasks such as LiDAR-based D object detection and pedestrian detection.
References
GrooMeD-NMS: Grouped Mathematically Differentiable NMS for Monocular 3D Object Detection Supplementary Material
Appendix A1 Detailed Explanation of NMS as a Matrix Operation
The rescoring process of the classical NMS is greedy set-based and calculates the rescore for a box (Line of Alg. 1) as
where is defined as the box indices sampled from having higher scores than box . For example, let us consider that . Then, for while for with denoting the empty set. This is possible since we had sorted the scores and in decreasing order (Lines - of Alg. 2) to remove the non-differentiable hard operation of the classical NMS (Line of Alg. 1).
Classical NMS only takes the overlap with unsuppressed boxes into account. Therefore, we generalize (15) by accounting for the effect of all (suppressed and unsuppressed) boxes as
The presence of on the RHS of (16) prevents suppressed boxes from influencing other boxes hugely. Let us say we have a box with a high overlap with an unsuppressed box . The classical NMS with a threshold pruning function assigns while (16) assigns a small non-zero value with a threshold pruning.
Although (16) keeps , getting a closed-form recursion in is not easy because of the product operation. To get a closed-form recursion with addition/subtraction in , we first carry out the polynomial multiplication and then ignore the higher-order terms as
Dropping the in the second term of (17) helps us get a cleaner form of (22). Moreover, it does not change the nature of the NMS since the subtraction keeps the relation intact as and are both between $$.
We can also reach (17) directly as follows. Classical NMS suppresses a box which has a high IoU overlap with any of the unsuppressed boxes () to zero. We consider any as a logical non-differentiable OR operation and use logical OR operator’s differentiable relaxation as . We next use this relaxation with the other expression .
When a box shows overlap with more than two unsuppressed boxes, the term in (17) or when a box shows high overlap with one unsuppressed box, the term . In both of these cases, . So, we lower bound (17) with a operation to ensure that . Thus,
We write the rescores in a matrix formulation as
We next write the above two equations compactly as
However, for a differentiable NMS layer, we need to avoid the recursion. Therefore, we first solve (21) assuming the operation is not present which gives us the solution . In general, this solution is not necessarily bounded between and . Hence, we clip it explicitly to obtain the approximation
Appendix A2 Loss Functions
We now detail out the loss functions used for training. The losses on the boxes before NMS, , is given by
is the predicted self-balancing confidence of each box , while and are its orientation bins . denotes the ground-truth. is the rolling mean of most recent losses per mini-batch , while denotes the weight of the orientation bins loss. CE and Smoooth-L1 denote the Cross Entropy and Smooth L1 loss respectively. Note that we apply D and D regression losses as well as the confidence losses only on the foreground boxes.
As explained in Sec. 4.3, the loss on the boxes after NMS, , is the Imagewise AP-Loss, which is given by
Let be the weight of the term. Then, our overall loss function is given by
We keep following and . Clearly, all our losses and their weights are identical to except .
Appendix A3 Additional Experiments and Results
We now provide additional details and results evaluating our system’s performance.
Training images are augmented using random flipping with probability . Adam optimizer is used with batch size , weight-decay and gradient clipping of . Warmup starts with a learning rate following a poly learning policy with power . Warmup and full training phases take and mini-batches respectively for Val 1 and Val 2 Splits while take and mini-batches for Test Split.
A3.2 KITTI Val 1 Oracle NMS Experiments
As discussed in Sec. 1, to understand the effects of an inference-only NMS on D and D object detection, we conduct a series of oracle experiments. We create an oracle NMS by taking the Val Car boxes of KITTI Val 1 Split from the baseline Kinematic (Image) model before NMS and replace their scores with their true IoU or IoU with the ground-truth, respectively. Note that this corresponds to the oracle because we do not know the ground-truth boxes during inference. We then pass the boxes with the oracle scores through the classical NMS and report the results in Tab. 7.
The results show that the AP increases by a staggering AP on Mod cars when we use oracle IoU as the NMS score. On the other hand, we only see an increase in AP by AP on Mod cars when we use oracle IoU as the NMS score. Thus, the relative effect of using oracle IoU NMS scores on D detection is more significant than using oracle IoU NMS scores on D detection. In other words, the mismatch is greater between classification and D localization compared to the mismatch between classification and D localization.
A3.3 KITTI Val 1 3D Object Detection
Comparisons with other NMS. We compare our method with the other NMS—classical, Soft and Distance-NMS and report the detailed results in Tab. 8. We use the publicly released Soft-NMS code and Distance-NMS code from the respective authors. The Distance-NMS model uses the class confidence scores divided by the uncertainty in (the most erroneous dimension in D localization ) of a box as the Distance-NMS input. Our model does not predict the uncertainty in of a box but predicts its self-balancing confidence (the D localization score). Therefore, we use the class confidence scores multiplied by the self-balancing confidence as the Distance-NMS input.
The results in Tab. 8 show that NMS inclusion in the training pipeline benefits the performance, unlike , which suggests otherwise. Training with GrooMeD-NMS helps because the network gets an additional signal through the GrooMeD-NMS layer whenever the best-localized box corresponding to an object is not selected. Moreover, Tab. 8 suggests that we can replace GrooMeD-NMS with the classical NMS in inference as the performance is almost the same even at IoU.
How good is the classical NMS approximation? GrooMeD-NMS uses several approximations to arrive at the matrix solution (22). We now compare how good these approximations are with the classical NMS. Interestingly, Tab. 8 shows that GrooMeD-NMS is an excellent approximation to the classical NMS as the performance does not degrade after changing the NMS in inference.
A3.4 KITTI Val 1 Sensitivity Analysis
There are a few adjustable parameters for the GrooMeD-NMS, such as the NMS threshold , valid box threshold , the maximum group size , the weight for the , and . We carry out a sensitivity analysis to understand how these parameters affect performance and speed, and how sensitive the algorithm is to these parameters.
Sensitivity to NMS Threshold. We show the sensitivity to NMS threshold in Tab. 9. The results in Tab. 9 show that the optimal . This is also the in .
Sensitivity to Valid Box Threshold. We next show the sensitivity to valid box threshold in Tab. 10. Our choice of performs close to the optimal choice.
Sensitivity to Maximum Group Size. Grouping has a parameter group size . We vary this parameter and report AP and AP at two different IoU thresholds on Moderate Cars of KITTI Val 1 Split in Fig. 5. We note that the best AP performance is obtained at and we, therefore, set in our experiments.
Sensitivity to Loss Weight. We now show the sensitivity to loss weight in Tab. 11. Our choice of is the optimal value.
Sensitivity to Best Box Threshold. We now show the sensitivity to the best box threshold in Tab. 12. Our choice of is the optimal value.
Conclusion. Our method has minor sensitivity to and , which is common in object detection. Our method is not as sensitive to since it only decides a box’s validity. Our parameter choice is either at or close to the optimal. The inference speed is only affected by . Other parameters are used in training or do not affect inference speed.
A3.5 Qualitative Results
We next show some qualitative results of models trained on KITTI Val 1 Split in Fig. 6. We depict the predictions of GrooMeD-NMS in image view on the left and the predictions of GrooMeD-NMS, Kinematic (Image) , and ground truth in BEV on the right. In general, GrooMeD-NMS predictions are more closer to the ground truth than Kinematic (Image) .
A3.6 Demo Video of GrooMeD-NMS
We next include a short demo video of our GrooMeD-NMS model trained on KITTI Val 1 Split. We run our trained model independently on each frame of the three KITTI raw sequences - 2011_10_03_drive_0047, 2011_09_29_drive_0026 and 2011_09_26_drive_0009. None of the frames from these three raw sequences appear in the training set of KITTI Val 1 Split. We use the camera matrices available with the raw sequences but do not use any temporal information. Overlaid on each frame of the raw input videos, we plot the projected D boxes of the predictions and also plot these D boxes in the BEV. We set the frame rate of this demo at fps. The demo is also available in HD at https://www.youtube.com/watch?v=PWctKkyWrno. In the demo video, notice that the orientation of the boxes are stable despite not using any temporal information.
Acknowledgements
This research was partially sponsored by Ford Motor Company and the Army Research Office under Grant Number W911NF-18-1-0330. This document’s views and conclusions are those of the authors and do not represent the official policies, either expressed or implied, of the Army Research Office or the U.S. Government.
We thank Mathieu Blondel and Quentin Berthet from Google Brain, Paris, for several useful discussions on differentiable ranking and sorting. We also discussed the logical operators’ relaxation with Ashim Gupta from the University of Utah. Armin Parchami from Ford Motor Company suggested the learnable NMS paper . Enrique Corona and Marcos Paul Gerardo Castro from Ford Motor Company provided feedback during the development of this work. Shengjie Zhu from the Computer Vision Lab at Michigan State University proof-read our manuscript and suggested several changes. We also thank Xuepeng Shi from University College London for sharing the Distance-NMS code for bench-marking. We finally acknowledge anonymous reviewers for their feedback that helped in shaping the final manuscript.