Improving Object Localization with Fitness NMS and Bounded IoU Loss

Lachlan Tychsen-Smith, Lars Petersson

Introduction

Multiclass object detection is defined as the joint task of localizing bounding boxes of instances and classifying their contents. In this paper, we address the problem of bounding box localization which has become increasingly important as the number of objects in view rises. In particular, by identifying a better fitting bounding box, current methods can better resolve the position, scale and number of unique instances in view. Modern object detection methods most often utilize a CNN based bounding box regression and classification stage followed by a Non-Max Suppression method to identify unique object instances.

Introduced in R-CNN, bounding box regression enables each region of interest (RoI) to estimate an updated bounding box with the goal of better matching the nearest true instance. Prior work has demonstrated that this task can be improved with multiple bounding box regression stages, increasing the number of (or carefully selecting) RoI anchors in the region proposal network, and increasing the input image resolution (or using image pyramids). Alternatively, the DeNet method demonstrated enhanced precision by improving the localization of the sampling RoIs before bounding box regression. This was achieved via a novel corner-based RoI estimation method, replacing the typical region proposal network (RPN) used in many other methods. This approach operates in real-time, demonstrated improved fine object localization and has the additional benefit of not requiring user-defined anchor bounding boxes. Despite only demonstrating our results with DeNet, we believe the novelties presented are equally applicable to other detectors.

Experimental Setup

Here we introduce some important parameters and definitions used throughout the paper:

Sampling Region-of-Interest (RoI): a bounding box generated before classification. Only applicable to two-stage detectors (e.g., R-CNN variants, etc).

Intersection-over-Union (IoU): the intersection area divided by the union area for two bounding boxes.

2 Datasets, Training and Testing

For validation experiments, we combine Pascal VOC 2007 trainval and VOC 2012 trainval datasets to form the training data and test on Pascal VOC 2007 test. For testing, we train on MSCOCO trainval and use the test-dev dataset for evaluation. Following DeNet and SSD, all timing results are provided for an Nvidia Titan X (Maxwell) GPU with cuDNN v5.1 and a batch size of 8. Furthermore, unless stated otherwise, the same learning schedule, hyper-parameters, augmentation methods, etc were used as in the DeNet paper.

Fitness Non-Max Suppression

In this section, we highlight a flaw in the Non-Max Supression based instance estimation, namely, the reliance on a single matching IoU. Following this analysis, we propose and demonstrate a novel Fitness NMS algorithm.

2 Detection Clustering

To indicate the score of a bounding box bjb_{j}, many detection models (including DeNet) apply:

3 Novel Fitness NMS Method

To address the issues described in the previous sections we propose augmenting Equation 1 with an additional expected fitness term:

In our implementation, the fitness fjf_{j} can take on FF values (F=5F=5 in this paper) and is mapped via:

where 0≤ρj≤10\leq\rho_{j}\leq 1 is the maximum IoU overlap between bjb_{j} and the set of groundtruth bounding boxes BTB_{T}. If ρj\rho_{j} is less than 0.5, the bounding box bjb_{j} is assigned the null class with no associated fitness. From these definitions the expected value is given by:

The Joint Fitness NMS method demonstrates a clear lead at high matching IoUs with no observable loss for low matching IoUs. In Figure 3, we provide the recall for the Joint Fitness NMS method and the recall delta between Joint Fitness NMS and Baseline for the same model. These results directly demonstrate the improved recall at various operating points obtained by the Fitness NMS methods at fine localization accuracies with negligible losses for coarse localization.

2 Bounded IoU Loss

Here we propose a novel bounding box loss and compare it to the R-CNN method used so far. This new loss aims to maximise the IoU overlap between the RoI and the associated groundtruth bounding box, while providing good convergence properties for gradient descent optimizers.

Given a sampling RoI bs=(xs,ys,ws,hs)b_{s}=(x_{s},y_{s},w_{s},h_{s}), an associated groundtruth target bt=(xt,yt,wt,ht)b_{t}=(x_{t},y_{t},w_{t},h_{t}), and an estimated bounding box β=(x,y,w,h)\beta=(x,y,w,h) the widely used R-CNN formulation provides the following cost functions:

where Δx=x−xt\Delta x=x-x_{t} and L1(z)L_{1}(z) is the Huber Loss (also known as smooth L1 loss). Note that we restrict this analysis to the X position and width for the sake of brevity, the Y position and height equations can be identified with suitable substitutions. The Huber loss is defined by:

In Table 7, we compare the corner clustering method, standard NMS method (with a 0.7 threshold), and no clustering when the number of RoIs is reduced to 576 for the DeNet wide models. The results demonstrate a 30% to 80% improvement in evaluation rate due to the decreased number of RoI classifications needed. We found both clustering methods improved upon no clustering and obtained very similar MAP@[0.5:0.95] results, however, the standard NMS method was slightly slower due to an increased CPU load. Relative to the NMS method, the corner clustering method appears to perform slightly better at high matching IoU and worse at low matching IoUs.

Input Image Scaling

So far our model has been demonstrated at low input image size, i.e., rescaling the largest input dimension to 512512 pixels. In comparison, state-of-the-art detectors typically use an input image with the smallest side resized to 600 or 800 pixels. This smaller input image has provided our model with an improved evaluation rate at the cost of object localization accuracy. In the following, we relax the constraint on evaluation rate to demonstrate localization accuracies when computational resources are less constrained.

In Table 8, we provide MSCOCO results for the DeNet-101 (wide) model with Fitness NMS, Bounded IoU Loss and Corner Clustering with the largest input image dimension varied from 384 to 1536. Note that the model was not retrained with these settings, only tested. Since the benchmarked methods use different input image scaling methods, we provide the mean input pixels per sample (ignoring black borders) calculated over the MSCOCO test-dev dataset. These results demonstrate that input image size is very important for small and medium sized object localization in MSCOCO, however, we observed a loss in precision for large objects as scale is increased. This precision asymmetry suggests that multi-scale evaluation is likely optimal with the current design.

Conclusion

We highlighted an issue in common detector designs and propose a novel Fitness NMS method to address it. This method significantly improves MAP at high localization accuracies without a loss in evaluation rate. Following this we derive a novel bounding box loss better suited to IoU maximisation while still providing convergence properties suitable to gradient descent. Combining these results with a simple RoI clustering method, we obtain highly competitive MAP vs Inference Times on MSCOCO (see Figure 6). These results highlight that with these modification a two-stage detector can be made highly competitive with single-stage methods in terms of MAP vs inference time. Though not demonstrated empirically in this paper, we believe the novelties presented in this paper are equally applicable to other bounding box detector designs.

References