Generalized Intersection over Union: A Metric and A Loss for Bounding Box Regression

Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian Reid, Silvio Savarese

Introduction

IoUIoU, also known as Jaccard index, is the most commonly used metric for comparing the similarity between two arbitrary shapes. IoUIoU encodes the shape properties of the objects under comparison, e.g. the widths, heights and locations of two bounding boxes, into the region property and then calculates a normalized measure that focuses on their areas (or volumes). This property makes IoUIoU invariant to the scale of the problem under consideration. Due to this appealing property, all performance measures used to evaluate for segmentation , object detection , and tracking rely on this metric.

In this paper, we explore the calculation of IoUIoU between two axis aligned rectangles, or generally two axis aligned n-orthotopes, which has a straightforward analytical solution and in contrast to the prevailing belief, IoUIoU in this case can be backpropagated , i.e. it can be directly used as the objective function to optimize. It is therefore preferable to use IoUIoU as the objective function for 2D object detection tasks. Given the choice between optimizing a metric itself vs. a surrogate loss function, the optimal choice is the metric itself. However, IoUIoU as both a metric and a loss has a major issue: if two objects do not overlap, the IoUIoU value will be zero and will not reflect how far the two shapes are from each other. In this case of non-overlapping objects, if IoUIoU is used as a loss, its gradient will be zero and cannot be optimized.

In this paper, we will address this weakness of IoUIoU by extending the concept to non-overlapping cases. We ensure this generalization (a) follows the same definition as IoUIoU, i.e. encoding the shape properties of the compared objects into the region property; (b) maintains the scale invariant property of IoUIoU, and (c) ensures a strong correlation with IoUIoU in the case of overlapping objects. We introduce this generalized version of IoUIoU, named GIoUGIoU, as a new metric for comparing any two arbitrary shapes. We also provide an analytical solution for calculating GIoUGIoU between two axis aligned rectangles, allowing it to be used as a loss in this case. Incorporating GIoUGIoU loss into state-of-the art object detection algorithms, we consistently improve their performance on popular object detection benchmarks such as PASCAL VOC and MS COCO using both the standard, i.e. IoUIoU based , and the new, GIoUGIoU based, performance measures.

The main contribution of the paper is summarized as follows:

We introduce this generalized version of IoUIoU, as a new metric for comparing any two arbitrary shapes.

We provide an analytical solution for using GIoUGIoU as loss between two axis-aligned rectangles or generally n-orthotopesExtension provided in supp. material.

We incorporate GIoUGIoU loss into the most popular object detection algorithms such as Faster R-CNN, Mask R-CNN and YOLO v3, and show their performance improvement on standard object detection benchmarks.

Related Work

Object detection accuracy measures: Intersection over Union (IoUIoU) is the defacto evaluation metric used in object detection. It is used to determine true positives and false positives in a set of predictions. When using IoUIoU as an evaluation metric an accuracy threshold must be chosen. For instance in the PASCAL VOC challenge , the widely reported detection accuracy measure, i.e. mean Average Precision (mAP), is calculated based on a fixed IoUIoU threshold, i.e. 0.50.5. However, an arbitrary choice of the IoUIoU threshold does not fully reflect the localization performance of different methods. Any localization accuracy higher than the threshold is treated equally. In order to make this performance measure less sensitive to the choice of IoUIoU threshold, the MS COCO Benchmark challenge averages mAP across multiple IoUIoU thresholds.

Most popular object detectors utilize some combination of the bounding box representations and losses mentioned above. These considerable efforts have yielded significant improvement in object detection. We show there may be some opportunity for further improvement in localization with the use of GIoUGIoU, as their bounding box regression losses are not directly representative of the core evaluation metric, i.e. IoUIoU.

Optimizing IoUIoU using an approximate or a surrogate function: In the semantic segmentation task, there have been some efforts to optimize IoUIoU using either an approximate function or a surrogate loss . Similarly, for the object detection task, recent works have attempted to directly or indirectly incorporate IoUIoU to better perform bounding box regression. However, they suffer from either an approximation or a plateau which exist in optimizing IoUIoU in non-overlapping cases. In this paper we address the weakness of IoUIoU by introducing a generalized version of IoUIoU, which is directly incorporated as a loss for the object detection problem.

Generalized Intersection over Union

Two appealing features, which make this similarity measure popular for evaluating many 2D/3D computer vision tasks are as follows:

IoUIoU as a distance, e.g. LIoU=1−IoU\mathcal{L}_{IoU}=1-IoU, is a metric (by mathematical definition) . It means LIoU\mathcal{L}_{IoU} fulfills all properties of a metric such as non-negativity, identity of indiscernibles, symmetry and triangle inequality.

If ∣A∩B∣=0|A\cap B|=0, IoU(A,B)=0IoU(A,B)=0. In this case, IoUIoU does not reflect if two shapes are in vicinity of each other or very far from each other.

To address this issue, we propose a general extension to IoUIoU, namely Generalized Intersection over Union GIoUGIoU.

GIoUGIoU as a new metric has the following properties: Their proof has been provided in supp. material.

Similar to IoUIoU, GIoUGIoU as a distance, e.g. LGIoU=1−GIoU\mathcal{L}_{GIoU}=1-GIoU, holding all properties of a metric such as non-negativity, identity of indiscernibles, symmetry and triangle inequality.

Similar to IoUIoU, GIoUGIoU is invariant to the scale of the problem.

Similar to IoUIoU, the value 11 occurs only when two objects overlay perfectly, i.e. if ∣A∪B∣=∣A∩B∣|A\cup B|=|A\cap B|, then GIoU=IoU=1GIoU=IoU=1

GIoUGIoU value asymptotically converges to -1 when the ratio between occupying regions of two shapes, ∣A∪B∣|A\cup B|, and the volume (area) of the enclosing shape ∣C∣|C| tends to zero, i.e. lim⁡∣A∪B∣∣C∣→0GIoU(A,B)=−1\displaystyle\lim_{\frac{|A\cup B|}{|C|}\to 0}GIoU(A,B)=-1 .

In summary, this generalization keeps the major properties of IoUIoU while rectifying its weakness. Therefore, GIoUGIoU can be a proper substitute for IoUIoU in all performance measures used in 2D/3D computer vision tasks. In this paper, we only focus on 2D object detection where we can easily derive an analytical solution for GIoUGIoU to apply it as both metric and loss. The extension to non-axis aligned 3D cases is left as future work.

So far, we introduced GIoUGIoU as a metric for any two arbitrary shapes. However as is the case with IoUIoU, there is no analytical solution for calculating intersection between two arbitrary shapes and/or for finding the smallest enclosing convex object for them.

Fortunately, for the 2D object detection task where the task is to compare two axis aligned bounding boxes, we can show that GIoUGIoU has a straightforward solution. In this case, the intersection and the smallest enclosing objects both have rectangular shapes. It can be shown that the coordinates of their vertices are simply the coordinates of one of the two bounding boxes being compared, which can be attained by comparing each vertices’ coordinates using min and max functions. To check if two bounding boxes overlap, a condition must also be checked. Therefore, we have an exact solution to calculate IoUIoU and GIoUGIoU.

Since back-propagating min, max and piece-wise linear functions, e.g. Relu, are feasible, it can be shown that every component in Alg. 2 has a well-behaved derivative. Therefore, IoUIoU or GIoUGIoU can be directly used as a loss, i.e. LIoU\mathcal{L}_{IoU} or LGIoU\mathcal{L}_{GIoU}, for optimizing deep neural network based object detectors. In this case, we are directly optimizing a metric as loss, which is an optimal choice for the metric. However, in all non-overlapping cases, IoUIoU has zero gradient, which affects both training quality and convergence rate. GIoUGIoU, in contrast, has a gradient in all possible cases, including non-overlapping situations. In addition, using property 3, we show that GIoUGIoU has a strong correlation with IoUIoU, especially in high IoUIoU values. We also demonstrate this correlation qualitatively in Fig. 2 by taking over 10K random samples from the coordinates of two 2D rectangles. In Fig. 2, we also observe that in the case of low overlap, e.g. IoU≤0.2IoU\leq 0.2 and GIoU≤0.2GIoU\leq 0.2, GIoUGIoU has the opportunity to change more dramatically compared to IoUIoU. To this end, GIoUGIoU can potentially have a steeper gradient in any possible state in these cases compared to IoUIoU. Therefore, optimizing GIoUGIoU as loss, LGIoU\mathcal{L}_{GIoU} can be a better choice compared to LIoU\mathcal{L}_{IoU}, no matter which IoUIoU-based performance measure is ultimately used. Our experimental results verify this claim.

Loss Stability: We also investigate if there exist any extreme cases which make the loss unstable/undefined given any value for the predicted outputs.

LGIoU\mathcal{L}_{GIoU} behaviour when IoU = 0: For GIoUGIoU loss, we have LGIoU=1−GIoU=1+Ac−UAc−IoU\mathcal{L}_{GIoU}=1-GIoU=1+\frac{A^{c}-\mathcal{U}}{A^{c}}-IoU. In the case when BgB^{g} and BpB^{p} do not overlap, i.e. I=0\mathcal{I}=0 and IoU=0IoU=0, GIoUGIoU loss simplifies to LGIoU=1+Ac−UAc=2−UAc\mathcal{L}_{GIoU}=1+\frac{A^{c}-\mathcal{U}}{A^{c}}=2-\frac{\mathcal{U}}{A^{c}}. In this case, by minimizing LGIoU\mathcal{L}_{GIoU}, we actually maximize the term UAc\frac{\mathcal{U}}{A^{c}}. This term is a normalized measure between 0 and 1, i.e. 0≤UAc≤10\leq\frac{\mathcal{U}}{A^{c}}\leq 1, and is maximized when the area of the smallest enclosing box AcA^{c} is minimized while the union U=Ag+Ap\mathcal{U}=A^{g}+A^{p}, or more precisely the area of predicted bounding box ApA^{p}, is maximized. To accomplish this, the vertices of the predicted bounding box BpB^{p} should move in a direction that encourages BgB^{g} and BpB^{p} to overlap, making IoU≠0IoU\neq 0.

Experimental Results

All detection baselines have also been evaluated using the test set of the MS COCO 2018 dataset, where the annotations are not accessible for the evaluation. Therefore in this case, we are only able to report results using the standard performance measure, i.e. IoUIoU.

Training protocol. We used the original Darknet implementation of YOLO v3 released by the authors Available at: https://pjreddie.com/darknet/yolo/. For baseline results (training using MSE loss), we used DarkNet-608 as backbone network architecture in all experiments and followed exactly their training protocol using the reported default parameters and the number of iteration on each benchmark. To train YOLO v3 using IoUIoU and GIoUGIoU losses, we simply replace the bounding box regression MSE loss with LIoU\mathcal{L}_{IoU} and LGIoU\mathcal{L}_{GIoU} losses explained in Alg. 2. Considering the additional MSE loss on classification and since we replace an unbounded distance loss such as MSE distance with a bounded distance, e.g. LIoU\mathcal{L}_{IoU} or LGIoU\mathcal{L}_{GIoU}, we need to regularize the new bounding box regression against the classification loss. However, we performed a very minimal effort to regularize these new regression losses against the MSE classification loss.

PASCAL VOC 2007. Following the original code’s training protocol, we trained the network using each loss on both training and validation set of the dataset up to 50K50K iterations. Their performance using the best network model for each loss has been evaluated using the PASCAL VOC 2007 test and the results have been reported in Tab. 1.

Considering both standard IoUIoU based and new GIoUGIoU based performance measures, the results in Tab. 1 show that training YOLO v3 using LGIoU\mathcal{L}_{GIoU} as regression loss can considerably improve its performance compared to its own regression loss (MSE). Moreover, incorporating LIoU\mathcal{L}_{IoU} as regression loss can slightly improve the performance of YOLO v3 on this benchmark. However, the improvement is inferior compared to the case where it is trained by LGIoU\mathcal{L}_{GIoU}.

MS COCO. Following the original code’s training protocol, we trained YOLO v3 using each loss on both the training set and 88% of the validation set of MS COCO 2014 up to 502k502k iterations. Then we evaluated the results using the remaining 12% of the validation set and reported the results in Tab. 2. We also compared them on the MS COCO 2018 Challenge by submitting the results to the COCO server. All results using the IoUIoU based performance measure are reported in Tab. 3.

Similar to the PASCAL VOC experiment, the results show consistent improvement in performance for YOLO v3 when it is trained using LGIoU\mathcal{L}_{GIoU} as regression loss. We have also investigated how each component, i.e. bounding box regression and classification losses, contribute to the final AP performance measure. We believe the localization accuracy for YOLO v3 significantly improves when LGIoU\mathcal{L}_{GIoU} loss is used (Fig. 3 (a)). However, with the current naive tuning of regularization parameters, balancing bounding box loss vs. classification loss, the classification scores may not be optimal, compared to the baseline (Fig. 3 (b)). Since AP based performance measure is considerably affected by small classification error, we believe the results can be further improved with a better search for regularization parameters.

2 Faster R-CNN and Mask R-CNN

PASCAL VOC 2007. Since there is no instance mask annotation available in this dataset, we did not evaluate Mask R-CNN on this dataset. Therefore, we only trained Faster R-CNN using the aforementioned bounding box regression losses on the training set of the dataset for 20k iterations. Then, we searched for the best-performing model on the validation set over different parameters such as the number of training iterations and bounding box regression loss regularizer. The final results on the test set of the dataset have been reported in Tab. 4.

MS COCO. Similarly, we trained both Faster R-CNN and Mask R-CNN using each of the aforementioned bounding box regression losses on the MS COCO 2018 training dataset for 95K iterations. The results for the best model on the validation set of MS COCO 2018 for Faster R-CNN and Mask R-CNN have been reported in Tables 5 and 7 respectively. We have also compared them on the MS COCO 2018 Challenge by submitting their results to the COCO server. All results using the IoUIoU based performance measure are also reported in Tables 6 and 8.

Conclusion

In this paper, we introduced a generalization to IoUIoU as a new metric, namely GIoUGIoU, for comparing any two arbitrary shapes. We showed that this new metric has all of the appealing properties which IoUIoU has while addressing its weakness. Therefore it can be a good alternative in all performance measures in 2D/3D vision tasks relying on the IoUIoU metric.

We also provided an analytical solution for calculating GIoUGIoU between two axis-aligned rectangles. We showed that the derivative of GIoUGIoU as a distance can be computed and it can be used as a bounding box regression loss. By incorporating it into the state-of-the art object detection algorithms, we consistently improved their performance on popular object detection benchmarks such as PASCAL VOC and MS COCO using both the commonly used performance measures and also our new accuracy measure, i.e. GIoUGIoU based average precision. Since the optimal loss for a metric is the metric itself, our GIoUGIoU loss can be used as the optimal bounding box regression loss in all applications which require 2D bounding box regression.

In the future, we plan to investigate the feasibility of deriving an analytic solution for GIoUGIoU in the case of two rotating rectangular cuboids. This extension and incorporating it as a loss could have great potential to improve the performance of 3D object detection frameworks.

References