Boundary IoU: Improving Object-Centric Image Segmentation Evaluation

Bowen Cheng, Ross Girshick, Piotr Dollár, Alexander C. Berg, Alexander Kirillov

Introduction

The Common Task Framework , in which standardized tasks, datasets, and evaluation metrics are used to track research progress, yields impressive results. For example, researchers working on the instance segmentation task, which requires an algorithm to delineate objects with pixel-level binary masks, have improved the standard Average Precision (AP) metric on COCO by an astonishing 86% (relative) from 2015 to 2019 .

However, this progress is not equal across all error modes, because different evaluation metrics are sensitive (or insensitive) to different types of errors. If a metric is used for a prolonged time, as in the Common Task Framework, then the corresponding sub-field most rapidly resolves the types of errors to which this metric is sensitive. Research directions that improve other error types typically advance more slowly, as such progress is harder to quantify.

This phenomenon is at play in instance segmentation, where, among the multitude of papers contributing to the impressive 86% relative improvement in AP (e.g\onedot, ), only a few address mask boundary quality.

Note that mask boundary quality is an essential aspect of image segmentation, as various downstream applications directly benefit from more precise object segmentations . However, the dominant family of Mask R-CNN-based methods are well-known to predict low-fidelity, blobby masks (see Figure 1). This observation suggests that the current evaluation metrics may have limited sensitivity to mask prediction errors near object boundaries.

To understand why, we start by analyzing Mask Intersection-over-Union (Mask IoU), the underlying measure used in AP to compare predicted and ground truth masks. Mask IoU divides the intersection area of two masks by the area of their union. This measure values all pixels equally and, therefore, is less sensitive to boundary quality in larger objects: the number of interior pixels grows quadratically in object size and can far exceed the number of boundary pixels, which only grows linearly. In this paper we aim to identify a measure for image segmentation that is sensitive to boundary quality across all scales.

Towards this goal we start by studying standard segmentation measures like Mask IoU and boundary-focused measures such as Trimap IoU and F-measure . We study error-sensitivity characteristics of each measure by generating a variety of error types on top of the high-quality ground truth masks from the LVIS dataset . Our analysis confirms that Mask IoU is less sensitive to errors in larger objects. In addition, the analysis reveals limitations of existing boundary-focused measures, such as asymmetries and instability to small changes in mask quality.

Based on these insights we propose a new Boundary IoU measure. Boundary IoU is simple and easy to compute. Instead of considering all pixels, it calculates the intersection-over-union for mask pixels within a certain distance from the corresponding ground truth or prediction boundary contours. Our analysis demonstrates that Boundary IoU measures boundary quality of large objects well, unlike Mask IoU, and it does not over-penalize errors on small objects. An illustrative examples compares Boundary IoU to Mask IoU in Figure 1.

Boundary IoU enables new task-level evaluation metrics. For the task of instance segmentation , we propose Boundary Average Precision (Boundary AP), and for panoptic segmentation , we propose Boundary Panoptic Quality (Boundary PQ).

Boundary AP assesses all relevant aspects of instance segmentation, simultaneously taking into account categorization, localization, and segmentation quality, unlike prior boundary-focused metrics for instance segmentation like AF that ignore false positive rates. We test Boundary AP on three common datasets: COCO , LVIS , and Cityscapes . With real predictions from recent instance segmentation methods that directly aim to improve boundary quality , we verify that Boundary AP tracks improvements better than Mask AP. With synthetic predictions, we show that Boundary AP is significantly more sensitive to large-object boundary quality than Mask AP.

For panoptic segmentation, we apply Boundary PQ to the COCO and Cityscapes panoptic datasets. We test the new metric with synthetic predictions and show that it is more sensitive than the previous metric based on Mask IoU. Finally, we evaluate the performance of various recent instance and panoptic segmentation models with the new evaluation metrics to ease comparison for future research.

These new metrics reveal improvements in boundary quality that are generally ignored by Mask IoU-based evaluation metrics. We hope that the adoption of these new boundary-sensitive evaluations can enable faster progress towards segmentation models with better boundary quality.

Related Work and Preliminaries

Image segmentation tasks like semantic, instance, or panoptic segmentation are evaluated by comparing segmentation masks predicted by a system to ground truth masks provided by annotators. Modern evaluation metrics for these tasks are based on segmentation quality measures that evaluate consistency between ground truth object shape GG and prediction shape PP represented by binary masks of a fixed resolution. We define the most common segmentation quality measures and the new Boundary IoU measure in Table 1 using the unified notation presented in Table 2. We split the measures into mask- and boundary-based types and discuss their differences next.

take into account all pixels of an object mask. The first PASCAL VOC semantic segmentation track in 2007 used Pixel Accuracy measure to evaluate predictions. For each class it calculates the ratio of correctly labeled ground truth pixels (see Table 1). Pixel accuracy is not symmetric and biased toward prediction masks that are larger than ground truth masks. Subsequently, PASCAL VOC switched its evaluation to the Mask Intersection-over-Union (Mask IoU) measure.

Mask IoU segmentation consistency measure divides the number of pixels in the intersection of the prediction and ground truth masks by the number of pixels in their union (see Table 1). The measure is widely used in the evaluation metrics for most popular semantic, instance, and panoptic segmentation tasks and datasets . Unlike Pixel Accuracy, Mask IoU is symmetric, however, as we will show in this paper, it demonstrates unbalanced responsiveness to the boundary quality across object sizes.

Boundary-based segmentation measures

evaluate segmentation quality by estimating contour alignment between predicted and ground truth masks. Unlike mask-based measures, these measures only evaluate the pixels that lie directly on the masks’ contours or in their close proximity.

Trimap IoU is a boundary-based segmentation measure that calculates IoU in a narrow band of pixels within a pixel distance dd from the contour of the ground truth mask (see Table 1). In contrast to Mask IoU, Trimap IoU reacts similarly to comparable pixel errors across object scales because it calculates IoU only for pixels around the contour. However, unlike Mask IoU, the measure is not symmetric and favors predictions whose masks are larger than the corresponding ground truth masks. Moreover, the measure ignores prediction errors that appear outside the band around the ground truth contour.

Trimap and F-measure are often used to evaluate boundary quality for semantic segmentation tasks in an ad-hoc fashion. For example, Trimap IoU is used as an extra evaluation to show boundary quality improvement , but it is not reported by most segmentation methods. In the next section we will study both measures in detail and analyze their behavior across different error types and object sizes.

Sensitivity Analysis

In §4 and §5 we will compare several mask consistency measures by observing how a measure’s value changes in response to errors of different magnitudes. We will observe and interpret these curves to draw conclusions about the behavior of these measures, a methodology that we refer to as sensitivity analysis.

To enable a systematic comparison, we simulate a set of common segmentation errors across different mask sizes by generating pseudo-predictions from ground truth annotations. This approach allows us explicitly control the type and severity of the errors used in the analysis. Moreover, the use of pseudo-predictions avoids any bias toward specific segmentation models which makes the analysis more robust and general. A potential limitation of this approach is that simulated errors may not fully represent errors created by real models. We aim to counteract this limitation by using a diverse set of error types. Figure 2 depicts an example of each error type we consider: ∙\bullet Scale error. Dilation/erosion are applied to the ground truth masks. The error severity is controlled by the kernel radius of the morphological operations. ∙\bullet Boundary localization error. Random Gaussian noise is added to the coordinate of each vertex in polygons that represent ground truth masks. The error severity is controlled by the standard deviation (std) of the Gaussian noise. ∙\bullet Object localization error. Ground truth masks are shifted with random offsets. The error severity is controlled by the pixel length of the shift. ∙\bullet Boundary approximation error. The simplify function from Shapely removes vertices from the polygons that represent ground truth masks while keeping the simplified polygons as close to the original ones as possible. The error severity is controlled by the error tolerance parameter of the simplify function. ∙\bullet Inner mask error. Holes of random shape are added to ground truth masks. The error severity is controlled by the number of holes added. While this error type is not common for modern segmentation approaches, we include it to assess the effect of interior mask errors.

For the analysis, we randomly sample instance masks from the LVIS v0.5 validation set. The dataset is selected due to its high-quality annotations. Using these masks, for each segmentation error type we create multiple sets of pseudo-predictions by varying the severity of the error. To analyze a segmentation measure, we report its mean and standard deviation across a set of pseudo-predictions that represent a given error type of a fixed severity. We will also compare segmentation measures across different object sizes by generating a separate set of pseudo-predictions using ground truth objects within a specific mask area range. For all boundary-based measures that use pixel distance parameter, dd, we set it to 2%2\% of the image diagonal for fair comparison.

Analysis of Existing Segmentation Measures

First, we analyze the standard Mask IoU segmentation consistency measure from both theoretical and empirical perspectives. Then, we study two existing alternatives – Trimap IoU and F-measure boundary-based measures.

Mask IoU is scale-invariant w.r.t\onedotobject area. For a fixed Mask IoU value, a larger object will have more incorrect pixels and the change in incorrect pixel count grows in proportional to the change in object area (as Mask IoU is a ratio of areas). However, when scaling up a typical object, the number of interior pixels grows quadratically, whereas the number of contour pixels only grows linearly. These different asymptotic growth rates cause Mask IoU to tolerate a larger number of misclassified pixels per each unit of contour length for a larger object.

Empirical analysis.

This property corresponds to an assumption that boundary localization error in ground truth annotations (i.e\onedot, intrinsic annotation ambiguity) also grows with the object size. However, a classic study on multi-region segmentation shows that the pixel distance between two contours of the same object labeled by different annotators seldomly exceeds 1% of the image diagonal, irrespective of the object size. We confirm this observation by exploring double annotations that are provided for a subset of images in LVIS . In Figure 3 we present a random pair of objects with significant size difference. While one of the objects is 100 times larger, the boundary discrepancy within the cropped part, which has the same resolution, is similar between the two objects. Observed results suggest that boundary ambiguity is fixed and independent of objects area. This is likely a consequence of the annotation tool, which includes the ability to zoom while drawing contours.

Using simulated scale errors (described in §3) we confirm Mask IoU’s bias in favor of large objects. The dilation/erosion of the ground truth mask by a fixed number of pixels significantly decreases Mask IoU for small objects while Mask IoU grows as object area increases (see Figure 4(a)). Note that Mask IoU’s insensitivity to boundary errors on large objects cannot be addressed by simply increasing the lowest Mask IoU threshold in evaluation metrics like AP or PQ. Such a change does not remedy the bias and will lead to relative over-penalization for smaller objects.

2 Boundary-Based Measures

Next, we will analyze the boundary-based measures Trimap IoU and F-measure. These measures focus on pixels within a distance dd from object contours. The parameter dd is usually fixed on the dataset or image level which results in these measures treating boundary localization errors independently of the size of the object. By matching the natural characteristic of the ground truth segmentation data, these boundary-based measures are better suited to evaluate improvements in boundary quality across object sizes.

computes IoU for a region around the ground truth boundary only (i.e\onedot, the region is independent of the prediction), and therefore is not symmetric: swapping the prediction and ground truth masks will give a different score. In Figure 4(b) we show that this asymmetry favors predictions that are larger than ground truth masks. For larger (dilated) pseudo-predictions Trimap IoU does not drop below some positive value irrespective of the error severity, whereas for smaller (eroded) pseudo-predictions it drops to 0. Moreover, the measure ignores any errors outside of the ground truth boundary region, penalizing inner mask errors less than Mask IoU (see the appendix for details).

F-Measure

matches the pixels of the predicted and ground truth contours if they are within the pixels distance threshold dd. Hence, it ignores small contour misalignments that can be attributed to ambiguity. While robustness to ambiguity is good in principle, in Figure 4(c) we observe that F-measure can be nearly discontinuous, rapidly stepping from 1 to 0 when the error severity changes by a small amount. Sharp response curves can lead to task metrics with high variance. In comparison, Mask IoU is more continuous. Further, dd may be large relative to small objects, causing F-measure to award significant errors a perfect score.

Discussion.

Given the limitations presented above, we conclude that neither Trimap IoU nor F-measure can replace Mask IoU as the main segmentation consistency measure for a broad range of evaluation metrics. At the same time, Mask IoU is biased towards large objects in a way that discourages improvements to boundary segmentations. Next, we propose Boundary IoU as a new measure to evaluate segmentation boundary-quality that does not have any of the previously mentioning limitations.

Boundary IoU

1 Evaluation on Synthetic Predictions

Using synthetic predictions we evaluate the segmentation quality aspect of instance segmentation in isolation without a bias that any particular model can have. We simulate predictions by capping the effective resolution of each mask. First, we downscale cropped ground truth masks to a 28×2828\times 28 resolutionThis is a popular prediction resolution used in practice . mask with continuous values, we then upscale it back using bilinear interpolation, and finally binarize it. Such synthetic masks are close to the ground truth masks for smaller objects, however the discrepancy grows with object size. In Table 3 we report overall AP and APSS, APMM, and APLL for object size splits defined in COCO . Mask AP follows the behavior of Mask IoU, showing little sensitivity to the error growth between APS and APL. In contrast, Boundary AP successfully captures the difference with significantly lower score for larger objects. In the appendix, we provide an example of the synthetic predictions and more results using different effective resolutions.

2 Evaluation on Real Predictions

We use outputs of existing segmentation models to further study Boundary AP. Unless specified, to isolate the segmentation quality from categorization and localization errors for purposes of analysis, we supply ground truth boxes to these methods and assign a random confidence score to each box. We use Detectron2 with a ResNet-50 backbone unless otherwise specified.

Table 4(a) shows both Mask AP and Boundary AP for the standard Mask R-CNN model . Mask R-CNN is well-known to predict blobby masks with significant visual defects for larger objects (see Figure 1). Nevertheless, Mask APLL is larger than Mask APSS. In contrast, we observe that Boundary APLL is smaller than Boundary APSS for Mask R-CNN suggesting that the new measure is more sensitive to the boundary quality of the large objects. Note that in this experiment the use of ground truth boxes removes any categorization and localization errors that are usually larger for small objects.

Segmentation vs\onedotcategorization and localization.

A general evaluation metric for instance segmentation should track the improvements in all aspects of the task including segmentation, categorization, and localizations. In Table 4(b) we first evaluate Mask R-CNN with several backbones (ResNet-50, ResNet-101, and ResNeXt-101-32×\times8d ), again supplying ground truth boxes. Note that both Mask AP and Boundary AP do not change significantly with different backbones, suggesting that more powerful backbones do not directly influence the segmentation quality. Next, we evaluate Boundary AP requiring each model to predict its own boxes as is standard. We observe that Boundary AP is able to track improvements from better localization and categorization similarly to Mask AP.

Mask quality improvements.

We explore Boundary AP’s ability to capture the improvements in segmentation quality by the methods designed for this purpose in Tables 4(c) and 4(d). To compare the segmentation quality aspect across models we again supply ground truth boxes to each model.

PointRend was developed to improve pixel-level prediction quality of models like Mask R-CNN and can produce predictions of varying resolution. PointRend significantly improves mask quality, while this can be measured via mask AP, it is more pronounced in Boundary AP, especially for large objects and for a higher resolution PointRend variant. See Table 4(c) for details.

Boundary-preserving Mask R-CNN (BMask R-CNN) improves boundary quality by adding a direct boundary supervision loss and increasing the resolution of feature maps used in its mask head. In Table 4(d) Boundary AP shows that BMask R-CNN with its 28×2828\times 28 output resolution outperforms PointRend for small objects, whereas for larger objects the 224×224224\times 224 resolution output of PointRend is preferable, which matches a subjective visual quality assessment (see an example in Figure 1). We hope that the improved sensitivity of the new Boundary AP metric will lead to a rapid progress in the methods that improve boundary quality for instance segmentation.

Unlike the standard Mask IoU, Boundary IoU segmentation quality measure provides a clear, quantitative gradient that rewards improvements to boundary segmentation quality. We hope that the new measure will challenge our community to develop new methods with high-fidelity mask predictions. In addition, Boundary IoU allows a more granular analysis of segmentation-related errors for the complex multifaceted tasks like instance and panoptic segmentation. Incorporation of the measure in the performance analysis tools like TIDE can provide better insights into specific error types of instance segmentation models.

computes IoU for a band around the ground truth boundary and, therefore, it ignores errors away from the ground truth boundary (e.g\onedotinner mask prediction errors). We generate pseudo-predictions with such errors by adding holes of random shapes to ground truth masks. In Figure 8 we show that Trimap IoU penalizes inner mask prediction errors less than Mask IoU.

F-measure

matches the pixels of the predicted and ground truth contours if they are within the pixels distance threshold dd. In the experiments presented in the main text we observe that this strategy makes F-measure ignore scale type errors for smaller objects. The Mean F-measure (mF-measure) modification ameliorates this limitation by averaging several F-measures with different threshold parameters dd. Figure 9 demonstrates the sensitivity curves of this measure for the scale (dilation) error type. For mF-measure we use dd from 0.1%0.1\% to 2.1%2.1\% image diagonal with 0.4%0.4\% increment (from 1 pixel to 17 pixels on average) to compare it with Boundary IoU that uses single dd set to 2%2\% image diagonal. We observe that mF-measure behaves similarly to Boundary IoU for large objects, however it under-penalizes errors in small objects where Boundary IoU matches Mask IoU behavior. Furthermore, mF-measure is substantially slower than Boundary IoU as it requires to perform the matching of prediction/ground truth pairs several times for different thresholds dd.

Boundary IoU

can award a perfect score for two non-identical masks (see Figure 10). As discussed in the main text, we observe that Boundary IoU is smaller or equal to Mask IoU in the absolute majority of cases and the inequality is violated when prediction misses interior part of an object (similar to the toy example in Figure 10). To mitigate this limitation, we propose a simple combination of Mask IoU and Boundary IoU by taking their minimum for real-world segmentation evaluation metrics.

Appendix B Application

We evaluate instance segmentation on three datasets: COCO , LVIS and Cityscapes .

COCO is the most popular instance segmentation benchmark for common objects. It contains 80 categories. There are 118k images for training, 5k images for validation and 20k images for testing.

LVIS is a federated dataset with more than 1000 categories. It shares the same set of images as COCO but the dataset has higher quality ground truth masks. We use LVISv0.5 version of the dataset. Following , we construct the LVIS∗v0.5 dataset which keeps only the 80 COCO categories from LVISv0.5. LVIS∗v0.5 allows us to compare models trained on COCO using higher quality mask annotations from LVIS (i.e\onedotAP∗ in ).

Cityscapes is a street-scene high-resolution dataset. There are 5k images annotated with high quality pixel-level annotations and 8 classes with instance-level segmentation.

Evaluation on synthetic predictions.

We simulate predictions by capping the effective resolution of each mask. First, we downscale cropped ground truth masks to a fixed resolution mask with continuous values, we then upscale it back using bilinear interpolation, and finally binarize it. Figure 11 show visualization of the synthetic predictions with different effective resolutions. In Table 5 we compare Mask AP and Boundary AP for the synthetic predictions with different synthetic scales across different datasets.

Evaluation on real predictions.

In addition to the experiments with COCO in the main text, we evaluate Mask R-CNN , PointRend , and Boundary-preserving Mask R-CNN (BMask R-CNN) on LVIS∗v0.5 and Cityscapes in Table 6. For each method we feed ground truth boxes to isolate the segmentation quality aspect of the instance segmentation task. On all datasets we observe that Boundary AP better captures improvements in the mask quality.

Reference Boundary AP evaluation.

We provide Boundary AP evaluation for various recent and classic models on COCO (Table 7), LVIS (Table 8), and Cityscapes (Table 9) datasets. We do not train any models ourselves and use the Detectron2 framework or official implementations instead. These results can be used as a reference to simplify the comparison for future methods.

B.2 Panoptic Segmentation

The standard evaluation metric for panoptic segmentation is panoptic quality (PQ or Mask PQ) , defined as: PQ = ⏟∑(p, g) ∈TPIoU(p, g)12—TP—_segmentation quality (SQ) ×⏟—TP——TP— + 12—FP— + 12—FN—_recognition quality (RQ)

Mask IoU is presented in two places: (1) calculating the average Mask IoU for true positives in Segmentation Quality (SQ) component and (2) matching prediction and ground truth masks to split them into true positives, false positives, and false negatives. Similarly to Boundary AP, we replace Mask IoU with min(Mask IoU,Boundary IoU)\text{min}(\text{Mask IoU},\text{Boundary IoU}) in both places and refer the new metric as Boundary PQ.

We use two popular datasets with panoptic annotation: COCO panoptic and Cityscapes .

COCO panoptic combines annotations from COCO instance segmentation and COCO stuff segmentation into a unified panoptic format with no overlaps. COCO panoptic has 80 things and 53 stuff categories.

Cityscapes has 8 thing and 11 stuff categories.

Similar to the instance segmentation task, we set dilation width to 2%2\% image diagonal for COCO panoptic and 0.5%0.5\% image diagonal for Cityscapes.

Analysis with synthetic predictions.

Following our experimental setup for instance segmentation, we evaluate Boundary PQ on low-fidelity synthetic predictions generated from ground truth annotations to avoid any potential bias toward a specific model. The synthetic predictions are generated by downscaling ground truth panoptic segmentation maps for each image and then upscaling it back using nearest neighbor interpolation in both cases. This image-level generation process ensures a unified treatment of both things and stuff segments following the idea behind the panoptic segmentation task.

In Table 11, we report Panoptic Quality and its two components: Segmentation Quality (SQ) and Recognition Quality (RQ) for synthetic predictions with various downscaling ratios across different datasets. Similar to our findings for AP, Boundary PQ better tracks boundary quality improvements than Mask PQ for panoptic segmentation. Furthermore, we find that the difference between Boundary PQ and Mask PQ is mainly caused by the difference in SQ. This observation confirms that Boundary IoU better tracks the mask quality of predictions and does not significantly change other aspects like the matching procedure between prediction and ground truth segments.

References Boundary PQ evaluation.

We provide Boundary PQ evaluation for various models on COCO panoptic and Cityscapes datasets in Table 10. We do not train any models ourselves and use models trained by their authors. These results can be used as a reference to simplify the comparison for future methods.

References