Boundary-preserving Mask R-CNN

Tianheng Cheng, Xinggang Wang, Lichao Huang, Wenyu Liu

Introduction

Instance segmentation, a fundamental but challenging task in computer vision, aims to assign a pixel-level mask to localize and categorize each object in images, driving numerous vision applications such as autonomous driving, robotics and image editing. With the rapid development of deep convolutional neural networks (DCNN), various methods based on DCNN were proposed for instance segmentation. Prevalent methods for instance segmentation are based on object detection, which provides box-level localization information for instance-level segmentation, among which Mask R-CNN is the most successful one. It extends Faster R-CNN by adding a simple fully convolutional network (FCN) to predict the mask of each detected instance. Due to the great effectiveness and flexibility, Mask R-CNN serves as a state-of-the-art baseline and has facilitated most recent instance segmentation research, such as .

In the Mask R-CNN framework, state-of-the-art instance segmentation networks obtain instance masks by performing pixel-level classification via FCN. It treats all pixels in the proposal equally and ignores the object shape and boundary information. However, the pixels near boundaries are hard to be classified. Evidently, it is hard for pixel-level classifier to guarantee precise masks. We find that fine boundaries can provide better localization performance and make the object masks more distinct and clear. As illustrated in Fig. 2, Mask R-CNN (the first row) without consideration about boundaries is prone to output coarse and indistinct segmentation results with unreasonable overlaps between objects in comparison with the one that involves boundaries (the second row).

To address this issue, we leverage instance boundary information to enhance the mask prediction. Instance boundary is a dual representation of instance mask and it can guide the mask prediction network to output masks that are well-aligned with their groundtruths. Thus, the masks are more distinct and give more precise object location. Based on this motivation, we propose a conceptually simple and novel Boundary-preserving Mask R-CNN (BMask R-CNN) that unifies instance-level mask prediction and boundary prediction in one network.

Specifically, based on Mask R-CNN, we replace the original mask head with the proposed boundary-preserving mask head which contains two sub-networks for jointly learning object masks and boundaries. We insert two feature fusion blocks to strengthen the connection between boundary feature learning and mask feature learning. At last, mask prediction is guided by boundary features which contain abundant shape and localization information. The main purpose of learning boundaries is to capture features for precise object localization. Nevertheless, learning boundaries is non-trivial, because boundary groundtruths that generated from sparse annotated polygons (i.e.\emph{i.e}.\hbox{}, in the COCO dataset ) are noisy and boundary classification has less training pixels than that for mask classification. To solve this problem, we further dive into the optimization for boundary learning by performing studies about boundary loss and exploit a boundary classification loss by combining binary cross-entropy loss and the dice loss .

We perform extensive experiments to evaluate the performance of BMask R-CNN. On the challenging COCO dataset , BMask R-CNN achieves considerably significant improvements compared with Mask R-CNN regardless of the backbones. Note that our BMask R-CNN provides larger gains if it requires more precise mask localization, as shown in Fig. 2. On the fine-annotated Cityscapes dataset , BMask R-CNN brings larger improvements with better mask annotations.

The main contributions of this paper can be summarized as follows.

We present a novel Boundary-preserving Mask R-CNN (BMask R-CNN), which is the first work that explicitly exploits object boundary information to improve mask-level localization accuracy in the state-of-the-art Mask R-CNN framework.

BMask R-CNN is conceptually simple yet effective. Without bells and whittles, BMask R-CNN outperforms Mask R-CNN by 1.7%1.7\% AP and 2.2%2.2\% AP on the COCO val set and the Cityscapes test set respectively. Further, BMask R-CNN obtains higher AP gains when the mask IoU threshold becomes higher, as shown in Fig. 2.

We perform ablation studies on the components of BMask R-CNN, e.g., feature fusion blocks, boundary features, boundary losses and the Sobel mask head, which are helpful to interpret how BMask R-CNN works and provide some thoughts for further research on instance segmentation.

Related Work

Existing methods can be divided into two categories, i.e. detection-based methods and segmentation-based methods. Detection-based methods employ object detectors to generate region proposals and then predict their masks after RoI pooling/align . Based on CNN, predict masks for object proposals. FCIS extends InstanceFCN by exploiting position-sensitive inside/outside score maps and fully convolutional networks for instance segmentation. MNC using end-to-end cascade networks divides instance segmentation into three tasks: regress bounding boxes, estimate instance masks and categorize instances. BAIS uses boundary-based distance transform to predict mask pixels that are beyond bounding boxes. Mask R-CNN extends Faster R-CNN by adding a mask prediction branch in parallel with the existing box regression and classification branches, demonstrating competitive performance on both object detection and instance segmentation. PANet based on Mask R-CNN introduces the bottom-up path augmentation for FPN to enhance information flow and adaptive feature pooling for better mask features. MaskLab , built on Faster R-CNN, obtains instance-level masks by combining cropped semantic segmentation and direction prediction which separate objects of different semantic classes and same semantic classes respectively. Mask scoring R-CNN addresses the misalignment between mask quality and mask score in Mask R-CNN by explicitly learning the quality of predicted masks. further improves Cascade Mask R-CNN by interweaving box and mask branches in a multi-stage cascade manner and providing spatial context through semantic segmentation. Huang et al. apply a criss-cross attention module to capture the full-image contextual information for instance segmentation. draws on the idea of rendering and adaptively selects key points to recover fine details for high-quality image segmentation.

Segmentation-based methods first exploit pixel-level segmentation over the image and then group the pixels together for each object. InstanceCut adopts boundaries to partition semantic segmentation into instance-level segmentation. SGN groups pixels along rows and columns by line segments. utilizes predicted instance centers and pixel-wise directions to group instances with fully convolutional networks. Recently, several methods take the advantage of deep metric learning to learn the embedding for instances and then group pixels to form instance-level segmentation.

0.2 Boundary, Edge and Segmentation:

Deep fully convolutional neural networks has achieved great progress in edge detection. Xie et al. propose the fully convolutional holistically-nested edge detector HED which performs in an image-to-image manner and end-to-end training. CASENet presents a novel challenging task semantic boundary detection, aiming to detect category-aware boundaries. investigate the label misalignment problem caused by noisy labels in semantic boundary detection. proposes geometric aware loss function for object skeleton detection in nature images. In semantic segmentation, Chen et al. propose fully connected conditional random field (CRF) to capture spatial details and refine boundaries. Recent semantic segmentation methods leverage predicted boundaries or edges to facilitate semantic segmentation. refine segmentation results with direction fields learned from predicted boundaries. Zimmermann et al. propose edge agreement head to focus on boundaries of instances with an auxiliary edge loss. Different from these previous methods, BMask R-CNN explicitly predicts instance-level boundaries, from which we obtain instance shape information for better mask localization. Compared to semantic segmentation, boundaries in instance segmentation have dual relations to the masks. Therefore, we build fusion blocks to mutually learn boundary and mask features and improve the representations for mask localization and lead the mask prediction focus more on the boundaries.

Boundary-preserving Mask R-CNN

In Mask R-CNN, instance segmentation is performed based on pixel-level predictions. To learn a translation invariant predictor, predictions are made based on the local information. Though the local features extracted using deep network have large receptive fields, the shape information of object is ignored. Thus, the predicted masks often contain coarse and indistinct as well as some false positive predictions. For better understanding this problem, we analyze and visualize some raw mask prediction from Mask R-CNN with ResNet-50 and FPN. As shown in Fig. 3, some mask predictions are rough and imprecise. Obviously, employing object boundaries will be helpful to address this issue by providing better localization and guidance. Therefore, we propose a Boundary-preserving Mask R-CNN to exploit boundary information to guide more precise mask prediction.

2 Boundary-preserving Mask Head

BMask R-CNN improves the mask head in Mask R-CNN with boundary features and boundary prediction, as illustrated in Fig. 4. The new mask head is termed as boundary-preserving mask head. In the first stage, FPN extracts and constructs pyramidal features for RPN and the heads in the second stage. After RPN generates proposals, the box head inherited from Mask R-CNN utilizes these proposals and extracted RoI features for classification and bounding box regression. The boundary-preserving mask head performs RoIAlign to acquire RoI features for both boundary and mask prediction.

Boundary-preserving mask head jointly learns object boundaries and masks in an end-to-end manner. Note that object boundary and object mask have a close relation and we can easily convert either one to another. Features from the mask sub-network can provide high-level semantic information for learning boundaries. After obtaining boundaries, the shape information and abundant location information in boundary features can guide more precise mask predictions.

We define Rm\mathcal{R}_{m} and Rb\mathcal{R}_{b} as Region of Interest (RoI) features for mask prediction and boundary prediction respectively. Following , Rm\mathcal{R}_{m} is extracted from the specific feature pyramid level (P2∼P5P2\sim P5) according to the scale of the proposal, while Rb\mathcal{R}_{b} is obtained from the finest-resolution feature level P2P2, containing abundant spatial information. To preserve spatial information better for boundary prediction, the resolution of Rb\mathcal{R}_{b} is set to be larger than that of Rm\mathcal{R}_{m} when performing RoIAlgin. Then, it is downsampled by a strided 3×33\times 3 convolution and the output feature is denoted as Rb~\widetilde{\mathcal{R}_{b}}. Rb~\widetilde{\mathcal{R}_{b}} has the same resolution as Rm\mathcal{R}_{m} and is used for feature fusion.

The feature fusion scheme in BMask R-CNN is illustrated in Fig. 4. Mask RoI features Rm\mathcal{R}_{m} is fed into 4 consecutive 3×33\times 3 convolutions and the output feature is denoted as Fm\mathcal{F}_{m}. Boundary features Rb~\widetilde{\mathcal{R}_{b}} is fused with Fm\mathcal{F}_{m} and then fed into two consecutive 3×33\times 3 convolutions.

2.2 Mask →→\rightarrow Boundary (M2B) Fusion:

Mask features Fm\mathcal{F}_{m} contain rich high-level information, i.e., the pixel-wise object category information, which is beneficial to predict object boundaries. Hence, we propose a simple fusion block to integrate boundary features and mask features for boundary prediction. The fusion block can be formulated as follows.

where Fb\mathcal{F}_{b} denotes the output boundary features and ff means a 1×11\times 1 convolution with ReLU.

2.3 Boundary →→\rightarrow Mask (B2M) Fusion:

We fuse the final boundary features with mask features; thus, boundary information can be used to enrich mask features and guide precise mask prediction. The fusion block is the same as the M2B fusion block.

2.4 Predictor:

Following Mask R-CNN, our predictor is a 2×22\times 2 deconvolution followed by the 1×11\times 1 convolution as output layer. Both masks and boundaries are class-specific.

3 Learning and Optimization

Following the common practice in edge detection , we regard boundary prediction as a pixel-level classification problem. The learned boundary features are fused with mask features to provide shape information for mask prediction.

We use the Laplacian operator to generate soft boundaries from the binary mask groundtruths. The Laplacian operator is a second-order gradient operator and can produce thin boundaries. The produced boundaries are converted into binary maps by a threshold as the final groundtruths.

3.2 Boundary Loss:

Most boundary or edge detection methods take the advantage of weighted cross-entropy to alleviate the class-imbalance problem in edge/boundary prediction. However, weighted binary cross-entropy leads to thick and coarse boundaries . Following , we use dice loss and binary cross-entropy to optimize the boundary learning. Dice loss measures the overlap between predictions and groundtruths and is insensitive to the number of foreground/background pixels, thus alleviating the class-imbalance problem. Our boundary loss Lb\mathcal{L}_{b} is formulated as follows.

where ii denotes the ii-th pixel and ϵ\epsilon is a smooth term to avoid zero division (We set ϵ=1\epsilon=1.). In ablation experiments, we will analyze and evaluate different loss functions with quantitative results and qualitative results.

3.3 Multi-Task Learning:

Multi-task learning has been proved effective in many works , which achieves better performance for different tasks comparing with separate training. Since boundary and mask are crossed linked by two fusion blocks, jointly training can enhance the feature representation for both boundary and mask predictions. We define a multi-task loss for each sample as follows.

where the classification loss Lcls\mathcal{L}_{cls} and regression loss Lbox\mathcal{L}_{box} are inherited from Mask R-CNN. Mask loss Lmask\mathcal{L}_{mask} is a category-specific pixel-level binary cross-entropy loss for each instance, which is directly taken from Mask R-CNN. The boundary loss Lb\mathcal{L}_{b} has been introduced in detail in Equation (2).

Experiments

We perform extensive experiments on the challenging COCO dataset and the Cityscapes dataset to demonstrate the effectiveness of Boundary-preserving Mask R-CNN. To better understanding each component of our method, we provide detailed ablation experiments on COCO.

COCO contains 115k images for training, 5k images for validation and 20k images for testing. Our models are trained on the training set (train2017). We report the results on the validation set (val2017) for ablation studies and the results on testing set (test-dev2017) to compare with other methods. The Cityscapes dataset is collected in urban scenes which contains 2975 training, 500 validation and 1525 testing images. As for instance segmentation, Cityscapes involves 8 object categories and provides more precise instance-level segmentation annotations than COCO. We train our models on the training set and report our performance on the validation set and the testing set. For both COCO and Cityscapes, we use the same evaluation metric (i.e., COCO AP), which is the average precision over different IoU thresholds (from 0.5 to 0.95).

0.2 Implementation details:

We adopt Mask R-CNN as our baseline and our method is developed based on it. All hyper-parameters are kept the same. Unless specified, we use ResNet-50 with FPN as our backbone network. We initialize our backbone networks with ImageNet pre-trained weights and all Batch Normalization layers are frozen due to small mini-batch. The input images are resized such that the shorter side is 600 pixels and the longer is less than 1000 pixels for ablation experiments. As for main experiments, we adopt 800 pixels for shorter side (the longer is less than 1333 pixels). Following the standard practice, we train all models on 4 NVIDIA GPUs using Synchronized SGD with initial learning rate 0.02 and 16 images per mini-batch for 90,000 iterations and reduce the learning rate by a factor of 0.1 and 0.01 after 60,000 and 80,000 iterations respectively. For larger backbones, we follow the linear scaling rule to adjust the learning rate schedule when decreasing mini-batch size.

1 Overall Results

We first evaluate our BMask R-CNN with different backbones on COCO and compare it with Mask R-CNN. As shown in Table 1, our method outperforms Mask R-CNN by remarkable APs in spite of different backbones and input sizes. Compared with Mask R-CNN, BMask R-CNN significantly achieves 1.41.4, 1.71.7 and 1.51.5 AP improvements using ResNet-50-FPN, ResNet-101-FPN and HRNetV2-W32-FPN respectively. Exploiting boundary information contributes to more precise mask localization due to the observation that our method yields noteworthy and stable improvements (≈2.3\approx 2.3 AP) on AP75. APb shows AP for bounding box, on which BMask R-CNN very slightly improves over Mask R-CNN.

In Table 2, we compare BMask R-CNN with some state-of-the-art instance segmentation methods. All models are trained on COCO train2017 and evaluated on COCO test-dev2017. Without bells and whistles, BMask R-CNN with ResNet-101-FPN can surpass these methods.

Fig. 2 illustrates the AP curves of BMask R-CNN and Mask R-CNN under different IoU thresholds. Note that our method obtains larger gain when the IoU threshold increases, which shows the better localization performance of BMask R-CNN.

2 Ablation Experiments

In order to comprehend how BMask R-CNN works, we perform exhaustive experiments to analyze the components in BMask R-CNN. Table 3 shows the results of gradually adding components to the Mask R-CNN baseline. Each component of our proposed BMask R-CNN will be investigated in the following sections.

To validate the effect of boundaries for mask prediction, we use mask targets to replace boundary targets and also evaluate the performance without bounadry supervision and loss with the architecture kept the same. Table 4 indicates that boundary supervision with our proposed boundary-preserving mask head improves mask results by 0.80.8 and 0.70.7 AP compared with mask supervision and no supervision respectively. Notably, using boundary can improve the mask localization performance (AP75) by a significant margin.

2.2 RoI Feature Extraction:

Compared with mask prediction, predicting boundaries requires more precise spatial information due to boundaries are spatially sparse. Therefore, we explore several strategies and present two considerations to extract better RoI features for boundaries. The first aspect is the source of RoI features. Lin et al. propose that RoI features are extracted from the different levels (P2∼P5P2\sim P5) in FPN depending on the scales of the corresponding proposals. Features of high levels in FPN lacks spatial information which are inappropriate for boundaries. Fig. 5 illustrates different sources for mask features Rm\mathcal{R}_{m} and boundary features Rb\mathcal{R}_{b}. Fig. 4.2.2 shows that boundary features are directly extracted from P2P2 while mask features are from from P2∼P5P2\sim P5 according to the scale of the proposal. Fig. 4.2.2 shows that both boundary features and mask features are extracted from the same feature level from P2∼P5P2\sim P5. The other aspect is the feature resolution. Higher-resolution features preserve more spatial information which is beneficial to boundaries. Therefore, we explore the effects of 28×2828\times 28 resolution and 14×1414\times 14 resolution RoI features for learning boundaries. As Table 4.2.2 shows, directly extracting boundary RoI features from P2P2 is more effective with larger resolution. We employ boundary features extracted from P2P2 with 28×2828\times 28 resolution in other experiments.

2.3 Feature Fusion:

In Section 3.2, we have emphasized the relation between boundary features and mask features. Fusion blocks in our boundary-preserving mask head build explicit links to enrich both feature representation. Table 6 shows more results: if there is no fusion, it improves Mask R-CNN by 0.50.5 AP, which is the gain of multi-task learning; with both M2B and B2M fusion blocks, BMask R-CNN has 1.51.5 AP improvement over Mask R-CNN. We further investigate the influence of adding more subsequent fusion blocks. Keeping the overall computation cost substantially unchanged, adding more B2M or M2B fusions brings negligible improvements.

2.4 Loss Functions:

In order to obtain precise boundaries, we evaluate the impacts of different loss functions for optimizing boundary learning. Table 4.2.4 shows that the combination of BCE and Dice loss leads to better performance compared with individual BCE or Dice loss. Weighted BCE brings less gain than BCE in boundary prediction.

To investigate how the Dice-BCE combined loss provides such competitive improvements, we present detailed analysis on the visualization results of these experiments. As shown in Fig. 6, different loss functions have different impacts on learning boundaries. BCE loss provides considerably precisely-localized but unclear boundaries due to the class-imbalance problem. Weighted BCE solves this problem by applying balancing weights but this hard balancing leads to thick and coarse boundaries which exceed their corresponding masks. Dice loss also solves the class-imbalance problem without thick boundaries but lacks precise localization. Consequently, combining Dice loss and BCE can provide better-localized boundaries and avoid the class-imbalance problem.

2.5 Computation Cost:

Compared with Mask R-CNN, our method involves four 3×33\times 3 and two 1×11\times 1 convolutional layers for boundary prediction and two fusion blocks which increase the computation cost. To clarify the improvements of BMask R-CNN are not from extra computation cost, we form a larger mask head by adding 4 more 3×33\times 3 convolutional layers as a comparison. Table 8 shows that BMask R-CNN still achieves a significant gain compared with Mask R-CNN with equal computation cost.

3 Experiments on Cityscapes

To further explore the effects of BMask R-CNN on the fine-annotated Cityscapes dataset, we only use images with fine annotations to train and evaluate our models. For fair comparisons, we use ResNet-50-FPN as our backbone and resize images with shorter edge randomly selected from for training. For inference, input images are kept the original size 1024×\times2048. Models are trained by SGD on 4 GPUs with mini-batch size 4 for 48,000 iterations. The learning rate is 0.005 at the beginning and reduced to 0.0005 after 36,000 iterations. Other settings are the same with experiments on COCO.

We report the results evaluated on Cityscapes val and test in Table 9. BMask R-CNN achieves 29.429.4 AP on test and obtains a remarkable 2.22.2 AP gain compared with the baseline Mask R-CNN. BMask R-CNN outperforms previous methods without extra data.

4 Discussions

When datasets become larger and larger, obtaining precise mask annotations is unavoidably time-consuming. Though the COCO dataset provides abundant instance-level annotations, the mask and boundary annotations (represented by sparse polygons) are coarse, which limits the performance of our method BMask R-CNN. Nevertheless, BMask R-CNN can output more precise and smooth boundaries with fewer mask overlap between instances; some selected examples are shown in Fig. 7.

4.2 Sobel mask head:

Instead of predicting boundaries using an extra branch, we also design a simple Sobel mask head to predict boundaries from masks, which is a improved version of . As illustrated in Fig. 7, it has a Sobel operator and two 3×33\times 3 convolutions following the mask predictions. We adopt the same Dice-BCE loss function for training. Using ResNet-50-FPN backbone and keep the rest settings the same, this Sobel mask head method obtains 34.034.0 AP which improves Mask R-CNN by 0.80.8 AP but is 0.70.7 AP worse than our main method.

5 Qualitative Results

We provide representative visualization results on COCO to compare our method with Mask R-CNN and further prove the effectiveness of our method. Fig. 8 shows the qualitative results on COCO val. Mask R-CNN is more prone to generate masks with coarse boundaries which contain much background along with some false positive areas. Our proposed BMask R-CNN can alleviate this issue with the help of preserving boundaries. We further visualize our raw boundary and mask results in Fig. 8. It can be easily observed that predicted masks are more clear and highly coincident with their boundaries. Furthermore, utilizing predicted boundaries to refine masks brings minor improvement and the refinement is vulnerable to the noises.

Conclusion

We address the issue that coarse boundaries and imprecise localization in instance segmentation and propose a novel Boundary-preserving Mask R-CNN. It incorporates boundary information to guide the mask learning for better boundaries and localization. Our experiments demonstrate that our method achieves remarkable and stable improvements on both COCO and Cityscapes especially in terms of localization performance. Extensive studies and visualization results provide a deep understanding of how our method BMask R-CNN works. Our method could also be plugged into Cascade Mask R-CNN and etc. for higher performance. We hope it can be a strong baseline and sheds light on this fundamental research topic.

Acknowledgements

This work was in part supported by NSFC (No. 61733007 and No. 61876212), Zhejiang Lab (No. 2019NB0AB02), and HUST-Horizon Computer Vision Research Center.

References