Deep Occlusion-Aware Instance Segmentation with Overlapping BiLayers

Lei Ke, Yu-Wing Tai, Chi-Keung Tang

Introduction

State-of-the-art approaches in instance segmentation often follow the Mask R-CNN paradigm with the first stage detecting bounding boxes, followed by the second stage to segment instance masks. Mask R-CNN and its variants have demonstrated notable performance, and most of the leading approaches in the COCO instance segmentation challenge have adopted this pipeline. However, we note that most incremental improvement comes from better backbone architecture designs, with little attention paid in the instance mask regression after obtaining the ROI (Region-of-Interest) features from object detection. We observe that a lot of segmentation errors are caused by overlapping objects, especially for object instances belonging to the same class. This is because each instance mask is individually regressed, and the regression process implicitly assumes the object in an ROI has almost complete contour, since most objects in the training data in COCO do not exhibit significant occlusions.

We propose the Bilayer Convolutional Network (BCNet). As illustrated in Figure 1, BCNet simultaneously regresses both occluding region (occluder) and partially occluded object (occludee) after ROI extraction, which groups the pixels belonging to the occluding region and treat them equally as the pixels of the occluded object but in two separate image layers, and thus naturally decouples the boundaries for both objects and considers the interaction between them during the mask regression stage.

Previous approaches resolve the mask conflict between neighboring objects through non-maximum suppression or additional post-processing . Consequently, their results are over-smooth along boundaries or exhibit small gaps between neighboring objects. Furthermore, since the receptive field in the ROI observes multiple objects that belong to the same class, when the occluding regions were included as part of the occluded object, traditional mask head design falls short of resolving such conflict, leaving a large portion of error as shown in Figure 2. We compare BCNet with recent amodal segmentation methods , which predict complete object masks, including the occluded region. However, these amodal methods only regress single occluded target in the ROI, thus lacking occluder-occludee interaction reasoning, making their specially designed decoupling structure suffer when handling mask conflict between highly-overlapping objects. Correspondingly, Figure 3 compares the architecture of our BCNet with previous mask head designs .

Our BCNet consists of two GCN layers with a cascaded structure, each respectively regresses the mask and boundaries of the occluding and partially occluded objects. We utilize GCN in our implementation because GCN can consider the non-local relationship between pixels, allowing for propagating information across pixels despite the presence of occluding regions. The explicit bilayer occluder-occludee relational modeling within the same ROI also makes our final segmentation results more explainable than previous methods. For object detector, we use the FCOS owing to its efficient memory and running time, while noting that other state-of-the-art object detectors can also be used as demonstrated in our experiments.

Since our paper focuses on occlusion handling in instance segmentation, in addition to the original COCO evaluation, we extract a subset of COCO dataset containing both occluding objects and partially occluded objects to evaluate the robustness of our approach in comparison with other instance segmentation methods in occlusion handling. In this paper we also contribute the first large-scale occlusion aware instance segmentation datasets with ground-truth, complete object contours for both occluding and partially occluded objects. Extensive experiments show that our approach outperforms state-of-the-art methods in both the modal and amodal instance segmentation tasks.

Related Work

Two stage instance segmentation methods achieve state-of-the-art performance by first detecting bounding boxes and then performing segmentation in each ROI region. FCIS introduces the position-sensitive score maps within instance proposals for mask segmentation. Mask R-CNN extends Faster R-CNN with a FCN branch to segment objects in the detected box. PANet further integrates multi-level feature of FPN to enhance feature representation. MS R-CNN mitigates the misalignment between mask quality and score. CenterMask is built upon the anchor free detector FCOS with a SAG-Mask branch. In contrast, our BCNet is a bilayer mask prediction network for addressing the issues of heavy occlusion and overlapping objects in two-stage instance segmentation. Experiments validate that our approach leads to significant performance gain on overall instance segmentation performance not limited to heavily occluded cases.

One-stage instance segmentation methods remove the bounding box detection and feature re-pooling steps. AdaptIS produces masks for objects located on point proposals. PolarMask models instance masks in polar coordinates by instance center classification and dense distance regression. YOLOACT introduces prototype masks with per-instance coefficients. SOLO applies the “instance categories” concept to directly output instance masks based on the location and size. Grouping-based approaches regard segmentation as a bottom-up grouping task by first producing pixel-wise predictions followed by grouping object instances in the post-processing stage. These one-stage methods, with simpler procedures than their two-stage counterparts, are more efficient but tend to be less accurate.

Occlusion Handling

Methods for occlusion handling have been proposed . A layout consistent random field is used in to segment images of cars and faces by imposing asymmetric local spatial constraints. Ghiasi et al. model occlusion by learning deformable models with local templates for human pose estimation while reconstructs dense 3D shape for vehicle pose. Tighe et al. build a histogram to predict occlusion overlap scores between two classes for inferring occlusion order in the scene parsing task. Chen et al. handle occlusion by incorporating category specific reasoning and exemplar-based shape prediction for instance segmentation. For pedestrian detection with occlusion, bi-box regression is proposed in for both full body and visible part estimation while repulsion loss and aggregation loss are designed to improve the detection accuracy. SeGAN learns occlusion patterns by segmenting and generating the invisible part of an object. Recently, OCFusion uses an additional branch to model instances fusion process for replacing detection confidence in panoptic segmentation. A self-supervised scene de-occlusion method is proposed in by recovering the occlusion ordering and completing the mask and content for the invisible object parts.

Compared to these methods, our BCNet tackles occlusion by explicitly modeling occlusion patterns in shape and appearance. This equips the segmentation model with strong occlusion perception and reasoning capability. Our bi-layer approach can be smoothly integrated into state-of-the-art segmentation framework for end-to-end training.

Amodal Instance Segmentation

Different from traditional segmentation which only focuses on visible regions, amodal instance segmentation can predict the occluded parts of object instances. Li and Malik first propose a method by extending , which iteratively enlarges the modal bounding box following the direction of high heatmap values and synthetically adds occlusion. Zhu et al. propose a COCO amodal dataset with 5000 images from the original COCO and use AmodalMask as a baseline, which is SharpMask trained on amodal ground truth. COCOA cls augments this dataset by assigning class-labels to the objects while SAIL-VOS dataset in is targeted for video object segmentation. In autonomous driving, Qi et al. establish the large-scale KITTI InStance segmentation dataset (KINS) and present ASN to improve amodal segmentation performance.

Comparing to most of the amodal and occlusion reasoning methods which regress single occluded object boundary directly on the input (single-layered) image, our BCNet decouples overlapping objects in the same ROI into two disjoint graph layers by predicting the complete object segments (Figure 1), where the occludee is segmented under the guidance from the shape and location of the occluder.

Occlusion-Aware Instance Segmentation

We first give an overview to the overall instance segmentation framework, and then describe the proposed Bilayer Graph Convolutional Network (BCNet) with explicit occluder-occludee modeling. Finally, we specify the objective functions for the whole network optimization, and provide details of training and inference process.

For images with heavy occlusion, multiple overlapping objects in the same bounding box may result in confusing instance contours from both real objects and occlusion boundaries. The mask head design of Mask R-CNN and its variants in Figure 3 directly regress the occludee with a fully convolutional network, which neglects both the occluding instances and the overlapping relations between objects. To mitigate this limitation, BCNet extends existing two stage instance segmentation methods, by adding an occlusion perception branch parallel to the traditional target prediction pipeline. Thus, the interactions between objects within the ROI region can be well considered during the mask regression stage.

Figure 4 gives the overall architecture of BCNet for addressing occlusion in instance segmentation. Following typical models for instance segmentation, our model has three parts: (1) Backbone with FPN for ROI feature extraction; (2) Object detection head in charge of predicting bounding boxes as instance proposals. We employ FCOS as the object detector owing to its anchor-free efficiency though our method is flexible and can deploy any existing fully supervised object detectors ; (3) The occlusion-aware mask head, BCNet, uses bilayer GCN structure for decoupling overlapping relations and segments the instance proposals obtained from the object detection branch. BCNet reformulates the traditional class-agnostic segmentation as two complementary tasks: occluder modeling using the first GCN and occludee prediction with the second GCN, where the auxiliary predictions from the first GCN provide rich occlusion cues, such as shape and positions of occluding regions, to guide target (occludee) object segmentation.

Work Flow

Given an input image, the backbone network equipped with FPN first extracts intermediate convolutional features for downstream processing. Then, the object detection head predicts bounding boxes with positions as well as categories for potential instances, and prepares the cropped ROI feature for BCNet to produce segmentation masks. The occlusion perception branch consists of the first GCN layer followed by FCN (two convolution layers), which is targeted for modeling occluding regions by jointly detecting contours and masks. Forming a residual connection, the distilled occlusion feature is element-wise added to the original input ROI feature and passed to second GCN. Finally, the second GCN, which has a similar structure to the first GCN, segments the occludee guided by this occlusion-aware feature and outputs contours and masks for the partially occluded instance.

2 Bilayer Occluder-Occludee Modeling

Recently, Graph Convolutional Network (GCN) has been adopted to model long-range relationships in images and videos . Given highly-overlapping objects, pixels belonging to the same partially occluded object may be separated into disjoint subregions by the occluder. Thus, we adopt GCN as our basic block due to its non-local property , where each graph node represents a single pixel on the feature map. To explicitly model the occluding region, we further extend the single GCN block to the bilayer GCN structure as shown in Figure 4, which constructs two orthogonal graphs in a single general framework.

Following , given an adjacency graph G=⟨V,E⟩\mathcal{G=\langle\mathcal{V},\mathcal{E}}\rangle with edges E\mathcal{E} among nodes V\mathcal{V}, we represent the graph convolution operation as,

To construct the adjacency matrix A\mathbf{A}, we define the pairwise similarity between every two graph nodes xi,xj\mathbf{x}_{i},\mathbf{x}_{j} by dot product similarity as,

where θ\theta and ϕ\phi are two trainable transformation function implemented by 1×11\times 1 convolution as shown in the non-local operator part of Figure 4, so that high confidence edge between two nodes corresponds to larger feature similarity.

In our bilayer GCN structure, we further define Gi\mathcal{G}^{i} to indicate the iith graph, XroiX_{roi} for the input ROI feature and Wf\mathbf{W}_{f} for weights in FCN layers, then the complete formulae are:

For connecting the two GCN blocks, the output feature Z0\mathbf{Z}^{0} of the occluder from the first GCN is directly added to Xroi\mathbf{X}_{roi} to obtain the fused occlusion-aware feature Xf\mathbf{X}_{f}, which is the input for the second GCN layer to output Z1\mathbf{Z}^{1} for occludee mask prediction.

Compared to previous class-agnostic mask head with single layer structure, where there is only binary label (foreground/background) per pixel, the bilayer GCN additionally constructs a new semantic graph space for occluding region. Thus a pixel node in overlapping areas in ROI can concurrently correspond to two different states in bilayer graph. While other choices may exist, we believe modeling GCN as a dual-layered structure as shown in Figure 4 is a natural choice for handling occlusion.

Occluder-occludee Modeling

We explicitly model occlusion patterns by detecting both contours and masks for the occluders using the first GCN layer. Since the second GCN layer jointly predicts contours for the occludee, the overlap between the two layers can be directly identified as occlusion boundary which can thus be distinguished from real object contour (e.g., the occluder and occludee prediction on the rightmost of Figure 4). The rationale behind this design is that such irregular occlusion boundary unrelated to the occludee is confusing, which in turn provides essential cues for decoupling occlusion relations. Besides, accurate boundary localization explicitly contributes to segmentation mask prediction.

The module for occluder modeling is designed in a simple yet effective way: one 3×\times3 convolutional layer followed by one GCN layer and one FCN layer. Then we feed the output to the up-sampling layer and one 1×\times1 convolutional layer to obtain one channel feature map for joint boundary and mask predictions. The boundary detection for occluder is trained with loss L′Occ-B\mathcal{L^{\prime}}_{\text{Occ-B}}:

where LBCE\mathcal{L}_{\text{BCE}} denotes the binary cross-entropy loss, Focc\mathcal{F}_{occ} denotes the nonlinear transformation function of the occlusion modeling module, WBW_{B} is the boundary predictor weight, Xroi\mathbf{X}_{roi} is the cropped FPN feature map given by RoIAlign operation for the target region, and GTB\mathcal{GT}_{B} is the off-the-shelf occluder boundary that can be readily computed from mask annotations.

For occluder mask prediction, it utilizes the shared feature Focc(Xroi)\mathcal{F}_{occ}(\mathbf{X}_{roi}), which is jointly optimized by boundary prediction. The segmentation loss L′Occ-S\mathcal{L^{\prime}}_{\text{Occ-S}} for occluder modeling is designed as

where WSW_{S} denotes the trainable weight of segmentation mask predictor by 1×11\times 1 convolutional layer, and GTS\mathcal{GT}_{S} is the mask annotations for the occluder.

3 End-to-end Parameter Learning

The whole instance segmentation framework can be trained in an end-to-end manner defined by a multi-task loss function L\mathcal{L} as,

where LOcc-B\mathcal{L}_{\text{Occ-B}} and LOcc-S\mathcal{L}_{\text{Occ-S}} denote respectively the boundary detection and mask segmentation losses in the second GCN layer for the occludee, which are similar to Eq. 7 and Eq. 8. LDetect\mathcal{L}_{\text{Detect}} supervises both the position prediction and the category classification borrowed from the FCOS detector,

and λ1\lambda_{1}, λ2\lambda_{2}, λ3\lambda_{3}, λ4\lambda_{4} and λ5\lambda_{5} are hyper-parameter weights to balance the loss functions, which are tuned to be {1,0.5,0.25,0.5,1.0}\{1,0.5,0.25,0.5,1.0\} respectively on the validation set.

For training the first GCN layer of BCNet, since partial occlusion cases only occupy a small fraction compared to the complete objects in COCO, we filter out part of the non-occluded ROI proposals to keep occlusion cases taking up 50% for balance sampling. SGD with momentum is employed for training 90K iterations which starts with 1K constant warm-up iterations. The batch size is set to 16 and initial learning rate is 0.01. In ablation study, ResNet-50-FPN is used as backbone and the input images are resized without changing the aspect ratio by keeping the shorter side and longer side of no more than 600 and 900 pixels respectively. For leaderboard comparison, we adopt the scale-jitter where the shorter image side is randomly sampled from following .

Inference:

During inference, the mask head predicts masks for the occluded target object in the high-score box proposals (no more than 50) generated by the FCOS detector, where the first GCN layer only produces occlusion-aware feature as input for the second GCN.

Experiments

We conduct experiments on COCO dataset , where we train on 2017train (115k images) and evaluate results on both 2017val and 2017test-dev using the standard metrics. For further investigating segmentation performance with occlusion handling, we propose a subset split, called COCO-OCC, which contains 1,005 images extracted from the validation set (5k images) where the overlapping ratio between the bounding boxes of objects is at least 0.2. Segmenting COCO-OCC with highly overlapping objects is much more difficult than 2017val, where we observe a performance gap around 3.0APAP for the same model in the experiment section.

KINS and COCOA

We also evaluate BCNet on two amodal instance segmentation benchmarks: (1) KINS , built on the original KITTI , is the largest amodal segmentation benchmark for traffic scenes with both annotated amodal and modal masks for instances. BCNet is trained on the training split (7,474 images and 95,311 instances) and tested on the testing split (7,517 images and 92,492 instances) following the setting in . (2) COCOA is a subpart of COCO , where we train BCNet on the official training split (2,500 images) and test on the validation split (1,323 images). Note that each instance has no class label and we only use the modal and amodal mask labels for the COCOA dataset.

Synthetic Occlusion Dataset

Since most objects in COCO do not exhibit significant occlusions, we synthesize a large-scale instance segmentation dataset which contains 100k images following uniform class distribution for instances among the 80 categories in COCO. Each synthetic image has true and complete object contours for both occluding and partially occluded objects, thus allowing the explicit modeling of occlusion relationship between the occlusion regions and occluded objects. On the other hand, COCOA , which has only 5,000 images, relies on user annotation on a given training image for “guessing” occluded object boundaries. More details on our occlusion dataset synthesis process are provided in the supplementary file.

2 Ablation Study

We validate the efficacy of different components proposed for explicit occlusion modeling on the first GCN layer. Table 1 tabulates the quantitative comparison: 1) Baseline: BCNet with no explicit occlusion modeling targets; 2) modeling segmentation masks for occluding regions (occluder); 3) modeling contours of the occluding regions; 4) joint occlusion modeling on both masks and contours. Compared to the baseline, joint occlusion modeling produces the most obvious improvement especially for the heavy occlusion cases, which promotes mask APAP on the standard validation set from 32.65 to 33.43, and the APAP on the proposed COCO-OCC split is increased from 29.04 to 30.37.

Effect of Bilayer Occluder-occludee Modeling

Built on the first GCN layer with explicit occlusion modeling, we further validate the second GCN layer in Table 2, which demonstrates the importance of occlusion-aware feature guidance for the second GCN layer to segment target object (occludee) by boosting 1.23 APAP on COCO-OCC, and 1.06 APAP on COCO respectively. Table 3 shows the results comparison on adopting the proposed bilayer structure and existing direct regression model with single layer. On the COCO-OCC split, bilayer GCN improves APAP from 29.63 to 30.68 compared to single GCN, and bilayer FCN boosts the performance of single FCN from 28.43 to 30.12.

Using FCN or GCN?

Table 3 also reveals the advantage of GCN over FCN, where GCN achieves consistent superior performance both in the singe layer and bilayer structure. We also compute the number of parameters of each model and find that although GCN has more trainable parameters, the increased model size is acceptable compared to performance gain, because the feature size of input ROI has been down-sampled to only 14×\times14 (spatial size) with 256 channels.

Influence of Object Detector

To investigate the influence of object detectors to BCNet, besides using one-stage detector FCOS , we also use representative two-stage detector Faster R-CNN to perform experiments. As shown in Table 4, the performance gain brought by BCNet is consistent, with an improvement of 2.23 (for FCOS) and 2.04 (for Faster R-CNN) mask APAP on COCO-OCC respectively. Here, baseline denotes mask head design in Mask R-CNN.

3 Performance Comparison and Analysis

Table 8 compares BCNet with state-of-the-art instance segmentation methods on COCO dataset. BCGN achieves consistent improvement on different backbones and object detectors, demonstrating its effectiveness by outperforming both PANet and Mask Scoring R-CNN by 1.5 AP using Faster R-CNN, and exceeding CenterMask by 1.3 AP using FCOS. Our single model achieves comparable result with HTC , which uses a 3-stage cascade refinement with multiple object detectors and mask heads, and far more parameters.

Comparison with Amodal Segmentation Methods

Table 7 and Table 7 compare BCNet with other SOTA amodal segmentation methods on both the COCOA and KINS datasets, where: 1) AmodalMask directly predicts amodal masks from image patches; 2) Occlusion RCNN (ORCNN) is an extension of Mask R-CNN with both amodal and modal mask heads; 3) ASN module contains additional occlusion classification branch and multi-level coding. Compared to these occlusion handling approaches, our bilayer GCN with cascaded structure still performs favorably against the state-of-the-art methods, which shows the effectiveness of BCNet in decoupling overlapping objects and mask completion under the amodal segmentation setting. Figure 5 and Figure 6 show the qualitative comparison on COCOA and KINS respectively.

Evaluation on Occluded Images

We adopt COCO-OCC split to compare the occlusion handling ability of BCNet with other methods on images with highly overlapping objects. As shown in Table 7, our BCNet with Faster R-CNN detector has 31.71 AP vs. 30.32 for the Mask Scoring R-CNN . By further training BCNet on the synthetic occlusion dataset, the performance of AP and AP50 is significantly promoted to 32.89 and 53.25 respectively, which shows the advantage brought by this new dataset.

Qualitative Evaluation.

Figure 7 shows qualitative comparison of CenterMask and BCNet on images with overlapping objects. In each ROI region, GCN-1 detects occluding regions while GCN-2 models the partially occluded instance by directly regressing the contours and masks. For example, BCNet decouples the occluding and occluded baseball players in similar clothes into GCN-1 and GCN-2 respectively, and detects the left leg missed by CenterMask. See supplementary file for more visual comparisons.

Conclusion

We propose BCNet, an effective mask prediction network for addressing instance segmentation in the presence of highly-overlapping objects in two-stage instance segmentation. BCNet achieves consistent gains on overall segmentation performance using different backbones and object detectors in both the modal and amodal settings. With explicit occluder-occludee modeling, occluding and occluded instances are decoupled into two disjoint graph spaces, where the interaction between objects within each ROI region are explicitly considered. This effective approach will benefit future research in both occlusion handling and instance segmentation.

References