Mask Encoding for Single Shot Instance Segmentation
Rufeng Zhang, Zhi Tian, Chunhua Shen, Mingyu You, Youliang Yan
Introduction
Instance segmentation enables various visual applications like autonomous driving and robot navigation, to name a few. Instead of separately detecting objects or assigning category labels to pixels, instance segmentation unifies these tasks together, thus being one of the most challenging tasks in computer vision.
Recent advances in deep convolutional neural networks (CNNs) have enabled tremendous progress in instance segmentation, e.g., . One of the mainstream methods employs a two-stage pipeline that first generates proposals and then performs pixel classification within each proposal, as popularized by Mask R-CNN . Almost all the methods in the top rank on the challenging COCO benchmark are built upon Mask R-CNN thus far. One drawback of these two-stage solutions is not sufficiently efficient as their runtime is constrained by the number of instances in an image. On the other hand, one-stage paradigms process the full image straightforward, making the speed stable no matter how many objects present.
Several works have attempted to incorporate mask prediction into fully convolutional networks (FCNs) , resulting in single shot instance segmentation frameworks. These algorithms share a common insight, i.e., encoding the object shape with a set of contour coefficients. Specifically, ESE-Seg designs an “inner-center radius” shape signature for each instance and fits it with Chebyshev polynomials. Concurrently, PolarMask regresses the dense distance of rays between mass-center and contours. These contour-based methods enjoy the advantages of easy optimization and fast inference. The major issue of these methods is that the predicted masks may exhibit “hollow decay” inevitably, since they can only depict instances with a single contour, as shown in Figure 1.
Alternatively, a non-parametric mask representation is more natural for mask prediction as traditionally done, with the price of increasing both design and computation complexity. As natural object masks are not random and akin to natural images, instance masks reside in a much lower intrinsic dimension than that of the pixel space. This inspires us to ask a question, “Is it possible to predict the object mask in the intrinsic low-dimensional space and still achieve competitive accuracy?” Here we provide an affirmative answer: we propose to encode instance masks using a learned dictionary such that only a few scalar coefficients are needed to represent each mask. We demonstrate that such an approach is robust to noise, and efficient, easy to decode for reconstruction.
Then an one-stage detector such as RetinaNet , FCOS can be easily extended by adding a branch for predicting these fixed-dimensional mask coefficients, along with the bounding box regression and category classification branches. We build our method on top of FCOS for its simplicity and good detection performance.
We demonstrate that our method can outperform recent one-stage algorithms with this simple design. In particular, experiments on the COCO val2017 show that MEInst achieves a large gain compared to ESE-Seg , outperforming by in AP50 and in AP75, respectively. Our model beats PolarMask in accuracy with similar computational complexity, owing to the lower reconstruction error and more effective reconstruction. This is expected, as the mask representation of our method is more powerful than the parametric representation of .
Additionally, we take a closer look at how the object detector influences the performance of instance segmentation based on extensive qualitative experiments. With a careful design based on our finding, MEInst achieves comparable performance with Mask R-CNN with the advantage of being much simpler and flexible.
It is noteworthy that our method is compatible with most one-stage detection frameworks including the anchor-free paradigm. We demonstrate its generality using the FCOS detector, and evaluate the performance on the COCO benchmark . Other anchor-based methods such as YOLO , RetinaNet may be used here with minimum modification. Moreover, the vanilla detectors can also benefit from the paralleled mask prediction branch, improving the bounding box detection accuracy.
The main contributions of this work can be summarized as follows.
We propose to encode a two-dimensional instance mask into a compact representation vector. The compressed vector, takes advantages of the redundancy in the original mask and proves to be effective and efficient for reconstruction.
Encoding can be done with a few dictionary learning methods, including PCA, sparse coding, and auto-encoders. Here we show that even the simplest PCA already suffices for mask encoding.
With this mask representation, a new framework is introduced for single shot instance segmentation, termed mask encoding based instance segmentation (MEInst), by extending FCOS with a mask branch for mask coefficient regression. Actually, our mask encoding is completely independent of the mechanism of detectors, and it may be easily incorporated into other detectors.
We demonstrate a simple and flexible one-stage instance segmentation method. Our best model, attains mask AP of on COCO test-dev, achieving a good balance between accuracy and speed.
Related Work
We review a few works that are most relevant to ours.
Two-stage Instance Segmentation The mainstream approaches to instance segmentation inherit the pipeline of two-stage object detectors, as pioneered by Mask R-CNN . These methods typically detect instance bounding boxes and then perform binary-class segmentation in boxes. Compared with segmentation-driven ones , this group of paradigms lead on most benchmarks in accuracy. In particular, Mask R-CNN replaces ROIPool with ROIAlign to better align features. Following Mask R-CNN, Liu et al. present bottom-up path augmentation and adaptive feature pooling for further feature optimization. Mask Scoring R-CNN extends Mask R-CNN with an extra MaskIoU branch, aiming to calibrate the mismatch between mask’s quality and the corresponding confidence. The above methods consistently advance the performance.
One-stage Instance Segmentation The second family of solutions are built upon the success of semantic segmentation, i.e., generating pixel-wise classification maps firstly and then clustering them into instances. Specifically, InstanceCut addresses the problem with two paralleled sub-tasks, instance-agnostic segmentation and instance-specific boundaries.
In the meantime, dense object segmentation has not witnessed remarkable progress. Impressively, several works have attempted to fill in the gaps lately. For example, TensorMask can be viewed as a precursor to this group of algorithms, in which a structured 4D tensor is introduced to represent the mask over a spatial domain. It achieves similar performance with two-stage methods with the cost of heavy computation overhead in training and testing. In YOLACT , a series of global prototypes and individual linear coefficients are assembled for masks, achieving a real-time speed. BlendMask improves YOLACT in both accuracy and speed. Recently, Xie et al. propose a general framework named PolarMask , which is capable to directly predict the mask without bounding box using a parametric representation of masks. More recently, SOLO and its improved version SOLOv2 demonstrate promosing results with a simple FCN-like framework .
Our Method
In this section, we first present the overall architecture of MEInst. We then introduce the instance representation with mask encoding and its optimization. Finally, we explore the correlation between detection quality and mask generation to further improve the performance of MEInst.
The object detection modules in our method mainly inherit the pipeline from FCOSWe use the improved version, including sharing the features between center-ness and regression branch, central sampling and so on. Please refer to for further details. for its flexibility and simplicity, including a backbone module , a feature pyramid module , and two task-specific heads for classification, box regression and center-ness (they share the same head). Then a parallel branch is included for predicting encoded mask coefficients. Additionally, we carefully re-design some parts of the framework, which further boosts the performance. Details are discussed in the following subsection. The overall framework is illustrated in Figure 2.
2 Mask Encoding
Given a structured instance mask, we can easily figure out the redundancy in its representation. An example can be seen in Figure 3(b). The discriminative pixels are mainly distributed along the object boundaries while most pixels in its body hold the properties of being category-continuous and category-consistent. In other words, the existing mask representations contains redundant information and it may be highly compressed with negligible loss. In this subsection, we describe how to encode the two-dimensional geometry into a much more compact representation vector in detail.
We follow the strategy in DUpsampling and optimize this objective by using principal component analysis (PCA). The overall process is illustrated in Figure 3. Please refer to DUpsampling for details. There may be alternative options to minimize the reconstruction loss, e.g., sparse coding or non-linear auto-encoder.
Loss Function We define our mask loss function as follows:
Here is the loss for detection, consisting of for classification, for bounding box regression and for center-ness. In particular, is focal loss as in , is the GIoU loss following FCOS . denotes the binary cross entropy (BCE) loss for center-ness. All the balance weights in are set to for simplicity in our experiments.
3 Correlation Between Boxes and Masks
In general, instance segmentation and object detection are inseparable in detection-driven pipelines. Intuitively, better bounding boxes improves the overall performance in the mask branch. Here we carry out several experiments to validate our assumptions empirically.
Take Mask R-CNN as an example. The inference flow is as follows: 1) A backbone module is used to extract semantic feature from the input image. 2) The extracted feature is then sent to the following modules for classification and object regression. 3) Afterwards, the mask stage computes features using ROIAlign from each detected box. 4) Finally, the regional representation is performed pixel-wise segmentation. It only predicts a binary mask.
In our experiments, the Mask-R-50-FPN model pre-trained by He et al. is used as the main backbone. The step-2 in the above process is replaced with a series of pre-acquired detection results predicted by different detectors, in which case all the variables are kept the same except the boxes. Here we choose Mask R-CNN (two-stage) and FCOS (one-stage) with different backbones as object detectors. In the sequel, AP means mask AP and box AP is denoted as APbb. The quantitative results are shown in Table 1 and Figure 4.
As for the same architecture, the detector brings consistent and noticeable gain in mask when the network goes deeper. However, the results of instance segmentation fall below our expectations with different pipelines. Compared with Mask R-CNN, FCOS achieves better detection performances among all backbones under the metric APbb, measuring , , , respectively. Nevertheless, the corresponding segmentation has not been witnessed equivalent improvement, and even performs worse ( vs. ). It seems counter-intuitive.
We observe that FCOS performs better under all the general metrics except AP, which indicates that the boxes predicted by FCOS are location-accurate but with more false-positive (FP). Figure 4(b) shows the average number of bounding boxes predicted by different models. FCOS predicts significantly more bounding boxes than Mask R-CNN with the same confidence threshold (e.g., ), which may degrade the performance under the metric AP. Mask R-CNN employs a two-stage pipeline, i.e., first proposes candidates and then refines the boxes, in which case most mis-proposed boxes can be filtered out effectively. However, one-stage paradigm such as FCOS outputs results directly for faster inference, resulting in the redundant boxes. Actually almost all the one-stage methods suffer from this dilemma.
We hypothesize that the issue may be related to the effective receptive field (ERF). Zhou et al. declare that the effective receptive field is much smaller than the theoretical receptive field, since CNN tends to capture information from central regions. The insufficient ERF may lead to many false-positive (FP) boxes as the network can not “see” the objects. To tackle this issue, we simply employ deformable convolution that has the capacity to focus on salient regions and enlarge the ERF to some extent. Specifically, we replace the last vanilla convolutional layer in multi-head branches respectively. Note that other modules such as dilated convolution and Large Kernel , which are beneficial to ERF, may also boost the performance. We provide further comparisons in the experimental section.
Experiments
Our experiments are conducted on the challenging MS COCO benchmark using the standard metrics for instance segmentation. All models are trained on the COCO train2017 split (118k images) and evaluated with val2017 (5k images). The final results are reported on test-dev (20k images). Moreover, we adopt the training strategy , single scale training and testing unless otherwise specified.
Training Details ResNet-50 is used as the backbone network and all hyper-parameters are kept consistent with FCOS unless specified. Specifically, we use the stochastic gradient descent (SGD) optimizer, weight decay 0.0001, momentum 0.9 with 90K iterations in all. The initial learning rate is set to and divided by 10 at iteration 60K and 80K, respectively. We use a mini-batch of 16 images and all models are trained with 8 GPUs. The backbone is initialized with the pre-trained weights on ImageNet and other newly added layers are initialized as in . The shorter side of images is fixed as 800 pixels with the longer side being 1333 or less. Moreover, we sum up all the losses directly, i.e., in Eq. (4). We expect that the performance may be better with a careful parameter tuning.
Inference Details The inference process is kept the same as FCOS since we only append one more prediction to the predicted boxes. An input image goes through the network and then predicts boxes with several attributes, such as categories and mask coefficients. We peform mask reconstruction after non-maximum suppression (NMS) to avoid unnecessary computational overhead (the highest scoring 100 samples). Since the matrix multiplication is fast, MEInst introduces slight overhead to its FCOS counterpart.
Analysis of Upper Bound We first reshape all the annotations into binary-class masks. Afterwards, these masks are encoded and recovered to two-dimensional matrices with Eq. (1). Finally we use the metric of mIoU to evaluate the quality of reconstructed masks. The reconstruction error on the COCO train2017 split is shown in Figure 5. It is evident that the reconstruction error goes down consistently with the increase of the number of components kept, and can even reach an extremely low level when the dimension goes to 100 (only ). Moreover, we observe that the class-agnostic matrix achieves a similar result to class-specific one (up to times in dimensions). Thus, the former is a better choice for memory-conserving consideration.
Dimension of Encoding Representation It plays a very fundamental role in MEInst. As shown in Table 2, the performance grows steadily with the increase of dimension and reaches saturation at last. For example, there is an improvement of from 20 to 60 and it remains stable beyond 60. The reconstruction has a great influence at the beginning. However, when adequate components can reconstruct the mask well, it is no longer the main factor constraining the performance. We choose in our experiments unless otherwise specified.
Learning without Explicit Encoding Alternatively, the mask can be learned without explicit encoding. That is, instead of compressing the redundant label into a fix-dimensional vector, we recover the predicted mask with the reconstruction matrix and perform pixel-wise classification on it. This projecting process is essentially identical to employing a convolution along the spatial dimensions, with convolutional kernels stored in . Note that these parameters are frozen during training. Moreover, we also explore the potential of learning without mask encoding, i.e., the network straightly outputs the high-dimensional masks (e.g., ). The results are shown in Table 3. The over-high dimension makes it hard to optimize, resulting in a performance drop. Particularly, AP75 and APL decrease considerably, measuring by and , respectively. The relatively compact vector is not only for faster inference, but also beneficial for optimization. With the same dimension, our method still performs better under all the metrics, which further proves the effectiveness of mask encoding.
Loss Function As discussed above, mask encoding converts the task of instance segmentation into a set of coefficient regression problems. We try several popular losses in our experiments to supervise the regression problems, more specifically, smooth- loss, loss and loss. in Eq. (4) is set to 1 for simplicity. As shown is Table 4, loss performs better than others. We also consider the case to view the mask vector as a whole, so we apply cosine similarity loss. However, the performance goes worse, which indicates that mask encoding has already eased the redundancy in original representation, and now the elements in vectors are independent.
Large Receptive Field Here we demonstrate the importance of large receptive field. Firstly, we apply large kernel (LK) in the mask prediction layer. The LK layer is a combination of and convolutions. is set to 9 in our experiments. Compared with convolution, it introduces negligible overhead. As shown in Table 5, LK in prediction layer achieves 0.7% AP gains. We also explore the potential of deformable convolution (DCN). Specifically, we only use it in the last layer of head to keep our model efficient. With the ability of capturing more meaningful and larger receptive features, it obtains 1.5% improvement in AP.
Learning Masks boosts Object Detection As mentioned in , learning with instance mask prediction can usually boost the performance of one-stage detectors. We also find the similar phenomenon in our experiments, i.e., our MEInst outperforms FCOS by AP in box, as demonstrated in Table 6. Compared with RetinaMask which employs a few tricks, our method is simpler yet achieving the same performance.
Mask-Based vs. Contour-Based We compare MEInst against the recent contour-based method termed ESE-Seg . To make this a fair comparison, we do not apply any deformable convolutions in our model. As shown in Table 7, MEInst shows a large gain compared to the ESE-Seg method. Additionally, when the input scale becomes smaller (e.g., 400), our model still achieves a better performance at a real-time speed. Note that we do not specifically train a new model here. It indicates that MEInst can not only achieve good performance in mask AP, but also shows promises for real-time applications. Besides the performance, our mask-based method also shows a detail-preserving advantage that ESE-Seg lacks, which is illustrated in Figure 1. Experiments demonstrate that the proposed method enjoys desirable properties comparing with contour-based algorithms such as PolarMask and ESE-Seg .
2 Comparison with State-of-the-art Methods
We evaluate MEInst on COCO test-dev and compare our results with some state-of-the-art methods, including both one-stage and two-stage models. The results are shown in Table 8 and Figure 6. Without bells and whistles, MEInst achieves a mask AP of , which outperforms most one-stage methods by a large margin. Note that we do not use any tricks in our experiments, e.g., auxiliary semantic segmentation supervision. Our performance may be further improved with those tricks. Moreover, the gap between TensorMask and ours is mainly because 1) Tensormask uses a very long training schedule, as well as 2) bipyramid and aligned representation. Considering that these modules are time- and memory-consuming, we do not plug them into our model.
3 Advantages and Limitations
MEInst has the capacity to better deal with “disjointed” objects. An example can be found in Figure 6 (row 3 column 1).
An interesting phenomenon is that, MEInst surpasses Mask R-CNN when the detected object is small ( vs. ) while performs worse when the object becomes larger ( vs. ). We argue that the main reasons are two folds:
For small objects, the capacity of the single feature vector in our work is not a problem. While in Mask R-CNN, it requires the mask prediction head to label each pixel of a small object, which is challenging when the object is very small. That is why we outperform Mask R-CNN for small objects.
As for large objects, a compact representation vector is difficult to accommodate all the details of the mask. In this case, non-parametric pixel labelling shows advantages. Additional modules to encode details are needed in this case.
Conclusion
In this work, we have introduced a new, simple single-shot instance segmentation framework termed MEInst. Different from previous works that typically solve mask prediction as binary classification in a spatial layout, MEInst represents the mask with a fixed-dimensional and compact vector, and casts the task into a regression task. The reformation allows the challenging task to be solved by appending a parallel regression branch to existing one-stage object detectors. Experimental analyses demonstrate that the proposed framework achieves competitive accuracy and speed among one-stage paradigms. In the future, we will explore the possibility of using other dictionary learning methods for encoding instance masks, and the possibility of applying this idea to other instance recognition tasks.