Real-time Instance Segmentation with Discriminative Orientation Maps
Wentao Du, Zhiyu Xiang, Shuya Chen, Chengyu Qiao, Yiman Chen, Tingming Bai
Introduction
Instance segmentation aims at pixel-wise predictions for every individual object. It integrates instance-level object detection and pixel-level semantic segmentation , formulating a more fine-grained visual perception task. Currently there are two dominating types of solutions, namely detection-based and segmentation-based methods. The former extends an object detector with additional foreground dense predictions while the latter deploys specific per-pixel attributes or embeddings to separate instances of the same category in a bottom-up way.
Both paradigms have apparent drawbacks. Conventional detection-based methods like Mask R-CNN rely on the features pooling operation to project all regions of interest (RoIs) into a fixed size. Since the subsequent mask head should be applied to abundant feature maps of every region proposal, the speed is largely constrained especially when objects densely appear. Moreover, the constant mask resolution brings in unnecessary computation for small objects and loses valuable details for large targets. On the contrary, segmentation-based methods retain the fine-grained appearance and geometry in a pixel-to-pixel manner. They can acquire satisfactory results at elementary scenarios but often fall behind detection-based approaches in accuracy. When the scale of objects varies and the number of categories increases, the generalization of pixel-level clustering adopted in segmentation-based methods is still in doubt.
For the requirement of real-time inference, YOLACT is proposed along with a special mask construction scheme, which linearly combines shared non-local prototypes with instance-wise coefficients. It discards the RoI pooling operation that commonly adopted in earlier detection-based methods and directly assemblies masks from fine-grained feature maps. Following this paradigm, an improved approach named BlendMask is put forward. It replaces the 1D instance-specific coefficients with a set of attention maps, which supply additional spatial-adaptive information to enrich fine-grained details of masks. The success of these solutions shows great potential to incorporate informative global features into detection-based methods. However, one noticeable flaw of these approaches lies in the dependence of RoI cropping operations when generating the assembled masks, which may bring in some mask incompleteness due to inaccurate bounding box predictions.
In this work, we attempt to integrate fine-grained expressions with an one-stage detector in another way. To be specific, we focus on compact mask representation and efficient integration with the anchor-based detector YOLOv3 to achieve real-time performance. First, a novel discriminative orientation map is proposed to encode multiple masks independently, where pixels are assigned with centripetal or centrifugal vectors according to their positive or negative labels. This design is totally free from any other semantic segmentation or foreground predictions and is lightweight to decode the complete masks. In addition, considering objects of diverse scales vary the magnitude distributions of orientation vectors, multi-scale design is also taken into account. We assign different orientation maps for instances matching with certain anchor sizes so that the completeness of mask representation is guaranteed. OrienMask merely extends an extra head to the object detector and its function is tightly combined with the box assignment and pre-defined anchor sizes. During inference, for each predicted bounding box, its instance mask can be quickly constructed based on the discriminative vectors in the corresponding orientation map, as shown in Figure 1. This process is simple and direct, consisting of nothing but determining binary labels for all pixels by the spatial destinations that orientation vectors indicate. The main contributions of our work can be summarized as follows:
We put forward a light and discriminative orientation-based mask representation for real-time instance segmentation. By defining opposite orientation vectors for foreground and background pixels, we are able to effectively encode multiple instance masks in a fine-grained two-channel map without the need for explicit foreground segmentation. In inference, given the target regions of instances, their masks can be easily constructed from orientation maps in parallel.
To deal with objects with various sizes, we propose an instance grouping mechanism derived from anchor-based detectors. Each group of instances with similar sizes are assigned to share a common class-agnostic orientation map. We also expand the annotated bounding boxes to provide sufficient supervision for background. The enlarged valid training area not only balances the number of positive and negative samples but also helps distinguishing them around the boarders.
We integrate the discriminative orientation maps into a fast anchor-based detector YOLOv3 and implement the resulting model OrienMask end-to-end. Experiments show that it is able to achieve 34.8 mask AP at a speed of 42.7 fps on COCO benchmark, which is quite competitive among state-of-the-art real-time methods.
Related Work
Detect-then-segment Methods With the help of proposals generated by object detectors, detect-then-segment methods first extract reliable RoIs from feature maps, and then obtain fine-grained instance representations. Driven by the success of two-stage detector Faster R-CNN , Mask R-CNN adds a branch for mask prediction in parallel with bounding box regression, and employs RoIAlign to fix the misalignment caused by spatial quantization. After that, PANet is presented to strengthen message passing in the bottom-up path and fuse pooling features from all levels. HTC extends Mask R-CNN into a cascade structure by interleaving the mask and box branches while maintaining semantic features fusion. Instead of using the confidence of the detector, Mask Scoring R-CNN predicts an extra score to accurately represent the mask quality. However, due to the heavy computations in the second stage, these approaches can hardly satisfy the requirement of real-time inference.
Detect-and-segment Methods Benefiting from those compact architectures of one-stage object detection , detect-and-segment methods customize masks from the global feature maps jointly with implicit instance-specific representations. In YOLACT , acknowledged as the milestone of this paradigm, a series of mask coefficients are produced along with the box predictions. Then they are multiplied with a set of high-resolution prototypes to generate instance masks. Chen et al. rethink the trade-off between feature resolution and coefficients dimension, and propose BlendMask, which blends some attention maps of each instance with a group of shared bases. Inspired by conditionally parameterized convolutions , CondInst predicts instance-aware convolutional kernel weights and applies them on high-resolution feature maps. Thanks to the flexible framework and fine-grained representation, these methods keep a good balance between speed and accuracy. Compared with them, our OrienMask employs an explicit and discriminative features sharing scheme for mask representation rather than implicit parameterized forms, which is more concise and provides strong interpretability.
Other Mask Representations Apart from bounding boxes and foreground probability maps, some compact mask representations also contribute to instance segmentation. For example, Jetley et al. deploy an auto-encoder to compress masks into low-dimensional vectors which can be incorporated in a detector. Xu et al. describe an instance as a series of inner-center radii and encode them into Chebyshev polynomial coefficients. To achieve better precision, polar centerness and polar IoU loss are proposed in PolarMask . Peng et al. implement the snake algorithm in a learning-based form and the circular convolution is proposed to iteratively regress sampling points towards contour positions. Serving as effective representation, pixel offset and its variants are popular in separating instances within the segmented foreground. Uhrig et al. deploy template matching by predicted depth classes and discrete directions to assign pixels of the same semantic label to different instance centers. Box2Pix matches the foreground pixels with predicted box centers according to offset vectors. Similarly, Li et al. merge foreground pixels based on adaptive voting zones around detection centers. Neven et al. propose a joint optimization scheme for class-specific seed and sigma maps along with dense offset vectors. Given the learned clustering bandwidth, masks are sequentially recovered. PersonLab utilizes the short-range and mid-range offset vectors to decode person poses, and then clusters foreground pixels by long-range offsets. Novotny et al. propose a semi-convolutional operator, which adds coordinates to part of learned embedding. PointGroup extends offset descriptor to 3D instance segmentation, where points of the same category are grouped step by step. Our method also draws inspiration from spatial offset descriptor. However, unlike above methods which mostly use spatial offset to assign segmented foreground pixels to instances, our orientation map is self-discriminative. It is able to filter out background regions and separate instances at the same time. To achieve this goal, special valid training areas are defined and distinctive orientation vectors for both positive and negative samples are considered. Besides, these fine-grained orientation maps are tightly bonded with anchors of the detector, which retains the mask completeness of different sizes and eases the regression.
OrienMask
The network architecture of OrienMask is mainly built upon anchor-based detector YOLOv3 with the backbone Darknet-53. As illustrated in Figure 2, we add an extra OrienHead to predict orientation maps, which is the key part of the whole framework. The obtained bounding boxes and orientation maps are combined for mask construction.
YOLOv3 deploys 9 bounding box priors and evenly assigns them across 3 scales. Given an image with height and width as input, the BoxHead at the scale of output stride produces bounding box predictions, where denotes the number of anchors per grid cell. OrienHead takes the largest feature maps P2 after feature pyramid network (FPN) as input and then predicts orientation maps of fixed resolution with two channels, matching with different anchor sizes respectively. Given that handling feature maps of high resolutions is time-consuming, our OrienHead is designed to be light-weight. It is interleaved with three and convolution layers of input channels 128 and 256 alternatively. After the standard non-maximum suppression, each bounding box is paired with an orientation map according to its anchor size. As demonstrated in the right part of Figure 2, all pixels whose orientation vectors end within a contracted bounding box form a foreground mask.
2 Orientation-based Mask Representation
Valid Training Area We first expand regions enclosed by annotated bounding boxes to form valid training areas. Pixels outside of any expanded region are ignored during training, which means they are not involved in the loss calculation. All remaining valid pixels are divided into two parts, namely positive and negative samples, based on whether they are covered by instance masks or not. Since the number of positive samples is constant and more negative samples will be counted when the valid training area expands, a proper expand ratio should be determined to balance the number of positive and negative samples. Meanwhile, this expansion also provides sufficient guidance for distinguishing pixels nearby the instance boarders.
Orientation Vectors To easily distinguish positive and negative samples in mask construction process, the orientation vectors of them are defined to point at opposite directions. To be specific, a base position is first specified for each instance and the bounding box centroid is a preferable choice in our experiments. The positive samples on the orientation map are defined pointing to the base position while the negative ones should point to the boarder of the valid training area in centrifugal directions. Denoting the base position as , the target orientation vector for pixel at location can be expressed as
Since orientation vectors are locally defined as spatial offsets pointing to some neighboring destinations, they vary smoothly in both directions within the interior places. The harmony may be slightly disturbed when valid training areas of different instances overlap but the overall unity is still maintained. For each pair of neighboring pixels locating across the mask boundary, the positive is pulled to the base position while the negative is pushed towards outside to the expanded boarder. Thus one direction of gradients at these positions equals to the distance between the base position and the corresponding boarder of valid training area, which is significantly larger than one pixel difference in other interior places. Our experiments in Section 4.3 do prove that the learned orientation maps retain this property, which is helpful to accurately delineate instance masks.
Instance Grouping Although simply stacking all instance masks onto a two-channel orientation map is fascinating for real-time considerations, it could fail in handling some instance overlapping scenarios, such as a person wearing a tie. To alleviate this problem, we introduce an instance grouping mechanism. Noticing that YOLOv3 assigns objects to different anchor sizes based on the intersection over unions with those bounding box priors, we naturally transfer this assignment to our mask representation. To be specific, instances are divided into several groups according to the anchor sizes that they are matched, and each group of instance masks are assigned to an independent orientation map.
Besides solving the overlapping problem caused by objects of different aspect ratios or scales, the instance grouping mechanism has more additional advantages. Since bigger objects often require larger receptive field, this arrangement is beneficial to adapt each orientation map to appropriate scale. Meanwhile, the orientation vectors of grouped instances can be normalized within a small interval so that the magnitude distribution of each group does not vary significantly, which is good for the network training. Moreover, our design also complies with the observation that an image may contain many small instances but only a few large ones. Hence it can preserve as many objects as possible.
3 Mask Construction
The mask construction process involves two elements: a predicted bounding box and an orientation map . Recall that each bounding box prediction has an anchor size and each anchor size is associated with an orientation map. Therefore, each bounding box is bound to be matched with an orientation map. Supposing that and have been paired, we take the centroid of as the base position according to the definition in Eq. (1). Then a rectangular target region centered on the base position is defined, whose size is proportional to the width and height of . If we denote the base position as and the size of bounding box as , the constructed mask can be expressed as
4 Loss Function
The loss function consists of two components, providing supervisions for object detection and orientation maps respectively. It can be formulated as
where is a hyper-parameter to balance these two terms. is completely copied from official YOLOv3 without any tricks in the literature.
With regard to , we compute smooth-l1 loss per pixel in valid training areas, and then take the average over positive and negative samples respectively. In addition, we multiply them by the number of instances in accordance with . The complete expression is written as
Experiments
We conduct experiments on the challenging MS COCO dataset and evaluate the predictions with standard metrics. Following the common practice, all models are trained with 118k images of train2017 and tested on 5k images of val2017 or 20k images of test-dev subset.
Training Details For the network structure, we retain the official implementation of YOLOv3 and extend a fully convolutional OrienHead as described above. The backbone Darknet-53 is initialized with a pretrained detector and the network is trained end-to-end. We employ stochastic gradient descent (SGD) optimizer with momentum 0.9 and weight decay 0.0005. The batch size is 16 and the synchronized batch normalization is used in our ultimate model but not in ablation study. The initial learning rate is 0.001 and divided by 10 at iterations 520k and 660k respectively. All models are trained for 100 epochs with the input resolution . Multiple data augmentations are applied, such as color jitter, random resizing, and horizontal flipping.
Inference Details Similar to YOLACT , input images are directly resized to without test-time augmentation. The inference speed is evaluated on RTX 2080 Ti by default and measured with frames per second (FPS).
In our ablation experiments, implementations of the object detector is fixed. We adjust other hyper-parameters of our method to obtain the best configurations. To integrate OrienHead with the detector more tightly, some additional refinements will also be applied to the base model.
Valid Training Area For orientation maps, the definition of negative samples is closely related to the boarder of valid training area. Given that the size of each valid training area is proportional to its bounding box, we vary the expand ratio from 1.0 to 1.6 with the stride 0.2. As the experimental results shown in Table 1, our model achieves best performance when and either smaller or larger expand ratio brings some drops in AP. We notice the amount of negative samples and their magnitudes of orientation vectors increase simultaneously as valid training areas expand. A moderate expand ratio keeps a balance between the number of positive and negative samples while maintains enough differentiation around boundaries. It also keeps a proper numeric distribution of negative samples. These two aspects both contribute to better convergence of our model.
Orientation Loss Weight In the loss function of Eq. (4), and are used to provide box-level and pixel-level supervisions respectively. In order to associate these two terms together, we explore the orientation loss weight from 5 to 20 and obtain the results in Table 2. The mask AP metrics gradually increase as becomes larger. However, we find the training process gets unstable and the performance saturates when applying weights larger than 20. Hence is adopted in our subsequent experiments. Moreover, we observe that the adjustment of the orientation loss weight does not disturb the box-level performance remarkably and even has a bit positive effects, which indicates OrienHead maintains the network stability to some extent.
Orientation Target Area For any bounding box that survives NMS, the mask is simply constructed by collecting all pixels whose orientation vectors point to somewhere close to its base position, without any other operations like RoI cropping. The bounding box is contracted by a scale factor to form an orientation target area so that it is compatible with objects of different aspect ratios or scales. Here we choose contraction ratio from 0.4 to 0.8. From Table 3 we find the performance is sensitive to the contraction ratio and the best performance is obtained when . More specific parameter tuning for each group of instances is optional to achieve even higher AP.
Other Improvements We further explore some measures to better integrate OrienHead with the detector and improve the overall performance. These measures are applied step by step and the results are illustrated in Table 4. For the orientation definition, we initially choose the grid centers as base positions, which keeps consistent with the rule of box regression in the detector. Since the box centroid is more accurate to locate an instance, we adopt it as the base position. This refinement improves the mask AP metric from 32.5 to 33.3. Then we are inspired by the next generation of YOLO framework to adopt larger anchors, which are proved more suitable for the given input resolution. It should be noted that no extra training tricks are taken and we still maintain other settings as pure YOLOv3. Benefiting from the larger anchors, the performance is boosted by 0.5 AP. In earlier experiments, we follow the standard FPN to produce P2 for OrienHead but the predicted orientation maps are related with box predictions from multiple scales. To closely associate these two outputs, we merge multi-scale pyramid features to predict orientation maps while keeping the same network structure in subsequent layers, as dashed lines in Figure 2 indicate. The resulting model surpasses the previous once again with negligible extra computing cost.
2 Comparison with State-of-the-art Methods
We first evaluate our OrienMask on the canonical COCO test-dev benchmark and select a series of representative frameworks for comparison. From the quantitative results displayed in Table 5, we find OrienMask is both faster and more accurate when compared against those having similar input resolutions like YOLACT and CenterMask, and some methods that aim at simplified mask representation like PolarMask and MEInst. We admit that OrienMask falls behind some non-real-time approaches with either higher input image resolutions or more complicated pipelines. When considering the twice or even three times faster in inference, it seems reasonable to accept some sacrifice in accuracy.
Since real-time instance segmentation is the main motivation of our work, we further compare OrienMask with state-of-the-art methods capable of real-time inference on COCO val2017. All models adopt relatively shallow backbones along with small input resolutions, which makes the comparison fair and persuasive. As shown in Table 6, Our method surpasses YOLACT with 5.6 AP at the cost of 4.0 fps slower. Except for that, OrienMask serves as the leading method in speed comparison and outperforms most counterparts in the mask AP metric. It reaches a good balance between efficiency and accuracy. We also calculate the memory occupation of top feature maps that used for mask construction, and record the results in the ‘space’ column. The input resolution of all methods is assumed to be fixed as for brevity. The statistics manifest that our method occupies the least memory resources to construct masks, which proves its success in reducing redundancy while maintaining good mask quality.
3 Discussions
In this subsection, we analyze some underlying properties of our method from a qualitative view. The advantages and limitations are both covered.
Orientation Maps We pick two predicted orientation maps to unearth the mechanism of our mask representation. As shown in Figure 4, the attention of each orientation map is grabbed by objects with the specific anchor size. Two kids and a skateboard are assigned to two orientation maps based on their sizes instead of categories. We display the gradient maps of both directions and their pixel-wise sum, which exhibit the obvious difference around instance boundaries. It can also be observed that two components of gradient maps concentrate on different regions of instances, i.e., the left and right parts are highlighted in while the top and bottom parts are accentuated in . Combining these two complementary orientation maps together, the complete contours of objects can be depicted and then all inner pixels for each instance are safely gathered. These visualized patterns along with high-quality predicted masks verify the effectiveness of our orientation-based mask representation.
Qualitative Results As displayed in Figure 5, our method performs well in separating adjacent instances and precisely delineating their masks. Getting rid of RoI cropping and directly collecting foreground pixels according to the vectors in orientation maps, our mask construction procedure has much tolerance to inaccurate bounding box predictions. Meanwhile, our method also performs well in some complex object overlapping scenarios, especially when one or more small objects locates on a large one. This is illustrated in several images of Figure 5, such as persons wearing ties, food placed on the table, baseball gloves at the front of people and so on. Thanks to the instance grouping mechanism, objects matched with different anchor sizes do not disturb each other and their masks are completely preserved. Moreover, we do not introduce any pixel-level categorical information for OrienHead. The class-agnostic orientation maps are proved to be qualified for recovering masks with satisfactory quality, no matter what categories they belong to.
Failure Cases Although OrienMask works well for most cases, we observe two typical failures as shown in the last column of Figure 5. The first case appears when two instances with the same category and similar size heavily overlap. Orientation maps cannot distinguish them because pixels of both masks point to almost the same base position. The second failure happens due to the severe confrontation of some background pixels between instances, especially when the base positions locate close to their mask boundaries. For example, two giraffes in the bottom right corner of Figure 5 both tend to push the intermediate part outwards, which accidentally makes some background pixels intrude into another target region by mistake. Overall, these errors caused by the incompleteness of mask representation are unusual and only occur in limited cases.
Conclusion
In this work, a real-time instance segmentation framework termed OrienMask is proposed, which integrates discriminative orientation maps with an anchor-based detector. Apart from those centripetal vectors for foreground pixels, we further consider negative samples in orientation maps so that both background filtering and instances separation can be accomplished at the same time. An instance grouping mechanism is also presented and each orientation map specializes in grouped objects with the same anchor size. Given the target regions indicated by predicted boxes, masks can be efficiently constructed from corresponding orientation maps, without the need for explicit foreground predictions. Experiments on COCO show that the proposed OrienMask can reach competitive accuracy under real-time conditions.
This work is supported by National Natural Science Foundation of China-Zhejiang Joint Fund for the Integration of Industrialization and Informatization (U1709214), and Key Research & Development Plan of Zhejiang Province (2021C01196).