Deep Snake for Real-Time Instance Segmentation
Sida Peng, Wen Jiang, Huaijin Pi, Xiuli Li, Hujun Bao, Xiaowei Zhou
Introduction
Instance segmentation is the cornerstone of many computer vision tasks, such as video analysis, autonomous driving, and robotic grasping, which require both accuracy and efficiency. Most of the state-of-the-art instance segmentation methods perform pixel-wise segmentation within a bounding box given by an object detector , which may be sensitive to the inaccurate bounding box. Moreover, representing an object shape as dense binary pixels generally results in costly post-processing.
An alternative shape representation is the object contour, which is a set of vertices along the object silhouette. In contrast to pixel-based representation, a contour is not limited within a bounding box and has fewer parameters. Such a contour-based representation has long been used in image segmentation since the seminal work by Kass et al. , which is well known as snakes or active contours. Given an initial contour, the snake algorithm iteratively deforms it to match the object boundary by optimizing an energy functional defined with low-level features, such as image intensity or gradient. While many variants have been developed in literature, these methods are prone to local optima as the objective functions are handcrafted and typically nonconvex.
Some recent learning-based segmentation methods also represent objects as contours and try to directly regress the coordinates of contour vertices from an RGB image. Although such methods are fast, most of them do not perform as well as pixel-based methods. Instead, Ling et al. adopt the deformation pipeline of traditional snake algorithms and train a neural network to evolve an initial contour to match the object boundary. Given a contour with image features, it regards the input contour as a graph and uses a graph convolutional network (GCN) to predict vertex-wise offsets between contour points and the target boundary points. It achieves competitive accuracy compared with pixel-based methods while being much faster. However, the method proposed in is designed to help annotation and lacks a complete pipeline for automatic instance segmentation. Moreover, treating the contour as a general graph with a generic GCN does not fully exploit the special topology of a contour.
In this paper, we propose a learning-based snake algorithm, named deep snake, for real-time instance segmentation. Inspired by previous methods , deep snake takes an initial contour as input and deforms it by regressing vertex-wise offsets. Our innovation is introducing the circular convolution for efficient feature learning on a contour, as illustrated in Figure 1. We observe that the contour is a cycle graph that consists of a sequence of vertices connected in a closed cycle. Since every vertex has the same degree equal to two, we can apply the standard 1D convolution on the vertex features. Considering that the contour is periodic, deep snake introduces the circular convolution, which indicates that an aperiodic function (1D kernel) is convolved in the standard way with a periodic function (features defined on the contour). The kernel of circular convolution encodes not only the feature of each vertex but also the relationship among neighboring vertices. In contrast, the generic GCN performs pooling to aggregate information from neighboring vertices. The kernel function in our circular convolution amounts to a learnable aggregation function, which is more expressive and results in better performance than using a generic GCN, as demonstrated by our experimental results in Section 5.2.
Based on deep snake, we develop a pipeline for instance segmentation. Given an initial contour, deep snake can iteratively deform it to match the object boundary and obtain the object shape. The remaining question is how to initialize a contour, whose importance has been demonstrated in classic snake algorithms. Inspired by , we propose to generate an octagon formed by object extreme points as the initial contour, which generally encloses the object tightly. Specifically, we integrate deep snake with an object detector. The detected bounding box initializes a diamond contour defined by four center points on the edges. Then, deep snake takes the diamond as input and outputs offsets that point from diamond vertices to object extreme points, which are used to construct an octagon following . Finally, deep snake deforms the octagon contour to match the object boundary.
Our approach exhibits competitive performances on Cityscapes , KINS , SBD and COCO datasets, while being efficient for real-time instance segmentation, 32.3 fps for images on a GTX 1080ti GPU. The following two facts make learning-based snake fast and accurate. First, our approach can deal with errors in the object localization stage and thus allows a light detector. Second, the contour representation has fewer parameters than the pixel-based representation and does not require costly post-processing, e.g., mask upsampling.
In summary, this work has the following contributions:
We propose a learning-based snake algorithm for real-time instance segmentation and introduce the circular convolution for feature learning on the contour.
We propose a two-stage pipeline for instance segmentation: initial contour proposal and contour deformation. Both stages can deal with errors in the initial object localization.
We demonstrate state-of-the-art performances of our approach on Cityscapes, KINS, SBD and COCO datasets. For images, our algorithm runs at 32.3 fps, which is efficient for real-time applications.
Related work
Most methods perform instance segmentation on the pixel level within a region proposal, which works particularly well with standard CNNs. A representative instantiation is Mask R-CNN . It first detects objects and then uses a mask predictor to segment instances within the proposed boxes. To better exploit the spatial information inside the box, PANet fuses mask predictions from fully-connected layers and convolutional layers. Such proposal-based approaches achieve state-of-the-art performance. One limitation of these methods is that they cannot resolve errors in localization, such as too small or shifted boxes. In contrast, our approach deforms the detected boxes to the object boundaries, so the spatial extension of object shapes will not be limited.
There exist some pixel-based methods that are free of region proposals. In these methods, every pixel produces the auxiliary information, and then a clustering algorithm groups pixels into object instances based on their information. The auxiliary information and grouping algorithms could be various. predicts the boundary-aware energy for each pixel and uses the watershed transform algorithm for grouping. differentiates instances by learning instance-level embeddings. consider the input image as a graph and regress pixel affinities, which are then processed by a graph merge algorithm. Since the mask is composed of dense pixels, the post-clustering algorithms tend to be time-consuming.
Contour-based methods.
In these methods, the object shape comprises a sequence of vertices along the object boundary. Traditional snake algorithms first introduced the contour-based representation for image segmentation. They deform an initial contour to the object boundary by optimizing a handcrafted energy with respect to the contour coordinates. To improve the robustness of these methods, proposed to learn the energy function in a data-driven manner. Instead of iteratively optimizing the contour, some recent learning-based methods try to regress the coordinates of contour points from an RGB image, which is much faster. However, their reported accuracy is not on par with state-of-the-art pixel-based methods.
In the field of semi-automatic annotation, have tried to perform the contour labeling using other networks instead of standard CNNs. predict the contour points sequentially using a recurrent neural network. To avoid sequential inference, follows the pipeline of snake algorithms and uses a graph convolutional network to predict vertex-wise offsets for contour deformation. This strategy significantly improves the annotation speed while being as accurate as pixel-based methods. However, lacks a pipeline for instance segmentation and does not fully exploit the special topology of a contour. Instead of treating the contour as a general graph, deep snake leverages the cycle graph topology and introduces the circular convolution for efficient feature learning on a contour.
Proposed approach
Inspired by , we perform object segmentation by deforming an initial contour to match object boundary. Specifically, deep snake takes a contour as input and predicts per-vertex offsets pointing to the object boundary. Features on contour vertices are extracted from the input image with a CNN backbone. To fully exploit the contour topology, we propose the circular convolution for efficient feature learning on the contour, which facilitates deep snake to learn the deformation. Based on deep snake, we also develop a pipeline for instance segmentation.
Given an initial contour, traditional snake algorithms treat the coordinates of the vertices as a set of variables and optimize an energy functional with respect to these variables. By designing proper forces at the contour coordinates, the algorithms could drive the contour to the object boundary. However, since the energy functional is typically nonconvex and handcrafted based on low-level image features, the optimization tends to find local optimal solutions.
In contrast, deep snake directly learns to evolve the contour in an end-to-end manner. Given a contour with vertices , we first construct feature vectors for each vertex. The input feature for a vertex is a concatenation of learning-based features and the vertex coordinate: , where denotes the feature maps. The feature maps are obtained by applying a CNN backbone on the input image. The CNN backbone is shared with the detector in our instance segmentation pipeline, which will be discussed later. The image feature is computed using the bilinear interpolation at the vertex coordinate . The appended vertex coordinate is used to encode the spatial relationship among contour vertices. Since the deformation should not be affected by the translation of the contour in the image, we subtract each dimension of by the minimum value over all vertices.
and propose to encode the periodic features by the circular convolution defined as:
Similar to the standard convolution, we can construct a network layer based on the circular convolution for feature learning, which is easy to be integrated into a modern network architecture. After the feature learning, deep snake applies three 11 convolution layers to the output features for each vertex and predicts vertex-wise offsets between contour points and the target points, which are used to deform the contour. In all experiments, the kernel size of circular convolution is fixed to be nine.
As discussed in the introduction, the proposed circular convolution better exploits the circular structure of the contour than the generic graph convolution. We will show the experimental comparison in Section 5.2. An alternative method is to use standard CNNs to regress a pixel-wise vector field from the input image to guide the evolution of the initial contour . We argue that an important advantage of deep snake over the standard CNNs is the object-level structured prediction, i.e., the offset prediction at a vertex depends on other vertices of the same contour. Therefore, deep snake will predict a more reasonable offset for a vertex located far from the object. Standard CNNs may have difficulty in this case, as the regressed vector field may drive this vertex to another object which is closer.
Figure 3(a) shows the detailed schematic. Following ideas from , deep snake consists of three parts: a backbone, a fusion block, and a prediction head. The backbone is comprised of 8 “CirConv-Bn-ReLU” layers and uses residual skip connections for all layers, where “CirConv” means circular convolution. The fusion block aims to fuse the information across all contour points at multiple scales. It concatenates features from all layers in the backbone and forwards them through a 11 convolution layer followed by max pooling. The fused feature is then concatenated with the feature of each vertex. The prediction head applies three 11 convolution layers to the vertex features and output vertex-wise offsets.
2 Deep snake for instance segmentation
Figure 3(b) overviews the proposed pipeline for instance segmentation. We combine deep snake with an object detector. The detector first produces object bounding boxes that are used to construct diamond contours. Then deep snake shifts the diamond vertices to object extreme points, which are used to construct octagon contours. Finally, deep snake takes octagons as initial contours and performs iterative contour deformation to obtain the object shape.
Most active contour models require precise initial contours. Since the octagon proposed in tightly encloses the object, we choose it as the initial contour, as shown in Figure 3(b). This octagon is formed by four extreme points, which are top, leftmost, bottom, rightmost pixels in an object, respectively, denoted by . Given a detected object box, we extract four center points at the top, left, bottom, right box edges, denoted by , and then connect them to get a diamond contour. Deep snake takes this contour as input and outputs four offsets that point from each vertex to the extreme point , namely . In practice, to consider more context information, the diamond contour is uniformly upsampled to 40 points, and deep snake correspondingly outputs 40 offsets. The loss function only supervises the offsets at .
We construct the octagon by generating four line segments based on extreme points and connecting their endpoints. Specifically, the four extreme points define a new bounding box. From each extreme point, a line is extended along the corresponding box edge in both directions by 1/4 of the edge length and truncated if it meets the box corner. Then, the endpoints of the four line segments are connected to form the octagon.
Contour deformation.
We first uniformly sample points along the octagon contour starting from the top extreme points . Similarly, the ground-truth contour is generated by uniformly sampling vertices along the object boundary and defining the first vertex as the one nearest to . Deep snake takes the initial contour as input and outputs offsets that point from each vertex to the target boundary point. We set as in all experiments, which can uniformly cover most object shapes.
However, regressing the offsets in one pass is challenging, especially for vertices far away from the object. Inspired by , we deal with this problem in an iterative optimization fashion. Specifically, our approach first predicts offsets based on the current contour and then deforms this contour by vertex-wise adding the offsets to its vertex coordinates. The deformed contour can be used for the next iteration. In experiments, the number of inference iteration is set as 3 unless otherwise stated.
Note that the contour is an alternative representation for the spatial extension of an object. By deforming the initial contour to match the object boundary, deep snake could address the localization errors from the detector.
Multi-component detection.
Some objects are split into several components due to occlusions, as shown in Figure 4. However, a contour can only outline one component. To overcome this problem, we propose to use another detector to find the object components within the object box. Figure 4 shows the basic idea. Specifically, using the detected box, our approach performs RoIAlign to extract a feature map and adds a detector branch on the feature map to produce the component boxes. For the detected components, we use deep snake to segment each of them and then merge the segmentation results.
Implementation details
Detector.
We adopt CenterNet as the detector for all experiments. CenterNet reformulates the detection task as a keypoint detection problem and achieves an impressive trade-off between speed and accuracy. For the object box detector, we adopt the same setting as , which outputs class-specific boxes. For the component box detector, a class-agnostic CenterNet is adopted. Specifically, given an feature map, the class-agnostic CenterNet outputs an tensor representing the component center and an tensor representing the box size.
Experiments
contains training, 500 validation and testing images with high quality annotations. Besides, it has 20k images with coarse annotations. The performance is evaluated in terms of the average precision (AP) metric averaged over eight semantic classes of the dataset.
KINS [35]
was created by additionally annotating Kitti dataset with instance-level semantic annotation. This dataset is used for amodal instance segmentation, which aims to recover complete instance shapes even under occlusion. KINS consists of training images and testing images. Following its setting, we evaluate our approach on seven object categories in terms of the AP metric.
SBD [16]
re-annotates images from the PASCAL VOC dataset with instance-level boundaries. The reason that we don’t evaluate on PASCAL VOC is that its annotations contain holes, which is not suitable for contour-based methods. SBD is split into training images and testing images. We report our results in terms of 2010 VOC APvol , AP50, AP70 metrics. APvol is the average of AP with nine thresholds from 0.1 to 0.9.
COCO [24]
is one of the most challenging datasets for instance segmentation. It consists of 115k training , 5k validation and 20k testing images. We report our results in terms of the AP metric.
2 Ablation studies
We conduct ablation studies on the SBD dataset as it has 20 semantic categories which could fully evaluate the ability to handle various object shapes. The three proposed components are evaluated, including our network architecture, initial contour proposal, and circular convolution. In these experiments, the detector and deep snake are trained end-to-end for 160 epochs with multi-scale data augmentation. The learning rate starts from and decays by half at 80 and 120 epochs.
Table 1 summarizes the results of ablation studies. The row “Baseline” lists the result of a direct combination of Curve-gcn with CenterNet . Specifically, the detector produces object boxes, which gives ellipses around objects. Then ellipses are deformed towards object boundaries through Graph-ResNet. Note that, this baseline method represents the contour as a graph and uses a graph convolution network for contour deformation.
To validate the advantages of our network, the model in the second row keeps the convolution operator as graph convolution and replaces Graph-ResNet with our proposed architecture, which yields 1.4 APvol improvement. The main difference between the two networks is that our architecture appends a global fusion block before the prediction head.
When exploring the influence of the contour initialization, we add the initial contour proposal before the contour deformation. Instead of directly using the ellipse, the proposal step generates an octagon initialization by predicting four object extreme points, which not only compensates for the detection errors but also encloses the object more tightly. The comparison between the second and the third row shows a 1.3 improvement in terms of APvol.
Finally, the graph convolution is replaced with the circular convolution, which achieves 0.8 APvol improvement. To fully validate the importance of circular convolution, we further compare models with different convolution operators and different inference iterations, as shown in table 2. Circular convolution outperforms graph convolution across all inference iterations. Circular convolution with two iterations outperforms graph convolution with three iterations by 0.6 APvol. Figure 5 shows qualitative results of graph and circular convolution on SBD, where circular convolution gives a sharper boundary. Both the quantitative and qualitative results indicate that models with the circular convolution have a stronger ability to deform contours.
3 Comparison with the state-of-the-art methods
Since fragmented instances are very common in Cityscapes, the proposed multi-component detection strategy is adopted. Our network is trained with multi-scale data augmentation and tested at a single resolution of . No testing tricks are used. The detector is first trained alone for 140 epochs, and the learning rate starts from and drops by half at 80, 120 epochs. Then the detection and snake branches are trained end-to-end for 200 epochs, and the learning rate starts from and drops by half at 80, 120, 150 epochs. We choose a model that performs best on the validation set.
Table 3 compares our results with other state-of-the-art methods on the Cityscapes validation and test sets. All methods are tested without tricks. Using only the fine annotations, our approach achieves state-of-the-art performances on both validation and test sets. We outperform PANet by 0.9 AP on the validation set and 1.3 AP50 on the test set. Our approach achieves 28.2 AP on the test set when the strategy of handling fragmented instances is not adopted. Visual results are shown in Figure 6.
Performance on KINS.
The KINS dataset is for amodal instance segmentation, where objects are all annotated as single-component, so the multi-component detection strategy is not adopted. We train the detector and snake end-to-end for 150 epochs. The learning rate starts from and decays with 0.5 and 0.1 at 80 and 120 epochs, respectively. We perform multi-scale training and test the model at a single resolution of .
Table 4 shows the comparison with on the KINS dataset in terms of the AP metric. Our approach achieves the best performance across all methods. We find that the snake branch can improve the detection performance. When CenterNet is trained alone, it obtains 30.5 AP on detection. When trained with the snake branch, its performance improves by 2.3 AP. For an image resolution of on the KINS dataset, our approach runs at 7.6 fps on a 1080 Ti GPU. Figure 6 shows some qualitative results on KINS.
Performance on SBD.
Since annotations of objects on SBD are mostly single-component, the multi-component detection strategy is not adopted. For fragmented instances, our approach detects their components separately instead of detecting the whole object. We train the detection and snake branches end-to-end for 150 epochs with multi-scale data augmentation. The learning rate starts from and drops by half at 80 and 120 epochs. The network is tested at a single scale of .
In Table 5, we compare with other contour-based methods on the SBD dataset in terms of the VOC AP metrics. predict the object contours by regressing shape vectors. STS defines the object contour as a radial vector, and ESE approximates object contour with the Chebyshev polynomial. We outperform these methods by a large margin of at least 19.1 APvol. Note that, our approach yields 21.4 AP50 and 36.2 AP70 improvements, demonstrating that the improvement increases as the IoU threshold gets smaller. This indicates that our method outlines object boundaries more precisely. For images on the SBD dataset, our approach runs at 32.3 fps on a 1080 Ti. Some qualitative results are illustrated in Figure 7.
Performance on COCO.
Similar to the experiment on SBD, the multi-component detection strategy is not adopted. The network is trained with multi-scale data augmentation and tested at the original image resolution without tricks (e.g., flip augmentation). The detection and snake branches are trained end-to-end for 160 epochs, where the detector is initialized with the pretrained model released by . The learning rate starts from and drops by half at 80 and 120 epochs. We choose a model that performs best on the validation set. Table 6 compares our method with other real-time methods. Our method achieves 30.3 segm AP and 33.2 bbox AP on COCO test-dev set with 27.2 fps.
4 Running time
Table 7 compares our approach with other methods in terms of running time on the PASCAL VOC dataset. Since the SBD dataset shares images with PASCAL VOC, the running time on the SBD dataset is technically the same as the one on PASCAL VOC. We obtain the running time of other methods from .
For images on the SBD dataset, our algorithm runs at 32.3 fps on a desktop with an Intel i7 3.7GHz and a GTX 1080 Ti GPU, which is efficient for real-time instance segmentation. Specifically, CenterNet takes 18.4 ms, the initial contour proposal takes 3.1 ms, and each iteration of contour deformation takes 3.3 ms. Since our approach outputs the object boundary, no post-processing like upsampling is required. If the multi-component detection strategy is adopted, the detector additionally takes 3.6 ms.
Conclusion
We proposed a learning-based snake algorithm for real-time instance segmentation, which introduces the circular convolution for efficient feature learning on the contour and regresses vertex-wise offsets for the contour deformation. Based on deep snake, we developed a two-stage pipeline for instance segmentation: initial contour proposal and contour deformation. We showed that this pipeline gained a superior performance than direct regression of the coordinates of the object boundary points. To overcome the limitation of the contour representation that it can only outline one connected component, we proposed the multi-component detection strategy and demonstrated the effectiveness of this strategy on Cityscapes. The proposed model achieved competitive results on the Cityscapes, Kins, Sbd and COCO datasets with a real-time performance.
Acknowledgements: The authors would like to acknowledge support from NSFC (No. 61806176) and Fundamental Research Funds for the Central Universities (2019QNA5022).