Attention-guided Unified Network for Panoptic Segmentation
Yanwei Li, Xinze Chen, Zheng Zhu, Lingxi Xie, Guan Huang, Dalong Du, Xingang Wang
Introduction
Scene understanding is a fundamental yet challenging task in computer vision, which has a great impact on other applications such as autonomous driving and robotics. Classic tasks for scene understanding mainly include object detection, instance segmentation and semantic segmentation. This paper considers a recently proposed task named panoptic segmentation , which aims at finding all foreground (FG) objects (named things, mainly including countable targets such as people, animals, tools, etc.) at the instance level, meanwhile parsing the background (BG) contents (named stuff, mainly including amorphous regions of similar texture and/or material such as grass, sky, road, etc.) at the semantic level. The benchmark algorithm and MS-COCO panoptic challenge winners dealt with this task by directly combining FG instance segmentation models and BG scene parsing algorithms, which ignores the underlying relationship and fails to borrow rich contextual cues between things and stuff.
In this paper, we present a conceptually simple and unified framework for panoptic segmentation. To facilitate information flow between FG things and BG stuff, we combine conventional instance segmentation and semantic segmentation networks, leading to a unified network with two branches. This strategy brings an immediate improvement in segmentation accuracy as well as higher efficiency in computation (because the network backbone can be shared). This implies that panoptic segmentation benefits from complementary information provided by FG objects and BG contents, which lays the foundation of our approach.
Going one step further, we explore the possibility of integrating higher-level visual cues (i.e., beyond the features extracted from the end of the backbone) towards the more accurate segmentation. This is achieved via two attention-based modules working at the object level and the pixel level, respectively. For the first module, we refer to the regional proposals, each of which indicates a possible FG thing, and adjusts the probability of the corresponding region to be considered as FG things and BG stuff. For the second module, we take out the FG segmentation mask, and use it to refine the boundary between FG things and BG stuff. In the context of deep networks, these two modules, named the Proposal Attention Module (PAM) and Mask Attention Module (MAM), respectively, are implemented as additional connections across FG and BG branches. Within MAM, a new layer named RoIUpsample is designed to define an accurate mapping function between pixels in the fixed-shape FG mask and the corresponding feature map. In practice, all additional connections go from the FG branch to the BG branch, mainly due to the observation that FG segmentation is often more accurateWe find the pixel accuracy of things is much higher (6.7% absolute gap) than that of stuff, when considering instance with the same semantic as one category, e.g., all individuals are evaluated as person in testing. We evaluate them on the same MS-COCO semantic evaluation metric.. Furthermore, BG stuff, while being refined by FG things, also gives feedback via gradients. Consequently, both FG and BG segmentation accuracies are considerably improved.
The overall approach, named Attention-guided Unified Network (AUNet), can be easily instantiated to various network backbones, and optimized in an end-to-end manner. We evaluate AUNet in two popular segmentation benchmarks, namely, the MS-COCO and Cityscapes datasets, and claim the state-of-the-art performance in terms of PQ, a standard metric integrating accuracies of both things and stuff . In addition, the benefits brought by joint optimization and two attention-based modules are verified through an extensive ablation study 4.2.
The major contribution of this research is to present a simple and unified framework for both FG and BG segmentation, which reaches the top performance in MS-COCO and Cityscapes datasets. Furthermore, this work also investigate the complementary information delivered by FG objects and BG contents. While panoptic segmentation serves as a natural scenario of studying this topic, its application lies in a wider range of visual tasks. Our solution, AUNet, is a preliminary exploration in this field, yet we look forward to more efforts along this direction.
The remainder of this paper is organized as follows. Section 2 briefly reviews related work. Section 3 elaborates the proposed AUNet, including two attention-based modules. After experiments are shown in Section 4, we conclude this work in Section 5.
Related Work
Traditional deep learning based scene understanding researches often focused on foreground or background targets . Recently, the rapid progress in object detection and instance segmentation made it possible to achieve object localization and segmentation at a finer level. Meanwhile, the development of semantic segmentation boosted the performance of scene parsing. Despite their effectiveness, the separation of these tasks caused the lack of contextual cues in instance segmentation as well as the confusion brought by individuals in semantic segmentation. To bridge this gap, recently, researchers proposed a new task named panoptic segmentation , which aims at accomplishing both tasks (FG instance and BG semantic segmentation) simultaneously.
Panoptic Segmentation: In , the author gave a benchmark of panopic segmentation by combining instance and semantic segmentation models. Later, a weakly-supervised method was proposed on top of initialized semantic results, and an end-to-end approach was designed to combine both FG and BG cues. However, their performance is far from the benchmark . Different from them, our proposed AUNet achieves the top performance in an end-to-end framework. Furthermore, we also establish the bond between proposal-based instance and FCN based semantic segmentation. Most recently works include .
Instance Segmentation: Instance segmentation aims at discriminating different instances of the same object. There are mainly two streams of methods to solve this task, namely, proposal-based methods and segmentation-based methods. Proposal-based methods, with the help of accurate regional proposals, often achieved higher performance. Recent examples include MNC , FCIS , Mask R-CNN and PANet . Moreover, segmentation-based methods aggregated pixel-level cues to compose instances combined with semantic segmentation or depth ordering results.
Semantic Segmentation: With the development of so-called encoding-decoding networks such as FCN , rapid progress has been made in semantic segmentation . In segmentation, capturing contextual information plays a vital role, for which various approaches were proposed including ASPP used in DeepLab for multi-scale contexts, DenseASPP for global contexts, and PSPNet which collected contextual priors. There were also efforts to use attention modules for spatial feature selection, such as , which will be detailed discussed next.
Attention-based Modules: Attention-based modules have been widely applied in visual tasks, including image processing, video understanding, and object tracking . In particular, SENet formulated channel-wise relationships via an attention-and-gating mechanism, non-local network bridged self-attention for machine translation to video classification using non-local filters. In the scope of scene understanding, and aggregated global contextual information as well as class-dependent features by channel-attention operations. More recently, self-attention and channel attention were adopted by to model long-range contexts in the spatial and channel dimensions, respectively. In this work, we establish the relationship between foreground things and background stuff in panoptic segmentation with a series of coarse-to-fine attention blocks.
Attention-guided Unified Network
Panoptic segmentation task aims at understanding everything visible in one view, which means each pixel of an image must be assigned a semantic label and an instance ID. To address this issue, the existing top algorithms directly combined the instance and semantic results from separate models, such as Mask R-CNN and PSPNet .
We formulate the problem of panoptic segmentation as recognizing and segmenting all FG things and understanding all BG stuff. In this way, we solve the problem from two aspects, namely foreground branch and background branch in a unified network (Figure 2). In detail, given an input image , our goal is to generate FG things result and BG stuff result simultaneously. Thus, the panoptic result can be generated from and directly using the fusion method in Section 3.4. The performance of panoptic results is evaluated by panoptic quality (PQ) as described in Section 4.1. For this purpose, we firstly introduce our unified framework for panoptic segmentation in this section. Then, key elements in our designed attention-guided modules are elaborated, including proposal attention module (PAM) and mask attention module (MAM). Finally, we give our implementation details.
2 Unified Framework
3 Attention-guided Modules
Considering the complementary relationship between FG things and BG stuff, we introduce features from foreground branch to background branch for more contextual cues. From another perspective, the attention operation connecting two branches also establishes a bond between proposal-based method and FCN-based method segmentation. To this end, two spatial attention modules are proposed, namely proposal attention module (PAM) and mask attention module (MAM).
Based on the above formulation of PAM, we highlight the background regions in the shared feature maps via attention operation and background reweight function. It also facilitates the learning of things in turn by enhancing the weights of activated foreground regions during backpropagation (see Section 4.2).
3.2 Mask Attention Module
With the introduction of contextual cues by PAM, background branch is encouraged to focus more on the regions of stuff. However, the predicted coarse areas from RPN branch lack enough cues for precise BG representations. Unlike RPN features, the fixed-shape masks generated from foreground branch encode finer FG layouts. Thus, we propose Mask Attention Module (MAM) to further model the relationship, as illustrated in Figure 5. Consequently, the shape FG segmentation mask is needed for similar attention operations as before. Now, the problem is: how to reproduce the shape FG feature map from masks?
RoIUpsample: In order to solve the size mismatching problem, we propose a new differentiable layer called RoIUpsample. Specifically, RoIUpsample is designed similar to the inverse process of RoIAlign , as clearly illustrated in Figure 4. In the RoIUpsample layer, the mask ( equals to 14 or 28 in Mask R-CNN) is firstly reshaped to the same size of RoIs (generated from RPN). Then we utilize the designed inverse bilinear interpolation to compute values of the output features at four regularly sampled locations (same with RoIAlign) in each mask bin, and then sum up the final results as the generated mask feature map. To meet the requirement of bilinear interpolation , in which near points are given more contributions, an operation for inverse bilinear interpolation is formulated:
where denotes the result of point after inverse bilinear interpolation, here equals to one quarter of the corresponding value in the input mask, and normalized weights , are defined as:
in which and indicate the distance between grid point and generated in two axes respectively, as presented in Figure 4(b). Note that with the Equation 5 and 6, the mask can also be reverted from the generated feature map with the forward bilinear interpolation.
Then, the generated feature map is assigned to four different scales according to the size of RoIs, which is similar with that in FPN . Consequently, the generated FG feature map is achieved for the following operations.
Attention Operation: Different from traditional instance segmentation tasks, the predicted FG masks are utilized to give background branch more contextual guidance in pixel-level. We firstly aggregate them together to the feature map using RoIUpsample, as presented in Figure 5. Then, the finer activated BG regions can be produced, similar with that in PAM. With the introduction of attention, the FG masks is also supervised by semantic loss function, which enables a further improvement in scene understanding (both for things and stuff), as discussed in Section 4.2. A similar background reweight function is adopted to aggregate useful highlighted background features. Consequently, we model the complementary relationship between FG things and BG stuff with the proposed PAM and MAM.
4 Implementation Details
In this section, we give more implementation details on the training and inference stage of our proposed AUNet.
Training: As well elaborated in Section 3.2, all of our proposed methods are trained in a unified framework. The whole network is optimized via a joint loss function during training stage:
In details, we adopt ResNet-FPN as our backbone. And the hyperparameters in the foreground branch are set following Mask R-CNN . The backbone is pre-trained on ImageNet , and the remaining parameters are initialized following . As standard practice , 8 GPUs are used to train all the models. Each mini-batch has 2 images per GPU for ResNet-50 and ResNet-101 based networks and 1 image per GPU for the others. The networks are optimized for several epochs (18 for MS-COCO and 100 for Cityscapes) using mini-batch stochastic gradient descent (SGD) with a weight decay of 4e-5 and a momentum of 0.9. Batch Normalization in the backbone is fixed and Group Normalization is added to all of the branches in our final results. For MS-COCO , the learning rate is initialized with 0.02 for the first 13 epochs and divided by 10 at 15-th and 18-th epoch respectively. Input images are horizontally flipped and reshaped to the scale with a 600 pixels short edge during training. Multi-scale testing is adopted for final results 4.3. For Cityscapes , the learning rate is initialized with 0.01 and divided by 10 at 68-th and 88-th epoch respectively. We construct each minibatch for training from 16 random 5121024 image crops (2 crops per GPU) after randomly flipping and scaling each image by 0.5 to 2.0. Multi-scale testing is dropped in 4.3.
Inference: The panoptic results are produced in inference stage by fusing the results of FG things and BG stuff in a similar way with that in . In this stage, the overlaps of things are first resolved in a NMS-like procedure which predicts the segments with higher confidence scores. And the relationships among categories are also considered during this procedure. For example, ties should not be overlapped by person in the final result. Then, the non-overlapping instance segments are combined with stuff results by assigning instance label first in favor of the things.
Experiments
In this section, our approach is evaluated on Microsoft COCO and Cityscapes datasets. We first give description of the datasets as well as the evaluation metrics. Then we evaluate our method and give detailed analyses. Comparison with the state-of-the-art methods in panoptic segmentation are presented at last.
Dataset: Due to the novelty of panoptic task itself, there are few datasets with detailed panoptic annotations as well as public evaluation metrics. Microsoft COCO is the most suitable and challenging one for the new panoptic segmentation task, for the detailed annotations and high data complexity. It consists of 115k images for training and 5k images for validation, as well as 20k images for test-dev and 20k images for test-challenge. MS-COCO panoptic annotations includes 80 thing categories and 53 stuff categories. We train our models on train set with no extra data and reports results on val set and test-dev set for comparison. Cityscapes dataset is adopted to further illustrate the effectiveness of the proposed method. In detail, it contains 2975 images for training, 500 images for validation and 1525 images for testing with fine annotations. It has another 20k coarse annotations for training, which are not used in our experiment. We report our results on val set with 19 semantic label and 8 annotated instance categories.
Evaluation Metrics: We adopt the evaluation metrics introduced by , which computes panoptic quality (PQ) metric for evaluation. PQ can be explained as the multiplication of a segmentation quality (SQ) and a recognition quality (RQ) term:
where means the intersection-over-union between predicted object and ground truth , true positives () denotes matched pairs of segments (), false positives () represents unmatched predicted segments, and false negatives () means unmatched ground truth segments. PQ, SQ, and RQ of both thing and stuff are also reported in our results.
2 Component-wise Analysis and Diagnosis
In this section, we will decompose our approach step-by-step to reveal the effect of each component. All experiments in this section are trained and evaluated on MS-COCO dataset in a single model with no extra data. Here, we adopt ResNet-50-FPN as our backbone. For fair comparison, we strictly follow the merging method in with no trick or multi-scale data augmentation in training and inference stage when doing component-wise analyses. As presented in Table 1, our proposed AUNet achieve an absolute improvement of in PQ when compared with separate training method.
As elaborated in Section 3.2, our proposed unified framework deals with FG things and BG stuff in parallel branches. As shown in Table 1, the unified framework boosts up the performance both in PQSt and PQTh, which brings absolute improvements in PQ. This can be attributed to the shared backbone and joint optimization, with which the network is supervised to focus on more discriminative features for both things and stuff. With the shared backbone, the misclassification in stuff are effectively reduced and the things are given more details.
2.2 Proposal Attention Module
The proposed PAM builds the complementary relationship between things and stuff from different scales. By this way, the binary-classified RPN branch is optimized under the supervision of semantic labels. With the bond between stuff and things established, the network performs consistent gain in PQSt and PQTh, as presented in Table 1. The background reweight function proves its effectiveness in PQSt. This can be resulted from the global contextual features introduced by global average pooling in Equation 3, which means it chooses to aggregate highlighted BG features under the guidance of global context. As shown in Figure 6, the activated feature map emphasize the background areas with context cues. It is worth noting that we have tried other fusion methods for FG and BG feature fusion, such as concatenation and direct summary after feature transformation. But these strategies have minor contributions, which means the attention is more appropriate for relationship establishment.
2.3 Mask Attention Module
While the PAM establishes the bond between FG objects and BG contents, the MAM gives background finer representations, as elaborated in Section 3.3.2 and Figure 6. As that in PAM, MAM also achieves better performance over the raw method in both PQSt and PQTh. However, the contribution of MAM is slightly lower than PAM. We guess this is caused by the lack of contextual cues in the generated FG segmentation mask.We adopt zero padding for vacant areas in RoIUpsample layer, resulting in blank BG context. This needs to be investigated in the future works. In fact, we also evaluate the performance when adopting different resolution masks for RoIUpsample, namely the mask and the one. The result shows the high resolution mask features bring a further gain ( absolute improvement in PQ) over the smaller one. This is reasonable, because RoIUpsample layer generates finer layouts if given higher resolution masks. With the help of background reweight function, MAMr achieves in PQ.
3 Comparison to State-of-the-arts
We compare our proposed network with other state-of-the-art methods on MS-COCO test-dev and Cityscapes val set.
MS-COCO: As shown in Table 2, the proposed AUNet achieves the leading PQ performance in MS-COCO dataset without bells-and-whistles. In details, winners of COCO2018 panoptic challenge adopt numerous additional network enhancements during training and inference stage, e.g., abundant extra data (110k external annotated MS-COCO images), multi-scale training, model ensemble. Moreover, considering the network enhancements adopted by the winner teams, cascade R-CNN is adopted for things and extra blocks or label bank are added for stuff as well. Different from them, the proposed AUNet achieves the top performance in a unified framework with no extra data or additional network enhancements for both things and stuff. To be more specific, only one single model based on the ResNeXt-152-FPNWe use the d variant of ResNeXt with deformable conv and non-local blocks . is adopted in the AUNet.
Filtering out the improvement bring by model ensemble, we compare the AUNet with “PKU_360” team who adopted a similar backbone but with additional skills. The result shows that our algorithm perform better than them especially in PQSt, for about absolute improvements. Furthermore, the AUNet overpasses the former end-to-end method, namely JSIS-Net , with a absolute gap, which proves the effectiveness of the proposed method. In Table 2, it is clear that the AUNet have a great balance between things and stuff, even when comparing with the challenge winners (no extra data). This is due to the introduction of unified framework and attention-guided modules for complementary relationship establishment, as well elaborated in Section 4.2. Figure 7 gives intuitive presentations of the top performance using our proposed AUNet.
Cityscapes: We compare our proposed method with the leading bottom-up methods and Mask R-CNN in Table 3. Firstly, we adopt the same training strategy with that in MS-COCO, which means all things are considered as one category in background branch, denoted as Oursequ. However, the strategy is inferior to that when using all 19 semantic labels, as illustrated in Table 3. Additionally, the MAM, which is proved to decrease the PQ in Cityscapes, is disabled in the final results. We guess the decline is caused by the inconsistency with prior information 1, which means the relatively worse things prediction may give wrong cues to stuff. Overall, the proposed method surpass previous state-of-the-art , with a 5.2% absolute gap.
Conclusions
This paper presents AUNet, a unified framework for panoptic segmentation. The key difference from prior approaches lies in that we unify FG (instance-level) and BG (semantic-level) segmentation into one model, so that the FG branch, often being better optimized, can assist the BG branch via two sources of attention (i.e., proposal attention module and mask attention module), which offer object-level and pixel-level guidance, respectively. In experiments, we observe consistent accuracy gain in MS-COCO, based on which new state-of-the-arts are achieved.
Our research delivers an important message: in visual tasks, it is often beneficial to partition targets into a few subclasses according to their properties, so that complementary information can be propagated across subclasses to assist scene understanding. Panoptic segmentation, being a new task, offers a natural partition between FG things and BG stuff, yet more possibilities remain unexplored and to be studied in the future.
Acknowledgement
We would like to thank Jiagang Zhu and Yiming Hu for valuable discussions. This work was supported by the National Key Research and Development Program of China No. 2018YFD0400902 and National Natural Science Foundation of China under Grant 61573349.