Panoptic SegFormer: Delving Deeper into Panoptic Segmentation with Transformers

Zhiqi Li, Wenhai Wang, Enze Xie, Zhiding Yu, Anima Anandkumar, Jose M. Alvarez, Ping Luo, Tong Lu

Introduction

Semantic segmentation and instance segmentation are two important and related vision tasks. Their underlying connections recently motivated panoptic segmentation as a unification of both the tasks . In panoptic segmentation, image contents are divided into two types: things and stuff. Things refer to countable instances (e.g., person, car) and each instance has a unique id to distinguish it from the other instances. Stuff refers to the amorphous and uncountable regions (e.g., sky, grassland) and has no instance id .

Recent works attempt to employ transformers to handle both things and stuff through a query set. For example, DETR simplifies the workflow of panoptic segmentation by adding a panoptic head on top of an end-to-end object detector. Unlike previous methods , DETR does not require additional handcrafted pipelines . While being simple, DETR also causes some issues: (1) It requires a lengthy training process to converge; (2) Because the computational complexity of self-attention is squared with the length of the input sequence, the feature resolution of DETR is limited. So that it uses an FPN-style panoptic head to generate masks, which always suffer low-fidelity boundaries; (3) It handles things and stuff equally, yet representing them with bounding boxes, which may be suboptimal for stuff . Although DETR achieves excellent performance on the object detection task, its superiority on panoptic segmentation has not been well demonstrated. In order to overcome the defects of DETR on panoptic segmentation, we propose a series of novel and effective strategies that improve the performance of transformer-based panoptic segmentation models by a large margin.

Our approach. In this work, we propose Panoptic SegFormer, a concise and effective framework for panoptic segmentation with transformers. Our framework design is motivated by the following observations: 1) Deep supervision matters in learning high-qualities discriminative attention representations in the mask decoder. 2) Treating things and stuff with the same recipe is suboptimal due to the different properties between things and stuff . 3) Commonly used post-processing such as pixel-wise argmax tends to generate false-positive results due to extreme anomalies. We overcome these challenges in Panoptic SegFormer framework as follows:

We propose a mask decoder that utilizes multi-scale attention maps to generate high-fidelity masks. The mask decoder is deeply-supervised, promoting discriminative attention representations in the intermediate layers with better mask qualities and faster convergence.

We propose a query decoupling strategy that decomposes the query set into a thing query set to match things via bipartite matching and another stuff query set to process stuff with class-fixed assign. This strategy avoids mutual interference between things and stuff within each query and significantly improves the qualities of stuff segmentation. Kindly refer to Sec. 3.3.1 and Fig. 3 for more details.

We propose an improved post-processing method to generate results in panoptic format. Besides being more efficient than the widely used pixel-wise argmax method, our method contains a mask-wise merging strategy that considers both classification probability and predicted mask qualities. Our post-processing method alone renders a 1.3% PQ improvement to DETR .

We conduct extensive experiments on COCO dataset. As shown in Fig. 1, Panoptic SegFormer significantly surpasses priors arts such as MaskFormer and K-Net with much fewer parameters. With deformable attention and our deeply-supervised mask decoder, our method requires much fewer training epochs than previous transformer-based methods (24 vs. 300++) . In addition, our approach also achieves competitive performance with current methods on the instance segmentation task.

Related Work

Panoptic Segmentation. Panoptic segmentation becomes a popular task for holistic scene understanding . The panoptic segmentation literature mainly treats this problem as a joint task of instance segmentation and semantic segmentation where things and stuff are handled separately . Kirillov et al. proposed the concept of and benchmark of panoptic segmentation together with a baseline that directly combines the outputs of individual instance segmentation and semantic segmentation models. Since then, models such as Panoptic FPN , UPSNet and AUNet have improved the accuracy and reduced the computational overhead by combining instance segmentation and semantic segmentation into a single model. However, these methods approximate the target task by solving the surrogate sub-tasks, therefore introducing undesired model complexities and suboptimal performance.

Recently, efforts have been made to unify the framework of panoptic segmentation. Li et al. proposed Panoptic FCN where the panoptic segmentation pipeline is simplified with a “top-down meets bottom-up” two-branch design similar to CondInst . In their work, things and stuff are jointly modeled by an object/region-level kernel branch and an image-level feature branch. Several recent works represent things and stuff as queries and perform end-to-end panoptic segmentation via transformers. DETR predicts the bounding boxes of things and stuff and combines the attention maps of the transformer decoder and the feature maps of ResNet to perform panoptic segmentation. Max-Deeplab directly predicts object categories and masks through a dual-path transformer regardless of the category being things or stuff. On top of DETR, MaskFomer used an additional pixel decoder to refine high spatial resolution features and generated the masks by multiplying queries and features from the pixel decoder. Due to the computational complexity of self attention , both DETR and MaskFormer use feature maps with limited spatial resolutions for panoptic segmentation, which hurts the performance and requires combining additional high-resolution feature maps in final mask prediction. Unlike the methods mentioned above, our query decoupling strategy deals with things and stuff with separate query sets. Although thing and stuff queries are designed for different targets, they are processed by the mask decoder with the same workflow. Prediction results of these queries are in the same format so that we can process them in an equal manner during the post-processing procedure. One concurrent work employs a similar line of thinking to use dynamic kernels to perform instance and semantic segmentation, and it aims to utilize unified kernels to handle various segmentation tasks. In contrast to it, we aim to delve deeper into the transformer-based panoptic segmentation. Due to the different nature of various tasks, whether a unified pipeline is suitable for these tasks is still an open problem. In this work, we utilize an additional location decoder to assist things to learn location clues and get better results.

End-to-end Object Detection. The recent popular end-to-end object detection frameworks have inspired many other related works . DETR is arguably the most representative end-to-end object detector among these methods. DETR models the object detection task as a dictionary lookup problem with learnable queries and employs an encoder-decoder transformer to predict bounding boxes without extra post-processing. DETR greatly simplifies the conventional detection framework and removes many hand-crafted components such as Non-Maximum Suppression (NMS) and anchors . Zhu et al. proposed Deformable DETR, which further reduces the memory and computational cost through deformable attention layers. In this work, we adopt deformable attention for the improved efficiency and convergence over DETR .

Methods

As illustrated in Fig. 2, Panoptic SegFormer consists of three key modules: transformer encoder, location decoder, and mask decoder, where (1) the transformer encoder is applied to refine the multi-scale feature maps given by the backbone, (2) the location decoder is designed to capturing location clues of things, and (3) the mask decoder is for final classification and segmentation.

During inference, we adopt a mask-wise merging strategy to convert the predicted masks from final mask decoder layer into the panoptic segmentation results, which will be introduced in detail in Sec. 3.5.

2 Transformer Encoder

High-resolution and the multi-scale features maps are important for the segmentation tasks . Since the high computational cost of self-attention layer, previous transformer-based methods can only process low-resolution feature maps (e.g., ResNet C5C_{5}) in their encoders, which limits the segmentation performance. Different from these methods, we employ the deformable attention to implement our transformer encoder. Due to the low computational complexity of the deformable attention, our encoder can refine and involve positional encoding to high-resolution and multi-scale feature maps FF.

3 Decoder

In this section, we introduce our query decoupling strategy firstly, and then we will explain the details of our location decoder and mask decoder.

We argue that using one query set to handle both things and stuff equally is suboptimal. Since there many different properties between them, things and stuff is likely to interfere with each other and hurt the model performance, especially for PQst. To prevent things and stuff from interfering with each other, we apply a query decoupling strategy in Panoptic SegFormer, as shown in Fig. 3. Specifically, NthN_{\rm th} thing queries are used to predict things results, and NstN_{\rm st} stuff queries target stuff only. Using this form, things and stuff queries can share the same pipeline since they are in the same format. We can also customize private workflow for things or stuff according to the characteristics of different tasks. In this work, we use an additional location decoder to detect individual instances with thing queries, and this will assist in distinguishing between different instances . Mask decoder accepts both thing queries and stuff queries and generates the final masks and categories. Note that, for thing queries, ground truths are assigned by bipartite matching strategy. For stuff, We use a class-fixed assign strategy, and each stuff query corresponds to one stuff category.

Thing and stuff queries will output results in the same format, and we handle these results with a uniform post-processing method.

3.2 Location Decoder

Location information plays an important role in distinguishing things with different instance ids in the panoptic segmentation task . Inspired by this, we employ a location decoder to introduce the location information of things into the learnable queries. Specifically, given NthN_{\rm th} randomly initialized thing queries and the refined feature tokens generated by transformer encoder, the decoder will output NthN_{\rm th} location-aware queries.

In the training phase, we apply an auxiliary MLP head on top of location-aware queries to predict the bounding boxes and categories of the target object, We supervise the prediction results with a detection loss Ldet\mathcal{L}_{\rm det}. The MLP head is an auxiliary branch, which can be discarded during the inference phase. The location decoder follows Deformable DETR . Notably, the location decoder can learn location information by predicting the mass centers of masks instead of bounding boxes. This box-free model can still achieve comparable results to our box-based model.

3.3 Mask Decoder

Similar to methods , we directly perform classification through a FC layer on top of the refined query QrefineQ_{\rm refine} from each decoder layer. Each thing query needs to predict probabilities over all thing categories. Stuff query only predicts the probability of its corresponding stuff category.

At the same time, to predict the masks, we first split and reshape the attention maps AA into attention maps A3A_{3}, A4A_{4}, and A5A_{5}, which have the same spatial resolution as C3C_{3}, C4C_{4}, and C5C_{5}. This process can be formulated as:

where Split(⋅){\rm Split}(\cdot) denotes the split and reshaping operation. After that, as illustrated in Eq. 2, we upsample these attention maps to the resolution of H/8 ⁣× ⁣W/8H/8\!\times\!W/8 and concatenate them along the channel dimension,

where Up×2(⋅){\rm Up}_{\times 2}(\cdot) and Up×4(⋅){\rm Up}_{\times 4}(\cdot) mean the 2 times and 4 times bilinear interpolation operations, respectively. Concat(⋅){\rm Concat}(\cdot) is the concatenation operation. Finally, based on the fused attention maps AfusedA_{\rm fused}, we predict the binary mask through a 1 ⁣× ⁣11\!\times\!1 convolution.

Previous literature argues that the reason for slow convergence of DETR is that attention modules equally pay attention to all the pixels in the feature maps, and learning to focus on sparse meaningful locations requires plenty of effort. We use two key designs to solve this problem in our mask decoder: (1) Using an ultra-light FC head to generate masks from the attention maps, ensuring attention modules can be guided by ground truth mask to learn where to focus on. This FC head only contains 200200 parameters, which ensures the semantic information of attention maps is highly related to the mask. Intuitively, the ground truth mask is exactly the meaningful region on which we expect the attention module to focus. (2) We employ deep supervision in the mask decoder. Attention maps of each layer will be supervised by the mask, the attention module can capture meaningful information in the earlier stage. This can highly accelerate the learning process of attention modules.

4 Loss Function

During training, our overall loss function of Panoptic SegFormer can be written as:

where Lthings{L}_{\rm things} and Lstuff{L}_{\rm stuff} are loss for things and stuff, separately. λthings\lambda_{\rm things} and λstuff\lambda_{\rm stuff} are hyperparameters.

Things Loss. Following common practices , we search the best bipartite matching between the prediction set and the ground truth set. Specifically, we utilize Hungarian algorithm to search for the permutation with the minimum matching cost, which is the sum of the classification loss Lcls\mathcal{L}_{\rm cls}, detection loss Ldet\mathcal{L}_{\rm det} and the segmentation loss Lseg\mathcal{L}_{\rm seg}. The overall loss function for the thing categories is accordingly defined as follows:

where λcls\lambda_{\rm cls}, λseg\lambda_{\rm seg}, and λloc\lambda_{\rm loc} are the weights to balance three losses. DmD_{m} is the number of layers in the mask decoder. Lclsi\mathcal{L}_{\rm cls}^{i} is the classification loss that is implemented by Focal loss , and Lsegi\mathcal{L}_{\rm seg}^{i} is the segmentation loss implemented by Dice loss . Ldet\mathcal{L}_{\rm det} is the loss of Deformable DETR that used to perform detection.

Stuff Loss. We use a fixed matching strategy for stuff. Thus there is a one-to-one mapping between stuff queries and stuff categories. The loss for the stuff categories is similarly defined as:

where Lclsi\mathcal{L}_{\rm cls}^{i} and Lsegi\mathcal{L}_{\rm seg}^{i} are the same as those in Eq. 4.

5 Mask-Wise Merging Inference

Panoptic Segmentation requires each pixel to be assigned a category label (or void) and instance id (ignored for stuff) . One challenge of panoptic segmentation is that it requires generating non-overlap results. Recent methods directly use pixel-wise argmax to determine the attribution of each pixel, and this can solve the overlap problem naturally. Although pixel-wise argmax strategy is simple and effective, we observe that it consistently produces false-positive results due to the abnormal pixel values.

Unlike pixel-wise argmax resolves conflicts on each pixel, we propose the mask-wise merging strategy by resolving the conflicts among predicted masks. Specifically, we use the confidence scores of masks to determine the attribution of the overlap region. Inspired by previous NMS methods , the confidence scores take into both classification probability and predicted mask qualities. The confidence score of the i-th result can be formulated as:

where pip_{i} is the most likely class probability of i-th result. mi[h,w]m_{i}[h,w] is the mask logit at pixel [h,w][h,w], α,β\alpha,\beta are used to balance the weight of classification probability and segmentation qualities.

As illustrated in Algorithm 1, mask-wise merging strategy takes cc, ss, and mm as input, denoting the predicted categories, confidence scores, and segmentation masks, respectively. It outputs a semantic mask SemMsk\tt SemMsk and an instance id mask IdMsk\tt IdMsk, to assign a category label and an instance id to each pixel. Specifically, SemMsk\tt SemMsk and IdMsk\tt IdMsk are first initialized by zeros. Then, we sort prediction results in descending order of confidence score and fill the sorted predicted masks into SemMsk\tt SemMsk and IdMsk\tt IdMsk in order. Then we discard the results with confidence scores below tcls{\rm t}_{\rm cls} and remove the overlaps with lower confidence scores. Only remained non-overlap part with a sufficient fraction tkeep{\rm t}_{\rm keep} to origin mask will be kept. Finally, the category label and unique id of each mask are added to generate non-overlap panoptic format results.

Experiments

We evaluate Panoptic SegFormer on COCO and ADE20K dataset , comparing it with several state-of-the-art methods. We provide the main results of panoptic segmentation and instance segmentation. We also conduct detailed ablation studies to verify the effects of each module. Please refer to Appendix for implementation details.

We perform experiments on COCO 2017 datasets without external data. The COCO dataset contains 118K training images and 5k validation images, and it contains 80 things and 53 stuff. We further demonstrate the generality of our model on the ADE20K dataset , which contains 100 things and 50 stuff.

2 Main Results

Panoptic segmentation. We conduct experiments on COCO val set and test-dev set. In Tab. 1 and Tab. 2, we report our main results, comparing with other state-of-the-art methods. Panoptic SegFormer attains 49.6% PQ on COCO val with ResNet-50 as the backbone and single-scale input, and it surpasses previous methods K-Net and DETR over 2.5% PQ and 6.2% PQ, respectively. Except for the remarkable performance, the training of Panoptic SegFormer is efficient. Under 1 ⁣×1\!\times training strategy (12 epochs), Panoptic SegFormer (R50) achieves 48.0% PQ that outperforms MaskFormer that training 300 epochs by 1.5% PQ. Enhanced by vision transformer backbone Swin-L , Panoptic SegFormer attains a new record of 56.2% PQ on COCO test-dev without bells and whistles, surpassing MaskFormer over 2.9% PQ. Our method even surpasses the previous competition-level method Innovation over 2.7 % PQ. We also obtain comparable performance by employing PVTv2-B5 , while the model parameters and FLOPs are reduced significantly compared to Swin-L. Panoptic SegFormer also outperforms MaskFormer by 1.7% PQ on ADE20K dataset , see Tab. 3.

Instance segmentation. Panoptic SegFormer can be converted to an instance segmentation model by just discarding stuff queries. In Tab. 4, we report our instance segmentation results on COCO test-dev set. We achieve results comparable to the current state-of-the-art methods such as QueryInst and HTC , and 1.8 AP higher than K-Net . Using random crops during training boosts the AP by 1.3 percentage points.

3 Ablation Studies

First, we show the effect of each module in Tab. 5. Compared to baseline DETR, our model achieves better performance, faster inference speed and significantly reduces the training epochs. We use Panoptic SegFormer (R50) to perform ablation experiments by default.

Effect of Location Decoder. Location decoder assists queries to capture the location information of things. Tab. 6 shows the results with varying the number of layers in the location decoder. With fewer location decoder layers, our model performs worse on things, and it demonstrates that learning location clues through the location decoder is beneficial to the model to handle things better. * notes we predict mass centers rather than bounding boxes in our location decoder, and this box-free model achieves comparable results (49.2% PQ vs. 49.6% PQ).

Mask-wise Merging. As shows in Tab. 7, we compare our mask-wise merging strategy against pixel-wise argmax strategy on various models. We use both Mask PQ and Boundary PQ to make our conclusions more credible. Models with mask-wise merging strategy always performs better. DETR with mask-wise merging outperforms origin DETR by 1.3% PQ . In addition, our mask-wise merging is 20% less time-consuming than DETR’s pixel-wise argmax since DETR uses more tricks in its code, such as merging stuff with the same category and iteratively removing masks with small areas. Fig. 4 shows one typical fail case of using pixel-wise argmax.

Mask Decoder. Our proposed mask decoder converges faster since the ground truth masks guide the attention module to focus on meaningful regions. Fig. 5 shows the convergence curves of several models. We only supervise the last layer of the mask decoder while not employing deep supervision. We can observe that our method achieves 49.6% PQ with training for 24 epochs, and longer training has little effect. However, D-DETR-MS needs at least 50 epochs to achieve better performance. Deep supervision is vital for our mask decoder to perform better and converge faster. Fig. 6 shows the attention maps of different layers in the mask decoder, and the attention module focuses on the target car in the previous layer when using deep supervision. The attention maps are very similar to the final predicted masks, since masks are generated by attention maps with a lightweight FC head.

Since our mask decoder can generate masks from each layer, we evaluate the performance of each layer in the mask decoder, see Tab. 10. During inference, using the first two layers of mask decoder will be on par with the whole mask decoder. It also inferences faster because the computational cost decreases. PQth is hardly affected by the number of layers, PQst performs a little poorly in the first layer. The reason is that the location decoder has made additional refinements to the thing queries.

Effect of Query Decoupling Strategy. We compare our proposed query decoupling strategy with previous DETR’s matching method (described here as “joint matching”) , as shown in Tab. 8. Following DETR, joint matching uses a set of queries to target both things and stuff and feeds all queries to both location decoder and mask decoder. For our proposed query decoupling strategy, we use thing queries to detect things through bipartite matching and use location decoder to refine them. Stuff queries are assigned through class-fixed assign strategy. For a fair comparison, both the joint matching strategy and our query decoupling strategy employ 353 queries. We can observe that our proposed strategy highly boost PQst. In addition, panoptic segmentation model can perform instance segmentation by utilizing its thing results only. However, previous panoptic segmentation methods always perform poorly on instance segmentation task even though the two tasks are closely related. Tab. 8 shows both panoptic segmentation and instance segmentation performance of various methods. Our query decoupling strategy can achieve sota performance on panoptic segmentation task while obtaining a competitive instance segmentation performance.

In short, query decoupling strategy achieves higher PQst and APseg compared to joint matching. We analyze the experimental results of joint matching and find that if one query prefers things more, the precision of stuff results detected by it will be lower, see Fig. 7. Each point represents the Thing-Preference and Stuff-Precision corresponding to each query, and the specific definitions are in Appendix. The red line is the linear regression of these points. When using one query set to detect things and stuff together, it will cause interference within each query. Our query decoupling strategy prevents things and stuff from interfering within the same query.

4 Robustness to Natural Corruptions

Panoptic segmentation has promising applications in many fields, such as autonomous driving. Model robustness is one of the top concerns of autonomous driving. In this experiment, we evaluate the robustness of our model to disturbed data. We follow and generate COCO-C, which extends the COCO validation set to include disturbed data generated by 16 algorithms from blur, noise, digital and weather categories. We compare our model to Panoptic FCN , D-DETR-MS and MaskFormer . The results are shown in Tab. 11. We calculated the mean results of disturbed data on COCO-C. Using the same backbone, our model always performs better than others. Previous literature found that transformer-based model has stronger robustness on image classification and semantic segmentation tasks. Our experimental results also show that the transformer-based backbone (Swin-L and PVTv2-B5) can bring better robustness to the model. However, for tasks requiring a more complex pipeline, such as panoptic segmentation, we argue that the design of the task head also plays an important role for the robustness of the model. For example, Panoptic SegFormer (Swin-L) has an average result of 47.2% PQ on COCO-C, outperforming MaskFormer (Swin-L) by 5.5% PQ, higher than their gap (2.9% PQ) on clean data. We posit it is due to our transformer-based mask decoder having stronger robustness than the convolution-based pixel decoder of MaskFormer.

Conclusion

Limitation. This work relies on deformable attention to process multi-scale features, and the speed is a little slow. Our model is still hard to handle features with a larger spatial shape and does not perform well for small targets.

Discussion. Recently, the segmentation field attempted to use a uniform pipeline to process various tasks, including semantic segmentation, instance segmentation, and panoptic segmentation. However, we think that complete unification is conceptually exciting but not necessarily a suitable strategy. Given the similarities and differences among the various segmentation tasks, “seek common ground while reserving differences” is a more reasonable guiding ideology. With query decoupling strategy, we can handle things and stuff in the same paradigm since they are represented as queries. In addition, we can also design customized pipelines for things or stuff. Such a flexible strategy is more suitable for various segmentation tasks. At present, task-specific designs still bring better performance. We encourage the community to further explore the unified segmentation frameworks and expect that Panoptic SegFormer can inspire future works.

Acknowledge

This work is supported by the Natural Science Foundation of China under Grant 61672273 and Grant 61832008. Ping Luo is supported by the General Research Fund of HK No.27208720 and 17212120. Wenhai Wang and Tong Lu are corresponding authors.

Appendix A Implementation Details

Our settings mainly follow DETR and Deformable DETR for simplicity. The hyper-parameters in deformable attention are the same as Deformable DETR . We use Channel Mapper to map dimensions of the backbone’s outputs to 256. The location decoder contains 6 deformable attention layers, and the mask decoder contains 6 vanilla cross-attention layers . The spatial positional encoding is the commonly used fixed absolute encoding that is the same as DETR. The window size of Swin-L we used is 7. Since we equally treat each query. λthings\lambda_{\rm things} and λstuff\lambda_{\rm stuff} are dynamically adjusted according to the relative proportion of things and stuff in each image, and their sum is 1. λcls\lambda_{\rm cls}, λseg\lambda_{\rm seg}, and λdet\lambda_{\rm det} in Eq. 4 are set to 2, 1, 1, respectively.

During the training phase, the predicted masks that be assigned ∅\varnothing will have a weight of zero in computing Lseg\mathcal{L}_{seg}. While using the mass center of instance to replace the bounding box, we only use L1 loss to supervise the mass center of predicted mask and mass center of ground truth. We employ a threshold 0.5 to obtain binary masks from soft masks. Threshold tcnf{\rm t}_{\rm cnf} and tkeep{\rm t}_{\rm keep} are 0.25 (0.3) and 0.6, respectively. α\alpha and β\beta in Eq. 6 are 1 and 2, respectively. All experiments are trained on one NVIDIA DGX node with 8 Tesla V100 GPUs.

By default, for COCO dataset , We train our models with 24 epochs, a batch size of 1 per GPU, a learning rate of 1.4 ⁣× ⁣10−41.4\!\times\!10^{-4} (decayed at the 18th epoch by a factor of 0.1, learning rate multiplier of the backbone is 0.1). We use a multi-scale training strategy with the maximum image-side not exceeding 1333 and the minimum image size varying from 480 to 800, and random crop augmentations is applied during training. The number of thing queries NthN_{\rm th} is set to 300. Stuff queries have tge equal number of stuff classes, and it is 53 in COCO.

For the ADE20K dataset , we train our model with 100 epochs (decayed at 80th epoch), image size varying from 512 to 2048. Since ADE20K contains 50 stuff, we use 50 stuff queries. Other settings are the same to COCO dataset.

FPS and FLOPs. FPS in Tab.5 is measured on a V100 GPU with a batch size of 1. ”DETR” and ”DETR+mask wise merging” are from Detectron2 and DETR’s implementation. Others are from Mmdet and our own implementation. Our framework is slightly more efficient than DETR. FLOPs of DETR are measured from Detectron2 on an average of 100 images.

A.2 Deformable DETR for Panoptic Segmentation

Following DETR for panoptic segmentation, we transplanted the panoptic head of DETR to Deformable DETR. To ensure consistency, we only generate the attention maps with the spatial shape of 32s. When using single scale deformable DETR, the process of generating attention maps is the same as DETR. When using multi-scale deformable DETR, we only multiply queries and the features (from C5) to generate attention maps. Other settings of deformable DETR for object detection are kept unchanged. We apply iterative bounding box refinement as the default setting for Deformable DETR. We use 300 queries and this brings huge computation costs, although this model achieves pretty good performance.

Appendix B Discussion

We will deliver more ablation studies, more detailed analysis in this section.

Effect of Deformation Attention. To ablate the effect of deformable attention, we extend Deformable DETR on panoptic segmentation with the panoptic head of DETR. For more implementation details, please refer to the Appendix A. As shown in Tab. B.1, multi-scale deformable attention improves 2.9% PQ compared to DETR. Multi-scale attention outperforms single-scale attention by 5.7% PQ, highlighting the important role of multi-scale features for segmentation task.

Defects of Pixel-wise Argmax. Pixel-wise argmax only considers the mask logits of each pixel. It has multiple issues that may lead to incorrect results. First of all, the pixel value generated from argmax may be extremely small, as shown in Fig. B.1, which will generate plenty of false-positive results. The second issue is that the pixel with max mask logit may be the suboptimal result, as shown in Fig. 4 of the paper. This kind of error frequently appears in the segmentation maps generated by pixle-wise argmax. MaskFormer alleviates this problem by multiplying the classification probability by the masks logits. But this kind of error will still exist.

Heuristic Procedure. The heuristic procedure was the first proposed post-processing method of panoptic segmentation. It uses different strategies to handle things and stuff separately. Pixel-wise argmax was still used in its stuff workflow. One apparent defect of this method is that it solves the overlap problem of stuff and things by always preferring things. This is an unfair way of dealing with stuff. Tab. B.2 shows that PQst of using heuristic procedure is lower than other methods because all stuff are treated unfairly.

Our mask-wise merging needs two thresholds to filter out undesirable results. Tab. B.4 shows that our algorithm is not very dependent on the choice of threshold tcnf and tkeep. Tab. B.4

Although our proposed mask-wise merging strategy has achieved better results than other post-processing methods, it also has several shortcomings. First of all, we binaries the mask through a fixed threshold. This may cause one pixel to be easily assigned a void label because the values of all candidate instances at this pixel are below the threshold. Secondly, our strategy highly depends on the accuracy of confidence scores. If the confidence scores are not accurate, it will produce a low-quality panoptic format mask.

B.2 Location Decoder

Although we use the location decoder to detect the bounding boxes of things, our workflow is still very different from the previous box-based panoptic segmentation. For example, Panoptic FPN performs instance segmentation with Mask R-CNN style. The two-stage method usually needs to extract regions from the feature based on the bboxes and then use these regions to perform segmentation. The quality of segmentation is heavily dependent on the quality of detection. However, our location decoder is used to assist in learning the location clues of the query and distinguishing different instances. Mask will not have the wrong boundary due to the wrong boundary prediction of the bbox since the bbox does not constrain the mask. We also show that using mass centers of masks to replace bboxes can still learn location clues.

Another valuable function of the location decoder is to help filter out low-quality thing queries during the training and inference phase. This can greatly save memory. Current transformer-based panoptic segmentation methods always consume a lot of GPU memory. For example, MaskFormer takes up more than 20G of GPU memory with a batch size of 1 and R50 backbone. Although these methods have achieved excellent results, they also require high hardware resources. However, our Panoptic SegFormer can be trained with taking up less than 12G memory by using a location decoder to filter out low-quality thing queries. In particular, we use bipartite matching for multiple rounds of matching in the detection phase. The thing query that already be matched will not participate in the next round of matching. After several rounds, we can select partial promising thing queries. Only these promising thing queries will be fed to the mask decoder. with this strategy, the mask decoder usually only needs to handle less than 100 thing and stuff queries.

B.3 Mask Decoder

Fig. B.3 shows the architecture of DETR’s panoptic head. Although it only contains 1.2M parameters, it has a huge computational burden (about 150150G FLOPs). DETR adds ResNet features to each attention map, and this process repeats 100 times since there are 100 attention maps. Fig. B.4 shows the model architecture of our mask decoder. Fig. B.5 shows the process of converting multi-scale multi-head attention maps to mask. We found that discarding the self-attention in the decoder does not affect the effectiveness of the model. The computational cost of our mask decoder is around 30G FLOPs.

Multi-head attention maps. Fig. B.6 shows some samples of multi-head attention maps. Through a multi-head attention mechanism, different heads of one query learn their own attention preference. We observe that some heads pay attention to foreground regions, some prefer boundaries, and others prefer background regions. This shows that each mask is generated by considering various comprehensive information in the image. Tab. B.5 shows that utilizing a multi-head attention mechanism will outperform single-head attention by 0.4% PQ.

B.4 Advantage of Query Decoupling Strategy

DETR uses the same recipe to predict boxes of things and stuff (To facilitate the distinction between the query decoupling strategy we proposed, we refer to the DETR’s strategy as a joint matching strategy.). However, detecting bboxes for stuff like DETR is suboptimal. We counted the ratio of the area of masks to the area of bboxes on the COCO train2017. The ratios of things and stuff are 52.5% and 9.2%, separately. This shows that bounding boxes can not represent stuff well since the stuff is amorphous and dispersed. We also observe that bbox AP of DETR drops from 42.0 to 38.8 after training on panoptic segmentation. This may be due to the interference of stuff on things, since predicting stuff bboxes needs to re-adapt the model.

Fig. B.7 shows that DETR seems to learn automatic segregation between things and stuff, and each query either prefers things or stuff. However, we argue that this automatically learned segregation is not ideal. If one query prefers things, it will perform poorly when it generates stuff results. This situation is very common, and our following experiments based on Panoptic SegFormer will give detailed data. Following DETR, we use 353 queries to predict things and stuff with the same recipe. Specifically, the input of the location decoder is 353 queries, which will detect both things and stuff. The refined queries are fed to the mask decoder to predict category labels and masks. We define a query’s preference for things as Pt, which can be calculated by:

where Nthings{\rm N}_{\rm things} and Nstuff{\rm N}_{\rm stuff} are the number of things and stuff masks that ii-th query predicted on COCO val set. Pti>0.5{\rm P}_{t}^{i}>0.5 means that ii-th query prefers things more than stuff. The predicted mask is a true positive (TP) if IoU between it and one ground truth mask is larger than 0.5 and the category of them is the same. Then we can calculate the precision of queries’ predicted masks. Tab. B.6 shows relevant statistical results. First, we can observe that the queries that own lower Pt basically have higher precision. The stuff precision of the queries that have the highest Pt ([0.9,1.0]\left[0.9,1.0\right]) only is 0.30, which is much lower than the average stuff precision (0.60) on all queries. These erroneous results are mainly due to errors in the predicted category. Queries that have no obvious preference for stuff and things( PtP_{t} in [0.4,0.6)\left[0.4,0.6\right) ) performs poorly both on stuff and things. These results demonstrate that using one query set to predict things and stuff simultaneously is flawed. This joint matching strategy is suboptimal for stuff.

In order to avoid mutual interference between stuff and things, we propose the query decoupling strategy to handle things and queries with a separate query set. Compared to stuff query, thing query will go through an additional location decoder. However, all queries will produce the outputs in the same format. Things and stuff use the same loss for training, except that things use an additional detection loss. During inference, we can use our mask-wise merging strategy to merge them uniformly. This is different from the previous methods that modeled panoptic segmentation into instance segmentation and semantic segmentation. For example, Panoptic FPN uses one branch to generate things and one branch to generate stuff. The things and stuff generated by Panoptic FPN are in different formats and need different training strategies and post-processing methods. PQst with query decoupling outperforms joint matching strategy by 2.9% PQ and experimental results verify the effectiveness of our method. The stuff precision by using query decoupling is 0.66, better than the joint matching strategy.

Appendix C Visualization

Fig. B.8 shows our visualization result against DETR and MaskFormer. We use the original codes that they officially implemented. First of all, compared with other methods, we can observe that our results are more consistent with ground truths. Due to the defects of pixel-wise argmax we discussed in Sec. B.1, DETR always generates results with artifacts. MaskFormer performs better because they improved pixel-wise argmax by considering classification probabilities. However, it may still fail in hard cases. For example, it recognizes the billboard as a car in the fourth row. Fig. B.9 shows some failure cases of our model. Firstly, our model may have lower recall when facing crowded scenarios filled with the same things, especially for the small targets. Another typical failure mode is that large stuff with a high confidence score occupies most of the space, causing other things not to be added to the canvas. Fig. B.10 shows the results on some complex scenes.

Appendix D Various Backbones

We give all the panoptic segmentation results under various backbones, as shown in Tab. D.1. Fig. D.1 shows two training curves with backbone ResNet-101 and Swin-L. With Swin-L, Panoptic SegFormer with training for 24 epochs even performs better than training for 50 epochs.

Appendix E Code and Data

We use the official implementations of DETRhttps://github.com/facebookresearch/detr, MaskFormerhttps://github.com/facebookresearch/MaskFormer, Panoptic FCNhttps://github.com/dvlab-research/PanopticFCN to perform additional experiments. The models they provide all can reproduce the same scores they reported in their literature. Deformable DETR is from Mmdethttps://github.com/open-mmlab/mmdetection.

References