CMT-DeepLab: Clustering Mask Transformers for Panoptic Segmentation
Qihang Yu, Huiyu Wang, Dahun Kim, Siyuan Qiao, Maxwell Collins, Yukun Zhu, Hartwig Adam, Alan Yuille, Liang-Chieh Chen
Introduction
Panoptic segmentation , a recently proposed challenging segmentation task, aims to unify semantic segmentation and instance segmentation . Due to its complicated nature, most panoptic segmentation frameworks decompose the problem into several manageable proxy tasks, such as box detection , box-based segmentation , and semantic segmentation .
Recently, the paradigm has shifted from the proxy-based approaches to end-to-end systems, since the pioneering work DETR , which introduces the first end-to-end object detection method with transformers . In their framework, the image features, extracted by a convolutional network , are enhanced by transformer encoders. Afterwards, a set of fixed size of positional embeddings, named object queries, interact with the extracted image features through several transformer decoders, consisting of cross-attention and self-attention modules . The object queries, transformed into output embeddings by the decoders, are then directly used for bounding box predictions.
Along the same direction, end-to-end panoptic segmentation framework has been proposed to simplify the panoptic segmentation procedure, avoiding manually designed modules. The core idea is to exploit a set of object queries conditioned on the inputs to predict a set of pairs, each containing a class prediction and a mask embedding vector. The mask embedding vector, multiplied by the image features, yields a binary mask prediction. Notably, unlike the box detection task, where the prediction is based on object queries themselves, segmentation mask prediction requires both object queries and pixel features to interact with each other to obtain the results, which consequently incurs different needs when updating the object queries. To have a deeper understanding towards the role that object queries play, we particularly look into the cross-attention module in the mask transformer decoder, where object queries interact with image features.
Our investigation finds that the update and usage of object queries are performed differently in the transformer-based method for segmentation tasks . Specifically, when updating the object queries, a softmax operation is applied to the image dimension, allowing each query to identify its most similar pixels. On the other hand, when computing the segmentation output, a softmax is performed among the object queries so that each pixel finds its most similar object queries. The formulation may potentially cause two issues: sparse query updates and infrequent pixel-query communication. First, the object queries are only sparsely updated due to the softmax being applied to a large image resolution, so it tends to focus on only a few locations (top row in Fig. 1). Second, the pixels only have one chance to communicate with the object queries in the final output. The first issue is particularly undesired, since segmentation tasks require dense predictions, and ideally a query should densely activate all the pixels that belong to the same target. This is different from the box detection task, where object extremities are sufficient (see Fig. 6 of DETR paper ).
To alleviate the issues, we draw inspiration from the traditional clustering algorithms . In the current end-to-end panoptic segmentation system , the final segmentation output is obtained by assigning each pixel to the object queries based on the feature affinity, similar to pixel-cluster assignment step in . The observation motivates us to rethink the transformer-based methods from the clustering perspective by considering the object queries as cluster centers. We therefore propose to additionally perform the cluster-update step, where the centers are updated by pooling pixel features based on the clustering assignment, when updating the cluster centers (i.e., object queries) in the cross-attention module. As a result, our model generates denser attention maps (bottom row in Fig. 1). We also utilize the pixel-cluster assignment to update the pixel features within each transformer decoder, enabling frequent communication between pixel features and cluster centers.
Additionally, we notice that in the cross-attention module, pixel features are treated as in “bag of words” , while the location information is not well utilized. To resolve the issue, we propose to adopt a dynamic position encoding conditioned on the inputs for location-sensitive clustering. We explicitly predict a reference mask consisting of a few points for each cluster center. The location-sensitive clustering is then achieved by adding location information to pixel features and cluster centers via the coordinate convolution at the beginning of each transformer decoder.
Combining all the proposed components results in our CMT-DeepLab, which reformulates and further improves the previous end-to-end panoptic segmentation system from the traditional clustering perspective. The panoptic segmentation result is naturally obtained by assigning each pixel to its most similar cluster center based on the feature affinity (Fig. 2). In the Clustering Mask Transformer (CMT) module, the pixel features, cluster centers, and pixel-cluster assignments are updated in a manner similar to the clustering algorithms . As a result, without bells and whistles, our proposed CMT-DeepLab surpasses its baseline MaX-DeepLab by 4.4% PQ and achieves 55.7% PQ on COCO panoptic test-dev set .
Related Works
Transformers. Transformer variants have advanced the state-of-the-art in many natural language processing tasks by capturing relations across modalities or in a single context (self-attention) . In computer vision, transformers are either combined with CNNs or used as standalone models . Both classes of methods have boosted various vision tasks, such as image classification , object detection , semantic segmentation , video recognition , image generation , and panoptic segmentation .
Proxy-based Panoptic Segmentation. Most panoptic segmentation methods rely on proxy tasks, such as object bounding box detection. For example, Panoptic FPN follows a box-based approach that detects object bounding boxes and predicts a mask for each box, usually with a Mask R-CNN and FPN . Then, the instance segments (‘thing’) and semantic segments (‘stuff’) are fused by merging modules to generate panoptic segmentation. Other proxy-based methods typically start with semantic segments and group ‘thing’ pixels into instance segments with various proxy tasks, such as instance center regression , Watershed transform , Hough-voting , or pixel affinity . DetectoRS achieved the state-of-the-art in this category with recursive feature pyramid and switchable atrous convolution. Recently, DETR extended the proxy-based methods with its transformer-based end-to-end detector.
End-to-end Panoptic Segmentation. Along the same direction, MaX-DeepLab proposed an end-to-end strategy, in which class-labeled object masks are directly predicted and are trained by Hungarian matching the predicted masks with ground truth masks. In this work, we improve over MaX-DeepLab by approaching the pixel assignment task from a clustering perspective. Concurrent with our work, Segmenter and MaskFormer formulated an end-to-end strategy from a mask classification perspective, same as MaX-DeepLab , but extends from panoptic segmentation to semantic segmentation.
Method
Herein, we firstly introduce recent transformer-based methods for end-to-end panoptic segmentation. Our observation reveals a difference between the cross-attention and final segmentation output regarding the way that they utilize object queries. We then propose to resolve it with a clustering approach, resulting in our proposed Clustering Mask Transformer (CMT-DeepLab), as shown in Fig. 3 and Fig. 4. In the following parts, object queries and cluster centers refer to the same learnable embedding vectors and we use them interchangeably for clearer representation.
The ground truth masks do not overlap with each other, i.e., , and denotes the ground truth class label of mask .
Inspired by DETR , several transformer-based end-to-end panoptic segmentation methods have been proposed recently, which directly predict masks and their semantic classes. is a fixed number and .
where denotes the predicted semantic class confidence for the corresponding mask, including ‘thing’ classes, ‘stuff’ classes, and the void class .
To predict these masks, object queries are utilized to aggregate information from the image features through a transformer decoder, which consists of self-attention and cross-attention modules. The object queries and image features interact with each other in the cross-attention module:
2 Current Issues and New Clustering Perspective
We propose to extend the formulation to a transformer decoder module, whose query, key, and value are obtained by linearly projecting the image features and cluster centers:
In the following subsection, we detail how the clustering perspective alleviates the issues of current transformer-based methods. In the discussion, we use object queries and cluster centers interchangeably.
3 Clustering Mask Transformers
In this subsection, we redesign the cross-attention in the transformer decoder from the clustering perspective, aiming to resolve the issues raised in Sec. 3.2.
Residual Path between Cluster Assignments. Similar to other designs , we stack the transformer decoder multiple times. To facilitate the learning of pixel-cluster assignment, we add a residual connection between clustering results including the final segmentation result. That is,
Solution to Sparse Query Update. We propose a simple and effective solution to avoid the sparse query update by combining the proposed clustering center update (i.e., Eq. (5)) with the original cross-attention (i.e., Eq. (3)), resulting in
where is obtained from Eq. (6). The update is shown in the center panel of Fig. 4, while the effect of densified attention could be found in Fig. 1.
Solution to Infrequent Pixel Updates. We propose to also utilize the clustering result to perform an update on the pixel features using the features of cluster centers, i.e.,
To this end, we have improved the transformer cross-attention module by simultaneously updating the clustering result (i.e., pixel-cluster assignment), pixel features, and cluster centers. However, we notice that during the interaction between pixel features and cluster centers, pixel features are treated as bag of words , while the location information is not well utilized. Although learnable positional encodings (i.e. object queries ) are used for the cluster center embeddings, the positional encodings are fixed for all input images, which is suboptimal when an object query predicts masks at different locations in different input images. To resolve the issue, we propose to adopt a dynamic positional encoding conditioned on the inputs for location-sensitive clustering.
Location-Sensitive Clustering. To inject dynamic location information to cluster centers, we explicitly predict a reference mask that consists of points for each cluster center. In particular, a MLP is used to predict the reference mask out of cluster center features, followed by a sigmoid activation function. That is, we have:
We add location information to pixel features and cluster centers through a coordinate convolution . Specifically, we apply coordinate convolutions at the beginning of each transformer layer to ensure location information is considered during the clustering process, as shown below.
We note that compared to the reference point used in the Deformable DETR , the proposed reference mask provides a rough mask shape prior for the whole object mask. Besides, we adopt a much simpler way to incorporate the location information via coordinate convolution.
In order to learn meaningful reference mask predictions, we optimize the reference masks towards ground truth masks by proposing a mask approximation loss.
Mask Approximation Loss. We propose a loss to minimize the distance between the distribution of predicted reference points and that of points of ground-truth object masks. In detail, we utilize the Hungarian matching result to assign the ground-truth mask for each cluster center. Given the predicted points for each cluster center, we infer their extreme points and mask center. We then apply an loss to push them to be closer to their ground-truth extreme points and center. Specifically, we have
where are pixels on ground-truth masks and predicted reference masks have been filtered and re-ordered based on Hungarian matching results.
Finally, combining all the proposed designs results in our Clustering Mask Transformer, or CMT-DeepLab, which rethinks the current mask transformer design from the clustering perspective.
4 Network Instantiation
We instantiate CMT-DeepLab on top of MaX-DeepLab-S (abbreviated as MaX-S). We first refine its architecture design. Afterwards, we enhance it with the proposed Clustering Mask Transformers.
Base Architecture. We use MaX-S as our base architecture. To better align it with other state-of-the-art architecture designs , we use GeLU activation to replace the original ReLU activation functions. Besides, we remove all transformer blocks in the pretrained backbones, which reverts the backbone from MaX-S back to Axial-ResNet-50 . On top of the backbone, we append six dual-path axial-transformer blocks (three at stage-5 w/ channels 2048, and the other three at stage-4 w/ channels 1024), yielding totally six axial self-attention and six cross-attention modules, which aligns with the number of attention operations used in other works . Additionally, we obtain a larger network backbone by scaling up the number of blocks in stage-4 of the backbone . As a result, two different model variants are used: one built upon Axial-ResNet-50 backbone with number of blocks . See the supplementary material for a detailed illustration.
Loss Functions. Following , we use the PQ-style loss and three other auxiliary losses for the model training, including the instance discrimination loss, mask-ID cross-entropy, and semantic segmentation loss. However, we note that the instance discrimination loss proposed in aims to push pixel features to be close to the feature center computed based on the ground-truth mask, instead of directly to the cluster centers. Therefore, we adopt the pixel-wise instance discrimination loss, which learns closely aligned representations for all pixels from the same class, allowing better clustering results.
Formally, we sample a set of pixels from the image, where we add bias to pixels’ sampling probability based on the size of object mask they belong to. Thus, final sampled pixels are more balanced from objects with different scales. Afterwards, we directly perform contrastive loss on top of these pixels with multiple positive targets :
where is a subset of pixels of that belongs to the same cluster (i.e., object mask) with , and is its cardinally. We use to denote a pixel feature vector, and is the temperature.
Recursive Feature Network. Motivated by DetectoRS and CBNet , we adopt a simple strategy, named Recursive Feature Network (RFN), to increase the network capacity by stacking twice the whole model (including the backbone and added transformer blocks). There are two main differences. First, since we do not employ an FPN (as in ), we simply connect the features at stride 4 (i.e., same stride as the segmentation output). Second, we do not use the complicated fusion module proposed in , but simply average the features between two stacked networks, which we empirically found to be better by around 0.2% PQ.
Experimental Results
We report main results on COCO along with state-of-the-art methods, followed by ablation studies on the architecture variants, clustering mask transformers, pretrained weights, post-processing, and scaling strategies. Finally, we analyze the working mechanism behind CMT-DeepLab with visualizations.
Implementation Details. We build CMT-DeepLab on top of MaX-DeepLab with the official code-base . The training strategy mainly follows MaX-DeepLab. If not specified, the model is trained with 64 TPU cores for 100k iterations with the first 5k for warm-up. We use batch size = 64, Adam optimizer, a poly schedule learning rate of . The ImageNet-pretrained backbone has a learning rate multiplier 0.1. Weight decay is set to 0 and drop-path rate to 0.2. The input images are resized and padded to for training and inference. We use for pixel-wise contrastive loss and for reference masks, we also tried other values but did not observe significant difference. Loss weight is 1.0 for the mask approximation loss. Other losses employ the same setting as . During inference, we adopt a mask-wise merging scheme to obtain the final results.
Our main results on the COCO panoptic segmentation val set and test-dev set are summarized in Tab. 1.
Val Set. We compare our validation set results with box-based, center-based, and end-to-end panoptic segmentation methods. It is noticeable that CMT-DeepLab, built upon a smaller backbone Axial-ResNet-50, already surpasses all other box-based and center-based methods by a large margin. More importantly, when compared with its end-to-end baseline MaX-DeepLab-S , we observe a significant improvement of 4.6% PQ. Our small model even surpasses previous state-of-the-art method MaX-DeepLab-L , which has more than parameters, by 1.9% PQ. Compared to recently proposed MaskFormer , CMT-DeepLab still shows a significant advantage of 1.2% PQ and 1.4% PQ while being more light-weight over the small and large model variant, respectively. The significant improvement illustrates the importance of introducing the concept of clustering into transformer, which leads to a denser attention preferred by the segmentation task. Our CMT-DeepLab with a deeper backbone Axial-ResNet-104 improves the single-scale performance to 54.1% PQ, outperforming multi-scale Axial-DeepLab by 10.2% PQ. Moreover, we enhance the model with the proposed RFN, which further improves the PQ to 55.3%.
Test-dev Set. We verify the transfer-ability of CMT-DeepLab on test-dev set, which shows consistently better results compared to other methods. Especially, the small version of CMT-DeepLab with Axial-R50 backbone outperforms DETR by 7.4% PQ, MaX-DeepLab-S by 4.4% PQ, and MaX-DeepLab-L by 2.1% PQ. Additionally, employing a deeper backbone Axial-R104 can boost the PQ score by 1.1% PQ. On top of it, using the proposed RFN further improves PQ to 55.7%, surpassing MaskFormer with Swin-L backbone by 2.4% PQ.
2 Ablation Studies
Herein, we evaluate the effectiveness of different components of the proposed CMT-DeepLab. For all the following experiments, we use MaX-DeepLab-S with GeLU activation function as our baseline. This improved baseline has a 0.3% higher PQ compared to the original MaX-DeepLab-S. If not specified, we perform all ablation studies with the Axial-R50 backbone , ImageNet-1K pretrained, crop size , and k training iterations.
Clustering Mask Transformer. We start with adding the design variants of Clustering Mask Transformer step by step, as summarized in Tab. LABEL:tab:cmt_appearnace. Regarding the object queries as cluster centers, and adding a clustering-style update can improve the PQ by 0.9%, illustrating the effectiveness of the cluster center perspective and the importance of including more pixels into the cluster center updates. Next, we utilize pixel-wise contrastive loss instead of the original instance-wise contrastive loss, resulting in another 0.4% PQ improvement, as it provides a better supervision signal from a clustering perspective. In short, re-designing the transformer layer from a clustering perspective leads to a 1.3% PQ improvement overall.
Location-Sensitive Clustering. Location information plays an important role in the clustering process, as shown in Tab. LABEL:tab:cmt_location. Each cluster center needs to predict a reference mask without using pixel features (i.e., appearance information), which requires cluster centers to include more location information in the feature embedding and thus benefits clustering. Adding reference masks prediction alone brings a gain of 0.4% PQ. Using the coordinate convolution (coord-conv) to include the reference mask information yields another 0.3% PQ improvement. In sum, the location-sensitive clustering brings up the PQ score by 0.7%.
Stronger Decoder. We study the effect of using a stronger decoder design . We remove all transformer layers from the pretrained backbone, which reverts the MaX-S backbone to Axial-ResNet-50 . Then we stack more axial-blocks with transformer module in the decoder part. More specifically, we use six self-attention modules and six cross-attention modules in total for the decoder, which aligns to the design of DETR . As shown in Tab. LABEL:tab:cmt_arch, this stronger decoder brings 0.9% PQ improvement (47.1% 46.2%).
As shown in Tab. LABEL:tab:cmt_arch, these improvements are complementary to each other, while combining them together can further boost the performance. Adding all of them leads to CMT-DeepLab, which improves 2.2% PQ over the MaX-DeepLab-S-GeLU baseline. We note that the major cost comes from the stronger decoder, which accounts for the increase of 29.1M parameters, while clustering update and location-sensitive clustering improve the PQ by 1.3% and 0.7%, respectively, with neglectable extra parameters.
Pretraining, Post-processing, and Scaling. We further verify the effect of better pretraining, post-processing, and scaling-up, with results summarized in Tab. LABEL:tab:cmt_arch2 and Tab. 3. Specifically, we find that using ImageNet-22K for pretraining can improve the performance by 0.9% PQ. Furthermore, we empirically find that using the mask-wise merge strategy to obtain panoptic results, compared to the simple per-pixel strategy , improves PQ by 0.5%. Next, we scale up CMT-DeepLab from different dimensions. With a longer training strategies (from 100k to 200k iterations), we observe a consistent 0.5% PQ improvement over various settings, where the improvement mainly comes from PQTh (i.e., thing classes), indicating that the model needs a longer training schedule to better segment thing objects. We also find that using a larger input resolution (from 641 to 1281) significantly boosts the performance by more than 2% PQ. Besides, increasing the model size by using a deeper backbone or stacking the model with RFN can improve the performance by 1.6% and 1.0%, respectively.
Visualization. In Fig. 5, we visualize the clustering results in each stage as well as the learned reference masks. As shown in the figure, the clustering results, starting with a close-to-random assignment, gradually learn to focus on the target instances. For example, in the last two rows of Fig. 5, the clustering results firstly focus on all the ‘person’ instances and the background ‘snow’, and then they start to concentrate on the specific person instance, showing a refinement from “semantic segmentation” to “instance segmentation”. Moreover, as shown in the last column of Fig. 5, the learned reference mask provides a reasonable prior for the object mask.
Conclusion
In this work, we have introduced CMT-DeepLab, which rethinks object queries, used in the current mask transformers for panoptic segmentation, from a clustering perspective. Considering object queries as cluster centers, our framework additionally incorporates the proposed cluster center update in the cross-attention module, which significantly enriches the learned cross-attention maps and further facilitates the segmentation prediction. As a result, CMT-DeepLab achieves new state-of-the-art performance on the COCO dataset, and sheds light on the working mechanism behind mask transformers for segmentation tasks.
We thank Jun Xie for the valuable feedback on the draft. This work was supported in part by ONR N00014-21-1-2812.
References
More Technical Details
Backbones. In Fig. 6, we provide an architectural comparison of MaX-DeepLab-S and CMT-DeepLab built upon Axial-R50/104 . Specifically, we simplify the backbone from MaX-DeepLab-S by removing transformer modules in the backbone (light blue), and stacking more blocks in the decoder module (light orange). The Axial-R104 backbone is obtained by scaling up Axial-R50 (i.e., four times more layers in the stage-4).
Recursive Feature Network. We construct Recursive Feature Network (RFN) in a manner similar to . More specifically, we stack two models together, with a skip-connection from the decoder features at stride 4 in the first network to the encoder features at stride 4 in the second network. Instead of using the complicated fusion module proposed in , we simply average the features for fusion. Moreover, the two networks share the same set of cluster centers (i.e., object queries), which are sequentially updated from the first network to the second one. We also add supervision for the first network but use the Hungarian matching results based on the final output.
More Results and In-depth Analysis
Effect of frequent pixel update (our second solution). As discussed in the main paper, the clustering results will be also used to update pixel features besides cluster centers to ensure a frequent pixel update. We tried removing the pixel feature updates from clustering transformer, which leads to a degradation of 0.4% PQ.
Comparison with more concurrent works. Also shown in Tab. 4, we compare our CMT-DeepLab with the baseline MaX-DeepLab , and concurrent works MaskFormer and K-Net on the test-dev set. As shown in the table, our best model (using 200K iterations and RFN) attains the performance of 55.7% PQ on the test-dev set, which is 4.4% and 2.4% better than MaX-DeepLab-L and MaskFormer . Our best model is 0.5% PQ better than K-Net , which adopts a different framework (i.e., dynamic kernels) than mask-transformer-based approaches. In addition to PQ, we further look into RQ and SQ for performance analysis. We observe that with a similar performance to K-Net in RQ, our best model performs better in SQ. Specifically, our best model yields 83.6% SQ, which is 1.2%, 1.6%, and 1.1% better than K-Net, MaskFormer, and MaX-DeepLab-L, respectively. Interestingly, our lightweight variant, CMT-DeepLab with Axial-R50, achieves 83.0% SQ, which is still better than all the other methods. We attribute our better performance in SQ to the proposed clustering mask transformer layer, which yields denser attention maps to facilitate segmentation tasks.
Accuracy-cost Trade-off Comparison. We provide a comprehensive comparison of training cost (epochs, memory), model size (params, FLOPs, FPS), and performance (PQ) in Tab. 5. The training memory is measured on a TPU-v4, while other statistics are measured with a Tesla V100-SXM2 GPU. We use TensorFlow 2.7, cuda 11.0, input size , and batch size . For MaskFormer (PyTorch-based), we cite the numbers from their paper. As shown in the table, our CMT-DeepLab-S (Axial-R50) outperforms MaskFormer-SwinB by 1.2% PQ with comparable model size and inference cost. Our CMT-DeepLab-S also outperforms MaskFormer-SwinL while using much fewer model parameters and running faster. All our models outperform MaX-DeepLab. Notably, our best model CMT-DeepLab-L-RFN (Axial-R104-RFN) outperforms MaX-DeepLab-L by 4.2% PQ while using only 60% model parameters and 33.6% FLOPs.
Backbone Differences. As different backbones are adopted for different methods (e.g., MaX-S/L , Swin ), it hinders a direct and fair comparison across different methods. To this end, we provide results based on a ResNet-50 backbone across different models on COCO val set. As shown in Tab. 6, our CMT-DeepLab significantly outperforms MaX-DeepLab and concurrent works (MaskFormer and K-Net).
Results on Cityscapes. We provide additional results on Cityscapes in Tab. 7. For a fair comparison, we adopt the same setting, including pretrain weights (IN-1k), training hyper-parameters (e.g., iterations 60k, learning rate 3e-4, crop size ), and post-processing scheme (pixel-wise argmax as in MaX-DeepLab). As shown in the table, our CMT-DeepLab-S significantly outperforms MaX-DeepLab-S by 2.9% PQ and 1.6% mIoU.
Visual Comparison
Visualization Details. To visually compare the clustering results/attention maps, we firstly follow DETR to average values across multi-heads to obtain a single attention map, which is then transformed into a heatmap in a manner similar to CAM by normalizing the values to the range . Note that we do not apply any smoothing techniques (e.g., square root), which in fact adjust the learned attention values. These differences make the visualization differ from those in the paper of MaX-DeepLab . All visualizations are done with CMT-DeepLab based on Axial-R50, and MaX-DeepLab-S, with input size .
Clustering results. In Fig. 7, Fig. 8, Fig. 9, and Fig. 10, we provide more clustering visualization results. We observe the same trend as we presented in the main paper that the clustering results, providing denser attention maps, are close-to-random at the beginning and are gradually refined to focus on different objects. Interestingly, we also observe some exceptions (see Fig. 7, Fig. 8, Fig. 9), where the clustering results start with a good semantic-level clustering, indicating that some cluster centers can embed semantic information and thus specialize in some classes.
Attention map comparison with MaX-DeepLab. In Fig. 11, Fig. 12, and Fig. 13, we show more attention map comparison with MaX-DeepLab. As shown in those figures, CMT-DeepLab provides a much denser attention map than MaX-DeepLab.
Limitations
Motivated from a clustering perspective, CMT-DeepLab generates denser attention maps and thus leads to a superior performance in the segmentation task. However, the proposed clustering mask transformer, though significantly improves the segmentation quality (SQ), does not bring the same-level performance boost on the recognition ability (RQ). Specifically, we have adopted some simple scaling-up strategies, including increasing model size, input size, or training iterations. Those strategies result in a large performance gain in RQ as a compensation, but with a cost at parameters, computation, or training time. It thus remains an interesting problem to explore in the future that how to improve its recognition ability efficiently and effectively.
Potential Negative Impacts
In this paper, we present a new panoptic segmentation framework, inspired by the traditional clustering-based algorithm, generates denser attention maps and further achieves new state-of-the-art performance. The findings described in this paper can potentially help advance the research in developing stronger, faster, and more elegant end-to-end segmentation methods. However, we also note that there is a long-lasting debate on the impacts of AI on human world. As a method improving the fundamental task in computer vision, our work also advances the development of AI, which means there could be both beneficial and harmful influences depending on the users.
License of used assets. COCO dataset : CC-by 4.0. ImageNet : https://image-net.org/download.php.