K-Net: Towards Unified Image Segmentation
Wenwei Zhang, Jiangmiao Pang, Kai Chen, Chen Change Loy
Introduction
Image segmentation aims at finding groups of coherent pixels . There are different notions in groups, such as semantic categories (e.g., car, dog, cat) or instances (e.g., objects that coexist in the same image). Based on the different segmentation targets, the tasks are termed differently, i.e., semantic and instance segmentation, respectively. There are also pioneer attempts to joint the two segmentation tasks for more comprehensive scene understanding.
Grouping pixels according to semantic categories can be formulated as a dense classification problem. As shown in Fig. 1-(a), recent methods directly learn a set of convolutional kernels (namely semantic kernels in this paper) of pre-defined categories and use them to classify pixels or regions . Such a framework is elegant and straightforward. However, extending this notion to instance segmentation is non-trivial given the varying number of instances across images. Consequently, instance segmentation is tackled by more complicated frameworks with additional steps such as object detection or embedding generation . These methods rely on extra components, which must guarantee the accuracy of extra components to a reasonable extent, or demand complex post-processing such as Non-Maximum Suppression (NMS) and pixel grouping. Recent approaches generate kernels from dense feature grids and then select kernels for segmentation to simplify the frameworks. Nonetheless, since they build upon dense grids to enumerate and select kernels, these methods still rely on hand-crafted post-processing to eliminate masks or kernels of duplicated instances.
In this paper, we make the first attempt to formulate a unified and effective framework to bridge the seemingly different image segmentation tasks (semantic, instance, and panoptic) through the notion of kernels. Our method is dubbed as K-Net (‘K’ stands for kernels). It begins with a set of convolutional kernels that are randomly initialized, and learns the kernels in accordance to the segmentation targets at hand, namely, semantic kernels for semantic categories and instance kernels for instance identities (Fig. 1-(b)). A simple combination of semantic kernels and instance kernels allows panoptic segmentation naturally (Fig. 1-(c)). In the forward pass, the kernels perform convolution on the image features to obtain the corresponding segmentation predictions.
The versatility and simplicity of K-Net are made possible through two designs. First, we formulate K-Net so that it dynamically updates the kernels to make them conditional to their activations on the image. Such a content-aware mechanism is crucial to ensure that each kernel, especially an instance kernel, responds accurately to varying objects in an image. Through applying this adaptive kernel update strategy iteratively, K-Net significantly improves the discriminative ability of the kernels and boosts the final segmentation performance. It is noteworthy that this strategy universally applies to kernels for all the segmentation tasks.
Second, inspired by recent advances in object detection , we adopt the bipartite matching strategy to assign learning targets for each kernel. This training approach is advantageous to conventional training strategies as it builds a one-to-one mapping between kernels and instances in an image. It thus resolves the problem of dealing with a varying number of instances in an image. In addition, it is purely mask-driven without involving boxes. Hence, K-Net is naturally NMS-free and box-free, which is appealing to real-time applications.
To show the effectiveness of the proposed unified framework on different segmentation tasks, we conduct extensive experiments on COCO dataset for panoptic and instance segmentation, and ADE20K dataset for semantic segmentation. Without bells and whistles, K-Net surpasses all previous state-of-the-art single-model results on panoptic (54.6% PQ) and semantic segmentation benchmarks (54.3% mIoU) and achieves competitive performance compared to the more expensive Cascade Mask R-CNN . We further analyze the learned kernels and find that instance kernels incline to specialize on objects at specific locations of similar sizes.
Related Work
Semantic Segmentation. Contemporary semantic segmentation approaches typically build upon a fully convolutional network (FCN) and treat the task as a dense classification problem. Based on this framework, many studies focus on enhancing the feature representation through dilated convolution , pyramid pooling , context representations , and attention mechanisms . Recently, SETR reformulates the task as a sequence-to-sequence prediction task by using a vision transformer . Despite the different model architectures, the approaches above share the common notion of making predictions via static semantic kernels. Differently, the proposed K-Net makes the kernels dynamic and conditional on their activations in the image.
Instance Segmentation. There are two representative frameworks for instance segmentation – ‘top-down’ and ‘bottom-up’ approaches. ‘Top-down’ approaches first detect accurate bounding boxes and generate a mask for each box. Mask R-CNN simplifies this pipeline by directly adding a FCN in Faster R-CNN . Extensions of this framework add a mask scoring branch or adopt a cascade structure . ‘Bottom-up’ methods first perform semantic segmentation then group pixels into different instances. These methods usually require a grouping process, and their performance often appears inferior to ‘top-down’ approaches in popular benchmarks . Unlike all these works, K-Net performs segmentation and instance separation simultaneously by constraining each kernel to predict one mask at a time for one object. Therefore, K-Net needs neither bounding box detection nor grouping process. It focuses on refining kernels rather than refining bounding boxes, different from previous cascade methods .
Recent attempts perform instance segmentation in one stage without involving detection nor embedding generation. These methods apply dense mask prediction using dense sliding windows or dense grids . Some studies explore polar representation, contour , and explicit shape representation of instance masks. These methods all rely on NMS to eliminate duplicated instance masks, which hinders end-to-end training. The heuristic process is also unfavorable for real-time applications. Instance kernels in K-Net are trained in an end-to-end manner with bipartite matching and set prediction loss, thus, our methods does not need NMS.
Panoptic Segmentation. Panoptic segmentation combines instance and semantic segmentation to provide a richer understanding of the scene. Different strategies have been proposed to cope with the instance segmentation task. Mainstream frameworks add a semantic segmentation branch on an instance segmentation framework or adopt different pixel grouping strategies based on a semantic segmentation method. Recently, DETR tries to simplify the framework by transformer but need to predict boxes around both stuff and things classes in training for assigning learning targets. These methods either need object detection or embedding generation to separate instances, which does not reconcile the instance and semantic segmentation in a unified framework. By contrast, K-Net partitions an image into semantic regions by semantic kernels and object instances by instance kernels through a unified perspective of kernels.
Concurrent to K-Net, some recent attempts apply Transformer for panoptic segmentation. MaskFormer reformulates semantic segmentation as a mask classification task, which is commonly adopted in instance-level segmentation. From an inverse perspective, K-Net tries to simplify instance and panoptic segmentation by letting a kernel to predict the mask of only one instance or a semantic category, which is the essential design in semantic segmentation. In contrast to K-Net that directly uses learned kernels to predict masks and progressively refines the masks and kernels, MaX-DeepLab and MaskFormer rely on queries and Transformer to produce dynamic kernels for the final mask prediction.
Dynamic Kernels. Convolution kernels are usually static, i.e., agnostic to the inputs, and thus have limited representation ability. Previous works explore different kinds of dynamic kernels to improve the flexibility and performance of models. Some semantic segmentation methods apply dynamic kernels to improve the model representation with enlarged receptive fields or multi-scales contexts . Differently, K-Net uses dynamic kernels to improve the discriminative capability of the segmentation kernels more so than the input features of kernels.
Recent studies apply dynamic kernels to generate instance or panoptic segmentation predictions directly. Because these methods generate kernels from dense feature maps, enumerate kernels of each position, and filter out kernels of background regions, they either still rely on NMS or need extra kernel fusion to eliminate kernels or masks of duplicated objects. Instead of generated from dense grids, the kernels in K-Net are a set of learnable parameters updated by their corresponding contents in the image. K-Net does not need to handle duplicated kernels because its kernels learn to focus on different regions of the image in training, constrained by the bipartite matching strategy that builds a one-to-one mapping between the kernels and instances.
Methodology
We consider various segmentation tasks through a unified perspective of kernels. The proposed K-Net uses a set of kernels to assign each pixel to either a potential instance or a semantic class (Sec. 3.1). To enhance the discriminative capability of kernels, we contribute a way to update the static kernels by the contents in their partitioned pixel groups (Sec. 3.2). We adopt the bipartite matching strategy to train instance kernels in an end-to-end manner (Sec. 3.3). K-Net can be applied seamlessly to semantic, instance, and panoptic segmentation as described in Sec. 3.4.
Despite the different definitions of a ‘meaningful group’, all segmentation tasks essentially assign each pixel to one of the predefined meaningful groups . As the number of groups in an image is typically assumed finite, we can set the maximum group number of a segmentation task as . For example, there are pre-defined semantic classes for semantic segmentation or at most objects in an image for instance segmentation. For panoptic segmentation, is the total number of stuff classes and objects in an image. Therefore, we can use kernels to partition an image into groups, where each kernel is responsible to find the pixels belonging to its corresponding group. Specifically, given an input feature map of images, produced by a deep neural network, we only need kernels to perform convolution with to obtain the corresponding segmentation prediction as
where , , and are the number of channels, height, and width of the feature map, respectively. The activation function can be softmax function if we want to assign each pixel to only one of the kernels (usually used in semantic segmentation). Sigmoid function can also be used as activation function if we allow one pixel belong to multiple masks, which results on binary masks by setting a threshold like 0.5 on the activation map (usually used in instance segmentation).
This formulation has already dominated semantic segmentation for years . In semantic segmentation, each kernel is responsible to find all pixels of a similar class across images. Whereas in instance segmentation, each pixel group corresponds to an object. However, previous methods separate instances by extra steps instead of by kernels.
This paper is the first study that explores if the notion of kernels in semantic segmentation is equally applicable to instance segmentation, and more generally panoptic segmentation. To separate instances by kernels, each kernel in K-Net only segments at most one object in an image (Fig. 1-(b)). In this way, K-Net distinguishes instances and performs segmentation simultaneously, achieving instance segmentation in one pass without extra steps. For simplicity, we call these kernels as semantic and instance kernels in this paper for semantic and instance segmentation, respectively. A simple combination of instance kernels and semantic kernels can naturally preform panoptic segmentation that either assigns a pixel to an instance ID or a class of stuff (Fig. 1-(c)).
2 Group-Aware Kernels
Despite the simplicity of K-Net, separating instances directly by kernels is non-trivial. Because instance kernels need to discriminate objects that vary in scale and appearance within and across images. Without a common and explicit characteristic like semantic categories, the instance kernels need stronger discriminative ability than static kernels.
To overcome this challenge, we contribute an approach to make the kernel conditional on their corresponding pixel groups, through a kernel update head, as shown in Fig. 2. The kernel update head contains three key steps: group feature assembling, adaptive kernel update, and kernel interaction. Firstly, the group feature for each pixel group is assembled using the mask prediction . As it is the content of each individual groups that distinguishes them from each other, is used to update their corresponding kernel adaptively. After that, the kernel interacts with each other to comprehensively model the image context. Finally, the obtained group-aware kernels perform convolution over feature map to obtain more accurate mask prediction . As shown in Fig. 3, this process can be conducted iteratively because a finer partition usually reduces the noise in group features, which results in more discriminative kernels. This process is formulated as
Notably, the kernel update head with the iterative refinement is universal as it does not rely on the characteristic of kernels. Thus, it can enhance not only instance kernels but also semantic kernels. We detail the three steps as follows.
Group Feature Assembling. The kernel update head first assembles the features of each group, which will be adopted later to make the kernels group-aware. As the mask of each kernel in essentially defines whether or not a pixel belongs to the kernel’s related group, we can assemble the feature for by multiplying the feature map with the as
where is the batch size, is the number of kernels, and is the number of channels.
Adaptive Feature Update. The kernel update head then updates the kernels using the obtained to improve the representation ability of kernels. As the mask may not be accurate, which is more common the case, the feature of each group may also contain noises introduced by pixels from other groups. To reduce the adverse effect of the noise in group features, we devise an adaptive kernel update strategy. Specifically, we first conduct element-wise multiplication between and as
The gate learned here plays a role like the self-attention mechanism in Transformer , whose output is computed as a weighted summation of the values. In Transformer, the weight assigned to each value is usually computed by a compatibility function dot-product of the queries and keys. Similarly, adaptive kernel update essentially performs weighted summation of kernel features and group features . Their weight and are computed by element-wise multiplication, which can be regarded as another kind of compatibility function.
3 Training Instance Kernels
While each semantic kernel can be assigned to a constant semantic class, there lacks an explicit rule to assign varying number of targets to instance kernels. In this work, we adopt bipartite matching strategy and set prediction loss to train instance kernels in an end-to-end manner. Different from previous works that rely on boxes, the learning of instance kernels is purely mask-driven because the inference of K-Net is naturally box-free.
Loss Functions. The loss function for instance kernels is written as , where is Focal loss for classification, and and are CrossEntropy (CE) loss and Dice loss for segmentation, respectively. Given that each instance only occupies a small region in an image, CE loss is insufficient to handle the highly imbalanced learning targets of masks. Therefore, we apply Dice loss to handle this issue following previous works .
Mask-based Hungarian Assignment. We adopt Hungarian assignment strategy used in for target assignment to train K-Net in an end-to-end manner. It builds a one-to-one mapping between the predicted instance masks and the ground-truth (GT) instances based on the matching costs. The matching cost is calculated between the mask and GT pairs in a similar manner as the training loss.
4 Applications to Various Segmentation Tasks
Panoptic Segmentation. For panoptic segmentation, the kernels are composed of instance kernels and semantic kernels as shown in Fig. 3. We adopt semantic FPN for producing high resolution feature map , except that we add positional encoding used in to enhance the positional information. Specifically, given the feature maps produced by FPN , positional encoding is computed based on the feature map size of , and it is added with . Then semantic FPN is used to produce the final feature map.
As semantic segmentation mainly relies on semantic information for per-pixel classification, while instance segmentation prefers accurate localization information to separate instances, we use two separate branches to generate the features and to perform convolution with and for generating instance and semantic masks and , respectively. Notably, it is unnecessary to produce ‘thing’ and ‘stuff’ masks initially from different branches to produce a reasonable performance. Such a design is consistent with previous practices and empirically yields better performance (about 1% PQ).
We then construct , , and as the inputs of kernel update head to dynamically update the kernels and refine the panoptic mask prediction. Because ‘things’ are already separated by instance masks in , while contains the semantic masks of both ‘things’ and ‘stuff’, we select , the masks of stuff categories from , and directly concatenate it with to form the panoptic mask prediction . Due to similar reason, we only select and concatenate the kernels of stuff classes in with to form the panoptic kernels . To exploit the complementary semantic information in and localization information in , we add them together to obtain as the input feature map of the kernel update head. With , , and , the kernel update head can produce group-aware kernels and mask . Then kernels and masks are iteratively by times and finally we can obtain the mask prediction .
To produce the final panoptic segmentation results, we paste thing and stuff masks in a mixed order following MaskFormer . We also find it necessary in K-Net to firstly sort the pasting order of masks based on their classification scores for further filtering out lower-confident mask predictions. Such a method empirically performs better (about 1% PQ) than the previous strategy that pasting thing and stuff masks separately .
Instance Segmentation. In the similar framework, we simply remove the concatenation process of kernels and masks to perform instance segmentation. We did not remove the semantic segmentation branch as the semantic information is still complementary for instance segmentation. Note that in this case, the semantic segmentation branch does not use extra annotations. The ground truth of semantic segmentation is built by converting instance masks to their corresponding class labels.
Semantic Segmentation. As K-Net does not rely on specific architectures of model representation, K-Net can perform semantic segmentation by simply appending its kernel update head to any existing semantic segmentation methods that rely on semantic kernels.
Experiments
Dataset and Metrics. For panoptic and instance segmentation, we perform experiments on the challenging COCO dataset . All models are trained on the train2017 split and evaluated on the val2017 split. The panoptic segmentation results are evaluated by the PQ metric . We also report the performance of thing and stuff, noted as PQTh, PQSt, respectively, for thorough evaluation. The instance segmentation results are evaluated by mask AP . The AP for small, medium and large objects are noted as APs, APm, and APl, respectively. The AP at mask IoU thresholds 0.5 and 0.75 are also reported as AP50 and AP75, respectively. For semantic segmentation, we conduct experiments on the challenging ADE20K dataset and report mIoU to evaluate the segmentation quality. All models are trained on the train split and evaluated on the validation split.
Implementation Details. For panoptic and instance segmentation, we implement K-Net with MMDetection . In the ablation study, the model is trained with a batch size of 16 for 12 epochs. The learning rate is 0.0001, and it is decreased by 0.1 after 8 and 11 epochs, respectively. We use AdamW with a weight decay of 0.05. For data augmentation in training, we adopt horizontal flip augmentation with a single scale. The long edge and short edge of images are resized to 1333 and 800, respectively, without changing the aspect ratio. When comparing with other frameworks, we use multi-scale training with a longer schedule (36 epochs) for fair comparisons . The short edge of images is randomly sampled from $$ .
For semantic segmentation, we implement K-Net with MMSegmentation and train it with 80,000 iterations. As AdamW empirically works better than SGD, we use AdamW with a weight decay of 0.0005 by default on both the baselines and K-Net for a fair comparison. The initial learning rate is 0.0001, and it is decayed by 0.1 after 60000 and 72000 iterations, respectively. More details are provided in the appendix.
Model Hyperparameters. In the ablation study, we adopt ResNet-50 backbone with FPN . For panoptic and instance segmentation, we use for Focal loss following previous methods , and empirically find , , work best. For efficiency, the default number of instance kernels is 100. For semantic segmentation, equals to the number of classes of the dataset, which is 150 in ADE20K and 133 in COCO dataset. The number of rounds of iterative kernel update is set to three by default for all segmentation tasks.
Panoptic Segmentation. We first benchmark K-Net with other panoptic segmentation frameworks in Table 1. K-Net surpasses the previous state-of-the-art box-based method and box/NMS-free method by 1.7 and 1.5 PQ on val split, respectively. On the test-dev split, K-Net with ResNet-101-FPN backbone even obtains better results than that of UPSNet , which uses Deformable Convolution Network (DCN) . K-Net equipped with DCN surpasses the previous method by 1.1 PQ. Without bells and whistles, K-Net obtains new state-of-the-art single-model performance with Swin Transformer serving as the backbone.
We also compare K-Net with concurrent work Max-DeepLab , MaskFormer , and Panoptic SegFormer . K-Net surpasses these methods with the least training epochs (36), taking only about 44 GPU days (roughly 2 days and 18 hours with 16 GPUs). Note that only 100 instance kernels and Swin Transformer with window size 7 are used here for efficiency. K-Net could obtain a higher performance with more instance kernels (Sec. 4.2), Swin Transformer with window size 12 (used in MaskFormer ), as well as an extended training schedule with aggressive data augmentation used in previous work .
Instance Segmentation. We compare K-Net with other instance segmentation frameworks in Table 2. More details are provided in the appendix. As the only box-free and NMS-free method, K-Net achieves better performance and faster inference speed than Mask R-CNN , SOLO , SOLOv2 and CondInst , indicated by the higher AP and frames per second (FPS). We adopt 256 instance kernels (K-Net-N256 in the table) to compare with Cascade Mask R-CNN . The performance of K-Net-N256 is on par with Cascade Mask R-CNN but enjoys a 92.2% faster inference speed (19.8 v.s 10.3).
On COCO test-dev split, K-Net with ResNet-101-FPN backbone obtains performance that is 0.9 AP better than Mask R-CNN . It also surpasses previous kernel-based approach CondInst and SOLOv2 by 1.2 AP and 0.6 AP, respectively. With ResNet-101-FPN backbone, K-Net surpasses Cascade Mask R-CNN with 100 and 256 instance kernels in both accuracy and speed by 0.2 AP and 6.7 FPS, and 0.7 AP and 6 FPS, respectively.
We also compare the number of parameters of these models in Table 2. Though K-Net does not have the least number of parameters, it is more lightweight than Cascade Mask R-CNN by approximately half number of the parameters (37.3 M vs. 77.1 M).
Semantic Segmentation. We apply K-Net to existing frameworks that rely on static semantic kernels in Table 3. K-Net consistently improves different frameworks. Notably, K-Net significantly improves FCN (6.6 mIoU). This combination surpasses PSPNet and UperNet by 0.7 and 0.9 mIoU, respectively, and achieves performance comparable with DeepLab v3. Furthermore, the effectiveness of K-Net does not saturate with strong model representation, as it still brings significant improvement (1.4 mIoU) over UperNet with Swin Transformer . The results suggest the versatility and effectiveness of K-Net for semantic segmentation.
In Table 3, we further compare K-Net with other state-of-the-art methods with test-time augmentation on the validation set. With the input of 512512, K-Net already achieves state-of-the-art performance. With a larger input of 640640 following previous method during training and testing, K-Net with UperNet and Swin Transformer achieves new state-of-the-art single model performance, which is 0.8 mIoU higher than the previous one.
2 Ablation Study on Instance Segmentation
We conduct an ablation study on COCO instance segmentation dataset to evaluate the effectiveness of K-Net in discriminating instances. The conclusion is also applicable to other segmentation tasks since the design of K-Net is universal to all segmentation tasks.
Positional Information. We study the necessity of positional information in Table 4. The results show that positional information is beneficial, and positional encoding works slightly better than coordinate convolution. The combination of the two components does not bring additional improvements. The results justify the use of just positional encoding in our framework.
Number of Stages. We compare different kernel update rounds in Table 4. The results show that FPS decreases as the update rounds grow while the performance saturates beyond three stages.
Such a conclusion also holds for semantic segmentation as shown in Table 5. The performance of FCN K-Net on ADE20K dataset gradually increases as the increase of iteration number but also saturates after four iterations.
Number of Kernels. We further study the number of kernels in K-Net. The results in Table 4 reveal that 100 kernels are sufficient to achieve good performance. The observation is expected for COCO dataset because most of the images in the dataset do not contain many objects (7.7 objects per image in average ). K-Net consistently achieves better performance given more instance kernels since they improve the models’ capacity in coping with complicated images. However, a larger may lead to small performance gains and then get saturated (when , we all get 34.9% mAP). Therefore, we select in other experiments for efficiency if without further specification.
3 Visual Analysis
Overall Distribution of Kernels. We carefully analyze the properties of instance kernels learned in K-Net by analyzing the average of mask activations of the 100 instance kernels over the 5000 images in the val split. All the masks are resized to have a similar resolution of 200 200 for the analysis. As shown in Fig. 4, the learned kernels are meaningful. Different kernels specialize on different regions of the image and objects with different sizes, while each kernel attends to objects of similar sizes at close locations across images.
Masks Refined through Kernel Update. We further analyze how the mask predictions of kernels are refined through the kernel update in Fig. 4. Here we take K-Net for panoptic segmentation to thoroughly analyze both semantic and instance masks. The masks produced by static kernels are incomplete, e.g., the masks of river and building are missed. After kernel update, the contents are thoroughly covered by the segmentation masks, though the boundaries of masks are still unsatisfactory. The boundaries are refined after more kernel update. The classification confidences of instances also increase after kernel update. More results are given in the appendix.
Conclusion
This paper explores instance kernels that can learn to separate instances during segmentation. Thus, extra components that previously assist instance segmentation can be replaced by instance kernels, including bounding boxes, embedding generation, and hand-crafted post-processing like NMS, kernel fusion, and pixel grouping. Such an attempt, for the first time, allows different image segmentation tasks to be tackled through a unified framework. The framework, dubbed as K-Net, first partitions an image into different groups by learned static kernels, then iteratively refines these kernels and their partition of the image by the features assembled from their partitioned groups. K-Net obtains new state-of-the-art single-model performance on panoptic and semantic segmentation benchmarks and surpasses the well-developed Cascade Mask R-CNN with the fastest inference speed among the recent instance segmentation frameworks. We wish K-Net and the analysis to pave the way for future research on unified image segmentation frameworks.
Acknowledgements. This study is supported under the RIE2020 Industry Alignment Fund Industry Collaboration Projects (IAF-ICP) Funding Initiative, as well as cash and in-kind contribution from the industry partner(s). It is also partially supported by the NTU NAP grant. Jiangmiao Pang and Kai Chen are also supported by the Shanghai Committee of Science and Technology, China (Grant No. 20DZ1100800). The authors would like to thank the valuable suggestions and comments by Jiaqi Wang, Rui Xu, and Xingxing Zou.
We first provide more implementation details for the three segmentation tasks of K-Net (Appendix A). Then we provide more benchmark details and discussion of comparison between K-Net and other methods (Appendix B). We further analyze K-Net about its results of kernel update and failure cases (Appendix C). Last but not the least, we discuss the broader impact of K-Net (Appendix D).
Appendix A Implementation Details
Training details for semantic segmentation. We implement K-Net based on MMSegmentation for experiments on semantic segmentation. We use AdamW with a weight decay of 0.0005 and train the model by 80000 iterations by default. The initial learning rate is 0.0001, and it is decayed by 0.1 after 60000 and 72000 iterations, respectively. This is different from the default training setting in MMSegmentation that uses SGD with momentum by 160000 iterations. But our setting obtains similar performance as the default one. Therefore, we apply AdamW with 80000 iterations to all the models in Table 3a of the main text for efficiency while keeping fair comparisons. For data augmentation, we follow the default settings in MMSegmentation . The long edge and short edge of images are resized to 2048 and 512, respectively, without changing the aspect ratio (described as 512 512 in the main text for short). Then random crop, horizontal flip, and photometric distortion augmentations are adopted.
Appendix B Benchmark Results
Accuracy comparison. In Table 2 of the main text, we compare both accuracy and inference speed of K-Net with previous methods. For fair comparison, we re-implement Mask R-CNN and Cascade Mask R-CNN with the multi-scale 3 training schedule , and submit their predictions to the evaluation serverhttps://competitions.codalab.org/competitions/20796 for obtaining their accuracies on the test-dev split. For SOLO , SOLOv2 , and CondInst , we test and report the accuracies of the models released in their official implementation , which are trained by multi-scale 3 training schedule. This is because the papers of SOLO and SOLOv2 only report the results of multi-scale 6 schedule, and the APs, APm, and APl of CondInst are calculated based on the areas of bounding boxes rather than instance masks due to implementation bug. The performance of TensorMask is reported from Table 3 of the paper. The results in Table 2 show that K-Net obtains better APm and APl but lower APs than Cascade Mask R-CNN. We hypothesize this is because Cascade Mask R-CNN rescales the regions of small, medium, and large objects to a similar scale of 28 28, and predicts masks on that scale. On the contrary, K-Net predicts all the masks on a high-resolution feature map.
Inference Speed. We use frames per second (FPS) to benchmark the inference speed of the models. Specifically, we benchmark SOLO , SOLOv2 , CondInst , Mask R-CNN , Cascade Mask R-CNN and K-Net with an NVIDIA V100 GPU. We calculate the pure inference speed of the model without counting in the data loading time, because the latency of data loading depends on the storage system of the testing machine and can vary in different environments. The reported FPS is an average FPS obtained in three runs, where each run measures the FPS of a model through 400 iterations . Note that the inference speed of these models may be updated due to better implementation and specific optimizations. So we present them in Table 2 only to verify that K-Net is fast and effective.
B.2 Semantic Segmentation
In Table 3b of the main text, we compare K-Net on UperNet using Swin Transformer with the previous state-of-the-art obtained by Swin Transformer . We first directly test the model in the last row of Table 3a of the main text (52.0 mIoU) with test-time augmentation and obtain 53.3 mIoU, which is on-par with the current state-of-the-art result (53.5 mIoU). Then we follow the setting in Swin Transformer to train the model with larger scale, which resize the long edge and short edge of images to 2048 and 640, respectively, during training and testing. The model finally obtains 54.3 mIoU on the validation set, which achieves new state-of-the-art performance on ADE20K.
Appendix C Visual Analysis
Masks Refined through Kernel Update. We analyze how the mask predictions change before and after each round of kernel update as shown in Figure A1. The static kernels have difficulties in handling the boundaries between masks, and the mask prediction cannot cover the whole image. The mask boundaries are gradually refined and the empty holes in big masks are finally filled through kernel updates. Notably, the mask predictions after the second and the third rounds look very similar, which means the discriminative capabilities of kernels start to saturate after the second round kernel update. The visual analysis is consistent with the evaluation metrics of a similar model on the val split, where the static kernels before kernel update only achieve 33.0 PQ, and the dynamic kernels after the first update obtain 41.0 PQ. The dynamic kernels after the second and the third rounds obtain 46.0 PQ and 46.3 PQ, respectively.
Failure Cases. We also analyze the failure cases and find two typical failure modes of K-Net. First, for the contents that have very similar texture appearance, K-Net sometimes have difficulties to distinguish them from each other and results in inaccurate mask boundaries and misclassification of contents. Second, as shown in Figure A2, in crowded scenarios, it is also challenging for K-Net to recognize and segment all the instances given limited number of instance kernels.
Appendix D Broader Impact
Simplicity and effectiveness are two significant properties pursued by computer vision algorithms. Our work pushes the boundary of segmentation algorithms through these two aspects by providing a unified perspective that tackles semantic, instance, and panoptic segmentation tasks consistently. The work could also ease and accelerate the model production and deployment in real-world applications, such as in autonomous driving, robotics, and mobile phones. The model with higher accuracy proposed in this work could also improve the safety of its related applications. However, due to limited resources, we do not evaluate the robustness of the proposed method on corrupted images and adversarial attacks. Therefore, the safety of the applications using this work may not be guaranteed. To mitigate that, we plan to analyze and improve the robustness of models in the future research.