You Only Segment Once: Towards Real-Time Panoptic Segmentation

Jie Hu, Linyan Huang, Tianhe Ren, Shengchuan Zhang, Rongrong Ji, Liujuan Cao

Introduction

Panoptic segmentation is a task that involves assigning a semantic label and an instance identity to each pixel of an input image. The semantic labels are typically classified into two types, i.e., stuff including amorphous and uncountable concepts (such as sky and road), and things consisting of countable categories (such as persons and cars). This division of label types naturally separates panoptic segmentation into two sub-tasks: semantic segmentation for stuff and instance segmentation for things. Thus, one of the major challenges for achieving real-time panoptic segmentation is the requirement for separate and computationally intensive branches to perform semantic and instance segmentation respectively. Typically, instance segmentation employs boxes or points to distinguish between different things, while semantic segmentation predicts distribution maps over semantic categories for stuff. As shown in Fig. 1, numerous efforts have been made to unify panoptic segmentation pipelines for improved speed and accuracy. However, achieving real-time panoptic segmentation still remains an open problem. On the one hand, heavy necks, e.g., the multi-scale feature pyramid network (FPN) used in , and heads, e.g., the Transformer decoder used in , are required to ensure accuracy, making real-time processing unfeasible. On the other hand, reducing the model size leads to a decrease in model generalization. Therefore, developing a real-time panoptic segmentation framework that delivers competitive accuracy is challenging yet highly desirable.

In this paper, we present YOSO, a real-time panoptic segmentation framework. YOSO predicts panoptic kernels to convolute image feature maps, with which you only need to segment once for the masks of background stuff and foreground things. To make the process lightweight, we design a feature pyramid aggregator for extracting image feature maps, and a separable dynamic decoder for generating panoptic kernels. In the aggregator, we propose convolution-first aggregation (CFA) to re-parameterize the interpolation-first aggregation (IFA), resulting in an approximately 2.6×\times speedup in GPU latency without compromising performance. Specifically, we demonstrate that the order, i.e., interpolation-first or convolution-first, of applying bilinear interpolation and 1×\times1 convolution (w/o bias) does not affect results, but the convolution-first way provides a considerable speedup to the pipeline. In the decoder, we propose separable dynamic convolution attention (SDCA) to perform multi-head cross-attention in a weight-sharing way. SDCA achieves better accuracy (+1.0 PQ) and higher efficiency (approximately 1.2×\times faster GPU latency) than traditional multi-head cross-attention.

In general, YOSO has three notable advantages. First, CFA reduces computational burden without re-training the model or compromising performance. CFA can be adapted to any task that uses the combination of bilinear interpolation and 1×\times1 convolution operations. Second, SDCA performs multi-head cross-attention with better accuracy and efficiency. Third, YOSO runs faster and has competitive accuracy compared to state-of-the-art panoptic segmentation models, and its generalization is validated on four popular datasets: COCO (46.4 PQ, 45.6 FPS), Cityscapes (52.5 PQ, 22.6 FPS), ADE20K (38.0 PQ, 35.4 FPS), and Mapillary Vistas (34.1 PQ, 7.1 FPS).

Related Work

Real-Time Panoptic Segmentation. Panoptic segmentation aims to jointly perform semantic and instance segmentation, where each pixel in an input image is assigned both a semantic label and a unique instance identity. Many studies have been conducted for fast panoptic segmentation . For instance, UPSNet utilizes a deformable convolution based semantic segmentation head and a Mask R-CNN style instance segmentation head. FPSNet proposes a fast architecture for panoptic segmentation, avoiding instance mask prediction and merging outputs via soft attention masks. Recently, PanopticDeepLab , LPSNet , and RealTimePan generate semantic masks for all categories first, then locate masks of instances via boxes or points, enabling efficient object segmentation. Meanwhile, PanopticFCN , K-Net , and MaskFormer attempt to predict masks for both things and stuff simultaneously through dynamic convolutions. Despite significant advances in this field, achieving real-time panoptic segmentation remains an open problem. In this paper, YOSO enables real-time panoptic segmentation with competitive accuracy by utilizing the proposed feature pyramid aggregator and separable dynamic decoder.

Real-Time Instance Segmentation. Instance segmentation aims to predict masks and classes for each instance in an image. To achieve real-time instance segmentation, various approaches have been proposed in recent literature. YOLACT proposes to multiply the predicted mask coefficients with prototype masks, and SipMask utilizes spatial mask coefficients for more accurate segmentation. CenterMask employs an efficient anchor-free framework, and DeepSnake explores the use of object contours for fast segmentation of instances. OrienMask designs discriminative orientation maps that recover the masks without additional foreground segmentation, and SOLO segments objects by locations, with a decoupled branch to speed up the framework. Recently, SparseInst introduced a sparse set of instance activation maps that highlight informative regions for each object in an image, constructing a real-time instance segmentation framework. As complex operations are required to distinguish different things, solving instance segmentation efficiently is the key to real-time panoptic segmentation. In this paper, YOSO predicts unified panoptic kernels for stuff and things. Implementing bipartite matching loss for fast discrimination of different things, YOSO avoids time-consuming object localization operations such as RoIAlign and post-processes such as non-maximum suppression. The output masks naturally represent independent instances for the categories of things. Moreover, the experimental results in our supplementary material show that YOSO can also achieve competitive performance on real-time instance segmentation.

Real-Time Semantic Segmentation. Semantic segmentation aims to predict pixel-wise categories for input images. In recent years, many approaches have been developed to enable real-time semantic segmentation. For example, E-Net proposes a lightweight architecture for high-speed segmentation, and SegNet combines a small network architecture with skip connections to achieve fast segmentation. ICNet uses an image cascade algorithm to speed up the pipeline, and ESPNet introduces an efficient spatial pyramid dilated convolution. Additionally, BiSeNet separates spatial details and categorical semantics to enable both high accuracy and high efficiency in semantic segmentation. More recently, SegFormer employs Transformers with a lightweight multi-layer perceptron decoder for fast semantic segmentation. In contrast to traditional methods that predict distribution maps over classes for semantic masks, YOSO predicts kernels with their corresponding categories for segmentation. This enables an efficient way to jointly solve semantic and instance segmentation for panoptic segmentation.

Method

YOSO Framework. As shown in Fig.1, YOSO is a compact framework designed for real-time panoptic segmentation, which consists of a feature pyramid aggregator and a separable dynamic decoder. The backbone network, such as ResNet, extracts multi-level feature maps from input images. The feature pyramid aggregator compresses and aggregates the multi-level feature maps into single-level. The separable dynamic decoder then generates panoptic kernels with the single-level feature maps for both mask prediction and classification.

2 Feature Pyramid Aggregator

Interpolation/Convolution-First Aggregation. In IFA, the pyramid feature maps are first upsampled to the scale of h×wh\times w via bilinear interpolation. Then, the feature maps are concatenated and fused using a 1×\times1 convolutional layer. In CFA, the pyramid feature maps are first fed to different 1×\times1 convolutional layers. Then, the feature maps are bilinearly interpolated to the scale of h×wh\times w and summed.

Observation I: The output of IFA is exactly equal to that of CFA when using 1×\times1 convolution without bias.

This observation is attributed to the homogeneity and additivity properties of the bilinear interpolation function f(⋅)f(\cdot), where f(∑iwivx,yi)=∑iwif(vx,yi)f(\sum_{i}{w^{i}\boldsymbol{v}_{x,y}^{i}})=\sum_{i}{w^{i}f(\boldsymbol{v}_{x,y}^{i})} for the constant wiw^{i} and the value vector vx,yi\boldsymbol{v}_{x,y}^{i} from the four positions (x1,y1),(x1,y2),(x2,y1),(x2,y2)(x_{1},y_{1}),(x_{1},y_{2}),(x_{2},y_{1}),(x_{2},y_{2}) of the ii-th feature map. Specifically, the bilinear interpolation estimates the value at (x0,y0)(x_{0},y_{0}) using the four positions by:

where wiw^{i} can be interpreted as a 1×\times1 kernel that convolutes the values of the original positions in the feature maps. Eq. 1 implies that applying 1×\times1 convolution (without bias) before or after bilinear interpolation does not affect the final results. Hence, it can be inferred that when using 1×\times1 convolution without bias in the aggregators, IFA and CFA produce identical outputs.

Observation II: CFA requires significantly fewer floating point operations (FLOPs) than IFA.

The reduction ratio of FLOPs between IFA and CFA is:

where dd is the channel dimension of output feature maps. In the numerator of Eq. 2 (i.e., the FLOPs of IFA), the first term represents the number of FLOPs used in bilinear interpolation, and the second term represents the number of FLOPs used in 1×\times1 convolution. In the denominator of Eq. 2 (i.e., the FLOPs of CFA), the terms represent the number of FLOPs for 1×\times1 convolution, bilinear interpolation, and accumulation, respectively.

Given the above two observations, we adopt CFA in the proposed feature pyramid aggregator. It is noteworthy that the learned weights of the 1×\times1 convolutional layer in IFA can be readily re-parameterized to CFA by dividing the weights into four 1×\times1 convolutional layers. This can accelerate the pipeline without incurring any additional costs.

3 Separable Dynamic Decoder

In order to generate accurate kernels for segmentation, previous methods typically relied on dense predictors or heavy Transformer decoders . In contrast, we propose a lightweight kernel generator called the separable dynamic decoder, which speeds up kernel generation while maintaining high accuracy. The separable dynamic decoder, shown in Fig. 3, consists of three modules: a pre-attention module, a separable dynamic convolution module, and a post-attention module. Specifically, the separable dynamic convolution efficiently performs multi-head cross-attention and achieves better accuracy. We describe each module in detail below.

where r(⋅)r(\cdot) reshapes A\boldsymbol{A} and S\boldsymbol{S} to the sizes of (n,hw)(n,hw) and (d,hw)(d,hw), respectively, ∗* denotes 2D convolution operation.

Although increasing the model capacity enhances performance, it also results in high computational burden. Intuitively, multi-head cross-attention involves three fundamental operations: multi-head projection, cross-token interaction, and cross-dimension interaction. In Eq. 4, the cross-dimension interaction, represented by Wo\boldsymbol{W}^{o}, learns to re-weight the importance of every hidden dimension. In Eq. 5, the mutli-head projection, represented by VWiv\boldsymbol{V}\boldsymbol{W}_{i}^{v}, maps the hidden dimensions dd into tt different spaces with the size of d/td/t, and the cross-token interaction, represented by KiVi\boldsymbol{K}_{i}\boldsymbol{V}_{i}, uses the correlation matrix to interact between tokens. These operations motivated us to perform multi-head cross-attention using 1D convolution to make the process lightweight, which is defined as:

Correspondingly, the basic operations of multi-head cross-attention are also performed in 1D convolution in a weight-sharing manner. For the multi-head projection, the sliding window in 1D convolution densely splits the hidden dimensions into dd groups of size tt, and the tt successive hidden dimensions in each group are projected with shared kernels. For the cross-token interaction, the first accumulation term in Eq. 7 interacts nn tokens by nn different kernels. For the cross-dimension interaction, the second accumulation term in Eq. 7 locally incorporates the information from tt successive hidden dimensions, instead of using all hidden dimensions globally in multi-head cross-attention. Furthermore, inspired by , we employ a dynamic approach to generate the kernel K\boldsymbol{K} conditioned on Q\boldsymbol{Q}, which introduces the cross-attention mechanism into 1D convolution and defines a dynamic convolution attention as follows:

Separable Dynamic Convolution. The standard convolution can be further decomposed into a depthwise convolution and a pointwise convolution as in . Following this approach, we propose a separable form for vanilla dynamic convolution as follows:

In the numerator of Eq. 10 (i.e., the FLOPs of MHCA), the first term denotes the number of FLOPs used in multi-head projection, while the second term represents the number of FLOPs used in cross-attention. In the denominator of Eq. 10 (i.e., the FLOPs of SDCA), the terms correspond to the number of FLOPs for 1D convolution operation and linear projection, respectively.

Post-Attention. In the post-attention module, we use a multi-head self-attention layer and a feed-forward network to generate the panoptic kernels. Subsequently, we produce the masks by means of 2D convolution and predict the associated classes using additional feed-forward networks. Inspired by , we utilize the panoptic kernels to update the proposal kernels iteratively for improved accuracy.

Experiments

Datasets. We evaluated the effectiveness and efficiency of YOSO on four widely used panoptic segmentation datasets: the COCO dataset , the Cityscapes dataset , the ADE20K dataset, and the Mapillary Vistas dataset. The COCO dataset gathers images of complex everyday scenes with common objects, containing 80 things categories and 53 stuff categories in 118k images for training and 5k images for validation. The Cityscapes dataset contains images of urban street-view scenes, which has 8 things categories and 11 stuff categories in 2.9k images for training and 0.5k images for validation. The ADE20K dataset is annotated in an open-vocabulary setting with 50 things categories and 100 stuff categories, including 20k images for training and 2k images for validation. The Mapillary Vistas dataset is a large-scale urban street-view dataset with 37 things categories and 28 stuff categories in 18k and 2k images for training and validation.

2 Implementation Details

For the COCO dataset, we used a batch size of 16 and set the learning rate to 0.0001. The models were trained for 370k iterations with large-scale jitter augmentation . For the Cityscapes and Mapillary datasets, we trained the models with a batch size of 16, the learning rate set to 0.0001, and the training schedule set to 180k iterations. For the ADE20K dataset, we set the batch size to 16, the learning rate to 0.0001, and the models were trained for 30k iterations. The ResNet50 (denoted as R50) pre-trained on the ImageNet dataset is employed as our backbone for the four datasets. We used ResNet50 (denoted as R50) pre-trained on ImageNet as the backbone for all four datasets, and set the hidden dimension dd to 256 through experimentation. Since the COCO dataset contains images from various scenes, ranging from indoor to outdoor, we conducted ablation studies on this dataset. In the ablation studies, the models were trained with 270k iterations.

3 Main Results

The results of panoptic segmentation on the COCO dataset are presented in Tab. 1. We have the following findings. First, YOSO is significantly faster than predominant efficient panoptic segmentation models such as PanopticFPN and RealTimePan . Specifically, YOSO achieves a PQ of 48.4 and an FPS of 23.6 with an input image scale of (800, 1333). This PQ is 11.3 higher than that of RealTimePan, while the speed is approximately 1.5×\times faster. Furthermore, when scaling the input image to (512, 800), YOSO is around 2.3×\times faster than the previous fastest model, PanopticDeepLab, while achieving an approximately 11.0-point higher PQ. Second, YOSO achieves comparable accuracy with state-of-the-art models such as MaskFormer , Mask2Former , Max-DeepLab , and K-Net . For example, YOSO outperforms MaskFormer and K-Net by 1.9 and 1.3 PQ, respectively, and achieves the same PQ performance as Max-DeepLab. Although the PQ of YOSO is 3.5 points lower than that of Mask2Former, YOSO is 2.7×\times faster than Mask2Former.

The results of panoptic segmentation on the Cityscapes dataset are presented in Tab. 2. YOSO is the fastest model with competitive accuracy among state-of-the-art approaches. For example, YOSO achieves 59.7 PQ and 11.1 FPS with the input image size of (1024, 2048), which is 4.7 points higher than FPSNet. Moreover, when reducing the input image scale to (512, 1024), YOSO achieves an accuracy of 52.5 PQ with 22.6 FPS.

In Tab. 3 and Tab. 4, we show the panoptic segmentation results on the ADE20K and the Mapillary Vistas datasets, respectively, to evaluate the model generalization of YOSO. On the ADE20K dataset, YOSO outperforms most previous methods such as PanSegFormer and MaskFormer in terms of both speed and accuracy. On the Mapillary Vistas dataset, although YOSO has a good PQs, the performance of PQt lags behind that of state-of-the-art models. This suggests that YOSO still has the potential to be improved on the Mapillary Vistas dataset.

Additionally, we plot the PQ w.r.t. FPS results on the COCO and the Cityscapes datasets in Fig. 4, which shows that YOSO runs faster and achieves competitive accuracy among state-of-the-art models. In summary, the main results on the four datasets validate the good generalization and well-balanced speed-accuracy of YOSO.

4 Ablation Study

In order to examine the impact of different components on the speed and accuracy of YOSO, we conducted several ablation studies focusing on the feature pyramid aggregator and the separable dynamic decoder. Specifically, we evaluated the effectiveness of the aggregation modules and the attention modules, which resulted in several interesting findings. Moreover, we analyzed how variations in the number of attention blocks, kernel size, iteration stages, and proposal kernels affected the performance. The details of our investigations are discussed below.

Comparison of different aggregators. The results of using different aggregators are presented in Tab. 5. Specifically, we trained YOSO with IFA and CFA, respectively. The PQ results indicate that IFA achieves higher accuracy than CFA, with PQ values of 47.5 and 47.0, respectively. However, IFA has much larger FLOPs than CFA, with 16.6G compared to 2.1G, and a slower GPU latency, with 4871μ\mus compared to 1877μ\mus. Considering that the learned parameters of IFA can be directly re-parameterized to CFA, we can train YOSO with IFA and infer with the re-parameterized CFA for better speed and accuracy.

In Fig. 5 (left), we investigate the reduction ratio of FLOPs between CFA and IFA. Specifically, we set the channel dimensions of the input feature maps c5c_{5}, c4c_{4}, c3c_{3}, and c2c_{2} to be equal to cc, and analyze how the input channel dimension cc and the output hidden dimension dd affect the FLOPs reduction ratio in Eq. 2. Our results indicate that increasing the input channel dimension and decreasing the output hidden dimension lead to an increase in the reduction ratio. This implies that CFA will be more efficient when the dimension of the input channel is large.

Comparison of different attention modules. We present a comparison of the effectiveness of different attention modules in Tab. 6. In terms of PQ performance, we have made two interesting findings. First, we were surprised to find that the MHCA module did not perform better than DCA and SDCA. The reason for this may be the difference between the basic operations in these two types of attention modules: DCA densely splits the hidden dimension into dd groups, while MHCA sparsely splits it into tt groups. Second, PDCA showed inferior performance compared to the other modules, with a PQ of 43.7. This is likely due to the fact that PDCA is the only module that does not apply cross-dimension interaction, which implies that interactions between hidden dimensions are significant for the attention module. This observation is supported by the performance of DDCA, which only performs cross-token interaction and achieved a PQ of 46.7. These two results suggest that the cross-dimension interaction may be more significant than the cross-token interaction in panoptic kernel generation, as the panoptic kernels are expected to be independent to represent different stuff or things.

In terms of FLOPs, we observed that the attention modules based on dynamic convolution require fewer FLOPs. Specifically, DDCA exhibits the lowest computational cost, requiring only 0.3M FLOPs. In terms of GPU latency, we made two interesting observations. First, although the time complexity of MHCA, DCA, SDCA, and PDCA are all O(n2d)O(n^{2}d), the speedup on the GPU latency is remarkable. Second, we found that SDCA runs slower than DCA, contrary to what the FLOPs analysis suggests. We speculate that the additional convolutional and fully connected layers in SDCA are executed serially rather than in parallel, which may result in longer execution times in practice.

Fig. 5 (right) presents the analysis of the FLOPs reduction ratio, as given by Eq. 10. We fix the kernel size tt to 3, and investigate the effect of the token size nn and the hidden dimension dd on the reduction ratio. The results indicate that the reduction ratio enlarges with the token size decreases and the hidden dimension increases. Specifically, DCA performs better when the token size is smaller than the hidden dimension, indicating its suitability for vision tasks where the token size is smaller than the hidden dimension.

Number of attention blocks. To assess the effectiveness of the separable dynamic convolution attention in the separable dynamic decoder, we varied the number of blocks and analyzed the speed-accuracy trade-off, as presented in Tab. 7. The results show that the PQ accuracy improves with additional blocks, but at the expense of lower FPS. Consequently, we selected N=2N=2 as a compromise between speed and accuracy for YOSO.

Kernel size of dynamic convolution. Tab. 8 shows the effectiveness of different kernel sizes for the separable dynamic convolution attention module. The PQ performance confirms our observation in Tab. 6, suggesting that the cross-dimension interaction plays a significant role in the attention modules. Specifically, reducing the kernel size from 5 to 1 leads to a degradation in the PQ performance from 47.3 to 46.8. Furthermore, the performance reaches saturation when the kernel size is increased from 5 to 7.

Iteration stages. To assess the impact of the number of stages, we conducted experiments with different numbers of stages and report the results in Tab. 9. Our results show that increasing the number of stages improves the PQ performance, but it comes at the cost of decreasing the FPS performance. We observed that the trade-off between speed and accuracy is best achieved when using T=2T=2 stages. Hence, we selected this configuration for YOSO.

Number of proposal kernels. We investigate the impact of the number of proposal kernels in Tab. 10. The results indicate that the PQ performance improves when increasing the number of proposal kernels from 50 to 100, and saturates at 150. Meanwhile, the speed decreases when increasing the number of proposal kernels. The setting of nn=100 well balances the accuracy and speed for YOSO.

Conclusion

In this paper, we propose a real-time panoptic segmentation framework, termed YOSO. With YOSO, you only need to segment once for the masks of both foreground things and background stuff. YOSO includes a feature pyramid aggregator and a separable dynamic decoder to accelerate the pipeline. The CFA module in the feature pyramid aggregator re-parameters the IFA module, reducing FLOPs without extra costs. The SDC module in the separable dynamic decoder performs weight-sharing multi-head cross-attention, enhancing both speed and accuracy. Our extensive experiments demonstrate that YOSO is significantly faster than other prominent panoptic segmentation methods while maintaining competitive PQ performance. Given its effectiveness and simplicity, we hope YOSO can serve as a strong baseline and bring fresh insights for future research on real-time panoptic segmentation.

This work was supported by National Key R&D Program of China (2022ZD0118202), National Science Fund for Distinguished Young Scholars (No.62025603), National Natural Science Foundation of China (No.U21B2037, No.U22B2051, No.62176222, No.62176223, No.62176226, No.62072386, No.62072387, No.62072389, No.62002305 and No.62272401), and Natural Science Foundation of Fujian Province of China (No.2021J01002, No.2022J06001).

References

Appendix A Qualitative Results for Panoptic Segmentation

Appendix B Quantitive Results on Real-Time Instance Segmentation

Solving the instance segmentation task efficiently is one of the keys to achieving real-time panoptic segmentation. Therefore, we study the performance of YOSO for real-time instance segmentation in Tab. 11. Note that the model is not specifically trained for instance segmentation. We show the results from the model trained on the COCO training set for panoptic segmentation. From the results shown in Tab. 11, we find that YOSO also achieves competitive performance for instance segmentation on the COCO validation set. When scaling the input images to 550, YOSO achieves 38.7 FPS and 34.7 mAP. The speed is only 5.9 FPS lower than the current state-of-the-art model, i.e., SparseInst, while the mAP is 0.3 points higher. Specifically, the performance of YOSO on large objects, i.e., APl, is better than all the state-of-the-art models. For example, when scaling the input images to 448, YOSO still achieves 57.6 APl for large objects, which is approximately 2.2 points higher than the performance of SOLOv2. The result suggests that YOSO is good at segmenting large objects for instance segmentation.