TopFormer: Token Pyramid Transformer for Mobile Semantic Segmentation

Wenqiang Zhang, Zilong Huang, Guozhong Luo, Tao Chen, Xinggang Wang, Wenyu Liu, Gang Yu, Chunhua Shen

Introduction

Vision Transformers (ViTs) have shown considerably stronger results for a few vision tasks, such as image classification , object detection , and semantic segmentation . Despite the success, the Transformer architecture with the full-attention mechanism requires powerful computational resources beyond the capabilities of many mobile and embedded devices. In this paper, we aim to explore a mobile-friendly Vision Transformer specially designed for dense prediction tasks, e.g., semantic segmentation.

To adapt Vision Transformers to various dense prediction tasks, most recent Vision Transformers such as PVT , CvT , LeViT , MobileViT adopt a hierarchical architecture, which is generally used in Convolution Neural Networks (CNNs), e.g., AlexNet , ResNet . These Vision Transformers apply the global self-attention and its variants on the high-resolution tokens, which bring heavy computation cost due to the quadratic complexity in the number of tokens.

To improve the efficiency, some recent works, e.g., Swin Transformer , Shuffle Transformer , Twins and HR-Former , compute self-attention within the local/windowed region. However, the window partition is surprisingly time-consuming on mobile devices. Besides, Token slimming and Mobile-Former decrease calculation capacity by reducing the number of tokens, but sacrifice their recognition accuracy.

Among these Vision Transformers, MobileViT and Mobile-Former are specially designed for mobile devices. They both combine the strengths of CNNs and ViTs. For image classification, MobileViT achieves better performance than MobileNets with a similar number of parameters. Mobile-Former achieves better performance than MobileNets with a fewer number of FLOPs. However, they do not show advantages in actual latency on mobile devices compared to MobileNets, as reported in . It raises a question: Is it possible to design mobile-friendly networks which could achieve better performance on mobile semantic segmentation tasks than MobileNets with lower latency?

Inspired by MobileViT and Mobile-Former, we also make use of the advantages of CNNs and ViTs. A CNN-based module, denoted as Token Pyramid Module, is used to process high-resolution images to produce local featuresWe use ‘features’ and ‘tokens’ interchangeably here. pyramid quickly. Considering the very limited computing power on mobile devices, here we use a few stacked light-weight MobileNetV2 blocks with a fast down-sampling strategy to build a token pyramid. To obtain rich semantics and large receptive field, the ViT-based module, denoted as Semantics Extractor, is adopted and takes the tokens as input. To further reduce the computational cost, the average pooling operator is used to reduce tokens to an extremely small number, e.g., \nicefrac1(64×64)\nicefrac{{1}}{{(64\times 64)}} of the input size. Different from ViT , T2T-ViT and LeViT use the last output of the embedding layer as input tokens, we pool the tokens from different scales (stages) into the very small numbers (resolution) and concatenate them along the channel dimension. Then the new tokens are fed into the Transformer blocks to produce global semantics. Due to the residual connections in the Transformer block, the learned semantics are related to scales of tokens, denoted as scale-aware global semantics.

To obtain powerful hierarchical features for dense prediction tasks, scale-aware global semantics is split by the channels of tokens from different scales, then the scale-aware global semantics are fused with the corresponding tokens to augment the representation. The augmented tokens are used as the input of the segmentation head.

To demonstrate the effectiveness of our approach, we conduct experiments on the challenging segmentation datasets: ADE20K , Pascal Context and COCO-Stuff . We examine the latency on hardware, i.e., an off-the-shelf ARM-based computing core. As shown in Figure 1, our approach obtains better results than MobileNets with lower latency. To demonstrate the generalization of our approach, we also conduct experiments of object detection on the COCO dataset. To summarize, our contributions are as follows.

The proposed TopFormer takes tokens from different scales as input, and pools the tokens to the very small numbers, in order to obtain scale-aware semantics with very light computation cost.

The proposed Semantics Injection Module can inject the scale-aware semantics into the corresponding tokens to build powerful hierarchical features, which is critical to dense prediction tasks.

The proposed base model can achieve 5% mIoU better than that of MobileNetV3, with lower latency on an ARM-based mobile device on the ADE20K dataset. The tiny version can perform real-time segmentation on an ARM-based mobile device, with competitive results.

Related Work

In this section, we review recent approaches in terms of three aspects: 1) Light-weight Vision Transformer, 2) Efficient Convolutional Neural Networks, 3) Mobile Semantic Segmentation.

There are many explorations for the use of transformer structures in image recognition. ViT is the first work to apply a pure transformer to image classification, achieving state-of-the-art performance. Following that, DeiT introduces token-based distillation to reduce the amount of data necessary for training the transformer. T2T-ViT structures the image to tokens by recursively aggregating neighboring tokens into one token to reduce tokens length. Swin Transformer computes self-attention within each local window, resulting in linear computational complexity in the number of input tokens. However, these Vision Transformers and the follow-ups are often of a large number of parameters and heavy computation complexity.

To build a light-weight Vision Transformer, LeViT designs a hybrid architecture that uses stacked standard convolution layers with stride-2 to reduce the number of tokens, then appends an improved Vision Transformer to extract semantics. For classification task, LeViT can significantly outperform EfficientNet on CPU. MobileViT adopts the same strategy and uses the MobilenetV2 block instead of the standard convolution layer for downsampling the feature maps. Mobile-Former takes parallel structure with a bidirectional bridge and leverages the advantage of both MobileNet and transformer. However, the MobileViT and other ViT-based networks are significantly slower than MobileNets on mobile devices, as reported in . For the segmentation task, the input images are always of high-resolution. Thus it is even more challenging for ViT-based networks to execute faster than MobileNets. In this paper, we aim to design a light-weight Vision Transformer which can outperform MobileNets with lower latency for the segmentation task.

2 Efficient Convolutional Neural Networks

The increasing need of deploying vision models on mobile and embedded devices encourages the study on efficient Convolutional Neural Networks designs. MobileNet proposes an inverted bottleneck structure which stacks depth-wise and point-wise convolutions. IGCNet and ShuffleNet use channel shuffle/permutation operators to the channel to make cross-group information flow for multiple group convolution layers. GhostNet uses the cheaper operator, depth-wise convolutions, to generate more features. AdderNet utilizes additions to trade massive multiplications. MobileNeXt flips the structure of the inverted residual block and presents a building block that connects high-dimensional representations instead. EfficientNet and TinyNet study the compound scaling of depth, width and resolution.

3 Mobile Semantic Segmentation

The most accurate segmentation networks usually require computation at billions of FLOPs, which may exceed the computation capacity of the mobile and embedded devices. To speed up the segmentation and reduce the computational cost, ICNet uses multi-scale images as input and a cascade network to be more efficient. DFANet utilizes a light-weight backbone to speed up its network and proposes a cross-level feature aggregation to boost accuracy. SwiftNet uses lateral connections as the cost-effective solution to restore the prediction resolution while maintaining the speed. BiSeNet introduces the spatial path and the semantic path to reduce computation. AlignSeg and SFNet align feature maps from adjacent levels and further enhances the feature maps using a feature pyramid framework. ESPNets save computation by decomposing standard convolution into point-wise convolution and spatial pyramid of dilated convolutions. AutoML techniques are used to search for efficient architectures for scene parsing. NRD dynamically generates the neural representations with dynamic convolution filter networks. LRASPP adopts MobileNetV3 as encoder and proposes a new efficient segmentation decoder Lite Reduced Atrous Spatial Pyramid Pooling (LR-ASPP), and it still serves as a strong baseline for mobile semantic segmentation.

Architecture

Our overall network architecture is shown in Figure 2. As we can see, our network consists of several parts: Token Pyramid Module, Semantics Extractor, Semantics Injection Module and Segmentation Head. The Token Pyramid Module takes an image as input and produces the token pyramid. The Vision Transformer is used as a semantics extractor, which takes the token pyramid as input and produces scale-aware semantics. The semantics are injected into the tokens of the corresponding scale for augmenting the representation by the Semantics Injection Module. Finally, Segmentation Head uses the augmented token pyramid to perform the segmentation task. Next, we present the details of these modules.

Inspired by MobileNets , the proposed Token Pyramid Module consists of stacked MobileNet blocks . Different from MobileNets, the Token Pyramid Module does not aim to obtain rich semantics and large receptive field, but uses fewer blocks to build a token pyramid. We show the layer settings of the Token Pyramid Module in Subsection 3.4.

2 Vision Transformer as Scale-aware Semantics Extractor

The Scale-aware Semantics Extractor consists of a few stacked Transformer blocks. The number of Transformer blocks is LL. The Transformer block consists of the multi-head Attention module, the Feed-Forward Network (FFN) and residual connections. To keep the spatial shape of tokens and reduce the numbers of reshape, we replace the Linear layers with a 1×11\times 1 convolution layer. Besides, all of TopFormer’s non-linear activations are ReLU6 instead of GELU function in ViT.

For the Multi-head Attention module, we follow the settings of LeViT , and set the head dimension of keys KK and queries QQ to have D=16D=16, the head of values VV to have 2D=322D=32 channels. Decreasing the channels of KK and QQ will reduce computational cost when calculating attention maps and output. Meanwhile, we also drop the Layer normalization layer and append a batch normalization to each convolution. The batch normalization can be merged with the preceding convolution during inference, which can run faster over layer normalization.

For the Feed-Forward Network, we follow to enhance the local connections of Vision Transformer by inserting a depth-wise convolution layer between the two 1×11\times 1 convolution layers. The expansion factor of FFN is set to 2 to reduce the computational cost. The number of Transformer blocks is LL and then the number of heads will be given in subsection 3.4.

As shown in Figure 2, the Vision Transformer takes the tokens from different scales as input. To further reduce the computation, the average pooling operator is used to reduce the numbers of tokens from different scales to 164×64\frac{1}{64\times 64} of the input size. The pooled tokens from different scales have the same resolution, and they are concatenated together as the input of the Vision Transformer. The Vision Transformer can obtain full-image receptive field and rich semantics. To be more specific, the global self-attention exchanges information among tokens along the spatial dimension. The 1×11\times 1 convolution layer will exchange information among tokens from different scales. In each Transformer block, the residual mapping is learned after exchanging information of tokens from all scales, then residual mapping is added into tokens to augment the representation and semantics. Finally, the scale-aware semantics are obtained after passing through several transformer blocks.

3 Semantics Injection Module and Segmentation Head

After obtaining the scale-aware semantics, we add them with the other tokens TNT^{N} directly. However, there is a significant semantic gap between the tokens {T1,...,TN}\{\mathbf{T}^{1},...,\mathbf{T}^{N}\} and the scale-aware semantics. To this end, the Semantics Injection Module is introduced to alleviate the semantic gap before fusing these tokens. As shown in Fig. 2, the Semantics Injection Module (SIM) takes the local tokens of the Token Pyramid module and the global semantics of the Vision Transformer as input. The local tokens are passed through the 1×11\times 1 convolution layer, followed by a batch normalization to produce the feature to be injected. The global semantics are fed into the 1×11\times 1 convolution layer followed by a batch normalization layer and a sigmoid layer to produce semantics weights, meanwhile, the global semantics also passed through the 1×11\times 1 convolution layer followed by a batch normalization. The three outputs have the same size. Then, the global semantics are injected into the local tokens by Hadamard production and the global semantics are also added with the feature after the injection. The outputs of the several SIMs share the same number of channels, denoted as MM.

After the semantics injection, the augmented tokens from different scales capture both rich spatial and semantic information, which is critical for semantic segmentation. Besides, the semantics injection alleviates the semantic gap among tokens. The proposed Segmentation head firstly up-samples the low-resolution tokens to the same size as the high-resolution tokens and element-wise sums up the tokens from all scales. Finally, the feature is passed through two convolutional layers to produce the final segmentation map.

4 Architecture and Variants

To customize the network of various complexities, we introduce TopFormer-Tiny (TopFormer-T), TopFormer-Small (TopFormer-S) and TopFormer-Base (TopFormer-B), respectively.

The model size and FLOPs of the base, small and tiny models are given in the Table 1. The base, small and tiny models have 8, 6 and 4 heads in each multi-head self-attention module, respectively, and have M=256M=256, M=192M=192 and M=128M=128 as the target numbers of channels. For more details of network configures, please refer to the supplementary materials.

To achieve better trade-offs between accuracy and the actual latency, we choose the tokens from last three scales T2,T3T^{2},T^{3} and T4T^{4} as the inputs of SIM and the segmentation head.

Experiments

In this section, we first conduct experiments on several public datasets. We describe implementation details and compare results with other works for semantic segmentation tasks. We then conduct ablation studies to analyze the effectiveness and efficiency of different parts. Finally, we report the performance on object detection task to show the generalization ability of our method.

We perform experiments over three datasets, ADE20K , PASCAL Context and COCO-Stuff . The mean of class-wise intersection over union (mIoU) is set as our evaluation metric. The full-precision TopFormer models are converted to TNN , the latency is then measured on an ARM-based computing core. ADE20K: The ADE20K dataset contains 25K images in total, covering 150 categories. All images are split into 20K/2K/3K for training, validation, and testing. PASCAL Context: The Pascal Context dataset has 4998 scene images for training and 5105 images for testing. There are 59 semantic labels and 1 background label. COCO-Stuff: The COCO-Stuff dataset augments COCO dataset with pixel-level stuff annotations. There are 10000 complex images selected from COCO. The training set and test set consist of 9K and 1K images respectively.

1.2 Implementation Details

Our implementation is built upon MMSegmentation and Pytorch. It utilizes ImageNet pre-trained TopFormerThe TopFormer-base achieve 75.3% Top-1 acc. on ImageNet-1K. as the backbone. The standard BatchNorm layer is replaced by the Synchronize BatchNorm to collect the mean and standard-deviation of BatchNorm across multiple GPUs during training. For ADE20K dataset, we follow Segformer to use 160K scheduler and batch size is 16. The training iteration of COCO-Stuff and PASCAL Context is 80K. For all methods and datasets, the initial learning rate is set as 0.00012 and weight decay is 0.01. A “poly” learning rate scheduled with factor 1.0 is used. On ADE20K, we adopt the same data augmentation strategy as for fair comparison. The training images are augmented by first randomly scaling and then randomly cropping out the fixed size patches from the resulting images. In addition, we also apply random resize, random horizontal flipping, random cropping etc. On COCO-Stuff and PASCAL Context, we use the default augmentation strategy of . We resize and crop the images to 480×480480\times 480 for PASCAL Context and 512×512512\times 512 for COCO-Stuff during training. Finally, we report the single scale results on validation set to compare with other methods. During inference, we follow the common strategy to rescale the short side of images to training cropping size for ADE20K and COCO-Stuff. As for PASCAL Context, the images are resized to 480×480480\times 480 and then fed into our network.

1.3 Experiments on ADE20K

We compare our TopFormer with the previous approaches on the ADE20K validation set in Table 1. Actual latency is measured on the mobile device with a single Qualcomm Snapdragon 865 processor. Here, we choose light-weight vision transformers (ViT) and efficient convolution neural networks (CNNs) as the encoder. Besides, the various decoders are also included in Table 1. Among all methods in Table 1, Deeplabv3+ based on MobilenetV2 achieve best mIoU (38.1%), however, the latency is more than 1000 ms, which restrict its application on mobile devices.

Among these CNNs based baselines, the approach which takes mobilenetV3-large as encoder and LR-ASPP as decoder, achieves good trade-off between computation (2.0 GFLOPs) and accuracy (33.1 mIoU). Following , we also reduce the channel counts of all feature layers in the last stage for further reducing the computation cost, denoted as MobileNetV3-Large-reduce. Based on the lighter backbone, LR-ASPP could achieve 32.3% mIoU with the lower latency (81 ms). Our small version of TopFormer is 3.8% more accurate compared to a LR-ASPP model with comparable latency. The tiny version of TopFormer could achieve comparable performance compared to LR-ASPP with 2×2\times less computation (0.6G vs. 1.3G). Lite-ASPP is the reduced channel version of Deeplabv3+.

Among these ViT based baselines, HR-NAS-B uses search techniques to introduce Transformer block into the HRNet design, also achieves good trade-off between computation amount (2.2 GFLOPs) and accuracy (34.9 mIoU). Our small version of TopFormer is 1.2% more accurate compared to HR-NAS-B model with fewer computation. SegFormer achieves great performance (37.4 mIoU) with fewer parameters (3.8M), although SegFormer adopts the efficient multi-head self-attention, the computation is still heavy due to a large number of tokens. Our base version of TopFormer could achieve comparable performance compared to SegFormer with more than 4×4\times less computation (1.8 GFLOPs vs. 8.4 GFLOPs).

To achieve real-time segmentation on the ARM-based mobile device, we resize the input image to 448×448448\times 448 and feed it into TopFormer-tiny, the inference time is reduced to 32ms with a slight performance drop. To the best of our knowledge, it is the first ViT based method could achieve real-time segmentation on the ARM-based mobile device with competitive results.

1.4 Ablation Study

We first conduct ablation experiments to discuss the influence of different components, including the token pyramid, Semantic Injection Module and Segmentation Head. Without loss of generality, all results are obtained by training on the training set and evaluation on the validation (val) set.

Here, we discuss the Token Pyramid from two aspects, the influence of taking Token Pyramid as input and the influence of choosing Tokens from different scales as output. As reported in Table 2, we conduct the experiments that take stacked tokens from different scales as input of the Semantics Extractor, and take the last token as input of the Semantics Extractor, respectively. For fair comparison, we append a 1×11\times 1 convolution layer to expand the channels as same as the stacked tokens. The experimental results demonstrate the effectiveness of using the token pyramid as input.

After obtaining scale-aware semantics, the SIM will inject the semantics into the local Tokens. To pursuit better trade-off between the accuracy and the computation cost, we try to choose tokens from different scales for injecting. As shown in Table 3, using tokens from {14,18,116,132}\{\frac{1}{4},\frac{1}{8},\frac{1}{16},\frac{1}{32}\} could achieve the best performance with the heaviest computation. Using tokens from {116,132}\{\frac{1}{16},\frac{1}{32}\} achieves worse performance with the lightest computation. To achieve a good trade-off between the accuracy and the computation cost, we choose to use the tokens from {18,116,132}\{\frac{1}{8},\frac{1}{16},\frac{1}{32}\} in all other experiments.

Here, we have conducted experiments on Topformer-T to check the SASE. The results are shown in the Table 4. Here, we use Topformer without SASE as the baseline. Adding SASE brings about 10% mIoU gains, which is a significant improvement. To verify the multi-head self-attention module (MHSA) in the Transformer block, we remove all MHSA modules and add more FFNs for a fair comparison. The results demonstrate the MHSA could bring about 2.4% mIoU gains, which is an efficient and effective module under the careful architecture design. Meanwhile, we compare the SASE with the popular contextual models, such as ASPP and PPM, on the top of TPM. As shown in Table 4, “+SASE” could achieve better performance with much less computation cost than “+PSP” and “+ASPP”. The experimental results demonstrate that the SASE is more appropriate for use in mobile devices.

Due to the close relationship of Semantic Injection Module and Segmentation Head, we discuss these two together. Here, we discuss the design of Semantic Injection Module at first. As shown in Table 5, multiplying the local tokens and the semantics after a Sigmoid layer, denoted as “SigmoidAttn”. Adding the semantics from Semantics Extractor into the corresponding local tokens, denoted as “SemInfo”. Compared with “SigmoidAttn” and “SemInfo”, adding “SigmoidAttn” and “SemInfo” simultaneously could bring pretty improvement with a little extra computation.

Here, we discuss the design of Segmentation Head. After passing the feature into Semantic Injection Module, the output hierarchical features are with both strong semantics and rich spatial details. The proposed segmentation head simply adds them together and then uses two 1×11\times 1 convolution layers to predict the segmentation map. We also design the other two Segmentation Heads, as shown in Figure 4. The “Sum Head” is identical to only adding “SemInfo” in SIM. The “Concat Head” uses a 1×11\times 1 convolution layer to reduce the channels of the outputs of SIM, then the features are concatenated together. Compared with “Concat Head” and “Sum Head”, the current segmentation head could achieve better performance.

In the paper, we donate the number of channels in SIM as MM. Here, we study the influence of different M in SIM and find a suitable MM to achieve a good trade-off. As shown in Table 7, the M=256,192,128M=256,192,128 achieve similar performance with very close computation. Thus, we set M=128,192,256M=128,192,256 in tiny, small and base model, respectively.

The tokens from different stages are pooled to fixed resolution, namely the output stride. The results of different resolutions are shown in Table 8. s32, s64, s128 are denoted that the resolution of pooled tokens are 132×32,164×64,1128×128\frac{1}{32\times 32},\frac{1}{64\times 64},\frac{1}{128\times 128} of input size. Considering the trade-off of computation and accuracy, we choose s64 as the output stride of input tokens of the Semantics Extractor.

Here, we make the statistics of the computation, parameters and latency of the proposed TopFormer-Tiny. As shown in Figure 3, although the Semantics Extractor has most parameters (74%), the FLOPs and actual latency of the Semantics Extractor is relatively low (about 10%).

1.5 Experiments on Pascal Context

We compare our TopFormer with the previous approaches on the Pascal Context test set in Table 9. We evaluate the performance over 59 categories and 60 categories (including background), respectively. It is obvious that our approach achieves better performance than all previous approaches based on CNNs or ViT with fewer computation. For better understanding, the FLOPs of the backbone and the head are measured, respectively. The proposed method can achieve best performance with the lightest backbone and head.

1.6 Experiments on COCO-Stuff

We compare our TopFormer with the previous approaches on the COCO-Stuff validation set in Table 10. The FLOPs of the backbone and the head are measured, respectively. It can be seen that our approach achieves the best performance, and the base version of TopFormer is 8% more accurate compared to a MobileNetV3 model with comparable computation.

2 Object Detection

To further demonstrate the generalization ability of the proposed TopFormer, we conduct object detection task on COCO dataset. COCO consists of 118K images for training, 5K for validation and 20K for testing. We train all models on train2017 split and evaluate all methods on val2017 set. We choose RetinaNet as object detection methods and adopt different backbones to produce feature pyramid. Our implementation is built on MMdetection and Pytorch. For the proposed TopFormer, we replace the segmentation head with the detection head in RetinaNet. As shown in Table 11, The RetinaNet based on TopFormer could achieve better performance than MobileNetV3 and ShuffleNet with lower computation.

Conclusion and Limitations

In this paper, we present a new architecture for mobile vision tasks. With a combination of the advantages of CNNs and ViT, the proposed TopFormer achieves a good trade-off between accuracy and the computational cost. The tiny version of TopFormer could yield real-time inference on an ARM-based mobile device with competitive result. The experimental results demonstrate the effectiveness of the proposed method. The major limitation of TopFormer is the minor improvements on object detection. We will continue to promote the performance of object detection. Besides, we will explore the application of TopFormer in dense prediction in the future work.

Acknowledgement

This work was in part supported by NSFC (No. 61733007, No. 61876212, No. 62071127 and No. 61773176) and the Zhejiang Laboratory Grant (No. 2019NB0AB02 and No. 2021KH0AB05).

References

ImageNet Pre-training

For fair comparison, we also use the ImageNet pre-trained parameters as initialization. As shown in Figure 5, the classification architecture of the proposed TopFormer appends the average pooling layer and Linear layer on the global semantics for producing class scores. Due to the small resolution (224×224224\times 224) of input images, we set the target resolution of input tokens of the Semantics Extractor is 132×32\frac{1}{32\times 32} of input size. The classification results are shown in Table 12. Because our target task is mobile semantic segmentation, we do not explore more technologies, e.g. more epochs and distillation uesd in LeViT, to further improve the accuracy. In the future work, we will continue to improve the classification accuracy.

Network Structure

The detailed network structures are given in Table 14. Although the Token Pyramid Module have the most layers, as the statistics of the computation and parameters in the paper, the ViT-based Semantics Extractor accounts for the vast majority of parameters.

The Performance on Cityscapes

Our implementation is based on MMSegmentation and Pytorch. We perform 80K iterations. The initial learning rate is 0.0003 and weight decay is 0.01. A poly learning rare scheduled with factor 1.0 is used. For full-resolution version, the training images are randomly scaling and then cropping to fixed size of 1024×10241024\times 1024. As for the half-resolution version, the training images are resized to 1024×5121024\times 512 and randomly scaling, the crop size is 1024×5121024\times 512. We follow the data augmentation strategy of Segformer for fair comparison.

To validate the performance of the proposed method, we directly fed a full-resolution input and a half-resolution input into the trained segmentation models for testing, respectively. As shown in Table 13, the proposed method with a full-resolution, denoted as Ours(f), achieves about 2.6% higher accuracy in mIoU than L-ASPP based on MobileNetV2 with lower computation. The experimental results demonstrate that TopFormer could achieve good trade-off between accuracy and computation even if the input image is with large resolution.

1 Visualization

We present some visualization comparisons among the proposed TopFormer and other CNNs- and ViT-based methods on the ADE20K validation (val) set in Figure 6. Here, we choose deeplabv3+ based on mobilenetV2 as a representative of CNNs-based methods and Segformer as a representative of ViT-based methods. These two methods both have larger model size and computational cost. As shown in Figure 6, the proposed method could achieve better segmentation results than these two methods.