SeMask: Semantically Masked Transformers for Semantic Segmentation
Jitesh Jain, Anukriti Singh, Nikita Orlov, Zilong Huang, Jiachen Li, Steven Walton, Humphrey Shi
Introduction
Semantic Segmentation aims to perform dense prediction for labeling each pixel in an image corresponding to the class that the pixel represents. Transformer-based vision networks have outperformed Convolutional Neural Networks on the image-classification task . In modern times, transformer backbones have shown impressive performance when transferred to downstream tasks like semantic segmentation .
Most of the architectural designs in vision transformers approach the problem in either of the two ways: (i) Use an existing pretrained backbone as an encoder and transfer it to downstream tasks using pre-existing standard decoders such as, Semantic FPN or UperNet ; OR (ii) design a new encoder-decoder network where the encoder is pretrained on ImageNet for the semantic segmentation task. Both of these ways, as mentioned earlier, involve finetuning the encoder backbone on the segmentation task. Finetuning from a large-scale dataset help early attention layers to incorporate local information at lower layers of the transformers . However, it can still not harness the semantic context during finetuning due to the relatively smaller size of the dataset and a change in the number and nature of semantic classes from classification to the segmentation task. Hierarchical vision transformers tackle the problem with progressive downsampling of features along the stages, although they still lack the semantic context of the image.
Liu et al. introduced the Swin Transformer, which constructs hierarchical feature maps making it compatible as a general-purpose backbone for major downstream vision tasks. proposed to use two attention: globally sub-sampled and locally sub-samples on top of PVT and CPVT for effective segmentation. Xie et al. further modified the hierarchical transformer encoder by making it free from positional-encoding and thus robust to different resolutions as generally found in the segmentation task. All these works modified the encoders to make them work better for downstream tasks like segmentation and achieved success to an impressive extent. Still, they did not pay attention to capturing the semantic-level contextual information of the whole image. A lack of semantic contextual information leads to sub-optimal segmentation performance, especially in the case of small objects where those get merged with the boundaries of the larger categories, leading to wrong predictions. Recently, tried to tackle this issue by designing a pure transformer-based decoder that jointly processes the patch and class embedding. However, it does not perform efficiently for tiny variants and fails with hierarchical architectures leading to sub-optimal performance when used with major transformer backbones like Swin , and Twins transformers.
Jin et al. in proposed ISNet to model the image level contextual information along with semantic level contextual information by introducing the SLCM and ILCM modules in the decoder structure. However there is still a caveat: ISNet is a CNN based method and only focuses on the decoder part of the network, leaving out the encoder unchanged.
To address the issues mentioned above, we propose the SeMask framework that incorporates semantic information into hierarchical vision transformer architectures and augments the global feature information captured by the transformers with the semantic context. The existing frameworks formulate the architecture as an encoder-decoder structure with transformers pretrained on ImageNet acting as the encoders and using a specialized decoder for semantic segmentation. In contrast to directly using the hierarchical transformers as a backbone, we insert a Semantic Layer after the Transformer Layer at each stage in the backbone, giving us the SeMask version of the backbone as illustrated in Fig. 1. We use a lightweight semantic decoder to accumulate the semantic maps from all the stages, and a standard decoder like Semantic-FPN for the main per-pixel prediction. The added semantic modeling with feature modeling throughout the encoder helps us improve the performance of the semantic segmentation task. In Sec. 4, we integrate the proposed SeMask block into the Swin Transformer and Mix Transformer backbones. Our experimental results show considerable improvement in semantic segmentation for both backbones on two different datasets. To summarize, our contributions are three fold:
To the best of our knowledge, we are the first to study the effect of adding semantic context to pretrained transformer backbones for the semantic segmentation task. Furthermore, we introduce a SeMask Block which can be plugged into any existing hierarchical vision transformer. We provide empirical evidence by integrating SeMask into Swin-transformer and Mix-Transformer , and achieving considerable performance improvement.
We also propose to use a simple semantic decoder for aggregating the semantic priors from different stages of the encoder. The semantic priors receive supervision from the ground truth using a per-pixel cross-entropy loss.
Lastly, we provide an in-depth analysis of the SeMask Block’s effect on two different datasets: ADE20K and Cityscapes. We achieve the new state-of-the-art performance on the ADE20K dataset and an improvement above on the Cityscapes dataset.
Related Work
Semantic segmentation broadly formulates to a dense per-pixel classification task. The seminal work of FCN introduced the use of deep CNNs, removing fully connected layers to tackle the segmentation task. Several following works were built upon the same idea of using the encoder-decoder architecture. introduced the use of atrous convolutions inside the DCNN to tackle the signal downsampling issue. Later, various works focused on the aggregating long-range context in the final feature map: ASPP uses atrous convolutions with different dilation rates; PPM uses pooling operators with different kernel sizes.
The recent DCNN based models focus on efficiently aggregating the hierarchical features from a pretrained backbone based encoder with specially designed modules: introduce attention modules in the decoder; use different forms of non-local blocks ; proposes a novel FAM module to solve the misalignment issue using semantic flow; AlignSeg proposes aligned feature aggregation module and aligned context modeling module to make contextual features be better aligned. uses a segmentation shelf for better information flow. In this work, we also follow the established direction to use a pretrained backbone and aggregating the hierarchical features using the Semantic-FPN decoder.
2 Transformers for Segmentation
After being heavily used in Natural Language Processing field, transformer based models have gained popularity for various computer vision tasks since the introduction of ViT for image classification . SETR used ViT as an encoder and two decoders based upon progressive upsampling and multi-level feature aggregation. SegFormer proposed to use a hierarchical pyramid vision transformer network as an encoder with an MLP based decoder to obtain the segmentation mask. Segmenter designed mask transformer as a decoder, which uses learnable class-map tokens to enhance decoding performance. MaskFormer defines the problem of per-pixel classification from a mask classification point of view, creating an all-in-one module for all segmentation tasks. Mask2Former further evolves masked attention to solve panoptic, instance and semantic segmentation tasks in one framework. Most recent transformer-based segmentation frameworks are based on finetuning a pretrained hierarchical backbone as an encoder, and standard decoders like Semantic-FPN and UperNet to the segmentation task. In this work, we follow the same paradigm and, in addition, propose a framework to enhance the finetuning ability of the pretrained vision transformer backbone. Note that there is also recent concurrent work like SwinV2 that reaches new state-of-the-art performance on ADE20k benchmark by using improved and giant backbones (e.g. SwinV2-G with 3.0 billion parameters). That is out of the scope of this work and we follow the current practice mainly based on Swin-L backbone. Theoretically, we can get even better performance if we apply our approach to such giant models.
3 Semantic Context in Segmentation
Zhang et al. proposed the Context Encoding Module in which captures the global semantic context along with a feedback loop to balance the importance of classes in the features extracted by a ResNet backbone . More recently, focus on capturing and integrating the semantic-level contextual information along with the image-level context with specially designed decoders which shows significant improvement in DCNN based methods. Each of these works captures the semantic context after the encoding stage based on the extracted features and not the encoder’s ability to capture the semantic features.
In this work, we argue that semantic information is lost during the encoding stage and hence, propose a framework to capture semantic information which can be plugged into any pretrained vision transformer backbone network.
Method
An overview of our architecture with Swin-Transformer backbone is shown in Fig. 2. The RGB input image, size , is first split into non-overlapping patches of size . The smaller size of the patch supports dense prediction in segmentation. These patches act as tokens and are given as input to the hierarchical vision transformer encoder, which is the Swin-Transformer in our architecture. The encoding step consists of four different stages of hierarchical feature modeling. Every stage during the encoding step consists of two layers: The transformer layer, which is number of Swin Transformer blocks (Fig. 3(a)) stacked together and Semantic Layer with number of SeMask Attention blocks (Fig. 3(b)). We collectively refer to the Transformer Layer and Semantic Layer at each stage as our SeMask Block. The patch tokens pass through each stage at of the original image resolution for the feature maps and intermediate semantic-prior maps extraction.
In the encoder part of the network, the Semantic Layer takes in features from the Transformer Layer as inputs and returns the intermediate semantic-prior maps and semantically masked features (Fig. 3(b)). When we plug the SeMask Attention Block into other hierarchical vision transformers, the Transformer Layer consists of attention blocks corresponding to the specific backbone, like Efficient-Self Attention-based Transformer Layer for the Mix Transformer backbone. The semantically masked features from each stage are aggregated using the semantic-FPN decoder for producing the final dense-pixel prediction. Moreover, the semantic-prior maps from all the stages are aggregated using a lightweight upsample & sum operation-based semantic decoder to predict the semantic-prior for the network during training. Both decoders’ outputs are supervised using a weighted per-pixel cross-entropy loss. These additional semantic-prior maps greatly assist the feature extraction and eventually improve the performance on the semantic segmentation task.
The resulting feature from the Transformer Layer after the last Swin Transformer block then acts as an input to the subsequent semantic layer in the same stage as shown in Fig. 3.
Semantic Layer. The Semantic Layer follows the Transformer Layer at each stage of our hierarchical vision transformer. Unlike the Transformer Layer, the Semantic Layer’s significance is in modeling the semantic context, which is used as a prior for calculating a segmentation score to update the feature maps based on guidance from the semantic nature present in the image. Inside each semantic layer, there are SeMask attention blocks (Fig. 3(b)). Inspired by the shifted window-based division of the tokens for efficient computation cost, we also divide the input to our SeMask blocks into windows with cross-window connections before calculating the segmentation score using a single-head self-attention operation. The SeMask block is responsible for capturing the semantic context in our encoder. It updates the features from the transformer layer from the segmentation score providing guidance and giving a semantic-prior map for efficient supervision of the semantic modeling during training. SeMask attention block divides the features from the preceding transformer layer into three entities: Semantic Query (), Semantic Key (), and Feature Value (). We get and by projecting the features onto the semantic space. The dimension of both and is where is equal to the number of classes, and the dimension of is where is the embedding dimension, is the length of the sequence with window size which we set as equal to that used inside the transformer layer. returns the semantic map, and a segmentation score is calculated using and . The score is passed through a softmax and is used to update as shown in Fig. 3(b). This SeMask attention equation is expressed as follows:
We perform a matrix multiplication between the feature values and the segmentation score. The matrix product is later passed through a linear layer and multiplied with a learnable scalar constant , used for smooth finetuning. After a residual connection , we finally get the modified features, rich with semantic information which we call the Semantically Masked features. The semantic queries are later used to predict the semantic-prior map.
2 Decoder
We use two decoders to aggregate the features and the semantic-prior maps respectively from the different stages in the encoder.
For aggregating the semantically masked features, we employ the popular Semantic-FPN decoder . The Semantic-FPN fuses the features from different stages with a series of convolution, bilinear upsampling, and sum operations, making it efficient and straightforward as a segmentation decoder for our purpose. In addition, we use a lightweight semantic decoder during training to provide ground truth supervision to the semantic-prior maps at every stage of the encoder. As the semantic-prior maps have the channel dimension of in each stage, we only employ a series of upsampling and sum operations to aggregate the maps with being equal to the number of classes in the dataset. Lastly, the output from both the decoders is up-scaled to the resolution of the original image for the final predictions as shown in Fig. 2.
3 Loss function
To train our model’s parameters, we calculate the total loss as a summation of two per-pixel cross-entropy losses: and . The loss is calculated on the main prediction from the Semantic-FPN decoder and loss is calculated on the semantic-prior prediction from our light-weight decoder. contains the main prediction of the network and denotes the semantic-prior prediction. We define our losses on and as follows:
Here, denotes for converting the ground truth class label stored in into one-hot format, denotes that the summation is carried out over all the pixels of the , and is the cross-entropy loss. We empirically set (check appendix for more details).
Experiments
We compare our approach with Swin Transformer , and Mix-Transformer with extensive experiments to demonstrate the effectiveness of the SeMask framework. We also ablate the SeMask structure and confirm that providing a semantic-prior to mask out the features improves semantic segmentation performance. The experiments are performed on two widely used datasets: ADE20K and Cityscapes . We include more experimental results in the appendix proving that our method is dataset agnostic.
ADE20K. ADE20K is a scene parsing dataset covering 150 fine-grained semantic concepts and it is one of the most challenging semantic segmentation datasets. The training set contains 20,210 images with 150 semantic classes. The validation and test set contain 2,000 and 3,352 images respectively.
Cityscapes. Cityscapes is an urban street driving dataset for semantic segmentation consisting of 5,000 images from 50 cities with 19 semantic classes. There are 2,975 images in the training set, 500 images in the validation set and 1,525 images in the test set.
Metrics. We report mean Intersection-over-Union () over all classes.
2 Implementation details
Transformer models. For the encoder, we build upon the Swin Transformer and consider the Tiny, Small, Base and Large variants as described in Tab. 1. The variation in number of parameters among the baselines is due to the number of transformer blocks () (Fig. 3(a)) and the embedding dimension () for each stage of the model. The number of heads () of a shifted window based multi-headed self-attention (SW-MSA) or Swin Transformer block varies from stage to stage. The hidden size of the MLP following SW-MSA is four times the embedding dimension at the corresponding stage. We also experiment with the MiT-B4 backbone variant of the Mix-Transformer on the ADE20K dataset.
In the following sections, we use an abbreviation to describe the model variant. For example, Swin-T denotes the Tiny variant. The backbones pretrained on ImageNet-22k and with resolution are denoted with a †: Swin-B†. All the other models are pretrained on ImageNet-1k and with resolution.
Network Initialization. Our SeMask models are initialized with publicly available models. The Tiny and Small variants are pre-trained on ImageNet-1k with an image resolution of . The Base and Large variants are pretrained on ImageNet-22k with a resolution of . We keep the window size fixed as in the pretrained models and fine-tune the models for the semantic segmentation task at higher resolution depending on the dataset. Following , we include relative position bias while calculating the attention scores. The decoders, described in Sec. 3.2 are initialized with random weights from a normal distribution .
Data augmentation. During training, we perform mean subtraction, scaling the image to a ratio randomly sampled from , random left-right flipping, and color jittering. We randomly crop large images and pad small images to a fixed size of for ADE20K and for Cityscapes. On ADE20K, we train our largest model Semask-L† FPN with a resolution, matching the resolution used by the Swin-Transformer .
Training Settings. To fine-tune the pre-trained models on the semantic segmentation task, we employ the AdamW optimizer with a base learning rate . Following the seminal work of DeepLab we adopt the poly learning rate decay where and represent the current iteration number and the total iteration number. We use a linear warmup strategy for 1,500 iterations.
For ADE20K, we set the base learning rate to , weight decay to and train for iterations with a batch size of .
For Cityscapes, we set to , a weight decay of and train for iterations with a batch size of .
Inference. To handle varying image sizes during inference, we keep the aspect ratio intact and resize the image to a resolution with the smaller edge resized to the training resolution and consequently rescaled to the original dimensions before calculating the metric score. For multi-scale inference, following standard practice we use rescaled versions of the image with scaling factors of .
3 Ablation Studies
In this section, we ablate different variants of our SeMask framework. We investigate the model size, semantic attention, number of SeMask blocks (), effect of the learnable scalar constant () inside the SeMask block and the pretraining dataset as well as image resolution. Unless stated otherwise, we use the Semantic-FPN as our decoder for the main prediction and report results using single-scale (s.s.) inference on the ADE20K val dataset.
Transformer size. We study the impact of transformers size on performance in Tab. 2 by experimenting with the four different Swin variants: Tiny, Small, Base and Large with for all the experiments. Our method gives improvement consistently over all the baseline variants with the improvement on the Cityscapes dataset being more impressive due to the fewer number of classes in the segmentation dataset creating a stronger prior.
We evaluate and record the mIoU scores for the baseline Swin models by training our networks using their publicly released code based on the MMSegmentation Library .
Semantic Attention. We study the impact of the semantic attention operation calculated inside the SeMask Block on performance in Tab. 3 by replacing the SeMask Block with a simple single-head self-attention block on the Swin-Tiny variant. It is evident that simple attention does not help improve the results proving the validity and effectiveness of our SeMask Block.
Learnable Constant (). We study the impact of on performance in Tab. 4, by removing it for the Tiny and Small variants. We observe that the inclusion of is potent to the success of the SeMask block as it acts as a tuning factor for the modified features, keeping the noise from weights’ initialization in check. We also observe that during inference for different stages in the encoder.
Number of SeMask Blocks (). In Tab. 5 we study the impact of number of SeMask attention blocks on performance by changing the values of inside each semantic layer on the Swin-Tiny variant. We observe that is the best setting. Interestingly, when stacking multiple blocks in a layer, we observe that inputting the from the previous SeMask block into the later one gives better performance than obtaining from the features. This shows that extracting semantic features using a single semantic attention operation is the optimum setting.
Pretraining Dataset. We study the impact of the pretraining dataset (ImageNet-1k v/s ImageNet-22k) on performance in Tab. 6 by training and evaluating the Base variant pretrained on various settings. Our framework is agnostic to the pretraining setting showing improvement for all combinations mainly used for the ImageNet pretraining: (i) ImageNet-1k and image resolution; (ii) ImageNet-22k and image resolution; and (iii) ImageNet-22k and image resolution.
4 Main Results
ADE20K. Using SeMask Swin-L† as the encoder and Mask2Former-MSFaPN as our decoder for the main prediction, we achieve a new state-of-the-art performance with scores of and on the single-scale and multi-scale mIoU metric, respectively. Following , our models were trained on images. We also achieve competitive results with our SeMask Swin-L† backbone with Semantic-FPN to the Swin-L† based UPerNet model as shown in Tab. 8.
We also integrate our SeMask into the MiT-B4 based SegFormer model as shown in Tab. 8 and achieve an improvement of on the single scale mIoU and improvememt on the multi-scale mIoU metric scores. This supports our claim that SeMask can be plugged into any existing hierarchical vision transformer and show performance improvement.
Cityscapes. Tab. 7 reports the performance of SeMask on Cityscapes. Semask Swin-L† is competitive with other state-of-the-art methods with SeMask Swin-L† Mask2Former achieving mIoU. We train our SeMask-L Mask2Former on images following Mask2Former .
Qualitative results. Fig. 4 shows a qualitative comparison of Swin-T FPN and SeMask-T FPN on the Cityscapes dataset generated using the MMSegmentation library . It is evident that SeMask-T FPN is able to generate better class-wise predictions than the Swin-T FPN. As shown in the second row in Fig. 4, we are able to segment the pole with out SeMask-T FPN, while Swin-T FPN fails to do so. Similarly in the third row, we are better able to segment the boundary.
Conclusion
This paper argues that directly finetuning off-the-shelf pretrained transformer backbone networks as encoders for semantic segmentation does not consider the semantic context tied up with the images. We claim that adding a semantic prior to guide the encoder’s feature modeling enhances the finetuning process for semantic segmentation. Furthermore, to support our claim, we propose the SeMask Block, which can be plugged into any existing hierarchical vision transformer and uses a semantic attention operation to capture the semantic context and augment the semantic representation of the feature maps. We train and evaluate the proposed framework building on the Swin-Transformer and Mix-Transformer backbones based networks and show a considerable improvement in the semantic segmentation performance on the Cityscapes and ADE20K dataset, with improvements above on the Cityscapes dataset. We provide a comprehensive experimental analysis applying SeMask to different backbone variants and achieving considerable performance improvement in every setting. Our method also achieves the new state-of-the-art performance on the ADE20K dataset. As a direction for future research, it will be interesting to observe the effect of adding similar priors for other vision downstream tasks like object detection and instance segmentation.
References
Appendix A Tuning the hyperparameter α𝛼\alpha
We weigh the loss () calculated on the semantic-prior prediction with a hyperparameter as formulated in Eq. 6. Using weighted supervision for the semantic-prior maps is critical so that the model treats the semantic context as an additional signal for feature modeling and not as the main prediction.
We study the impact of on performance in Tab. I by changing the values of on the Swin-Tiny variant. is the optimum setting for modeling the network’s image feature level and semantic level context.
Appendix B Experiments on COCO-Stuff 10k
COCO-Stuff 10k comprises of a total of 10k images with dense pixel-level annotations, selected from the COCO dataset. The training set contains 9k images with 171 semantic classes and the test set contains 1k images.
We set the base learning rate to , weight decay to and train for 80K iterations with a batch size of 16.
We provide our experimental results in Tab. II. Our SeMask framework shows impressive improvement on the COCO-Stuff 10k dataset proving its dataset-agnostic ability.
Appendix C Analysis on SeMask
In order to confirm our hypothesis that adding semantic context inside the encoder with the help of the semantic attention operation helps in improving the semantic quality of the features, we analyze the pixel-wise attention quality of the intermediate features of our SeMask-T FPN model on the Cityscapes val dataset as shown in Fig. I.
Specifically, we analyze pixel-wise attention for the pre-SeMask () and post-SeMask () features (Fig. II) for Stage-3 and Stage-4 which are downsampled by and , respectively. We calculate the pixel-wise attention maps corresponding to the target pixel (red cross sign), and we observe that post-SeMask features have more similar features for the same semantic category region with better boundaries than the pre-SeMask features. It reflects that the semantic prior maps help increase similarity between the pixels belonging to the same semantic category and improve the semantic segmentation performance.
Appendix D Qualitative results
We provide qualitative results on the COCO-Stuff 10k test set in Fig. III where SeMask-L FPN produces better per-pixel predictions compared to Swin-L FPN. It is evident in as the Swin-L FPN network fails to label the pole correctly and completely mislabels the sky region in .
We show more qualitative results on the ADE20K validation set in Fig. IV. Swin-L FPN mislabels mirror as curtain in due to the reflection of the curtain. On the other hand, SeMask-L FPN classifies the regions accurately.