Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers
Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip H. S. Torr, Li Zhang
Introduction
Since the seminal work of , existing semantic segmentation models have been dominated by those based on fully convolutional network (FCN). A standard FCN segmentation model has an encoder-decoder architecture: the encoder is for feature representation learning, while the decoder for pixel-level classification of the feature representations yielded by the encoder. Among the two, feature representation learning (i.e., the encoder) is arguably the most important model component . The encoder, like most other CNNs designed for image understanding, consists of stacked convolution layers. Due to concerns on computational cost, the resolution of feature maps is reduced progressively, and the encoder is hence able to learn more abstract/semantic visual concepts with a gradually increased receptive field. Such a design is popular due to two favorable merits, namely translation equivariance and locality. The former respects well the nature of imaging process which underpins the model generalization ability to unseen image data. Whereas the latter controls the model complexity by sharing parameters across space. However, it also raises a fundamental limitation that learning long-range dependency information, critical for semantic segmentation in unconstrained scene images , becomes challenging due to still limited receptive fields.
To overcome this aforementioned limitation, a number of approaches have been introduced recently. One approach is to directly manipulate the convolution operation. This includes large kernel sizes , atrous convolutions , and image/feature pyramids . The other approach is to integrate attention modules into the FCN architecture. Such a module aims to model global interactions of all pixels in the feature map . When applied to semantic segmentation , a common design is to combine the attention module to the FCN architecture with attention layers sitting on the top. Taking either approach, the standard encoder-decoder FCN model architecture remains unchanged. More recently, attempts have been made to get rid of convolutions altogether and deploy attention-alone models instead. However, even without convolution, they do not change the nature of the FCN model structure: an encoder downsamples the spatial resolution of the input, developing lower-resolution feature mappings useful for discriminating semantic classes, and the decoder upsamples the feature representations into a full-resolution segmentation map.
In this paper, we aim to provide a rethinking to the semantic segmentation model design and contribute an alternative. In particular, we propose to replace the stacked convolution layers based encoder with gradually reduced spatial resolution with a pure transformer , resulting in a new segmentation model termed SEgmentation TRansformer (SETR). This transformer-alone encoder treats an input image as a sequence of image patches represented by learned patch embedding, and transforms the sequence with global self-attention modeling for discriminative feature representation learning. Concretely, we first decompose an image into a grid of fixed-sized patches, forming a sequence of patches. With a linear embedding layer applied to the flattened pixel vectors of every patch, we then obtain a sequence of feature embedding vectors as the input to a transformer. Given the learned features from the encoder transformer, a decoder is then used to recover the original image resolution. Crucially there is no downsampling in spatial resolution but global context modeling at every layer of the encoder transformer, thus offering a completely new perspective to the semantic segmentation problem.
This pure transformer design is inspired by its tremendous success in natural language processing (NLP) . More recently, a pure vision transformer or ViT has shown to be effective for image classification tasks. It thus provides direct evidence that the traditional stacked convolution layer (i.e., CNN) design can be challenged and image features do not necessarily need to be learned progressively from local to global context by reducing spatial resolution. However, extending a pure transformer from image classification to a spatial location sensitive task of semantic segmentation is non-trivial. We show empirically that SETR not only offers a new perspective in model design, but also achieves new state of the art on a number of benchmarks.
The following contributions are made in this paper: (1) We reformulate the image semantic segmentation problem from a sequence-to-sequence learning perspective, offering an alternative to the dominating encoder-decoder FCN model design. (2) As an instantiation, we exploit the transformer framework to implement our fully attentive feature representation encoder by sequentializing images. (3) To extensively examine the self-attentive feature presentations, we further introduce three different decoder designs with varying complexities. Extensive experiments show that our SETR models can learn superior feature representations as compared to different FCNs with and without attention modules, yielding new state of the art on ADE20K (50.28%), Pascal Context (55.83%) and competitive results on Cityscapes. Particularly, our entry is ranked the place in the highly competitive ADE20K test server leaderboard.
Related work
Semantic image segmentation has been significantly boosted with the development of deep neural networks. By removing fully connected layers, the fully convolutional network (FCN) is able to achieve pixel-wise predictions. While the predictions of FCN are relatively coarse, several CRF/MRF based approaches are developed to help refine the coarse predictions. To address the inherent tension between semantics and location , coarse and fine layers need to be aggregated for both the encoder and decoder. This leads to different variants of the encoder-decoder structures for multi-level feature fusion.
Many recent efforts have been focused on addressing the limited receptive field/context modeling problem in FCN. To enlarge the receptive field, DeepLab and Dilation introduce the dilated convolution. Alternatively, context modeling is the focus of PSPNet and DeepLabV2 . The former proposes the PPM module to obtain different region’s contextual information while the latter develops ASPP module that adopts pyramid dilated convolutions with different dilation rates. Decomposed large kernels are also utilized for context capturing. Recently, attention based models are popular for capturing long range context information. PSANet develops the pointwise spatial attention module for dynamically capturing the long range context. DANet embeds both spatial attention and channel attention. CCNet alternatively focuses on economizing the heavy computation budget introduced by full spatial attention. DGMN builds a dynamic graph message passing network for scene modeling and it can significantly reduce the computational complexity. Note that all these approaches are still based on FCNs where the feature encoding and extraction part are based on classical ConvNets like VGG and ResNet . In this work, we alternatively rethink the semantic segmentation task from a different perspective.
Transformer
Transformer and self-attention models have revolutionized machine translation and NLP . Recently, there are also some explorations for the usage of transformer structures in image recognition. Non-local network appends transformer style attention onto the convolutional backbone. AANet mixes convolution and self-attention for backbone training. LRNet and stand-alone networks explore local self-attention to avoid the heavy computation brought by global self-attention. SAN explores two types of self-attention modules. Axial-Attention decomposes the global spatial attention into two separate axial attentions such that the computation is largely reduced. Apart from these pure transformer based models, there are also CNN-transformer hybrid ones. DETR and the following deformable version utilize transformer for object detection where transformer is appended inside the detection head. STTR and LSTR adopt transformer for disparity estimation and lane shape prediction respectively. Most recently, ViT is the first work to show that a pure transformer based image classification model can achieve the state-of-the-art. It provides direct inspiration to exploit a pure transformer based encoder design in a semantic segmentation model.
The most related work is which also leverages attention for image segmentation. However, there are several key differences. First, though convolution is completely removed in as in our SETR, their model still follows the conventional FCN design in that spatial resolution of feature maps is reduced progressively. In contrast, our sequence-to-sequence prediction model keeps the same spatial resolution throughout and thus represents a step-change in model design. Second, to maximize the scalability on modern hardware accelerators and facilitate easy-to-use, we stick to the standard self-attention design. Instead, adopts a specially designed axial-attention which is less scalable to standard computing facilities. Our model is also superior in segmentation accuracy (see Section 4).
Method
In order to contrast with our new model design, let us first revisit the conventional FCN for image semantic segmentation. An FCN encoder consists of a stack of sequentially connected convolutional layers. The first layer takes as input the image, denoted as with specifying the image size in pixels. The input of subsequent layer is a three-dimensional tensor sized , where and are spatial dimensions of feature maps, and is the feature/channel dimension. Locations of the tensor in a higher layer are computed based on the locations of tensors of all lower layers they are connected to via layer-by-layer convolutions, which are defined as their receptive fields. Due to the locality nature of convolution operation, the receptive field increases linearly along the depth of layers, conditional on the kernel sizes (typically ). As a result, only higher layers with big receptive fields can model long-range dependencies in this FCN architecture. However, it is shown that the benefits of adding more layers would diminish rapidly once reaching certain depths . Having limited receptive fields for context modeling is thus an intrinsic limitation of the vanilla FCN architecture.
Recently, a number of state-of-the-art methods suggest that combing FCN with attention mechanism is a more effective strategy for learning long-range contextual information. These methods limit the attention learning to higher layers with smaller input sizes alone due to its quadratic complexity \wrtthe pixel number of feature tensors. This means that dependency learning on lower-level feature tensors is lacking, leading to sub-optimal representation learning. To overcome this limitation, we propose a pure self-attention based encoder, named SEgmentation TRansformers (SETR).
2 Segmentation transformers (SETR)
A straightforward way for image sequentialization is to flatten the image pixel values into a 1D vector with size of . For a typical image sized at , the resulting vector will have a length of 691,200. Given the quadratic model complexity of Transformer, it is not possible that such high-dimensional vectors can be handled in both space and time. Therefore tokenizing every single pixel as input to our transformer is out of the question.
Transformer
3 Decoder designs
To evaluate the effectiveness of SETR’s encoder feature representations , we introduce three different decoder designs to perform pixel-level segmentation. As the goal of the decoder is to generate the segmentation results in the original 2D image space , we need to reshape the encoder’s features (that are used in the decoder), , from a 2D shape of to a standard 3D feature map . Next, we briefly describe the three decoders.
This naive decoder first projects the transformer feature to the dimension of category number (e.g., 19 for experiments on Cityscapes). For this we adopt a simple 2-layer network with architecture: conv + sync batch norm (w/ ReLU) + conv. After that, we simply bilinearly upsample the output to the full image resolution, followed by a classification layer with pixel-wise cross-entropy loss. When this decoder is used, we denote our model as SETR-Naïve.
(2) Progressive UPsampling (PUP)
Instead of one-step upscaling which may introduce noisy predictions, we consider a progressive upsampling strategy that alternates conv layers and upsampling operations. To maximally mitigate the adversarial effect, we restrict upsampling to 2. Hence, a total of 4 operations are needed for reaching the full resolution from with size . More details of this process are given in Figure 1(b). When using this decoder, we denote our model as SETR-PUP.
(3) Multi-Level feature Aggregation (MLA)
The third design is characterized by multi-level feature aggregation (Figure 1(c)) in similar spirit of feature pyramid network . However, our decoder is fundamentally different because the feature representations of every SETR’s layer share the same resolution without a pyramid shape.
Specifically, we take as input the feature representations () from layers uniformly distributed across the layers with step to the decoder. streams are then deployed, with each focusing on one specific selected layer. In each stream, we first reshape the encoder’s feature from a 2D shape of to a 3D feature map . A 3-layer (kernel size , , and ) network is applied with the feature channels halved at the first and third layers respectively, and the spatial resolution upscaled by bilinear operation after the third layer. To enhance the interactions across different streams, we introduce a top-down aggregation design via element-wise addition after the first layer. An additional conv is applied after the element-wise additioned feature. After the third layer, we obtain the fused feature from all the streams via channel-wise concatenation which is then bilinearly upsampled to the full resolution. When using this decoder, we denote our model as SETR-MLA.
Experiments
We conduct experiments on three widely-used semantic segmentation benchmark datasets.
densely annotates 19 object categories in images with urban scenes. It contains 5000 finely annotated images, split into 2975, 500 and 1525 for training, validation and testing respectively. The images are all captured at a high resolution of . In addition, it provides 19,998 coarse annotated images for model training.
ADE20K
is a challenging scene parsing benchmark with 150 fine-grained semantic concepts. It contains 20210, 2000 and 3352 images for training, validation and testing.
PASCAL Context
provides pixel-wise semantic labels for the whole scene (both “thing” and “stuff” classes), and contains 4998 and 5105 images for training and validation respectively. Following previous works, we evaluate on the most frequent 59 classes and the background class (60 classes in total).
Implementation details
Following the default setting (e.g., data augmentation and training schedule) of public codebase mmsegmentation , (i) we apply random resize with ratio between 0.5 and 2, random cropping (768, 512 and 480 for Cityscapes, ADE20K and Pascal Context respectively) and random horizontal flipping during training for all the experiments; (ii) We set batch size 16 and the total iteration to 160,000 and 80,000 for the experiments on ADE20K and Pascal Context. For Cityscapes, we set batch size to 8 with a number of training schedules reported in Table 2, 6 and 7 for fair comparison. We adopt a polynomial learning rate decay schedule and employ SGD as the optimizer. Momentum and weight decay are set to 0.9 and 0 respectively for all the experiments on the three datasets. We set initial learning rate 0.001 on ADE20K and Pascal Context, and 0.01 on Cityscapes.
Auxiliary loss
As we also find the auxiliary segmentation loss helps the model training. Each auxiliary loss head follows a 2-layer network. We add auxiliary losses at different Transformer layers: SETR-Naïve (), SETR-PUP (), SETR-MLA (). Both auxiliary loss and main loss heads are applied concurrently.
Multi-scale test
We use the default settings of mmsegmentation . Specifically, the input image is first scaled to a uniform size. Multi-scale scaling and random horizontal flip are then performed on the image with a scaling factor (0.5, 0.75, 1.0, 1.25, 1.5, 1.75). Sliding window is adopted for test (e.g., for Pascal Context). If the shorter side is smaller than the size of the sliding window, the image is scaled with its shorter side to the size of the sliding window (e.g., 480) while keeping the aspect ratio. Synchronized BN is used in decoder and auxiliary loss heads. For training simplicity, we do not adopt the widely-used tricks such as OHEM loss in model training.
Baselines
We adopt dilated FCN and Semantic FPN as baselines with their results taken from . Our models and the baselines are trained and tested in the same settings for fair comparison. In addition, state-of-the-art models are also compared. Note that the dilated FCN is with output stride 8 and we use output stride 16 in all our models due to GPU memory constrain.
SETR variants
Three variants of our model with different decoder designs (see Sec. 3.3), namely SETR-Naïve, SETR-PUP and SETR-MLA. Besides, we use two variants of the encoder “T-Base” and “T-Large” with 12 and 24 layers respectively (Table 1). Unless otherwise specified, we use “T-Large” as the encoder for SETR-Naïve, SETR-PUP and SETR-MLA. We denote SETR-Naïve-Base as the model utilizing “T-Base” in SETR-Naïve.
Though designed as a model with a pure transformer encoder, we also set a hybrid baseline Hybrid by using a ResNet-50 based FCN encoder and feeding its output feature into SETR. To cope with the GPU memory constraint and for fair comparison, we only consider ‘T-Base” in Hybrid and set the output stride of FCN to . That is, Hybrid is a combination of ResNet-50 and SETR-Naïve-Base.
Pre-training
We use the pre-trained weights provided by ViT or DeiT to initialize all the transformer layers and the input linear projection layer in our model. We denote SETR-Naïve-DeiT as the model utilizing DeiT pre-training in SETR-Naïve-Base. All the layers without pre-training are randomly initialized. For the FCN encoder of Hybrid, we use the initial weights pre-trained on ImageNet-1k. For the transformer part, we use the weights pre-trained by ViT , DeiT or randomly initialized.
We use patch size for all the experiments. We perform 2D interpolation on the pre-trained position embeddings, according to their location in the original image for different input size fine-tuning.
Evaluation metric
Following the standard evaluation protocol , the metric of mean Intersection over Union (mIoU) averaged over all classes is reported. For ADE20K, additionally pixel-wise accuracy is reported following the existing practice.
2 Ablation studies
Table 2 and 3 show ablation studies on (a) different variants of SETR on various training schedules, (b) comparison to FCN and Semantic FPN , (c) pre-training on different data, (d) comparison with Hybrid, (e) compare to FCN with different pre-training. Unless otherwise specified, all experiments on Table 2 and 3 are trained on Cityscapes train fine set with batch size 8, and evaluated using the single scale test protocol on the Cityscapes validation set in mean IoU (%) rate. Experiments on ADE20K also follow the single scale test protocol.
From Table 2, we can make the following observations: (i) Progressively upsampling the feature maps, SETR-PUP achieves the best performance among all the variants on Cityscapes. One possible reason for inferior performance of SETR-MLA is that the feature outputs of different transformer layers do not have the benefits of resolution pyramid as in feature pyramid network (FPN) (see Figure 5). However, SETR-MLA performs slightly better than SETR-PUP, and much superior to the variant SETR-Naïve that upsamples the transformers output feature by 16 in one-shot, on ADE20K val set (Table 3 and 4). (ii) The variants using “T-Large” (e.g., SETR-MLA and SETR-Naïve) are superior to their “T-Base” counterparts, i.e., SETR-MLA-Base and SETR-Naïve-Base, as expected. (iii) While our SETR-PUP-Base (76.71) performs worse than Hybrid-Base (76.76), it shines (78.02) when training with more iterations (80k). It suggests that FCN encoder design can be replaced in semantic segmentation, and further confirms the effectiveness of our model. (iv) Pre-training is critical for our model. Randomly initialized SETR-PUP only gives 42.27% mIoU on Cityscapes. Model pre-trained with DeiT on ImageNet-1K gives the best performance on Cityscapes, slightly better than the counterpart pre-trained with ViT on ImageNet-21K. (v) To study the power of pre-training and further verify the effectiveness of our proposed approach, we conduct the ablation study on the pre-training strategy in Table 3. For fair comparison with the FCN baseline, we first pre-train a ResNet-101 on the Imagenet-21k dataset with a classification task and then adopt the pre-trained weights for a dilated FCN training for the semantic segmentation task on ADE20K or Cityscapes. Table 3 shows that with ImageNet-21k pre-training FCN baseline experienced a clear improvement over the variant pre-trained on ImageNet-1k. However, our method outperforms the FCN counterparts by a large margin, verifying that the advantage of our approach largely comes from the proposed sequence-to-sequence modeling strategy rather than bigger pre-training data.
3 Comparison to state-of-the-art
Table 4 presents our results on the more challenging ADE20K dataset. Our SETR-MLA achieves superior mIoU of 48.64% with single-scale (SS) inference. When multi-scale inference is adopted, our method achieves a new state of the art with mIoU hitting 50.28%. Figure 2 shows the qualitative results of our model and dilated FCN on ADE20K. When training a single model on the train+validation set with the default 160,000 iterations, our method ranks place in the highly competitive ADE20K test server leaderboard.
Results on Pascal Context
Table 5 compares the segmentation results on Pascal Context. Dilated FCN with the ResNet-101 backbone achieves a mIoU of 45.74%. Using the same training schedule, our proposed SETR significantly outperforms this baseline, achieving mIoU of 54.40% (SETR-PUP) and 54.87% (SETR-MLA). SETR-MLA further improves the performance to 55.83% when multi-scale (MS) inference is adopted, outperforming the nearest rival APCNet with a clear margin. Figure 3 gives some qualitative results of SETR and dilated FCN. Further visualization of the learned attention maps in Figure 6 shows that SETR can attend to semantically meaningful foreground regions, demonstrating its ability to learn discriminative feature representations useful for segmentation.
Results on Cityscapes
Tables 6 and 7 show the comparative results on the validation and test set of Cityscapes respectively. We can see that our model SETR-PUP is superior to FCN baselines, and FCN plus attention based approaches, such as Non-local and CCNet ; and its performance is on par with the best results reported so far. On this dataset we can now compare with the closely related Axial-DeepLab which aims to use an attention-alone model but still follows the basic structure of FCN. Note that Axial-DeepLab sets the same output stride 16 as ours. However, its full input resolution () is much larger than our crop size , and it runs more epochs (60k iteration with batch size 32) than our setting (80k iterations with batch size 8). Nevertheless, our model is still superior to Axial-DeepLab when multi-scale inference is adopted on Cityscapes validation set. Using the fine set only, our model (trained with 100k iterations) outperforms Axial-DeepLab-XL with a clear margin on the test set. Figure 4 shows the qualitative results of our model and dilated FCN on Cityscapes.
Conclusion
In this work, we have presented an alternative perspective for semantic segmentation by introducing a sequence-to-sequence prediction framework. In contrast to existing FCN based methods that enlarge the receptive field typically with dilated convolutions and attention modules at the component level, we made a step change at the architectural level to completely eliminate the reliance on FCN and elegantly solve the limited receptive field challenge. We implemented the proposed idea with Transformers that can model global context at every stage of feature learning. Along with a set of decoder designs in different complexity, strong segmentation models are established with none of the bells and whistles deployed by recent methods. Extensive experiments demonstrate that our models set new state of the art on ADE20, Pascal Context and competitive results on Cityscapes. Encouragingly, our method is ranked the place in the highly competitive ADE20K test server leaderboard on the day of submission.
Acknowledgments
This work was supported by Shanghai Municipal Science and Technology Major Project (No.2018SHZDZX01), ZJLab, and Shanghai Center for Brain Science and Brain-Inspired Technology.
References
Appendix
Appendix A Visualizations
Visualization of the learned position embedding in Figure 7 shows that the model learns to encode distance within the image in the similarity of position embeddings.
Features
Figure 9 shows the feature visualization of our SETR-PUP. For the encoder, 24 output features from the 24 transformer layers namely are collected. Meanwhile, 5 features () right after each bilinear interpolation in the decoder head are visited.
Attention maps
Attention maps (Figure 10) in each transformer layer catch our interest. There are 16 heads and 24 layers in T-large. Similar to , a recursion perspective into this problem is applied. Figure 8 shows the attention maps of different selected spatial points (red).